A Multi-Label Printed Image Retrieval Method Based on Deep Hashing
Through a multi-label printed image retrieval method based on deep hash, the printed image features are extracted using convolutional neural network and attention mechanism, and the hash layer encoding is used to solve the problem of insufficient search accuracy and speed of printed image retrieval in the prior art, achieving a more efficient image retrieval effect.
Patent Information
- Application Number
- CN202211300811.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-10-24
AI Technical Summary
The existing printing image retrieval methods cannot meet the needs of enterprises in terms of accuracy and speed, especially for large-scale printing image retrieval work.
A multi-label printed image retrieval method based on deep hash is used to extract image features through a convolutional neural network, combine semantic attention network and graph convolutional neural network to learn interdependent features, and encode images through a hash layer to improve the real-time and accuracy of retrieval.
It significantly improves the accuracy and speed of printed image retrieval, and can more effectively process large-scale printed image datasets.
Smart Images

Figure CN115510254B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of printed image retrieval, and in particular to a multi-label printed image retrieval method based on deep hashing. Background Art
[0002] The analysis of printed images has great potential in commercial applications and industrial production. A single local textile enterprise can design 100,000 printed images in a year. The increasing number of printed pattern images has caused the accumulation of local image libraries, bringing corresponding conventional values and challenges of general image retrieval, while there is basically no retrieval method for printed images at present.
[0003] In the field of image retrieval, it is roughly divided into text-based image retrieval and content-based image retrieval. Text-based image retrieval uses the attached information of images such as titles, descriptions, keywords, etc. for search and matching. However, this method does not adapt to the current large-scale image retrieval work because it ignores the visual content contained in the printed images themselves and overly relies on the labeled content. For content-based image retrieval methods, currently, they mainly rely on deep learning methods, using the hierarchical structure of convolution, pooling, and non-linearity of convolutional neural networks to extract the semantic features of images, calculating the extracted semantic features into a compact vector representation, and using the Euclidean distance or approximate nearest neighbor search algorithm for retrieval.
[0004] For printed images, the method of simply using a convolutional neural network to extract semantic features and performing retrieval through the Euclidean distance or approximate nearest neighbor search algorithm cannot meet the requirements of current enterprise printed image retrieval in terms of retrieval accuracy and speed. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a multi-label printed image retrieval method based on deep hashing, which greatly improves the accuracy and speed of printed image retrieval.
[0006] To solve the above technical problem, the present invention provides a multi-label printed image retrieval method based on deep hashing, including the following steps:
[0007] Step 1: Collect and organize printed images of various styles, annotate the printed images from multiple angles, and establish a multi-label printed image data set;
[0008] Step 2: Enhance the data by using image processing methods such as flipping, image shearing, color transformation, and contrast transformation;
[0009] Step 3: Split the printed image data set into a training set and a query set according to a certain ratio;
[0010] Step 4: Construct a multi-label printed image retrieval network model based on deep hashing;
[0011] Step 5: Use the printed image training set to train the model, then extract features from the printed image dataset through the trained model, construct the hash representation of the printed image, and build an image hash database;
[0012] Step 6: Use the printed image query set to perform retrieval on the trained model, and evaluate the retrieval accuracy and generalization ability of the model.
[0013] Preferably, in Step 4, constructing a multi-label printed image retrieval network model based on deep hashing specifically includes the following steps:
[0014] Step 41: Load a convolutional neural network as the backbone network for extracting image features;
[0015] Step 42: Add a semantic attention network after the backbone network to obtain the activated regions based on specific label features from the high-level of the deep network using soft attention mapping;
[0016] Step 43: Add a graph convolutional neural network to learn interdependent features for each printed image label;
[0017] Step 44: Add a supervised hashing network to reduce the high-dimensional feature vectors extracted by the deep network to a string of short bitcodes, enhancing the real-time performance of retrieval.
[0018] Preferably, in Step 41, load the convolutional neural network ResNet-50 as the backbone network, load the weights pre-trained on the ImageNet dataset, remove the last fully connected layer, and retain the remaining layers; use the visual feature map of the "res4b6relu" layer of the backbone network as the input visual features for the subsequent two sub-networks of the model, with an input size of 3×224×224 and an output size of the backbone network of 14×14×1024.
[0019] Preferably, in Step 42, add a semantic attention network after the backbone network to obtain the activated regions based on specific label features from the high-level of the deep network using soft attention mapping; the attention network includes three learnable convolutional layers with kernel sizes of 1×1×512, 3×3×512, and 1×1×C, where C is the number of all label types in the database, and all three convolutional layers use ReLU as the activation function; the channel dimension of the feature map X ∈ R 14×14×1024 is reduced to 512→512→C respectively after three convolutions, and the unnormalized attention value map A' ∈ R H×W×C is obtained after convolution, where H = 14, W = 14, and the map A' of each channel k ∈ R H×Wcorresponds to a label, that is, 1 ≤ k ≤ C; then the mapping graph of the attention value corresponding to the label k is normalized using the softmax function to obtain the final visual attention mapping graph A ∈ R H×W×C , where H = 14 and W = 14; the feature map X output in the backbone network stage is weighted and summed using the attention map of each label to generate the label representation V k ∈ R 1024 , where 1 ≤ k ≤ C; this semantic attention network uses the soft attention mapping to obtain C label representations V = [V1, V2,..., V C ∈ R C×1024 , and each feature representation can selectively aggregate relevant features regarding a specific label; for each label feature vector V k , a linear fully connected layer is used to evaluate the confidence score of label k; finally, the confidence of the input image for all labels is obtained Using the true label y = [y1, y2,..., y C and y att The cross-entropy loss between them is used to learn the evaluation parameters of the attention map.
[0020] Preferably, in step 43, a graph convolutional neural network is added to learn interdependent features for each printed image label; the content-aware label representation V obtained by the semantic attention network is transformed through the training learning method of the graph convolutional neural network to change the relevance of these label representations in the multi-label image retrieval task; the graph convolutional neural network uses the content-aware label representation V as the input node feature, and uses a non-negative correlation matrix M ∈ R C×C to represent the label relationship, then M is symmetrized to obtain M', and M' and the learnable transformation matrix W are used to update the label representation V. After repeatedly stacking the graph convolutional network, the feature representation of V → Z ∈ R C×D is obtained, and finally a linear fully connected layer is used to obtain the embedded representation E ∈ R 2048 of the picture; the graph convolutional neural network uses the ranking loss function triplet loss as the loss value to update the graph convolutional neural network, which involves the relationship between triplet samples and provides a concurrent ranking with the anchor sample as the core for negative sample embeddings and positive sample embeddings.
[0021] Preferably, in step 44, a supervised hashing network is added to reduce the high-dimensional feature vectors extracted by the deep network into a string of short bitcodes, enhancing the real-time performance of retrieval; after the last embedding layer of the graph convolutional neural network, a new fully-connected hashing layer with q hidden nodes is added. This hashing layer converts the deep representation of the printed image into a low-dimensional hashing code representation, and then an activation function is used to control the output within the interval (-1, 1) to more conveniently implement the hashing code; pairwise loss is introduced to ensure that the learned hashing codes can retain the fine and complex similarity information between a pair of images, and a new triplet metric loss is introduced to learn the ranking relationship between samples. Finally, a quantization loss is used to control the quality of the mapped hash.
[0022] Preferably, in step 5, the model is trained using the printed image training set, and then the trained model is used to extract features from the printed image dataset, construct the hashing representation of the printed image, and construct an image hashing database; there are four stages in training this model. In the first stage, the attention loss is used to train the parameters of the semantic attention network and the backbone network; in the second stage, the weights of the backbone network and the semantic attention network are fixed, and the ranking loss is used to focus on training the parameters in the graph convolutional network; in the third stage, the attention loss and the ranking loss are used to alternately train the backbone network, the semantic attention network, and the graph convolutional network; finally, the ranking loss, pairwise loss, and quantization loss of the joint hashing layer are used to fine-tune the backbone network, the semantic attention network, the graph convolutional network, and the second-to-last fully-connected layer.
[0023] The beneficial effects of the present invention are as follows: The present invention uses a convolutional network to extract image features, uses an attention mechanism to obtain the activated regions based on specific label features from the high levels of the deep network through soft attention mapping, then uses a graph convolutional neural network to learn the interdependent features for each printed image label, and finally encodes the printed image through a hashing layer, greatly improving the accuracy and speed of printed image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic flowchart of the method of the present invention.
[0025] Figure 2 It is a schematic diagram of the model architecture of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0026] As Figure 1 shown, a multi-label printed image retrieval method based on deep hashing includes the following steps:
[0027] Step 1: Collect and organize printed images of various styles, annotate the printed images from multiple angles, and establish a multi-label printed image dataset;
[0028] Step 2: Enhance the data using image processing methods such as flipping, image shearing, color transformation, and contrast transformation;
[0029] Given that the printed image has perceptual consistency under spatial transformation, that is, when the input image undergoes translation, rotation, flipping, and scaling transformations, the human visual perception of the picture remains unchanged. In order to fully train the deep learning model and reduce the overfitting problem, the present invention also uses the following image processing methods to enhance the data:
[0030] (a) Flipping: Randomly flip the printed image horizontally;
[0031] (b) Image shearing: Use a sliding window in the opencv library to crop the image at multiple scales;
[0032] (c) Color transformation: Multiply the hue of each pixel in the image by a random value in [0,1] to enhance the image pixels;
[0033] (d) Contrast transformation: Directly adjust the histogram or pixel values.
[0034] Finally, after sorting, data enhancement, and screening, the scale of the printed image dataset reached the order of 30,000, and the semantic labels reached 29 kinds, including 20 elements such as flowers, branches and leaves, birds, butterflies, animals, bottles and jars, words, mountains, oceans, people, portraits, buildings, animal skin patterns, ethnic patterns, plant patterns, grids, stripes, circles, triangles, and stars, 6 styles such as pastoral style, Chinese style, ethnic style, European style, geometric abstraction, and cartoon, and 3 layouts such as single pattern, repeated pattern, and regular continuity. Each picture is at least one of them. The present invention randomly selects 3,000 of them as the query set, and the rest as the training set to construct a printed image dataset.
[0035] Step 3: Split the printed image dataset into a training set and a query set according to a ratio of 9:1;
[0036] Step 4: As Figure 2 shown, construct a multi-label printed image retrieval network model based on deep hashing;
[0037] Step 41: Load a convolutional neural network as the backbone network to extract image features;
[0038] Step 42: Add a semantic attention network after the backbone network to obtain the region activated by specific label features from the high layer of the deep network using soft attention mapping;
[0039] Step 43: Add a graph convolutional neural network to learn interdependent features for each printed image label;
[0040] Step 44: Add a supervised hashing network to reduce the high-dimensional feature vectors extracted by the deep network into a string of short bit codes, enhancing the real-time performance of retrieval.
[0041] In step 41, load the convolutional neural network ResNet-50 as the backbone network, load the weights pre-trained on the ImageNet dataset, remove the last fully connected layer, and retain the remaining layers. In this method, the visual feature map of the "res4b6relu" layer of the backbone network is used as the input visual features for the subsequent two sub-networks of the model. In this method, the input size is 3×224×224, and the output size of the backbone network is 14×14×1024. Thus, in this method, there is sufficient resolution to learn spatial attention.
[0042] In step 42, add a semantic attention network after the backbone network to obtain the regions activated by specific label features from the high-level of the deep network using soft attention mapping. The semantic attention network is constructed by three learnable convolutional layers, and their convolutional kernel sizes are 1×1×512, 3×3×512, 1×1×C (C is the number of all label types in the database), and ReLU is used as the activation function for all three convolutional layers.
[0043] Suppose an input printed image to the deep network is denoted by I, and its true label is denoted by y = [y1, y2,..., y C , where y l ∈{0,1}, 1≤k≤C, and C is the number of all label types in the database. If the image I contains the k label, then y k = 1, and if the image I does not contain the k label, then y k = 0.
[0044] The channel dimension of the feature map X ∈ R 14×14×1024 is reduced to 512→512→C respectively after three convolutions. After convolution, the unnormalized attention value map A' ∈ R H×W×C is obtained, where H = 14, W = 14, and each channel map A' k ∈ R H×W corresponds to a label, that is, 1≤k≤C. Then, the softmax function is used to normalize the attention value map corresponding to the label k to obtain the final visual attention map A ∈ R H×W×C .
[0045] Next, use the attention map of each label to perform weighted summation on the feature map X output in the backbone network stage to generate the label representation V k ∈ R 1024 , where 1≤k≤C, and V k can be expressed as follows:
[0046]
[0047] where X i,j ∈R 1024 represents the feature vector at the (i, j) position in the feature map X.
[0048] Then, the attention mechanism uses the semantic attention network to obtain C label representations V = [V1, V2,..., V C ∈ R C ×1024 , and each feature representation can selectively aggregate relevant features regarding a specific label. For each label feature vector V k , a linear fully connected layer is used to evaluate the confidence score of label k. Finally, the confidence of the input image for all labels is obtained
[0049] Using the cross-entropy loss between the true label y = [y1, y2,..., y C of the image and y att to learn the evaluation parameters of the attention map, as shown in the following formula:
[0050]
[0051] where σ(·) is the Sigmoid function.
[0052] In step 43, a graph convolutional neural network is added to learn the interdependent features for each printed image label. Using the content-aware label representation V obtained by the semantic attention network, the relevance of these label representations in the multi-label image retrieval task is transformed through the training learning method of the graph convolutional neural network. The graph convolutional neural network takes the content-aware label representation V as the input node feature, and uses a non-negative value correlation matrix M ∈ R C×C to represent the label relationship (edge). Then, M is symmetrized to obtain M', which can be expressed as the following formula:
[0053]
[0054] Each layer of the convolutional neural network can be described in the following form:
[0055] V u = δ(M'VW)
[0056] where M' ∈ R C×C is globally shared, W ∈ R d×D is the learnable transformation matrix in the network training, and δ(·) is the non-linear activation function LeakyReLU(·) function.
[0057] As shown in the above formula, during the forward propagation process, the correlation matrix M' first spreads the correlation information among all nodes, and then each node receives all the necessary information and updates its state through the linear transformation W. Finally, the feature representation of V→Z∈R is obtained by repeatedly stacking the graph convolutional network, and finally a linear fully connected layer is used to obtain the embedded representation E∈R of the image. C×D That is, each image will obtain a 2048-dimensional feature representation after the forward propagation of the graph convolutional neural network. The graph convolutional neural network uses the triplet loss function as the loss value to update the graph convolutional network. It involves the relationship between triplet samples and provides a concurrent sorting with the anchor sample as the core for the negative sample embedding and the positive sample embedding, which can be shown by the following formula: 2048
[0058]
[0059] a p
[0060]
[0061]
[0062] a where b is the number of samples in a mini-batch, the set Γ is the sampling set of anchor, positive and negative samples, and d(·) is the standard Euclidean distance formula d(x,y)=||x - y||2. E a is the embedded representation of the anchor sample, and E p is the embedded representation of the positive sample, where there is a part of the label shared between the anchor sample a and the positive sample p and the sample that does not meet the positive sample condition is the negative sample. γ is a hyperparameter, which is set to 0.2 - 0.5 in the present invention to ensure that the embedded vectors will not be projected too far from each other.
[0060]
[0061]
[0062] In step 44, a supervised hashing neural network is added to reduce the high-dimensional feature vector extracted by the deep network to a string of short bitcodes, enhancing the real-time performance of retrieval. A new fully connected hashing layer with q hidden nodes is added after the last embedding layer of the graph convolutional neural network. This hashing layer converts the deep representation of the printed image into a low-dimensional hashing code representation, and then uses f a (x)=x / (|x| + 1) as the activation function to control the output within the interval (-1,1), making it more convenient to implement the hashing code. A pairwise loss is introduced to ensure that the learned hashing codes can retain the fine and complex similarity information between a pair of images, a new triplet metric loss is introduced to learn the sorting relationship between samples, and finally a quantization loss is used to control the quality of the mapped hash.
[0061] Given the hashing codes B of all printed images, the conditional probability P(s ij |B) of the similarity relationship between a printed image I i and another I j is:
[0062]
[0063] Among them, σ(·) is the Sigmoid function. In the present invention, such a sigmoid function is used to transform the Hamming distance into a form of similarity probability measure. In addition, the present invention implements the inner product calculation Ω ij =<b i ,b j > = b i T b j to measure the Hamming distance state of paired printed samples in the Hamming space.
[0064] Next, the cross-entropy loss function is implemented to calculate the paired loss in the case of s ij = 0 or 1, as shown in the following formula:
[0065]
[0066] Then, the present invention uses the sigmoid function to replace σ(Ω ij ) and simplifies to obtain:
[0067]
[0068] s ij is a continuous value. When s ij ∈(-1,1), that is, in the case of partial similarity, the present invention uses the following formula, that is, the root mean square error, to maintain the similarity between hash codes:
[0069]
[0070] Among them, the inner product value >b i ,b j > will be in the range of [-q,q], so that The value of will be in the interval [0,q], which is consistent with the value range of s ij ·q.
[0071] In the continuous interval of s ij ∈[0,1], L1 + ζL2 forms the paired loss, as shown in the following formula, where M ij = 0 or 1 represents partial similarity and complete similarity or not respectively. In the multi-label printed dataset with complex semantic relationships and many shared labels, the paired loss can play its role. Among them, ζ is a hyperparameter that weighs the root mean square error loss and the cross-entropy loss, and is set here By dividing by q (q represents the number of bits), the gradient of the root mean square error can be adaptively adjusted within a reasonable range.
[0072]
[0073] However, directly optimizing the above formula is still difficult. Since q represents the changes in different bit positions, this method needs to threshold the network output to reduce the possibility of the vanishing gradient problem during training. First, this method uses f to replace the binary encoding b as the output of the deep hashing network, and then the present invention redefines Ω in the above formula. ij as αf i T f j , where α is a control constraint hyperparameter, and let so that the value of Ω ij is within a relatively appropriate range of [-5, 5].
[0074] The pairwise loss within a batch will be rewritten in the following form:
[0075]
[0076] The ranking loss of the present invention is a new triplet loss for multi-label printing data. It considers the degree of similarity and its ranking, so that the learned model can infer the complex similarity relationships between multi-labels. This ranking loss approximates the ratio of label distances by learning the ratio of distances in the hash space, and is calculated as follows:
[0077]
[0078] where y is the multi-label vector of the printing image, f replaces the binary encoding b and is the output of the deep hashing network, D(·) represents the square of the Euclidean distance. And (a, p, n) represents a triplet, where a is the anchor sample, and p and n are samples with a certain distance relationship with the a sample.
[0079] Since the network does not directly output binary codes, the present invention uses quantization loss to drive the deep network to output features close to binary codes.
[0080] Q(k) = |||f k | - 1||1
[0081] where 1 represents a q-bit vector all of whose elements are 1, ||·||1 is the L1 norm, and |·| represents the absolute value.
[0082] By combining the ranking loss, pairwise loss, and quantization loss, the total loss of this method in a batch of the dataset is:
[0083]
[0084] where λ and ω are the weight coefficients for controlling the quantization loss and ranking loss respectively.
[0085] Finally, the sgn(·) function is used to generate the hash code from the result of the network forward propagation.
[0086] For the ranking loss of the hash layer, the existing data sampling methods are not applicable because it deals with images with a multi-label structure. Therefore, the present invention adopts a data sampling strategy for this ranking loss.
[0087] First, in a batch of training samples B constructed in the present invention, there is an anchor sample a. Then, k nearest neighbor samples are selected according to the distance from the anchor label and marked as P, and the remaining |B| - k - 1 samples are randomly sampled from the remaining training set and marked as N. Therefore, a very dense triple is constructed in the finally formed set B and satisfies the following conditions:
[0088] Γ(B) = {(a, p, n)|D(y a , y p ) < D(y a , y n ), p ∈ P, n ∈ N}
[0089] Step 5: Use the printed image training set to train the model, then extract features from the printed image dataset through the trained model, construct the hash representation of the printed image, and construct an image hash database;
[0090] There are four stages to train this model. In the first stage, the attention loss is used to train the parameters of the semantic attention neural network and the backbone network; in the second stage, the weights of the backbone network and the semantic attention network are fixed, and the ranking loss is used to focus on training the parameters in the graph convolutional network; in the third stage, the attention loss and the ranking loss are used to alternately train the backbone network, the semantic attention neural network, and the graph convolutional network; finally, the ranking loss, pairwise loss, and quantization loss of the joint hash layer are used to fine-tune the backbone network, the semantic attention network, the graph convolutional network, and the penultimate fully connected layer.
[0091] All experiments are implemented based on Pytorch. The experimental environment is the Linux operating system, and a total of 4 Nvidia GeForce RTX2080ti with 11G are deployed. During the training process, the data augmentation method of random cropping provided in Pytorch is adopted in the present invention to avoid overfitting, and the size is adjusted to 224×224×3 for input into the network. Stochastic gradient descent is selected as the optimizer of the present invention, with a momentum of 0.9 and a weight decay of 0.0001. The batch size is set to {18, 32, 64, 96}. The initial learning rate of the attention network and the graph convolutional neural network is set to 0.5, and that of the backbone network is 0.01. The training model is iterated 80 times in total, 20 times in the semantic attention stage, 20 times in the graph convolutional neural network stage, and there are 40 iterations in the alternating training stage. The learning rate of the backbone network is reduced to 0.1 times the previous value every 10 iterations, and that of other sub-networks starts to be reduced to 0.1 times every 10 iterations from the 30th iteration. All learning rates stop decreasing when reaching 0.00001. Finally, the backbone network, the semantic attention network, the graph convolutional network, and the second last fully connected layer are fine-tuned using the ranking loss, pairwise loss, and quantization loss of the hash layer. The initial learning rate is 0.001, and the last layer, i.e., the hash fully connected layer, is trained through backpropagation. The Adam method is used for stochastic optimization, with a mini-batch size of 128, and the learning rate decay rate after every 10 iterations is 0.5. In the test stage, all input images are adjusted to a size of 224×224 for evaluation in the present invention.
[0092] Step 6: Use the printed image query set to retrieve on the trained model, and evaluate the retrieval accuracy and generalization ability of the model.
[0093] Finally, in the current mainstream deep hashing algorithms including LSH, ITQ, KSH, BDE, DHN, ADSH, etc., we verified the performance of each network on the multi-label printed dataset. As can be seen from Table 1, where the horizontal rows are the MAP values under different bits and the vertical columns are different hashing algorithms. It can be observed that the present method is much better than all the comparison methods to a large extent. Among the traditional hashing methods, KSH achieved the best result. Compared with KSH, the present method improved the MAP values by 8% - 9% and 6% - 7% respectively on the multi-label printed dataset. This shows the effectiveness and advantages of the present invention on the multi-label printed data with fine-grained similarity.
[0094] Table 1 Comparison table of test results
[0095]
Claims
1. A multi-label printed image retrieval method based on deep hashing, characterized in that, It includes the following steps: Step 1: Collect and organize printed image samples of various patterns, annotate the printed images from multiple perspectives, and establish a multi-label printed image dataset; Step 2: Use methods such as flipping, image cropping, color transformation, and contrast transformation for image processing to enhance the data; Step 3: Split the printed image dataset into a training set and a query set according to a certain ratio; Step 4: Construct a multi-label printed image retrieval network model based on deep hashing; specifically including the following steps: Step 41: Load a convolutional neural network as the backbone network for extracting image features; load the convolutional neural network ResNet-50 as the backbone network, load the weights pre-trained on the ImageNet dataset, remove the last fully connected layer, and retain the remaining layers; use the visual feature map of the "res4b6relu" layer of the backbone network as the input visual features for the subsequent two sub-networks of the model, with an input size of 3×224×224 and an output size of 14×14×1024 for the backbone network; Step 42: Add a semantic attention network after the backbone network to obtain regions activated by specific label features from the high-level of the deep network using soft attention mapping; The attention network consists of three learnable convolutional layers with kernel sizes of 1×1×512, 3×3×512, and 1×1×C respectively, where C is the number of all label categories in the database. ReLU is used as the activation function for all three convolutional layers; the channel dimension of the feature map X ∈ R 14×14×1024 is reduced to 512→512→C after three convolutions respectively, and the unnormalized attention value map A' ∈ R H×W×C is obtained after convolution, where H = 14, W = 14, and each channel map A' k ∈ R H×W corresponds to a label, that is, 1 ≤ k ≤ C; then the softmax function is used to normalize the attention value map corresponding to label k to obtain the final visual attention map A ∈ R H×W×C where H = 14, W = 14; the feature map X output in the backbone network stage is weighted and summed using the attention map of each label to generate the label representation V k ∈ R 1024 where 1 ≤ k ≤ C; this attention mechanism uses the semantic attention network to obtain C label representations V = [V1, V2,..., V C ∈ R C×1024 and each feature representation can selectively aggregate relevant features regarding a specific label; for each label feature vector V k , a linear fully connected layer is used to evaluate the confidence score of label k; finally, the confidence of the input image for all labels is obtained The cross - entropy loss between the true label y = [y1, y2,..., y C of the image and y att is used to learn the evaluation parameters of the attention map; Step 43: Add a graph convolutional neural network to learn interdependent features for each printed image label; Step 44: Add a supervised hashing network to reduce the high-dimensional feature vectors extracted by the deep network to a string of short bitcodes, enhancing the real-time performance of retrieval; Step 5: Use the printed image training set to train the model, then extract features from the printed image dataset through the trained model, construct a hashing representation of the printed images, and build an image hashing database; Step 6: Use the printed image query set to perform retrieval on the trained model, and evaluate the retrieval accuracy and generalization ability of the model.
2. The multi-label printed image retrieval method based on deep hashing according to claim 1, wherein, In step 43, a graph convolutional neural network is added to learn the interdependent features for each printed image label; the content-aware label representation V obtained by the attention network is transformed through the training learning method of the graph convolutional neural network for the relevance in the multi-label image retrieval task; the graph convolutional neural network takes the content-aware label representation V as the input node feature, and uses the non-negative value correlation matrix M ∈ R C×C to represent the label relationship, then symmetrizes M to obtain M', updates the label representation V using M' and the learnable transformation matrix W, and repeats the superposition using the graph convolutional network to obtain the feature representation V → Z ∈ R C×D Finally, a linear fully connected layer is used to obtain the embedding representation E ∈ R 2048 of the picture; the graph convolutional network updates the graph convolutional network using the ranking loss function triplet loss as the loss value, which involves the relationship between triplet samples and provides a concurrent ranking with the anchor sample as the core for the negative sample embedding and the positive sample embedding.
3. The multi-label printed image retrieval method based on deep hashing according to claim 1, characterized in that In Step 44, add a supervised hashing network to reduce the high-dimensional feature vectors extracted by the deep network to a string of short bitcodes, enhancing the real-time performance of retrieval; add a new fully connected hashing layer with q hidden nodes after the embedding layer of the last layer of the graph convolutional network. This hashing layer converts the deep representation of the printed image into a low-dimensional hashing code representation, and then uses an activation function to control the output within the interval (-1, 1) to more conveniently implement hashing coding; introduce pairwise loss to ensure that the learned hash codes can retain the fine and complex similarity information between a pair of images, introduce a new triplet metric loss to learn the ranking relationship between samples, and finally use a quantization loss to control the quality of the mapped hash.
4. The multi-label printed image retrieval method based on deep hashing according to claim 1, wherein, In step 5, the model is trained using the printed image training set, and then the trained model is used to extract features from the printed image dataset, construct a hash representation of the printed image, and construct an image hash database; there are four stages in training the model. In the first stage, the parameters of the semantic attention network and the backbone network are trained using the attention loss. In the second stage, the weights of the backbone network and the semantic attention network are fixed, and the ranking loss is used to focus on training the parameters in the graph convolutional network. In the third stage, the backbone network, the semantic attention network, and the graph convolutional network are alternately trained using the attention loss and the ranking loss. Finally, the ranking loss, pairwise loss, and quantization loss of the joint hash layer are used to adjust the backbone network, the semantic attention network, the graph convolutional network, and the penultimate fully connected layer.
Citation Information
Patent Citations
Multi-label image retrieval method based on deep hash
CN110457514A
Fine-grained bird image retrieval method based on graph neural network and deep hash
CN114329031A