Image retrieval method and medium
By performing channel weighting and feature fusion on the depth feature map in the image retrieval method, and using hash-guided metric loss method, the problem of semantic information loss during binarization is solved, and the accuracy of image retrieval is improved.
Patent Information
- Application Number
- CN202510592146.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
In supervised hash image retrieval, important semantic information may be lost during the binarization process, affecting the retrieval accuracy.
An image retrieval method is proposed, by obtaining the depth feature map of the image to be retrieved, performing channel weighting processing, extracting local and global features for fusion, and computing hash encoding based on the hash-guided metric loss method to reduce information loss during the binarization process.
Effectively prevent the loss of important semantic information during the binarization process and improve the accuracy of image retrieval.
Smart Images

Figure CN120104824A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image retrieval technology, and in particular to an image retrieval method and medium. Background Art
[0002] In the field of supervised hash image retrieval, the feature representation of the image is usually processed separately from the quantization step; this processing method mainly relies on extracting the feature representation of the image first, and then binarizing these features to generate hash codes. However, this separate processing method may cause some important semantic information to be lost when generating binary codes, thereby affecting the accuracy of retrieval. Summary of the invention
[0003] The present invention aims to solve one of the technical problems in the related art at least to a certain extent. To this end, one object of the present invention is to propose an image retrieval method that can effectively prevent the loss of important semantic information in the binarization process, thereby improving the accuracy of image retrieval.
[0004] In a first aspect, an embodiment of the present invention proposes an image retrieval method, comprising the following steps: obtaining an image to be retrieved, and performing feature extraction on the image to be retrieved to obtain a depth feature map corresponding to the image to be retrieved; performing weighted processing on each channel of the depth feature map to generate a weighted feature map; extracting local features and global features of the weighted feature map, and performing feature fusion based on the local features and the global features to generate a fusion feature; calculating a hash code to be retrieved corresponding to the fusion feature based on a hash-guided metric loss method; calculating the similarity between the image to be retrieved and a database image based on the hash code to be retrieved and the hash code of the database image, and determining a target image based on the similarity.
[0005] Furthermore, each channel of the depth feature map is weighted to generate a weighted feature map, including: performing global average pooling on each channel of the depth feature map to generate a global spatial compression representation; calculating a weight vector according to the global spatial compression representation based on an activation function; and performing principal element multiplication operation on the depth feature map according to the weight vector to generate a weighted feature map.
[0006] Furthermore, the global space compression representation is calculated by the following formula: ; in, represents the global space compressed representation, represents the height of the feature map, represents the width of the feature map, Represents the spatial position in the feature map Place The characteristic value of each channel is the spatial compression of the feature map along the channel dimension; ; in, express Activation function, express Activation function, The element value representing the feature map comes from the output feature map of the previous layer in the image feature extraction network and is used to enhance the nonlinear expression ability of the feature; ; ; ; in, It represents the generation of an intermediate representation containing global information. After this process, the feature map is compressed and converted into information that is more suitable for downstream tasks. Then, Z enters the second fully connected layer to further generate the channel weight vector W. express Activation function, and represents the bias vector, and represents the weight matrix, represents the weight vector, represents the weighted feature map, Represents the component in the weight vector for each channel.
[0007] Furthermore, local features and global features of the weighted feature map are extracted, and feature fusion is performed based on the local features and the global features to generate fused features, including: using a convolutional layer to process the weighted feature map to generate local features and global features; reshaping the local features into a first matrix, and reshaping the global features into a second matrix; transposing the second matrix, and performing matrix multiplication of the transposed second matrix with the first matrix, and using a softmax layer to calculate a spatial attention map; performing matrix multiplication of the spatial attention map and the second matrix to obtain a new feature map; combining the new feature map and the global features based on element-by-element addition to generate a fused feature.
[0008] Furthermore, the spatial attention map is calculated by the following formula: ; in, Indicates the local feature The position and global feature The correlation between the positions, represents the number of pixels, represents the first matrix, represents the second matrix; The fusion feature is calculated by the following formula: ; in, represents the fusion feature, Represents the weight coefficient that is dynamically adjusted during the learning process. In the initial stage of the model, it is usually initialized to 0, but during the training process, it will gradually update according to the gradient to adjust the fusion method of the final feature map; Represents the local feature map The eigenvalues at the positions, Represents the first The feature value at each position.
[0009] Furthermore, a hash code to be retrieved corresponding to the fused feature is calculated based on a hash-guided metric loss method, including: performing a global average pooling operation on the fused feature to obtain a corresponding global feature vector; inputting the global feature vector into a fully connected layer to map the global feature vector to a dimension consistent with the hash code length to obtain a mapping feature; and converting the mapping feature into a binary hash code based on a sign function.
[0010] Furthermore, the global eigenvector is calculated by the following formula: ; in, represents the global eigenvector, represents the height of the feature map, represents the width of the feature map, represents the fusion feature, and are the spatial indexes in the height and width directions of the feature map, respectively. is the channel index; ; in, represents the mapping feature, represents the weight of the fully connected layer, Represents the bias parameter of the fully connected layer; ; in, Represents a binary hash code.
[0011] Furthermore, in the process of generating the hash code, a hash-guided metric loss method is used to limit the learning range of the metric item, specifically, including: calculating the metric loss and correcting the error component in the metric loss based on a preset threshold to alleviate the conflict between the metric loss and the quantization loss; combining the metric loss and the quantization loss to reduce the information loss in the binarization process; and using the hash-guided metric loss function to optimize the model parameters.
[0012] In some embodiments, the metric loss is calculated by the following formula: ; in, represents the metric loss function, represents the set of proxy vectors, represents the embedding vector representing the input, represents the positive proxy vector, represents the boundary value, represents the scaling factor; The error component in the metric loss is corrected by the following formula: ; in, represents the improved hash-guided metric loss function, represents the set of all negative sample pairs, A hash code representing the proxy anchor. Represents the hash code of the current sample, The proxy anchor hash code representing the negative sample, express and The cosine similarity between represents the margin hyperparameter, Indicates the preset threshold; ; in, represents the minimum distance to the best global solution, Indicates the length of the hash code; In the process of combining metric loss and quantization loss, the learning objective is expressed by the following formula: ; ; in, Set to 1, represents the quantization loss, represents the predicted value of the i-th sample, represents the true value of the i-th sample, Represents the square L2 norm of the predicted value and the target value, and its calculation formula is: , is the dimension of the hash code, and Respectively The sample in The predicted and target values of the dimension.
[0013] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium on which an image retrieval program is stored. When the image retrieval program is executed by a processor, the image retrieval method described above is implemented.
[0014] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention.
[0015] According to the image retrieval method of the embodiment of the present invention, first, the image to be retrieved is obtained, and the feature extraction of the image to be retrieved is performed to obtain the depth feature map corresponding to the image to be retrieved; then, each channel of the depth feature map is weighted to generate a weighted feature map; then, the local features and global features of the weighted feature map are extracted, and feature fusion is performed based on the local features and global features to generate fusion features; then, the hash code to be retrieved corresponding to the fusion feature is calculated based on the hash-guided metric loss method; then, the similarity between the image to be retrieved and the database image is calculated based on the hash code to be retrieved and the database image hash code, and the target image is determined according to the similarity. Thereby, the loss of important semantic information in the binarization process is effectively prevented, thereby improving the accuracy of image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic flow chart of an image retrieval method according to an embodiment of the present invention; Figure 2 is a schematic diagram of the ResNet-50 model architecture according to an embodiment of the present invention; Figure 3 is a schematic diagram of a channel weighting process according to an embodiment of the present invention; Figure 4 is a schematic diagram of a multi-scale contextual collaboration process according to an embodiment of the present invention; Figure 5 is a schematic diagram of the ImageNet dataset retrieval effect according to an embodiment of the present invention; Figure 6 It is a schematic diagram of the CIFAR-10 dataset retrieval effect according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0018] The image retrieval method according to an embodiment of the present invention is described below with reference to the accompanying drawings.
[0019] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of an image retrieval method according to an embodiment of the present invention. Figure 1 As shown, the image retrieval method includes the following steps: S101, obtaining an image to be retrieved, and performing feature extraction on the image to be retrieved to obtain a depth feature map corresponding to the image to be retrieved.
[0020] As an example, Figure 2 As shown in the figure, first, in the feature extraction model architecture, the pre-trained ResNet-50 model is used as the backbone network. This model has been extensively trained on the ImageNet dataset and has strong generalization capabilities. Through residual learning in the deep network, the problem of gradient disappearance during deep neural network training is reduced, ensuring the efficiency of model optimization and the robustness of features. Figure 2 middle, Represents the output feature map of the third residual block (Res3) of ResNet-50, where 1024 represents the number of feature channels and 14×14 represents the spatial dimension of the feature map. Represents the feature map after processing by the Res4 block, with 2048 channels and a spatial dimension of 7 × 7. These parameter designs enable the model to maintain sufficient feature expression capabilities at deeper network levels while controlling the computational complexity.
[0021] Next, a multi-branch structure is adopted, including global branch and local branch. After the third residual block (Res3) of ResNet-50, a local and global dual-branch structure is introduced to process and refine the multi-dimensional features of the image, respectively. Among them, the global branch maintains the basic structure of ResNet-50, removes all pooling layers after the Res4 block, and retains the complete feature map spatial information, thereby improving the quality of global context features. In addition, in the global branch, the Feature Pyramid Pooling (FPP) technology is added to perform multi-scale pooling operations on the feature map, and the feature pooling results in different windows are integrated into a feature vector to ensure the diversity and scale invariance of global features, and to assist the network in making more accurate global semantic predictions. In the output of the global branch, F g[512,7,7] represents the output feature map after feature pyramid pooling. The number of channels is reduced to 512, while maintaining the 7×7 spatial dimension. This dimensionality reduction design effectively reduces the number of model parameters while retaining key feature information.
[0022] The local branch is dedicated to capturing and enhancing fine-grained local features. The Atrous Spatial Pyramid Pooling (ASPP) module is used, which contains multi-scale atrous convolution layers to cope with the problem of object size changes in the image. The parallel dilated convolutions of the ASPP module can capture spatial contexts of different scales at one time, strengthening the model's understanding of local information. After fusing feature maps with different dilation rates, the convolution processing obtains more representative and discriminative local attributes. In the local branch, Concat represents the feature connection operation, which connects the feature maps generated by the convolution layers with different dilation rates in the ASPP module in the channel dimension to form a richer feature representation. SelfAtt represents the self-attention mechanism, which is used to calculate the correlation between the positions within the feature map and enhance the model's attention to important local features. The final output F of the local branch l [512,7,7] has the same dimension as the global branch output, which is convenient for subsequent feature fusion. In addition, the local branch also contains a self-attention module (Self-AttentionModule) to deeply explore the correlation of each local feature point and improve the model's understanding and ability to distinguish image details. The multi-scale context collaboration module fuses and collaborates the features of the global and local branches, and finally generates a final descriptor with a dimension of [512×1] through a fully connected layer. This descriptor has strong distinguishing ability and robustness and can be used for efficient image retrieval tasks. This dual extraction and fusion strategy of global and local features enables the model to focus on the overall semantics and detail features of the image at the same time, thereby achieving better performance in image retrieval tasks. Especially for images with complex backgrounds or subtle differences, this structural design shows obvious advantages.
[0023] S102, performing weighted processing on each channel of the depth feature map to generate a weighted feature map.
[0024] In some embodiments, each channel of the depth feature map is weighted to generate a weighted feature map, including: performing global average pooling on each channel of the depth feature map to generate a global spatial compression representation; calculating a weight vector according to the global spatial compression representation based on an activation function; and performing principal element multiplication operations on the depth feature map according to the weight vector to generate a weighted feature map.
[0025] In some embodiments, the global spatial compressed representation is calculated by the following formula: ; in, represents the global space compressed representation, represents the height of the feature map, represents the width of the feature map, Represents the spatial position in the feature map Place The characteristic value of each channel is the spatial compression of the feature map along the channel dimension; ; in, express Activation function, express Activation function, The element value representing the feature map comes from the output feature map of the previous layer in the image feature extraction network and is used to enhance the nonlinear expression ability of the feature; ; ; ; in, It represents the generation of an intermediate representation containing global information. After this process, the feature map is compressed and converted into information that is more suitable for downstream tasks. Then, Z enters the second fully connected layer to further generate the channel weight vector W. express Activation function, and represents the bias vector, and represents the weight matrix, represents the weight vector, represents the weighted feature map, Represents the component in the weight vector for each channel.
[0026] As an example, Figure 3 As shown in Figure 1, channel weighting aims to dynamically adjust the weights of each channel in the feature map, thereby strengthening the channel features that contribute significantly to the image retrieval task and suppressing irrelevant or interfering features. First, the deep feature map is used as input, and the feature map of each channel is globally averaged and pooled to generate a global spatial compression representation. The global spatial compression representation is calculated by the following formula: ; in, represents the global space compressed representation, represents the height of the feature map, Indicates the width of the feature map; In terms of weight generation mechanism, and The activation function constructs the weight generation mechanism, As a smooth nonlinear activation function, the function can reduce the gradient vanishing problem and optimize the training of deep networks. The function is expressed by the following formula: ; Next, the generated global space compressed representation is subjected to two layers of full connection operations and used Activation function: ; ; in, It represents the generation of an intermediate representation containing global information. After this process, the feature map is compressed and converted into information that is more suitable for downstream tasks. Then, Z enters the second fully connected layer to further generate the channel weight vector W. express Activation function, and represents the bias vector, and represents the weight matrix, represents the weight vector.
[0027] Then, the generated weight vector is element-wise multiplied with the input deep feature map: ; in, represents the weighted feature map, represents the component of the weight vector for each channel. It represents the channel The weights on the channel are used to multiply the corresponding channel features element by element to highlight the features that contribute to the image retrieval task. The dot multiplication operation ensures that the global information in the weight vector can directly affect each element of the corresponding feature map, not just a single dimension. In this way, the channel weighting module not only amplifies or suppresses the features, but also enhances the semantic consistency of the feature map and the reliability of the information.
[0028] Next, output: weighted feature map. The final output is the feature map after channel weighting processing, which has stronger representation power and decision accuracy and will be used in the subsequent feature fusion and retrieval process.
[0029] Through the above steps, the channel weighting module can dynamically adjust the weight of each channel, highlight the features that are beneficial to the current image retrieval task, and thus enhance the overall representation ability of the network. This mechanism realizes the weighted adjustment of feature maps through the combination of global average pooling, Swish and Sigmoid activation functions, and optimizes the extraction of image features and subsequent retrieval effects.
[0030] S103, extracting local features and global features of the weighted feature map, and performing feature fusion based on the local features and the global features to generate fused features.
[0031] In some embodiments, local features and global features of a weighted feature map are extracted, and feature fusion is performed based on the local features and the global features to generate fused features, including: using a convolutional layer to process the weighted feature map to generate local features and global features; reshaping the local features into a first matrix, and reshaping the global features into a second matrix; transposing the second matrix, and performing matrix multiplication of the transposed second matrix with the first matrix, and using a softmax layer to calculate a spatial attention map; performing matrix multiplication of the spatial attention map and the second matrix to obtain a new feature map; combining the new feature map and the global features based on element-by-element addition to generate a fused feature.
[0032] In some embodiments, the spatial attention map is calculated by the following formula: ; in, Indicates the local feature The position and global feature The correlation between the positions, represents the number of pixels, represents the first matrix, represents the second matrix; The fusion feature is calculated by the following formula: ; in, represents the fusion feature, Represents the weight coefficient that is dynamically adjusted during the learning process. In the initial stage of the model, it is usually initialized to 0, but during the training process, it will gradually update according to the gradient to adjust the fusion method of the final feature map. Represents the local feature map The eigenvalues at the positions, Represents the first The feature value at each position.
[0033] It should be noted that in deep feature learning, multi-scale fusion of image features is crucial to capture rich semantic information. However, in the process of feature fusion, dealing with semantic differences between features has always been a challenge that cannot be ignored. Therefore, in order to effectively narrow the semantic gap between feature maps of different scales, an innovative Multi-scale Contextual Collaboration Module (MSCCM) is proposed to calculate the correlation between pixels in different feature maps through matrix multiplication. Subsequently, it uses the correlation as the weight vector of the high-level feature map. The greater the similarity between the feature representations of pixels at two locations, the higher their correlation.
[0034] As an example, Figure 4 As shown in the figure, first, the input feature map is processed using a convolutional layer to perform channel compression to reduce the computational load, while generating feature maps A (local features) and B (global features). Next, the local feature map A and the global feature map B are reshaped into matrices L and G, respectively, where N = H × W represents the number of pixels. L and G represent the reshaped local and global features, respectively.
[0035] Then, to measure the spatial attention, G is transposed and matrix multiplied with L, and a softmax layer is applied to compute the spatial attention map P. The spatial relationship matrix map P maps the spatial correlation between pixels: ; in, Indicates the local feature The position and global feature The correlation between the positions, represents the number of pixels, represents the first matrix, represents the second matrix; Next, a matrix multiplication operation is performed by multiplying G and the spatial attention map P to obtain a new feature map T. Each vector in T contains the integrated information of the correlation between the features of each position in the entire image and the current position.
[0036] Then, B and T are combined by element-wise addition to get the final output: ; in, represents the fusion feature, Represents the weight coefficient that is dynamically adjusted during the learning process. In the initial stage of the model, it is usually initialized to 0, but during the training process, it will gradually update according to the gradient to adjust the fusion method of the final feature map. Represents the local feature map The eigenvalues at the positions, Represents the first The feature value at each position.
[0037] It should be noted that α starts at 0 at the time of initialization. However, the weight distribution will be gradually adjusted as learning progresses.
[0038] Each position in the final feature is the weighted sum of all position features in the global feature. Since the final feature is generated from the top-level features, high-level semantic information is well preserved in the final output, greatly improving the representation ability of the deep network descriptor. This will undoubtedly have important practical value for image retrieval tasks in complex scenes.
[0039] S104, calculating the to-be-retrieved hash code corresponding to the fused feature based on the hash-guided metric loss method.
[0040] In some embodiments, a hash code to be retrieved corresponding to a fused feature is calculated based on a hash-guided metric loss method, including: performing a global average pooling operation on the fused feature to obtain a corresponding global feature vector; inputting the global feature vector into a fully connected layer to map the global feature vector to a dimension consistent with the hash code length to obtain a mapping feature; and converting the mapping feature into a binary hash code based on a sign function.
[0041] In some embodiments, the global feature vector is calculated by the following formula: ; in, represents the global eigenvector, represents the height of the feature map, represents the width of the feature map, represents the fusion feature, and are the spatial indexes in the height and width directions of the feature map, respectively. is the channel index; ; in, represents the mapping feature, represents the weight of the fully connected layer, Represents the bias parameter of the fully connected layer; ; in, Represents a binary hash code.
[0042] In some embodiments, in the process of generating hash codes, a hash-guided metric loss method is used to limit the learning range of metric items, specifically, including: calculating metric loss and correcting the error component in the metric loss based on a preset threshold to alleviate the conflict between metric loss and quantization loss; combining metric loss and quantization loss to reduce information loss in the binarization process; and optimizing model parameters using a hash-guided metric loss function.
[0043] In some embodiments, the metric loss is calculated by the following formula: ; in, represents the metric loss function, represents the set of proxy vectors, represents the embedding vector representing the input, represents the positive proxy vector, represents the boundary value, represents the scaling factor; The error component in the metric loss is corrected by the following formula: ; in, represents the improved hash-guided metric loss function, represents the set of all negative sample pairs, A hash code representing the proxy anchor. Represents the hash code of the current sample, The proxy anchor hash code representing the negative sample, express and The cosine similarity between represents the margin hyperparameter, Indicates the preset threshold; ; in, represents the minimum distance to the best global solution, Indicates the length of the hash code; In the process of set metric loss and quantization loss, the learning objective is expressed by the following formula: ; ; in, Set to 1, represents the quantization loss, represents the predicted value of the i-th sample, represents the true value of the i-th sample, Represents the square L2 norm of the predicted value and the target value, and its calculation formula is: , is the dimension of the hash code, and Respectively The sample in The predicted and target values of the dimension.
[0044] It should be noted that the design of metric loss function is crucial in image retrieval tasks. The traditional proxy-anchor method uses the cosine similarity estimated in small batches of data to design the metric loss term. Therefore, a hash-guided metric loss (HGM-Loss) is proposed to improve the linear calculation part of the cosine similarity in the proxy anchor metric loss to limit the scope of metric item learning, thereby improving the accuracy and robustness of feature extraction.
[0045] As an example, the proxy anchor is metrically calculated using the following formula: ; in, represents the metric loss function, represents the set of proxy vectors, represents the embedding vector representing the input, Represents a positive proxy vector. Each category has a proxy vector, which is continuously updated during training to maximize the distance between different categories while minimizing the distance within the same category. represents the boundary value, Represents the scaling factor.
[0046] It should be noted that is the main metric loss function, which is used to ensure that the hash code H of the sample is consistent with its corresponding proxy anchor point H p It has a good similarity relation in the hash space. Specifically, it achieves this goal by maximizing the similarity between similar samples and minimizing the similarity between different samples.
[0047] Next, we introduce the preset threshold , by introducing a specific threshold To correct the calculation error component in the proxy anchor metric loss, so as to alleviate the conflict between metric loss and quantization loss. The improved hash-guided metric loss function is expressed by the following formula: ; in, represents the improved hash-guided metric loss function, represents the set of all negative sample pairs, A hash code representing the proxy anchor. Represents the hash code of the current sample, The proxy anchor hash code representing the negative sample, express and The cosine similarity between represents the margin hyperparameter, Indicates the preset threshold; For the first term of the formula, the similarity between the positive sample pairs is calculated and passed To control the influence of similarity, if the similarity is lower than , it does not contribute to the loss; otherwise, the loss increases according to the similarity.
[0048] The second item calculates the similarity between negative sample pairs and passes To control the influence of similarity, contrary to the first item, we hope that the similarity of negative sample pairs is as low as possible, so the negative sign is introduced.
[0049] It should be noted that the preset threshold The setting depends on the length of the binary code and the number of categories, which represents the minimum distance to the best global solution The purpose of HGM loss is to push the distance between different categories to the minimum distance Determining the appropriate threshold , preset threshold Calculated by the following formula: ; in, represents the minimum distance to the best global solution, Indicates the length of the hash code; In the process of combining metric loss and quantization loss, the learning objective is expressed by the following formula: ; ; in, Set to 1, represents the quantization loss, represents the predicted value of the i-th sample, Represents the true value of the i-th sample.
[0050] The model parameters are optimized using a hash-guided metric loss function, which includes the following steps: 1) Initialization parameters: At the beginning of training, all parameters of the network (including proxy vectors and other model weights) are randomly initialized.
[0051] 2) Forward propagation: The input image passes through the network’s feature extraction module to generate feature representation, and then is binarized through the hash coding module.
[0052] 3) Using hash-guided metric loss function And the quantization loss function Calculate the total loss.
[0053] 4) Back propagation: The model parameters are updated by calculating the gradient of the loss function, so that the representation capabilities of the feature extraction and hash coding modules are gradually optimized.
[0054] 5) Parameter update: Use the optimization algorithm (SGD) to update the model parameters and iterate the training until the model converges.
[0055] S105, calculating the similarity between the image to be retrieved and the database image based on the hash code to be retrieved and the hash code of the database image, and determining the target image according to the similarity.
[0056] As an example, this step is mainly completed by comparing the hash codes of the image to be retrieved with the image in the database and calculating the similarity between the two. Specifically, it includes: first, input feature coding: input the image to be retrieved into the trained deep network, extract its feature representation, and generate the corresponding binary hash code through the hash coding module. Suppose the hash code of the image to be retrieved is Each image in the database has also been extracted features and hashed by the same network. Let the hash code of the 𝑖th image in the database be . Next, the similarity of the hash code is calculated. The hash code of the image to be retrieved is calculated using the Hamming distance The hash code of each image in the database The Hamming distance is defined as the number of different positions between two binary codes of the same length, that is: ; in, The length of the hash code; and Respectively and In the The binary value of the bit.
[0057] Then, perform similarity sorting: calculate the Hamming distance between the image to be retrieved and all images in the database, and sort these distances in ascending order. The smaller the distance, the higher the similarity. The sorting result is a list of images sorted in descending order of similarity, that is, from the most similar to the image to be retrieved to the least similar.
[0058] Finally, the search results are returned: according to the sorting results, the top N most similar images are selected as search results and returned to the user. These images are the images in the database with the highest similarity to the image to be searched.
[0059] In addition, in order to better illustrate the retrieval effect of the image retrieval method proposed in the present invention, Figure 5 and Figure 6 For example, Figure 5 is the retrieval effect of the image retrieval method proposed in the embodiment of the present invention on the ImageNet dataset; and Figure 6 The retrieval effect of the image retrieval method proposed in the embodiment of the present invention on the CIFAR-10 dataset. Figure 5 and Figure 6 In the figure, the left column Query is the input image, and the five columns on the right are the top five most similar images retrieved.
[0060] In summary, according to the image retrieval method of the embodiment of the present invention, first, the image to be retrieved is obtained, and the feature extraction of the image to be retrieved is performed to obtain the depth feature map corresponding to the image to be retrieved; then, each channel of the depth feature map is weighted to generate a weighted feature map; then, the local features and global features of the weighted feature map are extracted, and feature fusion is performed based on the local features and global features to generate fusion features; then, the hash code to be retrieved corresponding to the fusion feature is calculated based on the hash-guided metric loss method; then, the similarity between the image to be retrieved and the database image is calculated based on the hash code to be retrieved and the database image hash code, and the target image is determined according to the similarity. Thereby, the loss of important semantic information in the binarization process is effectively prevented, thereby improving the accuracy of image retrieval.
[0061] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium on which an image retrieval program is stored. When the image retrieval program is executed by a processor, the image retrieval method described above is implemented.
[0062] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0063] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0064] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0065] In the description of the present invention, it is to be understood that the terms “center”, “longitudinal”, “lateral”, “length”, “width”, “thickness”, “up”, “down”, “front”, “back”, “left”, “right”, “vertical”, “horizontal”, “top”, “bottom”, “inside”, “outside”, “clockwise”, “counterclockwise”, “axial”, “radial”, “circumferential”, etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0066] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0067] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0068] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, a first feature being "above", "above" or "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below", "below" or "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.
[0069] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. An image retrieval method, characterized in that: The following steps are involved: Acquire an image to be retrieved, and perform feature extraction on the image to be retrieved to obtain a depth feature map corresponding to the image to be retrieved; Performing weighted processing on each channel of the depth feature map to generate a weighted feature map; Extracting local features and global features of the weighted feature map, and performing feature fusion based on the local features and the global features to generate fused features; Calculate the to-be-retrieved hash code corresponding to the fused feature based on a hash-guided metric loss method; Calculating the similarity between the image to be retrieved and the database image based on the hash code to be retrieved and the hash code of the database image, and determining the target image according to the similarity; The process of weighting each channel of the depth feature map to generate a weighted feature map includes: Performing global average pooling on each channel of the deep feature map to generate a global spatial compressed representation; Calculate a weight vector according to the global space compressed representation based on an activation function; A principal element multiplication operation is performed on the depth feature map according to the weight vector to generate a weighted feature map.
2. The image retrieval method according to claim 1, wherein: The global space compression representation is calculated by the following formula: ; in, represents the global space compressed representation, represents the height of the feature map, represents the width of the feature map, Represents the spatial position in the feature map Place The characteristic value of each channel; ; in, express Activation function, express Activation function, Represents the element value of the feature map; ; ; ; in, It represents the generation of an intermediate representation containing global information. After this process, the feature map is compressed and converted into information that is more suitable for downstream tasks. Then, Z enters the second fully connected layer to further generate the channel weight vector W. express Activation function, and represents the bias vector, and represents the weight matrix, represents the weight vector, represents the weighted feature map, Represents the component in the weight vector for each channel.
3. The image retrieval method according to claim 1, wherein: Extracting local features and global features of the weighted feature map, and performing feature fusion based on the local features and the global features to generate fused features, including: Processing the weighted feature map using a convolutional layer to generate local features and global features; Reshaping the local features into a first matrix, and reshaping the global features into a second matrix; Transposing the second matrix, performing matrix multiplication on the transposed second matrix and the first matrix, and calculating the spatial attention map using a softmax layer; Performing matrix multiplication on the spatial attention map and the second matrix to obtain a new feature map; The new feature map and the global feature are combined based on element-by-element addition to generate a fused feature.
4. The image retrieval method according to claim 3, characterized in that: The spatial attention map is calculated by the following formula: ; in, Indicates the local feature The position and global feature The correlation between the positions, represents the number of pixels, represents the first matrix, represents the second matrix; The fusion feature is calculated by the following formula: ; in, represents the fusion feature, Represents the weight coefficient that is dynamically adjusted during the learning process; Represents the local feature map The eigenvalues at the positions, Represents the first The feature value at each position.
5. The image retrieval method according to claim 1, wherein: Calculating the to-be-retrieved hash code corresponding to the fused feature based on a hash-guided metric loss method includes: Performing a global average pooling operation on the fused features to obtain a corresponding global feature vector; Inputting the global feature vector into a fully connected layer to map the global feature vector to a dimension consistent with the hash code length to obtain a mapping feature; The mapping feature is converted into a binary hash code based on a sign function.
6. The image retrieval method according to claim 5, characterized in that: The global eigenvector is calculated by the following formula: ; in, represents the global eigenvector, represents the height of the feature map, represents the width of the feature map, represents the fusion feature, and are the spatial indexes in the height and width directions of the feature map, respectively. is the channel index; ; in, represents the mapping feature, represents the weight of the fully connected layer, Represents the bias parameter of the fully connected layer; ; in, Represents a binary hash code.
7. The image retrieval method according to claim 1, wherein: In the process of generating the hash code, a hash-guided metric loss method is used to limit the learning range of the metric item, specifically, including: Calculate the measurement loss and correct the error component in the measurement loss based on a preset threshold to alleviate the conflict between the measurement loss and the quantization loss; Combining metric loss and quantization loss to reduce information loss in the binarization process; The model parameters are optimized using a hash-guided metric loss function.
8. The image retrieval method according to claim 7, characterized in that: The metric loss is calculated by the following formula: ; in, represents the metric loss function, represents the set of proxy vectors, represents the embedding vector of the input, represents the positive proxy vector, represents the boundary value, represents the scaling factor; The error component in the metric loss is corrected by the following formula: ; in, represents the improved hash-guided metric loss function, represents the set of all negative sample pairs, A hash code representing the proxy anchor. Represents the hash code of the current sample, The proxy anchor hash code representing the negative sample, express and The cosine similarity between represents the margin hyperparameter, Indicates the preset threshold; ; in, represents the minimum distance to the best global solution, Indicates the length of the hash code; In the process of combining metric loss and quantization loss, the learning objective is expressed by the following formula: ; ; in, Set to 1, represents the quantization loss, represents the predicted value of the i-th sample, represents the true value of the i-th sample, Represents the square L2 norm of the predicted value and the target value, and its calculation formula is: , is the dimension of the hash code, and Respectively The sample in The predicted and target values of the dimension.
9. A computer-readable storage medium, characterized in that: An image retrieval program is stored thereon, and when the image retrieval program is executed by a processor, the image retrieval method as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Pedestrian hash retrieval based on loss measurement in depth learning networks
CN109241317A
Deep hash image retrieval method based on feature pyramid under attention mechanism
CN111625675A
Image hash retrieval method based on hierarchical feature complementation
CN112084362A
Fine-grained bird image retrieval method based on graph neural network and deep hash
CN114329031A
No-reference screen content image quality evaluation method based on edge feature guidance
CN115797304A
Cited By
Deep hash image retrieval method and device
CN120929631A
Deep hashing image retrieval method and apparatus
CN120929631B