An image retrieval method integrating fully connected layers and deep convolutional features
By fusing the fully connected layer and deep convolution features in image retrieval technology and using the MCR-MAC feature extraction module, the problems of loss of feature information and neglected associated information caused by ordinary pooling operations are solved, and more efficient image retrieval accuracy is achieved.
Patent Information
- Application Number
- CN202310606770.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-05-26
AI Technical Summary
When existing image retrieval technology uses deep convolution features, ordinary pooling operations lead to loss of feature information and neglect of associated information, making it difficult to effectively distinguish similar and unrelated image samples.
A method of image retrieval that fuses fully connected layer and deep convolutional features is proposed. Through the MCR-MAC feature extraction module, the maximum value correction and region weighted hybrid pooling is combined to enhance the robustness of features, and the weight vector weighted features of the fully connected layer is improved to improve the accuracy of image retrieval.
It significantly improves the accuracy of image retrieval, can more effectively utilize the local features and channel weight information of the image, and improves the performance of image retrieval.
Smart Images

Figure CN116701689B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and deep learning technology, and specifically relates to an image retrieval method integrating a fully connected layer and deep convolution features. Background Art
[0002] The image retrieval task is mainly to perform similarity search on specific images and obtain target images similar to the query image, which provides a reliable solution for many software applications involving image management and image search functions. At present, image retrieval technology has been widely used in shopping software product search, digital library, satellite remote sensing image search and other fields, providing many conveniences for practical problems such as commodity trading, digital image information management, military facility construction and agricultural landform exploration. Therefore, image retrieval technology has important research significance for various fields and various practical application scenarios.
[0003] Image retrieval technology originated in the 1970s. Due to the development of image processing technology, its research hotspot has shifted from the initial text-based image retrieval method to the current mainstream content-based image retrieval method. In the process of image processing technology development, the features used in the field of image retrieval are mainly low-level features defined by people based on the visual meanings expressed numerically according to the pixel values themselves and the correlation between them. After the emergence of the deep neural network AlexNet in 2012, people began to use the deep convolution features extracted by the convolutional neural network for image retrieval tasks, and explored many methods for post-processing deep convolution features to improve the accuracy of retrieval. Due to the single sampling method of common pooling operations such as maximum pooling, mean pooling, and sum pooling, more feature information will be lost, and the correlation information between the feature values inside the feature map will be ignored, which is not enough to distinguish the similarities between related samples and the differences between unrelated samples. In order to make up for the shortcomings of ordinary pooling, relevant researchers have proposed a variety of methods, which focus on the following three aspects: (1) extracting local features on the convolutional feature map, mainly to obtain local correlation information on the feature map, usually implemented by sliding window pooling operation; (2) evaluating the importance of each channel of the convolutional feature map, mainly to eliminate some redundant features and noise interference, usually using some statistical information of feature values and dimensionality reduction methods; (3) fine-tuning the model on a specific dataset, mainly to improve the impact of dataset differences and to solve the inconsistency problem between classification tasks and retrieval tasks to a certain extent.
[0004] At present, Yue-Hei et al. found that the fully connected layer performs worse than the shallow convolutional features in image retrieval tasks due to the loss of spatial correlation. In order to evaluate the importance of feature channels, Kalantidis et al. introduced the number of non-zero response values in the feature map as statistical information, assigned weights to each feature channel, and constructed spatial weights based on the independent responses of the feature map to improve the model's attention to the target area. The disadvantage is that the calculation is complex and the association of local information is ignored. Tolias et al. used the idea of sliding windows, performed overlapping integral pooling operations on the feature map, and accumulated the response values of each region as the feature value of the current channel. They proposed a convolution feature post-processing method called Regional Maximum Activation of Convolutions (R-MAC), which solved the problem of target positioning in instance-level image retrieval and utilized the local correlation information of the feature map to a certain extent. However, the disadvantage is that it did not pay attention to the importance of feature channels. Kim et al. introduced a regional attention network based on the R-MAC feature to eliminate the influence of background clutter and changes in regional importance, and further improved the accuracy of image retrieval.
[0005] Weinzaepfe et al. used iterative attention modules to construct super features, forming an ordered set with local distinguishing properties, trained the super features with contrast loss as a constraint, and achieved excellent retrieval results on landmark datasets. In order to build more robust features, Cao et al. embedded global features and local features in a deep network, introduced generalized mean pooling and attention selection mechanisms for both, and performed end-to-end model training by balancing the gradient flows of the two.
[0006] In summary, the current image retrieval work based on deep convolutional features can be summarized as follows: (1) Consider the acquisition of local features to enhance the robustness of features; (2) Consider the importance of feature channels and assign weights to feature channels to eliminate the impact of redundant information and noise interference on image retrieval tasks; (3) Retrain the model parameters to extract more effective convolutional features.
[0007] Based on the above analysis, we hope to propose an image retrieval method that takes into account global features, local features and channel weights. Summary of the invention
[0008] In view of the above-mentioned deficiencies in the prior art, the present invention provides an image retrieval method that integrates fully connected layers and deep convolutional features, which can greatly improve the accuracy of image retrieval.
[0009] An image retrieval method integrating a fully connected layer and deep convolution features, characterized in that it comprises the following steps:
[0010] Step 1: construct a training set consisting of multiple triplets, where the triplets include a query sample, and positive samples and negative samples corresponding to the query sample;
[0011] Step 2: Construct an image retrieval framework, specifically: first construct a deep convolutional neural network model that does not include a fully connected layer, use the last maximum pooling layer as the output layer, and connect the MCR-MAC (Maximum correction and region-weighted mixing pooled activation of convolution) feature extraction module after the output layer; connect the adaptive maximum pooling layer and the fully connected layer in sequence between the last convolutional layer and the output layer of the deep convolutional neural network model, and finally activate the output of the fully connected layer with the Softmax function;
[0012] Among them, the output result of the output layer of the deep convolutional neural network model is a deep convolution feature map of size w×h×c, which is composed of c feature maps of size w×h;
[0013] The MCR-MAC feature extraction module includes a maximum correction module and a regional weighted mixed pooling convolution activation module;
[0014] The maximum correction module calculates the sampling frequency distribution of each eigenvalue position at L scales for the deep convolution feature map of size w×h×c, and then calculates the sampling frequency distribution of each eigenvalue position under the superposition of L scales to obtain the maximum sampling frequency value and the sampling frequency value corresponding to the position of the maximum eigenvalue in the deep convolution feature map of size w×h×c; the ratio of the maximum sampling frequency value to the sampling frequency value corresponding to the position of the maximum eigenvalue is used as the first correction coefficient, and multiplied by the maximum pooling response value of each w×h feature map to obtain the correction value of the maximum pooling response value of each w×h feature map;
[0015] The regional weighted mixed pooling convolution activation module adopts sliding window mean pooling and maximum pooling at L scales to calculate the weighted mixed pooling response value of each window area, specifically: for each feature map of size w×h, the ratio of the mean pooling value of each window area at the lth, l=1, 2, .., L scales to the maximum pooling value is used as the second correction coefficient, the second correction coefficient is added by 1 and multiplied by the maximum pooling value to obtain the weighted mixed pooling response value of each window area at the lth, l=1, 2, .., L scales;
[0016] For each w×h feature map, the correction value of its maximum pooling response value is summed with the weighted mixed pooling response value of each window area under the superposition of L scales to obtain its final response value; then the feature vector of the w×h×c deep convolutional feature map composed of the final response values of c w×h feature maps is obtained, that is, the MCR-MAC feature vector;
[0017] The output result of the last convolutional layer of the deep convolutional neural network model is a deep convolutional feature map of size 2w×2h×c;
[0018] The adaptive maximum pooling layer compresses the deep convolution feature map of size 2w×2h×c into a deep convolution feature map of a preset fixed size;
[0019] The fully connected layer converts the output of the adaptive maximum pooling layer into an output vector of the same dimension as the MCR-MAC feature vector through linear projection, and then activates it with the Softmax function to obtain a weight vector;
[0020] Finally, the MCR-MAC feature vector is multiplied by the weight vector to obtain the weighted MCR-MAC feature vector as the output of the image retrieval framework;
[0021] Step 3: Based on the training set, first train the convolutional layer in the deep convolutional neural network model to obtain the trained deep convolutional neural network model; then freeze the convolutional layer parameters of the trained deep convolutional neural network model, train the fully connected layer to obtain the trained fully connected layer; and then obtain the trained image retrieval framework;
[0022] Step 4: Input the query image to be tested into the trained image retrieval framework, and output the image retrieval results after sorting according to the Euclidean distance between the feature vector of the query image and the feature vectors of all images in the training set.
[0023] Furthermore, the process of calculating the sampling frequency value of each eigenvalue at the lth, l=1,2,..,Lth scale in step 2 is specifically as follows:
[0024] First, determine the left and right boundary values of the sliding window at the lth, l=1, 2, .., Lth scale to obtain the corresponding interval range, and then calculate the prefix sum vector of each w×h feature map by the interval prefix sum calculation method; according to the horizontal and vertical coordinates of the position of each eigenvalue, query the corresponding element values in the prefix sum vector of the feature map where it is located, and multiply them to obtain the sampling frequency value of each eigenvalue at the lth, l=1, 2, .., Lth scale;
[0025] Then, we get the maximum sampling frequency value And the sampling frequency value corresponding to the location of the maximum eigenvalue in the deep convolution feature map According to the nth, n=1,2,..,c w×h size feature map F (n) The maximum pooling response value of Calculate the nth w×h feature map F (n) The corrected value of the maximum pooled response is Where K is the length of the prefix sum vector; The horizontal coordinate of the position of the maximum eigenvalue in the deep convolution feature map is the element value queried in the prefix sum vector of the feature map where it is located; The ordinate of the position of the maximum eigenvalue in the deep convolutional feature map is the element value queried in the prefix sum vector of the feature map where it is located.
[0026] Furthermore, the deep convolutional neural network model in step 2 specifically adopts VGG-16.
[0027] Furthermore, the value of L is a positive integer greater than 1.
[0028] Furthermore, the preset fixed size is 7×7×c.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] The present invention proposes an image retrieval method integrating a fully connected layer and deep convolution features. On the basis of the existing R-MAC convolution feature post-processing method, an MCR-MAC convolution feature post-processing method is proposed through sampling frequency correction based on interval prefix sum calculation and weight correction of regional response. While considering feature information at different scales, more attention is paid to the aggregation of local features, and the weight vector output by the fully connected layer is further integrated to measure the importance of each feature map. In addition, when performing model training based on a training set in the form of triple data, the convolution layer parameters are frozen to retain the MCR-MAC features and give the fully connected layer the feature combination capability. In summary, the present invention can significantly improve the accuracy of the image retrieval algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a schematic diagram of the structure of the image retrieval framework proposed in Example 1 of the present invention;
[0032] Figure 2 This is a schematic diagram of feature extraction of the MCR-MAC feature extraction module in Example 1 of the present invention;
[0033] Figure 3 This is a schematic diagram of the interval prefix sum calculation method in Example 1 of the present invention. DETAILED DESCRIPTION
[0034] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and embodiments.
[0035] Example 1
[0036] This embodiment proposes an image retrieval method that integrates a fully connected layer and deep convolution features, including the following steps:
[0037] Step 1: Construct a training set consisting of multiple triplets, where the triplets include a query sample, and positive and negative samples corresponding to the query sample, all of which are images with a resolution of 724×724.
[0038] Step 2: Build an image retrieval framework with the following structure: Figure 1 As shown in the figure, specifically: first build a deep convolutional neural network model that does not include a fully connected layer (specifically using VGG-16), take the last maximum pooling layer as the output layer, and connect the MCR-MAC feature extraction module after the output layer; connect the adaptive maximum pooling layer and the fully connected layer (three layers in total) in sequence between the last convolutional layer and the output layer of the deep convolutional neural network model, and finally use the Softmax function to activate the output of the fully connected layer.
[0039] Among them, the output result of the output layer of the deep convolutional neural network model is a deep convolutional feature map of size 22×22×512, which is composed of 512 feature maps of size 22×22.
[0040] like Figure 2 As shown, the MCR-MAC feature extraction module includes a maximum correction module and a regional weighted mixed pooling convolution activation module.
[0041] The maximum correction module (i.e., the maximum correction based on the two-dimensional prefix sum matrix) calculates the sampling frequency distribution of each eigenvalue at three scales for the 22×22×512 size deep convolution feature map. Specifically, the pooling area size of each sliding window at the lth, l=1, 2, and 3 scales is S×S, where Therefore, the number of pooling areas at the three scales is 1, 4, and 9 respectively; determine the left and right boundary values of the sliding window at the lth, l=1, 2, and 3rd scales to obtain the corresponding interval range, and then calculate the prefix sum vector of each 22×22 feature map by the interval prefix sum calculation method, with a length of K=23; according to the horizontal and vertical coordinates of the position of each eigenvalue, query the corresponding element values in the prefix sum vector of the feature map where it is located, and multiply them to obtain the sampling frequency value of each eigenvalue at the lth scale;
[0042] Take the interval prefix sum calculation of a 7×7 feature map at the third scale as an example, Figure 3As shown in the figure, the intervals of each sliding window are [1,3], [3,5], and [5,7], and different depths of colors are used to represent the sampling frequencies of 1, 2, and 4, respectively. Then, the prefix sum vector (1,1,2,1,2,1,1) is obtained. According to the horizontal and vertical coordinates of the position of each eigenvalue, the corresponding element values are queried in the prefix sum vector of the feature graph where they are located, and the sampling frequency values of each eigenvalue at the third scale are obtained by multiplying them.
[0043] Then calculate the sampling frequency distribution of each eigenvalue location under the superposition of three scales to obtain the maximum sampling frequency value And the sampling frequency value corresponding to the position of the maximum eigenvalue in the deep convolution feature map of size w×h×c According to the nth, n=1,2,..,c 22×22 feature map F (n) The maximum pooling response value of Calculate the nth 22×22 feature map F (n) Corrected value of the maximum pooled response value in, The horizontal coordinate of the position of the maximum eigenvalue in the deep convolution feature map is the element value queried in the prefix sum vector of the feature map where it is located; The ordinate of the position of the maximum eigenvalue in the deep convolutional feature map is the element value queried in the prefix sum vector of the feature map where it is located.
[0044] The regional weighted hybrid pooling convolution activation module adopts sliding window mean pooling and maximum pooling at three scales to calculate the weighted hybrid pooling response value of each window area. Specifically, for each 22×22 feature map, the total number of pooling areas at three scales is 1+4+9=14, and the mean pooling value avgPooling(R (i) ), i = 1, 2, ..., 14 and the maximum pooling value maxPooling (R (i) ),i=1,2,...,14 as the second correction coefficient, the second correction coefficient plus 1 is equal to the maximum pooling value maxPooling(R (i) ) to obtain the weighted mixed pooling response value of each window area under each pooling area.
[0045]
[0046] Among them, ε is to ensure that the zero division error that causes the maximum pooling response to be zero when pooling the all-zero area is avoided, so a smaller number ε needs to be added to the denominator, set to 0.1; when l = 1,
[0047] For each 22×22 feature map, the correction value R of the maximum pooling response value c The weighted mixed pooling response value (R m ) and get the final response value R mcr-mac ; Then we get the feature vector of the deep convolution feature map of size 22×22×512, which is composed of the final response values of 512 feature maps of size 22×22, namely the MCR-MAC feature vector
[0048] The output result of the last convolutional layer of the deep convolutional neural network model is a deep convolutional feature map of size 45×45×512;
[0049] The adaptive maximum pooling layer compresses the 45×45×512-sized deep convolutional feature map into a 7×7×512-sized deep convolutional feature map by automatically adjusting the size and step size of the pooling kernel;
[0050] The fully connected layer converts the output of the adaptive maximum pooling layer into an output vector of the same dimension as the MCR-MAC feature vector through linear projection, and then activates it with the Softmax function to obtain the weight vector
[0051] Finally, the MCR-MAC feature vector With the weight vector Perform point multiplication to obtain the weighted MCR-MAC feature vector As the output of the image retrieval framework.
[0052] Step 3: Based on the training set, first train the convolutional layer in the deep convolutional neural network model to obtain the trained deep convolutional neural network model; then freeze the convolutional layer parameters of the trained deep convolutional neural network model, train the fully connected layer to obtain the trained fully connected layer; and then obtain the trained image retrieval framework;
[0053] Among them, the training loss function selects the ternary contrast loss function.
[0054] Step 4: Input the query image to be tested into the trained image retrieval framework, and output the image retrieval results after sorting according to the Euclidean distance between the feature vector of the query image and the feature vectors of all images in the training set.
[0055] The image retrieval method (Weighted-MCR-MAC[M]+Fine-tune+QE) proposed in this embodiment that integrates the fully connected layer and the deep convolutional features is used for verification tests on the Oxford5k dataset and the Paris6k dataset (medium resolution 724×724), and the retrieval accuracy (mAP) is compared with that of the existing image retrieval method. The data is shown in Table 1. It can be seen that the method proposed in this embodiment can further improve the retrieval accuracy. Among them, the existing image retrieval methods compared include: R-MAC, VGG-16-GeM, Crow, GeM+SOLAR, R101-DELG, Shen et al. (Shen X, Lin Z, Brandt J, et al. Spatially-constrained similarity measure for large-scale object retrieval [J]. IEEE transactions on pattern analysis and machine intelligence, 2013, 36 (6): 1229-1241), Deng et al. (Deng C, Ji R, Liu W, et al. Visual reranking through weakly supervised multi-graph learning [C] / / Proceedings of the IEEE International Conference on Computer Vision. 2013: 2600-2607).
[0056] Table 1. Comparison of retrieval accuracy
[0057]
[0058] Although the above describes the illustrative specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations using the concept of the present invention are protected.
Claims
1. An image retrieval method that combines fully connected layers with deep convolutional features. It is characterized in that The following steps are involved: Step 1: Construct a training set consisting of multiple triplets, wherein the triplets include query samples, positive samples, and negative samples; Step 2: Construct an image retrieval framework, specifically: first construct a deep convolutional neural network model that does not include a fully connected layer, use the last maximum pooling layer as the output layer, and connect the MCR-MAC feature extraction module after the output layer; connect the adaptive maximum pooling layer and the fully connected layer in sequence between the last convolutional layer and the output layer of the deep convolutional neural network model, and finally use the Softmax function to activate the output of the fully connected layer; Among them, the output result of the output layer of the deep convolutional neural network model is a deep convolution feature map of size w×h×c; The MCR-MAC feature extraction module includes a maximum correction module and a regional weighted mixed pooling convolution activation module; The maximum correction module calculates the sampling frequency distribution of each eigenvalue at L scales for the deep convolution feature map of size w×h×c, and then calculates the sampling frequency distribution of each eigenvalue at the superposition of L scales to obtain the maximum sampling frequency value and the sampling frequency value corresponding to the position of the maximum eigenvalue; the ratio of the maximum sampling frequency value to the sampling frequency value corresponding to the position of the maximum eigenvalue is used as the first correction coefficient, and multiplied by the maximum pooling response value of each w×h size feature map to obtain the correction value of its maximum pooling response value; The regional weighted mixed pooling convolution activation module adopts sliding window mean pooling and maximum pooling at L scales to calculate the weighted mixed pooling response value of each window area, specifically: for each feature map of size w×h, the ratio of the mean pooling value of each window area at the lth, l=1, 2, .., L scales to the maximum pooling value is used as the second correction coefficient, the second correction coefficient is added by 1 and multiplied by the maximum pooling value to obtain the weighted mixed pooling response value of each window area at the lth, l=1, 2, .., L scales; For each w×h feature map, the correction value of its maximum pooling response value is summed with the weighted mixed pooling response value of each window area under the superposition of L scales to obtain its final response value; then the feature vector of the w×h×c deep convolutional feature map composed of the final response values of c w×h feature maps is obtained, that is, the MCR-MAC feature vector; The output result of the last convolutional layer of the deep convolutional neural network model is a deep convolutional feature map of size 2w×2h×c; The adaptive maximum pooling layer compresses the deep convolution feature map of size 2w×2h×c into a deep convolution feature map of a preset fixed size; The fully connected layer converts the output of the adaptive maximum pooling layer into an output vector of the same dimension as the MCR-MAC feature vector through linear projection, and then activates it with the Softmax function to obtain a weight vector; Finally, the MCR-MAC feature vector is multiplied by the weight vector to obtain the weighted MCR-MAC feature vector as the output of the image retrieval framework; Step 3: Based on the training set, first train the convolutional layer in the deep convolutional neural network model to obtain the trained deep convolutional neural network model; then freeze the convolutional layer parameters of the trained deep convolutional neural network model, train the fully connected layer to obtain the trained fully connected layer; and then obtain the trained image retrieval framework; Step 4: Input the query image to be tested into the trained image retrieval framework, and output the image retrieval results after sorting according to the Euclidean distance between the feature vector of the query image and the feature vectors of all images in the training set.
2. According to the image retrieval method of integrating the fully connected layer and the deep convolution feature according to claim 1, It is characterized in that The specific process of calculating the sampling frequency value of each eigenvalue at the lth, l=1,2,..,Lth scale in step 2 is: First, determine the left and right boundary values of the sliding window at the lth, l=1, 2, .., Lth scale to obtain the corresponding interval range, and then calculate the prefix sum vector of each w×h feature map through the interval prefix sum calculation method; according to the horizontal and vertical coordinates of the position of each eigenvalue, query the corresponding element values in the prefix sum vector of its feature map, and multiply them to obtain the sampling frequency value of each eigenvalue at the lth, l=1, 2, .., Lth scale.
3. According to the image retrieval method of integrating the fully connected layer and the deep convolutional features of claim 1, It is characterized in that The deep convolutional neural network model in step 2 specifically uses VGG-16.
4. According to the image retrieval method of integrating the fully connected layer and the deep convolution feature according to claim 1, It is characterized in that The preset fixed size is 7×7×c.
5. According to the image retrieval method of integrating the fully connected layer and the deep convolution feature according to claim 1, It is characterized in that The value of L is a positive integer greater than 1.
Citation Information
Patent Citations
Cross-media image retrieval method and system
CN113536013A
Similarity-based detection of prominent objects using deep CNN pooling layers as features
US20170083792A1