A UAV Image Hashing Retrieval Method Based on a Salience Capture Mechanism
Through the combination of significance capture mechanism and local fine-grained information, a new drone image hash retrieval method was designed, which solves the problem of high storage space and retrieval complexity in the existing methods, improves the retrieval accuracy and optimizes hash code learning.
Patent Information
- Application Number
- CN202310007898.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-01-04
AI Technical Summary
The existing drone image retrieval methods require a large amount of storage space, have high spatial and temporal complexity, and pay too much attention to global information and ignore the key information of fine-grained significance.
By learning the semantic information of the drone image data, using the significance capture mechanism, distribution smoothing terms and local fine-grained information, a new hash code learning method is designed, combining the information extraction module and the significance capture module, and using the objective function composed of similarity maintenance terms, distribution smoothing terms and quantization errors to optimize the hash code learning process.
It improves the accuracy of drone image retrieval, reduces the space-time complexity of retrieval and reduces the storage space requirements, and optimizes the hash code learning process.
Smart Images

Figure CN116089646B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for retrieving drone images by a saliency capture mechanism, which is specifically applicable to improving the retrieval progress. Background Art
[0002] With the rapid development of drone technology, the retrieval of images taken by drones has received extensive attention in the field of image processing. Compared with satellites, drones usually have a real-time streaming function and can enable quick decision-making. In addition, drones can significantly reduce the dependence on weather conditions and provide higher flexibility in dealing with various problems. As the number of drones increases, the number of images taken by drones will also increase significantly. Therefore, how to mine effective drone image information has become increasingly important. In order to mine useful information, many researchers have paid great attention to the research on the retrieval of drone image data. Because drone data retrieval can quickly retrieve useful information and has been applied in many aspects such as agriculture and military. Drone image retrieval is a branch under general image retrieval and focuses more on the image data taken by drones in terms of the content to be retrieved.
[0003] With the explosive growth of drone-shot data, efficient analysis techniques for ground images have received urgent attention in dealing with drone data. The task of drone image retrieval is to retrieve relevant drone images using drone image data. Due to the large amount of data and the significant differences in information at different scales of the data, it is difficult for users to quickly obtain favorable information. How to solve the multi-scale problem of drone image data is an important challenge for the task of drone image retrieval.
[0004] In recent years, many scholars have used deep learning methods to solve the problem of drone image data retrieval. A common approach is to encode all drone image data into their corresponding features and then calculate the similarity of different images in a common feature space. Although existing drone image retrieval methods have made certain progress, there are still several deficiencies: 1) They require a large amount of storage space, and the spatio-temporal complexity of retrieval is relatively low; 2) Existing hashing methods pay too much attention to global information and ignore the key information of fine-grained saliency. Summary of the Invention
[0005] The object of the present invention is to address the above deficiencies, learn the semantic information of drone image data, use a saliency capture mechanism, a distribution smoothing term, global information, and local fine-grained information to learn effective hash codes, and finally use similarity calculation to retrieve a given number of drone image items. The present invention fully utilizes the fine-grained key information of drone images to provide a method for retrieving drone images by a saliency capture mechanism that further improves the retrieval performance.
[0006] To achieve the above object, the technical solution of the present invention is:
[0007] A method for retrieving drone images by hashing based on a saliency capture mechanism, the method comprising the following steps:
[0008] Step 1, divide the photos in the drone image library into a training data set and a test data set;
[0009] Step 2, information extraction, improve based on the pre-trained ResNet50 network, and use the pictures in the training data set to perform information extraction training on the ResNet50 network;
[0010] Use the pictures in the training data set to train and extract features by the ResNet50 network. The ResNet50 network performs four stages of feature mapping processing on the pictures in the training data set. First, upsample the feature map output generated by the first stage of ResNet50, and then connect it with the feature mapping output of the second stage of ResNet50 to form the local low-level feature F low ; After that, upsample the feature mapping output of the third stage of ResNet50, and then connect it with the feature mapping output of the fourth stage of ResNet50 to form the local high-level feature F high ; Finally, process the local low-level feature with a 3×3 convolution, and process the local high-level feature with a 1×1 convolution so that these two features have the same size, and connect these two features to become the connected feature; in order to avoid the loss of high-level feature semantics, use the residual structure to connect the average value of the local high-level feature and the connected feature to obtain the local joint feature F j ; In addition, performing fine-grained transformation on the local joint feature can reduce redundant information;
[0011] Step 3, saliency capture, after generating the local fine-grained feature F l To enhance the effectiveness of the feature, perform saliency capture processing; first capture the information interaction attention, and then capture the visual enhancement attention;
[0012] The capture mechanism of the information interaction attention is to make the global feature and the local fine-grained feature learn from each other to obtain the feature embedding vector F captured by the information interaction attention ia ; The capture mechanism of the visual enhancement attention is to enhance the visual representation of the extracted effective features to obtain the saliency feature F output by the saliency module va ;
[0013] Step 4, perform hashing learning training, and use the saliency feature F output by the saliency module obtained in step 3 vaAfter that, it is input into the hash learning module for training, that is, the fully connected hash layer of k nodes. The hash used the tanh function as the activation function; in the training stage, k-bit hash class codes are generated and learned through an objective function composed of a similarity maintenance term, a distribution smoothing term, and a quantization error; in the testing stage, the sign function is used to quantize the k-bit class hash code into a k-bit hash code.
[0014] Step 5: Train the saliency capture model. Use the training dataset to loop through steps 2 to 4 to train the network model. When the training reaches 100 iterations or the loss of the final objective function no longer decreases, end the algorithm operation, and then obtain the trained overall network model to calculate the hash codes of the samples in the test dataset.
[0015] Step 6: Use the trained overall network model to calculate the hash codes of the samples in the test dataset. Sort the Hamming distances between the query sample and the hash codes of each sample in the training dataset from large to small, and calculate the top n precisions of the ranking list to obtain the mean average precision (MAP) metric and the top n retrieval results. At this time, the retrieval results are output and the retrieval is completed.
[0016] In step 2, the pictures in the training dataset are trained to extract features using the ResNet50 network. The ResNet50 network performs four stages of feature mapping processing on the pictures in the training dataset. The pictures in the training dataset are processed in the first stage in the ResNet50 network to obtain the projection of the first stage and the network parameters of the first stage The projection of the first stage and the network parameters of the first stage are processed in the second stage in the ResNet50 network to obtain the projection of the second stage and the network parameters of the second stage The projection of the second stage and the network parameters of the second stage are processed in the third stage in the ResNet50 network to obtain the projection of the third stage and the network parameters of the third stage The projection of the third stage and the network parameters of the third stage are processed in the fourth stage in the ResNet50 network to obtain the projection of the fourth stage and the network parameters of the fourth stage The features output by connecting the four stages of the ResNet50 network in sequence are the global feature projections.
[0017] Input the UAV image while considering global feature extraction and feature extraction from different convolutional layers; upsample the feature map output of the first stage of ResNet50, and then connect the feature map output of the second stage of ResNet50 as the local low-level feature F low , the specific formula is as follows:
[0018]
[0019] Among them, F low is the local low-level feature, represents the concatenation operation, represents the projection of the first stage, represents the network parameters of the first stage, represents the projection of the second stage, represents the network parameters of the second stage;
[0020] After that, upsample the feature map output of the third stage of ResNet50, and then connect the feature map output of the fourth stage of ResNet50 as the local high-level feature F high , the specific formula is as follows:
[0021]
[0022] Among them, F high is the local high-level feature, represents the concatenation operation, represents the projection of the third stage, represents the network parameters of the third stage, represents the projection of the fourth stage, represents the network parameters of the fourth stage;
[0023] After that, use 3×3 convolution and 1×1 convolution to process and concatenate the local low-level feature and the local high-level feature respectively to make them have the same size; then use the residual structure to connect the average value of the local high-level feature and the concatenated feature to obtain the local joint feature F j , the specific formula is:
[0024]
[0025] Among them, F j is the local joint feature, ρ is the average value calculation, is the summation operation, ψ is the parametric rectified linear unit function, is the 3×3 convolution, is the 1×1 convolution;
[0026] In order to reduce the redundant information of the local joint feature, for the local joint feature F jPerform fine-grained transformation to obtain local fine-grained feature F l , and the specific formula is as follows:
[0027]
[0028] Among them, F l is the local fine-grained feature, ⊙ represents element-wise multiplication, and δ is the sigmoid function;
[0029] At this time, the information extraction is completed.
[0030] In step 3 mentioned above,
[0031] Step 3.1, Capture of information interaction attention. Project the global feature onto the Query of the attention mechanism through different fully connected layers to obtain Q ia , and project the local fine-grained feature F l onto Key and Value respectively to obtain and V ia , and the correlation S ia between the global feature and the local fine-grained feature is as follows:
[0032]
[0033] Among them, φ represents the softmax function, represents the set scaling parameter, Q ia is the Query in the attention mechanism, is the transposed Key in the attention mechanism;
[0034] To perform information interaction, calculate the similarity using multi-head attention and splice and fuse the similarities of different heads. The specific process is as follows:
[0035]
[0036]
[0037] Among them, L is the number of attention heads, represents the output of the l-th head, is the learnable parameter matrix, is the Dropout operation, represents the splicing operation, S l is the similarity of the l-th head, is the Value projected by the local fine-grained feature of the l-th head;
[0038] To enhance the visual representation and further obtain efficient feature embedding, combine the global feature and T ia , and the specific formula is as follows:
[0039]
[0040] F ia Namely, it is the feature embedding vector of the information interaction attention module. Represents layer normalization operation. Is a multi-layer perceptron; at this time, the feature embedding vector F captured by the information interaction attention is obtained. ia ;
[0041] Step 3.2, capture of visual enhancement attention: To enhance the visual performance, first project the feature embedding vector F of the information interaction attention module ia onto Query, Key, and Value of the attention mechanism to obtain Q va , and V va ; The similarity S of different tokens va is calculated as follows:
[0042]
[0043] where S va is the embedding matrix of different features, φ is the softmax function, is the set proportional parameter,
[0044] After that, the multi-head attention mechanism is used to calculate the similarity. The specific process is as follows:
[0045]
[0046]
[0047] where m is the head number of the enhanced visual attention module, is the output of the m-th head, W va is the learnable parameter of the enhanced visual attention module, represents the concatenation operation, S m is the similarity of the m-th head, is the feature embedding vector F of the m-th head ia projected Value;
[0048] Finally, it is processed through layer normalization to generate the saliency feature F va , and the specific formula is as follows:
[0049]
[0050] where F va is the saliency feature, is the layer normalization process.
[0051] The specific formula of the hash function in step 4 is as follows:
[0052] b = sign(h) = sign(τ(F va ,W h ))
[0053]
[0054] where F va is the output of the saliency capture module, W h is the weight of the approximation function, τ is the approximation function, h is the class hash code, and b is the generated hash code;
[0055] The objective function consists of a similarity maintenance term, a distribution smoothing term, and a quantization error;
[0056] The calculation formula of the similarity maintenance term is as follows:
[0057]
[0058] where ε is the marginal parameter, max is the maximum value function, H( ) calculates the Hamming distance, is the pairwise label of the samples (similar is 1, dissimilar is 0);
[0059] Introducing the distribution smoothing term can smooth the distribution center at the theoretical value, and the calculation formula is:
[0060]
[0061] where is the smoothing term, γ is the hyperparameter, θ is the label smoothing function, y n1 represents the label of the n1th input, b n is the generated hash code, and y n is the true label of the sample;
[0062] However, the above objective function is difficult to optimize during the training process. Therefore, the Euclidean distance D is used instead of the Hamming distance, that is:
[0063]
[0064] However, the above hash code will generate quantization errors. Therefore, a quantization error term is added, and the final objective function is:
[0065]
[0066] where represents the L2 norm result of the generated hash code and the true hash code, and λ is the hyperparameter.
[0067] In step 5, when training the overall network model, the Adam algorithm is used for optimization, and the learning rate is set to 10 -4 , the size of the input image is adjusted to 256×256; the batch size is set to 64, the length k of the hash code is set to 16, 24, 32, 48, 64, the margin parameter ε is set to 2k, and the initial weights of the convolutional neural network ResNet50 are initialized using the pre-trained weight parameter matrix W and bias parameter matrix B. Repeat steps 2 to 4 to iteratively train the network model, thereby optimizing the weight parameter matrix W and bias parameter matrix B to reduce the loss of the objective function L. When the training reaches 100 iterations or the loss of the final objective function no longer decreases, the algorithm runs ends, and then the trained overall network model is obtained to calculate the hash code of the samples in the test dataset.
[0068] In step 6, the query sample is the input of the drone image in the test dataset or prediction scenario.
[0069] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0070] 1. In the method for retrieving drone image hashes based on a saliency capture mechanism of the present invention, first, a new drone image retrieval framework is designed, which uses an information extraction module and a saliency capture module to solve the problem of efficient information of drone images in the process of hash code learning. Secondly, a new objective function composed of a similarity maintenance term, a distribution smoothing term, and a quantization error is designed, which not only maintains the similarity of the hash code, but also smooths the distribution of the drone image dataset, and reduces the quantization error between the hash code and the class hash code.
[0071] 2. The method for retrieving drone image hashes based on a saliency capture mechanism of the present invention mainly includes three implementation steps: extraction, learning, and selection. Given a drone image to be queried, first extract the representation features of the drone image; then use the similarity relationship of fixed same-class drone images for hash code learning; finally, use similarity calculation to obtain K similar images, effectively improving the retrieval accuracy. Through the comparative test results of the retrieval average precision index of two datasets, it can be seen that the retrieval effect of the designed drone image retrieval method is better than the existing method.
[0072] 3. The method for retrieving drone image hashes based on a saliency capture mechanism of the present invention learns the semantic information of drone image data, and uses the saliency capture mechanism, distribution smoothing term, global information, and local fine-grained information to learn effective hash codes to improve the retrieval accuracy. At the same time, this method uses the deep hashing method to reduce the spatio-temporal complexity of the retrieval, thereby reducing the storage space required by the retrieval method. Description of the Drawings
[0073] Figure 1 It is a schematic diagram of the network structure of the present invention.
[0074] Figure 2 It is a retrieval result diagram of the present invention.
[0075] Figure 3 It is a visualization effect diagram of the saliency capture module of the present invention. Detailed implementation manners
[0076] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0077] See Figure 1 , a method for retrieving drone image hashes based on a saliency capture mechanism, the method comprising the following steps:
[0078] Step 1, divide the photos in the drone image library into a training data set and a test data set;
[0079] Step 2, information extraction, which is improved based on the pre-trained ResNet50 network, and the ResNet50 network is trained for information extraction using the pictures in the training data set;
[0080] The pictures in the training data set are trained using the ResNet50 network to extract features. The ResNet50 network performs four-stage feature mapping processing on the pictures in the training data set. First, the feature map output generated in the first stage of ResNet50 is upsampled, and then connected to the feature mapping output of the second stage of ResNet50 to form the local low-level feature F low ; After that, the feature mapping output of the third stage of ResNet50 is upsampled, and then connected to the feature mapping output of the fourth stage of ResNet50 to form the local high-level feature F high ; Finally, the local low-level feature is processed using a 3×3 convolution, and the local high-level feature is processed using a 1×1 convolution so that these two features have the same size, and these two features are connected to form a connected feature; in order to avoid the loss of high-level feature semantics, the average value of the local high-level feature and the connected feature are connected using a residual structure to obtain the local joint feature F j ; In addition, performing fine-grained transformation on the local joint feature can reduce redundant information;
[0081] Step 3, saliency capture, after generating the local fine-grained feature F l , in order to enhance the effectiveness of the feature, saliency capture processing is performed; first, information interaction attention is captured, and then visual enhancement attention is captured;
[0082] The capture mechanism of information interaction attention enables the global features and local fine-grained features to learn from each other and interact, resulting in the feature embedding vector F captured by information interaction attention. ia The capture mechanism of visual enhancement attention is to enhance the visual representation of the effective features extracted, resulting in the saliency feature F output by the saliency module. va ;
[0083] Step 4: Conduct hash learning training. Input the saliency feature F output by the saliency module obtained in Step 3 va into the hash learning module for training, which is a fully connected hash layer with k nodes. The hash layer uses the tanh function as the activation function; during the training phase, k-bit hash class codes are generated and learned through an objective function composed of a similarity maintenance term, a distribution smoothing term, and a quantization error; during the testing phase, the k-bit class hash codes are quantized into k-bit hash codes using the sign function.
[0084] Step 5: Train the saliency capture model. Use the training dataset to loop through Steps 2 to 4 to train the network model. When 100 rounds of iteration are completed or the loss of the final objective function no longer decreases, end the algorithm operation, and thus obtain the trained overall network model to calculate the hash codes of the samples in the test dataset.
[0085] Step 6: Use the trained overall network model to calculate the hash codes of the samples in the test dataset. Sort the Hamming distances between the hash codes of the query sample and each sample in the training dataset from largest to smallest, calculate the top-n precision of the ranking list, and obtain the mean average precision (MAP) metric and the top-n retrieval results. At this time, the retrieval results are output, and the retrieval is completed.
[0086] In Step 2, the images in the training dataset are used to train and extract features using the ResNet50 network. The ResNet50 network performs four stages of feature mapping processing on the images in the training dataset. The images in the training dataset are processed in the first stage in the ResNet50 network to obtain the projection of the first stage and the network parameters of the first stage The projection of the first stage and the network parameters of the first stage are processed in the second stage in the ResNet50 network to obtain the projection of the second stage and the network parameters of the second stage The projection of the second stage and the network parameters of the second stage are processed in the third stage in the ResNet50 network to obtain the projection of the third stage and the network parameters of the third stage The projection of the third stage and the network parameters of the third stage Perform the fourth-stage processing in the ResNet50 network to obtain the projection of the fourth stage and the network parameters of the fourth stage The features output by connecting the four stages of the ResNet50 network in sequence are global feature projections;
[0087] Input the drone image while considering global feature extraction and feature extraction of different convolutional layers; Upsample the feature map output of the first stage of ResNet50, and then connect the feature map output of the second stage of ResNet50 as the local low-level feature F low , and the specific formula is as follows:
[0088]
[0089] Among them, F low is the local low-level feature, represents the concatenation operation, represents the projection of the first stage, represents the network parameters of the first stage, represents the projection of the second stage, represents the network parameters of the second stage;
[0090] After that, upsample the feature map output of the third stage of ResNet50, and then connect the feature map output of the fourth stage of ResNet50 as the local high-level feature F high , and the specific formula is as follows:
[0091]
[0092] Among them, F high is the local high-level feature, represents the concatenation operation, represents the projection of the third stage, represents the network parameters of the third stage, represents the projection of the fourth stage, represents the network parameters of the fourth stage;
[0093] After that, use 3×3 convolution and 1×1 convolution to process and concatenate the local low-level feature and the local high-level feature respectively to make them have the same size; Then use the residual structure to connect the average value of the local high-level feature and the concatenated feature to obtain the local joint feature F j , and the specific formula is:
[0094]
[0095] Among them, F jis the local joint feature, ρ is the average value calculation, is the summation operation, ψ is the parametric rectified linear unit function, is the 3×3 convolution, is the 1×1 convolution;
[0096] To reduce the redundant information of the local joint feature, the local joint feature F j is subjected to a fine-grained transformation to obtain the local fine-grained feature F l , and the specific formula is as follows:
[0097]
[0098] where F l is the local fine-grained feature, ⊙ represents element-wise multiplication, and δ is the sigmoid function;
[0099] At this time, the information extraction is completed.
[0100] In step 3 described above,
[0101] Step 3.1, capturing the information interaction attention, projecting the global feature onto the Query of the attention mechanism through different fully connected layers to obtain Q ia , projecting the local fine-grained feature F l onto Key and Value respectively to obtain and V ia , and the correlation S ia between the global feature and the local fine-grained feature is as follows:
[0102]
[0103] where φ represents the softmax function, represents the set scaling parameter, Q ia is the Query in the attention mechanism, is the transposed Key in the attention mechanism;
[0104] To perform information interaction, calculate the similarity using multi-head attention and splice and fuse the similarities of different heads. The specific process is as follows:
[0105]
[0106]
[0107] where L is the number of attention heads, represents the output of the l-th head, W ia is the learnable parameter matrix, is the Dropout operation, represents the splicing operation, Sl is the similarity of the l-th head, is the Value of the local fine-grained feature projection of the l-th head;
[0108] To enhance the visual representation and further obtain efficient feature embeddings, the global feature and T ia are combined, and the specific formula is as follows:
[0109]
[0110] F ia is the feature embedding vector of the information interaction attention module, represents the layer normalization operation, is the multi-layer perceptron; at this time, the feature embedding vector F captured by the information interaction attention is obtained ia ;
[0111] Step 3.2, Capture of visual enhancement attention: To enhance the visual performance, first project the feature embedding vector F of the information interaction attention module ia onto the Query, Key, and Value of the attention mechanism to obtain Q va , and V va ; The similarity S of different tokens va is calculated as follows:
[0112]
[0113] where S va is the embedding matrix of different features, φ is the softmax function, is the set ratio parameter,
[0114] After that, the multi-head attention mechanism is used to calculate the similarity, and the specific process is:
[0115]
[0116]
[0117] where m is the head number of the visual enhancement attention module, is the output of the m-th head, W va is the learnable parameter of the visual enhancement attention module, represents the concatenation operation, S m is the similarity of the m-th head, is the feature embedding vector F of the m-th head ia projected Value;
[0118] Finally, it is processed through layer normalization to generate the saliency feature Fva , the specific formula is as follows:
[0119]
[0120] Among them, F va is the significant feature, is layer normalization processing.
[0121] The specific formula of the hash function in step 4 is:
[0122] b = sign(h) = sign(τ(F va , W h ))
[0123]
[0124] Among them, F va is the output of the significant feature capture module, W h is the weight of the approximation function, τ is the approximation function, h is the class hash code, and b is the generated hash code;
[0125] The objective function consists of a similarity maintenance term, a distribution smoothing term, and a quantization error;
[0126] The calculation formula of the similarity maintenance term is as follows:
[0127]
[0128] Among them, ε is the edge parameter, max is the maximum value function, H() calculates the Hamming distance, is the paired label of the samples (similar is 1, dissimilar is 0);
[0129] Introducing the distribution smoothing term can smooth the distribution center at the theoretical value, and the calculation formula is:
[0130]
[0131] Among them, is the smoothing term, γ is the hyperparameter, θ is the label smoothing function, represents the label of the n1th input, b n is the generated hash code, y n is the true label of the sample;
[0132] However, the above objective function is difficult to optimize during training, so the Euclidean distance D is used instead of the Hamming distance, that is:
[0133]
[0134] However, the above hash code will generate a quantization error, so a quantization error term is added, and the final objective function is:
[0135]
[0136] Among them, represents the L2-norm result of the generated hash code and the true hash code, and λ is a hyperparameter.
[0137] In step 5, when training the overall network model, the Adam algorithm is used for optimization, and the learning rate is set to 10 -4 , the input image size is adjusted to 256×256; the batch size is set to 64, the length k of the hash code is set to 16, 24, 32, 48, 64, the margin parameter ε is set to 2k, and the initial weights of the convolutional neural network ResNet50 are initialized using the pre-trained weight parameter matrix W and bias parameter matrix B. Repeat steps 2 to 4 to iteratively train the network model, so as to optimize the weight parameter matrix W and bias parameter matrix B to reduce the loss of the objective function L. When the training reaches 100 iterations or the loss of the final objective function no longer decreases, the algorithm runs end, and then the trained overall network model is obtained to calculate the hash code of the samples in the test dataset.
[0138] In step 6, the query sample is the input of the drone image in the test dataset or the prediction scenario.
[0139] Example 1:
[0140] A drone image hashing retrieval method based on a saliency capture mechanism, the method comprising the following steps:
[0141] The environment adopted in this embodiment is GeForce RTX 3090 GPU, Inter Xeon(R) Silver 4210R CPU@2.40GHz×40, 62.6G RAM, Linux operating system, and is developed using Python and the open-source library Pytorch.
[0142] Step 1, divide the photos in the drone image library into a training dataset and a test dataset; use the Era and Drone-Action datasets, and select 80% of the dataset as the training dataset I train , and the remaining 20% as the test dataset I test ;
[0143] Step 2, information extraction, based on the improvement of the pre-trained ResNet50 network, use the pictures in the training dataset to perform information extraction training on the ResNet50 network;
[0144] The images in the training dataset are used to train and extract features using the ResNet50 network. The ResNet50 network performs four stages of feature mapping processing on the images in the training dataset. First, the feature map output generated in the first stage of ResNet50 is upsampled, and then connected to the feature mapping output of the second stage of ResNet50 to form the local low-level feature F low ; After that, the feature mapping output of the third stage of ResNet50 is upsampled, and then connected to the feature mapping output of the fourth stage of ResNet50 to form the local high-level feature F high ; Finally, the local low-level feature is processed using a 3×3 convolution, and the local high-level feature is processed using a 1×1 convolution so that these two features have the same size. These two features are connected to become the connected feature; To avoid the loss of high-level feature semantics, a residual structure is used to connect the average value of the local high-level feature and the connected feature to obtain the local joint feature F j ; In addition, performing fine-grained transformation on the local joint feature can reduce redundant information;
[0145] Step 3, saliency capture. After generating the local fine-grained feature F l , to enhance the effectiveness of the feature, saliency capture processing is performed; First, information interaction attention is captured, and then visual enhancement attention is captured;
[0146] The capture mechanism of information interaction attention is to enable the global feature and the local fine-grained feature to interact and learn from each other to obtain the feature embedding vector F captured by information interaction attention ia ; The capture mechanism of visual enhancement attention is to enhance the visual representation of the extracted effective features to obtain the saliency feature F output by the saliency module va ;
[0147] Step 4, perform hashing learning training. After the saliency feature F output by the saliency module obtained in Step 3 va , is input into the hashing learning module for training, that is, a fully connected hashing layer with k nodes. The hashing layer uses the tanh function as the activation function; In the training stage, k-bit hashing class codes are generated and learned through an objective function composed of a similarity maintenance term, a distribution smoothing term, and a quantization error; In the test stage, the k-bit class hashing code is quantized into a k-bit hashing code using the sign function;
[0148] Step 5, train the saliency capture model. Use the training dataset to loop through Steps 2 to 4 to train the network model. When 100 rounds of iteration are trained or the loss of the final objective function no longer decreases, the algorithm operation ends, and then the trained overall network model is obtained to calculate the hashing codes of the samples in the test dataset;
[0149] Step 6: Use the trained overall network model to calculate the hash codes of the samples in the test dataset. Sort the Hamming distances between the query sample and the hash codes of each sample in the training dataset from largest to smallest, calculate the top-n precision of the ranking list, and obtain the mean average precision (MAP) and the top-n retrieval results. At this time, the retrieval results are output and the retrieval is completed.
[0150] Example 2:
[0151] Example 2 is basically the same as Example 1, except that:
[0152] In step 2, the pictures in the training dataset are used to train and extract features using the ResNet50 network. The ResNet50 network performs four stages of feature mapping processing on the pictures in the training dataset, and the pictures in the training dataset are processed in the ResNet50 network in the first stage to obtain the projection of the first stage and the network parameters of the first stage The projection of the first stage and the network parameters of the first stage are processed in the ResNet50 network in the second stage to obtain the projection of the second stage and the network parameters of the second stage The projection of the second stage and the network parameters of the second stage are processed in the ResNet50 network in the third stage to obtain the projection of the third stage and the network parameters of the third stage The projection of the third stage and the network parameters of the third stage are processed in the ResNet50 network in the fourth stage to obtain the projection of the fourth stage and the network parameters of the fourth stage The features output by connecting the four stages of the ResNet50 network in sequence are the global feature projection;
[0153] When inputting the UAV image, both global feature extraction and feature extraction of different convolutional layers are considered; the feature map output of the first stage of ResNet50 is upsampled, and then the feature map output of the second stage of ResNet50 is connected to form the local low-level feature F low , and the specific formula is as follows:
[0154]
[0155] where F low is the local low-level feature, represents the concatenation operation, represents the projection of the first stage, Represents the network parameters of the first stage, Represents the projection of the second stage, Represents the network parameters of the second stage;
[0156] After that, the feature map output of the third stage of ResNet50 is upsampled, and then the feature map output of the fourth stage of ResNet50 is concatenated into the local high-level feature F high , and the specific formula is as follows:
[0157]
[0158] Among them, F high is the local high-level feature, represents the concatenation operation, represents the projection of the third stage, represents the network parameters of the third stage, represents the projection of the fourth stage, represents the network parameters of the fourth stage;
[0159] After that, the local low-level feature and the local high-level feature are processed and concatenated using a 3×3 convolution and a 1×1 convolution respectively to make them have the same size; then the average value of the local high-level feature and the concatenated feature are connected using a residual structure to obtain the local joint feature F j , and the specific formula is:
[0160]
[0161] Among them, F j is the local joint feature, ρ is the average value calculation, is the summation operation, ψ is the parametric rectified linear unit function, is the 3×3 convolution, is the 1×1 convolution;
[0162] In order to reduce the redundant information of the local joint feature, the local joint feature F j is subjected to a fine-grained transformation to obtain the local fine-grained feature F l , and the specific formula is as follows:
[0163]
[0164] Among them, F l is the local fine-grained feature, ⊙ represents element-wise multiplication, and δ is the sigmoid function;
[0165] At this time, the information extraction is completed.
[0166] In the said step 3,
[0167] Step 3.1, Capturing information interaction attention: Project the global features onto the Query of the attention mechanism through different fully connected layers to obtain Q ia , the local fine-grained feature F l Projected onto Key and Value respectively to obtain and V ia , the correlation S between the global feature and the local fine-grained feature ia is as follows:
[0168]
[0169] where φ represents the softmax function, represents the set scaling parameter, Q ia is the Query in the attention mechanism, is the transposed Key in the attention mechanism;
[0170] To perform information interaction, calculate the similarity using multi-head attention and splice and fuse the similarities of different heads. The specific process is as follows:
[0171]
[0172]
[0173] where L is the number of attention heads, represents the output of the l-th head, W ia is a learnable parameter matrix, is the Dropout operation, represents the splicing operation, S l is the similarity of the l-th head, is the Value projected from the local fine-grained feature of the l-th head;
[0174] To enhance the visual representation and further obtain efficient feature embeddings, combine the global feature and T ia The specific formula is as follows:
[0175]
[0176] F ia is the feature embedding vector of the information interaction attention module, represents the layer normalization operation, is the multi-layer perceptron; At this time, the feature embedding vector F ia captured by the information interaction attention is obtained;
[0177] Step 3.2, Capturing visual enhancement attention: To enhance the visual performance, first, the feature embedding vector F of the information interaction attention module iaThe Query, Key, and Value projected onto the attention mechanism respectively obtain Q va , and V va ; The similarity S of different tokens va is calculated as follows:
[0178]
[0179] where S va is the embedding matrix of different features, φ is the softmax function, is the set ratio parameter,
[0180] After that, the multi-head attention mechanism is used to calculate the similarity. The specific process is as follows:
[0181]
[0182]
[0183] where m is the head number of the enhanced visual attention module, is the output of the m-th head, W va is the learnable parameter of the enhanced visual attention module, represents the concatenation operation, S m is the similarity of the m-th head, is the feature embedding vector F of the m-th head ia projected Value;
[0184] Finally, it is processed through layer normalization to generate the saliency feature F va , and the specific formula is as follows:
[0185]
[0186] where F va is the saliency feature, is the layer normalization process.
[0187] In step 4, the specific formula of the hash function is:
[0188] b = sign(h) = sign(τ(F va , W h ))
[0189]
[0190] where F va is the output of the saliency capture module, W h is the weight of the approximation function, τ is the approximation function, h is the class hash code, and b is the generated hash code;
[0191] The objective function consists of a similarity maintenance term, a distribution smoothing term, and a quantization error.
[0192] The calculation formula for the similarity maintenance term is as follows:
[0193]
[0194] Among them, ε is the edge parameter, max is the maximum value function, H() calculates the Hamming distance, is the paired label of the samples (similar is 1, dissimilar is 0);
[0195] Introducing the distribution smoothing term can smooth the distribution center at the theoretical value, and the calculation formula is:
[0196]
[0197] Among them, is the smoothing term, γ is the hyperparameter, θ is the label smoothing function, represents the label of the n1th input, b n is the generated hash code, y n is the true label of the sample;
[0198] However, the above objective function is difficult to optimize during the training process. Therefore, the Euclidean distance D is used instead of the Hamming distance, that is:
[0199]
[0200] However, the above hash code will generate a quantization error. Therefore, a quantization error term is added, and the final objective function is:
[0201]
[0202] Among them, represents the L2 norm result of the generated hash code and the true hash code, and λ is the hyperparameter.
[0203] In step 5, when training the overall network model, the Adam algorithm is used for optimization, and the learning rate is set to 10 -4, the input image size is adjusted to 256×256; the batch size is set to 64, the length k of the hash code is set to 16, 24, 32, 48, 64, the margin parameter ε is set to 2k, and the initial weights of the convolutional neural network ResNet50 are initialized using the pre-trained weight parameter matrix W and bias parameter matrix B. Repeat steps 2 to 4 to iteratively train the network model, thereby optimizing the weight parameter matrix W and bias parameter matrix B to reduce the loss of the objective function L. When the training reaches 100 iterations or the loss of the final objective function no longer decreases, the algorithm runs end, and then the trained overall network model is obtained to calculate the hash codes of the samples in the test dataset.
[0204] In step 6, the query sample is the input of drone images in the test dataset or prediction scenario.
[0205] To evaluate the effectiveness of the method of the present invention, the method of the present invention is compared with several state-of-the-art methods in terms of retrieval performance, including DHN, DCH, DFH, DPH, DSHSD, GreedyHash, DSDH, DTSH, LCDSH, QSMIH. In this experiment, 16, 24, 32, 48, 64-bit hash codes are used, and the Drone-Action dataset and ERA dataset are used. DHN uses the Bayesian framework to perform deep hash learning in a supervised manner. The methods of DCH, DFH, DPH, DSHSD, GreedyHash, DSDH, DTSH, LCDSH, QSMIH are executed according to the original text.
[0206] Table 1
[0207]
[0208]
[0209] Table 1 shows the comparison experiment results of the present invention and other methods in the drone image retrieval task on the ERA dataset, where mAP is the average precision metric.
[0210] Table 2
[0211]
[0212] Table 2 shows the comparison experiment results of the present invention and other methods in the drone image retrieval task on the Drone-Action dataset, where mAP is the average precision metric.
[0213] From the comparison results of the average precision metrics of the two datasets above, it can be seen that the retrieval effect of the drone image retrieval method of this design is better than the existing methods.
Claims
1. A method for retrieving drone images by hash based on a saliency capture mechanism, characterized in that: The method includes the following steps: Step 1, divide the photos in the drone image library into a training data set and a test data set; Step 2, information extraction, improve based on the pre-trained ResNet50 network, and use the pictures in the training data set to train the ResNet50 network for information extraction; The images in the training dataset are used to train and extract features using the ResNet50 network. The ResNet50 network performs four stages of feature mapping processing on the images in the training dataset. First, the feature map output generated in the first stage of ResNet50 is upsampled, and then connected to the feature mapping output in the second stage of ResNet50 to form the local low-level feature F low ; After that, the feature mapping output in the third stage of ResNet50 is upsampled, and then connected to the feature mapping output in the fourth stage of ResNet50 to form the local high-level feature F high ; Finally, the local low-level feature is processed using a 3×3 convolution, and the local high-level feature is processed using a 1×1 convolution so that these two features have the same size. These two features are connected to form the connected feature; in order to avoid the loss of high-level feature semantics, the average value of the local high-level feature and the connected feature are connected using a residual structure to obtain the local joint feature F j ; In addition, performing fine-grained transformation on the local joint feature can reduce redundant information; Step 3, saliency capture. After generating the local fine-grained feature F l , in order to enhance the effectiveness of the feature, saliency capture processing is performed; first, information interaction attention is captured, and then visual enhancement attention is captured; The capture mechanism of information interaction attention enables the global features and local fine-grained features to learn from each other and interact, resulting in the feature embedding vector F captured by information interaction attention ia ; The capture mechanism of visual enhancement attention is to enhance the visual representation of the effective features extracted, resulting in the salient feature F output by the salient module va ; Step 4: Conduct hash learning training. Input the saliency features F output by the saliency module obtained in Step 3 va into the hash learning module for training, that is, a fully-connected hash layer with k nodes. The hash layer uses the tanh function as the activation function; during the training phase, k-bit hash class codes are generated and learned through an objective function composed of a similarity maintenance term, a distribution smoothing term, and a quantization error; In the test stage, use the sign function to quantize the k-bit class hash code into a k-bit hash code; Step 5, train the saliency capture model, use the training data set to loop through steps 2 to 4 to train the network model, and end the algorithm operation when 100 rounds of iteration are trained or the loss of the final objective function no longer decreases, and then obtain the trained overall network model to calculate the hash code of the samples in the test data set; Step 6, use the trained overall network model to calculate the hash codes of the samples in the test data set, sort the Hamming distances between the query sample and the hash codes of each sample in the training data set from largest to smallest, calculate the top n precisions of the ranking list, obtain the mean average precision (MAP) and the top n retrieval results, and at this time the retrieval results are output and the retrieval is completed.
2. A method for retrieving drone images by hash based on a saliency capture mechanism according to claim 1, characterized in that: In step 2, the pictures in the training data set are used to train and extract features by the ResNet50 network. The ResNet50 network performs four-stage feature mapping processing on the pictures in the training data set, and the pictures in the training data set are processed in the first stage in the ResNet50 network to obtain the projection of the first stage and the network parameters of the first stage The projection of the first stage and the network parameters of the first stage are processed in the second stage in the ResNet50 network to obtain the projection of the second stage and the network parameters of the second stage The projection of the second stage and the network parameters of the second stage are processed in the third stage in the ResNet50 network to obtain the projection of the third stage and the network parameters of the third stage The projection of the third stage and the network parameters of the third stage are processed in the fourth stage in the ResNet50 network to obtain the projection of the fourth stage and the network parameters of the fourth stage The features output by connecting the four stages of the ResNet50 network in sequence are the global feature projection; Input the UAV image while considering both global feature extraction and feature extraction from different convolutional layers; upsample the feature map output of the first stage of ResNet50, and then connect the feature map output of the second stage of ResNet50 as the local low-level feature F low , and the specific formula is as follows: Among them, F low is the local low-level feature, represents the splicing operation, represents the projection in the first stage, represents the network parameters in the first stage, represents the projection in the second stage, represents the network parameters in the second stage; After that, the feature map output of the third stage of ResNet50 is upsampled, and then the feature map output of the fourth stage of ResNet50 is concatenated as the local high-level feature F high , and the specific formula is as follows: Among them, F high is the local high-level feature, represents the splicing operation, represents the projection in the third stage, represents the network parameters in the third stage, represents the projection in the fourth stage, represents the network parameters in the fourth stage; Subsequently, the local low-level features and local high-level features are processed and concatenated using 3×3 convolution and 1×1 convolution respectively to make them have the same size; then the average value of the local high-level features and the concatenated features are connected using a residual structure to obtain the local joint feature F j , and the specific formula is: Among them, F j is the local joint feature, ρ is the average value calculation, is the summation operation, ψ is the parametric rectified linear unit function, is the 3×3 convolution, is the 1×1 convolution; To reduce the redundant information of the local joint feature, the local joint feature F j is subjected to a fine-grained transformation to obtain the local fine-grained feature F l , and the specific formula is as follows: Among them, F l is the local fine-grained feature, ⊙ represents element-wise multiplication, and δ is the sigmoid function; At this time, the information extraction is completed.
3. A method for retrieving drone images by hash based on a saliency capture mechanism according to claim 2, characterized in that: In the said step 3, Step 3.1, capturing information interaction attention, project the global feature onto the Query of the attention mechanism through different fully connected layers to obtain Q ia , local fine-grained feature F l Project onto Key and Value respectively to obtain and V ia , the correlation S between the global feature and the local fine-grained feature ia is as follows: where φ represents the softmax function, represents a set scaling parameter, Q ia is the Query in the attention mechanism, is the transposed Key in the attention mechanism; In order to perform information interaction, calculate the similarity using multi-head attention, and splice and fuse the similarities of different heads. The specific process is as follows: Among them, L is the number of attention heads, represents the output of the l-th head, W ia is a learnable parameter matrix, is the Dropout operation, represents the concatenation operation, S l is the similarity of the l-th head, is the Value of the local fine-grained feature projection of the l-th head; To enhance the visual representation and further obtain efficient feature embeddings, the global features and T ia are combined, and the specific formula is as follows: F ia Namely, it is the feature embedding vector of the information interaction attention module, indicating a layer normalization operation, is a multi-layer perceptron; at this time, the feature embedding vector F captured by the information interaction attention is obtained ia ; Step 3.2, Capture of visually enhanced attention: To enhance visual performance, first embed the feature vector F of the information interaction attention module ia into Query, Key, and Value of the attention mechanism respectively to obtain Q va , and V va ; The similarity S of different tokens va is calculated as follows: Among them, S va is the embedding matrix of different features, φ is the softmax function, is the set proportionality parameter, After that, calculate the similarity using the multi-head attention mechanism. The specific process is: Among them, m is the head number of the enhanced visual attention module, is the output of the m-th head, W va are the learnable parameters of the enhanced visual attention module, represents the concatenation operation, S m is the similarity of the m-th head, is the feature embedding vector F of the m-th head ia projected Value; Finally, it is processed by layer normalization to generate the saliency feature F va , and the specific formula is as follows: Among them, F va is the significant feature, and is layer normalization processing.
4. A method for retrieving drone images by hash based on a saliency capture mechanism according to claim 3, characterized in that: In the said step 4, the specific formula of the hash function is: b = sign(h) = sign(τ(F va ,W h )) Among them, F va is the output of the saliency capture module, W h is the weight of the approximation function, τ is the approximation function, h is the class hash code, and b is the generated hash code; The objective function consists of a similarity maintenance term, a distribution smoothing term and a quantization error; The calculation formula of the similarity maintenance term is as follows: where ε is the edge parameter, max is the maximum value function, and H( ) calculates the Hamming distance, which is the paired label of the samples (1 for similar and 0 for dissimilar); Introducing the distribution smoothing term can smooth the distribution center at the theoretical value. The calculation formula is: Among them, is the smoothing term, γ is the hyperparameter, is the label smoothing function, represents the label of the n1-th input, b n is the generated hash code, y n is the true label of the sample; However, the above objective function is difficult to optimize during the training process. Therefore, the Euclidean distance D is used instead of the Hamming distance, that is: However, the above hash code will generate a quantization error. Therefore, a quantization error term is added. The final objective function is: Among them, represents the L2-norm result of the generated hash code and the true hash code, and λ is a hyperparameter.
5. A method for retrieving drone images by hash based on a saliency capture mechanism according to claim 4, characterized in that: In step 5, when training the overall network model, the Adam algorithm is used for optimization, and the learning rate is set to 10 -4 , the input image size is adjusted to 256×256; the batch size is set to 64, the length k of the hash code is set to 16, 24, 32, 48, 64, the margin parameter ε is set to 2k, and the initial weights of the convolutional neural network ResNet50 are initialized using the pre-trained weight parameter matrix W and bias parameter matrix B. Repeat steps 2 to 4 to iteratively train the network model, so as to optimize the weight parameter matrix W and bias parameter matrix B to reduce the loss of the objective function L. When the training reaches 100 rounds of iteration or the loss of the final objective function no longer decreases, end the algorithm operation, and then obtain the trained overall network model to calculate the hash code of the samples in the test dataset.
6. A method for retrieving drone images by hash based on a saliency capture mechanism according to claim 5, characterized in that: In the said step 6, the query sample is the input of a drone picture in the test data set or a prediction scenario.
Citation Information
Patent Citations
Depth significance-based remote sensing image rapid retrieval method
CN106909924A
Image processing method and device, storage medium and equipment
CN114863138A