Airborne Image Vehicle Target Recognition Method Based on Feature Fusion and Attention Mechanism
By adopting feature fusion and attention mechanism methods in the recognition of vehicle targets in the image of UAV, the problems of large differences in vehicle target scales and fewer feature information are solved, and the accuracy of detection and network representation ability are improved.
Patent Information
- Application Number
- CN202111654301.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-30
AI Technical Summary
In the drone image, the vehicle target scales vary greatly, and the small targets are mostly small, and the feature information is less, which makes it easy to cause missed detection and missed detection problems.
The airborne image vehicle target recognition method based on feature fusion and attention mechanism is adopted to feature fusion through a double pyramid structure from top to bottom to top and bottom to top, and the channel attention and spatial attention modules are used to highlight the features of the target area.
It improves the network's representation ability, obtains richer semantic information, reduces the missed detection and misdetection problems of target detection, and improves the accuracy of detection without increasing excessive computing costs.
Smart Images

Figure CN114332620B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing in deep learning. Specifically, it relates to an airborne vehicle target detection method based on a feature pyramid and an attention mechanism. Background Art
[0002] In recent years, the continuous increase in the number of vehicles has led to problems such as road congestion, tight parking spaces, and an increase in traffic accidents. As a powerful supplement to remote sensing images and road monitoring, drone images can be widely applied to fields such as traffic emergency command, parking space monitoring, and traffic accident investigation and evidence collection. As an effective tool for acquiring images, drones provide people with a flexible perspective, a broad view, and more valuable information for obtaining and analyzing vehicle picture information. Therefore, the vehicle target detection method under drone images has gradually become a research hotspot in the field of target detection, and the vehicle target detection algorithm is even the top priority among them. However, the large variation in the flight altitude of drones results in the same target having different scales in the captured images, which poses a major challenge for vehicle target detection.
[0003] Traditional drone image vehicle target detection methods mainly consist of region selection, feature extraction, and a classifier. Previous research mainly focused on manually extracting features, such as Scale-Invariant Feature Transform (SIFT), Histogram of Oriented Gradients (HOG), Haar-like features, etc. Although traditional methods have made great contributions to vehicle detection in drone images, the method of manually extracting features is slow, requires a large amount of work, and has limitations in performance. Especially for drone images with complex backgrounds, the shapes of vehicles are similar to those of other objects (such as building roofs, road signs, trash cans, electrical units, etc.). These factors make it difficult for manually crafted descriptors to completely distinguish vehicles from the background.
[0004] In recent years, deep convolutional neural networks (CNNs) have achieved remarkable results in fields such as image classification, target detection, and semantic segmentation. In the field of vehicle detection in drone images, more and more methods based on convolutional neural networks have been proposed. The main methods are based on target detection and can be roughly divided into two categories. The first category is two-stage object detectors based on candidate regions (such as Faster R-CNN, Cascade R-CNN, etc.), and the second category is single-stage object detectors (such as SSD, YOLO, etc.). Currently, the vast majority of drone image target detection algorithms are proposed based on these two types of detectors and have achieved good detection results. However, in drone images, the scale differences of vehicle targets are large, there are many small targets, and the feature information is less, which is prone to problems of missed detection and false detection. Summary of the Invention
[0005] Aiming at the deficiencies in the prior art, the present invention provides an airborne image vehicle target recognition method based on feature fusion and attention mechanism. The method first performs bottom-up feature extraction on the input image to obtain the first-level features; then performs top-down upsampling on the extracted features, and fuses them with the features of the same size generated at the first level through lateral connection to obtain the second-level features; then performs top-down downsampling on the second-level features, and fuses them with the features of the same size generated at the first and second levels through lateral connection to obtain the third-level features; then passes the fused features through an attention module composed of a channel attention module and a spatial attention module in parallel; finally, sends the feature map to the detection head for classification and regression.
[0006] An airborne image vehicle target recognition method based on feature fusion and attention mechanism, comprising the following steps:
[0007] S1: Construct an airborne image vehicle target recognition model;
[0008] The airborne image vehicle target recognition model includes a feature extraction module, a feature fusion module, an attention mechanism module, and a detection module. The feature extraction module is used to extract features from the input original picture; the feature fusion module fuses the extracted features to obtain richer semantic information; the attention mechanism highlights the key information in the fused features; the detection module is used to obtain the category and location of the target.
[0009] S2: Through the feature extraction module, use a convolutional neural network to extract features from the original picture and output a feature map with multiple scales.
[0010] S3: Upsample the top-level feature map through the feature fusion module and perform lateral connection with the low-level features to construct a top-down pyramid structure, and output the preliminarily fused multi-scale features. Then process the preliminarily fused multi-scale features, downsample the bottom-level features, and perform lateral connection with the high-level feature map and the high-level feature map output in S2 respectively to form a bottom-up pyramid structure, and output the finally fused multi-scale features.
[0011] S4: Through the attention module, operate on the finally fused multi-scale features along two dimensions of space and channel. In the two sub-modules of space and channel, calculate the scaling factor and weight of the feature map respectively, and multiply them with the original feature map to adaptively adjust the features, so that the network learns to focus on the key information of the feature map. Finally, add the outputs of the two sub-modules of space and channel to the input feature map to obtain the final result.
[0012] S5: Generate candidate boxes of multiple different ratios and sizes, obtain a list of candidate boxes at each output feature map position, calculate the candidate boxes corresponding to the original image on each feature map, and calculate the positive and negative sample attributes of each candidate box using the ground truth information.
[0013] S6: The detection module selects a one-stage detector as needed, which includes two sub-networks: classification and regression. The feature map is sent into the classification and regression sub-networks respectively, which are used to judge the category of the target and the specific position of the target, obtain the confidence of each prediction box and its category, and use the regression network to correct the position. Finally, non-maximum suppression is used to remove redundant prediction boxes, and only the one with the best result is retained to obtain the final detection result.
[0014] Preferably, the detailed steps of S2 are as follows:
[0015] Use Resnet50 as the backbone of the feature extraction module to extract the features of the original image. Resnet50 has five stages, and each stage outputs a feature map, which is also the input of the next stage. The scales of the feature maps output by different stages are different. The higher the layer, the smaller the scale and the more channels.
[0016] First, process the input original image. The specific process is as follows:
[0017] Out = Conv 7×7 (C, W, H, k, s) (1)
[0018] Where Out represents the feature map output by the first stage, and Conv 7×7 represents a convolutional layer with a size of 7×7. C represents the number of channels of the input image. The number of channels of an RGB image is 3. W and H represent the width and height of the input image respectively. k represents the size of the convolutional kernel, and s represents the stride of the convolutional kernel movement. Where k = 64 and s = 2.
[0019] Secondly, perform feature extraction on the output feature map in sequence. The specific process is as follows:
[0020] P i = Conv 3×3 (C i , W, H, k i , s i ) (2)
[0021] Where P i is the feature map output by the i-th stage, and Conv 3×3 represents a convolutional layer with a size of 3×3. C i represents the number of channels of the feature map input in the i-th stage. W and H represent the width and height of the input feature map respectively. k i represents the size of the convolutional kernel, where ki =256×2 i-1 .s i Represents the step size of the convolution kernel movement, s i =2. Finally, four feature maps are output, from bottom to top, respectively: {P 2 ,P 3 ,P 4 ,P 5}, and the number of channels are {256,512,1024,2048} respectively.
[0022] Preferably, the detailed steps of S3 are:
[0023] S3.1: The feature map {P 2 ,P 3 ,P 4 ,P 5} for top-down enhancement. First, the channels of the feature map obtained in S2 are normalized, and the top feature map P 5 Translate to get N 5 The translated feature map is shuffled and the original feature map is enlarged by two times. Then the channels of the enlarged feature map are normalized so that it can be horizontally connected with the low-level feature map in S2 to obtain the new feature map N 4 . N 4 Repeat the enlargement and lateral connection operations until all the low-level feature maps in S2 are laterally connected.
[0024]
[0025] Where N i is the i-th layer feature map after feature fusion, Conv 1×1 is a convolution of size 1×1, PS(·) represents the pixel shuffle function to upsample the features, L(·) represents the channel normalization operation, and N i+1 is the feature map of the previous layer, k n is the number of convolution kernels, s n is the step size of the convolution kernel movement, where k n =256,s n =1. i Compared with N i+1 The final output is the initial fused feature map {N 2 ,N 3 ,N 4 ,N 5}
[0026] S3.2: The feature map {N 2 ,N 3 ,N 4 ,N5} perform bottom-up enhancement. First, translate the feature map of the bottom layer. After that, perform convolution operations on each layer of the feature map to reduce the feature map by a factor of two, and then horizontally connect it with the corresponding N i feature map and P i feature map, and perform convolution operations on the connected feature map to finally generate a new high-resolution feature map. Specifically, it is as follows:
[0027]
[0028] where F i is the i-th layer of the output high-resolution feature map, Conv 3×3 is a convolution with a size of 3×3, k is the number of convolution kernels, where k f = 2, 2 indicates that the convolution kernel moves with a stride of 2, N i is the i-th layer of the feature map in S3.1, P i is the feature map output in the i-th stage of S2, and 1 indicates the stride of the convolution kernel. The finally output feature map is {F 2 , F 3 , F 4 , F 5}.
[0029] Preferably, the detailed steps of S4 are as follows:
[0030] S4.1: Pass the finally fused feature map through the attention module to suppress insignificant features in the channels and space and enhance the ability to obtain significant features. The attention module is composed of two sub-modules, spatial attention and channel attention, in parallel. The input feature map passes through the spatial attention and channel attention modules in parallel, and finally, the output features of the two attention modules and the features of the input feature map are aggregated to obtain a better pixel-level prediction feature representation. Specifically, it is as follows:
[0031] Out a = SUM(SA(F i ), CA(F i ), F i ) (5)
[0032] where Out a refers to the finally output feature map, SUM(·) refers to the element summation function for completing feature fusion, SA(·) is the spatial attention module, CA(·) is the channel attention module, and F i is the input feature map.
[0033] S4.2: Input the feature map into the spatial attention module to improve its representation ability. First, batch-normalize each pixel in the feature map to obtain its scaling factor, calculate the weights through the scaling factor, pass through the activation function, and then add it to the input feature map to obtain the final output. Specifically, it is expressed as:
[0034] Out s = sigmoid(W α (BN s (F i ))) (6)
[0035] where Out s represents the output of the spatial attention module, sigmoid(·) represents the sigmoid activation function, W α represents the weight information of the channel attention module, BN s represents the scaling factor of the channel attention module, and F i represents the input feature map.
[0036] S4.3: Input the feature map into the channel attention module to summarize the spatial features. Use the scaling factor in batch normalization to reflect the variability and importance of each channel, and obtain the weights through the scaling factor. First, obtain the scaling factor of the feature map, then obtain the weight information through the scaling factor, and after passing through the activation function, add it to the information of the input feature map to obtain the final result. Specifically, it is expressed as:
[0037] Out c = sigmoid(W β (BN c (F i ))) (7)
[0038] where Out c represents the output of the channel attention module, sigmoid(·) represents the sigmoid activation function, W β represents the weight information of the channel attention module, BN c represents the scaling factor of the channel attention module, and F i represents the input feature map.
[0039] Preferably, the detailed steps of S5 are as follows:
[0040] On the feature map output by the attention module, multiple candidate boxes with different scales and sizes are generated on each feature map through the candidate region generation method. The number of candidate boxes is adjusted according to specific circumstances. And the size of the candidate boxes is changed according to the overall size of the targets in the dataset. Traverse each coordinate point on the feature map, map it to the coordinates of the original input image, and calculate the list of candidate boxes on the original image according to the size and position of the candidate boxes on the feature map. Calculate the intersection over union (IoU) between the ground truth and the candidate boxes to obtain the positive and negative sample attributes of each candidate box. If the IoU is greater than or equal to 0.5, the candidate box is a high-quality positive sample, and the possibility of having a target inside the candidate box is high. If the IoU is lower than 0.4, the candidate box is a negative sample, and the possibility of having a target inside the box is low. If the IoU is in the interval [0.4, 0.5), the candidate box is an ignored sample and is not calculated.
[0041] Preferably, the detailed steps of S6 are as follows:
[0042] Send the obtained positive samples into the detection module and input them into the classification and regression sub-networks respectively. The detection module selects a one-stage detector as needed. The classification network calculates which class each positive sample belongs to through a fully connected layer and a normalization function and outputs a probability vector of the class. The values in the probability vector correspond to different classes respectively, and the class corresponding to the largest value is the predicted class of the positive sample. The regression network obtains the position offset of each positive sample through bounding box regression to obtain a more accurate target detection box. Finally, non-maximum suppression is used to remove redundant prediction boxes and retain the best result. The final detection result is the position of the target detection box and the predicted class of the target detection box.
[0043] The beneficial effects of the present invention are as follows:
[0044] 1. The present invention uses a dual pyramid structure of top-down and bottom-up, and adds multiple lateral connections to fuse the output first-level features with the second- and third-level features, enhancing the network's representation ability and obtaining a feature map with richer semantic information. For problems such as small vehicles and unclear vehicle features in airborne images, it can be effectively solved without sacrificing a large computational cost. It can improve the problems of missed detection and false detection in target detection.
[0045] 2. The present invention addresses the problem of complex backgrounds in airborne images by using an attention module to highlight the features of the target area and enabling the network to focus on learning the areas in need. In the attention module, a parallel method of spatial attention and channel attention modules is used to obtain both spatial key information and channel key information, and add them to the original feature map to obtain a more accurate feature map. It can make subsequent detections more accurate without adding too much time. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of the method of the embodiment of the present invention;
[0047] Figure 2 Schematic diagram of the attention module of the embodiment of the present invention. Detailed implementation manners
[0048] The method of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0049] As Figure 1 shown, an airborne image vehicle target recognition method based on feature fusion and attention mechanism includes the following steps:
[0050] S1: Construct an airborne image vehicle target recognition model;
[0051] The airborne image vehicle target recognition model includes a feature extraction module, a feature fusion module, an attention mechanism module, and a detection module. The feature extraction module is used to extract features from the input original picture; the feature fusion module fuses the extracted features to obtain richer semantic information; the attention mechanism highlights the key information in the fused features; the detection module is used to obtain the category and location of the target.
[0052] S2: Through the feature extraction module, use a convolutional neural network to extract features from the original picture and output a feature map with multiple scales.
[0053] As Figure 1 shown at the far left, use Resnet50 as the backbone of the feature extraction module to extract the features of the original picture. Resnet50 has five stages, and each stage outputs a feature map, which is also the input of the next stage. The scales of the feature maps output by different stages are different. The higher the layer, the smaller the scale and the more channels of the feature map.
[0054] First, process the input original picture, and the specific process is as follows:
[0055] Out = Conv 7×7 (C, W, s, k, s) (1)
[0056] Among them, Out represents the feature map output by the first stage, and Conv 7×7 represents a convolutional layer with a size of 7×7. C represents the number of channels of the input picture. The number of channels of an RGB image is 3. W and H respectively represent the width and height of the input picture. k represents the size of the convolutional kernel, and s represents the step size of the convolutional kernel movement. Among them, k = 64 and s = 2.
[0057] Secondly, perform feature extraction on the output feature map in sequence, and the specific process is as follows:
[0058] Pi = Conv 3×3 (C i , W, H, k i , s i ) (2)
[0059] Among them, P i is the feature map output in the i-th stage, and Conv 3×3 represents a convolutional layer with a size of 3×3. C i represents the number of channels of the feature map input in the i-th stage, and W and H respectively represent the width and height of the input feature map. k i represents the size of the convolutional kernel, where k i = 256×2 i-1 . s i represents the stride of the convolutional kernel movement, and s i = 2. Finally, four feature maps are output, from bottom to top are {P 2 , P 3 , P 4 , P 5}, and the number of channels are {256, 512, 1024, 2048} respectively.
[0060] S3: Upsample the feature map at the top layer through the feature fusion module and perform horizontal connection with the low-level features to construct a top-down pyramid structure, and output the preliminarily fused multi-scale features. Then process the preliminarily fused multi-scale features, downsample the low-level features, and perform horizontal connection with the high-level feature map and the high-level feature map output in S2 respectively to form a bottom-up pyramid structure, and output the finally fused multi-scale features. The specific method is as follows
[0061] S3.1: As shown in the second column on the left, enhance the feature maps {P Figure 1 , P 2 , P 3 , P 4 , P 5} from top to bottom. First, normalize the channels of the feature map obtained in S2, and translate the topmost feature map P 5 to get N 5 . Perform pixel shuffling on the translated feature map to double the size of the original feature map. Then normalize the channels of the enlarged feature map for horizontal connection with the low-level feature map in S2 to obtain a new feature map N 4 . Repeat the operations of enlargement and horizontal connection until all horizontal connections with the low-level feature map in S2 are completed. The specific performance is as follows 4 Among them, N
[0062]
[0063] Among them, Ni is the i-th layer feature map after feature fusion, Conv 1×1 is a 1×1 convolution, PS(·) represents the pixel shuffle function to upsample the features, L(·) represents the channel normalization operation, N i+1 is the feature map of the previous layer, k n is the number of convolution kernels, s n is the stride of the convolution kernel movement, where k n = 256, s n = 1. P i is the feature map one layer lower than N obtained after Resnet50 feature extraction. The final output is the preliminarily fused feature map {N i+1 , N 2 , N 3 , N 4 , N 5}
[0064] S3.2: As shown in the third column on the left, enhance the feature map {N Figure 1 , N 2 , N 3 , N 4 , N 5} from bottom to top. First, translate the bottommost feature map. Thereafter, perform a convolution operation on each layer of the feature map to reduce the feature map by half, and then horizontally connect it with the corresponding N i feature map and P i feature map, and perform a convolution operation on the connected feature map to finally generate a new high-resolution feature map. Specifically, it is shown as:
[0065]
[0066] where F i is the output high-resolution i-th layer feature map, Conv 3×3 is a 3×3 convolution, k f is the number of convolution kernels, where k f = 2, 2 indicates that the stride of the convolution kernel movement is 2, N i is the i-th layer feature map in S3.1, P i is the feature map output in the i-th stage of S2, 1 indicates the stride of the convolution kernel movement. The final output feature map is {F 2 , F 3 , F 4 , F 5}.
[0067] S4: Through the attention module, the final fused multi-scale features are operated along the two dimensions of space and channel. In the two sub-modules of space and channel, the scaling factor and weight of the feature map are calculated respectively, and multiplied with the original feature map to adaptively adjust the features, so that the network learns to focus on the key information of the feature map. Finally, the output of the two sub-modules of space and channel is added to the input feature map to obtain the final result. The specific method is as follows:
[0068] S4.1: If Figure 1 The second column from the right and Figure 2 As shown in the figure, the final fused feature map is passed through the attention module to suppress the insignificant features in the channel and space, and enhance the ability to obtain significant features. The attention module is composed of two sub-modules, spatial attention and channel attention, in parallel. The input feature map passes through the spatial attention and channel attention modules in parallel, and finally the output features of the two attention modules and the features of the input feature map are summarized to obtain a better pixel-level prediction feature representation. The specific performance is:
[0069] Out a =SUM(SA(F i ),CA(F i ),F i ) (5)
[0070] Out a refers to the feature map of the final output, SUM(·) refers to the element summation function, which is used to complete feature fusion, SA(·) is the spatial attention module, CA(·) is the channel attention module, and F i is the input feature map.
[0071] S4.2: If Figure 2 As shown in the upper layer of , the feature map is input into the spatial attention module to improve its representation ability. First, each pixel in the feature map is batch normalized to obtain its scaling factor, and the weight is calculated by the scaling factor, and then added to the input feature map through the activation function to obtain the final output. The specific performance is:
[0072] Out s =sigmoid(W α (BN s (F i ))) (6)
[0073] Out s represents the output of the spatial attention module, sigmoid(·) represents the sigmoid activation function, and W α Represents the weight information of the channel attention module, BN s represents the scaling factor of the channel attention module, F iRepresents the input feature map.
[0074] S4.3: As Figure 2 shown in the lower layer of [], input the feature map into the channel attention module to summarize the spatial features. Use the scaling factor in batch normalization to reflect the variability and importance of each channel, and obtain the weights through the scaling factor. First, obtain the scaling factor of the feature map, then obtain the weight information through the scaling factor, and after passing through the activation function, add it to the information of the input feature map to obtain the final result. Specifically, it is as follows:
[0075] Out c = sigmoid(W β (BN c (F i ))) (7)
[0076] where Out c represents the output of the channel attention module, sigmoid(·) represents the sigmoid activation function, W β represents the weight information of the channel attention module, BN c represents the scaling factor of the channel attention module, F i represents the input feature map.
[0077] S5: Generate multiple candidate boxes with different ratios and sizes, obtain the candidate box list for each output feature map position, calculate the candidate boxes corresponding to the original image on each feature map, and calculate the positive and negative sample attributes of each candidate box using the ground truth information. The specific method is as follows
[0078] On the feature map output by the attention module, generate multiple candidate boxes with different ratios and sizes on each feature map through the candidate region generation method. The number of candidate boxes is adjusted according to the specific situation. In this embodiment, the number adopted is 9. And change the size of the candidate boxes according to the overall size of the targets in the dataset. Traverse each coordinate point on the feature map, map it to the coordinates of the input original image, and calculate the candidate box list on the original image according to the size and position of the candidate box on the feature map. Calculate the intersection over union (IoU) between the ground truth and the candidate boxes to obtain the positive and negative sample attributes of each candidate box. If the IoU is greater than or equal to 0.5, the candidate box is a high-quality positive sample, and the possibility of having a target inside the candidate box is high. If the IoU is lower than 0.4, the candidate box is a negative sample, and the possibility of having a target inside the box is low. If the IoU is in the interval [0.4, 0.5), the candidate box is an ignored sample and is not calculated.
[0079] S6: Feed the obtained positive samples into the detection module and input them into the classification and regression sub-networks respectively. The detection module selects a one-stage detector as needed. The classification network calculates which class each positive sample belongs to through a fully connected layer and a normalization function and outputs a probability vector of the class. The values in the probability vector correspond to different classes respectively, and the class corresponding to the largest value is the predicted class of the positive sample. The regression network obtains the position offset of each positive sample through bounding box regression to get a more accurate object detection box. Finally, non-maximum suppression is used to remove redundant prediction boxes and retain the one with the best result. The final detection result is the position of the object detection box and the predicted class of the object detection box.
Claims
1. An airborne image vehicle target recognition method based on feature fusion and attention mechanism, characterized in that, it includes the following steps: S1: Construct an airborne image vehicle target recognition model; The airborne image vehicle target recognition model includes a feature extraction module, a feature fusion module, an attention mechanism module, and a detection module; the feature extraction module is used to extract features from the input original picture; Feature The fusion module fuses the extracted features to obtain richer semantic information; the attention mechanism highlights the key information in the fused features; The detection module is used to obtain the category and location of the target; S2: Through the feature extraction module, use a convolutional neural network to extract features from the original picture and output a feature map with multiple scales; S3: Upsample the top-level feature map through the feature fusion module and perform horizontal connection with the low-level features to construct a top-down pyramid structure, and output the preliminarily fused multi-scale features; then process the preliminarily fused multi-scale features, downsample the low-level features, and perform horizontal connection with the high-level feature map and the high-level feature map output in S2 respectively to form a bottom-up pyramid structure, and output the finally fused multi-scale features; S4: Through the attention module, operate on the finally fused multi-scale features along two dimensions of space and channel; in the two sub-modules of space and channel, calculate the scaling factor and weight of the feature map respectively, and multiply them with the original feature map to adaptively adjust the features, so that the network learns to focus on the key information of the feature map; finally, add the outputs of the two sub-modules of space and channel to the input feature map to get the final result; S5: Generate multiple candidate boxes with different ratios and sizes, obtain a list of candidate boxes at each output feature map position, calculate the candidate boxes corresponding to the original picture on each feature map, and calculate the positive and negative sample attributes of each candidate box using the ground truth information; S6: The detection module selects a one-stage detector as needed, which includes two sub-networks for classification and regression. Send the feature map into the two sub-networks for classification and regression respectively, which are used to judge the category of the target and the specific location of the target, obtain the confidence of each prediction box and its category, and use the regression network to correct the position; finally, use non-maximum suppression to remove redundant prediction boxes and retain the best one to get the final detection result; The detailed steps of S3 are as follows: S3.1: Enhance the feature maps {P 2 , P 3 , P 4 , P 5} from top to bottom; first, normalize the channels of the feature maps obtained in S2, and translate the topmost feature map P 5 to get N 5 . Pixel shuffle the translated feature map to double the size of the original feature map; then normalize the channels of the enlarged feature map for horizontal concatenation with the low-level feature map in S2 to obtain a new feature map N 4 ; N 4 Repeat the operations of enlargement and horizontal concatenation until all horizontal concatenations with the low-level feature maps in S2 are completed; specifically manifested as: Where N i is the i-th layer feature map after feature fusion, Conv 1×1 is a 1×1 convolution, PS(·) represents the pixel shuffle function for upsampling the features, L(·) represents the channel normalization operation, N i+1 is the previous layer feature map, k n is the number of convolution kernels, s n is the step size of the convolution kernel movement, where k n = 256, s n = 1; P i is the feature map one layer lower than N obtained after Resnet50 feature extraction; finally, the initially fused feature maps {N i+1 , N 2 , N 3 , N 4 , N 5} are output S3.2: Enhance the feature maps {N 2 , N 3 , N 4 , N 5} from bottom to top; first, translate the bottommost feature map; thereafter, perform a convolution operation on each layer of the feature map to reduce the feature map by a factor of two, and then perform a horizontal connection with the corresponding N i feature map and the P i feature map, and perform a convolution operation on the connected feature map to finally generate a new high-resolution feature map; specifically: where F i is the high-resolution i-th layer feature map of the output, Conv 3×3 is a convolution with a size of 3×3, k f is the number of convolutional kernels, where k f = 2, and 2 represents that the convolutional kernel moves with a stride of 2, N i is the i-th layer feature map in S3.1, P i is the feature map output in the i-th stage in S2, and 1 represents the stride of the convolutional kernel; the finally output feature map is {F 2 , F 3 , F 4 , F 5}.
2. According to the airborne image vehicle target recognition method based on feature fusion and attention mechanism described in claim 1, characterized in that, The detailed steps of S2 are as follows: Use Resnet50 as the backbone of the feature extraction module to extract the features of the original picture; Resnet50 has five stages, and each stage outputs a feature map, which is also the input of the next stage; the scales of the feature maps output by different stages are different, and the higher the feature map, the smaller the scale and the more channels; First, process the input original picture, and the specific process is as follows: Out=Conv 7×7 (C, W, H, k, s) (1) Among them, Out represents the feature map output in the first stage, and Conv 7×7 represents a convolutional layer with a size of 7×7. C represents the number of channels of the input image. The number of channels of an RGB image is 3. W and H respectively represent the width and height of the input image. k represents the size of the convolutional kernel, and s represents the stride of the convolutional kernel movement; where k = 64 and s = 2; Secondly, perform feature extraction on the output feature map in turn, and the specific process is as follows: P i = Conv 3×3 (C i , W, H, k i , s i ) (2) Among them, P i is the feature map output in the i-th stage, and Conv 3×3 represents a convolutional layer with a size of 3×3. C i represents the number of channels of the feature map input in the i-th stage. W and H respectively represent the width and height of the input feature map; k i represents the size of the convolutional kernel, where k i = 256×2 i-1 ; s i represents the stride of the convolutional kernel movement, s i = 2; finally, four feature maps are output, which are {P 2 , P 3 , P 4 , P 5} from bottom to top, and the number of channels are {256, 512, 1024, 2048} respectively.
3. The airborne image vehicle target recognition method based on feature fusion and attention mechanism according to claim 2, characterized in that, the detailed steps of S4 are as follows: S4.1: Pass the finally fused feature map through the attention module to suppress insignificant features in the channels and space and enhance the ability to obtain significant features; the attention module is composed of two sub-modules, spatial attention and channel attention, in parallel; the input feature map passes through the spatial attention and channel attention modules in parallel, and finally the output features of the two attention modules and the features of the input feature map are aggregated to obtain a better pixel-level predicted feature representation; specifically: Out a = SUM(SA(F i ), CA(F i ), F i ) (5) Among them, Out a refers to the finally output feature map, SUM(·) refers to the element summation function for feature fusion, SA(·) is the spatial attention module, CA(·) is the channel attention module, and F i is the input feature map; S4.2: Input the feature map into the spatial attention module to improve its representation ability; first, batch-normalize each pixel in the feature map to obtain its scaling factor, calculate the weight through the scaling factor, pass through the activation function, and then add it to the input feature map to obtain the final output; specifically: Out s = sigmoid(W α (BN s (F i ))) (6) Among them, Out s represents the output of the spatial attention module, sigmoid(·) represents the sigmoid activation function, and W α represents the weight information of the channel attention module, and BN s represents the scaling factor of the channel attention module, and F i represents the input feature map; S4.3: Input the feature map into the channel attention module to aggregate spatial features; Use the scaling factor in batch normalization to reflect the variability and importance of each channel, and obtain the weight through the scaling factor; first obtain the scaling factor of the feature map, then obtain the weight information through the scaling factor, and add it to the information of the input feature map after passing through the activation function to obtain the final result; specifically: Out c = sigmoid(W β (BN c (F i ))) (7) Among them, Out c represents the output of the channel attention module, sigmoid(·) represents the sigmoid activation function, and W β represents the weight information of the channel attention module, and BN c represents the scaling factor of the channel attention module, and F i represents the input feature map.
4. The airborne image vehicle target recognition method based on feature fusion and attention mechanism according to claim 3, characterized in that, the detailed steps of S5 are as follows: On the feature map output by the attention module, generate multiple candidate boxes with different scales and sizes on each feature map through the candidate region generation method, and the number of candidate boxes is adjusted according to the specific situation; and change the size of the candidate boxes according to the overall size of the targets in the dataset; traverse each coordinate point on the feature map, map it to the coordinates of the original input image, and calculate the list of candidate boxes on the original image according to the size and position of the candidate boxes on the feature map; Calculate the intersection over union (IoU) between the ground truth and the candidate boxes to obtain the positive and negative sample attributes of each candidate box; if the IoU is greater than or equal to 0.5, the candidate box is a high-quality positive sample, and the possibility of there being a target inside the candidate box is high; if the IoU is lower than 0.4, the candidate box is a negative sample, and the possibility of there being a target inside the box is low; if the IoU is in the interval [0.4, 0.5), the candidate box is an ignored sample and no calculation is performed on it.
5. The airborne image vehicle target recognition method based on feature fusion and attention mechanism according to claim 4, characterized in that, the detailed steps of S6 are as follows: The obtained positive samples are sent into the detection module and input into the classification and regression sub-networks respectively; the detection module selects a one-stage detector as needed. The classification network calculates which category each positive sample belongs to through a fully connected layer and a normalization function and outputs a probability vector of the category; the values in the probability vector correspond to different categories respectively, and the category corresponding to the largest value is the predicted category of the positive sample; the regression network uses bounding box regression to obtain the position offset of each positive sample and obtains a more accurate object detection box; finally, non-maximum suppression is used to remove redundant prediction boxes and retain the one with the best result; the final detection result is the position of the object detection box and the predicted category of the object detection box.
Citation Information
Patent Citations
Remote sensing image multi-scale target detection method based on attention mechanism
CN111179217A
Remote sensing image vehicle target detection method based on multi-scale attention mechanism
CN111738110A