A Remote Sensing Aircraft Target Recognition Method Based on Deep Feature Fusion
By using deep feature fusion and prior box proportion adjustment methods in remote sensing aircraft target recognition, the problems of low efficiency and poor accuracy of remote sensing aircraft target detection in complex backgrounds are solved, and efficient identification of small targets is achieved.
Patent Information
- Application Number
- CN202110084548.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-01-21
AI Technical Summary
The prior art has low detection efficiency and poor accuracy of remote sensing aircraft targets in complex contexts, especially poor detection effect for small targets.
The remote sensing aircraft target recognition method based on deep feature fusion is adopted. By designing a feature fusion mechanism, the shallow, middle and deep feature maps are fused, and the proportion of the prior frame is adjusted to adapt to the characteristics of the remote sensing aircraft target.
It improves the recognition accuracy and detection efficiency of remote sensing aircraft targets in complex contexts, especially in small target detection.
Smart Images

Figure CN114821347B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image target recognition, and particularly relates to a method for remote sensing aircraft target recognition based on deep feature fusion. Background Technique
[0002] The technology of remote sensing image target detection is one of the important research directions in the field of computer vision. Compared with other types of remote sensing images such as synthetic aperture radar, high-resolution optical remote sensing images have their unique advantages. In both civilian and military fields, the detection of optical remote sensing images has been more and more widely applied. Traditional methods for remote sensing image target detection, such as scale-invariant feature transform, histogram of oriented gradients (HOG), deformable part model (DPM), etc., all rely on manually designed features, with low detection accuracy and poor real-time performance. With the excellent performance of deep learning in image classification, convolutional neural networks have been widely used in various fields of computer vision. Using deep learning to achieve target detection in the field of target detection has become a new direction.
[0003] Li Jianwei et al. proposed an algorithm based on the Faster R-CNN algorithm and the idea of feature fusion (Li Jianwei, Qu Changwen, Peng Shujuan, etc. Ship target detection in SAR images based on convolutional neural network [J]. Systems Engineering and Electronics, 2018, 40(9): 1953-1959). This algorithm uses targets such as ships and oil tanks in remote sensing images for verification experiments and achieves good results but has low detection efficiency. Xin Peng et al. proposed an aircraft target detection algorithm based on a multi-size convolutional neural network based on the SSD algorithm (Xin Peng, Xu Yuelai, Tang Hong, etc. Fast detection of aircraft with multi-layer feature fusion of fully convolutional networks [J]. Acta Optica Sinica, 2018, 38(3): 0315003), which improves the detection accuracy of multi-size aircraft targets in remote sensing images, but the detection speed is greatly reduced, and at the same time, the detection accuracy of small targets by this method is poor.
[0004] The above two-stage classes can obtain richer features and higher accuracy, but the detection speed is relatively slow; the single-stage algorithm has a simple structure and fast detection speed, but for the remote sensing target scene, the detection accuracy is poor, and neither can achieve the efficient detection of small remote sensing targets in complex backgrounds. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for remote sensing aircraft target recognition based on deep feature fusion to achieve the efficient detection of remote sensing aircraft targets in complex backgrounds.
[0006] The technical solution for achieving the purpose of the present invention is: A method for remote sensing aircraft target recognition based on deep feature fusion, including the following steps:
[0007] Step 1: Obtain remote sensing aircraft images, annotate aircraft targets using annotation software, and construct a VOC format dataset;
[0008] Step 2: Initialize the SSD network parameters: learning rate, the number of images fed into the SSD network in each batch, and the maximum number of iterations;
[0009] Step 3: Feed the images into the SSD network in batches to extract feature maps of 7 different sizes;
[0010] Step 4: For the feature maps obtained in Step 3, design a feature fusion mechanism. Each time, select 3 feature maps of different sizes, namely deep, middle, and shallow feature maps, and feed them into the fusion mechanism to fuse the shallow, middle, and deep features to obtain the fused feature map;
[0011] Step 5: For each feature map of different sizes including the fused feature map generated in Step 4, generate prior boxes of different sizes, and adjust the sizes of the prior boxes relative to the input images according to the characteristics of the remote sensing aircraft target sizes;
[0012] Step 6: Calculate the matching coefficient between the prior boxes obtained in Step 5 and the ground truth boxes, i.e., the true position boxes of the targets, and divide them into positive and negative samples. Then calculate the value of the loss function, i.e., the loss value, and optimize and improve the SSD network according to the loss value gradient; return to Step 3, feed the next batch of images into the improved SSD network until the maximum number of iterations is reached to obtain the improved SSD network model;
[0013] Step 7: Use the obtained improved SSD network model to complete the recognition of remote sensing aircraft targets.
[0014] Furthermore, in Step 4, the design of the feature fusion mechanism, each time selecting 3 feature maps of different sizes, namely deep, middle, and shallow feature maps, and feeding them into the fusion mechanism to fuse the shallow, middle, and deep features to obtain the fused feature map, includes the following steps:
[0015] Step 4.1: Extract three feature maps with different depths and sizes respectively: shallow feature map, middle feature map, and deep feature map;
[0016] Step 4.2: For the small-sized, high-semantic information deep feature maps extracted, perform upsampling using bilinear interpolation to the size of the middle feature map; then perform convolution operation using a 3×3 convolution kernel to eliminate the redundant information brought during the upsampling process; then perform 1×1 convolution to reduce the number of channels;
[0017] Step 4.3: For the large-sized, low-semantic information shallow feature maps extracted, perform convolution and upsampling to the size of the middle feature map, and then perform 1×1 convolution to reduce the number of channels;
[0018] Step 4.4: After upsampling and downsampling the deep and shallow feature maps to the size of the middle-layer feature map, perform vector splicing on the three feature maps with different depths but the same size, that is, merge the channels of the three-layer feature maps, and perform a 3×3 convolution on the feature map after channel merging to finally generate the fused feature map.
[0019] Further, generating prior boxes of different sizes in Step 5 includes the following steps:
[0020] Step 5.1: Divide the extracted feature maps of different sizes into an N×N grid structure by pixels. Use the center point of each grid as the center point of the prior box. Assuming the size of the feature map is d, it is equivalent to moving the prior box to detect the target at intervals of d / N.
[0021] Step 5.2: The prior boxes generated with the center point of each grid as the center have different sizes and different aspect ratios. There are 5 kinds of prior box aspect ratios: {1, 2, 3, 1 / 2, 1 / 3}.
[0022] According to the different sizes of the input feature maps, the sizes of the prior boxes are also different. For the size of the prior box, the calculation formula is:
[0023]
[0024] In the formula, m = 5, s k represents the ratio of the k-th prior box size to the input image, s min represents the minimum ratio of the prior box size to the original image, s max represents the maximum ratio of the prior box size to the original image;
[0025] Step 5.3: According to the characteristics of the target size in the remote sensing aircraft image, adjust the size of the prior box to match the target size, so that the receptive field size of the constructed feature is consistent with the target size, that is, adjust s min and s max values, so that the s k value of the shallow feature map decreases to adapt to the aircraft target size.
[0026] Further, in Step 6, the determination process of the improved SSD network model includes the following steps:
[0027] Step 6.1: Calculate the intersection over union (IOU) value between the prior box and the ground truth box:
[0028]
[0029] where B default represents the prior box, and B ground represents the ground truth box;
[0030] Select positive and negative samples according to a preset IOU threshold, and control the ratio of positive and negative samples to 1:3;
[0031] Step 6.2, calculate the loss function, and the loss function L is the weighted sum of the classification loss and the location loss:
[0032]
[0033] Among them, L conf (x, c) represents the loss that the prior box belongs to the aircraft target or the background, and L loc (x, l, g) represents the loss of coordinate regression of the prior box; x represents the prior box, c represents the matching probability, l represents the predicted value of the prior box position, g represents the true box position parameter, and N represents the number of positive sample sets; α is the weight and is set to 1;
[0034] Step 6.3, use the derivative of the loss function, and optimize the SSD network according to the stochastic gradient descent method; return to Step 3, send the next batch of pictures into the network until the maximum number of iterations is reached, and obtain the SSD network model.
[0035] Compared with the prior art, the significant advantages of the present invention are: (1) In view of the small size of the remote sensing aircraft target, a feature fusion module is designed to fuse the large-size shallow feature map responsible for detecting small targets with the deep feature map with rich semantic information. While fusing the rich semantic information from the deep feature map to the shallow feature map, the rich position information is also fused from the shallow feature map to the deep feature map, enriching the features of the feature map and improving the recognition accuracy of the remote sensing aircraft target; (2) According to the characteristic that the remote sensing aircraft target occupies few pixels, the ratio of the prior box in each size feature map to the original image is adjusted, making it better adapt to the small target recognition problem.
[0036] The present invention will be further described below in conjunction with the accompanying drawings of the specification. Brief Description of the Drawings
[0037] Figure 1 It is a flowchart of the remote sensing aircraft target recognition method based on deep feature fusion of the present invention.
[0038] Figure 2 It is a schematic diagram of the feature fusion mechanism.
[0039] Figure 3 It is a test result diagram of each method on three groups of images, where (a) is the test result diagram of the DSSD method, (b) is the test result diagram of the SSD method, and (c) is the test result diagram of the present invention. Detailed Embodiment
[0040] In view of the characteristics of remote sensing aircraft images with few target pixel sizes and insufficient information, the present invention proposes a remote sensing aircraft target recognition method based on deep feature fusion, designs a feature map fusion mechanism, fuses shallow feature maps with high resolution and deep feature maps with rich semantic information, and at the same time adjusts the ratio of the prior box to the original image, enabling the algorithm to better adapt to the problem of remote sensing aircraft target recognition. Experimental verification on the remote sensing aircraft dataset proves the effectiveness of the method.
[0041] A remote sensing aircraft target recognition method based on deep feature fusion according to the present invention includes the following steps:
[0042] Step 1: Obtain remote sensing aircraft pictures, use annotation software to annotate aircraft targets, and construct a VOC format dataset;
[0043] Step 2: Initialize the SSD network parameters: learning rate, the number of pictures sent into the SSD network in each batch, and the maximum number of iterations;
[0044] Step 3: Send the pictures into the SSD network in batches and extract 7 feature maps of different sizes;
[0045] Step 4: For the feature maps obtained in Step 3, design a feature fusion mechanism. Each time, select 3 feature maps of different sizes, namely deep, middle, and shallow, and send them into the fusion mechanism to fuse the shallow feature map, middle feature map, and deep feature to obtain a fused feature map;
[0046] Step 5: For each feature map of different sizes including the fused feature map generated in Step 4, generate prior boxes of different sizes, and adjust the size of the prior box relative to the input picture according to the characteristics of the remote sensing aircraft target size;
[0047] Step 6: Calculate the matching coefficient between the prior box obtained in Step 5 and the true box, that is, the true position box of the target, and divide the positive and negative samples. Then calculate the value of the loss function, that is, the loss value, and optimize and improve the SSD network according to the loss value gradient; return to Step 3, send the next batch of pictures into the improved SSD network until the maximum number of iterations is reached to obtain an improved SSD network model;
[0048] Step 7: Use the obtained improved SSD network model to complete the remote sensing aircraft target recognition.
[0049] Furthermore, in Step 4, the designed feature fusion mechanism, each time selects 3 feature maps of different sizes, namely deep, middle, and shallow, and sends them into the fusion mechanism to fuse the shallow feature map, middle feature map, and deep feature to obtain a fused feature map, including the following steps:
[0050] Step 4.1: Extract three feature maps with different depths and sizes respectively: shallow feature map, middle feature map, and deep feature map;
[0051] Step 4.2: For the extracted small-size, high-semantic information deep feature map, perform upsampling using the bilinear interpolation method to the size of the middle feature map; then perform convolution operation using a 3×3 convolution kernel to eliminate the redundant information brought by the upsampling process; then perform 1×1 convolution to reduce the number of channels;
[0052] Step 4.3: For the extracted large-size, low-semantic information shallow feature map, perform convolution to upsample it to the size of the middle feature map, and then perform 1×1 convolution to reduce the number of channels;
[0053] Step 4.4: After upsampling and downsampling the deep and shallow feature maps to the size of the middle feature map, perform vector splicing on the three feature maps with different depths and the same size, that is, merge the channels of the three feature maps, and perform 3×3 convolution on the feature map after channel merging to finally generate the fused feature map.
[0054] Further, generating prior boxes of different sizes in step 5 includes the following steps:
[0055] Step 5.1: Divide the extracted feature maps of different sizes into an N×N grid structure by pixels. Take the center point of each grid as the center point of the prior box. Suppose the size of the feature map is d, which is equivalent to moving the prior box to detect the target every d / N distance;
[0056] Step 5.2: The prior boxes generated with the center point of each grid as the center have different sizes and aspect ratios. There are 5 kinds of prior box aspect ratios: {1, 2, 3, 1 / 2, 1 / 3};
[0057] According to the different sizes of the input feature maps, the sizes of the prior boxes are also different. For the size of the prior box, the calculation formula is:
[0058]
[0059] In the formula, m = 5, s k represents the ratio of the k-th prior box size to the input image, s min represents the minimum ratio of the prior box size to the original image, s max represents the maximum ratio of the prior box size to the original image;
[0060] Step 5.3: According to the characteristics of the target size in the remote sensing aircraft image, adjust the size of the prior box to match the target size, so that the receptive field size of the constructed feature is consistent with the target size, that is, adjust s min 、s maxThe value such that s of the shallow feature map k The value decreases to adapt to the size of the aircraft target.
[0061] Furthermore, for the improved SSD network model in step 6, the determination process includes the following steps:
[0062] Step 6.1: Calculate the intersection over union (IOU) value between the prior box and the ground truth box:
[0063]
[0064] where B default represents the prior box, and B ground represents the ground truth box;
[0065] Select positive and negative samples according to a preset IOU threshold, and control the positive-negative sample ratio to 1:3;
[0066] Step 6.2: Calculate the loss function, where the loss function L is the weighted sum of the classification loss and the localization loss:
[0067]
[0068] where L conf (x, c) represents the loss of the prior box belonging to the aircraft target or the background, and L loc (x, l, g) represents the loss of the coordinate regression of the prior box; x represents the prior box, c represents the matching probability, l represents the predicted value of the prior box position, g represents the ground truth box position parameter, and N represents the number of positive sample sets; α is the weight, set to 1;
[0069] Step 6.3: Use the derivative of the loss function to optimize the SSD network according to the stochastic gradient descent method; return to step 3, send the next batch of images into the network until the maximum number of iterations is reached, and obtain the improved SSD network model.
[0070] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0071] Embodiment
[0072] Combined with Figure 1 , a remote sensing aircraft target recognition method based on deep feature fusion according to the present invention includes the following steps:
[0073] Step 1: Collect remote sensing aircraft images from the network, use annotation software to annotate the aircraft targets in each image, and generate xml files. Divide the images into a training set and a test set according to a certain ratio, and construct a remote sensing aircraft dataset in the VOC dataset format.
[0074] Step 2: Initialize the network parameter settings, set the learning rate lr and the number of images batch-size fed into the network each time during network training, as well as the maximum number of training times. Downsample the images to a size of 300*300, and perform data augmentation to increase the number of samples in the dataset. The image augmentation methods adopted include: expansion, cropping, mirroring, etc.
[0075] Step 3: Feed the generated images into the network to generate feature maps of different sizes, and extract seven feature maps of different sizes: Conv3_3, Conv4_3, fc7, Conv8_2, Conv9_2, Conv10_2, and Conv11_2. The sizes of each feature map are: 75×75, 38×38, 19×19, 10×10, 5×5, 3×3, and 1×1 respectively.
[0076] Step 4: For the feature maps generated in Step 3, design a feature fusion mechanism to fuse shallow features and deep features. The specific steps of Step 4 are as Figure 2 shown, including the following steps: Fuse shallow features and deep features.
[0077] Step 4.1: Extract feature maps of three sizes: 75×75, 38×38, and 10×10. First, perform convolution on the 75×75-sized feature map, downsample it to 38×38, and perform 1×1 convolution to change its number of channels. Second, simultaneously perform bilinear interpolation on the 10×10-sized feature map with rich semantic information to upsample it to 38×38, and then use 3×3 and 1×1 convolutions to remove redundancy and change its number of channels. Third, splice these three feature maps with the same size on the channels. Finally, perform smoothing on it using 1×1 convolution operation to generate a fused 38×38-sized feature map.
[0078] Step 4.2: Extract feature maps of three sizes: 38×38, 19×19, and 5×5. First, perform convolution on the 38×38-sized feature map, downsample it to 19×19, and perform 1×1 convolution to change its number of channels. Second, simultaneously perform bilinear interpolation on the 5×5-sized feature map with rich semantic information to upsample it to 19×19, and then use 3×3 and 1×1 convolutions to remove redundancy and change its number of channels. Third, splice these three feature maps with the same size on the channels. Finally, perform smoothing on it using 1×1 convolution operation to generate a fused 19×19-sized feature map.
[0079] Step 4.3: Extract feature maps of three sizes: 19×19, 10×10, and 3×3. First, perform convolution on the 19×19-sized feature map, downsample it to 10×10, and perform 1×1 convolution on it to change its number of channels. Second, simultaneously perform bilinear interpolation on the 3×3-sized feature map with rich semantic information, upsample it to 10×10, and then use 3×3 and 1×1 convolutions to remove redundancy and change the number of channels. Third, concatenate these three feature maps with the same size on the channels. Finally, perform smoothing on it using 1×1 convolution operation to generate a fused 10×10-sized feature map.
[0080] The new feature map generated by fusing deep, middle, and shallow feature maps of three different sizes has a large size and rich semantic information at the same time, making the detection of small remote sensing targets more accurate.
[0081] Step 5: Extract six different-sized feature maps including the fused feature map generated in Step 4: 38×38, 19×19, 10×10, 5×5, 3×3, and 1×1, and generate prior boxes of different sizes on these six feature maps. The said Step 4 includes the following steps:
[0082] Step 5.1: Divide the six different-sized feature maps extracted into an N×N grid structure by pixels, and use the center point of each grid as the center point of the prior box. It is equivalent to moving the prior box to detect the target at intervals of the feature map size / N.
[0083] Step 5.2: Generate prior boxes of different sizes and different ratios with the center point of each grid as the center.
[0084] Step 5.3: According to the different sizes of the input feature maps, the sizes of the prior boxes are also different. For the size of the prior box, the calculation formula for the size of the prior box is:
[0085]
[0086] In the formula, m = 5, s k represents the ratio of the prior box size to the original picture. s min represents the minimum ratio of the prior box size to the original picture, and s max represents the maximum ratio of the prior box size to the original picture
[0087] Step 5.4: The original s min , s maxThe values are adjusted from 0.2 and 0.9 to 0.12 and 0.88. Then, the sizes of the prior boxes generated on six different-sized feature maps are adjusted from the original 30, 60, 111, 162, 213, 264 to 18, 36, 111, 150, 207, 264. There are 5 types of aspect ratios for the prior boxes: {1, 2, 3, 1 / 2, 1 / 3}, and a total of 8732 prior boxes are generated on six different-sized feature maps.
[0088] Step 6: Divide the positive and negative samples according to the matching coefficient between the prior boxes obtained in Step 5 and the ground truth boxes, calculate the loss value, and adjust the network according to the loss value gradient to obtain the network model. Specifically as follows:
[0089] Step 6.1 Calculate the matching index, i.e., the IOU value, between the 8732 prior boxes generated in Step 5 and the ground truth boxes:
[0090]
[0091] where B default represents the default box, i.e., the prior box, and B ground represents the ground truth box, i.e., the most real box.
[0092] Step 6.2: Obtain the positive and negative samples. Obtaining the positive sample set: First, find a prior box with the largest IOU value for each ground truth box and put it into the positive sample set; second, select positive samples according to a pre-set IOU threshold, and those with an IOU greater than the threshold are put into the positive sample set.
[0093] Obtaining the negative sample set: Since the number of objects is small while the number of prior boxes is large, the number of prior boxes with an IOU less than the threshold is much larger than those greater than the threshold. To balance the number of positive and negative samples, hard negative mining is performed on the negative samples, that is, difficult negative samples that are likely to be misjudged are selected from the negative samples and put into the negative sample set to control the positive-negative sample ratio at approximately 1:3.
[0094] Step 6.3: Calculate the loss function. The loss function includes classification loss and localization loss. The softmax loss is used to calculate the classification loss:
[0095]
[0096] where, indicates whether the i-th prior box matches the j-th ground truth box of class p. The matching value is 1, and the non-matching value is 0; indicates the probability that the i-th prior box matches class p, indicates the probability that the i-th prior box matches the background class.
[0097] The smooth L1 loss is used to calculate the localization loss:
[0098]
[0099]
[0100]
[0101]
[0102] Among them, {cx, cy, w, h} represents the coordinates of the center point position and the width and height of the box. It is the encoded value of g, representing the predicted offset.
[0103] The loss function is the weighted sum of the classification loss and the location loss:
[0104]
[0105] Among them, α is the weight of the classification loss and the location loss, which is set to 1 through cross-validation.
[0106] Step 6.4: Calculate the gradient value of the loss function, and use the stochastic gradient descent method to update the network parameters until the maximum number of iterations is reached to obtain the network model.
[0107] Step 7: Use the obtained SSD network model to complete the remote sensing aircraft target recognition.
[0108] A total of 429 remote sensing aircraft images were collected from the network. The aircraft targets in each image were labeled using the annotation software labelImage to generate xml files. The images were divided into a training set and a test set according to the ratio of 0.8 and 0.2. Three text files, train.txt, test.txt, and trainval.txt, storing the names of the images in each part were generated using code. Finally, a remote sensing aircraft dataset was constructed in the VOC dataset format.
[0109] Using the remote sensing aircraft dataset, the method of the present invention was compared with traditional methods such as SSD and DSSD. The learning rate lr of all methods was set to 0.0001, the batch-size was set to 4, and the maximum number of training times was set to 50000. The experimental test results are as Figure 3 shown in (a) - (c).
[0110] Table 1 shows the recognition accuracy (mAP) of each method for each algorithm on the remote sensing aircraft dataset. The experimental results show that compared with other traditional methods, the method proposed in the present invention has higher recognition accuracy for remote sensing aircraft targets in complex backgrounds.
[0111] Table 1 Recognition accuracy (mAP) of each method on the remote sensing aircraft dataset
[0112]
[0113] In summary, the remote sensing aircraft target recognition method based on deep feature fusion proposed by the present invention can achieve efficient recognition of remote sensing aircraft targets under complex backgrounds.
Claims
1. A remote sensing aircraft target recognition method based on deep feature fusion, characterized in that, it includes the following steps: Step 1: Obtain remote sensing aircraft pictures, use annotation software to annotate aircraft targets, and construct a VOC format dataset; Step 2: Initialize the SSD network parameters: learning rate, the number of pictures sent into the SSD network in each batch, and the maximum number of iterations; Step 3: Send the pictures into the SSD network in batches to extract feature maps of 7 different sizes; Step 4: For the feature maps obtained in Step 3, design a feature fusion mechanism. Each time, select 3 feature maps of different sizes, namely deep, middle, and shallow, and send them into the fusion mechanism to fuse the shallow feature map, middle feature map, and deep feature map to obtain a fused feature map; Step 5: For each feature map of different sizes including the fused feature map generated in Step 4, generate prior boxes of different sizes, and adjust the size of the prior boxes relative to the input picture according to the characteristics of the remote sensing aircraft target size; Step 6: Calculate the matching coefficient between the prior boxes obtained in Step 5 and the true boxes, that is, the true position boxes of the targets, and divide the positive and negative samples. Then calculate the value of the loss function, that is, the loss value, and optimize the network according to the loss value gradient; return to Step 3, send the next batch of pictures into the SSD network until the maximum number of iterations is reached to obtain the SSD network model; Step 7: Use the obtained SSD network model to complete the remote sensing aircraft target recognition; The designed feature fusion mechanism in Step 4, each time selects 3 feature maps of different sizes, namely deep, middle, and shallow, and sends them into the fusion mechanism to fuse the shallow feature map, middle feature map, and deep feature map to obtain a fused feature map, including the following steps: Step 4.1: Extract three feature maps with different depths and sizes respectively: shallow feature map, middle feature map, and deep feature map; Step 4.2: For the small-size, high-semantic information deep feature map extracted, use bilinear interpolation to perform upsampling to the size of the middle feature map; then use a 3×3 convolution kernel to perform convolution operation to eliminate the redundant information brought during the upsampling process; then perform 1×1 convolution to reduce the number of channels; Step 4.3: For the large-size, low-semantic information shallow feature map extracted, perform convolution to upsample to the size of the middle feature map, and then perform 1×1 convolution to reduce the number of channels; Step 4.4: After upsampling the deep and shallow feature maps to the size of the middle feature map, perform vector splicing on the three feature maps with different depths and the same size, that is, merge the channels of the three feature maps, and perform 3×3 convolution on the feature map after channel merging to finally generate a fused feature map; The generation of prior boxes of different sizes in Step 5 includes the following steps: Step 5.1: Divide each feature map of different sizes extracted into an N×N grid structure by pixels. Take the center point of each grid as the center point of the prior box. Let the size of the feature map be d, which is equivalent to moving the prior box to detect the target every d / N distance; Step 5.2: The prior boxes generated with the center points of each grid as the centers have different sizes and different aspect ratios. There are 5 kinds of aspect ratios of the prior boxes: {1, 2, 3, 1 / 2, 1 / 3}. According to the different sizes of the input feature maps, the sizes of the prior boxes are also different. For the sizes of the prior boxes, the calculation formula is as follows: where m = 5, s k represents the ratio of the size of the k-th prior box to the input image, s min represents the minimum ratio of the prior box size to the original image, s max represents the maximum ratio of the prior box size to the original image; Step 5.3: According to the characteristics of the target size in the remote sensing aircraft images, adjust the size of the prior box to match the target size, so that the receptive field size of the constructed feature is consistent with the target size, that is, adjust s min , s max values, so that the s k value of the shallow feature map decreases to adapt to the aircraft target size; In step 6, for the improved SSD network model, the determination process includes the following steps: Step 6.1: Calculate the intersection over union (IOU) value between the prior box and the ground truth box: Among them, B default represents the prior box, and B ground represents the ground truth box; Select positive and negative samples according to the pre-set IOU threshold, and control the ratio of positive and negative samples to be 1:3; Step 6.2: Calculate the loss function. The loss function L is the weighted sum of the classification loss and the localization loss: Among them, L conf (x, c) represents the loss that the prior box belongs to the aircraft target or the background, and L loc (x, l, g) represents the loss of coordinate regression of the prior box; x represents the prior box, c represents the matching probability, l represents the predicted value of the prior box position, g represents the true box position parameter, and N represents the number of positive sample sets; α is the weight, set to 1; Step 6.3: Use the derivative of the loss function, and optimize the improved SSD network according to the stochastic gradient descent method; return to step 3, send the next batch of images into the network until the maximum number of iterations is reached, and obtain the improved SSD network model.
Citation Information
Patent Citations
Detection method for picture small target
CN111860587A
Target detection improved algorithm based on feature pyramid network and attention mechanism
CN111914917A