Remote sensing image target detection model based on improved RetinaNet
By introducing improved downsampling module and core selection module in the RetinaNet model, the problem of poor object detection effect in remote sensing images is solved, and higher detection accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202410169692.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-02-06
AI Technical Summary
The existing RetinaNet object detection model has poor detection effect in remote sensing images, mainly because the remote sensing image targets are small and dense, the scale changes are large, and distributed in any direction.
An improved downsampling module and a kernel selection module are introduced. The improved downsampling module enhances feature extraction capabilities through slice downsampling and branch processing. The kernel selection module enhances the fusion of multi-scale feature information through dynamic selection of convolution kernels.
The accuracy and accuracy of remote sensing image object detection are improved. The experimental results show that the average accuracy of the whole class on the DOTA dataset is better than that of the traditional RetinaNet model.
Smart Images

Figure CN117710827B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a remote sensing image target detection model based on improved RetinaNet. Background Art
[0002] Remote sensing image target detection is a key task in high-resolution image content analysis, which aims to accurately identify and locate specific target objects in remote sensing images, such as vehicles, ships, and aircraft. This technology plays an important role in the field of high-precision remote sensing image intelligent analysis and is widely used in many fields such as intelligent transportation, urban planning, and geographic information system updates.
[0003] In recent years, the rapid development of deep learning has made significant progress in the field of general object detection. However, in the specific field of remote sensing image analysis, due to the characteristics of remote sensing images such as small and dense objects, large scale changes, and distribution in arbitrary directions, general object detectors, such as the traditional RetinaNet object detection model, do not perform well when directly applied to remote sensing images. Summary of the invention
[0004] The purpose of the present invention is to provide a remote sensing image target detection model based on improved RetinaNet, which improves the target detection capability on remote sensing images in view of the characteristics of remote sensing images that the targets are small and dense, with large scale variations and distributed in arbitrary directions.
[0005] A remote sensing image target detection model based on improved RetinaNet, comprising a RetinaNet backbone network, a feature pyramid and a classification regression subnet, an improved downsampling module is introduced into the RetinaNet backbone network, and the model further comprises a kernel selection module;
[0006] The backbone network uses the improved downsampling module to perform downsampling when performing residual learning. The improved downsampling module copies the input image feature P as the image feature P 1 and image feature P 2 , where P∈ R H×W×C , R represents a real number, W , H and C Respectively represent the width, height and number of channels of the image feature. The improved downsampling module performs 1 Slice downsampling is performed, and after slicing processing, four spatial downsampled image features C are obtained 1 , C 2 , C 3 and C 4, the process of slicing downsampling, in the channel dimension, splicing image features C 1 , C 2 , C 3 and C 4 , get the new image features, after splicing, make the image features P 1 The number of channels of the new image feature is increased from C to 4C. Then, a 1×1 convolution operation with a step size of 1 is used to compress the number of channels of the new image feature to 2C, and the image feature Q is obtained. 1 ;
[0007] The improved downsampling module performs image feature P 2 Two branches are used for processing. In one branch, a group convolution GConv with a step size of 1 and a size of 3×3 is used for processing, and then a 3×3 convolution with a step size of 2 is used for downsampling, and a GELU activation function and a normalization layer are used to obtain the image feature Q 2 On the other branch, use the group convolution GConv with a step size of 1 and a size of 3×3, and perform maximum pooling and normalization to obtain the image feature Q 3 ;
[0008] Concatenate image features Q in the channel direction 1 , Q 2 and Q 3 , and use a 1×1 convolution layer on the concatenated result to get the image features ;
[0009] The kernel selection module dynamically selects a variety of different convolution kernel fusion features according to the characteristics of the input image.
[0010] According to the remote sensing image target detection model based on improved RetinaNet provided by the present invention, an improved downsampling module is introduced and embedded in the RetinaNet backbone network, and the three downsampling methods are fused to generate downsampled image features for the extracted features, thereby enhancing the model's ability to capture complex details. The convolution kernel selection mechanism of the kernel selection module is used to dynamically select the spatial receptive field, thereby enhancing the model's ability to extract and fuse multi-scale feature information, thereby modeling multi-scale information, and finally obtaining the classification and regression results of the target object. Experimental results show that the model of the present invention has a better average accuracy rate of all categories on the large-scale remote sensing image target detection dataset DOTA than the traditional RetinaNet target detection model, and can detect remote sensing targets more accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a schematic diagram of the structure of the ResNet50 network in the present invention;
[0012] Figure 2A schematic diagram of the downsampling process performed by the improved downsampling module in the present invention;
[0013] Figure 3 Schematic diagram of the process of slice downsampling;
[0014] Figure 4 Schematic diagram of the working principle of the nuclear selection module. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0016] An embodiment of the present invention provides a remote sensing image target detection model based on an improved RetinaNet, including a RetinaNet backbone network, a feature pyramid and a classification regression subnet. The present invention introduces an improved downsampling module into the RetinaNet backbone network to enhance the model's ability to capture complex details. In addition, the model also includes a kernel selection module for enhancing the network's ability to extract and fuse multi-scale feature information.
[0017] In remote sensing images, the scale of objects varies greatly, and the number of small objects accounts for a large proportion. The downsampling method used by the traditional RetinaNet object detection model mainly relies on the convolution layer, which may cause some key semantic information to be missed, and it is difficult to fully mine and retain fine-grained feature information. To solve this problem, the present invention introduces an improved downsampling module (IDM for short). Taking the ResNet50 network as an example, the position of the improved downsampling module in the network is as follows: Figure 1 As shown in Figure 1, an IDM is added at the input or output position of each bottleneck building block of the ResNet50 network.
[0018] The ResNet50 network consists of a series of stacked residual blocks, each of which contains multiple convolutional layers and identity mappings. IDM is used for downsampling during residual learning. The downsampling process is as follows: Figure 2 The present invention uses three branches to process the input features, realizes the extraction and fusion of multi-scale features, enhances the representation ability of features, and thus reduces the detail loss of the model during small target detection.
[0019] Specifically, the backbone network uses the improved downsampling module to perform downsampling when performing residual learning, and the improved downsampling module copies the input image feature P as the image feature P 1 and image feature P 2 , where P∈ R H×W×C , R represents a real number, W , H and C Respectively represent the width, height and number of channels of the image feature. The improved downsampling module performs 1 Slice downsampling is performed, and after slicing processing, four spatial downsampled image features C are obtained 1 , C 2 , C 3 and C 4 , the process of slice downsampling is as follows Figure 3 As shown, Figure 3 middle, x 11 , x 12 , x 13 , x 14 , x 21 , x 22 , x 23 , x 24 , x 31 , x 32 , x 33 , x 34 , x 41 , x 42 , x 43 , x 44 , x (H)(W) , x (H-1)(W-1) , x (H-1)(W) , x (H)(W-1) Represent the image features P 1Features at spatial positions (1, 1), (1, 2), (1, 3), (1, 4), (2, 1), (2, 2), (2, 3), (2, 4), (3, 1), (3, 2), (3, 3), (3, 4), (4, 1), (4, 2), (4, 3), (4, 4), (H, W), (H-1, W-1), (H-1, W), and (H, W-1).
[0020] In the channel dimension, concatenate the image features C 1 , C 2 , C 3 and C 4 , get the new image features, after splicing, make the image features P 1 The number of channels of the new image feature is increased from C to 4C. Then, a 1×1 convolution operation with a step size of 1 is used to compress the number of channels of the new image feature to 2C, and the image feature Q is obtained. 1 , halving the number of image feature channels can reduce the amount of computation required for the model.
[0021] The improved downsampling module performs image feature P 2 Two branches are used for processing. In one branch, a group convolution GConv with a step size of 1 and a size of 3×3 is used for processing, and then a 3×3 convolution with a step size of 2 is used for downsampling, and a GELU activation function and a normalization layer are used to obtain the image feature Q 2 On the other branch, use the group convolution GConv with a step size of 1 and a size of 3×3, and perform maximum pooling and normalization to obtain the image feature Q 3 .
[0022] Specifically, the improved downsampling module performs image feature P 1 During the slicing downsampling process, the following conditions are met:
[0023] Q 1 =Conv(CutD(P 1 ));
[0024] Q 2 =GELU(BN(DWConvD(GConv(P 2 ))));
[0025] Q 3 =BN(MaxP(GConv(P 2 ));
[0026] Among them, Conv, CutD, GELU, BN, DWConvD, GConv, and MaxP represent convolution, slicing, GELU activation function, batch normalization, depth convolution, grouped convolution, and maximum pooling operations, respectively.
[0027] Concatenate image features Q in the channel direction 1 , Q 2 and Q 3 , and use a 1×1 convolution layer on the concatenated result to obtain a set of image features with doubled channels and half the size .
[0028] Image features The following conditions are met:
[0029] ;
[0030] Among them, Concat represents the operation of connecting features in the channel direction.
[0031] Also, see Figure 4 In order to improve the model's ability to detect targets of different scales, the present invention adopts a kernel selection module, which dynamically selects a variety of different convolution kernel fusion features according to the characteristics of the input image, thereby improving the model's expression ability.
[0032] In the detection task head of the model, for the input image feature K, the kernel selection module uses three dilated convolutions with kernel sizes of 3×3, 5×5, and 7×7 to learn multi-scale spatial information and obtain image features X with three different scale receptive fields. 1 , X 2 , X 3 , where X 1 ∈ R H×W×C , X 2 ∈ R H×W×C , X 3 ∈ R H×W×C Then, use channel concatenation to merge X 1 , X 2 , X 3 , get the image feature X, and concatenate the average pooling and maximum pooling results of the image feature X in the channel direction, then use convolution and Sigmoid functions to obtain independent spatial selection masks, and then use the spatial selection mask to 1 , X 2 , X 3 Weighted separately, we get the image features F 1 、F 2 、F 3Finally, for F 1 、F 2 、F 3 Add element by element to get the fusion feature with attention, and multiply the fusion feature and the input image feature K element by element to get the image feature .
[0033] Among them, the image feature X 1 , X 2 , X 3 The following conditions are met:
[0034] X 2 =DWConv(X 1 );
[0035] X 3 =DWConv(X 2 );
[0036] Among them, DWConv represents dilated convolution.
[0037] Image features The following conditions are met:
[0038] .
[0039] The present invention is tested below:
[0040] 1. Experimental subjects
[0041] The DOTA dataset is selected as the object used for testing. The DOTA dataset is a large-scale public dataset of aerial images for target detection tasks. It consists of 2806 large-size images and contains 15 types of objects of different scales, directions and shapes. The DOTA dataset contains 15 object categories, including airplanes (PL), baseball fields (BD), bridges (BR), track and field fields (GTF), small vehicles (SV), large vehicles (LV), ships (SH), tennis courts (TC), basketball courts (BC), oil tanks (ST), football fields (SBF), loops (RA), ports (HA), swimming pools (SP) and helicopters (HC). The resolution of the image is between 800×800 and 4000×4000. The present invention crops the image to a size of 1024×1024 with a stride of 200. The training set and test set contain 21046 and 10833 images, respectively. The test results are submitted to the DOTA evaluation server.
[0042] 2. Experimental Setup
[0043] The experiment uses a GeForce RTX3090 graphics card with 24GB of video memory to train and test the algorithm. The batch size and epoch of training are set to 2 and 12 respectively. SGD is used as the optimizer, and the initial learning rate and momentum coefficient are 0.0025 and 0.9 respectively. Average Precision (AP) and meanAverage Precision (mAP) are used as detection evaluation indicators. In addition, Params (the total number of model parameters) and Flops (the number of floating-point operations) are used to measure the computational complexity and number of parameters of the model.
[0044] 3. Ablation experiment
[0045] The present invention analyzes the contribution of different downsampling modules to the model, as shown in Table 1.
[0046] Table 1 Comparison of ablation experimental results of downsampling module
[0047]
[0048] It can be seen from Table 1 that each downsampling strategy can improve the accuracy of the model to varying degrees. When these three downsampling strategies are used at the same time, the model has the highest mAP and the best performance.
[0049] In addition, the present invention also studies the influence of kernel composition on the experimental results. The feature map of large-scale receptive field can be obtained by directly processing with a large convolution kernel or by processing layer by layer with multiple small dilated convolution kernels. As shown in Table 2, when the feature map of receptive field size 29 is obtained after convolution operation, when the large-scale receptive field feature map is obtained by combining three small dilated convolution kernels, the computational complexity of the model is the lowest and the total number of parameters is the least.
[0050] Table 2 Comparison of experimental results of different core compositions
[0051]
[0052] The present invention also verifies the influence of the number of branches of the fusion feature in the kernel selection module on the model. The results are shown in Table 3. The present invention fuses receptive field feature maps of different scales under various settings. By comparing these experimental results, it can be found that the model shows the best performance when using 3×3, 5×5 and 7×7 combinations.
[0053] Table 3 Comparison of experimental results of different convolution kernel settings in the network
[0054]
[0055] 4. Comparative experiment
[0056] In order to verify the superiority of the present invention, an experiment was conducted to compare and analyze the present invention with other remote sensing image target detection models. As shown in Table 4, the present application achieved an mAP of 71.63%, which exceeded the models in the prior art. Compared with the baseline model, the detection accuracy of the AP indicator is significantly improved in target categories such as large vehicles (LV), ships (SH), seaports (HA), and roundabouts (RA). The experimental results show that the model proposed in the present invention can effectively improve the detection accuracy of objects with large scale changes.
[0057] Table 4 Comparison of average accuracy and average accuracy of all categories on the DOTA dataset
[0058]
[0059] In Table 4, prior art 1 is the paper: Azimi SM, Vig E, Bahmanyar R, et al. Towards multiclass object detection in unconstrained remote sensing imagery [C] / / Asian conference on computer vision. Ch-am: Springer International Publishing, 2018: 150-165. Prior art 2 is the paper: Lin TY, Goyal P, Girshick R, et al. Focal Lossf--or Dense Object Detection. IEEE Transactions on Pattern Analysis&MachineIntelligence, 2017, PP(99): 2999-3007. Prior art 3 is the paper: Yang X, Liu Q, Yan J, et al. R3Det: Refined Single Stage Detector with Feature Refinement for Rotating Object. 2019. Prior art 4 is a paper: Ding, Jian, et al. "Learning RoI transformer for oriented object detection in aerial images." Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition. 2019. Prior art 5 is a paper: Zhang G, Lu S, Zhang W. CAD-Net: A context-aware detection network for objects in remote sensing imagery. IEEE Transactions on Geoscience and Remote Sensing, 2019, 57(12): 10015-10024.Prior art 6 is the paper: Pan X, Ren Y, Sheng K, et al. Dynamic refinement network for oriented and densely packed object detection [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020: 11207-11216.
[0060] In addition, in order to qualitatively compare the effects of the baseline method and the present invention, 4 pictures were randomly selected from the data set, tested and visualized. The results show that the present invention is significantly better than the baseline model in detecting seaports and small vehicles (SVs). Compared with the baseline model, the model of the present invention can more accurately locate and identify targets with large scale changes such as seaports, ships, and airplanes, while the baseline model may miss detection or make false detections.
[0061] In summary, according to the remote sensing image target detection model based on improved RetinaNet provided by the present invention, an improved downsampling module is introduced and embedded in the RetinaNet backbone network, and the three downsampling methods are fused to generate downsampled image features for the extracted features, thereby enhancing the model's ability to capture complex details. The convolution kernel selection mechanism of the kernel selection module is used to dynamically select the spatial receptive field, thereby enhancing the model's ability to extract and fuse multi-scale feature information, thereby modeling multi-scale information, and finally obtaining the classification and regression results of the target object. Experimental results show that the model of the present invention has a better average accuracy rate than the traditional RetinaNet target detection model in all categories on the large-scale remote sensing image target detection dataset DOTA, and can detect remote sensing targets more accurately.
[0062] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0063] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A remote sensing image target detection system based on improved RetinaNet, including a RetinaNet backbone network, a feature pyramid and a classification regression subnet, characterized in that: An improved downsampling module is introduced into the RetinaNet backbone network, and the system further comprises a kernel selection module; The backbone network uses the improved downsampling module for downsampling when performing residual learning. The improved downsampling module copies the input image feature P into image feature P1 and image feature P2, where P∈ R H×W×C , R represents a real number, W , H and C Respectively represent the width, height and number of channels of the image feature. The improved downsampling module slices and downsamples the image feature P1. After slicing, four spatially downsampled image features C1, C2, C3 and C4 are obtained. In the process of slicing and downsampling, the image features C1, C2, C3 and C4 are spliced in the channel dimension to obtain a new image feature. After splicing, the number of channels of the image feature P1 is increased from C to 4C. Then, a 1×1 convolution operation with a step size of 1 is used to compress the number of channels of the new image feature to 2C, and the image feature Q1 is obtained. The improved downsampling module uses two branches to process the image feature P2. In one branch, a group convolution GConv with a step size of 1 and a size of 3×3 is used for processing, and then a 3×3 convolution with a step size of 2 is used for downsampling, and a GELU activation function and a normalization layer are used to obtain the image feature Q2; on the other branch, a group convolution GConv with a step size of 1 and a size of 3×3 is used for processing, and maximum pooling and normalization are performed to obtain the image feature Q3; Concatenate image features Q1, Q2, and Q3 in the channel direction, and use a 1×1 convolution layer on the concatenated result to obtain image features ; The kernel selection module dynamically selects a plurality of different convolution kernel fusion features according to the characteristics of the input image; In the process of slicing and downsampling the image feature P1 by the improved downsampling module, the following conditional formula is satisfied: Q1=Conv(CutD(P1)); Q2=GELU(BN(DWConvD(GConv(P2)))); Q3=BN(MaxP(GConv(P2)); Where, Conv, CutD, GELU, BN, DWConvD, GConv, and MaxP represent convolution, slicing, GELU activation function, batch normalization, depthwise convolution, grouped convolution, and maximum pooling operations, respectively; Image features The following conditions are met: ; Among them, Concat represents the operation of connecting features in the channel direction; In the detection task head of the system, for the input image feature K, the kernel selection module uses three dilated convolutions with kernel sizes of 3×3, 5×5, and 7×7 to learn multi-scale spatial information and obtain image features X1, X2, and X3 with three different scale receptive fields, where X1∈ R H×W×C , X2∈ R H×W×C , X3∈ R H×W×C Then, use channel splicing to fuse X1, X2, and X3 to obtain image feature X, and splice the average pooling and maximum pooling results of image feature X in the channel direction. Then, use convolution and Sigmoid functions to obtain independent spatial selection masks, and then use spatial selection masks to weight X1, X2, and X3 respectively to obtain image features F1, F2, and F3 respectively. Finally, add F1, F2, and F3 element by element to obtain fused features with attention, and multiply the fused features and the input image feature K element by element to obtain image features. ; Image features X1, X2, and X3 satisfy the following conditional expressions: X2=DWConv(X1); X3=DWConv(X2); Among them, DWConv represents dilated convolution; Image features The following conditions are met: 。
Citation Information
Patent Citations
Multi-scale single-stage target detection method based on RetinaNet
CN115861772A