Remote sensing image rotation small target detection method combined with super-resolution

By introducing feature enhancement reconstruction module and auxiliary super-resolution network in the remote sensing image rotation small object detection, combined with multi-scale feature association and alignment module, the problem of limited target feature enhancement effect and insufficient information interaction in the remote sensing image rotation small object detection is solved, and high-precision rotation small object detection is achieved.

CN120298877APending Publication Date: 2025-07-11NAT SPACE SCI CENT CAS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510283092.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the detection of small-spinning objects for rotation of remote sensing images, the problems of limited target feature enhancement effect, insufficient super-resolution network performance, and insufficient information interaction between tasks, it is difficult to meet the high-precision detection requirements of small-spinning objects under remote sensing low-resolution observation conditions.

Method used

Using the remote sensing image rotation small object detection method with joint super-resolution, a feature enhancement reconstruction module and an auxiliary super-resolution network are introduced into the backbone rotation small object detection network, and combining multi-scale feature association and alignment modules, the feature characterization capability and information interaction of the target detection network are improved.

Benefits of technology

It significantly improves the detection accuracy and efficiency of rotating small targets under low-resolution observation conditions, and improves the application value of low-resolution remote sensing systems in the fields of geological exploration, disaster monitoring and military reconnaissance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298877A_ABST
    Figure CN120298877A_ABST
Patent Text Reader

Abstract

The invention discloses a super-resolution combined remote sensing image rotation small target detection method, which comprises the following steps: receiving a to-be-detected remote sensing image, and inputting the remote sensing image into a pre-established and trained small target detection network model to obtain a target bounding box and category information; the model comprises a backbone rotation small target detection network, and the network introduces a feature enhancement reconstruction module on the basis of a backbone basic rotation target detection network for improving the feature characterization capability of the backbone rotation small target detection network for weak and small targets; in the training stage, the model further comprises an auxiliary super-resolution network and a multi-scale feature association and alignment module deployed between the trunk rotation small target detection network and the auxiliary super-resolution network. An auxiliary super-resolution network provides richer high-frequency features for a trunk rotation small target detection network branch so as to improve the detection precision of a weak and small target; and the multi-scale feature association and alignment module constructs deep information interaction between the super-resolution and the target detection task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method for detecting small rotating remote sensing images combined with super-resolution. Background Art

[0002] The resolution of space-based optical remote sensing images is usually limited, which significantly affects the ability to present image detail information and the identification accuracy of target features, posing a great challenge to the detection of small and weak targets. To overcome this limitation, the combination technology of super-resolution and downstream vision tasks such as target detection has become an important research field. The super-resolution technology restores the detail information of the high-resolution image from the low-resolution image by using the local features and prior knowledge of the image, improving the clarity and detail expression ability of the image. The target detection algorithm realizes the positioning and classification of the target by extracting features such as the contour shape of the target. The results of these studies can improve the accuracy of target detection in low-resolution images, thereby enhancing the interpretation ability of space-based optical remote sensing images and providing more accurate and reliable support for practical applications such as geological exploration, disaster monitoring, and military reconnaissance.

[0003] At present, some domestic and foreign scholars have carried out research on using image super-resolution technology to improve the performance of target detection. According to the different network architecture designs, it can be mainly divided into two categories, namely the traditional two-stage method and the method based on multi-task learning.

[0004] (1) Two-stage method

[0005] The idea of the two-stage method is to first input the low-resolution image into the super-resolution algorithm to output the super-resolution image, and then perform target detection on the obtained super-resolution image. At this time, the two tasks are carried out independently. For example, Li Hongyan et al. used the Enhanced Deep Super-Resolution Network (EDSR) to first super-resolve the low-resolution image to provide high-resolution data for the target detection network, and then used the improved Faster RCNN algorithm based on the attention mechanism to complete the target detection. The mAP of target detection tested on the public dataset NWPU_VHR-10 increased by 1.8%. Jacob et al. studied the impact of super-resolution in the field of remote sensing images on the performance of target detection. The study found that when the original image resolution is relatively high or the magnification factor executed by the super-resolution network is relatively small, using the super-resolution network to improve the image resolution can usually improve the performance of the target detection network. In other cases, the super-resolution network cannot improve or may even reduce the performance of target detection. The quality of the image generated by the super-resolution network in the two-stage method directly affects the target detection performance, and its distorted or inaccurate details will be amplified in the detection stage, thus interfering with target detection.

[0006] (2) Multitask learning-based methods

[0007] Subsequently, some scholars proposed some multitask learning-based methods by combining super-resolution and downstream vision tasks (classification, detection, etc.). The main idea is to connect the two tasks in series through multitask learning and optimize the entire network by calculating the losses of the two tasks in the reverse direction. Rabbi et al. used the Edge-Enhanced Super-Resolution Generative Adversarial Network (EESRGAN) to generate super-resolution images and input them into the object detection network. The detection loss of the object detection network and the discriminator loss of the generative adversarial network were backpropagated to fine-tune the parameters of the generative network, achieving the co-optimization of the super-resolution network and the object detection network. Yang et al. proposed a mutual feedback learning architecture. This architecture uses the Super-Resolution Generative Adversarial Network (SRGAN) to obtain super-resolution images before executing the object detection network Faster RCNN. A discriminator is used to distinguish the region-of-interest features extracted by the object detection region proposal network (RPN), the region-of-interest features cropped from the super-resolution images, and the region-of-interest features cropped from the high-resolution images, forming a mutual feedback between the super-resolution network and the object detection network to form a closed-loop structure, making SRGAN pay more attention to the regions where the objects may exist. Although such methods aim to achieve the co-optimization of super-resolution and object detection through joint losses, the super-resolution network only uses the detection loss as a low-level guidance and lacks the guidance of high-level features such as object information, so the improvement of object detection performance is limited. In addition, this cascaded method has a high computational complexity, is difficult to optimize, and the super-resolution process is cumbersome, further reducing the detection efficiency.

[0008] Different from the above work, the object detection network with joint super-resolution introduces a super-resolution network branch in the training stage as an auxiliary network to enhance the detection performance of small objects, which does not participate in the inference process and is a more promising architecture. Zhang et al. proposed a multi-modal remote sensing image super-resolution assisted object detection network, SuperYOLO. This network uses YOLOv5 as the baseline model of the object detection network and EDSR as the super-resolution auxiliary network. In the training stage, the two-layer backbone features of YOLOv5 are fused and input into the EDSR network. By jointly optimizing the backbone features of the object detection network with the object detection loss and the super-resolution loss, the detection performance is improved. However, this method lacks information interaction between tasks. Therefore, Liu et al. further proposed an end-to-end super-resolution enhanced real-time rotated object detector (ESRTMDet). This method designs a lightweight embedded feature map super-resolution module (ESRM), which is embedded in the detection model to enhance and amplify the backbone output features, thus improving the detection ability of small objects. In addition, by training a parallel super-resolution network branch (PSRB), the high-resolution image is restored using the backbone features of the detection network. Finally, combined with the feature alignment loss and the feature correlation layer, the PSRB effectively guides the enhancement of the ESRM feature map and promotes deep information interaction between tasks, significantly improving the detection accuracy of small objects. However, this algorithm still has deficiencies. First, the feature enhancement ability of the ESRM module used in the object detection network is weak, and the improvement of the detection performance is limited. Second, the auxiliary parallel super-resolution network branch PSRB uses the EDSR network with simple structure and limited performance, which is difficult to provide detailed feature guidance such as high-frequency textures. Finally, the PSRB only guides a single feature layer output by the ESRM module in the object detection network, and the guiding role of information interaction between tasks is relatively limited.

[0009] In summary, the existing methods for small object detection in remote sensing images with joint super-resolution have problems such as limited object feature enhancement effect, insufficient performance of the super-resolution network, and insufficient information interaction between tasks, making it difficult to meet the demand for high-precision detection of rotated small objects under low-resolution remote sensing observation conditions. Summary of the Invention

[0010] The object of the present invention is to overcome the defects of the prior art and provide a method for detecting small rotating targets in remote sensing images combined with super-resolution, so as to improve the detection accuracy and efficiency of small rotating targets under the condition of low-resolution remote sensing observation, thereby enhancing the application value of low-resolution remote sensing systems in fields such as geological exploration, disaster monitoring, and military reconnaissance.

[0011] In view of this, the present invention proposes a method for detecting small rotating targets in remote sensing images combined with super-resolution, including:

[0012] Receiving the remote sensing image to be detected, inputting it into a pre-established and trained small target detection network model, and obtaining the target bounding box and category information, so as to realize the detection of small targets;

[0013] The small target detection network model includes: a backbone rotating small target detection network, and the backbone rotating small target detection network introduces a feature enhancement and reconstruction module on the basis of the backbone basic rotating target detection network, which is used to improve the feature representation ability of the backbone rotating small target detection network for weak and small targets;

[0014] The small target detection network model in the training stage further includes: an auxiliary super-resolution network, and a multi-scale feature association and alignment module deployed between the backbone rotating small target detection network and the auxiliary super-resolution network; wherein,

[0015] The auxiliary super-resolution network is used to provide richer high-frequency features for the backbone rotating small target detection network branch, so as to improve the detection accuracy of weak and small targets;

[0016] The multi-scale feature association and alignment module is used to construct deep information interaction between the super-resolution and the target detection task.

[0017] Preferably, the input of the backbone rotating small target detection network is a remote sensing image, the output is the position information of all detected targets, that is, the target bounding box, and the target category information is predicted;

[0018] The backbone basic rotating target detection network includes: a basic convolution feature extraction module, a multi-scale feature extraction module, an orientation detection module, and a target detection post-processing module; wherein,

[0019] The input of the basic convolution feature extraction module is the remote sensing image I LR , and basic convolution feature extraction is performed to generate the image basic feature maps {C2, C3, C4};

[0020] The multi-scale feature extraction module performs multi-scale feature extraction on the image basic backbone features to generate the image multi-scale features {P2, P3, P4};

[0021] The orientation detection module performs orientation detection on the multi-scale features {P2, P3, P4} of the image respectively, and outputs the detection classification results {Reg i , Cls i} of all anchor boxes, where the subscript i = 2, 3, 4;

[0022] The target detection post-processing module decodes the 3 detection classification results to obtain the position information and category information of all bounding boxes.

[0023] Preferably, the feature enhancement and reconstruction module includes a feature enhancement part and a resolution improvement part, where

[0024] The feature enhancement part adopts a multi-branch convolution structure, including depthwise separable convolution and dilated convolution, and is used to extract various discriminative semantic information and expand the receptive field;

[0025] The resolution improvement part is implemented using sub-pixel convolution, and is used to magnify the backbone features of the backbone rotated small target detection network and improve the perception ability of the target.

[0026] Preferably, the input of the feature enhancement and reconstruction module is the image basic feature maps {C2, C3, C4}, and the output is three feature upsampling results {C2 up , C3 up , C4 up} and three multi-scale backbone enhancement features {C2 res , C3 res , C4 res}, satisfying the following formula:

[0027] C2 res , C3 res , C4 res = H enhance (C2, C3, C4)

[0028] C2 up , C3 up , C4 up = H upsample (C2 res , C3 res , C4 res )

[0029]

[0030] H DW = DWConv_5(Conv_3(C))

[0031] H Atr1 = AtrConv_3(Conv_3_1(Conv_1_3(Conv_1(C))))

[0032] H Atr2 = AtrConv_3(Conv_1_3(Conv_3_1(Conv_1(C))))

[0033] C ∈ {C2, C3, C4}

[0034] Among them, H enhance (·) represents feature enhancement, and H upsample (·) represents feature upsampling, Conv_1(·) represents a standard 1×1 convolution, Concat(·) represents channel dimension concatenation, Conv_3(·) represents a standard 3×3 convolution, Conv_3_1(·) represents a standard 3×1 convolution, Conv_1_3(·) represents a standard 1×3 convolution, DWConv_5(·) represents a 5×5 depthwise separable convolution, and H DW represents the output of the depthwise separable convolution, and AtrConv_3(·) represents a 3×3 dilated convolution. H A tr1 and H A tr2 respectively represent the outputs of one path of dilated convolution.

[0035] Preferably, the auxiliary super-resolution network adopts a super-resolution network based on large kernel attention distillation, including a feature extraction module, a feature mapping module, and a feature reconstruction module; among them,

[0036] The feature extraction module is used to map the image to a high-dimensional space;

[0037] The feature mapping module is used to improve the model's non-linearity and learn the mapping from the low resolution to the high resolution of the image;

[0038] The feature reconstruction module is used to enhance the resolution of the image.

[0039] Preferably, the input of the feature extraction module is the remote sensing image I LR , and shallow features F0 are extracted through 3 3×3 convolution modules;

[0040] The feature mapping module extracts deep features by stacking large kernel distillation blocks from the shallow features F0, and then performs multi-layer fusion to output the fused feature F map ; The large kernel distillation block includes: feature distillation, feature fusion, feature enhancement, and feature transformation;

[0041] The input of the feature reconstruction module is the fused feature F map and the shallow feature F0, adopts skip connection to enhance residual learning, and obtains the SR image I SR .

[0042] Preferably, the multi-scale feature association and alignment module includes: a multi-scale feature association part and a feature alignment part; where

[0043] The multi-scale feature association part performs multi-scale downsampling on the super-resolution feature F map +F0 to achieve the feature association between the super-resolution feature and the multi-scale backbone enhanced features C2 res 、C3 res 、C4 res By adjusting the channel dimension of the downsampled super-resolution feature through 1×1 convolution, ensuring that it is consistent with the multi-scale backbone enhanced features C2 res 、C3 res 、C4 res in the channel dimension, and obtaining the associated features F2, F3, F4;

[0044] The feature alignment part uses positive sample contrast loss to guide the model training to achieve multi-scale feature alignment.

[0045] Preferably, the positive sample contrast loss L pos_loss is:

[0046]

[0047] where sim_pos represents the cosine similarity of positive samples, used to measure the feature similarity of two features in the target area, sim_neg represents the cosine similarity of negative samples, used to measure the feature similarity of two features in the background area, temperature is the temperature parameter, used to adjust the sensitivity of the similarity, and ε is a small constant, used to avoid the denominator being zero.

[0048] Preferably, the method further includes the training steps of the small target detection network model; including:

[0049] Using the publicly available DOTA remote sensing image target detection dataset to establish a training set;

[0050] Adding an auxiliary super-resolution network outside the backbone rotation small target detection network, and adding a multi-scale feature association and alignment module between the backbone rotation small target detection network and the auxiliary super-resolution network;

[0051] Setting the initialization parameters, training parameters of the network, determining the loss function, and selecting an optimizer;

[0052] Training according to the set training parameters until the training requirements are met, obtaining the trained backbone rotation small target detection network, that is, the small target detection network model.

[0053] Preferably, the training parameters include: the batch size B, the training iteration period T, and the learning rate strategy; the learning rate strategy includes using a cosine annealing learning rate adjustment strategy and adopting a warm-up strategy with linear growth.

[0054] The loss function L Total is:

[0055] L Total = L Det + L SR + λ4L stem_align + λ5L pos_loss

[0056]

[0057] L SR = λ3L2(I SR - I LR , W gt )

[0058]

[0059] where N represents the number of positive samples, l i represents the correct classification result, g i represents the correct bounding box range, W gt represents the positive sample region mask, G(·) represents the normalized Gram matrix, C is the number of channels of the feature, and λ1, λ2, λ3, λ4, λ5 are loss balance parameters, which are default set to {1, 2, 1, 10, 1}.

[0060] Compared with the prior art, the advantages of the present invention are as follows:

[0061] 1. Under the framework of the super-resolution enhanced real-time rotating target detection ESRTMDet, the present invention improves the feature enhancement and reconstruction module in the backbone rotating small target detection network branch to enhance the perception ability of the target detection network for small and weak targets; improves the auxiliary super-resolution network branch to optimize the image super-resolution reconstruction quality and improve the network convergence speed; introduces a multi-scale feature association and alignment module to enhance the target region feature representation ability of the target detection network, thereby improving the accuracy of target detection;

[0062] 2. The present invention proposes a lightweight feature enhancement and reconstruction module combined with dilated convolution, which includes two parts: feature enhancement and resolution improvement. The feature enhancement part designs a multi-branch convolution structure. On the basis of the original depthwise separable convolution, two branches of dilated convolution are added to enhance the local context features of the target; the resolution improvement part is realized by sub-pixel convolution to magnify the backbone features of the target detection network, and finally multi-scale backbone enhanced features are obtained to improve the perception ability of the target detection network for the target;

[0063] 3. The present invention introduces a super-resolution network based on large kernel attention distillation (Large Kernel Distillation Network, LKDN) as an auxiliary super-resolution network branch. Through the design of large kernel distillation blocks, the super-resolution features of the input image are optimized, ensuring the image reconstruction quality while greatly improving the convergence speed of the super-resolution network.

[0064] 4. The present invention proposes a multi-scale feature correlation module and improves the positive sample contrast loss function based on the ground truth as a feature alignment loss to optimize the model performance. First, the feature correlation module performs multi-scale downsampling on the super-resolution features to construct feature correlations with multi-scale backbone enhanced features. Second, the ground truth information of object detection is used to replace the detection results to generate positive sample masks, effectively solving the problems of missed detection and false detection, and fundamentally improving the accuracy of the positive sample region masks. Finally, a positive sample contrast loss combined with cosine similarity is constructed as the optimization target for feature alignment to guide model training, enabling the high-frequency features learned by the super-resolution network to be transferred to the object detection network, enhancing the saliency and representation ability of the target features, and thus improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is a schematic flow diagram of the method for detecting small rotated objects in remote sensing images with joint super-resolution of the present invention;

[0066] Figure 2 is a structural diagram of the network model for detecting small rotated objects in remote sensing images with joint super-resolution of the present invention;

[0067] Figure 3 is a structural diagram of the network model of RTMDet;

[0068] Figure 4 is a structural diagram of the core module in RTMDet;

[0069] Figure 5 is a structural diagram of the target feature enhancement and reconstruction module FERM;

[0070] Figure 6 is a structural diagram of the auxiliary super-resolution network branch model;

[0071] Figure 7 is a structural diagram of the feature mapping module of the auxiliary super-resolution network branch;

[0072] Figure 8 are the model detection results of three remote sensing images. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] The present invention discloses a method for detecting small rotating targets in remote sensing images with joint super-resolution. Based on the ESRTMDet as the basic framework, the feature enhancement and reconstruction module and the auxiliary super-resolution network branch are improved. A multi-scale feature association and alignment module is introduced between the backbone small rotating target detection network branch and the auxiliary super-resolution network branch to construct the guidance of the super-resolution network for the target detection network, and then co-optimization is realized to improve the detection performance of small rotating targets under the low-resolution observation conditions of spaceborne remote sensing.

[0074] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0075] Embodiment 1

[0076] As Figure 1 shown, Embodiment 1 of the present invention proposes a method for detecting small rotating targets in remote sensing images with joint super-resolution. Specifically, it includes: receiving the remote sensing image to be detected, inputting it into the pre-established and trained network model for detecting small rotating targets in remote sensing images with joint super-resolution, and obtaining the target detection result, that is, the target bounding box and class information;

[0077] This embodiment includes four steps: establishing the network model for detecting small rotating targets in remote sensing images with joint super-resolution, constructing the sample data set, training the model, and testing and verifying the model.

[0078] The first step is to establish the model;

[0079] The structure of the network model for detecting small rotating targets in remote sensing images with joint super-resolution is as Figure 2 shown. The model specifically includes two network branches: the backbone small rotating target detection and the auxiliary super-resolution, as well as the multi-scale feature association and alignment module between the two. The backbone small rotating target detection network branch introduces a feature enhancement and reconstruction module on the basis of the backbone basic rotating target detection network, where the backbone basic rotating target detection network consists of four parts: basic convolutional feature extraction, multi-scale feature extraction, directional detection, and post-processing of target detection. The auxiliary super-resolution network branch consists of a feature extraction module, a feature mapping module, and a feature reconstruction module. A multi-scale feature association and alignment module is introduced between the backbone small rotating target detection network branch and the auxiliary super-resolution network branch. First, the feature association module performs multi-scale downsampling on the super-resolution features to construct the feature association with the multi-scale backbone enhanced features; secondly, the true value information of target detection is used to replace the detection result to generate a positive sample mask, effectively solving the problems of missed detection and false detection, and fundamentally improving the accuracy of the positive sample region mask; finally, a positive sample contrast loss combined with cosine similarity is constructed as the optimization target for feature alignment to guide the model training, realizing the transfer of the high-frequency features learned by the super-resolution network to the target detection network, so as to enhance the saliency and representation ability of the target features, thereby improving the detection accuracy.

[0080] In the present invention, it aims to detect the location information (x i , y i , w i , h i , θ i )(i.e., the rotated bounding box) of all targets for the input remote sensing image I and predict the target category information (cls i ).

[0081] The method specifically includes:

[0082] S1, the backbone-based rotated object detection network.

[0083] The backbone-based rotated object detection network of the present invention is the core component of the model and operates independently in the model inference stage. This branch adopts a single-stage real-time rotated object detection network (Real-Time Models for Object Detection, RTMDet) with excellent performance. The overall structure of the RTMDet model is as Figure 3 shown, and it consists of three modules: basic convolutional feature extraction ( Figure 3 Backbone in Figure 3 ), multi-scale feature extraction ( Figure 3 PAFPN in

[0084] 1.1 Basic convolutional feature extraction module

[0085] Perform basic convolutional feature extraction on the input remote sensing image I LR to generate the basic convolutional feature maps {C2, C3, C4} of the image, as Figure 3 shown; specifically including:

[0086] The basic convolutional feature extraction module uses the CSPNeXt network for basic convolutional feature extraction. CSPNeXt contains 1 Stem Layer and 4 Stage Layers. The Stem Layer consists of 3 ConvModules with a size of 3×3. The first 3 Stage Layers are composed of 1 ConvModule and 1 CSPLayer. The 4th Stage Layer adds an SPPFBottleneck between the ConvModule and the CSPLayer.

[0087] The CSPLayer consists of 3 ConvModules, n CSPNeXtBlocks, and 1 ChannelAttention module, as shown in (b) of Figure 4 ; the ConvModule consists of 1 layer of 3×3 Conv2d, BatchNorm, and the SiLU activation function, as shown in (a) of Figure 4 ; the ChannelAttention module consists of 1 layer of AdaptiveAvgPool2d, 1 layer of 1×1 Conv2d, and the Hardsigmoid activation function, as shown in (d) of Figure 4 ; the CSPNeXtBlock is the innovation of RTMDet, introducing a depthwise separable convolution with a large kernel, and introducing a 5×5 depthwise separable convolution after the 3×3 ConvModule to enhance the feature extraction ability of the basic unit, as shown in (c) of Figure 4 ; the SPPFBottleneck consists of 1 layer of 1×1 ConvModule, 3 layers of 5×5 MaxPool2d, Concat, and 1 layer of 1×1 ConvModule connected in series, as shown in (e) of Figure 4 ; Therefore, the present invention uses the CSPNeXt network to perform basic feature extraction on the input remote sensing image I LR to generate the image basic convolutional features {C2, C3, C4},

[0088] C2, C3, C4 = CSPNeXt(I LR )

[0089] where CSPNeXt(·) represents the basic convolutional feature extraction module, and I LR represents the input low-resolution image.

[0090] 1.2 Multi-scale feature extraction.

[0091] Perform multi-scale feature extraction on the image basic convolutional features to generate the image multi-scale feature maps {P2, P3, P4}, as shown in Figure 3 ; specifically including:

[0092] The multi-scale feature extraction module uses the PAFPN network for multi-scale feature extraction. PAFPN adds a bottom-up path on the basis of the Feature Pyramid Network FPN to make up for the lack of low-level feature details in the high-level features of FPN. Each layer of PAFPN fuses information from the upper layer and the lower layer. This cross-level feature fusion helps to improve the detection ability of the target detection network for small targets. The input of PAFPN is the basic convolutional features C2 up , C3 up , C4 up, the outputs are P2, P3, and P4, which can be expressed as:

[0093] P2, P3, P4 = PAFPN(C2 up , C3 up , C4 up )

[0094] Among them, PAFPN(·) represents the multi-scale feature extraction module.

[0095] 1.3 Orientation detection.

[0096] For the input features P at each scale i , orientation detection is performed separately, and finally the detection and classification results of all anchor boxes are output; specifically including:

[0097] The orientation detection module uses the SepBNHead detection head ( Figure 3 Heads in it), sharing the detection head parameters across scales, but using different batch normalization layers, effectively reducing the number of parameters of the detection head while maintaining the accuracy. The input of SepBNHead is P2, P3, P4, and the outputs are three classification results Cls2, Cls3, Cls4 and three bounding box regression results Reg2, Reg3, Reg4, which can be expressed as:

[0098] Cls i , Reg i = SepBNHead(P i ), i = 2, 3, 4

[0099] Among them, SepBNHead(·) represents the orientation detection module.

[0100] 1.4 Post-processing of object detection.

[0101] Decode the results Reg and Cls output by the detection heads of the 3-scale feature layers of the above RTMDet to obtain the position information and category information of all bounding boxes. Filter out the bounding boxes with category confidence scores higher than the set threshold according to the set threshold, and then perform non-maximum suppression processing on the bounding boxes to delete redundant bounding boxes, that is, retain the detection boxes with the highest confidence, and suppress those detection boxes that overlap with it and have lower confidence.

[0102] S2, Feature-Enhanced Reconstruction Module.

[0103] In the above backbone-based rotated object detection network of the present invention, a lightweight feature-enhanced reconstruction module (Feature-Enhanced Reconstruction Module, FERM) combined with dilated convolution is introduced to improve the feature representation ability of the object detection network for small and weak objects. The specific structure of FERM is asFigure 5 As shown, it consists of two parts: feature enhancement and resolution improvement, which are used to enhance and magnify the output features of the backbone of the object detection network, making it easier for the detection head to detect objects. The feature enhancement part of this module is designed as a multi-branch convolution structure, including depthwise separable convolution and dilated convolution, which can extract various discriminative semantic information and expand the receptive field; the resolution improvement part is implemented using sub-pixel convolution to magnify the backbone features of the object detection network, ultimately improving the object perception ability of the object detection network. Specifically, the input of FERM is the backbone features {C2, C3, C4} of the object detection network, and the output is the upsampling results of the three features {C2 up , C3 up , C4 up} and the results of enhancing the three features {C2 res , C3 res , C4 res}, which can be expressed as:

[0104] C2 res , C3 res , C4 res = H enhance (C2, C3, C4)

[0105] C2 up , C3 up , C4 up = H upsample (C2 res , C3 res , C4 res )

[0106]

[0107] H DW = DWConv_5(Conv_3(C))

[0108] H Atr1 = AtrConv_3(Conv_3_1(Conv_1_3(Conv_1(C))))

[0109] H Atr2 = AtrConv_3(Conv_1_3(Conv_3_1(Conv_1(C))))

[0110] C ∈ {C2, C3, C4}

[0111] Among them, H enhance (·) represents feature enhancement, H upsample(·) represents feature upsampling, Conv_1(·) represents a standard 1×1 convolution, Conv_3(·) represents a standard 3×3 convolution, Conv_3_1(·) represents a standard 3×1 convolution, Conv_1_3(·) represents a standard 1×3 convolution, DWConv_5(·) represents a 5×5 depthwise separable convolution, AtrConv_3(·) represents a 3×3 dilated convolution, and Concat(·) represents channel dimension concatenation.

[0112] S3, the auxiliary super-resolution network branch.

[0113] Based on the backbone rotating small target detection network branch, the present invention introduces an auxiliary super-resolution network branch. This branch performs super-resolution processing on the low-resolution image, generates a super-resolution image, and calculates the super-resolution loss to optimize the super-resolution features extracted in the network. The purpose of this auxiliary super-resolution network branch is to provide richer high-frequency features for the backbone rotating small target detection network branch, thereby improving the detection accuracy of small and weak targets. The auxiliary super-resolution network branch is only used in the training stage and does not participate in model inference to ensure the real-time detection ability of the model.

[0114] The auxiliary super-resolution network branch of the present invention adopts a super-resolution network based on large kernel attention distillation, which consists of a Feature Extraction Module (FEM), a Feature Mapping Module (FMM), and a Feature Reconstruction Module (FRM). First, the feature extraction module is used to map the image to a high-dimensional space; secondly, the feature mapping module is used to improve the model's non-linearity and learn the mapping from the low resolution to the high resolution of the image; finally, the feature reconstruction module is used to achieve the improvement of the image resolution. The overall structure of the auxiliary super-resolution network branch is as Figure 6 shown, and the components in the network structure are as Figure 7 shown. Specifically, it includes:

[0115] 3.1 Feature extraction.

[0116] The shallow feature extraction module FEM module proposed by the present invention is the same as the Stem Layer of the backbone network of RTMDet. The shallow features generated from the input LR image are:

[0117] F0 = h ext (I LR )

[0118] where h ext (·) represents the shallow feature extraction module, and F0 represents the shallow features. The structure of FEM consists of 3 3×3 convolution modules.

[0119] 3.2 Feature mapping.

[0120] The feature mapping module FMM proposed by the present invention is composed of multiple feature distillation blocks and multi-layer feature fusion. First, the shallow feature F0 is used to extract deep features through the stacking of large kernel distillation blocks (LKDBs), which can be expressed as:

[0121]

[0122] where represents the k-th LKDB, m is the number of LKDBs used, and m is 5 in the present invention. F k represents the output feature of the k-th LKDB. The LKDB consists of four parts: feature distillation, feature fusion, feature enhancement, and feature transformation. In the feature distillation stage, given the input F in , the feature distillation operation can be described as:

[0123]

[0124] where D i , R i represent the i-th distillation layer and the i-th refinement layer respectively, represent the i-th distillation feature and the i-th refinement feature respectively. In the feature fusion stage, all the distillation features generated in the previous distillation steps are concatenated together, and feature fusion is performed through 1×1 convolution, which can be expressed as:

[0125]

[0126] where H fusion (·) represents the 1×1 convolutional layer, and F fusion is the fused feature. In the feature enhancement stage, an efficient large kernel attention module (LKA) is introduced, which can be expressed as:

[0127] F enhance = H LKA (F fusion )

[0128] where H LKA (·) represents the LKA module, and F enhance is the enhanced feature. In the feature transformation stage, 1×1 convolution is adopted, and a pixel normalization module is added at the same time to ensure the stability of model training, which can be expressed as:

[0129] F trans = Norm pixel (Htrans (F enhance ))

[0130] Among them, H trans (·) represents a 1×1 convolution transformation operation, and F trans is the transformed feature, and Norm pixel represents the pixel normalization operation. Finally, long skip connections are used to enhance the residual learning ability of the model, which can be expressed as:

[0131] F out = F trans + F in

[0132] Among them, F out is the feature output after passing through an LKDB.

[0133] Secondly, after the LKDB is gradually refined, all intermediate features are fused and activated by a 1×1 convolutional layer and a GELU activation function, and a 3×3 BSConv layer is used to smooth the fused features. The process of multi-layer feature fusion can be expressed as:

[0134] F map = H fusion (Concat(F1,..., F k ))

[0135] Among them, H fusion (·) represents the feature fusion module, and F map is the feature output by the feature mapping module after feature fusion.

[0136] 3.3 Feature Reconstruction

[0137] The feature reconstruction module proposed by the present invention uses skip connections to enhance residual learning, and obtains the SR image through image reconstruction:

[0138] I SR = H rec (F map + F0)

[0139] Among them, H rec (·) represents the image reconstruction module, and I SR is the output of the model. The reconstruction process only includes a 3×3 convolution and a PixelShuffle operation. Since the original low-resolution image is downsampled by 2 times after FEM, FRM is used for upsampling by 4 times, and then the channel dimension is reduced to 3 through the final convolution, so as to obtain the final SR image. Finally, the auxiliary super-resolution network branch proposed by the present invention performs a super-resolution task of image ×2. Figure 7 In (a) is LKDB, (b) is BSConv, (c) is LKA, and (d) is RBSB.

[0140] S4, Multi-scale Feature Association and Alignment Module.

[0141] A multi-scale feature association and alignment module is introduced between the backbone rotating small target detection network branch and the auxiliary super-resolution network branch, as Figure 2 shown, to build deep information interaction between the super-resolution and target detection tasks. First, the super-resolution feature F map +F0 is downsampled at multiple scales using the feature association layer to achieve feature association between the super-resolution feature and the multi-scale backbone enhanced features C2 res 、C3 res 、C4 res ; then, the ground truth is used to generate a positive sample region mask, and the multi-scale feature alignment loss is calculated for the two sets of features to guide the overall model training. Feature alignment is achieved by transferring the high-frequency features learned by the super-resolution network to the target detection network, enhancing the saliency and representation ability of the target features, and thus improving the detection performance of rotating small targets under low-resolution observation conditions of spaceborne remote sensing. This module is only used in the training phase and does not participate in model inference.

[0142] The multi-scale feature association and alignment module proposed in the present invention consists of two parts: feature association and feature alignment, and specifically includes:

[0143] 4.1 Feature Association.

[0144] The Feature Association Layer (FAL) consists of a bilinear interpolation function and a 1×1 convolution. First, the super-resolution feature F map +F0 is downsampled at multiple scales through the bilinear interpolation function, and then the channel dimension of the downsampled super-resolution feature is adjusted through a 1×1 convolution to ensure its consistency with the multi-scale backbone enhanced features C2 res 、C3 res 、C4 res in the channel dimension, obtaining the associated features F2, F3, F4, which can be expressed as:

[0145] F2, F3, F4 = Conv_1(Interpolate(F map +F0))

[0146] where Interpolate(·) represents the bilinear interpolation function and Conv_1(·) represents the standard 1×1 convolution.

[0147] 4.2 Feature Alignment.

[0148] The feature alignment part designs a positive sample contrast loss to guide the model training and achieve multi-scale feature alignment. First, a positive sample region mask is generated. The ground truth information of the targets corresponding to the three detection heads of the object detection network is used to generate the positive sample region mask. Then, using this mask and the positive sample contrast loss, only the high-frequency texture features generated by the super-resolution network are passed to the object detection network in the target region, enhancing the saliency and feature representation ability of the target region in the object detection network and suppressing the attention degree of the background region. The positive sample contrast loss L pos_loss can be expressed as:

[0149]

[0150] where sim_pos represents the cosine similarity of positive samples, used to measure the feature similarity of two features in the target region, and sim_neg represents the cosine similarity of negative samples (background), used to measure the feature similarity of two features in the background region. Temperature is a temperature parameter used to adjust the sensitivity of the similarity. ε is a small constant used to avoid the denominator being zero.

[0151] Step 2: Construction of the sample dataset (including training and validation datasets);

[0152] The present invention uses the existing publicly available DOTA remote sensing image object detection dataset as the training sample dataset for the proposed remote sensing image rotation small object detection network model with joint super-resolution. The DOTA dataset is a large-scale remote sensing object detection benchmark dataset, which contains 2,806 remote sensing images with a resolution range of 800×800 to 4,000×4,000, and a total of 188,282 instances are labeled, covering 15 types of typical targets such as airplanes, baseball fields, bridges, athletic fields, small vehicles, large vehicles, ships, tennis courts, basketball courts, oil storage tanks, football fields, roundabouts, ports, swimming pools, and helicopters. The experiment adopts the strategy of joint training of the training set and the validation set and independent evaluation of the test set. The present invention performs slicing sampling of the original image at 1024×1024 pixels (with an overlapping area of 200 pixels) to generate a high-resolution experimental dataset, and uses bicubic interpolation to perform 2-fold downsampling on the high-resolution experimental dataset to generate a degraded image with a resolution of 512×512.

[0153] Based on the above constructed remote sensing dataset, the training set, validation set, and test set are divided according to the classification released by DOTA officially.

[0154] Step 3: Model training;

[0155] S1, initialize the remote sensing image rotation small object detection network model with joint super-resolution, including setting the initialization parameters and training parameters of the network model, determining the loss function, and selecting the optimizer.

[0156] S2. Train the model. Initialize the entire network model according to the above, and train the entire network model. Finally, select the optimal network model.

[0157] In this embodiment, a random initialization method is used to initialize the weights of all convolutional layers of the entire detection network model.

[0158] The training parameters that need to be set for the detection network model in this embodiment mainly include: specifying the paths of the training dataset and the validation dataset, the batch size B, the training iteration period T, the learning rate strategy, and the optimizer strategy. In this embodiment, the batch size B is set to 16; the training iteration period T is set to 36; the learning rate strategy includes using the cosine annealing learning rate adjustment strategy and adopting a linear growth warm-up strategy. The number of warm-up iterations is 1000 times, and the initial warm-up learning rate linearly grows from 1e-5 of the initial learning rate. Starting from the midpoint of the training period, the learning rate is updated in units of the period; the optimizer strategy includes using the AdamW optimizer with weight decay, an initial learning rate of 0.0025, and a weight decay of 0.5 to perform weight decay on all parameters to avoid overfitting.

[0159] The present invention designs a loss function to iteratively optimize the proposed remote sensing image rotation small target detection network model for joint super-resolution. The total loss function of the joint super-resolution remote sensing image small target detection model includes the object detection loss L Det , the image super-resolution loss L SR , the backbone feature alignment loss L stem_align , and the multi-scale feature alignment loss L pos_loss . The total loss function is expressed as:

[0160] L Total = L Det + L SR + λ4L stem_align + λ5L pos_loss

[0161]

[0162] L SR = λ3L2(I SR - I LR , W gt )

[0163]

[0164] Among them, N represents the number of positive samples, l i represents the correct classification result, g i represents the correct bounding box range, W gtLet \(M\) denote the positive sample region mask, \(G(\cdot)\) denote the normalized Gram matrix, \(C\) be the number of channels of the features, and \(\lambda_1,\lambda_2,\lambda_3,\lambda_4,\lambda_5\) be the loss balance parameters, which are default set to \(\{1,2,1,10,1\}\).

[0165] For the detection loss, the classification loss function uses Quality Focal Loss, and the regression loss function uses Rotated IOU Loss. The super-resolution loss uses the L2 loss function. The backbone feature alignment loss represents calculating the similarity between the shallow feature \(F_0\) output by the FEM module and the feature layer \(C_0\) output by the backbone network of the object detection network using the normalized Gram matrix, and using the Euclidean distance to measure the difference in the global structural relationship between different input features. The multi-scale feature alignment loss is as described in Section 4.2.

[0166] Step 4: Model test and verification;

[0167] For the above-mentioned jointly trained remote sensing image rotated small object detection network model, the test set in the above DOTA dataset is used to test the performance of the model. A set of validation data in the test set consists of 1 input image and the corresponding target annotation file. Inputting the input image in a set of test data into the above-trained model can predict and output the object detection result.

[0168] The evaluation metrics of current deep learning-based remote sensing image object detection algorithms mainly include mean average precision (mAP) and average precision for small objects (APs). Among them, mAP comprehensively quantifies the localization accuracy and classification confidence of the model for rotated objects by calculating the average of the object detection accuracies of multiple categories within the intersection over union (IoU) threshold range; APs defines objects with a pixel area less than \(32\times32\) as small objects according to the COCO standard, and evaluates the feature capture ability of the model for low-resolution and low-signal-to-noise small objects by statistically calculating the average precision of small objects within the IoU threshold interval. Based on this, this test and verification constructs an evaluation system using the two metrics of mAP and APs: the former is used to quantify the comprehensive detection performance of the model for all category objects, and the latter focuses on evaluating the detection performance for small objects. The combination of the two can comprehensively evaluate the technical advantages of the model in the rotated object detection task.

[0169] On the above DOTA test set, a performance comparison and verification experiment was conducted on the remote sensing image rotation small target detection method (SR-RTODet) with joint super-resolution proposed by the present invention and other advanced cascaded super-resolution target detection methods. The specific comparison methods are shown in the first column of Table 1. The experimental performance comparison and verification results are shown in Table 1, indicating that the small target detection APs of the method proposed by the present invention have been significantly improved, with a 3.4% improvement compared to the sub-optimal method; the mAP for all category targets has been improved by 1.67% compared to the sub-optimal method. The optimal results in Table 1 are highlighted in bold font, and the sub-optimal results are identified with italics and underlines.

[0170] Table 1

[0171]

[0172] In addition, for three typical targets of small vehicles, ships, and airplanes, the model target detection results of three remote sensing images of relevant scenes were selected for visual display, specifically as Figure 8 shown. Figure 8 The left column is the labeled image of the target true label, and the target is identified by bounding boxes of different colors and the target category; the right column is the model target detection result, and the target is identified by bounding boxes of different colors, the predicted target category, and the predicted confidence. Figure 8 The visualization results in

[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A remote sensing image rotation small target detection method combined with super-resolution, comprising: Receiving the remote sensing image to be detected, inputting it into a pre-established and trained small target detection network model, and obtaining the target bounding box and category information, so as to achieve small target detection; The small target detection network model includes: a backbone rotation small target detection network, and the backbone rotation small target detection network introduces a feature enhancement and reconstruction module on the basis of the backbone basic rotation target detection network, which is used to improve the feature representation ability of the backbone rotation small target detection network for weak and small targets; In the training stage, the small target detection network model also includes: an auxiliary super-resolution network, and a multi-scale feature association and alignment module deployed between the backbone rotation small target detection network and the auxiliary super-resolution network; wherein, The auxiliary super-resolution network is used to provide richer high-frequency features for the backbone rotation small target detection network branch to improve the detection accuracy of weak and small targets; The multi-scale feature association and alignment module is used to construct deep information interaction between the super-resolution and the target detection task.

2. The method for detecting small rotating targets in remote sensing images with joint super-resolution according to claim 1, characterized in that, The input of the backbone rotation small target detection network is the remote sensing image, and the output is the position information of all detected targets, that is, the target bounding box, and the target category information is predicted; The backbone basic rotation target detection network includes: a basic convolution feature extraction module, a multi-scale feature extraction module, an orientation detection module, and a target detection post-processing module; wherein, The input of the basic convolutional feature extraction module is the remote sensing image I LR , and basic convolutional feature extraction is performed to generate image basic feature maps {C2, C3, C4}; The multi-scale feature extraction module performs multi-scale feature extraction on the basic backbone features of the image to generate image multi-scale features {P2, P3, P4}; The orientation detection module performs orientation detection on the multi-scale features {P2, P3, P4} of the image respectively, and outputs the detection and classification results {Reg i , Cls i} of all anchor boxes, where the subscript i = 2, 3, 4; The target detection post-processing module decodes the 3 detection classification results to obtain the position information and category information of all bounding boxes.

3. The method for detecting small rotated remote sensing targets with joint super-resolution according to claim 2, characterized in that, The feature enhancement and reconstruction module includes a feature enhancement part and a resolution improvement part, wherein, The feature enhancement part adopts a multi-branch convolution structure, including depthwise separable convolution and dilated convolution, which is used to extract various discriminative semantic information and expand the receptive field; The resolution improvement part is implemented using sub-pixel convolution, which is used to magnify the backbone features of the backbone rotation small target detection network and improve the perception ability of the target.

4. The method for detecting small rotating targets in remote sensing images with joint super-resolution according to claim 3, wherein, The input of the feature enhancement and reconstruction module is the image base feature maps {C2, C3, C4}, and the outputs are three feature upsampling results {C2 up , C3 up , C4 up} and three multi-scale backbone enhancement features {C2 res , C3 res , C4 res}, satisfying the following formula: C2 res , C3 res , C4 res = H enhance (C2, C3, C4) C2 up , C3 up , C4 up = H upsample (C2 res , C3 res , C4 res ) H DW = DWConv_5(Conv_3(C)) H Atr1 = AtrConv_3(Conv_3_1(Conv_1_3(Conv_1(C)))) H Atr2 = AtrConv_3(Conv_1_3(Conv_3_1(Conv_1(C)))) C ∈ {C2, C3, C4} Among them, H enhance (·) represents feature enhancement, H upsample (·) represents feature upsampling, Conv_1(·) represents a 1×1 standard convolution, Concat(·) represents channel dimension concatenation, Conv_3(·) represents a 3×3 standard convolution, Conv_3_1(·) represents a 3×1 standard convolution, Conv_1_3(·) represents a 1×3 standard convolution, DWConv_5(·) represents a 5×5 depthwise separable convolution, H DW represents the output of the depthwise separable convolution, AtrConv_3(·) represents a 3×3 dilated convolution, H Atr1 and H Atr2 respectively represent the outputs of one path of dilated convolution.

5. The method for detecting small rotating targets in remote sensing images with joint super-resolution according to claim 4, characterized in that, The auxiliary super-resolution network adopts a super-resolution network based on large kernel attention distillation, including a feature extraction module, a feature mapping module, and a feature reconstruction module; wherein, The feature extraction module is used to map the image to a high-dimensional space; The feature mapping module is used to improve the non-linearity of the model and learn the mapping from the low resolution to the high resolution of the image; The feature reconstruction module is used to improve the resolution of the image.

6. The method for detecting small rotating targets in remote sensing images with joint super-resolution according to claim 5, wherein The input of the feature extraction module is the remote sensing image I LR , and the shallow feature F0 is extracted through three 3×3 convolution modules; The feature mapping module extracts deep features from the shallow features F0 through the stacking of large kernel distillation blocks, and then performs multi-layer fusion to output the fused feature F map ; The large kernel distillation block includes: feature distillation, feature fusion, feature enhancement, and feature transformation; The input of the feature reconstruction module is the fused feature F map and the shallow feature F0. Skip connections are used to enhance residual learning, and the SR image I is obtained through image reconstruction SR .

7. The method for detecting small rotating targets in remote sensing images with joint super-resolution according to claim 6, wherein The multi-scale feature association and alignment module includes: a multi-scale feature association part and a feature alignment part; wherein, The multi-scale feature correlation part performs multi-scale downsampling on the super-resolution feature F map +F0 to achieve the feature correlation between the super-resolution feature and the multi-scale backbone enhanced features C2 res , C3 res , C4 res . The channel dimension of the downsampled super-resolution feature is adjusted through 1×1 convolution to ensure its consistency with the multi-scale backbone enhanced features C2 res , C3 res , C4 res in the channel dimension, and the associated features F2, F3, and F4 are obtained; The feature alignment part uses positive sample contrast loss to guide the model training to achieve multi-scale feature alignment.

8. The method for detecting small rotating targets in remote sensing images with joint super-resolution according to claim 7, wherein The positive sample contrastive loss L pos_loss is as follows: Among them, sim_pos represents the cosine similarity of positive samples, which is used to measure the feature similarity of two features in the target region, sim_neg represents the cosine similarity of negative samples, which is used to measure the feature similarity of two features in the background region, temperature is a temperature parameter used to adjust the sensitivity of similarity, and ε is a small constant used to avoid a zero denominator.

9. The method for detecting small rotating targets in remote sensing images with joint super-resolution according to claim 8, wherein The method further includes the training steps of the small target detection network model, including: Using the publicly available DOTA remote sensing image target detection dataset to establish a training set; Adding an auxiliary super-resolution network outside the backbone rotation small target detection network, and adding a multi-scale feature association and alignment module between the backbone rotation small target detection network and the auxiliary super-resolution network; Setting the initialization parameters, training parameters of the network, determining the loss function, and selecting an optimizer; Training according to the set training parameters until the training requirements are met, and obtaining the trained backbone rotation small target detection network, that is, the small target detection network model.

10. The method for detecting small rotating remote sensing targets with combined super-resolution according to claim 9, wherein The training parameters include: batch data volume B, training iteration period T, and learning rate strategy; the learning rate strategy includes using a cosine annealing learning rate adjustment strategy and adopting a linear growth warm-up strategy. The loss function L Total is as follows: L Total = L Det + L SR + λ4L stem_align + λ5L pos_loss L SR = λ3L2(I SR - I LR , W gt ) Among them, N represents the number of positive samples, l i represents the correct classification result, g i represents the correct bounding box range, W gt represents the positive sample region mask, G(·) represents the normalized Gram matrix, C is the number of channels of the feature, and λ1, λ2, λ3, λ4, λ5 are loss balance parameters, with the default settings being {1, 2, 1, 10, 1}.

Citation Information

Cited By

  • A target detection method, an electronic device, and a computer-readable storage medium

    CN122574374A