Remote sensing image target detection method and system based on global local context sensing network

By applying a global local context-aware network in remote sensing images, the shortcomings of traditional object detection algorithms in multi-scale, complex background and extreme aspect ratio target detection are solved, and more accurate object detection effects are achieved.

CN119992043APending Publication Date: 2025-05-13Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411881888.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When traditional object detection algorithms are applied in remote sensing images, they face the problems of severe changes in multi-scale target sizes, complex backgrounds, and poor target detection results in extreme aspect ratios.

Method used

A remote sensing image object detection method based on global local context-aware network is proposed. Space-channel cross-fusion is performed through GLCCFNet, TSFFPN performs multi-scale feature fusion and background filtering, and MSSHead enlarges the receptive field to detect extreme aspect ratio targets.

Benefits of technology

The accuracy of multi-scale target and extreme aspect ratio target detection in complex contexts is achieved, exceeding the effects of existing feature extraction and feature fusion networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992043A_ABST
    Figure CN119992043A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing images, in particular to a remote sensing image target detection method and system based on a global and local context sensing network, and the method comprises the steps: firstly carrying out the cross fusion of global and local context information of an input remote sensing image through employing a space-channel cross fusion strategy through employing a GLCCFNet, and carrying out the cross fusion of the global and local context information in space and channel dimensions; the GLCCFNet fully considers the complementary characteristics of global and local context information in a channel and a space, so that efficient fusion of the global and local context information is realized; then, a TSFFPN is proposed to fuse the multi-scale feature maps and filter out complex background information, and the TSFFPN not only can reduce information loss of non-adjacent feature layers, but also can automatically select space attention weights required by feature layers with different scales, so that the complex background information is filtered out; and finally, the MSSHead is used for obtaining the features of the extreme length-width ratio target by increasing the receptive field of the fused feature map, and the features of the extreme length-width ratio target are concerned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image technology, and in particular to a remote sensing image target detection method and system based on a global local context perception network. Background Art

[0002] In recent years, thanks to the development of convolutional neural networks and some large-scale high-quality natural image datasets. General target detection methods based on convolutional neural networks have made great progress in detection accuracy and robustness. Therefore, many studies apply general target detection algorithms to remote sensing images to complete tasks such as emergency rescue, landslide detection, traffic management, and wildfire monitoring. However, there are differences in the sensor shooting angle between remote sensing images and natural images. There are some challenges in directly applying these general target detection algorithms to remote sensing images: (1) The challenge of multi-scale targets. For example, Figure 1 As shown in (a), in the same remote sensing image, different categories of targets (golf courses and bridges) have natural size differences, and the same category of targets also have size differences, resulting in dramatic size changes in the remote sensing image. (2) The challenge of complex background. Figure 1 As shown in (b), the target of interest in the remote sensing image occupies a very small area in the image, and most of it is complex background information. Complex background information will weaken the target features, resulting in missed detection and false detection. (3) Challenges of targets with extreme aspect ratios. Figure 1 As shown in (c), the length of the bridge in the remote sensing image is much larger than its width. Since the convolution kernel is regular, it cannot focus the receptive field on objects with extreme aspect ratios.

[0003] Since general target detection algorithms almost all have the same paradigm: a backbone network (backbone) is used for feature extraction, a feature fusion network (neck) is used to fuse the multi-scale feature maps output by the backbone network, and a detection head (head) is used for classification and regression.

[0004] Although the global context has rich semantic information and can improve the detection effect of the model. However, obtaining global context based on stacked or parallel convolutional layers and self-attention mechanisms will lead to the loss of some local context features, such as text, edges and shapes, which are key information for accurately detecting small objects. Therefore, some studies fuse global and local context information to achieve the complementarity of local details and global semantics. However, they ignore the different characteristics of global and local context information, which affects the effect of feature extraction. Therefore, the Pyramid Feature Attention Network uses channel attention and spatial attention to extract rich features based on the different characteristics of global and local context information. However, it ignores the complementary characteristics of global and local context information in channels and spaces.

[0005] In the feature fusion part, the commonly used feature fusion methods will lead to the loss and degradation of information in non-adjacent feature layers, weakening the effect of feature fusion. In addition, there is a huge semantic gap between non-adjacent feature maps, and these methods ignore the semantic gap between non-adjacent feature layers. Therefore, some studies solve these problems by gradually fusing from the bottom layer to the top layer and balancing the multi-layer semantic information. However, they ignore that multi-scale feature fusion will introduce a lot of background noise. Although adding an attention mechanism to multi-scale feature fusion can weaken the background noise and enhance the feature information of the target. However, feature layers of different scales require different attention weights, and their methods cannot select different attention weights for feature layers of different scales.

[0006] In the detection head part, good results are achieved by increasing the receptive field to obtain the features of targets with extreme aspect ratios. However, since the convolution kernel is regular, it cannot adapt to the shape of the target, which hinders the further improvement of the detection effect. Although Dyhead introduces deformable convolution in the detection head part to adapt to targets with extreme aspect ratios and achieves good results, the limited feature extraction capability of Dyhead also hinders the further improvement of the detection effect. Summary of the invention

[0007] The present invention aims to solve the problem that the traditional target detection algorithm is directly applied to remote sensing images. Due to the presence of a large number of complex backgrounds, drastic changes in target size and some targets with extreme aspect ratios in remote sensing images, the detection effect is not good. A remote sensing image target detection method and system based on a global local context perception network is proposed, which can accurately detect multi-scale targets and targets with extreme aspect ratios in remote sensing images under complex backgrounds.

[0008] To achieve the above purpose, the technical solution adopted is:

[0009] A remote sensing image target detection method based on a global local context perception network, comprising:

[0010] Firstly, GLCCFNet is used to cross-fuse the global and local context information in the spatial and channel dimensions of the input remote sensing image using a spatial-channel cross-fusion strategy to obtain multi-scale feature maps.

[0011] Then TSFFPN is proposed to fuse multi-scale feature maps and filter out complex background information;

[0012] Finally, MSSHead is used to obtain the features of targets with extreme aspect ratios by increasing the receptive field of the fused feature map, and focuses on the features of targets with extreme aspect ratios.

[0013] According to the remote sensing image target detection method based on the global local context-aware network of the present invention, further, GLCCFNet includes three parts: Stage n, MSDC Block and MSDC Module. The remote sensing image outputs a multi-scale feature map after Stage n. Each Stage includes a CSPBlock structure. The input features are divided into two parts after downsampling and convolution layers. One part is directly output, and the other part passes through multiple MSDC Blocks. The output results of the two parts are spliced ​​in the channel dimension and output the final result through a convolution layer; each MSDC Block includes a Conv and MSDC Module. The output results of the two parts are spliced ​​in the channel dimension and output the final result through a convolution layer.

[0014] According to the remote sensing image target detection method based on the global local context perception network of the present invention, further, GLCCFNet includes four stages, and the input and output of each stage are respectively expressed as and The structure of each stage is as follows: First, the input feature F n Output feature map after downsampling and convolution layer Then, X n Divide into two parts equally according to the channel direction X (1) n and X (2) n , X (1) n Output feature map directly without processing X (2) n After multiple MSDC blocks are connected in series, the feature map is output. Finally, X (1) n+1 and X (2) n+1 Concatenate in the channel dimension and fuse through a 1x1 convolution layer to output the feature map F n+1 .

[0015] According to the remote sensing image target detection method based on the global local context perception network of the present invention, further, the structure of the MSDC Block is: the input feature X (2) n Divide into two parts according to the channel direction and X (2) (1,n) Get global context information G through MSDC Module (1)n , X (2) (2,n) Model the local context information L through a 3x3 convolutional layer (1) n ; Then the global and local context information are fused through the space-channel cross fusion strategy.

[0016] According to the remote sensing image target detection method based on the global local context perception network of the present invention, further, the global and local context information are fused through the space-channel cross fusion strategy, which includes:

[0017] In the first branch, the local context information L (1) n Perform global maximum pooling to obtain feature map Z l ; then Z l The local context channel weight S is obtained through a Sigmoid function l , and finally S l With the global context information G (1) n Multiply point by point to get G (1) n+1 ;

[0018] In the second branch, the global context information G (1) n The spatial descriptors are obtained by global average pooling and global maximum pooling respectively, and the feature map Z is obtained by feature concatenation and convolution. g ; then Z g The global context space weight S is obtained through a Sigmoid function g , and finally S g With local context information L (1) n Multiply point by point to get L (1) n+1 ;

[0019] Finally, the local context L (1) n+1 and the global context G (1) n+1 Concatenate in the channel dimension and fuse through a 1x1 convolution to obtain the global-local feature map X (2) n+1 .

[0020] According to the remote sensing image target detection method based on the global local context perception network of the present invention, further, the structure of the MSDC Module is: first, input feature X (2) (1,n)After a convolutional layer, it is evenly divided into five parts G according to the channel direction. (1) (1,n) , G (1) (2,n) , G (1) (3,n) , G (1) (4,n) , G (1) (5,n) Then, they are respectively subjected to dilated convolutions with dilation rates of 1, 2, 4, 8, and 16, and the outputs of dilated convolutions with different dilation rates are stacked in sequence; finally, all the output results are concatenated in the channel dimension and fused through a 1x1 convolution to output the final result G (1) n .

[0021] According to the remote sensing image target detection method based on the global local context perception network of the present invention, further, the core of TSFFPN is the TSF module, and the TSF module includes four steps of merging, refining, selecting and filtering, specifically:

[0022] ① Merge: Select three feature maps of different sizes F3∈R output by GLCCFNet C×2H×2W , F4∈R C×H×W , As input; first, F3 and F5 are downsampled and upsampled to obtain feature maps and F5 (1) ∈R C×H×W ; Then, they are concatenated in the channel dimension and fused through a 1x1 convolutional layer, and finally the fused feature map U∈R is output C×H×W ;

[0023] ② Refinement: The multi-scale feature map U is refined through an MSDC Block to obtain the feature map Ω;

[0024] ③ Selection: First, the feature map Ω is subjected to channel-level average pooling and maximum pooling to extract the spatial feature descriptor SA avg ∈R 1×W×H and SA max ∈R 1×W×H ; Then, SA avg and SA max After splicing, the convolution layer is used to convert the spatial pooling features of the two channels into the spatial attention feature map SA of the three channels; finally, the Sigmoid activation function is applied to each spatial attention feature map SA i , i∈1,2,3, and obtain a multi-scale spatial mask

[0025] ④ Filtering: Multi-scale feature maps F3, F4, F5 are combined with the corresponding spatial masks Weighted, the weighted multi-scale feature maps are fused, and the spatial attention feature S∈R is obtained through the convolution layer C×W×H ; Then, the feature map F4 is weighted using the spatial attention feature S to obtain the feature map F4 (2) ; Finally, F4 (2) Upsample and downsample respectively to get feature map F3 (2) and F5 (2) ; F3 (2) 、F4 (2) , F5 (2) After feature fusion, output F i .

[0026] According to the remote sensing image target detection method based on the global local context perception network of the present invention, further, MSSHead adopts a decoupled detection head as the basic architecture, and the core part is a multi-scale multi-shape module. The structure of the multi-scale multi-shape module is as follows: first, the multi-scale feature F output from TSFFPN is obtained through strip pyramid attention. i Then, the receptive field is adjusted to the extreme shape objects through deformable spatial attention, and finally, task attention is used to focus on the bounding box regression and classification tasks.

[0027] Furthermore, the present invention also provides a remote sensing image target detection system based on a global local context awareness network, comprising:

[0028] The feature extraction module is used to use GLCCFNet to cross-fuse the global and local context information in the spatial and channel dimensions of the input remote sensing image using a spatial-channel cross-fusion strategy to obtain a multi-scale feature map;

[0029] Feature fusion module, used to propose TSFFPN to fuse multi-scale feature maps and filter out complex background information;

[0030] The detection module is used to use MSSHead to obtain the features of targets with extreme aspect ratios by increasing the receptive field of the fused feature map, and to focus on the features of targets with extreme aspect ratios.

[0031] The beneficial effects achieved by adopting the above technical solution are:

[0032] In order to detect multi-scale targets and extreme aspect ratio targets in complex backgrounds of remote sensing images, this paper innovatively proposes a global local context-aware network (BSPDet). (1) First, GLCCFNet is proposed to extract multi-scale features. It fully considers the complementary characteristics of global and local context information in space and channels, and realizes the efficient fusion of global and local context information. Its effect exceeds that of SOTA feature extraction networks PKINet, MobileNetv4, and StarNet. (2) Then TSFFPN is proposed to fuse multi-scale feature maps. It can not only eliminate the semantic gap of non-adjacent feature layers and reduce the loss and degradation of non-adjacent feature layer information, but also give different context weights to each feature layer, thereby filtering complex background information. Its effect exceeds that of SOTA feature fusion networks HSFPN, BiFPN, and AFPN. (3) Finally, MSSHead is designed to extract features of extreme aspect ratio targets. It can not only obtain rich context information from multi-scale feature layers, but also focus the receptive field on extreme aspect ratio targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings of the embodiments of the present invention, wherein the drawings are only used to illustrate some embodiments of the present invention, but not to limit all embodiments of the present invention thereto.

[0034] Figure 1 It is a challenge for target detection in remote sensing images;

[0035] Figure 2 It is a framework diagram of a remote sensing image target detection method based on a global local context awareness network according to an embodiment of the present invention;

[0036] Figure 3 is a structural diagram of GLCCFNet according to an embodiment of the present invention;

[0037] Figure 4 is a structural diagram of a TSF module according to an embodiment of the present invention;

[0038] Figure 5 is a structural diagram of MSSHead in an embodiment of the present invention;

[0039] Figure 6 Detection effects of BSPDet of the embodiment of the present invention on RSOD, DIOR, HRRSD, and SSDD datasets; (a) small-sized targets, (b) large-sized targets, (c) multi-scale targets, d) targets in complex backgrounds, and (e) targets with extreme aspect ratios;

[0040] Figure 7: This is a visualization analysis of the GLCCFNet and TSFFPN feature maps of an embodiment of the present invention; (a) original image, (b) feature map output by the original baseline, (c) feature map extracted by GLCCFNet, (d) feature map extracted after GLCCFNet and TSFFPN, (e) final detection result;

[0041] Figure 8 It is a visual comparative analysis of the detection results of BSPDet and other algorithms according to the embodiment of the present invention. DETAILED DESCRIPTION

[0042] The following will be combined with the drawings of specific embodiments of the present invention to clearly and completely describe the exemplary scheme of the embodiment of the present invention. Unless otherwise defined, the technical terms or scientific terms used in the present invention should be the common meanings understood by people with ordinary skills in the field.

[0043] like Figure 2 As shown, this embodiment discloses a remote sensing image target detection method based on a global local context perception network, which includes the following contents:

[0044] Step S101, using the Global-local contexts crossing fusion network (GLCCFNet) to cross-fuse the input remote sensing image with a space-channel cross-fusion strategy to cross-fuse the global and local context information in the space and channel dimensions to obtain a multi-scale feature map. GLCCFNet fully considers the complementary characteristics of global and local context information in space and channels. It not only guides the local context to correctly understand the semantic information through the spatial weight of the global context, but also constrains the global context through the channel weight of the local context to supplement the detail information.

[0045] like Figure 3 As shown in the figure, in order to extract multi-scale features, GLCCFNet is proposed, which mainly consists of three parts: Stage n, MSDCBlock and MSDC Module. Each Stage contains a CSPBlock structure. The input features are divided into two parts after downsampling and convolution layers. One part is directly output, and the other part passes through multiple MSDCBlocks. The output results of these two parts are spliced ​​in the channel dimension and passed through a convolution layer to obtain the final output result. Each MSDC Block contains a Conv and MSDC Module. The output results of these two parts are spliced ​​in the channel dimension and passed through a convolution layer to obtain the final output result. The structural details of GLCCFNet are as follows:

[0046] (1) Stage n: GLCCFNet has four stages in total. The input and output of each stage are represented as and The structure of Stage n is represented in Figure 3 (b), first input feature F n Output feature map after downsampling and convolution

[0047] X n =Conv 3×3 (DownSample(F n ))

[0048] Then, X n Divide into two parts equally according to the channel direction X (1) n and X (2) n , X (1) n Output feature map directly without any processing X (2) n After multiple MSDC blocks are connected in series, the feature map is output. Finally, X (1) n+1 and X (2) n+1 Concatenate in the channel dimension and fuse through a 1x1 convolution layer to output the feature map F n+1 :

[0049] F n+1 =Conv 1×1 (Concat(X (1) n+1 ,X (2) n+1 ))

[0050] (2)MSDC Block: The structure of MSDC Block is as follows Figure 3 As shown in (c), its main function is to fuse global-local context information and obtain the features of multi-scale targets. The specific structure is as follows:

[0051] Input features X (2) n Divide it into two parts evenly according to the channel direction to get the feature map and X (2) (1,n) Get global context information G through MSDC Module (1) n, mainly used to obtain the features of large-size targets. (2) (2,n) Model the local context information L through a 3x3 convolutional layer (1) n , which is mainly used to obtain the features of small-sized targets. Then, the global and local context information are efficiently fused through the spatial-channel cross-fusion strategy, as follows:

[0052] In the first branch, the global context information G (1) n The spatial dimension of the image contains highly abstract semantics, so there is no need to filter spatial information. However, in the channel dimension, it lacks detailed information and has a large number of redundant features. Although, using channel attention directly after the global context can filter out redundant features and emphasize key features. However, it does not make full use of the rich detailed information of the local context. Therefore, the channel weights of the local context are used to constrain the global context to supplement the detailed information and select the best semantic information. The specific process is as follows:

[0053] First, for L (1) n Perform global maximum pooling to obtain feature map Z l :

[0054]

[0055] Then, Z l The local context channel weight S is obtained through a Sigmoid function l , and finally S l With G (1) n Multiply point by point to get more detailed global context information G (1) n+1 :

[0056] G (1) n+1 =S l ·G (1) n

[0057] In the second branch, the local context information L (1) n There is almost no semantic difference between different channels, so there is no need to filter channel information. However, in the spatial dimension, it lacks semantic information, is easy to misunderstand the target, and contains a lot of background noise, which can easily interfere with the detection effect. Although, using spatial attention directly after local context can filter background noise and highlight the characteristics of key targets. However, it does not make full use of the rich semantic information of the global context. Therefore, through the global context information G (1)n The spatial weights of are used to guide the local context to correctly understand the semantic information and reduce the interference of complex background. The specific process is as follows:

[0058] First, the spatial descriptor of global semantic information is obtained through global average pooling and global maximum pooling, and the feature map is obtained through feature concatenation and convolution:

[0059] Z g =Conv 3×3 (Concat(P avg (G (1) n ),P max (G (1) n )))

[0060] Among them, P avg (·) is the global average pooling, P max (·) is the global max pooling.

[0061] Then, Z g The global context space weight S is obtained through a Sigmoid function g ; Finally, S g With local context information L (1) n Multiply point by point to get more semantically rich local context information L (1) n+1 :

[0062] L (1) n+1 =S g ·L (1) n

[0063] Finally, the local context L (1) n+1 and the global context G (1) n+1 Concatenate in the channel dimension and fuse through a 1x1 convolution to obtain the global-local feature map X (2) n+1 :

[0064] X (2) n+1 =Conv 1×1 (Concat(G (1) n+1 ,L (1) n+1 ))

[0065] (3) MSDC Module: In order to obtain a wider range of global context information and ensure that the amount of calculation and parameters is small, the MSDC Module is proposed, which fully utilizes the advantages of stacked convolutional layers and dilated convolution in obtaining a large range of receptive fields. The detailed structure of the MSDC Module is shown in Figure 2. Figure 3 As shown in (d), first, input feature X (2) (1,n) After a convolutional layer, it is evenly divided into five parts G according to the channel direction. (1) (1,n) , G (1) (2,n) , G (1) (3,n) , G (1) (4,n) , G (1) (5,n) Then, they are subjected to dilated convolutions with dilation rates of 1, 2, 4, 8, and 16 respectively, and the outputs of dilated convolutions with different dilation rates are stacked in sequence. The formula can be expressed as:

[0066] Z (1) n =Conv 1 3*3 (G (1) (1,n) )

[0067] Z (2) n =Conv 2 3*3 (Z (1) n +G (1) (2,n) )

[0068] Z (3) n =Conv 4 3*3 (Z (2) n +G (1) (3,n) )

[0069] Z (4) n =Conv 8 3*3 (Z (3) n +G (1) (4,n) )

[0070] Z (5) n =Conv 16 3*3(Z (4) n +G (1) (5,n) )

[0071] Among them, Z (1) n ,Z (2) n ,Z (3) n ,Z (4) n ,Z (5) n It is the feature map output by stacking different dilated convolutional layers.

[0072] Finally, all the output results are concatenated in the channel dimension and fused through a 1x1 convolution to output the final result G (1) n :

[0073] G (1) n =Conv(Concat(Z (1) n ,Z (2) n ,Z (3) n ,Z (4) n ,Z (5) n ))

[0074] Step S102: Propose a Triplet Select Focus Feature Pyramid Network (TSFFPN) to fuse multi-scale feature maps and filter out complex background information.

[0075] In order to fuse feature maps of different scales, TSFFPN is proposed, which can not only realize the interaction between feature layers of different scales and reduce the loss of information of non-adjacent feature layers, but also filter complex background information. Its core is the TSF module, and its structure is as follows: Figure 4 As shown in Figure 2, the TSF module includes four steps: merging, refining, selecting and filtering.

[0076] (1) Merge: After the input remote sensing image passes through GLCCFNet, it outputs five feature layers of different sizes F1, F2, F3, F4, and F5. Select F3∈R C×2H×2W , F4∈R C×H×W , As the input of the TSF module. First, F3 and F5 are downsampled and upsampled respectively to obtain the feature map F3 (1) ∈R C×H×W and F5(1) ∈R C×H×W , so that their size is the same as F4. Then, they are concatenated in the channel dimension and fused through a 1x1 convolution layer, and finally the fused feature map U∈R is output. C×H×W :

[0077] U=Conv 1x1 (Concat(F3 (1) ,F4,F5 (2) ))

[0078] (2) Refining: The multi-scale feature map U is refined through an MSDC Block to further extract multi-scale features and obtain a richer multi-scale feature map Ω:

[0079] Ω=MSDC(U)

[0080] (3) Selection: In order to obtain the spatial attention weights of feature maps of different scales, first, channel-based average pooling P avg (·) and the maximum pooling P max (·) Extract spatial relationships:

[0081] SA avg =P avg (Ω),SA max =P max (Ω)

[0082] Among them, SA avg ∈R 1×W×H and SA max ∈R 1×W×H It is the spatial feature descriptor after average pooling and maximum pooling.

[0083] Then, the two spatial pooling features are concatenated to realize the interaction of spatial feature descriptors, and the convolution layer is used to convert the pooling features of the two channels into the spatial attention feature map SA of the three channels:

[0084] SA=Conv 2→3 (Concat(SA avg ,SA max ))

[0085] Finally, the Sigmoid activation function is applied to each spatial attention feature map SA i , i∈1,2,3, and obtain a multi-scale spatial mask It contains the spatial context weights required for feature maps of different scales:

[0086]

[0087] (4) Filtering: In order to filter the background noise, the multi-scale feature map is combined with the corresponding spatial mask Weighted, the weighted multi-scale feature maps are fused, and the spatial attention feature S∈R is obtained through the convolution layer C×W×H :

[0088]

[0089] Then, the feature map F4 is weighted using the spatial attention feature S to filter out complex background information and enhance the features of the target of interest, resulting in the feature map F4 (2) :

[0090] F4 (2) =S·F4

[0091] Finally, F4 (2) Upsample and downsample respectively to get feature map F3 (2) and F5 (2) ; F3 (2) 、F4 (2) , F5 (2) After feature fusion, output F i .

[0092] Step S103: Use MSSHead to obtain the features of the extreme aspect ratio target by increasing the receptive field of the fused feature map, and focus on the features of the extreme aspect ratio target.

[0093] There are a large number of objects with extreme shapes and sizes in remote sensing images (bridges, ports, dams, overpasses, etc.). In order to detect objects with extreme aspect ratios in complex backgrounds, a multi-scale and multi-shape detection head is designed, which combines the advantages of large receptive field and dynamic convolution. The structure of MSSHead is as follows Figure 5 As shown in (a), the decoupled detection head is used as its basic architecture, and its core part is the multi-scale and multi-shape module.

[0094] The structure of the multi-scale and multi-shape module is as follows: Figure 5 As shown in (b), first, through a strip pyramid attention (π L ) The multi-scale features F output from TSFFPN i Then, through the deformable spatial attention (π S ) adjusts the receptive field to the extreme shape target, and finally, through the task attention (π C ) focuses on bounding box regression and classification tasks. The formula of the multi-scale multi-shape module is as follows:

[0095] W(Fi )=π C (π S (π L (F i )·F i )·F i )·F i

[0096] Among them, F i W×H×C (i∈3,4,5) is the multi-scale feature layer after feature fusion.

[0097] (1) Strip pyramid attention (π L ): In order to obtain the features of objects with extreme aspect ratios, a strip pyramid attention mechanism is proposed. First, a 3x5 convolution is used to obtain the features of objects with a larger length than width. Then, a 5x3 convolution is used to obtain the features of objects with a larger width than width. Finally, all features are concatenated in the channel dimension and fused through a convolution layer to obtain F output i , the formula is as follows:

[0098] F output i =Conv(Concat(F i ,Conv 3*5 (F i ),Conv 5*3 (Conv 3*5 (F i ))))

[0099] Then, in F output i Then use global average pooling, convolutional layer and hard-sigmoid (σ) to obtain channel attention weights and compare them with F i Weighted, the formula is as follows:

[0100]

[0101] (2) Deformable Spatial Attention (π S ): In order to focus the receptive field on objects with extreme aspect ratios, a deformable attention mechanism is used. Its formula is as follows:

[0102]

[0103] Where K is the number of sparse sampling locations, p k +Δp k is the position moved by the self-learned spatial offset, Δp k Learned by deformable convolution, focusing on ambiguous areas. Δm k Represents the position Δpk The importance of self-learning.

[0104] (4) Task attention (π C ): In order to let the model focus on the classification and bounding box regression tasks, a task attention mechanism is added. The formula is as follows:

[0105] π C (F i )·F i =max(α 1 (F i )·F C +β 1 (F i ),α2(F i )·F C +β 2 (F i ))

[0106] Among them, [α 1 ,α 2 ,β 1 ,β 2 ] T =θ(·) is a hyperparameter used to control the activation threshold, and θ(·) is similar to DyRelu.

[0107] Corresponding to the above method, this embodiment also proposes a remote sensing image target detection system based on a global local context awareness network, including:

[0108] The feature extraction module is used to use GLCCFNet to cross-fuse the global and local context information in the spatial and channel dimensions of the input remote sensing image using a spatial-channel cross-fusion strategy to obtain a multi-scale feature map.

[0109] The feature fusion module is used to propose TSFFPN to fuse multi-scale feature maps and filter out complex background information.

[0110] The detection module is used to use MSSHead to obtain the features of targets with extreme aspect ratios by increasing the receptive field of the fused feature map, and to focus on the features of targets with extreme aspect ratios.

[0111] In order to verify the effectiveness of this scheme, further explanation is given below in combination with experimental data.

[0112] (I) Test data

[0113] In order to fully verify the effectiveness and robustness of the method of the present invention, three optical image datasets DIOR, RSOD, and HRRSD were selected for quantitative and qualitative analysis. In addition, in order to further verify the effect of migrating the BSPDet proposed in the present invention to SAR images, a commonly used SAR dataset, SSDD, was selected.

[0114] DIOR dataset: It contains a total of 23,463 images, 190,288 instances, and 20 categories, namely airplane (AL), airport (AT), baseball field (BF), basketball court (BC), bridge (B), chimney (C), dam (D), expressway service area (ESA), expressway toll station (ETS), harbor (HB), golf course (GC), ground track field (GTF), overpass (O), ship (S), stadium (SD), storage tank (ST), tennis court (TC), train station (TS), vehicle (V) and windmill (WM). In order to facilitate the training and verification of the model, the size of all images is unified to 800x800 pixels, which is divided into three parts, 5862 images are used as training sets, 5863 images are used as verification sets, and 11738 images are used as test sets. The image resolution of the DIOR dataset ranges from 0.5-30m, the multi-scale characteristics of the target are obvious, and the background of the target is complex and diverse.

[0115] RSOD dataset: It contains four categories, including 4993 aircraft, 1586 oil tanks, 191 playgrounds and 180 overpasses. The image size ranges from 512x512 to 1961x1193 pixels. The training set and test set are divided according to a unified standard and 5:5.

[0116] HRRSD dataset: It contains 21761 images, 55740 instances, and the image resolution ranges from 0.15 to 12m. According to the original description, it is divided into three parts for training and validating the target detection algorithm, of which 5401 images are used as training sets, 5417 images are used as validation sets, and 10943 images are used as test sets. It contains a total of 13 categories, namely ship (S), bridge (B), ground track field (GTF), storage tank (ST), basketball court (BC), tennis court (TC), airplane (AL), baseball diamond (BD), harbor (HB), vehicle (V), crossroad (CR), T junction (TJ), and parking lot (PK). It contains a large number of extreme aspect ratio targets and multi-scale targets in complex backgrounds.

[0117] SSDD dataset: It is one of the most commonly used algorithms in the field of SAR image ship target detection. It contains 1160 images and 2456 ship targets, and the size of each image is between 500 and 600 pixels. Its image resolution is 1 to 15 meters, and there are ship targets of different sizes and ships against complex backgrounds near the coast. The dataset is divided into training set, validation set and test set according to a unified standard and 7:2:1.

[0118] (II) Evaluation indicators

[0119] In order to quantitatively evaluate the performance of BSPDet in detecting multi-scale and extreme aspect ratio targets in complex backgrounds of remote sensing images, this experiment uses common indicators such as precision (Precision, P), recall (Recall, R), average precision (Average Precision, AP), mean average precision (mAP), parameters, floating point operations (FLOPs) and F1 score.

[0120] (III) Experimental details

[0121] All experiments use the same hardware and software equipment to ensure the fairness of the experiments. The hardware equipment includes Intel(R) Xeon(R) Silver 4114 CPU, 128GB memory and NVIDIA GeForce 3090Ti GPU. The software equipment includes 64-bit Windows operating system, CUDA 11.7, PyTorch 1.13 and Python 3.8.

[0122] All deep learning models use the same hyperparameters during training, and the number of iterations during training is set to 300; the batch size is set to 8. The image input size for training BSPDet on the DIOR and HRRSD datasets is 800x800 pixels, the input image size on the SSDD dataset is 640x640 pixels, and the input image size on the RSOD dataset is 1024x1024 pixels.

[0123] (IV) Comparative experiment

[0124] In order to quantitatively analyze the performance of the BSPDet algorithm, it is compared with several state-of-the-art object detection algorithms on the DIOR, HRRSD, and RSOD datasets.

[0125] (1) Comparative experiments on the DIOR dataset: BSPDet is compared with 10 excellent target detection algorithms in the DIOR dataset, including RetinaNet, CenterNet, DETR, ASSD, FSoDNet, DIAG-TR, SRAF-Net, MSFC-Net, RSADe, and YOLOv8l. As shown in Table 1, BSPDet achieves the best results, with a mAP of 73.7%. It achieves the best results on 7 target categories: bridge (45.6%), dam (64.9%), golf course (91.5%), harbor (65.6%), overpass (60.9%), ship (91.2), and vehicle (56.1%). It is worth noting that the Transformer-based DETR and DIAG-TR do not perform well on the detection of small-sized vehicles in complex backgrounds. They achieve mAPs of 14.4% and 33.8% respectively on the vehicle category. However, the BSPDet proposed in the present invention achieves the best result of 56.1%. This is mainly due to the fact that the TSFFPN proposed in this invention can not only fuse multi-scale features, but also filter complex background noise, thereby enriching and highlighting the features of small targets.

[0126] The BSPDet proposed in this invention not only works well for small targets, but also has good detection effects on some extremely large-sized targets, such as golf courses (81.7%). This is mainly due to the fact that the GLCCFNet proposed in this invention not only has a large receptive field, but also can efficiently fuse global and local context features, so it can take into account the characteristics of both extremely large and extremely small-sized targets.

[0127] In addition, BSPDet also achieves the best results for some extreme aspect ratio targets, such as bridges (45.6%), dams (64.9%), ports (65.6%) and overpasses (60.9%). This is mainly because the MSSHead proposed in the present invention can not only increase the receptive field, but also focus the receptive field on the extreme aspect ratio targets.

[0128] Experiments show that the BSPDet proposed in the present invention can not only obtain the features of extremely large-sized targets, but also obtain the features of extremely small targets, and has good detection results for multi-scale targets and targets with extreme aspect ratios in complex backgrounds.

[0129] Table 1 Comparative experiments of BSPDet on DIOR

[0130]

[0131] (2) Comparative experiments on the HRRSD dataset: BSPDet is compared with six latest target detection algorithms on the HRRSD dataset, including MGCN, Cascade R-CNN, EGAT, AGMF-Net, SB-MSN, and TMAFNet. As shown in Table 2, the BSPDet proposed in the present invention achieves the best results, with an mAP of 94.40%, and achieves the best results in eight categories: ship (97.40%), baseball diamond (95.60%), ground track field (99.40%), harbor (99.30%), bridge (97.10%), vehicle (98.30%), crossroad (97.10%), and Tjunction (89.50%). Among them, ships and vehicles are small-sized targets, ground tracking stations, tennis courts, crossroads, and T-junctions are large-sized targets, and ports and bridges are targets with extreme aspect ratios. Moreover, these targets are in complex scenes. Experimental results further show that the BSPDet proposed in this invention has a good detection effect on multi-scale targets and targets with extreme aspect ratios in complex backgrounds.

[0132] Table 2 Comparative experiments of BSPDet on HRRSD dataset

[0133] Methods mAP AL S ST BD TC BC GTF HB B V CR TJ PK MGCN 86.27 93.27 88.96 92.10 91.73 83.98 55.79 93.46 92.96 89.79 89.85 94.95 79.87 74.69 Cascade R-CNN 89.10 98.00 91.40 95.60 91.30 92.60 72.20 98.20 93.90 89.60 95.50 92.30 79.70 68.40 EGAT 91.30 95.67 93.92 93.45 94.52 85.89 86.63 95.72 94.36 91.23 89.76 95.86 83.71 86.22 AGMF-Net 92.03 99.79 93.41 98.05 92.66 96.12 82.79 99.06 93.99 93.82 96.60 94.09 84.02 71.93 SB-MSN 92.60 99.30 94.50 97.80 93.90 96.00 80.20 99.00 96.00 92.70 96.60 95.30 86.40 76.50 TMAFNet 92.96 99.90 93.34 97.96 93.96 96.07 83.99 99.22 95.21 95.14 96.93 95.86 87.64 73.46 Our 94.40 99.30 97.40 97.20 95.60 94.90 81.40 99.40 99.30 97.10 98.30 97.10 89.50 80.70

[0134] (3) Comparative experiment on RSOD dataset: The detection results of BSPDet on RSOD dataset are compared with the six best object detection algorithms, including Sig-NMS, Cascade R-CNN, SSAFNet, ABNet, AGMF-Net, and DA 2 FNet. The detailed results are shown in Table 3. BSPDet achieves the best mAP, reaching 95.10%, and achieves the best detection result on the aircraft category, reaching 96.50%.

[0135] Table 3 Comparative experiments of BSPDet on RSOD dataset

[0136] Method Backbone mAP Aircraft Oil tank Overpass Playground Sig-NMS VGG16 89.40 80.60 90.60 87.40 99.10 Cascade R-CNN ResNet101 91.30 94.20 96.10 83.20 99.00 SSAFNet Hourglass104 92.82 95.75 98.39 84.66 92.50 ABNet ResNet50 94.17 91.49 96.14 89.61 99.44 AGMF-Net DarkNet53 94.30 96.02 99.02 82.43 99.70 <![CDATA[DA 2 FNet]]> Hourglass104 94.78 95.75 97.73 90.56 95.12 Our GLCCFNet 96.50 96.20 97.40 93.30 99.00

[0137] (V) Further verification on SAR remote sensing images

[0138] Due to differences in image resolution, natural differences in the shape and size of ships, and interference from nearshore objects, the use of target detection algorithms to automatically detect ship targets in SAR images still faces problems with complex backgrounds, multi-scale targets, and targets with extreme aspect ratios.

[0139] In order to test the effect of BSPDet proposed in this paper on migrating to SAR images, the detection results of BSPDet on the SSDD dataset are compared with the seven best SAR ship target detection algorithms, including DCMSNN, FBR-Net, CenterNet++, YOLOv5, YOLOv8, AFSar and DEPDet. The detailed results are shown in Table 4. BSPDet achieves the best results in mAP and F1 indicators, reaching 98.7% and 97.0% respectively. In addition, Figure 6 As shown in the figure, BSPDet has good detection results for small-sized ships and ships with extreme aspect ratios in SAR images, and still has good results for ships in complex backgrounds near the coast.

[0140] Experimental results show that the BSPDet proposed in this invention can also detect multi-scale targets and targets with extreme aspect ratios in complex backgrounds in SAR images.

[0141] Table 4 Comparative experiments of BSPDet on SSDD dataset

[0142] Method P R F1 mAP DCMSNN 90.5 89.1 89.8 89.4 FBR-Net 92.8 94.0 93.4 94.1 CenterNet++ 83.3 95.2 88.9 95.1 YOLOv5 91.9 90.4 91.1 96.1 YOLOv8 95.5 91.5 93.3 96.8 AFSar 94.1 98.2 96.1 97.7 DEPDet 97.9 92.3 93.3 98.2 Our 97.2 96.0 97.0 98.7

[0143] (VI) Qualitative analysis

[0144] In order to intuitively demonstrate the performance of the BSPDet proposed in the present invention, its detection results and feature graphs are qualitatively analyzed.

[0145] (1) Visual analysis of detection results: In order to intuitively demonstrate the performance of BSPDet, its detection results on the HRRSD, RSOD, DIOR, and SSDD datasets are visualized and analyzed. Figure 6 As shown in (a), BSPDet has excellent detection performance on small-sized targets, and can accurately detect small targets regardless of whether they are dense or sparse. Figure 6 As shown in (b), BSPDet has a good detection effect on extremely large-sized targets, such as overpasses, playgrounds, and golf courses. It is worth noting that their area basically occupies the entire image, and it is difficult for general target detection algorithms to obtain their complete features. Thanks to the large receptive field of the MSDC Module proposed in this invention, their complete features can be easily obtained. In addition, Figure 6 As shown in (c), BSPDet can detect objects of different sizes in the same image. Figure 6 The detection results of (a), (b) and (c) show that BSPDet can not only obtain the features of extremely large-sized targets, but also the features of small-sized targets, and has good detection results for multi-scale targets.

[0146] like Figure 6 As shown in (d), even in complex backgrounds, BSPDet can still accurately detect objects of various scales. This shows that BSPDet can not only extract the features of multi-scale objects, but also filter out a large amount of complex backgrounds. Figure 6 As shown in (e), BSPDet has a good detection effect on some targets with extreme aspect ratios, such as airports, bridges, ships, etc. This shows that BSPDet can obtain the features of targets with extreme aspect ratios.

[0147] From the above analysis, it can be seen that BSPDet can accurately detect targets with extreme aspect ratios and multi-scale targets in complex backgrounds of remote sensing images.

[0148] (2) Feature graph visualization analysis: In order to intuitively demonstrate the effects of GLCCFNet in extracting multi-scale features and TSFFPN in fusing multi-scale features and filtering complex background noise, their feature graphs are visualized and analyzed. Figure 7 As shown in (b), the baseline backbone multi-scale feature extraction capability is insufficient, resulting in incomplete large-scale targets and omission of small-scale targets in the final feature map. In addition, the features of multi-scale targets are interfered by complex background noise, and there are no clear and distinguishable boundaries, which affects the positioning effect of the target. Figure 7As shown in (c), the input image can extract rich multi-scale target features after GLCCFNet, but there is still some noise interference. Figure 7 As shown in (d), after TSFFPN fuses features of different scales and filters noise, the features of multi-scale targets are more obvious, and the complex background noise is basically suppressed. The experimental results show that GLCCFNet and TSFFPN are very effective in extracting multi-scale features and filtering complex background noise.

[0149] (3) Comparative analysis of detection results of BSPDet and other algorithms: In order to intuitively compare the differences in detection effects between BSPDet and other algorithms, the detection results of BSPDet are compared with some of the latest SOTA algorithms, including PKINet, RT-DETR and YOLOv10, in the same image area. Figure 8 As shown in (a), since the target and background features are very similar, PKINet and RT-DETR mistakenly identify the background as the target. Thanks to TSFFPN, it can filter the complex background and highlight the characteristics of the target. BSPDet can well identify the target in the complex background. Figure 8 As shown in (b), since PKINet, RT-DETR and YOLOv10 have weak expression capabilities for small-sized targets, it is easy to miss or misdetect small-sized aircraft targets. The BSPDet proposed in the present invention can improve the detection effect of small targets.

[0150] like Figure 8 As shown in (c), since PKINet and RT-DETR cannot adjust the receptive field of the target, it is easy to mistake a part of the features of a target as a complete target, resulting in repeated detection. Thanks to the receptive field adjustment function of TSFFPN, different receptive fields can be given to targets of different sizes, thus solving this problem well. Figure 8 As shown in (d), due to their weak feature expression ability, it is difficult to distinguish the features of similar objects, and the land is mistakenly identified as a basketball court. However, the CLCANet proposed in the present invention can solve this problem well.

[0151] In order to automatically detect multi-scale targets and extreme aspect ratio targets in complex backgrounds in remote sensing images, the present invention proposes a remote sensing image target detection method based on a global local context-aware network. The present invention proposes GLCCFNet to obtain multi-scale features, which fully considers the complementary characteristics of global and local context information in channels and spaces, and realizes the efficient fusion of global and local context information. Then, the present invention proposes TSFFPN fusion multi-scale feature layer, which can not only reduce the information loss of non-adjacent feature layers, balance the semantic gap of feature maps of different scales, but also automatically select the spatial attention weights required by feature layers of different scales, thereby filtering complex background information. Finally, the MSSHead proposed in the present invention obtains the features of extreme aspect ratio targets, which fully combines the advantages of large receptive field and dynamic convolution. Experimental results show that the BSPDet proposed in the present invention achieves the best results on three optical datasets and one SAR dataset, and it can accurately detect multi-scale targets and extreme aspect ratio targets in complex backgrounds in remote sensing images.

[0152] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A remote sensing image target detection method based on a global local context awareness network, characterized in that: include: Firstly, GLCCFNet is used to cross-fuse the global and local context information in the spatial and channel dimensions of the input remote sensing image using a spatial-channel cross-fusion strategy to obtain multi-scale feature maps. Then TSFFPN is proposed to fuse multi-scale feature maps and filter out complex background information; Finally, MSSHead is used to obtain the features of targets with extreme aspect ratios by increasing the receptive field of the fused feature map, and focuses on the features of targets with extreme aspect ratios.

2. The method for remote sensing image target detection based on global local context awareness network according to claim 1 is characterized in that: GLCCFNet consists of three parts: Stage n, MSDC Block and MSDC Module. The remote sensing image outputs a multi-scale feature map after Stage n. Each Stage contains a CSPBlock structure. The input features are divided into two parts after downsampling and convolution layers. One part is directly output, and the other part passes through multiple MSDC Blocks. The output results of these two parts are spliced ​​in the channel dimension and passed through a convolution layer to output the final result; each MSDCBlock contains a Conv and MSDC Module. The output results of these two parts are spliced ​​in the channel dimension and passed through a convolution layer to output the final result.

3. The method for remote sensing image target detection based on global local context awareness network according to claim 2 is characterized in that: GLCCFNet consists of four stages, and the input and output of each stage are represented as and The structure of each stage is as follows: First, the input feature F n Output feature map after downsampling and convolution layer Then, X n Divide into two parts equally according to the channel direction X (1) n and X (2) n , X (1) n Output feature map directly without processing X (2) n After multiple serial MSDC blocks, the feature map is output Finally, X (1) n+1 and X (2) n+1 Concatenate in the channel dimension and fuse through a 1x1 convolution layer to output the feature map F n+1 .

4. The method for remote sensing image target detection based on global local context awareness network according to claim 3 is characterized in that: The structure of MSDC Block is: Input feature X (2) n Divide into two parts according to the channel direction and X (2) (1,n) Get global context information G through MSDC Module (1) n , X (2) (2,n) Model the local context information L through a 3x3 convolutional layer (1) n ; Then the global and local context information are fused through the space-channel cross fusion strategy.

5. The method for remote sensing image target detection based on global local context awareness network according to claim 4 is characterized in that: The global and local context information are fused through the spatial-channel cross fusion strategy including: In the first branch, the local context information L (1) n Perform global maximum pooling to obtain feature map Z l ; then Z l The local context channel weight S is obtained through a Sigmoid function l , and finally S l With the global context information G (1) n Multiply point by point to get G (1) n+1 ; In the second branch, the global context information G (1) n The spatial descriptors are obtained by global average pooling and global maximum pooling respectively, and the feature map Z is obtained by feature concatenation and convolution. g ; then Z g The global context space weight S is obtained through a Sigmoid function g , and finally S g With local context information L (1) n Multiply point by point to get L (1) n+1 ; Finally, the local context L (1) n+1 and the global context G (1) n+1 Concatenate in the channel dimension and fuse through a 1x1 convolution to obtain the global-local feature map X (2) n+1 .

6. The method for remote sensing image target detection based on global local context awareness network according to claim 4 is characterized in that: The structure of MSDC Module is as follows: First, input feature X (2) (1,n) After a convolutional layer, it is evenly divided into five parts G according to the channel direction. (1) (1,n) , G (1) (2,n) , G (1) (3,n) , G (1) (4,n) , G (1) (5,n) Then, they are respectively subjected to dilated convolutions with dilation rates of 1, 2, 4, 8, and 16, and the outputs of dilated convolutions with different dilation rates are stacked in sequence; finally, all the output results are concatenated in the channel dimension and fused through a 1x1 convolution to output the final result G (1) n .

7. The method for remote sensing image target detection based on global local context awareness network according to claim 1, characterized in that: The core of TSFFPN is the TSF module, which includes four steps: merging, refining, selecting and filtering. Specifically: ① Merge: Select three feature maps of different sizes F3∈R output by GLCCFNet C×2H×2W , F4∈R C×H×W , As input; first, F3 and F5 are downsampled and upsampled respectively to obtain feature map F3 (1) ∈R C×H×W and F5 (1) ∈R C×H×W ; Then, they are concatenated in the channel dimension and fused through a 1x1 convolutional layer, and finally the fused feature map U∈R is output C×H×W ; ② Refinement: The multi-scale feature map U is refined through an MSDC Block to obtain the feature map Ω; ③ Selection: First, the feature map Ω is subjected to channel-level average pooling and maximum pooling to extract the spatial feature descriptor SA avg ∈R 1×W×H and SA max ∈R 1×W×H ; Then, SA avg and SA max After splicing, the convolution layer is used to convert the spatial pooling features of the two channels into the spatial attention feature map SA of the three channels; finally, the Sigmoid activation function is applied to each spatial attention feature map SA i , i∈1,2,3, and obtain a multi-scale spatial mask ④ Filtering: Multi-scale feature maps F3, F4, F5 are combined with the corresponding spatial masks Weighted, the weighted multi-scale feature maps are fused, and the spatial attention feature S∈R is obtained through the convolution layer C×W×H ; Then, the feature map F4 is weighted by the spatial attention feature S to obtain the feature map F4 (2) ; Finally, F4 (2) Upsample and downsample respectively to obtain feature map F3 (2) and F5 (2) ; F3 (2) 、F4 (2) , F5 (2) After feature fusion, output F i .

8. The method for remote sensing image target detection based on global local context awareness network according to claim 1, characterized in that: MSSHead uses a decoupled detection head as its basic architecture. The core part is a multi-scale multi-shape module. The structure of the multi-scale multi-shape module is as follows: First, the multi-scale feature F output by TSFFPN is obtained through strip pyramid attention. i Then, the receptive field is adjusted to the extreme shape objects through deformable spatial attention, and finally, task attention is used to focus on the bounding box regression and classification tasks.

9. A remote sensing image target detection system based on a global local context awareness network, characterized in that: include: The feature extraction module is used to use GLCCFNet to cross-fuse the global and local context information in the spatial and channel dimensions of the input remote sensing image using a spatial-channel cross-fusion strategy to obtain a multi-scale feature map; Feature fusion module, used to propose TSFFPN to fuse multi-scale feature maps and filter out complex background information; The detection module is used to use MSSHead to obtain the features of targets with extreme aspect ratios by increasing the receptive field of the fused feature map, and to focus on the features of targets with extreme aspect ratios.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.