Remote sensing target detection method, equipment and medium

Through the combination methods of multi-branch expansion convolution, dual attention and asymmetric decomposition convolution layers, the multi-scale feature fusion, complex background interference and rotary target characterization problems in remote sensing target detection are solved, and high-precision remote sensing target detection is achieved.

CN120431479APending Publication Date: 2025-08-05NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 17 Cited by

Patent Information

Application Number
CN202510513529.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing remote sensing object detection technology has shortcomings in multi-scale feature fusion, complex background interference suppression, rotary target characterization and downsampling information loss, making it difficult to achieve high-precision detection in remote sensing images.

Method used

The combination method of multi-branch expansion convolution structure, space and channel dual attention mechanism, cascade pooling module and asymmetric decomposition convolution layer is adopted to improve the accuracy of remote sensing target detection through multi-scale feature fusion, dynamic channel attention calibration and gated fusion network.

Benefits of technology

High-precision detection of multi-scale, rotating targets and high-density small targets in complex remote sensing scenarios is achieved, and the insufficient detection accuracy caused by fixed receptive fields, single-scale processing and information loss in traditional methods is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431479A_ABST
    Figure CN120431479A_ABST
Patent Text Reader

Abstract

The invention relates to a remote sensing target detection method and device and a medium, and the method comprises the steps: inputting a feature map to a backbone network for feature extraction, inputting an extracted feature tensor into a multi-branch expansion convolution structure, and extracting multi-scale features through convolution kernels with different expansion rates. Then, multi-scale features are fused through a space and channel double-path attention mechanism, and enhanced features are generated; and the enhanced features are further input into a cascade pooling module to generate multi-level reconstruction features, and weight coefficients are calculated through a gating fusion network to carry out weighted fusion, so that multi-scale fusion features are obtained. Next, these features are input into an asymmetric decomposition convolutional layer for downsampling, and dynamic channel attention calibration is performed to generate channel enhanced features. And finally, inputting the feature map processed by the backbone network and the neck network into a detection head network, and outputting a target bounding box and category prediction. According to the method, high-precision detection of multi-scale rotating targets and high-density small targets is realized in a complex remote sensing scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and remote sensing image analysis, and in particular to a remote sensing target detection method, device and medium. Background Art

[0002] In recent years, with the rapid development of satellite imaging technology, remote sensing target detection has demonstrated significant application value in fields such as environmental monitoring. However, due to the characteristics of remote sensing images, such as large variations in target scale, complex backgrounds, dense distribution of small targets, and arbitrary target orientation, traditional target detection technologies face significant challenges. While existing deep learning-based target detection frameworks have made significant progress in natural scene detection, they still face the following technical bottlenecks in the remote sensing field:

[0003] Imperfect multi-scale feature fusion mechanisms: Traditional Feature Pyramid Networks (FPNs) employ a fixed-level feature fusion strategy, making them incapable of coexisting objects with scales varying by tens of times in remote sensing scenarios. In particular, small objects are susceptible to downsampling during feature transfer, resulting in sub-pixel feature loss. While existing methods that expand the receptive field through dilated convolutions can improve global perception, their fixed dilation rate configuration lacks the ability to dynamically adapt to objects of varying scales.

[0004] Inadequate suppression of complex background interference: Remote sensing images contain complex textures and numerous target-like interferences (such as building shadows and cloud cover). Traditional attention mechanisms are prone to false activations when performing global weighting in the channel / spatial dimension, leading to confusion between background noise and target features. Research has shown that complex background interference can reduce the target's signal-to-noise ratio to below 0.5dB, seriously impacting detection reliability.

[0005] Downsampling algorithms are limited in their ability to capture detailed object features: The geometric symmetry of traditional convolution kernels results in a reduced angular sensitivity to rotated object features. While existing downsampling algorithms improve bounding box regression through an angle prediction branch, the underlying feature extraction stage still relies on conventional convolution operations, which cannot effectively capture directional edge features, resulting in reduced detection accuracy for rotated objects.

[0006] Severe information loss during downsampling: Traditional downsampling methods (such as max pooling and strided convolution) lack mechanisms to protect rotational features when compressing feature maps, resulting in the loss of directionally sensitive information. Experiments show that conventional downsampling can attenuate the feature response of small objects by over 60%, becoming a key factor limiting detection accuracy.

[0007] To address these issues, existing technologies have the following limitations: While methods based on multi-scale feature fusion (such as HRNet) can maintain high-resolution features, they fail to address the dynamic balance of cross-scale context perception; rotation-sensitive networks (such as the CSL method) lack a spatial decoupling mechanism during feature encoding, resulting in insufficient rotation invariance; and attention-enhanced models (such as DualAttention) are prone to over-suppression in complex backgrounds, destroying target details. Furthermore, existing downsampling modules generally use fixed kernel convolutions, which are difficult to adapt to the geometric characteristics of remote sensing targets. Summary of the Invention

[0008] The present invention provides a remote sensing target detection method, device and medium, which aims to solve the problems of existing remote sensing target detection technology in multi-scale feature fusion, complex background interference suppression, rotation target representation and downsampling information loss.

[0009] To achieve the above object, the present invention provides a remote sensing target detection method in a first aspect, comprising the following steps:

[0010] The input feature map is fed into the backbone network for feature extraction. During the feature extraction process, the feature tensor is fed into a multi-branch dilated convolution structure, and multi-scale features are extracted through convolution kernels with different parallel dilation rates.

[0011] The multi-scale features are fused with spatial and channel dual-path attention to generate enhanced features;

[0012] Input the enhanced features obtained in the backbone network for the last time into the cascade pooling module for multi-scale pyramid construction to generate multi-level reconstruction features;

[0013] Input the multi-level reconstruction features into the gated fusion network, dynamically calculate the weight coefficients of the reconstruction features of each level based on the learnable parameters, and perform weighted fusion through the calculated weight coefficients to generate multi-scale fusion features;

[0014] The multi-scale fusion features are input into the asymmetric decomposition convolution layer, and the horizontal and vertical convolution kernels are used to separate and process the spatial dimensions to obtain down-sampled features;

[0015] Perform dynamic channel attention calibration on the downsampled features, calculate the channel weight matrix through the adaptive compression ratio mechanism, and generate channel enhanced features;

[0016] The enhanced features obtained by the backbone network processing and the channel enhanced features after the neck network processing are input into the YOLOv10 detection head, and the target bounding box and category prediction results are output.

[0017] Furthermore, the multi-branch dilated convolution structure includes at least three parallel branches, the dilation rates of each parallel branch are 1, 3, and 5, respectively, and each parallel branch includes a cascaded D-Bottleneck module, which adds the output of the dilated convolution to the input features through a residual connection.

[0018] Furthermore, the fusion process of spatial and channel dual attention includes:

[0019] The residual features output by the multi-branch dilated convolution are added element-by-element to the original input features to generate a baseline feature matrix;

[0020] Perform local attention calculation on the benchmark feature matrix: compress the number of feature channels through channel compression convolution, reconstruct it to the original channel dimension after ReLU activation, and then generate the spatial attention weight matrix through the Sigmoid function;

[0021] Perform global attention calculation on the baseline feature matrix: extract channel-level global description through adaptive average pooling, reconstruct it to the original channel dimension after dimensionality reduction by fully connected layer and ReLU activation, and then generate channel attention weight matrix through Sigmoid function;

[0022] Add the spatial attention weight matrix and the channel attention weight matrix, and generate a mixed attention map through Sigmoid normalization;

[0023] The hybrid attention map is used to perform weighted fusion of the original input features and the residual features.

[0024] Furthermore, the method of inputting the enhanced features obtained for the last time in the backbone network into the cascade pooling module for multi-scale pyramid construction to generate multi-level reconstruction features includes:

[0025] Performing three cascaded maximum pooling operations on the input enhanced features, respectively using a pooling kernel with a stride of 1 and a kernel size of 5×5, to generate a multi-scale pyramid of three-level down-sampled features;

[0026] Perform channel reconstruction on each pooled feature in the multi-scale pyramid to generate channel reconstruction features;

[0027] Perform spatial reconstruction on the channel reconstruction features to generate a spatial attention map;

[0028] The spatial attention map is activated by Sigmoid and multiplied position by position with the channel reconstruction feature to output multi-level reconstruction features.

[0029] Furthermore, a method for performing channel reconstruction on each feature map in the multi-scale pyramid and generating channel reconstruction features includes:

[0030] Each pooled feature in the multi-scale pyramid is divided into four groups of sub-features along the channel dimension;

[0031] Perform 3×3 grouped depth convolution on each group of sub-features, and keep the output feature dimension the same as the input;

[0032] Perform global average pooling on the output features of the grouped depth convolution to generate a channel description vector;

[0033] The channel description vector is input into the fully connected layer and reduced to C / 16 dimensions. After ReLU activation, it is reconstructed to the original channel dimension to generate the channel attention weight.

[0034] Multiply the channel attention weights by the output features of the depth convolution channel by channel to obtain the sub-features after channel reconstruction;

[0035] The four groups of sub-features are spliced along the channel dimension to generate channel reconstruction features;

[0036] The method of spatially reconstructing the channel reconstruction features and generating a spatial attention map includes:

[0037] Calculate the average response map and the maximum response map along the channel dimension for the channel reconstruction features;

[0038] The average response map and the maximum response map are concatenated and input into a 7×7 convolutional layer to generate a spatial attention map.

[0039] Furthermore, the weight coefficient calculation method of the gated fusion network is:

[0040] Performing global average pooling on each level feature in the multi-level reconstruction features to generate a preliminary weight vector;

[0041] The preliminary weight vector is input into a fully connected layer consisting of a learnable parameter matrix and bias to generate a gating signal;

[0042] The gating signal is activated by the Sigmoid function to generate the normalized gating weight, where the calculation method of the normalized gating weight is:

[0043] g i =σ(W g ·e i +b g )

[0044] Among them, σ is the Sigmoid function, g i It represents the initial gating weight of the i-th level feature, which is generated by the fully connected layer and the Sigmoid function, reflecting the importance of the level feature. i Represents the channel-level global description vector of the i-th level feature, W g and b g These are all parameters that can be learned by the network;

[0045] Input the normalized gated weights into the Softmax function for cross-layer normalization and calculate the weight coefficients:

[0046]

[0047] Among them, w i represents the normalized fusion weight of the i-th level feature, exp represents the exponential function, which is used for numerical scaling before Softmax normalization; g j Represents the initial gating weight of the j-th level feature, used for denominator summation calculation, and i and j are both level indices.

[0048] Furthermore, the calculation method for downsampling of the asymmetric decomposition convolution layer is:

[0049] Performing a 1×1 convolution operation on the multi-scale fusion features, adjusting the channel dimension to the target number of channels, and generating channel alignment features;

[0050] Perform horizontal decomposition convolution on the channel alignment features: use a depthwise separable convolution kernel with a size of 3×1 and a step size of 2 to extract horizontal edge features and generate a horizontal feature map;

[0051] Performing vertical decomposition convolution on the horizontal feature map: using a depthwise separable convolution kernel with a size of 1×3 and a step size of 2 to extract vertical edge features and generate a vertical feature map;

[0052] The horizontal feature map is added to the vertical feature map element by element to obtain the downsampled features.

[0053] Furthermore, a method for performing dynamic channel attention calibration on the downsampled features, calculating a channel weight matrix through an adaptive compression ratio mechanism, and generating channel enhancement features includes:

[0054] Performing a global average pooling operation on the downsampled features to generate a global statistical vector of channel dimension;

[0055] Inputting the global statistical vector into a compression-excitation module consisting of a two-layer fully connected network for compression and excitation, wherein the first layer compresses the number of channels and the second layer restores the number of channels to the original number;

[0056] Perform ReLU activation and Sigmoid transformation on the compressed and stimulated global statistical vector in sequence to generate the channel attention weight matrix;

[0057] The channel attention weight matrix is multiplied by the downsampled features channel by channel to generate channel enhanced features.

[0058] Furthermore, a multi-branch dilated convolution structure is deployed in the backbone network and detection head of YOLOv10, the cascade pooling module is deployed in the feature fusion neck, and the asymmetric decomposition convolution layer replaces the standard downsampling module in the original network.

[0059] To achieve the above-mentioned purpose, the second aspect of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the remote sensing target detection method, and the processor is configured to execute the program stored in the memory.

[0060] To achieve the above-mentioned object, the third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the remote sensing target detection method are executed.

[0061] Beneficial effects of the present invention:

[0062] Compared with the existing technology, the remote sensing target detection method, device and medium provided by the present invention solve the core problems in remote sensing target detection through the synergistic effect of three innovative modules: to address the problem of multi-scale feature fusion, the MBAFF module adopts a mixed expansion rate multi-branch convolution structure to capture differentiated cross-scale contextual information, and combines the spatial-channel dual-path attention mechanism to dynamically enhance the local details and global semantic associations of small targets, effectively alleviating the problem of weakening features of low-pixel small targets; to address complex background interference, the GMAP module constructs a multi-level receptive field pyramid through cascade pooling, and introduces a channel-space dual reconstruction mechanism to suppress spectral redundancy and background noise, and combines gated weights to dynamically fuse multi-scale features, significantly improving the feature significance of the target area; to address the problem of detailed target representation and downsampling information loss, the ADDown module realizes independent perception of horizontal / vertical edges through asymmetric convolution kernel decomposition to enhance rotation sensitivity, and uses dynamic compression ratio channel attention to adaptively calibrate feature channels, so as to maximize the retention of geometric features of rotated targets during the downsampling process. The three modules form collaborative optimization from the three dimensions of feature extraction, multi-scale fusion and spatial compression, jointly solving the problem of insufficient detection accuracy caused by fixed receptive field, single-scale processing and information loss in traditional methods, and realizing high-precision detection of multi-scale, rotating targets and high-density small targets in complex remote sensing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments.

[0064] Figure 1 This is a flow chart of a remote sensing target detection method disclosed in an embodiment of the present invention.

[0065] Figure 2 This is a diagram of an improved YOLO V10 structure disclosed in an embodiment of the present invention.

[0066] Figure 3 This is a network structure diagram of an MBAFF module disclosed in an embodiment of the present invention.

[0067] Figure 4 This is a network structure diagram of a D-Bottleneck disclosed in an embodiment of the present invention.

[0068] Figure 5 This is a visual comparison diagram of the dilated convolution disclosed in an embodiment of the present invention in terms of improving the receptive field.

[0069] Figure 6 This is a network structure diagram of a BAFM disclosed in an embodiment of the present invention.

[0070] Figure 7 This is a network structure diagram of a GMAP disclosed in an embodiment of the present invention.

[0071] Figure 8 This is an ADDown network structure diagram disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0072] In order to enable those skilled in the art to fully understand the technical solutions and implementation details of the present invention, the steps described in the claims are specifically described below in conjunction with the accompanying drawings and examples. It should be noted that the core improvement modules of the present invention include multi-branch attention feature fusion (MBAFF), gated multi-scale attention pyramid (GMAP), asymmetric dynamic downsampling (ADDown), all of which have clear deployment locations in the YOLOv10 architecture (such as Figure 2 As shown in Figure 3, each module significantly improves the remote sensing target detection performance through a collaborative working mechanism.

[0073] Figure 2This article demonstrates the deployment and data flow of the core modules in the improved YOLOv10 architecture. This architecture densely embeds three improved modules across the backbone network, feature fusion neck, and detection head: Multi-branch Attention Feature Fusion (MBAFF) replaces the original C2f module with multi-branch dilated convolutions and an attention mechanism to enhance multi-scale object perception; Gated Multi-scale Attention Pyramid (GMAP) is deployed in the neck and dynamically balances multi-scale features through cascaded pooling and gated fusion; Asymmetric Dynamic Downsampling (ADDown) replaces the traditional downsampling layer and combines asymmetric convolution with dynamic attention to preserve rotated object information. Each module works together through the data flow indicated by the arrows: MBAFF iteratively processes features in the backbone network and detection head; GMAP performs multi-scale feature fusion in the neck; and ADDown optimizes rotated object representation during spatial dimensionality reduction. The following sections will elaborate on the implementation principles and technical results, combining technical details of each step with accompanying figures.

[0074] like Figure 1 As shown, the present invention provides a remote sensing target detection method. The core of the method is to solve the multi-scale, complex background and rotation sensitivity problems in remote sensing target detection through modular design. The technical implementation details of each step are as follows:

[0075] Step S100: Input the input feature map into the backbone network for feature extraction. During the feature extraction process, the feature tensor is input into the multi-branch dilated convolution structure, and multi-scale features are extracted through convolution kernels with differentiated parallel dilation rates.

[0076] In remote sensing target detection, target scales vary widely, small targets predominate, and they are often affected by background noise. To effectively extract feature information from targets of varying scales, the YOLO-MGA framework introduces a multi-branch dilated convolutional structure into the backbone network. This structure uses convolution kernels with varying dilation rates to enhance the receptive field, adapting to the detection needs of targets of varying scales.

[0077] The design goals of the multi-branch dilated convolutional structure are:

[0078] Enhanced small target detection: Because small targets account for a small percentage of pixels in remote sensing images, traditional convolutional networks tend to overlook their features. Parallel dilated convolutions can effectively increase the ability to express the features of small targets and improve the model's ability to capture fine-grained features.

[0079] Constructing multi-scale receptive fields: By using different expansion rates (e.g., 1, 3, 5) on different branches, contextual information of different scales can be obtained without increasing the amount of computation, enabling the model to have good detection capabilities for both large-scale and small-scale targets.

[0080] Alleviate the problem of feature loss: Traditional convolution kernels with fixed receptive fields may not be able to fully capture long-distance contextual information, while dilated convolution can increase the receptive field while maintaining image resolution, preventing small target features from being eroded by downsampling during network propagation.

[0081] Figure 3 The detailed structure of the multi-branch attention feature fusion (MBAFF) module is shown. The input feature map (Feature Map) is first channel-expanded through convolution (Conv), and then split (Split) into multiple parallel branches (Branch1, Branch2, etc.), where the basic branch retains the original features, and the deep branch extracts multi-scale receptive field features through a cascaded D-Bottleneck module (containing dilated convolutions with differentiated dilation rates). The features output by all branches are concatenated (Concat) to achieve multi-scale information fusion, and then the channel dimension is restored through compressed convolution (Conv), and finally input into the bilateral attention feature mixer (BAFM). This module captures local details and global context through multi-branch dilated convolution, and uses BAFM's dual-path attention mechanism (spatial-channel collaboration) to dynamically enhance key features, effectively alleviating the information loss problem caused by scale differences and low pixel ratios of remote sensing targets.

[0082] Figure 4 The cascade structure design of the dilated bottleneck module (D-Bottleneck) is presented. Its core is to achieve multi-scale feature extraction through two dilated convolutions with different dilation rates. The input features are firstly k After the first dilated convolution (dilation rate d = 3), the feature stability is enhanced by batch normalization (BatchNorm) and ReLU activation function; the processed features are further expanded by the second dilated convolution (dilation rate d = 3) to expand the receptive field, and then the feature amplitude is adjusted by weighted operation before being passed to the parallel branch Branch k+1 The module adopts a dual-branch cascade design (Branch k With Branch k+1 ), the exponential expansion of the receptive field is achieved by stacking dilated convolutional layers, while batch normalization and ReLU activation are used to ensure the stability of gradient propagation.

[0083] Figure 5The role of dilated convolution in improving the receptive field is demonstrated through visual comparison, where 3×3 convolution kernels with different dilation rates are used to gradually expand the effective receptive field. For example, when the dilation rate d = 2, the receptive field of a single convolution layer expands from 3×3 to 5×5; after multiple cascaded dilated convolution modules, the cumulative receptive field can be progressively expanded to 16×16. This design enables the network to capture cross-region contextual correlation features through differentiated dilation rate configurations, especially enhancing the model's perception of small targets and solving the problem of difficulty in detecting low-pixel targets due to weakened local features. The figure intuitively illustrates the core mechanism of the multi-branch dilated convolution structure in breaking through the scale limitations of traditional single-branch convolution.

[0084] When implementing it specifically, Figure 3 As shown, the input feature map X∈R B×C×H×W First, the channel is expanded by 1×1 convolution, and then it is divided into multiple parallel branches. Each branch uses a cascaded D-Bottleneck module (such as Figure 4 The feature processing is performed by using the dilation rate of each branch in an arithmetic sequence (e.g., the default is d = 1, 2, 3). Each D-Bottleneck module consists of two 3×3 dilated convolutions, whose dilation rate is dynamically adjustable, capturing local details and global context information through differentiated receptive fields. For example, when using dilated convolution with d = 2, the receptive field of a single layer is expanded to 5×5 (e.g., Figure 5 After three D-Bottleneck modules, the cumulative receptive field can reach 16 × 16. The output feature maps of each branch are fused through channel splicing to form enhanced features containing multi-scale information.

[0085] In the multi-branch dilated convolution structure, the present invention uses multiple parallel convolution branches, each with a different dilation rate, to extract features of different scales:

[0086] Branch 1: uses standard 3×3 convolution (d=1) to maintain the original resolution information;

[0087] Branch 2: 3×3 dilated convolution (d=3) is used to capture medium-scale target features;

[0088] Branch 3: Use 3×3 dilated convolution (d=5) to expand the receptive field and enhance the representation ability of large-scale objects.

[0089] After the dilated convolution process, the features of each branch are finally spliced along the channel dimension to form a fused feature, which is further input into the subsequent attention mechanism to improve the feature expression ability. The calculation formula is as follows:

[0090] Y = Conv 3×3 (B k,dilation=d)

[0091] B k+1 =Conv 3×3 (Y, dilation = d)

[0092] Among them, Y represents the intermediate variable, B k represents the input feature map of the kth D-Bottleneck module, B k+1 represents the output feature map after two dilated convolution operations. d represents the dilation rate of the dilated convolution, k is the number of convolution layers, and the effective receptive field gradually expands under different dilation rates for a convolution kernel size of 3×3. When the dilation rate d = 2, the effective receptive field expands to (d+1)(k-1)+1, that is, a single layer of dilated convolution can increase the receptive field from 3×3 to 5×5.

[0093] Furthermore, to improve feature selection, the output features of all branches are concatenated in the channel dimension and then subjected to a 1×1 compressed convolution to restore the number of channels. This strategy not only reduces computational effort but also enhances the expressive power of multi-scale features.

[0094] Compared with traditional single-scale convolution, the combination of different expansion rates enables the model to focus on the local details of small targets and the overall structure of large targets at the same time; maintain the resolution of feature maps, reduce information loss caused by pooling or strided convolution; adapt to complex background interference, and avoid target information loss due to scale mismatch.

[0095] Step S200: fusing multi-scale features with spatial and channel dual-path attention to generate enhanced features;

[0096] In remote sensing target detection tasks, complex backgrounds and the coexistence of multi-scale targets can lead to feature confusion, making small target information easily overwhelmed by larger targets or background noise. To address this issue, this study introduced a Bilateral Attention Feature Mixer (BAFM) in the feature fusion process of YOLO-MGA. This mechanism combines spatial attention (SA) and channel attention (CA) to enhance target features while suppressing background interference, improving the model's detection accuracy.

[0097] Figure 6This paper demonstrates the structure of a bilateral attention feature mixer, which achieves adaptive feature fusion through the collaboration of local and global attention pathways. The input feature map X and the residual feature map Y are first combined through a concatenation operation and then fed into the local and global attention pathways, respectively. The local pathway generates spatial attention weights through convolution (Conv), batch normalization (BN), ReLU activation, and a sigmoid function to highlight local detail features. The global pathway combines global average pooling (GAP), convolution, and ReLU activation to extract channel-level global contextual information and generates channel-wise attention weights through sigmoid. The two weights are then scaled and added together, and finally the original features are dynamically calibrated through an addition operation, achieving a refined fusion of local details and global semantics, enhancing the feature response of the target area.

[0098] Specifically, BAFM receives two input features:

[0099] Original input features X∈R B×C×H×W : Input features from the shallow layers of the network retain high-resolution local details (such as edge textures of small objects);

[0100] Residual feature Y∈R B×C×H×W : The multi-scale fusion features after the multi-branch dilated convolution processing in step S100 encode cross-region contextual semantic information.

[0101] First, perform element-by-element addition to generate the baseline feature matrix:

[0102] Z=X+Y

[0103] This operation fuses original details with multi-scale semantic information to form a feature base containing common features.

[0104] The dual-path attention weight calculation includes three parts: local attention path, global attention path, and hybrid attention map fusion, where:

[0105] The local attention path focuses on enhancing local features in the spatial dimension and generates spatial attention weights through four stages of processing:

[0106] Channel compression: Use 1×1 convolution to compress the number of channels of Z from C to C / r (default r = 4) to reduce redundant channel interference.

[0107] Non-linear activation: Introducing spatial non-linearity through the ReLU function to enhance the expressive power of local patterns.

[0108] Channel reconstruction: Restore the original number of channels C through 1×1 convolution to preserve the integrity of local features.

[0109] Spatial attention generation: Input the processing results into the Sigmoid function to generate the spatial attention weight matrix A local ∈R 1×H×W , the formula is:

[0110] A local =σ(Conv 1×1 (δ(Conv 1×1 (Z))))

[0111] Among them, A local is the spatial attention weight matrix, σ is the Sigmoid function, δ is the ReLU activation function, Conv 1×1 Refers to the 1x1 convolution operation, which is used for dimensionality reduction or dimensionality increase, and Z is the reference feature matrix.

[0112] The global attention path generates channel attention weights by modeling the channel-level global context. The process is as follows:

[0113] Global feature compression: Adaptive average pooling is performed on Z to generate a channel-level global description vector Z∈R C×1×1 , the calculation formula is:

[0114]

[0115] Among them, Z c represents the global average pooling value of the cth channel; H and W represent the height and width of the input feature map; h and w represent the spatial coordinate index (row, column) of the feature map; Z c,h,w Represents the pixel value of the cth channel, hth row, and wth column of the input feature map.

[0116] Dynamic channel interaction: A bottleneck structure is constructed by two fully connected layers (with the middle dimension of max(C / 16,4)) to capture nonlinear dependencies across channels:

[0117] q=W2·δ(W1·z)

[0118] Among them, q represents the intermediate feature vector after mapping through the fully connected layer, which is used to generate the channel attention weight, W1∈R max(C / 16,4)×C , W2∈R C×max(C / 16,4) is a learnable parameter and δ is the ReLU activation function.

[0119] Channel attention generation: Input the result into the Sigmoid function to generate the channel attention weight vector A global ∈R C ×1×1 .

[0120] The hybrid attention map fusion adds the local and global attention weights and normalizes them to generate the final hybrid attention map:

[0121] A fusion =σ(A local +A global )

[0122] Among them, A fusion A represents the mixed attention weight after adding local and global attention and normalized by Sigmoid; local Represents the local spatial attention weight matrix (focus on details); A global Represents the global channel attention weight vector (focusing on the overall semantics).

[0123] Based on the final mixed attention map, the original features and residual features are dynamically weighted fused:

[0124]

[0125] Among them, F enhanced Represents the enhanced feature map after mixed attention weighting, which combines the key information of the original features and residual features. This formula implements adaptive feature selection: high-response regions (such as small objects) tend to retain multi-scale semantic features (Y-dominated), while low-response regions (such as background) enhance original details (X-dominated).

[0126] Step S300: Input the enhanced features obtained last time in the backbone network into the cascade pooling module to perform multi-scale pyramid construction to generate multi-level reconstruction features;

[0127] This step constructs a multi-scale feature pyramid through a cascaded pooling structure and a channel-space dual reconstruction mechanism. Its core implementation method is as follows: Figure 7 The GMAP module architecture shown, Figure 7 The GMAP module is divided into two parts, the left side is the main processing flow, and the right side is the expanded details of the key submodules. The data flow starts from the input feature map (Feature Map) at the top of the left side and goes through the following stages in sequence:

[0128] Multi-scale pooling generates a four-level feature pyramid (Feature map_0 to Feature map_3).

[0129] Channel reconstruction and spatial reconstruction enhance the features of each scale respectively.

[0130] Gated fusion dynamically integrates multi-scale features and finally outputs the fused features through scaling.

[0131] Among them, in the multi-scale pooling stage, the input feature map is sent to the cascaded maximum pooling layer after the number of channels is adjusted by 1×1 convolution. Feature map_0: retains the original resolution and focuses on local details (such as the edges of small targets). Feature map_1~3: gradually increases the pooling kernel (such as 5×5, 9×9, etc.) to expand the receptive field to capture a wide range of semantics (such as the entire building); its purpose is to construct a feature pyramid covering different scales to adapt to the extreme size differences of remote sensing targets.

[0132] During the channel reconstruction phase, the feature maps at each scale are divided into multiple groups (e.g., four groups) along the channel dimension. Grouped convolution independently extracts local correlations for each feature group, reducing spectral redundancy (e.g., confusion between vegetation and vehicles). Channel attention generates channel statistics through global average pooling, dynamically adjusting channel weights using fully connected layers with ReLU / Sigmoid activation to enhance target-related channels (e.g., metallic reflective features of ships).

[0133] Spatial Reconstruction: Average and maximum pooling are performed on the channel-reconstructed features to generate a two-way spatial statistical map. After concatenating the two statistical maps, a 7×7 convolution is performed to capture local spatial dependencies. Sigmoid activation is used to generate a spatial weight map to highlight the target area (such as the vehicle outline). The spatial weight is multiplied element-wise with the input features to enhance the spatial localization of the target.

[0134] Gated fusion: Global average pooling is performed on features at each scale to initially extract scale importance. Dynamic weights are learned through fully connected layers and ReLU / Sigmoid activation, and finally softmax normalization is performed to ensure that the weight sum is 1. Weighted summation fuses multi-scale features according to weights, balancing details (Feature map_0) and semantics (Feature map_3). This avoids weight bias in fixed pyramids and adapts to the multi-scale requirements of remote sensing scenarios.

[0135] The specific process is as follows:

[0136] Input feature map F in ∈R B×C×H×W After three cascaded maximum pooling operations, a feature pyramid with four levels {F0, F1, F2, F3} is generated, where:

[0137] Basic feature F0: original input feature, retaining the highest resolution (H×W), used to capture small target details.

[0138] Multi-level pooling features F1, F2, and F3: By gradually increasing the pooling kernel size (default 5×5), the receptive field is gradually expanded.

[0139] First pooling: kernel size k = 5, step size s = 1, output feature map size

[0140] Second pooling: kernel size k = 5, step size s = 1, and the output size is further reduced.

[0141] The third pooling: Same as above, generating the highest level feature F3.

[0142] The receptive field between pooling levels expands exponentially (e.g., when the base receptive field r = 5×5, the sequence is 5×5→9×9→13×13→17×17), covering multi-scale contexts from local details to global semantics. This is particularly suitable for the characteristics of remote sensing images with significant differences in target size (such as the scale span of ships and airports in the DOTA dataset).

[0143] In the channel-space dual reconstruction mechanism, for each level of feature F i (i=0,1,2,3) performs channel reconstruction and spatial reconstruction in sequence to suppress complex background interference and enhance target saliency.

[0144] Channel reconstruction includes feature grouping, depth convolution and spatial reconstruction; feature grouping and depth convolution include:

[0145] F i Divided into two groups along the channel dimension Each group performs group depth convolution separately:

[0146]

[0147] in, GConv(F) represents the output of the kth group of features at the i-th level after group depth convolution. G represents the number of groups of group convolution (4 groups by default). Each group learns local features independently, reducing the amount of calculation while avoiding information confusion between channels. ik ) represents performing grouped depth-wise separable convolution on the kth group of features.

[0148] A Squeeze-and-Excitation (SE) module is introduced for each set of features to dynamically adjust the channel weights:

[0149] Global feature compression: Generate channel description vector z through global average pooling k ∈R C / 2 :

[0150]

[0151] Among them, z k,c H represents the global average pooling value of the cth channel of the kth group feature; i and W iRepresent the height and width of the i-th level feature map respectively; Represents the pixel value of the cth channel, hth row, and wth column of the kth feature group.

[0152] Channel weight generation: Calculate the channel attention weight s through two layers of fully connected layers (intermediate dimension max(C / 16,4)) k ∈R C / 2 :

[0153] s k =σ(W2·δ(W1·z k ))

[0154] Among them, s k is the channel attention weight, generated by the SE module, σ is the Sigmoid function, δ is the ReLU activation function, z k Represents the channel-level global description vector of the k-th group of features.

[0155] Channel weighted output: weight s k ∈R C / 2 Multiply the original features channel by channel to obtain channel reconstruction features

[0156] Splice the two sets of channel reconstruction features along the channel dimension to generate channel enhancement features

[0157] Spatial reconstruction includes three parts: cross-channel aggregation, spatial weighted output and multi-level reconstruction feature generation. Among them, cross-channel aggregation: i channel Perform cross-channel feature aggregation to generate a spatial attention map:

[0158] Average and maximum pooling: Calculate the mean along the channel dimension separately With the maximum value

[0159]

[0160] Among them, A avg (h,w) represents the mean of all channels at position (h,w); C represents the total number of channels in the feature map; A represents the pixel value of the cth channel, hth row, and wth column of the reconstructed feature of the i-th level channel; max (h,w) represents the maximum value of all channels at position (h,w).

[0161] Feature concatenation and convolution: A avg With A max Splicing along the channel dimension and generating spatial attention weights through 7×7 convolution

[0162]

[0163] Among them, Conv 7×7 Represents a 7×7 convolution operation, which is used to fuse the mean and maximum features to generate spatial attention weights.

[0164] Spatial weighted output: Multiply the spatial attention map with the channel reconstruction feature position by position to obtain the final reconstruction feature:

[0165] F i recon =A⊙F i channel

[0166] Among them, F i recon represents the feature map of the i-th level after spatial reconstruction; F i channel Represents the feature map of the i-th level after channel reconstruction; ⊙ represents element-by-element multiplication (Hadamard product).

[0167] Multi-level reconstruction feature generation includes: for each level F i After executing the above double reconstruction operation, four sets of reconstruction features {F0 recon ,F1 recon ,F2 recon ,F3 recon All features are upsampled to the original input size H×W by bilinear interpolation to ensure scale consistency for subsequent fusion.

[0168] Step S400: Input the multi-level reconstruction features into a gated fusion network, dynamically calculate the weight coefficients of the reconstruction features at each level based on learnable parameters, and perform weighted fusion using the calculated weight coefficients to generate multi-scale fusion features;

[0169] This step realizes the adaptive fusion of multi-level reconstruction features through a learnable gating mechanism. Its implementation method corresponds to Figure 7 The gated fusion unit of the GMAP module in the , the specific process is as follows:

[0170] For each level of reconstruction feature F i recon ∈R B×C×H×W (i=0,1,2,3), calculate the dynamic fusion weight according to the following steps:

[0171] The channel-level global description vector e of each feature is extracted by global average pooling i ∈R C :

[0172]

[0173] Among them, e i Represents the channel-level global description vector of the i-th level feature; e i,c represents the global average pooling value of the cth channel at the i-th level;

[0174] e i Input fully connected layer (parameter W g ∈R 4×C , bias b g ∈R 4 ) generates a preliminary gating signal and normalizes it into a weight coefficient through the Sigmoid function:

[0175] g i =σ(W g ·e i +b g )

[0176] Among them, σ is the Sigmoid function, g i represents the initial gating weight of the i-th level feature (i = 0, 1, 2, 3), which is generated by the fully connected layer and the Sigmoid function, reflecting the importance of the level feature, e i Represents the channel-level global description vector of the i-th level feature, W g and b g These are all parameters that can be learned by the network.

[0177] Perform Softmax normalization on the weights of the four layers to ensure that the sum of the weights is 1:

[0178]

[0179] Among them, w i represents the normalized fusion weight of the i-th level feature, exp represents the exponential function, which is used for numerical scaling before Softmax normalization; g j Represents the initial gating weight of the j-th level feature (j = 0, 1, 2, 3), used for denominator summation calculation, i and j are both level indices with the same value range (0 to 3).

[0180] Based on the normalized weights {w0, w1, w2, w3}, the four sets of reconstruction features are weighted summed to generate the final multi-scale fusion feature F fusion ∈R B×C×H×W :

[0181]

[0182] Among them, w i Represents the normalized fusion weight of the i-th level feature;

[0183] The following is an example of dynamic weight allocation to explain:

[0184] Small target dominance (F0 recon : If the input image contains dense small targets (such as a group of vehicles), the gated network will automatically increase the weight of w0 to enhance the high-resolution detail features.

[0185] Big goal dominance (F3 recon : If the scene is dominated by large objects (such as airport runways), the weight of w3 increases significantly, focusing on global semantic information.

[0186] It is understandable that the learnable gating mechanism is implemented by the fully connected layer parameters W g 、b g It autonomously learns the importance of features at each scale in different scenarios, avoiding the limitations of manually setting fixed fusion ratios. Softmax normalization constrains the sum of weight coefficients to 1, achieving a dynamic balance of feature contributions and preventing over-dominance of features at a single level. In multi-scale complementarity, high-resolution features (F0) preserve small object edges, while low-resolution features (F3) encode large object structures. The gating mechanism adaptively integrates these features based on the target distribution.

[0187] Step S500: input the multi-scale fusion features into an asymmetric decomposition convolution layer, and use horizontal and vertical convolution kernels to separate and process the spatial dimensions to obtain down-sampled features;

[0188] This step processes multi-scale fusion features through decomposition-type depth-separable convolution (Asymmetric Depthwise SeparableConvolution) to enhance the representation ability of rotating targets. Its core implementation method corresponds to Figure 8 The ADDown module in the code is as follows:

[0189] Input feature map F fusion ∈R B×C×H×W First, the number of channels is adjusted to the target number through 1×1 convolution to generate the channel alignment feature F′∈R B×C′×H×W (Default C′=2C) Then, the following asymmetric convolution operation is performed:

[0190] Perform horizontal depth convolution on the channel alignment features: use a depthwise separable convolution kernel with a size of 3×1 and a step size of 2 to extract horizontal edge features and generate a horizontal feature map:

[0191] F h =DepthwiseConv 3×1 (F′)

[0192] Among them, F h ″ represents the output feature of the 3×1 depth-wise separable convolution in the horizontal direction, and the output feature map size is DepthwiseConv 3×1 represents depth-wise separable convolution, which only performs independent convolution on each input channel; F′ represents the channel-aligned feature after 1×1 convolution adjustment.

[0193] Perform vertical depth convolution on the horizontal feature map: use a depthwise separable convolution kernel with a size of 1×3 and a step size of 2 to extract vertical edge features and generate a vertical feature map:

[0194] F v =DepthwiseConv 1×3 (F h ″)

[0195] Among them, F v ″ represents the output feature of the 1×3 depth-separable convolution in the vertical direction; the final output feature map X v ″∈R B ×C′×H′×W′ .

[0196] The horizontal feature map is added to the vertical feature map element by element to form the downsampled features.

[0197] The square structure of the traditional 3×3 convolution kernel has a decaying intensity in response to detail targets with angle θ (following the cosθ law), while the asymmetric decomposition convolution achieves rotation invariance optimization by independently learning the weights in the horizontal and vertical directions:

[0198] Horizontal convolution kernel W h : Sensitive to vertical edges (such as ship side walls, vehicle longitudinal axes).

[0199] Vertical convolution kernel W v : Sensitive to horizontal edges (such as building roofs and airport runways).

[0200] When the target rotates by angle θ, the characteristic response intensity is W h cosθ+W v The sinθ dynamic combination avoids the fixed geometric sensitivity attenuation of traditional convolution.

[0201] For the downsampled feature F v ″∈R B×C′×H′×W′ Perform dynamic channel attention to prevent information loss during downsampling:

[0202] Global feature compression: Generate channel description vector z∈R through global average pooling C′ :

[0203]

[0204] Among them, H′ and W′ represent the height and width of the downsampled feature map respectively; F″ v,c,h,wRepresents the pixel value of the cth channel, hth row, and wth column after vertical convolution.

[0205] Adaptive compression ratio attention: build a dynamic bottleneck structure and adaptively adjust the middle layer dimension d = max(C′ / 16,4) (when C′ < 64, force d = 4):

[0206] q=W2·δ(W1·z)

[0207] W1∈R d×C′

[0208] W2∈R C′×d

[0209] Among them, q represents the intermediate feature vector output by the dynamic bottleneck layer; W1 and W2 represent the weight matrices of the fully connected layer; z represents the channel description vector after global average pooling; δ is the ReLU activation function to avoid excessive compression at low channel numbers.

[0210] Channel weight generation: Generate channel attention weight A∈R through Sigmoid function C′×1×1 , and multiply it channel by channel with the input features:

[0211] A=σ(q)

[0212] F out =A⊙F v ″

[0213] Among them, A represents the channel attention weight matrix of Sigmoid normalization; F out represents the output characteristics after channel calibration; F v ″ represents the output feature of the 1×3 depthwise separable convolution in the vertical direction.

[0214] Step S600: Perform dynamic channel attention calibration on the downsampled features, calculate the channel weight matrix through the adaptive compression ratio mechanism, and generate channel enhancement features;

[0215] This step uses the dynamic compression ratio channel attention mechanism to adaptively enhance the channel dimension of the downsampled features output in step S500. Its core implementation method corresponds to Figure 8 The dynamic channel attention unit of the ADDown module in the figure has the following specific process:

[0216] Input feature map

[0217] F out ∈R B×C′×H′×W′ (Asymmetric convolution output from step S500) Channel attention calibration is achieved by the following steps:

[0218] (1) Global spatial information compression

[0219] Perform adaptive average pooling on each channel to generate a channel-level global description vector R C′ , the calculation formula is:

[0220]

[0221] Among them, F out,c,h,w Indicates the pixel value of the cth channel, hth row, and wth column of the feature map output in step S500;

[0222] This vector encodes the spatial response strength of each channel, such as the vertical edge channel of a ship target (activated by horizontal convolution) or the horizontal edge channel of a building target (activated by vertical convolution).

[0223] (2) Adaptive compression ratio bottleneck structure

[0224] Design a fully connected layer with a dynamic compression ratio r, and adaptively adjust the intermediate layer dimension according to the number of input channels C′:

[0225] Dynamic calculation of compression ratio:

[0226]

[0227] Where r represents the adaptive compression ratio (intermediate layer dimension); C′ represents the number of channels of the input feature map; when C′ < 64, r = 4 is forced to avoid excessive compression at low channel numbers and resulting in information loss.

[0228] Nonlinear mapping and reconstruction:

[0229] Through two layers of fully connected layers (parameters W1∈R r×C′ , W2∈R C′×r )Build the bottleneck structure:

[0230] q=W2·δ(W1·z)

[0231] Among them, q represents the intermediate feature vector after mapping by the fully connected layer, and δ is the ReLU activation function, which introduces nonlinear interaction.

[0232] (3) Channel weight generation

[0233] The reconstructed feature vector q∈R C′ Input Sigmoid function to generate channel attention weight A∈R C′×1×1 :

[0234] A=σ(q)

[0235] The weight matrix reflects the importance of each channel to rotating target detection, such as enhancing the ship vertical edge channel or suppressing the cloud interference channel.

[0236] Combine the channel attention weight A with the input feature F out Multiply channel by channel to generate the calibrated channel enhancement feature F enhanced ∈R B×C′×H′×W′ :

[0237] F enhanced,c,h,w =A c ·F out,c,h,w

[0238] Among them, F enhanced,c,h,w A represents the pixel value of the cth channel, hth row, and wth column after channel calibration; c Represents the attention weight value of the c-th channel.

[0239] This operation strengthens key channels (such as edge features of rotating targets) and weakens redundant channels (such as background noise) through dynamic weighting.

[0240] It's understandable that dynamically adjusting the compression ratio based on the number of channels C' avoids the over-compression that occurs with traditional fixed compression ratios (e.g., r = 16 in the SE module) at low channel counts (e.g., only two channels in the middle layer when C' = 32), which can lead to information loss. The feature channels in the asymmetric convolution output naturally encode rotation-sensitive information (horizontal and vertical edges), and the dynamic attention mechanism further selects channels that are sensitive to the current target orientation. For example, when the input is a tilted ship, the weights of the horizontal and vertical edge channels are dynamically adjusted to match its rotation angle. This dynamic compression ratio mechanism reduces the number of parameters (compared to a fixed compression ratio).

[0241] Step S700: Input the enhanced features obtained by the backbone network processing and the channel enhanced features after the neck network processing into the YOLOv10 detection head, and output the target bounding box and category prediction results.

[0242] In the remote sensing target detection task, the final detection head network is responsible for predicting the target’s bounding box and category based on the channel enhancement features extracted and enhanced by multiple layers of features. Figure 2 As can be seen, the input features include two enhanced features obtained from the backbone network and one channel-enhanced feature obtained from the neck network. First, the input feature map is passed through a regression network to regress the target bounding box, predicting the target's center position (x, y), width (w), and height (h). These bounding box predictions are normalized to a reasonable range using a sigmoid activation function. Ultimately, the calculated bounding box coordinates represent the target's location, providing candidate boxes for the subsequent non-maximum suppression (NMS) step.

[0243] Secondly, the detection head is responsible for object classification. In this step, each region of the input feature map is processed by the classification network, outputting the probability of the object belonging to each category. Using the Softmax activation function, the model calculates the probability of each object belonging to different categories and selects the category with the highest probability as the prediction result. Finally, combining the bounding box and category predictions, the detection head outputs complete object detection information, providing accurate positioning and classification results for subsequent object detection processing and decision-making, thus completing the object detection task.

[0244] According to another aspect of an embodiment of the present application, an electronic device is provided, including a processor and a memory, wherein the processor is configured to implement the steps of the method when executing a computer program stored in the memory.

[0245] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0246] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0247] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0248] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.

[0249] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A remote sensing target detection method, characterized in that: The steps include: The input feature map is fed into the backbone network for feature extraction. During the feature extraction process, the feature tensor is fed into a multi-branch dilated convolution structure, and multi-scale features are extracted through convolution kernels with different parallel dilation rates. The multi-scale features are fused with spatial and channel dual-path attention to generate enhanced features; Input the enhanced features obtained in the backbone network for the last time into the cascade pooling module for multi-scale pyramid construction to generate multi-level reconstruction features; Input the multi-level reconstruction features into the gated fusion network, dynamically calculate the weight coefficients of the reconstruction features of each level based on the learnable parameters, and perform weighted fusion through the calculated weight coefficients to generate multi-scale fusion features; The multi-scale fusion features are input into the asymmetric decomposition convolution layer, and the horizontal and vertical convolution kernels are used to separate and process the spatial dimensions to obtain down-sampled features; Perform dynamic channel attention calibration on the downsampled features, calculate the channel weight matrix through the adaptive compression ratio mechanism, and generate channel enhanced features; The enhanced features obtained by the backbone network processing and the channel enhanced features after the neck network processing are input into the YOLOv10 detection head, and the target bounding box and category prediction results are output.

2. The remote sensing target detection method according to claim 1, wherein: The multi-branch dilated convolution structure includes at least three parallel branches, the dilation rates of each parallel branch are 1, 3, and 5 respectively, and each parallel branch includes a cascaded D-Bottleneck module, which adds the output of the dilated convolution to the input features through a residual connection.

3. The remote sensing target detection method according to claim 1, wherein: The fusion process of spatial and channel dual attention includes: The residual features output by the multi-branch dilated convolution are added element-by-element to the original input features to generate a baseline feature matrix; Perform local attention calculation on the benchmark feature matrix: compress the number of feature channels through channel compression convolution, reconstruct it to the original channel dimension after ReLU activation, and then generate the spatial attention weight matrix through the Sigmoid function; Perform global attention calculation on the baseline feature matrix: extract channel-level global description through adaptive average pooling, reconstruct it to the original channel dimension after dimensionality reduction by fully connected layer and ReLU activation, and then generate channel attention weight matrix through Sigmoid function; Add the spatial attention weight matrix and the channel attention weight matrix, and generate a mixed attention map through Sigmoid normalization; The hybrid attention map is used to perform weighted fusion of the original input features and the residual features.

4. The remote sensing target detection method according to claim 1, wherein: The method of inputting the enhanced features obtained in the backbone network for the last time into the cascade pooling module for multi-scale pyramid construction to generate multi-level reconstruction features includes: Performing three cascaded maximum pooling operations on the input enhanced features, respectively using a pooling kernel with a stride of 1 and a kernel size of 5×5, to generate a multi-scale pyramid of three-level down-sampled features; Perform channel reconstruction on each pooled feature in the multi-scale pyramid to generate channel reconstruction features; Perform spatial reconstruction on the channel reconstruction features to generate a spatial attention map; The spatial attention map is activated by Sigmoid and multiplied position by position with the channel reconstruction feature to output multi-level reconstruction features.

5. The remote sensing target detection method according to claim 4, wherein: The method of performing channel reconstruction on each feature map in the multi-scale pyramid and generating channel reconstruction features includes: Each pooled feature in the multi-scale pyramid is divided into four groups of sub-features along the channel dimension; Perform 3×3 grouped depth convolution on each group of sub-features, and keep the output feature dimension the same as the input; Perform global average pooling on the output features of the grouped depth convolution to generate a channel description vector; The channel description vector is input into the fully connected layer and reduced to C / 16 dimensions. After ReLU activation, it is reconstructed to the original channel dimension to generate the channel attention weight. Multiply the channel attention weights by the output features of the depth convolution channel by channel to obtain the sub-features after channel reconstruction; The four groups of sub-features are spliced along the channel dimension to generate channel reconstruction features; The method of spatially reconstructing the channel reconstruction features and generating a spatial attention map includes: Calculate the average response map and maximum response map along the channel dimension for the channel reconstruction features; The average response map and the maximum response map are concatenated and input into a 7×7 convolutional layer to generate a spatial attention map.

6. The remote sensing target detection method according to claim 1, wherein: The weight coefficient calculation method of the gated fusion network is: Performing global average pooling on each level feature in the multi-level reconstruction features to generate a preliminary weight vector; The preliminary weight vector is input into a fully connected layer consisting of a learnable parameter matrix and bias to generate a gating signal; The gating signal is activated by the Sigmoid function to generate the normalized gating weight, where the calculation method of the normalized gating weight is: g i =σ(W g ·e i +b g ) Among them, σ is the Sigmoid function, g i It represents the initial gating weight of the i-th level feature, which is generated by the fully connected layer and the Sigmoid function, reflecting the importance of the level feature. i Represents the channel-level global description vector of the i-th level feature, W g and b g These are all parameters that can be learned by the network; Input the normalized gated weights into the Softmax function for cross-layer normalization and calculate the weight coefficients: Among them, w i represents the normalized fusion weight of the i-th level feature, exp represents the exponential function, which is used for numerical scaling before Softmax normalization; g j Represents the initial gating weight of the j-th level feature, used for denominator summation calculation, and i and j are both level indices.

7. The remote sensing target detection method according to claim 1, wherein: The calculation method for downsampling of the asymmetric decomposition convolution layer is: Performing a 1×1 convolution operation on the multi-scale fusion features, adjusting the channel dimension to the target number of channels, and generating channel alignment features; Perform horizontal decomposition convolution on the channel alignment features: use a depthwise separable convolution kernel with a size of 3×1 and a step size of 2 to extract horizontal edge features and generate a horizontal feature map; Performing vertical decomposition convolution on the horizontal feature map: using a depthwise separable convolution kernel with a size of 1×3 and a step size of 2 to extract vertical edge features and generate a vertical feature map; The horizontal feature map is added to the vertical feature map element by element to obtain the downsampled features.

8. The remote sensing target detection method according to claim 1, wherein: The method of performing dynamic channel attention calibration on the downsampled features, calculating the channel weight matrix through an adaptive compression ratio mechanism, and generating channel enhancement features includes: Performing a global average pooling operation on the downsampled features to generate a global statistical vector of channel dimension; Inputting the global statistical vector into a compression-excitation module consisting of a two-layer fully connected network for compression and excitation, wherein the first layer compresses the number of channels and the second layer restores the number of channels to the original number; Perform ReLU activation and Sigmoid transformation on the compressed and stimulated global statistical vector in sequence to generate the channel attention weight matrix; The channel attention weight matrix is multiplied by the downsampled features channel by channel to generate channel enhanced features.

9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the remote sensing target detection method according to any one of claims 1 to 8, and the processor is configured to execute the program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the remote sensing target detection method according to any one of claims 1 to 8 are executed.

Citation Information

Cited By

  • Lightweight target detection method

    CN120807955A

  • Insulator ultraviolet corona discharge target detection method based on YOLO-SM

    CN120912997A

  • Target tracking and counting method and system for electric power fittings

    CN121033106A

  • Dynamic multi-scale convolution and cross-level attention feature fusion method and device

    CN121053505A

  • Power transmission line iron tower detection method and device based on remote sensing image

    CN121121515A