Novel lightweight remote sensing weak and small target detection system

Through the collaborative optimization of DSSM, AMFE and EDAM modules, the problem of feature drowning and computational redundancy in remote sensing small target detection is solved, and efficient and accurate remote sensing small target detection is achieved.

CN120747729APending Publication Date: 2025-10-03ROCKET FORCE UNIV OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510757193.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional remote sensing small target detection methods are difficult to achieve accurate detection due to factors such as low image resolution, small target size, occlusion and background interference, and have serious computational redundancy, resulting in low detection accuracy.

Method used

The DSSM feature extraction module, AMFE feature enhancement module and EDAM two-stage routing processing module are used, combined with the CSP network and adaptive large-core attention mechanism to optimize feature extraction and calculation and enhance the small target detection capability.

Benefits of technology

It significantly improves the accuracy and computational efficiency of remote sensing small target detection, and can adaptively enhance small target feature representation, suppress background interference, and reduce redundant calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747729A_ABST
    Figure CN120747729A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image detection, in particular to a novel lightweight remote sensing weak and small target detection system which takes YOLO11 as a basic framework and comprises a DSSM feature extraction module, an AMFE feature enhancement module and an EDAM two-stage routing processing module. The DSSM feature extraction module is used for enhancing long-distance spatial dependency capture, the AMFE feature enhancement module is used for realizing multi-scale feature adaptive receptive field adjustment, and the EDAM two-stage routing processing module is used for reducing calculation complexity and maintaining feature expression capability at the same time. According to the method, three core modules of dynamic space state modeling (DSSM), adaptive large kernel attention processing (AMFE) and two-stage routing mechanism (EDAM) are fused, long-range dependence capture, multi-scale feature enhancement and lightweight calculation collaborative optimization are realized in a YOLO11 framework, and the technical bottlenecks of feature submerging, background interference and calculation redundancy in traditional remote sensing small target detection are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image detection, and in particular to a novel lightweight remote sensing dim target detection system. Background Art

[0002] With the rapid development of remote sensing technology, high-resolution remote sensing images have been widely used in many fields, including urban monitoring, aerospace reconnaissance, agricultural management, and disaster assessment. However, remote sensing small target detection, as a branch of target detection, remains a major research challenge due to the influence of multiple factors such as image resolution, target size, number, and orientation.

[0003] Traditional remote sensing small target detection methods are mainly based on manually designed feature extraction techniques, such as edge detection, texture analysis, and shape description. For example, the paper "Ship Detection Method in Remote Sensing Images Based on HOG Features and SVM Classifier" (State Key Laboratory of Shipping Technology and Safety, Shanghai Institute of Ship and Transportation Sciences, Shanghai, 200135) mentions a method for target recognition using template matching, sliding window, and feature clustering techniques, which has achieved certain results in a simple ocean background.

[0004] However, despite the maturity of target detection methods, they still face some difficulties and challenges. First, small targets in remote sensing images contain limited spatial and contextual information, making accurate detection difficult. When the targets are sparsely distributed and small in size, traditional detection methods find it difficult to establish effective long-range spatial dependencies. Secondly, problems such as low resolution, occlusion, background interference, and class imbalance further complicate the detection task. Finally, the core challenges in small target detection focus on insufficient information, inaccurate positioning, and scarcity of positive samples. In deep networks, small target features are easily overwhelmed by background information, seriously affecting detection accuracy. Therefore, in response to the problems existing in remote sensing small target detection, it is necessary to propose a new lightweight remote sensing weak small target detection system, which is optimized in many aspects based on the YOLO11 architecture to solve the technical bottlenecks of feature overwhelm, background interference, and computational redundancy in traditional remote sensing small target detection. Summary of the Invention

[0005] To solve the above problems, the present invention provides a novel lightweight remote sensing small target detection system. By integrating three modules: dynamic spatial state modeling (DSSM), adaptive large kernel attention processing (AMFE) and dual-stage routing mechanism (EDAM), the system realizes long-range dependency capture, multi-scale feature enhancement and lightweight computing collaborative optimization within the YOLO11 framework, thus solving the technical bottlenecks of feature drowning, background interference and computational redundancy in traditional remote sensing small target detection.

[0006] In order to achieve the above object, the technical solution of the present invention is as follows: a novel lightweight remote sensing small target detection system, comprising a DSSM feature extraction module, an AMFE feature enhancement module and an EDAM two-stage routing processing module;

[0007] The DSSM feature extraction module adopts the concept of CSP network and embeds SSBlock operation to identify the input image and extract global features. The DSSM feature extraction module consists of several SSB modules connected in series. When a feature is input, the input feature X is first channel-adjusted through 1×1 convolution and divided into X1 and X2. The main branch X2 is then further processed by the series of SSB modules. Finally, all processed features are cascaded and fused with the initial branch.

[0008] The AMFE feature enhancement module is used to perform multi-scale perception of the global features extracted by DSSM through an adaptive large-core attention mechanism, perform multi-scale enhancement and noise suppression, and use separable convolution to extract contextual information at different scales. Subsequently, dynamic weight allocation is used to enhance the feature response of small target areas while suppressing background noise interference. Finally, multi-level features are fused to enhance the model's ability to discriminate complex scenes, highlight the local features of small targets, and generate a multi-level fused feature map;

[0009] The EDAM two-stage routing processing module adopts a dual-path collaborative strategy of regional routing and feature routing. It first screens high-value candidate areas and establishes spatial associations through regional routing, and then dynamically aggregates channel dimensions based on feature routing. While reducing redundant calculations, it strengthens long-distance feature associations, and ultimately realizes small target feature modeling and cross-scale information fusion.

[0010] Furthermore, the mathematical expression of the SSB module is as follows:

[0011]

[0012] Among them, F cv1 represents a k×k convolutional layer (k is usually set to 3), F SSBlock Represents an SSBlock operation.

[0013] Furthermore, the SSBlock operation includes the LN layer normalization operation and the SS2D selective scanning attention operation. Its basic operations are as follows:

[0014] F SSBlock (X) = X + DropPath (F SS2D (F LN (X)))

[0015] Among them, X is the input feature map, F LN Representation layer normalization operation, F SS2DRepresents a 2D selective scan operation.

[0016] Furthermore, in the SS2D selective scanning attention operation, the input feature X is first mapped to a higher-dimensional feature space through a linear projection layer and decomposed into content features X' and gated features Z. The content features X' are further processed by 2D convolution; the processed features are reorganized into a sequence representation in four directions:

[0017] X seq =[X h ,X w ,X h_flip ,X w_flip ]

[0018] Among them, X h is the characteristic sequence expanded by row, X w is the characteristic sequence expanded by column, X h_flip and X w_flip They are the sequences after horizontal and vertical flipping respectively. For each direction, the SS2D operation is performed to model the sequence. For the input sequence x and the initial state h0, the state update and output are calculated as follows:

[0019]

[0020] Among them, A is the state transfer matrix, B is the input projection, C is the state output projection, and D is the skip connection coefficient.

[0021] Furthermore, the SS2D operation also introduces an adaptive time step mechanism:

[0022]

[0023] Among them, Δ t is the adaptive time step, is a time step control factor derived from the input features; the final processed features are fused with the original features through a gating mechanism and mapped back to the original dimension through the output projection layer:

[0024]

[0025] in, represents element-wise multiplication, and σ is the SiLU activation function; in the process of implementing the gating mechanism, the model can adaptively control the feature flow of each position and channel.

[0026] Furthermore, the AMFE feature enhancement module adopts a cascaded feature extraction and attention enhancement architecture. Given an input feature tensor X, it first undergoes a channel compression transformation X′:

[0027] X′=Τ C (X) = δ(WC ΘX+b C )

[0028] Among them, δ represents the nonlinear activation function, Θ represents the convolution operation, and W C and b C are the learnable convolution kernel parameters and bias vectors respectively; the AMFE feature enhancement module reduces the channel dimension of the input feature to half, and obtains T C ;

[0029] Then, through the progressive maximum pooling operation M p Construct multi-scale feature representation; define the progressive pooling process as:

[0030]

[0031] Among them, Mp(·,k) represents the maximum pooling operation with a kernel size of k×k, a stride of 1, and a padding of k / 2. Finally, the original features are concatenated with the pooled features at all levels in the channel dimension to obtain the attention map of the multi-scale fusion representation.

[0032] Furthermore, the AMFE feature enhancement module also adopts the SKA mapping operation. The SKA attention map A can be expressed as:

[0033] A(Z)=T1(T sv (T sh (T v (T h (Z)))))

[0034] The transformation operators are defined as follows:

[0035]

[0036] In the formula, Θ g represents grouped convolution, Θ g,d W represents the group convolution with dilation rate. h Represents the horizontal convolution kernel, W v Represents the vertical convolution kernel, W sh represents the horizontal convolution kernel with dilation rate d = 2, W sv represents the vertical convolution kernel with dilation rate d=2, and W1 represents the convolution kernel;

[0037] The attention map is applied to multi-scale features through residual connections, and finally, the enhanced features are mapped to the output space through channel fusion transformation.

[0038] Furthermore, the DAM two-stage routing processing module includes a HRABlock component, which includes a DSCA operation and an FFN feedforward network. Mathematically, the overall conversion process of the EDAM two-stage routing processing module is expressed as:

[0039]

[0040] Among them, F HRA Represents the conversion function of the HRABlock component;

[0041] The specific operation of DSCA operation is, first, the input feature X is linearly projected to generate query Q, key K and value V:

[0042] Q,K,V=F proj (X)

[0043] Among them, F proj Represents the convolutional layer, which maps the input features to the QKV space;

[0044] The average pooling operation is used to obtain region-level query and key representations:

[0045]

[0046] Among them, Q r and K r Represents the region-level query Q and key K, F avgpool represents average pooling, Q detach and K detach represents the gradient calculation of the separation;

[0047] Then, the similarity matrix between regions is calculated:

[0048]

[0049] in, and Denote the reshaped region query and key respectively. Based on the similarity matrix, the top-k most relevant regions of each region are selected:

[0050] idx r =TopK(A r ,k)

[0051] Based on the results of region routing, perform pixel-level attention calculation:

[0052]

[0053] Among them, M region The regional routing result idx r The generated sparse attention mask is used to restrict each pixel to interact only with pixels in the relevant region, d k is the dimension of the query vector.

[0054] Furthermore, the FEN feedforward network adopts a two-layer 1×1 convolutional structure for feature conversion and nonlinear modeling.

[0055] Furthermore, the HRABlock component also includes a LEPE supplementary enhancement operation, and the final attention output is obtained by merging the global attention result and the local enhancement part:

[0056]

[0057] The complete processing flow of the HRABlock component combines the attention mechanism and the FEN feedforward network, and processes the information flow X through the residual connection out =X+F FFN (X+F Attn (X)).

[0058] The above scheme has the following beneficial effects:

[0059] The design of the DSSM feature extraction module achieves efficient and comprehensive capture of small remote sensing target features by optimizing and replacing traditional convolution operations. Through multi-directional state space modeling, the DSSM module significantly expands the receptive field of feature extraction, enhances the ability to capture long-range dependencies, and provides richer and more discriminative feature representations for small target detection.

[0060] 2. This solution significantly enhances the model's ability to perceive small targets while maintaining computational efficiency through the design of the AMFE feature enhancement module and by optimizing the feature processing mechanism. The AMFE feature enhancement module adaptively strengthens the feature representation of small target areas while effectively suppressing background interference.

[0061] 3. This solution proposes a lightweight EDAM two-stage routing processing module. The two-stage routing attention mechanism solves the computational bottleneck problem of the traditional attention mechanism when processing large-size feature maps, and establishes more efficient long-distance feature association through a carefully designed regional routing strategy, which is suitable for the detection needs of small remote sensing targets.

[0062] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 This is a structural diagram of the DSSM module in an embodiment of the novel lightweight remote sensing small target detection system of the present invention;

[0064] Figure 2 This is a structural diagram of the SSB module in an embodiment of the novel lightweight remote sensing small target detection system of the present invention;

[0065] Figure 3 This is a structural diagram of the AMFE module in an embodiment of the novel lightweight remote sensing small target detection system of the present invention;

[0066] Figure 4 This is a structural diagram of the EDAM module in an embodiment of the novel lightweight remote sensing small target detection system of the present invention. DETAILED DESCRIPTION

[0067] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0068] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0069] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0070] The following is further described in detail through specific implementation methods:

[0071] Example 1:

[0072] A novel lightweight remote sensing small target detection system, including DSSM feature extraction module, AMFE feature enhancement module and EDAM two-stage routing processing module;

[0073] Specific combination of Figure 1 and Figure 2As shown in the figure, the DSSM feature extraction module adopts the idea of ​​CSP network and embeds SSBlock operation. The DSSM feature extraction module includes several SSB modules connected in series. When the feature is input, the input feature X is first channel-adjusted by 1×1 convolution and divided into X1 and X2. Then the main branch X2 is further processed by the series of SSB modules. Finally, all processed features are cascaded and fused with the initial branch. The mathematical formula of this process is as follows:

[0074] Y=F cv2 (Concat(X1,Y1,Y2,...,Y n ))

[0075] SSB is the core building block of DSSM, which combines the efficiency of the traditional Bottleneck structure and the powerful feature modeling capability of SSBlock. Its mathematical expression is shown below. cv1 represents a k×k convolutional layer (k is usually set to 3), and represents an SSBlock operation.

[0076]

[0077] SSBlock integrates layer normalization (LN) and selective scanning attention mechanism (SS2D). Its basic operation can be expressed by the following formula, where X is the input feature map, F LN Representation layer normalization operation, F SS2D Represents a 2D selective scan operation.

[0078] F SSBlock (X) = X + DropPath (F SS2D (F LN (X)))

[0079] SS2D efficiently captures long-range spatial dependencies by selectively scanning in multiple spatial directions. First, the input feature X is mapped to a higher-dimensional feature space via a linear projection layer and decomposed into content features X′ and gated features Z. The content features X′ are further processed by 2D convolution to enhance local features. The processed features are first reorganized into a sequence representation in four directions:

[0080] X seq =[X h ,X w ,X h_flip ,X w_flip ]

[0081] Among them, X h is the characteristic sequence expanded by row, X w is the characteristic sequence expanded by column, X h_flip and Xw_flip are the sequences after horizontal and vertical flipping, respectively. For each direction, SS2D uses an efficient state space model to model the sequence. Specifically, for the input sequence x and the initial state h0, the state update and output are calculated as follows:

[0082]

[0083] Among them, A is the state transfer matrix, B is the input projection, C is the state output projection, and D is the jump connection coefficient. In order to enhance the expressiveness of the model, SS2D introduces an adaptive time step mechanism:

[0084]

[0085] Among them, Δ t is the adaptive time step, and φ(xt) is the time step control factor derived from the input features. The final processed features are fused with the original features through a gating mechanism and mapped back to the original dimension through the output projection layer:

[0086]

[0087] in, represents element-wise multiplication, and σ is the SiLU activation function. The gating mechanism enables the model to adaptively control the feature flow of each position and channel, further enhancing the expressive power of the model.

[0088] The DSSM feature extraction module successfully addresses the limitations of the traditional C3k2 module in remote sensing small target detection by innovatively combining the CSP network architecture and a selective scanning mechanism. First, through its multi-directional selective scanning mechanism, the module effectively captures long-range spatial dependencies, significantly enhancing its ability to perceive small targets. Second, its adaptive time step mechanism enables the model to dynamically adjust the feature extraction process based on the input content, improving its adaptability to multi-scale targets.

[0089] The AMFE feature enhancement module is used to perform multi-scale perception of input features through an adaptive large-core attention mechanism, extract contextual information at different scales using separable convolution, and then enhance the feature response of small target areas through dynamic weight allocation while suppressing background noise interference. Finally, multi-level features are fused to enhance the model's ability to discriminate complex scenes.

[0090] Specific as Figure 3 As shown in the figure, the AMFE module adopts a cascaded feature extraction and attention enhancement architecture. Given an input feature tensor X, it first passes through the X' channel compression transformation T C :

[0091] X′=Τ C (X) = δ(WC ΘX+b C )

[0092] Among them, δ represents the nonlinear activation function, Θ represents the convolution operation, WC and bC are the learnable convolution kernel parameters and bias vector respectively. This operation reduces the channel dimension of the input feature to half, and obtains T C .

[0093] Then, the module performs a progressive max pooling operation M p Construct multi-scale feature representation. Define the progressive pooling process as:

[0094]

[0095] Here, Mp(·,k) represents a maximum pooling operation with a kernel size of k×k, a stride of 1, and a padding of k / 2. Finally, the original features are concatenated with the pooled features at all levels in the channel dimension to obtain a multi-scale fused representation.

[0096] AMFE uses SKA, which achieves efficient long-range dependency modeling through separable convolution kernel decomposition. Specifically, the SKA attention map A can be expressed as:

[0097] A(Z)=T1(T sv (T sh (T v (T h (Z)))))

[0098] The transformation operators are defined as follows:

[0099]

[0100] Here, Θ g represents grouped convolution, Θ g,d W represents the group convolution with dilation rate. h W represents a 1×3 horizontal convolution kernel. v W represents the 3×1 vertical convolution kernel. sh W represents a 1×5 horizontal convolution kernel with a dilation rate of d=2. sv represents a 5×1 vertical convolution kernel with a dilation rate of d = 2. W1 represents a 1×1 convolution kernel.

[0101] This design reduces the computational complexity from O(k 2 ) to O(k) while maintaining the feature capture capability of an equivalent 11×11 receptive field. The resulting attention map is applied to multi-scale features via residual connections. Finally, the enhanced features are mapped to the output space via a channel fusion transform, as shown in the following formula:

[0102] O=Τ Ο (Z′)=δ Ο (W O ΘZ′+b O )

[0103] Where W O and b O are learnable parameters.

[0104] The AMFE module guides the feature extraction process through an attention mechanism, enabling the network to adaptively strengthen the feature representation of small target areas while suppressing background interference. The large receptive field design of the SKA enables the module to perceive a wider range of contextual information, which is crucial for distinguishing morphologically similar but semantically distinct objects in remote sensing images.

[0105] Specific as Figure 4 As shown in the figure, the EDAM two-stage routing processing module adopts a dual-path collaborative strategy of region and feature. It first screens high-value candidate regions through regional routing and establishes spatial associations. Then, it dynamically aggregates channel dimensions based on feature routing. While reducing redundant calculations, it strengthens long-distance feature associations, and ultimately realizes small target feature modeling and cross-scale information fusion.

[0106] The EDAM module adopts a cascade structure design. Its core component is the Hierarchical Routing Attention Block (HRABlock), which consists of the Dual-Scale Context Attention (DSCA) and the Feedforward Network (FFN). Mathematically, the overall conversion process of the EDAM module can be expressed as:

[0107]

[0108] Among them, X1 and X2 are two parts of features obtained by segmenting the input feature map X after passing through the initial convolution layer. represents the output convolution layer, F HRA Represents the conversion function of HRABlock.

[0109] DSCA achieves efficient long-distance dependency modeling by cleverly combining region-level routing and pixel-level attention. First, the input feature X is linearly projected to generate the query Q, key H, and value V:

[0110] Q,K,V=F proj (X)

[0111] Among them, Fproj represents a 1×1 convolutional layer that maps the input features to the QKV space.

[0112] The average pooling operation is used to obtain region-level query and key representations:

[0113]

[0114] Among them, Q r and K r represents the query Q and key K at the region level, Favgpool represents the average pooling, Q detach and K detach The gradient calculation of the representation is separated, which ensures that the regional routing strategy only serves as an attention guide and does not directly participate in the gradient back propagation, effectively preventing the problem of training instability. Subsequently, the similarity matrix between regions is calculated:

[0115]

[0116] in, and Denote the reshaped region query and key respectively. Based on the similarity matrix, select the top-k most relevant regions for each region:

[0117] idx r =TopK(A r ,k)

[0118] Based on the results of region routing, perform pixel-level attention calculation:

[0119]

[0120] Among them, Mregion is the sparse attention mask generated by the region routing result idxr, which restricts each pixel to interact only with pixels in the relevant region. represents element-wise multiplication, and dk is the dimension of the query vector. This method reduces the computational complexity of the original self-attention from O(h 2 w 2 ) is reduced to Greatly improved computing efficiency.

[0121] To further enhance local feature representation, HRA designed Local Enhancement (LEPE), which uses depthwise separable convolution to capture local neighborhood information. The final attention output is obtained by merging the global attention result and the local enhancement part:

[0122]

[0123] The feedforward network in HRABlock uses a two-layer 1×1 convolutional structure for feature transformation and nonlinear modeling. HRABlock's complete processing flow combines the attention mechanism with the feedforward network, and ensures stable information flow through residual connections:

[0124] X out =X+FFFN (X+F Attn (X)

[0125] This design enables the network to effectively integrate global context information and local feature representation, while ensuring the stable propagation of gradients, which is conducive to the training convergence of the model.

[0126] Experiment 1:

[0127] Dataset establishment: Remote Sensing Airport-Plane Detection Dataset (RS-APD), which focuses on aircraft targets in airport environments. Its characteristics are small target size, variable posture and complex background.

[0128] The RS-APD dataset has been augmented using seven methods: random cropping and resizing (crop), rotation (rotate), horizontal flipping (flip), Gaussian noise (noise), RGB channel offset (color), brightness and contrast adjustment (brightness), and blurring (blur). The dataset contains 4,968 high-resolution remote sensing images with a resolution of 1024×1024. These images were acquired from multiple satellite platforms, including the GaoFen series (GF-1 / 2), WorldView-3 / 4, Pleiades-Neo, and SPOT-7, covering over 600 airports worldwide and encompassing diverse climate regions, seasonal variations, and weather conditions.

[0129] This proposal verifies the feasibility of the proposed improved model using the independently constructed RS-APD dataset and the VisDrone2019 dataset. The VisDrone2019 dataset is one of the most challenging and widely used benchmark datasets in the current remote sensing image field. This dataset contains more than 10,000 high-resolution aerial images. These images were collected by various drone platforms at altitudes ranging from 10 to 300 meters, at different time periods and under various weather conditions, covering a variety of complex scenarios.

[0130] Experimental Procedure: This experiment was conducted on the Ubuntu 20.04 operating system. The model was trained and tested on an NVIDIA GeForce RTX 3060 GPU. When training on the RS-APD dataset used in this article, the epoch number was set to 300, and the SGD optimizer was used to accelerate model convergence. Furthermore, an innovative input size of 1024 was used in the experiment, replacing the traditional 640.

[0131] Experimental data: Table 1 Experimental results of model performance comparison under different input sizes

[0132]

[0133] Experimental conclusion:

[0134] The comparative experiments in Table 1 show that the baseline algorithm significantly improves detection accuracy on the RS-APD dataset after changing the input size. At the same time, detection accuracy on the VisDrone2019 dataset also improves significantly.

[0135] Experiment 2:

[0136] Experimental purpose: To verify the performance advantage of the model proposed in this scheme in the remote sensing small target detection task.

[0137] Experimental steps:

[0138] 1. Dataset preparation:

[0139] Use the self-built remote sensing small target dataset RS-APD and the public dataset VisDrone2019.

[0140] Unified data preprocessing: the image size is adjusted to 640×640, normalized, and mosaic data augmentation is applied.

[0141] The training set, validation set, and test set are divided into 8:1:1 parts.

[0142] 2. Model configuration and training

[0143] The comparison models are: YOLOv5n, YOLOv8n, YOLO11n, YOLOv12n and DAE-YOLO (integrated DSSM, AMFE, EDAM modules).

[0144] The training parameters are unified: initial learning rate 0.01, SGD optimizer, momentum 0.937, weight decay 0.0005, and training for 300 epochs.

[0145] Hardware environment: NVIDIA A100 GPU, PyTorch 2.0 framework, mixed precision training enabled.

[0146] 3. Performance evaluation:

[0147] Test indicators: mAP50, mAP50:95, APs (small target detection accuracy), parameter count (Params) and computational complexity (FLOPs).

[0148] Use the COCO evaluation tool to calculate the average accuracy and resource consumption of each model on the test set.

[0149] Experimental data: Table 2

[0150]

[0151]

[0152] Experimental conclusion:

[0153] The proposed model achieves leading mAP50 and APs performance of 76.9% and 47.3%, respectively, compared to YOLO11n, achieving improvements of 3.3% (mAP50) and 6.1% (APs), demonstrating the effectiveness of multi-module collaborative design for small object detection. Despite a slight increase in parameters (+0.5M), the significant improvement in detection accuracy demonstrates the high parameter efficiency of the optimized model architecture.

[0154] Experiment 3:

[0155] Experimental purpose: To verify the performance advantages of the DSSM module in the C3k2 structure improvement.

[0156] Experimental steps:

[0157] 1. Module replacement and model construction

[0158] Baseline model: The C3k2 module of YOLO11n is replaced with the comparison modules (C3k2_Star, C3k2_RVB, C3k2_Faster) and DSSM modules.

[0159] Module adaptation: Ensure that the input and output channels of each module are consistent and keep other parts of the network structure unchanged.

[0160] 2. Training and Validation

[0161] Dataset: Only the RS-APD dataset is used, and the training strategy is consistent with Experiment 2.

[0162] Freeze the first 20 layers of Backbone and fine-tune the replacement module and subsequent network layers to avoid overfitting.

[0163] 3. Results Analysis

[0164] The APs and mAP50 indicators of the validation set are extracted to compare the performance differences of each module in small object detection.

[0165] Count the FLOPs and parameters to verify the effectiveness of the lightweight design of the module.

[0166] Experimental data: Table 3

[0167]

[0168] Experimental Conclusions: The proposed DSSM module significantly outperforms the comparison module with an APs of 44.7% (up to 3.4%). Its state-space modeling and multi-directional sequence processing effectively enhance the ability to capture small target features. While the number of parameters increases by only 0.2M, the computational efficiency (FLOPs) remains on par with similar modules, demonstrating its lightweight design advantage.

[0169] Experiment 4:

[0170] Experimental purpose: To verify the performance advantages of the AMFE module in improving the SPPF structure.

[0171] Experimental steps:

[0172] 1. Module integration and training

[0173] The SPPF layer of the baseline model YOLO11n is replaced by the comparison modules (SPPF_AIFI, SPPF_FPS, SPPF_FM) and AMFE modules.

[0174] Adjust the feature pyramid fusion strategy to ensure the consistency of multi-scale feature transfer.

[0175] 2. Optimization strategy

[0176] Enable dynamic learning rate adjustment (Cosine annealing strategy) to avoid oscillation in the late stages of training.

[0177] Added gradient clipping (max norm = 10) to improve training stability.

[0178] 3. Performance testing

[0179] For the sub-test set of cloud occlusion and low-contrast scenes, the APs improvement of each module is counted.

[0180] Visualize the heat map and analyze the AMFE module's suppression effect on background noise.

[0181] Experimental data: Table 4

[0182] Module APs (%) mAP50 (%) Params(M) FLOPs(G) SPPF_AIFI 38.9 69.8 1.7 4.3 SPPF_FPS 40.2 71.2 1.9 4.5 SPPF_FM 41.5 72.5 2.0 4.7 AMFE 45.1 76.3 2.2 4.9

[0183] Experimental Conclusions: The AMFE module achieves 45.1% APs through its large-core attention mechanism, a 3.6% improvement over SPPF_FM. This demonstrates the significant benefits of its multi-scale perception and noise suppression strategies for small object detection. The increase in FLOPs is only 0.2 GB, demonstrating its ability to balance computational efficiency and accuracy.

[0184] Experiment 5:

[0185] Experimental purpose: To verify the performance advantages of EDAM module in dual-stage routing design.

[0186] Experimental steps:

[0187] Experimental data: Table 5

[0188] Module APs (%) mAP50:95(%) Params(M) FLOPs(G) C2DA 42.8 49.1 2.1 5.0 C2CGA 43.6 50.3 2.3 5.2 EDAM 47.2 53.8 2.4 5.3

[0189] Experimental Conclusions: The proposed EDAM module leads the field with an APs score of 47.2% and an mAP50:95 score of 53.8%. Its dual-stage routing strategy significantly improves the efficiency of long-range feature association through region-feature collaborative optimization. The parameter count increases by only 0.1M, demonstrating the lightweight design's adaptability to small object detection tasks.

[0190] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A new lightweight remote sensing small target detection system, based on YOLO11, is characterized by: It includes DSSM feature extraction module, AMFE feature enhancement module and EDAM two-stage routing processing module; The DSSM feature extraction module adopts the concept of CSP network and embeds SSBlock operation to identify the input image and extract global features. The DSSM feature extraction module consists of several SSB modules connected in series. When a feature is input, the input feature X is first channel-adjusted through 1×1 convolution and divided into X1 and X2. The main branch X2 is then further processed by the series of SSB modules. Finally, all processed features are cascaded and fused with the initial branch. The AMFE feature enhancement module is used to perform multi-scale perception of the global features extracted by DSSM through an adaptive large-core attention mechanism, perform multi-scale enhancement and noise suppression, and use separable convolution to extract contextual information at different scales. Subsequently, dynamic weight allocation is used to enhance the feature response of small target areas while suppressing background noise interference. Finally, multi-level features are fused to enhance the model's ability to discriminate complex scenes, highlight the local features of small targets, and generate a multi-level fused feature map; The EDAM two-stage routing processing module adopts a dual-path collaborative strategy of regional routing and feature routing. It first screens high-value candidate areas and establishes spatial associations through regional routing, and then dynamically aggregates channel dimensions based on feature routing. While reducing redundant calculations, it strengthens long-distance feature associations, and ultimately realizes small target feature modeling and cross-scale information fusion.

2. The novel lightweight remote sensing small target detection system according to claim 1 is characterized in that: The mathematical expression of the SSB module is as follows: Among them, F cv1 represents a k×k convolutional layer (k is usually set to 3), F SSBlock Represents an SSBlock operation.

3. The novel lightweight remote sensing small target detection system according to claim 2 is characterized in that: The SSBlock operation includes the LN layer normalization operation and the SS2D selective scanning attention operation. Its basic operations are as follows: F SSBlock (X)=X+DropPath(F SS2D (F LN (X))) Among them, X is the input feature map, F LN Representation layer normalization operation, F SS2D Represents a 2D selective scan operation.

4. The novel lightweight remote sensing small target detection system according to claim 3 is characterized in that: In the SS2D selective scanning attention operation, the input feature X is first mapped to a higher-dimensional feature space through a linear projection layer and decomposed into a content feature X' and a gated feature Z. The content feature X' is further processed by 2D convolution; the processed features are reorganized into a sequence representation in four directions: X seq =[X h ,X w ,X h_flip ,X w_flip ] Among them, X h is the characteristic sequence expanded by row, X w is the characteristic sequence expanded by column, X h_flip and X w_flip They are the sequences after horizontal and vertical flipping respectively. For each direction, the SS2D operation is performed to model the sequence. For the input sequence x and the initial state h0, the state update and output are calculated as follows: Among them, A is the state transfer matrix, B is the input projection, C is the state output projection, and D is the skip connection coefficient.

5. The novel lightweight remote sensing small target detection system according to claim 4 is characterized in that: The SS2D operation also introduces an adaptive time step mechanism: Among them, Δ t is the adaptive time step, is a time step control factor derived from the input features; the final processed features are fused with the original features through a gating mechanism and mapped back to the original dimension through the output projection layer: in, represents element-wise multiplication, and σ is the SiLU activation function; in the process of implementing the gating mechanism, the model can adaptively control the feature flow of each position and channel.

6. The novel lightweight remote sensing small target detection system according to claim 5 is characterized in that: The AMFE feature enhancement module adopts a cascaded feature extraction and attention enhancement architecture. Given an input feature tensor X, it first undergoes a channel compression transformation X′: X′=Τ C (X)=δ(W C ΘX+b C ) Among them, δ represents the nonlinear activation function, Θ represents the convolution operation, and W C and b C are the learnable convolution kernel parameters and bias vectors respectively; the AMFE feature enhancement module reduces the channel dimension of the input feature to half, and obtains T C ; Then, through the progressive maximum pooling operation M p Construct multi-scale feature representation; define the progressive pooling process as: Among them, Mp(·,k) represents the maximum pooling operation with a kernel size of k×k, a stride of 1, and a padding of k / 2. Finally, the extracted global features are concatenated with the pooling features at all levels in the channel dimension to obtain the attention map of the multi-scale fusion representation.

7. The novel lightweight remote sensing small target detection system according to claim 6 is characterized in that: The AMFE feature enhancement module also uses the SKA mapping operation. The SKA attention map A can be expressed as: A(Z)=T1(T sv (T sh (T v (T h (Z))))) The transformation operators are defined as follows: In the formula, Θ g represents grouped convolution, Θ g,d W represents the group convolution with dilation rate. h Represents the horizontal convolution kernel, W v Represents the vertical convolution kernel, W sh represents the horizontal convolution kernel with dilation rate d = 2, W sv represents the vertical convolution kernel with dilation rate d=2, and W1 represents the convolution kernel; The attention map is applied to multi-scale features through residual connections, and finally, the enhanced features are mapped to the output space through channel fusion transformation.

8. The novel lightweight remote sensing small target detection system according to claim 7 is characterized in that: The EDAM two-stage routing processing module includes the HRABlock component, which includes the DSCA operation and the FFN feedforward network. Mathematically, the overall conversion process of the EDAM two-stage routing processing module is expressed as: Among them, F HRA Represents the conversion function of the HRABlock component; The specific operation of DSCA operation is, first, the input feature X is linearly projected to generate query Q, key K and value V: Q,K,V=F proj (X) Among them, F proj Represents the convolutional layer, which maps the input features to the QKV space; The average pooling operation is used to obtain region-level query and key representations: Among them, Q r and K r Represents the region-level query Q and key K, F avgpool represents average pooling, Q detach and K detach represents the gradient calculation of the separation; Then, the similarity matrix between regions is calculated: in, and Denote the reshaped region query and key respectively. Based on the similarity matrix, the top-k most relevant regions of each region are selected: idx r =TopK(A r ,k) Based on the results of region routing, perform pixel-level attention calculation: Among them, M region The regional routing result idx r The generated sparse attention mask is used to restrict each pixel to interact only with pixels in the relevant region, d k is the dimension of the query vector.

9. The novel lightweight remote sensing small target detection system according to claim 8, characterized in that: The FEN feedforward network adopts a two-layer 1×1 convolutional structure for feature conversion and nonlinear modeling.

10. The novel lightweight remote sensing small target detection system according to claim 9, characterized in that: The HRABlock component also includes a LEPE supplementary enhancement operation. The final attention output is obtained by merging the global attention result and the local enhancement part: The complete processing flow of the HRABlock component combines the attention mechanism and the FEN feedforward network, and processes the information flow X through the residual connection out =X+F FFN (X+F Attn (X)).