Ship detection method oriented to complex SAR (Synthetic Aperture Radar) scene

Through the DPM-YOLO ship detection network, using the DPSConv module, PromptFusionMod module and MscaleASFFHead detection head, the balance problem of high precision, high robustness and low latency in SAR ship detection is solved, and the detection performance in complex sea conditions is improved.

CN120747883APending Publication Date: 2025-10-03SHANGHAI MARITIME UNIVERSITY

Patent Information

Application Number
CN202511106030.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing SAR ship detection methods have difficulty achieving an effective balance between high precision, high robustness and low latency, especially in complex sea conditions, where the missed detection rate of small targets is high, cross-scale semantic alignment is insufficient, and lightweight deployment lacks robustness.

Method used

A DPM-YOLO ship detection network is constructed. The DPSConv module is introduced to improve fine-grained feature extraction, the PromptFusionMod module is used to enhance multimodal feature fusion, and the MscaleASFFHead detection head is used to coordinate the inconsistency between multi-scale features to achieve precise positioning.

Benefits of technology

It significantly improves detection accuracy and robustness in complex sea conditions, increases small target detection rate and large target positioning accuracy, while maintaining the computational efficiency of lightweight deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747883A_ABST
    Figure CN120747883A_ABST
Patent Text Reader

Abstract

The invention relates to a ship detection method for a complex synthetic aperture radar (SAR) scene, and belongs to the technical field of remote sensing image target detection.The method comprises the steps that an obtained high-resolution synthetic aperture radar image is input into a DPM-YOLO ship detection network, a ship target detection effect picture is output, the DPM-YOLO ship detection network is improved based on a YOLOv11 network, and the ship target detection effect picture is obtained; a DPSConv module is introduced into a backbone network to capture multi-scale context information by using a receptive field while keeping fine feature details, and a PromptFusionMod module is introduced into a neck network to perform multi-modal feature fusion through four processing stages of space compression and prompt fusion, an efficient attention mechanism, a lightweight multi-layer perceptron and output fine processing. And the original detection head is replaced by the MscaleASFHead detection head. Compared with the prior art, the method has the advantages that fine-grained feature extraction, cross-scale semantic alignment and lightweight deployment can be considered, and ship detection precision and robustness in a complex sea condition scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image target detection, and in particular to a ship detection method for complex SAR scenes. Background Art

[0002] As an active microwave imaging sensor, high-resolution synthetic aperture radar (SAR) images can achieve long-range, high-resolution ocean observation without being restricted by illumination or weather conditions. Therefore, they hold irreplaceable strategic value in key areas such as maritime search and rescue, oil spill monitoring, traffic management, fishery supervision, immigrant tracking, and coastal defense. The core goal of SAR ship detection is to accurately extract ship features under complex background interference, achieving high-confidence identification and sub-pixel positioning. However, the inherent coherent speckle noise, sea-land clutter coupling, and high heterogeneity of ship targets in scale, aspect ratio, orientation, and density in SAR images pose severe challenges to traditional deep learning methods in terms of feature representation and multi-scale modeling, severely restricting further improvements in detection accuracy and robustness.

[0003] In recent years, convolutional neural networks (CNNs) have made significant progress in SAR ship detection, but existing methods still face numerous limitations. GELAN-CMP, proposed by Dong et al., improves large-scale ship detection performance through a component model and a multi-layer, multi-pooled channel attention mechanism, but this inevitably leads to a significant increase in computational complexity. YOLO-FA, proposed by Zhang et al., introduces a frequency-domain adaptive weight module to suppress complex background clutter, but the frequency-domain transform operation introduces additional computational latency. FDI-YOLO, proposed by Wang et al., utilizes a reversible subnetwork, RCSPNet, and a dynamic detection head, DyHead, to achieve lightweight multi-scale detection. However, cross-layer feature consistency remains limited. The single-stage DB-YOLO, designed by Zhu et al., balances real-time performance with multi-scale fusion, but still suffers from a high rate of missed detection of small targets in extreme clutter. DS-YOLO, proposed by Hu et al., mitigates the effects of sample scarcity and noise by jointly training the despeckling and detection modules, but this results in a complex training process and significantly increased computational overhead. Chinese patent CN116844055A discloses a lightweight SAR ship detection method and system. The network used includes a lightweight backbone network that uses depthwise separable convolution to extract features and output multiple branch feature maps, an enhanced spatial pyramid that uses multi-scale pooling operations to process deep branch feature maps, a multi-scale feature fusion network that fuses each branch feature map from top to bottom, and a detection head that classifies and identifies the fused feature maps. However, in pursuit of lightweightness, the lightweight backbone network lacks robustness in feature extraction of tiny ships in complex backgrounds with high noise.

[0004] Therefore, the existing algorithms have not yet achieved an effective balance between "high precision, high robustness and low latency". There is an urgent need for a detection method that can take into account fine-grained feature extraction, cross-scale semantic alignment and lightweight deployment to improve the detection accuracy and robustness of SAR ship detection in complex sea conditions. Summary of the Invention

[0005] The purpose of this invention is to provide a ship detection method for complex SAR scenarios and construct a DPM-YOLO ship detection network. The network performs multi-scale fusion enhancement based on the YOLOv11 architecture. Through the three-level core innovative design, it effectively reduces the missed detection rate of small targets and the positioning deviation of large targets, so as to improve the detection accuracy and robustness in complex sea conditions.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] A ship detection method for complex SAR scenarios is proposed. The method inputs the acquired high-resolution synthetic aperture radar image into the DPM-YOLO ship detection network and outputs a ship target detection rendering. The DPM-YOLO ship detection network is improved based on the YOLOv11 network. The DPSConv module is introduced into the backbone network to capture multi-scale contextual information using the receptive field while retaining subtle feature details. The PromptFusionMod module is introduced into the neck network to perform multimodal feature fusion through four processing stages: spatial compression and prompt fusion, efficient attention mechanism, lightweight multi-layer perceptron, and output refinement processing. The original detection head is replaced by the MscaleASFFHead detection head.

[0008] In the DPM-YOLO ship detection network, the input features are preliminarily filtered and channel compressed by the CBS module and the C3K2 module to obtain low-level features, which are then processed by the first DPSConv module to obtain intermediate features. The intermediate features are normalized and activated by the CBS module and then input into the second DPSConv module. The output of the second DPSConv module is sent to the SPPF module for multi-scale maximum pooling parallel processing and feature aggregation. The features output by the SPPF module are sent to the C2PSA module, which uses the cross-stage partial attention mechanism to capture long-range dependencies and output high-level features.

[0009] The low-level features, intermediate features, and high-level features are iterated upward step by step in the neck network through the upsampling-splicing-fusion process to construct a multi-level semantic representation, wherein the upsampling-splicing-fusion process is specifically as follows: the relatively low-level features are processed by the CBS module, the C3K2 module, and the PromptFusionMod module, and then upsampled, spliced ​​with the relatively high-level features, and normalized by the CBS1×1 module after splicing;

[0010] The multi-scale features processed by the neck network are input into the corresponding MscaleASFFHead detection head respectively. Adaptive weights are dynamically assigned based on context information, and information from low, medium and high feature paths is integrated to achieve accurate positioning of targets of different sizes. The MscaleASFFHead detection heads of each scale output results in parallel and summarize them into the final ship prediction result.

[0011] The DPSConv module consists of two serially connected FConv submodules. The FConv submodule adopts a progressive receptive field expansion strategy, integrating the information bottleneck theory and the feature differentiation principle, and learns the optimal compressed representation of the input data through a dual-path structure. The dual-path structure operates as an implicit feature diversification mechanism, where each path is used to extract different subspace features.

[0012] The FConv submodule performs a dimensionality reduction projection transformation on its input feature map After that, the dual-path feature transformation is performed in sequence and Perform transformation and enhancement, and finally perform feature aggregation and dimension expansion transformation on the results of dual-path feature transformation The final output is obtained, and the process is expressed as:

[0013]

[0014] in, represents the FConv operation, Indicates the FConv operation with the bias set to 0, symbol represents the field of real numbers, represents the convolution operation, W D1 ~W D6 is the convolution kernel, b D is the bias term, BN(·) represents the batch normalization operation, φ(·) represents the activation function, Indicates the splicing operation, X D It is the input feature map of the FConv submodule, Y1 to Y5 represent the intermediate tensors in the hierarchical feature transformation process, and the feature Y1 after the dimensionality reduction projection transformation of the input feature map is used as the shared source for dual-path processing. Heterogeneous nonlinear transformations are performed separately through parallel branches. Y2 and Y3 correspond to the first-level outputs of path A and path B respectively. Y4 and Y5 represent the final enhanced representation of the dual paths. High-order features with complementary discriminant characteristics are generated through cascaded convolution-normalization-activation operations. After processing, the output Y of the FConv submodule is obtained by fusion D .

[0015] The spatial compression and prompt fusion processing stage of the PromptFusionMod module first introduces the learnable prompt parameter QP and expands it:

[0016]

[0017] Where QP∈R 1×C×1×1 represents the learnable hint parameter, QP expande is the tensor of the learnable hint parameter after the expansion operation, B P ,C P ,H P ,W P Represent the batch size, number of channels, feature map height and width after expansion, respectively. Represents the real number domain, expand represents the expansion operation;

[0018] Afterwards, the tensor after the expansion operation of the learnable prompt parameter and the input feature of the PromptFusionMod module are spliced ​​through the channel axis to generate the spliced ​​feature X concat :

[0019]

[0020] Among them, dim=1 means splicing operation along the channel dimension, X P Represents the input features of the PromptFusionMod module;

[0021] Perform the following operations on the concatenated features to generate the fused features F: fused :

[0022]

[0023] Among them, σ SiL Represents SiLU activation function, Conv 1×1 represents a 1×1 convolution operation, and BN represents a batch normalization operation.

[0024] The efficient attention mechanism processing stage of the PromptFusionMod module performs the following steps in sequence:

[0025] The fusion feature F output from the spatial compression and hint fusion stages fused Perform spatial downsampling and channel compression:

[0026]

[0027] Among them, AvgPool 2×2 represents a 2×2 average pooling operation, X down represents the downsampled feature map;

[0028] Perform feature refinement on the downsampled feature map:

[0029]

[0030] Among them, X attn Represents the output feature map refined by channel attention, that is, the tensor after the deep convolution-activation-channel reweighting operation, DWConv 3×3 represents a 3×3 depth-separable convolution operation, GELU represents the Gaussian error linear unit activation function, represents the channel attention function;

[0031] For the feature map X after feature refinement attn Perform spatial resolution restoration:

[0032]

[0033] Among them, A up Represents the attention refined feature map reconstructed by upsampling, ConvTranspose 2×2 It represents a 2×2 transposed convolution operation with a learnable kernel, and stride represents the step size parameter.

[0034] The lightweight multi-layer perceptron processing stage of the PromptFusionMod module processes the fusion feature F output by the spatial compression and prompt fusion processing stage. fused And the attention refined feature map A output by the efficient attention mechanism processing stage up Perform element-wise addition and then process it through a lightweight MLP network:

[0035]

[0036]

[0037] Among them, X add Represents the feature map after element-by-element addition, X mlp Represents the output feature map of the lightweight multilayer perceptron processing stage.

[0038] The output refinement processing stage of the PromptFusionMod module is used to process the output feature map X of the lightweight multi-layer perceptron processing stage. mlp Processing is performed and the final output is generated through residual connection and convolution refinement network:

[0039]

[0040] Among them, X res Represents the residual feature map, C Pout Indicates the number of output channels, Y PRepresents the output of the PromptFusionMod module.

[0041] The MscaleASFFHead detection head adopts a four-level feature fusion strategy based on the ASFF module, processing feature maps at four independent levels. Each level corresponds to a different target scale, and dynamically allocates fusion weights according to the importance of features at each level to achieve adaptive fusion of multi-scale features and output fused feature maps. Among them, level 0 processes the minimum-scale feature map to characterize the fine details of small targets; level 1 processes the medium-scale feature map to enhance the perception of medium-sized targets; level 2 focuses on large-scale feature maps to capture the contextual information of large targets; level 3 manages macro-scale feature maps to assist the network in aggregating cross-scale information to ensure the integrity of scene understanding; the process is expressed as:

[0042] W i =f weight (F i )

[0043]

[0044] Among them, f weight represents the learnable convolution operation, F i is the feature map of the i-th level, W i is the unnormalized weight tensor of the i-th level, α i is the normalized weight coefficient of the i-th level after cross-level normalization by the Softmax function, F Mfused is the fusion feature map.

[0045] The MscaleASFFHead detection head performs the following steps on the fused feature map to generate the output result:

[0046] For the fusion feature map F Mfused Perform dimension expansion to obtain the output feature map F out :

[0047] F out =f expand (F Mfused )

[0048] Among them, f expand represents the dimension expansion function;

[0049] Extract the local feature vector F corresponding to each candidate box from the output feature map out1 Based on the DFL function, a multi-position distributed regression strategy is used to assign the bounding box regression task to a discrete position set, and then the Softmax function is used to assign weight coefficients on the spatial position. Finally, the coordinate prediction is realized through probability distribution modeling to generate the bounding box prediction result:

[0050] box=DFL(F out1 )×norm

[0051] Among them, DFL is the distribution focus loss function, norm is the normalization factor to enhance numerical stability, and box represents the regression prediction of the bounding box;

[0052] At the same time, the output feature map F out Apply the final convolution operation:

[0053] output=Conv(F out )

[0054] Among them, output represents the prediction result of class probability, and Conv represents the final convolution operation;

[0055] The bounding box regression prediction and class probability prediction results are combined as the final output of the MscaleASFFHead detection head.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] (1) This paper proposes a dual-branch convolutional architecture (DPSConv) to enhance the fine-grained feature extraction capability of ship targets in SAR images. This architecture demonstrates enhanced feature discrimination and representation accuracy when addressing challenges such as the diverse shapes, significant scale differences, and complex textures of ship targets.

[0058] (2) The present invention constructs the PromptFusionMod module, which can effectively guide the fusion process of multi-source heterogeneous features, and combines the spatial compression and upsampling mechanisms to further enhance the spatial focusing ability and semantic consistency of the ship area.

[0059] (3) This paper constructs the MscaleASFFHead detection head, which specifically addresses the inconsistencies between feature levels for multi-scale ship targets. By coordinating the expression differences between features at each level, this module improves the detection rate of small targets and the detection accuracy of large targets, thereby effectively improving overall detection performance.

[0060] (4) A large number of experiments were carried out on the widely used high-resolution SAR image dataset HRSID. The experimental results show that the DPM-YOLO ship detection network of the present invention has more advantages in feature extraction and can effectively perform the ship detection task of SAR images. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 Flowchart of the DPM-YOLO ship detection network used in the present invention;

[0062] Figure 2 This is the overall architecture diagram of the DPM-YOLO ship detection network in the present invention;

[0063] Figure 3 Schematic diagram of the structure of the DPSConv module in the present invention;

[0064] Figure 4 Schematic diagram of the structure of the PromptFusionMod module in the present invention;

[0065] Figure 5 Schematic diagram of the structure of the MscaleASFFHead detection head in the present invention;

[0066] Figure 6 This is a visualization result diagram of an ablation experiment in one embodiment of the present invention;

[0067] Figure 7 This is a comparison diagram of the visual detection results of an embodiment of the present invention with other algorithms. DETAILED DESCRIPTION

[0068] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0069] This embodiment provides a ship detection method for complex SAR scenarios, such as Figure 1 As shown, the method includes the following steps:

[0070] Acquire high-resolution synthetic aperture radar imagery;

[0071] High-resolution synthetic aperture radar images are input into the DPM-YOLO ship detection network. The DPM-YOLO ship detection network is an improvement based on the YOLOv11 network. While inheriting the traditional feature extraction and fusion framework, it systematically enhances feature expression capabilities and multi-scale adaptability. The overall network still follows the feature progressive processing paradigm from bottom to top, but introduces the DPSConv module into the backbone network to capture multi-scale contextual information using the receptive field while retaining subtle feature details. The PromptFusionMod module is introduced into the neck network to perform multimodal feature fusion through four processing stages: spatial compression and prompt fusion, efficient attention mechanism, lightweight multi-layer perceptron, and output refinement. The original detection head is replaced with the MscaleASFFHead detection head to improve the richness of semantic expression and optimize the alignment accuracy of multi-scale features.

[0072] The DPM-YOLO ship detection network outputs a ship target detection rendering, thereby achieving accurate detection and positioning of ships in SAR images.

[0073] like Figure 2 As shown in the figure, in the DPM-YOLO ship detection network, the input features are preliminarily filtered and channel compressed by the CBS (Conv-BN-SiLU) module and the C3K2 (Cross Stage Partial with kernel size 2) module to obtain low-level features, which are then processed by the first DPSConv module to obtain intermediate features. The intermediate features are normalized and activated by the CBS module and then input into the second DPSConv module. The output of the second DPSConv module is sent to the SPPF (Spatial Pyramid Pooling-Fast) module for multi-scale maximum pooling parallel processing and feature aggregation. The features output by the SPPF module are sent to the C2PSA module, which uses the cross-stage partial attention mechanism to capture long-range dependencies and output high-level features, thereby improving the network's ability to focus on key target information.

[0074] At the intermediate and high-level feature processing levels, the network significantly improves the pertinence and discrimination of multi-scale fusion through the splicing and fusion of two PromptFusionMods and upsampled features. Specifically, low-level features, intermediate features, and high-level features are iterated upward step by step in the neck network through an upsampling-splicing-fusion process to construct a multi-level semantic representation. The upsampling-splicing-fusion process is specifically as follows: relatively low-level features are processed by the CBS module, C3K2 module, and PromptFusionMod module, then upsampled and spliced ​​with relatively high-level features. After splicing, they are normalized by the CBS1×1 module to ensure the numerical stability and distribution consistency of the fused features.

[0075] The multi-scale features processed by the neck network are input into the corresponding MscaleASFFHead detection head respectively. Adaptive weights are dynamically assigned based on context information, and information from low, medium and high feature paths is integrated to achieve accurate positioning of targets of different sizes. The MscaleASFFHead detection heads of each scale output results in parallel and summarize them into the final ship prediction result.

[0076] The DPSConv module is integrated into the YOLOv11 backbone network, aiming to improve the small object detection performance by retaining subtle feature details while using the receptive field to capture multi-scale context information. Figure 3As shown in the figure, the DPSConv module consists of two serially connected FConv submodules, which play a key role in capturing small-scale features in images. The FConv submodule adopts a progressive receptive field expansion strategy, integrates the information bottleneck theory and the feature differentiation principle, and learns the optimal compression representation of the input data through a dual-path structure. The dual-path structure operates as an implicit feature diversification mechanism, and each path is used to extract different subspace features.

[0077] Unlike traditional convolution, which generates all feature maps, the mathematical representation and transformation sequence of FConv is formally defined as follows:

[0078]

[0079] Among them, B D 、c D1 、c D2 、H D and W D They represent the corresponding batch size, number of input channels, number of output channels, and feature map height and width respectively.

[0080] The FConv submodule performs a dimensionality reduction projection transformation on its input feature map After that, the dual-path feature transformation is performed in sequence and Perform transformation and enhancement to improve inter-class separability and feature discriminability, and finally perform feature aggregation and dimension expansion transformation on the results of dual-path feature transformation The final output is obtained, and the process is expressed as:

[0081]

[0082]

[0083] in, represents the FConv operation, Indicates the FConv operation with the bias set to 0, symbol represents the field of real numbers, represents the convolution operation, W D1 ~W D6 is the convolution kernel, b D is the bias term, BN(·) represents the batch normalization operation, φ(·) represents the activation function, Indicates the splicing operation, X DIt is the input feature map of the FConv submodule, Y1 to Y5 represent the intermediate tensors in the hierarchical feature transformation process, and the feature Y1 after the dimensionality reduction projection transformation of the input feature map is used as the shared source for dual-path processing. Heterogeneous nonlinear transformations are performed separately through parallel branches. Y2 and Y3 correspond to the first-level outputs of path A and path B respectively. Y4 and Y5 represent the final enhanced representation of the dual paths. High-order features with complementary discriminant characteristics are generated through cascaded convolution-normalization-activation operations. After processing, the output Y of the FConv submodule is obtained by fusion D .

[0084] The design of this submodule combines the information bottleneck theory and the principle of feature differentiation, and learns the optimal compressed representation of the input data through a dual-path structure. Its dual-path architecture operates as an implicit feature diversification mechanism, with each path focusing on extracting different subspace features. Specifically, the FConv submodule adopts a progressive receptive field expansion strategy, with the initial layer focusing on a smaller receptive field, which gradually expands as the layers are stacked, thereby capturing increasingly fine detail features. The DPSConv module processes multi-scale targets through a serial-concatenated combination of two FConv submodules. FConv can expand the receptive field without increasing the number of parameters or computational cost, effectively capturing large-scale contextual information, while integrating conventional convolution and cross-layer connection operations to achieve multi-scale information fusion. Thanks to this structural advantage, DPSConv significantly improves the detection performance of targets with a small proportion in the dataset.

[0085] like Figure 4 As shown in the figure, the PromptFusionMod module is an efficient multimodal feature fusion network architecture, which achieves effective feature integration and enhancement through four processing stages: spatial compression and prompt fusion, efficient attention mechanism, lightweight multi-layer perceptron (MLP) and output refinement processing.

[0086] The spatial compression and prompt fusion processing stage of the PromptFusionMod module first introduces and expands the learnable prompt parameter QP to enhance the multimodal information representation capability:

[0087]

[0088] Where QP∈R 1×C×1×1 represents the learnable hint parameter, QP expande is the tensor of the learnable hint parameter after the expansion operation, B P ,C P ,H P ,W P Represent the batch size, number of channels, feature map height and width after expansion, respectively. "denotes the real domain, and expand denotes an expansion operation. The core of this design is to maintain full channel dimensionality to achieve cross-modal feature adaptability, enabling the model to capture channel-specific cues encoding complementary multimodal information. Simultaneously, extreme spatial compression is implemented to optimize parameter efficiency, eliminating redundant spatial computations and focusing on cross-modal channel interactions. Cross-sample parameter sharing is achieved through a unified batch dimension design to improve generalization. Combined with runtime dynamic expansion operations to adapt to arbitrary input scales, this ultimately creates a lightweight storage-dynamic adaptation paradigm.

[0089] Afterwards, the tensor after the expansion operation of the learnable prompt parameter and the input feature of the PromptFusionMod module are spliced ​​through the channel axis to generate the spliced ​​feature X concat , realize dimension expansion to fully retain the original input features X P Independence from the prompt feature:

[0090]

[0091] Among them, dim=1 means splicing operation along the channel dimension, X P Represents the input features of the PromptFusionMod module.

[0092] The following operations are performed on the concatenated features to achieve cross-channel interaction and dimensionality reduction processing to generate the fusion feature F fused :

[0093]

[0094] Among them, σ SiL Represents the SiLU activation function, which is used to significantly enhance the nonlinear expression ability, Conv 1×1 represents the 1×1 convolution operation for dimensionality reduction, and BN represents the batch normalization operation for stabilizing training dynamics. The core design of this fusion mechanism is to reduce the channel dimension from C P to 2C P Extension of , using channel axis splicing to retain the original input feature X P With extended hint QP expanded The independence of , forming a high-dimensional joint representation to enhance information capacity. Then 1×1 convolution is used to compress the channel dimension back to C P , with the help of linear projection learning, the optimal weights of cross-channel interactions are realized, while the computational complexity is controlled through the bottleneck structure, and feature distillation and discriminant information purification are achieved.

[0095] Afterwards, the efficient attention mechanism processing stage of the PromptFusionMod module performs the following steps in sequence:

[0096] The fusion feature F output from the spatial compression and hint fusion stages fusedPerform spatial downsampling and channel compression to significantly reduce computational costs while retaining key information:

[0097]

[0098] Among them, AvgPool 2×2 represents a 2×2 average pooling operation, X down Represents the downsampled feature map. The process is to downsample the feature map The generation process of the previous stage fusion feature F Pfused As input, 2×2 average pooling is first used to achieve 1 / 2 spatial downsampling, and then 1×1 convolution is used to compress the channel dimension to 1 / 2 of its original size. This cascade operation fully preserves the key information of the spatial structure while systematically optimizing the computational complexity.

[0099] Then, feature refinement is performed on the downsampled feature map:

[0100]

[0101] Among them, X attn Represents the output feature map refined by channel attention, that is, the tensor after the deep convolution-activation-channel reweighting operation, DWConv 3×3 represents a 3×3 depth-separable convolution operation, GELU represents the Gaussian error linear unit activation function, Represents the channel attention function. The calculation process performs three core operations in sequence. First, 3×3 depth-separable convolution is used to achieve lightweight spatial feature extraction, then nonlinear transformation is introduced through the Gaussian Error Linear Unit (GELU) activation function, and finally, after 1×1 convolution for feature integration, the input is the channel attention function F. CA Dynamically calibrate channel weights to improve characterization capabilities.

[0102] For the feature map X after feature refinement attn Perform spatial resolution restoration:

[0103]

[0104] Among them, A up Represents the attention refined feature map reconstructed by upsampling, ConvTranspose 2×2 Denotes the use of a 2×2 transposed convolution with a learnable kernel to achieve spatial dimension expansion, and stride represents the stride parameter. This operation reverses the downsampling effect of the feature refinement step while fully preserving the enhanced feature representation capability.

[0105] The lightweight multi-layer perceptron processing stage of the PromptFusionMod module processes the fusion feature F output by the spatial compression and prompt fusion processing stages. fused And the attention refined feature map A output by the efficient attention mechanism processing stage up Perform element-wise addition and then process it through a lightweight MLP network:

[0106]

[0107] Among them, X add Represents the feature map after element-by-element summation, and the attention-refined feature A restored by the spatial resolution up Fusion feature F from the previous stage Pfused Superposition generation, X mlp Represents the output feature map of the lightweight multi-layer perceptron processing stage, which is obtained by add The three-order processing is performed in sequence. First, the feature transformation is realized by 1×1 convolution, and then the nonlinear expression ability is introduced by the GELU activation function. Finally, the feature reconstruction is completed by another 1×1 convolution.

[0108] The output feature map X of the output refinement processing stage of the PromptFusionMod module to the lightweight multi-layer perceptron processing stage mlp Processing is performed and the final output is generated through residual connection and convolution refinement network:

[0109]

[0110] Among them, X res Represents the residual feature map, C Pout Indicates the number of output channels, Y P Represents the output of the PromptFusionMod module. Among them, the residual connection ensures the stability of gradient back propagation, while the 3×3 convolution extracts local spatial features, supplemented by batch normalization and SiLU activation function to enhance nonlinear representation capabilities. The output tensor dimension remains Under the premise of strictly maintaining the spatial resolution, the channel dimension is adjusted to C through the convolution operation Pout This process fuses multi-level features through residual connections, and combines lightweight convolution with activation functions to achieve feature refinement and optimization.

[0111] like Figure 5As shown in the figure, the MscaleASFFHead detection head adopts a four-level feature fusion strategy based on the ASFF module. This module significantly improves the model's detection ability for small objects by combining adaptive spatial feature fusion with distribution focal loss (DFL). Its four-level structure processes feature maps of different scales separately and dynamically assigns fusion weights based on the importance of features at each level, achieving adaptive fusion of multi-scale features. This mechanism enables the network to more effectively process multi-scale objects, especially enhancing the feature representation of small objects.

[0112] Among them, the four-level feature fusion strategy is specifically as follows: feature maps are processed at four independent levels, each level corresponds to a different target scale, and fusion weights are dynamically allocated according to the importance of features at each level to achieve adaptive fusion of multi-scale features and output fusion feature maps. Among them, level 0 processes the minimum scale feature map to characterize the fine details of small targets; level 1 processes the medium-scale feature map to enhance the perception of medium-sized targets; level 2 focuses on large-scale feature maps to capture contextual information of large targets; level 3 manages macro-scale feature maps to assist the network in aggregating cross-scale information to ensure the integrity of scene understanding. Each level learns independent weight coefficients of specific scales through convolutional layers to achieve feature fusion, and uses the Softmax function to normalize the fused feature map to ensure that features of different scales are weighted and fused according to their relative importance. The process is expressed as:

[0113] W i =f weight (F i )

[0114]

[0115] Among them, f weight represents the learnable convolution operation, F i is the feature map of the i-th level, W i is the unnormalized weight tensor of the i-th level, α i is the normalized weight coefficient of the i-th level after cross-level normalization by the Softmax function, F Mfused is the fusion feature map. First, we can learn the convolution operation f weight is the feature map F of the i-th level i Generate unnormalized weight tensor W i , and then normalize across levels through the Softmax function to obtain the normalized weight coefficient α i , and then the multi-scale feature dynamic fusion is realized through weighted sum operation, and the weight coefficient α i Adaptively adjust the contribution of features at each scale.

[0116] Afterwards, the fusion feature map F Mfused Perform dimension expansion and adjust the channel dimension under the premise of strictly maintaining the spatial resolution to obtain the output feature map F out :

[0117] F out =f expand (F Mfused )

[0118] Among them, f expand Represents the dimension expansion function.

[0119] DFL significantly improves the localization accuracy of small objects by assigning the regression task to multiple discrete spatial locations. Unlike the single-location regression mechanism of traditional focal loss, DFL employs a multi-location distributed regression strategy, enabling more accurate bounding box coordinate prediction and optimizing small object localization performance. This loss function achieves innovation through three core mechanisms: first, assigning the bounding box regression task to a set of discrete locations, then using the Softmax function to assign weight coefficients across spatial locations, and finally achieving coordinate prediction through probability distribution modeling.

[0120] Specifically, the local feature vector F corresponding to each candidate box is extracted from the output feature map out1 , based on the DFL function, generate bounding box prediction results:

[0121] box=DFL(F out1 )×norm

[0122] Among them, DFL is the distribution focus loss function, norm is the normalization factor to enhance numerical stability, and box represents the regression prediction of the bounding box.

[0123] At the same time, the output feature map F out Apply the final convolution operation:

[0124] output=Conv(F out )

[0125] Among them, output represents the prediction result of class probability, and Conv represents the final convolution operation.

[0126] The bounding box regression prediction and class probability prediction results are combined as the final output of the MscaleASFFHead detection head.

[0127] The core of the small object detection optimization mechanism lies in the fact that small objects in high-resolution images typically occupy only a limited area, making their fine features difficult to capture using traditional methods. The MscaleASFFHead detection head effectively integrates cross-scale contextual feature information through multi-scale feature fusion, significantly improving small object recognition. Its implementation involves a three-stage process: first, weighted aggregation of features at different scales is performed to achieve multi-scale fusion and output. The fusion result is then refined through a convolutional layer, ultimately generating a joint prediction of bounding box coordinates and classification probabilities.

[0128] In this example, the DPM-YOLO ship detection network was trained using the high-resolution SAR dataset (HRSID). The dataset was split into 70% training, 20% testing, and 10% validation sets, with input images fixed at 640×640 pixels. Mosaic data augmentation was employed during training and disabled after the 10th training epoch to prevent overfitting. Stochastic Gradient Descent (SGD) was used as the optimizer, with a fixed learning rate of 0.01 and a momentum coefficient of 0.937.

[0129] Figure 6 This is the visualization result of the ablation experiment in this invention. The letters a, b, and c in the picture represent the three modules DPSConv, PromptFusionMod, and MscaleASFFHead, respectively.

[0130] Table 1 Ablation experiment results

[0131]

[0132]

[0133] In order to more comprehensively evaluate the contribution of the DPSConv, PromptFusionMod, and MscaleASFFHead modules, this embodiment conducts an ablation experiment, and the results are summarized in Table 1. This table systematically presents the impact of introducing each module individually or in combination into the YOLOv11s model on detection performance and computational efficiency. The evaluation indicators include precision (P), recall (R), average precision (mAP) at a threshold of 0.5 intersection over union (IoU), and the average precision (mAP) at a threshold of 0.5. 0.5 ), mean average precision (mAP) between 0.5 and 0.95 intersection-over-union thresholds 0.5:0.95 ) and frames per second (FPS). The results in the table show that after integrating the DPSConv module into the baseline YOLOv11s model, P is improved to 92.3%, and all mean average precision (mAP) indicators are improved, but the inference speed drops to 370.37fps, indicating an increase in computational overhead. When the PromptFusionMod module is introduced alone, R is improved to 85.6%, and mAP0.5 and mAP 0.5:0.95After integrating the MscaleASFFHead module, the performance is further improved, with the precision reaching 92.1%, the recall rate rising to 87.6%, and the mAP 0.5 When the three modules are used in conjunction, the precision and recall rates reach 93.8% and 89.9% respectively, and the mAP 0.5 Improved to 94.9%, mAP 0.5:0.95 The accuracy increased to 72.9%, but the frame rate dropped to 156.25, which shows that the accuracy gain comes with higher resource requirements and slower inference speed. The experiment verified that the joint optimization strategy of DPSConv, PromptFusionMod and MscaleASFFHead can produce a synergistic effect and significantly improve the mAP of the YOLOv11s model. 0.5 and mAP 0.5:0.95 The performance of the above indicators is improved, thereby enhancing detection accuracy and generalization ability, and at the same time revealing the trade-off between improved accuracy and increased computational requirements.

[0134] Figure 7 This is the visual detection comparison result between the present invention and other algorithms, where the red box in the figure represents the ship target detected by the model.

[0135] Table 2 Comparative experimental results

[0136]

[0137]

[0138] As shown in Table 2, this embodiment systematically compares the performance of the proposed SAR image ship detection model with a variety of current advanced detection models, covering classic two-stage detectors based on the ResNet-50 backbone, including Faster-RCNN, Cascade-RCNN, Mask-RCNN, and lightweight single-stage detectors Rtmdet-tiny and Yolox-tiny using the CSPDarknet backbone. Overall, traditional two-stage detectors such as Cascade-RCNN and Mask-RCNN perform well in terms of accuracy, but their high computational complexity leads to slow inference speed, which is suitable for application scenarios where accuracy is prioritized. In contrast, lightweight networks such as Rtmdet-tiny and Yolox-tiny significantly improve inference efficiency while ensuring high accuracy, meeting real-time requirements. Further analysis of the tabular data shows that the model proposed in the present invention performs outstandingly in all indicators, with Precision reaching 94.3%, Recall being 90.0%, and mAP 0.5 and mAP0.5:0.95 The proposed method achieves 95.5% and 72.4% accuracy, respectively, with a computational complexity of 102.9 GFLOPs and a frame rate of 156.2 fps, demonstrating a good balance between accuracy and efficiency and possessing strong potential for practical application. This comparative analysis fully demonstrates the performance advantages and efficiency of the proposed method in the task of ship target detection in SAR images.

[0139] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A ship detection method for complex SAR scenes, characterized by: Abstract: In order to solve the problem of ship target detection failure, an experimental method based on the DPM-YOLO ship detection network was proposed. The high-resolution synthetic aperture radar image was input into the DPM-YOLO ship detection network and the ship target detection rendering was output. The DPM-YOLO ship detection network was improved based on the YOLOv11 network. The DPSConv module was introduced into the backbone network to capture multi-scale contextual information by using the receptive field while retaining subtle feature details. The PromptFusionMod module was introduced into the neck network to perform multimodal feature fusion through four processing stages: spatial compression and prompt fusion, efficient attention mechanism, lightweight multi-layer perceptron and output refinement. The original detection head was replaced by the MscaleASFFHead detection head.

2. The ship detection method for complex SAR scenes according to claim 1, characterized in that: In the DPM-YOLO ship detection network, the input features are preliminarily filtered and channel compressed by the CBS module and the C3K2 module to obtain low-level features, which are then processed by the first DPSConv module to obtain intermediate features. The intermediate features are normalized and activated by the CBS module and then input into the second DPSConv module. The output of the second DPSConv module is sent to the SPPF module for multi-scale maximum pooling parallel processing and feature aggregation. The features output by the SPPF module are sent to the C2PSA module, which uses the cross-stage partial attention mechanism to capture long-range dependencies and output high-level features. The low-level features, intermediate features, and high-level features are iterated upward step by step in the neck network through the upsampling-splicing-fusion process to construct a multi-level semantic representation, wherein the upsampling-splicing-fusion process is specifically as follows: the relatively low-level features are processed by the CBS module, the C3K2 module, and the PromptFusionMod module, and then upsampled, spliced ​​with the relatively high-level features, and normalized by the CBS1×1 module after splicing; The multi-scale features processed by the neck network are input into the corresponding MscaleASFFHead detection head respectively. Adaptive weights are dynamically assigned based on context information, and information from low, medium and high feature paths is integrated to achieve accurate positioning of targets of different sizes. The MscaleASFFHead detection heads of each scale output results in parallel and summarize them into the final ship prediction result.

3. The ship detection method for complex SAR scenes according to claim 1, characterized in that: The DPSConv module consists of two serially connected FConv submodules. The FConv submodule adopts a progressive receptive field expansion strategy, integrating the information bottleneck theory and the feature differentiation principle, and learns the optimal compressed representation of the input data through a dual-path structure. The dual-path structure operates as an implicit feature diversification mechanism, where each path is used to extract different subspace features.

4. The ship detection method for complex SAR scenes according to claim 3, characterized in that: The FConv submodule performs a dimensionality reduction projection transformation on its input feature map After that, the dual-path feature transformation is performed in sequence and Perform transformation and enhancement, and finally perform feature aggregation and dimension expansion transformation on the results of dual-path feature transformation The final output is obtained, and the process is expressed as: in, represents the FConv operation, Indicates the FConv operation with the bias set to 0, symbol represents the field of real numbers, represents the convolution operation, W D1 ~W D6 is the convolution kernel, b D is the bias term, BN(·) represents the batch normalization operation, φ(·) represents the activation function, ° represents the concatenation operation, X D It is the input feature map of the FConv submodule, Y1 to Y5 represent the intermediate tensors in the hierarchical feature transformation process, and the feature Y1 after the dimensionality reduction projection transformation of the input feature map is used as the shared source for dual-path processing. Heterogeneous nonlinear transformations are performed separately through parallel branches. Y2 and Y3 correspond to the first-level outputs of path A and path B respectively. Y4 and Y5 represent the final enhanced representation of the dual paths. High-order features with complementary discriminant characteristics are generated through cascaded convolution-normalization-activation operations. After processing, the output Y of the FConv submodule is obtained by fusion D .

5. The ship detection method for complex SAR scenes according to claim 1, characterized in that: The spatial compression and prompt fusion processing stage of the PromptFusionMod module first introduces the learnable prompt parameter QP and expands it: Where QP∈R 1×C×1×1 represents the learnable hint parameter, QP expande is the tensor of the learnable hint parameter after the expansion operation, B P ,C P ,H P ,W P Represent the batch size, number of channels, feature map height and width after expansion, respectively. Represents the real number domain, expand represents the expansion operation; Afterwards, the tensor after the expansion operation of the learnable prompt parameter and the input feature of the PromptFusionMod module are spliced ​​through the channel axis to generate the spliced ​​feature X concat : Among them, dim=1 means splicing operation along the channel dimension, X P Represents the input features of the PromptFusionMod module; Perform the following operations on the concatenated features to generate the fused features F: fused : Among them, σ SiL Represents SiLU activation function, Conv 1×1 represents a 1×1 convolution operation, and BN represents a batch normalization operation.

6. The ship detection method for complex SAR scenes according to claim 5, characterized in that: The efficient attention mechanism processing stage of the PromptFusionMod module performs the following steps in sequence: The fusion feature F output from the spatial compression and hint fusion stages fused Perform spatial downsampling and channel compression: Among them, AvgPool 2×2 represents a 2×2 average pooling operation, X down represents the downsampled feature map; Perform feature refinement on the downsampled feature map: Among them, X attn Represents the output feature map refined by channel attention, that is, the tensor after the deep convolution-activation-channel reweighting operation, DWConv 3×3 represents a 3×3 depth-separable convolution operation, GELU represents the Gaussian error linear unit activation function, represents the channel attention function; For the feature map X after feature refinement attn Perform spatial resolution restoration: Among them, A up Represents the attention refined feature map reconstructed by upsampling, ConvTranspose 2×2 It represents a 2×2 transposed convolution operation with a learnable kernel, and stride represents the step size parameter.

7. The ship detection method for complex SAR scenes according to claim 6, characterized in that: The lightweight multi-layer perceptron processing stage of the PromptFusionMod module processes the fusion feature F output by the spatial compression and prompt fusion processing stage. fused And the attention refined feature map A output by the efficient attention mechanism processing stage up Perform element-wise addition and then process it through a lightweight MLP network: Among them, X add Represents the feature map after element-by-element addition, X mlp Represents the output feature map of the lightweight multilayer perceptron processing stage.

8. The ship detection method for complex SAR scenes according to claim 7, characterized in that: The output refinement processing stage of the PromptFusionMod module is used to process the output feature map X of the lightweight multi-layer perceptron processing stage. mlp Processing is performed and the final output is generated through residual connection and convolution refinement network: Among them, X res Represents the residual feature map, C Pout Indicates the number of output channels, Y P Represents the output of the PromptFusionMod module.

9. The ship detection method for complex SAR scenes according to claim 1, characterized in that: The MscaleASFFHead detection head adopts a four-level feature fusion strategy based on the ASFF module, processing feature maps at four independent levels. Each level corresponds to a different target scale, and dynamically allocates fusion weights according to the importance of features at each level to achieve adaptive fusion of multi-scale features and output fused feature maps. Among them, level 0 processes the minimum-scale feature map to characterize the fine details of small targets; level 1 processes the medium-scale feature map to enhance the perception of medium-sized targets; level 2 focuses on large-scale feature maps to capture the contextual information of large targets; level 3 manages macro-scale feature maps to assist the network in aggregating cross-scale information to ensure the integrity of scene understanding; the process is expressed as: W i =f weight (F i ) Among them, f weight represents the learnable convolution operation, F i is the feature map of the i-th level, W i is the unnormalized weight tensor of the i-th level, α i is the normalized weight coefficient of the i-th level after cross-level normalization by the Softmax function, F Mfused is the fusion feature map.

10. The ship detection method for complex SAR scenes according to claim 9, characterized in that: The MscaleASFFHead detection head performs the following steps on the fused feature map to generate the output result: For the fusion feature map F Mfused Perform dimension expansion to obtain the output feature map F out : F out =f expand (F Mfused ) Among them, f expand represents the dimension expansion function; Extract the local feature vector F corresponding to each candidate box from the output feature map out1 Based on the DFL function, a multi-position distributed regression strategy is used to assign the bounding box regression task to a discrete position set, and then the Softmax function is used to assign weight coefficients on the spatial position. Finally, the coordinate prediction is realized through probability distribution modeling to generate the bounding box prediction result: box=DFL(F out1 )×norm Among them, DFL is the distribution focus loss function, norm is the normalization factor to enhance numerical stability, and box represents the regression prediction of the bounding box; At the same time, the output feature map F out Apply the final convolution operation: output=Conv(F out ) Among them, output represents the prediction result of class probability, and Conv represents the final convolution operation; The bounding box regression prediction and class probability prediction results are combined as the final output of the MscaleASFFHead detection head.

Citation Information

Patent Citations

  • Lightweight SAR ship detection method and system

    CN116844055A

Cited By

  • Synthetic aperture radar image target detection method and device, and medium

    CN121213893A

  • River reach dike personnel intrusion intelligent identification method and device based on target detection

    CN121482722A

  • Lightweight multi-scale SAR image target detection method based on improved YOLO architecture

    CN121746691A

  • SAR ship target detection method in complex environment

    CN122049706A

  • Ship target detection method based on Mama and YOLOv11 fusion

    CN122049707A