SAR image ship target detection method based on YOLOv8
By designing the SAR image ship object detection method based on YOLOv8, using multi-scale feature extraction, dynamic upsampling and lightweight detection heads, the problems of high false alarm rate and missed alarm rate in SAR image object detection are solved, and high-precision and efficient detection effects are achieved.
Patent Information
- Application Number
- CN202510210158.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
In SAR image object detection, speckle noise and natural clutter interference feature extraction lead to high false alarm rates and missed alarm rates, which are difficult to effectively solve in the prior art.
A SAR image ship object detection method based on YOLOv8 was designed. The overall architecture is divided into multi-scale feature extraction backbone network, lightweight neck module and lightweight dynamic detection head network. This method improves detection accuracy and efficiency through multi-scale feature extraction, dynamic upsampling and lightweight detection heads.
It significantly improves the accuracy and efficiency of ship object detection in SAR images, reduces false alarm rates and missed alarm rates, and is suitable for complex backgrounds and resource-constrained scenarios.
Smart Images

Figure CN120047672A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a SAR image ship target detection method based on YOLOv8. Background Art
[0002] Synthetic Aperture Radar (SAR) is widely used in the fields of ocean monitoring, disaster management, environmental protection, and intelligent driving, especially in ship target detection. Although SAR images are affected by speckle noise and complex coastlines, their all-weather characteristics are superior to visible light and infrared images under low-light and high-cloud conditions, and they still play an important role in port real-time monitoring, ship identification, and tracking.
[0003] SAR ship detection algorithms include Constant False Alarm Rate (CFAR) and detection techniques based on Convolutional Neural Network (CNN). CFAR can effectively detect ship targets and adapt to environmental changes, but it relies on expert experience and has low efficiency in complex backgrounds. In contrast, CNN-based algorithms have adaptive feature learning, high precision, adaptability to complex backgrounds, and parallel processing capabilities, and have become the preferred choice for ship detection.
[0004] CNN-based object detection algorithms are generally divided into single-stage and two-stage methods. Common two-stage detection algorithms include Region-Based Convolutional Neural Network (R-CNN) and Faster Region-Based Convolutional Neural Network (Faster R-CNN). These algorithms generate candidate boxes and perform classification and bounding box regression on them. In contrast, single-stage detection algorithms such as YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), and Retina-Net usually have lower computational complexity and are more suitable for real-time applications. The YOLO series of models have significantly improved the detection speed and accuracy of SAR ship images in complex backgrounds and high-noise environments through efficient detection, lightweight architectures, multi-scale feature fusion, and attention mechanisms. In recent years, Huang et al. proposed a lightweight SAR ship detection algorithm based on YOLOv5, which improved the detection accuracy and efficiency through techniques such as channel pruning, knowledge distillation, and bidirectional feature pyramid networks. Zhang et al. significantly improved the detection performance based on YOLOv7 by combining a random attention module and dynamic snake convolution. The model proposed by Tang et al. based on YOLOv7 further improved the detection accuracy through multi-scale receptive field convolution blocks.
[0005] YOLOv8 is a high-performance object detector in the YOLO series. It uses CSPDarknet53 as the backbone network, improves the feature extraction ability through residual connections and feature fusion, and optimizes the network structure to enhance the inference speed. Although YOLOv8 performs excellently in terms of detection accuracy and speed, in SAR image object detection, speckle noise and natural clutter will interfere with feature extraction, resulting in high false alarm rates and missed alarm rates. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for detecting ship targets in SAR images based on YOLOv8, aiming to solve the problems.
[0007] To solve the above technical problems, the present invention provides a method for detecting ship targets in SAR images based on YOLOv8. Its overall architecture is divided into three parts: a multi-scale feature extraction backbone network, a lightweight neck module, and a lightweight dynamic detection head network;
[0008] S1. Multi-scale feature extraction backbone network
[0009] It is composed of ordinary convolution, cross-stage fusion module, double dilation residual module, and deformable attention module;
[0010] S2, lightweight neck module
[0011] Using lighter dynamic upsampling, by learning the location of sampling points, it is possible to more accurately restore the feature information of low-resolution images, thereby improving the ability to detect small targets and enhancing the model's feature parsing ability in complex backgrounds. The core idea is to use the pixel-level dynamic offset mechanism to make the upsampling process more flexible, adapt to targets of different scales and shapes, and improve the stability and robustness of detection.
[0012] S3, lightweight dynamic detection head network
[0013] The lightweight dynamic detection head optimizes the original YOLOv8 detection head and combines it with lightweight dynamic convolution technology to effectively reduce computational costs while enhancing feature perception capabilities. Compared with the traditional dynamic detection head, the lightweight dynamic detection head cuts redundant channels and computational complexity to improve the inference speed while maintaining high detection accuracy. It is suitable for scenarios with limited resources. This module adopts a task, scale, and space perception interaction mechanism, where:
[0014] π L (Scale-aware) Adjust the information transfer between feature layers through average pooling and convolution operations, making the model more adaptable to objects of different scales;
[0015] π S (Spatial perception) A dynamic offset mechanism is used to adjust the spatial distribution of feature maps so that the detection head can capture the target area more accurately;
[0016] π T Extract high-level task features through global pooling and fully connected networks, and generate adaptive channel weights to improve the synergy between different tasks;
[0017] π S , π L and π T It can be described by the following formula:
[0018] Z(F)=π T (π S (π L (F)·F)·F)·F
[0019] In formula (6), the feature tensor F∈R L×S×T , π S (·),π L (·) and π T (·) The three perception blocks are scale enhancement perception block, space enhancement perception block and task enhancement perception block;
[0020]
[0021] where G is the number of sparse sampling positions, p g +Δp g is the position shifted by self-learned spatial offset, and Δp g is learned by deformable convolution to focus on the ambiguous region, and Δm g is the importance measure of the self-learned position p g ;
[0022] π T (F)·F = max(ε 1 (F)·F t +β 1 (F), ε 2 (F)·F t +β 2 )(F))
[0023] where F t is the feature in the t-th segmentation channel, ε 1 , ε 2 , β 1 , β 2 are hyperparameters for learning to control the activation threshold.
[0024] Further preferably, in step S1, the specific implementation process is as follows:
[0025] S1.1 Dual dilation residual module
[0026] The dual dilation residual module replaces the original residual structure by combining the extended dilation residual and the extended reparameterized convolution, improves the multi-scale feature extraction ability, adopts the parallel structure of the extended reparameterized convolution with different dilation rates and convolution kernel sizes, optimizes the residual connection, effectively enhances the detection ability for sparse patterns, and the extended dilation residual reparameterized convolution module further combines the efficient feature extraction of the extended dilation residual convolution and the flexible multi-scale perception ability of the extended reparameterized convolution, can dynamically adjust the receptive field, improves the adaptability to targets of different scales and shapes, and thus enhances the robustness and accuracy of SAR ship target detection;
[0027] S1.2 Deformable attention mechanism module
[0028] The deformable attention mechanism module introduces the deformable attention mechanism on the basis of fast spatial pyramid pooling, uses dynamic sampling points to focus on key regions, and improves the detection ability of the model for small targets in complex backgrounds.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] First, to address the problems of large variations in target size and insufficient ability to extract multi-scale context information, the C2f-DD module is designed, significantly improving the multi-scale feature extraction efficiency and enhancing the feature extraction ability of the convolutional kernel layer.
[0031] Second, the SPA module focuses on key regions in the image by integrating dynamic offset and multi-head attention mechanisms, significantly improving accuracy, reducing computational complexity, and enhancing the robustness of the model in complex backgrounds.
[0032] Third, replacing upsampling with a lighter upsampling module helps the model reduce computational overhead and model complexity. This ensures that DD-YOLO remains lightweight in applications and effectively addresses the problems of complex background noise and detail loss.
[0033] Fourth, the DyHead-Prue module replaces the detection head, enhancing the model's sensitivity to key feature extraction and improving its adaptability to various target features and transformations in complex backgrounds, thereby improving the accuracy of object detection in multi-faceted backgrounds. Description of the Drawings
[0034] Figure 1 is the DD-YOLO network structure;
[0035] Figure 2 is the C2f-DD module;
[0036] Figure 3 is the DWR_DRB module;
[0037] Figure 4 is the extended reparameterization module;
[0038] Figure 5 is the deformable attention module;
[0039] Figure 6 is the dynamic upsampling;
[0040] Figure 7 is the sampling point generator;
[0041] Figure 8 is the DyHead-Prune structure;
[0042] Figure 9 is the SAR ship picture;
[0043] Figure 10 is the relationship between mAP50 and epoch;
[0044] Figure 11 is the visual comparison between YOLOv8 and the DD-YOLO model;
[0045] Figure 12It is a heat map comparison between YOLOv8 and DD-YOLO models. Detailed implementation mode
[0046] The following further elaborates on the present invention a in combination with the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, only for the purpose of conveniently and clearly assisting in explaining the embodiments of the present invention. The same or similar reference numerals in the drawings represent the same or similar components.
[0047] Embodiment
[0048] 1 Network design
[0049] The overall design architecture of the proposed DD-YOLO model is as Figure 1 shown. DD-YOLO is based on YOLOv8 and mainly consists of three parts: the DDDAT backbone network, the neck network, and the detection part.
[0050] The DDDAT backbone network has enhanced feature extraction and multi-scale semantic perception capabilities. This backbone network includes a Conv module, a C2f-DD module, and an SPA block. The C2f-DD module combines the advantages of dynamic adjustment and enhanced features of DWR and DRB, and the SPA module adds a Dattention attention mechanism after SPPF. The DDDAT backbone network enriches the gradient flow information through enhanced feature extraction, multi-scale language, and deformable attention mechanisms.
[0051] The neck part replaces the original upsampling of YOLOV8 with Dysample, enhancing the detection ability of low-resolution images and small targets, and further improving the feature extraction ability of the model in complex backgrounds.
[0052] The detection head part uses the DyHead-Prune detection head, which effectively improves the model performance while maintaining high detection accuracy through pruning optimization and multi-dimensional attention mechanisms.
[0053] 1.1 C2f-DD
[0054] In the SAR ship images, the shapes, positions, and sizes of the targets vary greatly, with uneven sizes and shapes. The original C2f module in YOLOv8 is insufficient in extracting the features of these targets and cannot effectively extract multi-scale context information. Therefore, to enhance the network's ability to extract multi-scale context information from the SAR ship image dataset, this paper designs a brand-new module - the C2f-DD module. C2f-DD mainly replaces the Bottleneck module with the DWR_DRB module and adjusts the forward propagation method to adapt to the new module structure. These changes endow C2f-DD with different characteristics and behaviors when processing input tensors, strengthening its multi-scale feature extraction and fusion capabilities. The structure of the C2f-DD module is as shown in Figure 2 shown below.
[0055] The working principle of the module is divided into two stages. The module is divided into two stages. The first stage is regional residualization, which generates a concise feature map through 3x3 convolution, batch normalization (BN), and ReLU activation. The second stage is semantic residual, which applies depth convolutions with different dilation rates for morphological filtering to ensure the diversity and integrity of feature representations.
[0056] The DWR has a relatively low complexity. Although it uses multiple dilated convolutions to achieve multi-scale feature extraction, its structure is relatively simple, and the diversity and complexity of feature extraction are limited. At the same time, due to the fixed convolution structure, it lacks flexibility and is difficult to adapt to the requirements of different tasks. In addition, its relatively simple design results in limited expressive power and capacity of the model, making it insufficient to handle very complex image tasks.
[0057] Given the deficiencies of the DWR module in terms of complexity and flexibility in feature extraction, we designed the DWR_DRB module. By adding reparameterized dilated convolutions to the DWR module, the DWR_DRB module enhances the fineness and diversity of multi-scale feature extraction, while improving the flexibility of the module, enabling it to better adapt to various complex tasks. This improvement not only overcomes the limitations of DWR but also significantly enhances the model's performance in handling complex image tasks. The structure of the DWR_DRB module is as shown in Figure 3 shown below.
[0058] The DRB module (Dilated Weighted Residual Dilated Reparam Block) enhances the large kernel convolutional layer through multiple layers of compact convolutional kernels and different expansion rates. Its key hyperparameters include the size k of the large convolutional kernel, the size of the parallel convolutional layer, and the dilation rate. Figure 4Shows the case of four parallel layers: k, r = (1, 2, 3, 4), k = (5, 5, 3, 3). For larger k, larger kernel convolutional layers or dilation layers with higher dilation rates can be adopted. The only limitation is k - 1 < r - 1 < k. The DRB structure diagram is shown in Figure 5.
[0059] The large kernel layer is created by merging the BN layer, converting the dilation rate, and zero-padding the kernel. The dilation layer is equivalent to a large sparse kernel layer, allowing the entire block to be converted into a large kernel convolution. Parallel small kernels capture small-scale patterns, and the outputs are added after the BN layer and merged through reparameterization. During inference, the large kernel and small kernels are equivalently combined. This method enhances the ability to detect sparse patterns and improves the feature quality. The dilated convolutional layer scans the input channels through the dilation rate to identify spatial patterns.
[0060] 1.2 SPA
[0061] The SPA module consists of adding the Dattention attention mechanism after SPPF. Dattention attention significantly reduces the computational amount and improves the model performance by introducing deformable attention and dynamic sampling points, and only focuses on a small part of the key regions in the image. Adding Dattention can improve the detection accuracy, reduce the computational amount, and enhance the robustness requirements of the model. These improvements enable the model to more accurately identify and locate ship targets in noisy and complex environments.
[0062] Figure 5 Shows the information flow of deformable attention: The reference points are evenly placed on the feature map, and the offsets are learned by the query through the offset network. The deformed keys and values are projected based on the deformed points, and the relative position bias is calculated to enhance the multi-head attention. The example in the figure shows 4 reference points, and there are more reference points in the actual implementation.
[0063] The input feature map x ∈ RH×W×C, where H represents the height, W represents the width, and C represents the number of channels. The feature map is linearly projected onto the query token q, expressed as q = xW o . The reference points on the upper path are downsampled using the scaling factor r and arranged into a grid. The offset Δp is obtained from the offset network with the query as the input, expressed as Δp = θ dffset(q) , and added to the reference points to obtain the offset position information Δp. The deformed reference points are sampled through the bilinear interpolation method δ(.) to obtain as follows:
[0064]
[0065] where, is the projection matrix. The multi-head attention can be calculated by adding relative position encoding, as follows:
[0066]
[0067] z = Concat(z (1) ,...,z (M) )W 0 (5)
[0068] where d is the dimension of each head, δ(.) represents the softmax function, and z (m) represents the embedded output of the m-th attention head. The final output z is the concatenation of the multi-head outputs z (m) .
[0069] 1.3 Dysample
[0070] In object detection tasks, upsampling is used to adjust the size of the feature map to match the original image, thus effectively detecting various objects. The traditional bilinear interpolation method may cause loss of image details, and the upsampling based on convolutional kernels has a large computational cost, which is not conducive to lightweight networks. Aiming at the problems of complex background noise and detail loss in SAR ship detection images, the present invention introduces Dysample, a lightweight and effective dynamic upsampler. Dysample performs upsampling through a point-based sampling method and a learned sampling perspective, avoiding dynamic convolution operations, reducing computational resources, and improving image resolution and model performance.
[0071] The network structure of Dysample is as Figure 6 , 7 shown. Its sampling set S consists of the original sampling grid FN and the generated offset G. The offset is generated using the "linear + pixel shuffle" method, and the range of the offset can be determined by static and dynamic factors. Specifically, given a feature map of size and an upsampling factor s, the feature map first passes through a linear layer with c input channels and 2s output channels 2 . Then it is reshaped using the pixel shuffle method into where 2 represents the x and y coordinates. Finally, an upsampled feature map of size is generated. 1.4 DyHead-Prune
[0072] DyHead-Prune innovatively integrates different object detection heads by introducing an attention mechanism to achieve scale π L , spatial π S and task π C perception interactions. Specifically, π L facilitates scale perception between different feature layers, π S achieves spatial perception between spatial positions, and π CPromote task perception within the output channel. These mechanisms work together to form the DyHead-Prune dynamic detection head module, effectively improving the performance and accuracy of object detection.
[0073] Z(F) = π T (π S (π L (F)·F)·F)·F (6)
[0074] In Equation (6), given the feature tensor F ∈ R L×S×T , π L (·), π S (·) and π T (·) of the three perception blocks are the scale enhancement perception block, the spatial enhancement perception block, and the task enhancement perception block, respectively.
[0075]
[0076] Where G is the number of sparse sampling positions, p g +Δp g is the position moved by the self-learned spatial offset, and Δp g is learned by deformable convolution to focus on the ambiguous regions. Δm g is the importance measure of the self-learned position p g .
[0077]
[0078] Where F t is the feature in the t-th split channel. ε 1 , ε 2 , β 1 , β 2 are hyperparameters for learning to control the activation threshold. As Figure 8 shown, at the final detection stage of the model, it perceives the three feature maps output by the neck part in different dimensions to obtain the result. Its structure diagram is as Figure 8 shown.
[0079] This layer-by-layer processing method ensures that features of different scales, spaces, and channels are fully optimized. After pruning, the number of model parameters is reduced and the memory footprint is decreased, enabling DyHead-Prune to significantly improve the efficiency and adaptability of the model while maintaining a high detection accuracy. The modular design makes the processing of each attention mechanism more explicit and independent, facilitating subsequent optimization and improvement.
[0080] 2 Experimental Results and Analysis
[0081] By conducting extensive experiments on two SAR image datasets of different scales, the effectiveness and advancement of this method are verified.
[0082] 2.1 Dataset
[0083] The SSDD dataset is the first publicly available dataset in the field of SAR ship detection. It includes 1,160 images from Sentinel-1, TerraSAR, and RadarSat-2. Its images cover a variety of different marine environments and conditions.
[0084] The HRSID dataset is a more complex and high-resolution collection. It includes 5,604 images taken by two satellite sensors, TerraSAR-X, TanDEM-X, and Sentinel-B. These images have been finely processed and filtered to provide more detailed and accurate target information.
[0085] As Figure 9 shown, (a)-(c) are from the SSDD dataset, and (d)-(f) are from the HRSID dataset. Among them, (a), (c), (d), and (f) represent images in complex scenes, (a), (c), (d), and (f) represent images in complex scenes, (b) and (e) represent images in dense scenes, and Figures (a) and (c) represent medium and large-sized ship data.
[0086] 2.2 Evaluation Metrics
[0087] To fairly compare the algorithm performance, the evaluation metrics of the COCO dataset are used in the experiment. The main metric is the average precision (AP), which is obtained from the area under the P-R curve. The mAP at IoU = 0.5 is called mAP 50 . Similarly, mAP 55 , mAP 60 , mAP 65 , mAP 70 , mAP 80 , mAP 85 , mAP 90 and mAP 95 are calculated. mAP 50-95 represents the average of these ten values and is usually denoted as AP, which is used to evaluate the overall performance of the method. Precision and recall are calculated as follows:
[0088]
[0089] where TP, FP, and FN represent the number of truly detected ship targets, the number of missed ship targets, and the number of non-ship targets misdetected as ship targets, respectively.
[0090] F1 is introduced to more comprehensively reflect the detection performance. Its calculation method is as follows:
[0091]
[0092] AP is calculated based on precision and recall, and the calculation method is as follows:
[0093]
[0094] Among them, p represents precision, r is recall, and p is a function with r as a parameter.
[0095] The number of parameters shows the number of weights and biases in the model. The fewer the number of parameters, the lighter the model weight, the smaller the occupied space, and the less computational resources required. The experiment also uses GFLOPs as an auxiliary evaluation metric to test the efficiency of the model. The fewer GFLOPs, the less computational resources the model consumes during the inference process.
[0096] 2.3 Experimental Configuration
[0097] The experimental setup includes a computer with an Intel Core i5-12600KF Processor, 16GRAM, NVIDIA GeForce RTX 4060Ti GPU (16GB memory), equipped with the Ubuntu 18.04 operating system. The network structure is built based on Pytorch 1.9.0, the programming language is Python 3.8, and CUDNN 8.0 and CUDA 11.1 are used for accelerated training. All two datasets are randomly divided into training set, validation set, and test set with a ratio of 7:2:1. The model is trained with 300 epochs, the input image size of the model is 640×640 pixels, and the batch size is set to 8.
[0098] 2.4 Ablation Experiment
[0099] To fully verify the effectiveness of the designed basic module, ablation experiments were conducted and their effects were carefully analyzed. Since the HRSID dataset contains more samples and higher complexity, which helps to more comprehensively evaluate the performance of the module and can also ensure the reliability and stability of the experimental results, ablation experiments were selected to be carried out on the HRSID dataset.
[0100] The present invention proposes C2f-DD with a new residual connection structure. To verify its characteristics, we successively replaced C2f in the backbone part with C2f-DD starting from the last one, and a total of 4 groups of experiments were conducted, which were respectively recorded as Experiment 1, 2, 3, and 4, corresponding to the number of replaced C2f. The experimental results are shown in Table 1. Considering the model complexity and detection accuracy comprehensively, we only replaced the last C2f in the backbone part in the overall model, which significantly improved the detection ability while reducing the computational overhead and model complexity.
[0101] Table 1 C2f-DD Ablation Experiment on HRSID
[0102]
[0103] As can be seen from Table 2, the C2F-DD module improves P, R, mAP 50 and mAP 50-95 by 0.6%, 1.1%, 0.8% and 0.8% respectively. The Dattention attention mechanism improves the performance of R, mAP 50 and mAP 50-95 by 0.6%, 1.1%, 0.5% and 0.7% respectively. The Dysample upsampling improves P, R, mAP 50 and mAP 50-95 by 0.1%, 1.3%, 0.4% and 0.7% respectively. The detection head module improves P, R, mAP 50 and mAP 50-95 by 0.5%, 1.4%, 1.2% and 2.4% respectively. The SPA module shows the most significant improvements in P and R, increasing by 0.6% and 1.4% respectively. The C2f-DD module not only improves the ability to extract multi-scale context information but also reduces GFLOPs. The Dysample module solves the problem of reducing the computational overhead and model complexity of the model, and reduces the number of parameters while keeping the GFLOPs unchanged.
[0104] At the same time, the results of the multi-module ablation experiment show that DD-YOLO has a significant improvement in terms of precision P, R, mAP 50 and mAP 50-95 , increasing by 1.5%, 2.5%, 2.0% and 3.4% respectively.
[0105] Figure 10 Intuitively shows the relationship between the detection accuracy (mAP 50 ) of YOLOv8 and DD-YOLO and the number of training epochs on the HRSID dataset. As the number of training epochs increases, the detection accuracy of DD-YOLO is always higher than that of YOLOv8.
[0106] Figure 11 The visual comparison in
[0107] Figure 12The heatmap analysis further validates the advantage of combining each module of DD-YOLO in dealing with dense multi-scale targets under complex background noise. The heatmap clearly shows the efficiency of the combined model in identifying and distinguishing crowded targets, which is attributed to the multi-scale feature extraction and fusion of the C2f-DD module and the SPA module, significantly enhancing the model's detection ability in complex environments. The Dysample module has proven its advantage in improving efficiency, and the addition of the DyHead-Prue module further enhances the model's ability to handle complex scenarios and target detection accuracy. This combination can not only improve precision and recall but also ensure optimized performance while maintaining low parameter counts and computational complexity.
[0108] Table 2 Ablation experiments of DD-YOLO on HRSID
[0109]
[0110] 2.5 Comparative experiments
[0111] As shown in Table 3, the experimental results on the SSDD dataset show that although the accuracy of the model's recall rate is lower than that of Key-PointEstimation+Channnel Attion(2024), TWC-Net(2021), and ADERLNet-CW(2024), DD-YOLO is superior to other classical models in both accuracy and detection precision. The model's parameters and GFLOPs metrics are 13.94 and 9.8 respectively, lower than those of other classical models, which can greatly reduce the computational overhead and model complexity.
[0112] Table 3 Comparison of different object detection models on the SSDD dataset
[0113]
[0114]
[0115] As shown in Table 4, the experimental results on the HRSID dataset show that the accuracy of the model's recall rate is slightly lower than that of PPA-Net and Context-aware network, and the recall rate is slightly lower than that of Center Net, PPA-Net, and Context-awarenetwork, but DD-YOLO is superior to other classical models in detection precision. The model's parameters and GFLOPs metrics are 13.94 and 9.8 respectively, far lower than those of other classical models. In short, the DD-YOLO model achieves remarkable detection accuracy, and the detection results on multiple datasets verify the fine generalization ability of this method.
[0116] Comparison of Different Object Detection Models for the HRSID Dataset
[0117]
[0118] 3 Summary
[0119] In view of the deficiencies in the ship detection task of SAR images, such as small ship targets, large variations in target sizes, and complex background noise, the C2f-DD module enhances the multi-scale feature extraction and fusion capabilities by combining the DWR and DRB modules, improving the detection effect for targets of different scales. The Dattention module enhances the model's attention to key regions in complex backgrounds through dynamic offset and multi-head attention mechanisms, improving accuracy and reducing computational complexity. The Dysample module optimizes the upsampling process, reducing computational resource consumption while ensuring effective detection of small and medium-sized targets. The DyHead-Prune module enhances the feature extraction ability of the detection head in complex backgrounds through pruning optimization and multi-dimensional attention mechanisms, improving detection accuracy and ensuring the efficient operation of the model. The method proposed in this paper can effectively cope with the interference of complex background noise and detect ship targets of different scales in complex backgrounds.
[0120] It should also be supplemented and explained that all "settings" and similar descriptive words in this application (especially in the specification) express a connection relationship between two structures, but there is no excessive limitation on the specific means of connection between the two. And it is usually a conventional connection means, that is, it should be understood that this means is the prior art and does not need to be elaborated. For example, "n is provided on m" only expresses that n structure exists on m structure, and whether the two are connected by welding, riveting, adhesive bonding or integrally formed is within the protection scope of this application; another example, "y is rotatably provided on x" only expresses that y and x can rotate relative to each other, and as for whether the two are rotationally connected by a bearing, or y directly passes through x and is rotationally connected to x, or other achievable ways, they are all within the protection scope of this application.
[0121] The above description is only a description of the preferred embodiments of the present invention, and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the field of the present invention based on the above disclosure are within the protection scope of the claims.
Claims
1. A method for detecting ship targets in SAR images based on YOLOv8, characterized in that: Its overall architecture is divided into three parts: a multi-scale feature extraction backbone network, a lightweight neck module, and a lightweight dynamic detection head network; S1. Multi-scale feature extraction backbone network It consists of ordinary convolution, cross-stage fusion module, double dilation residual module and deformable attention module; S2, lightweight neck module Using lighter dynamic upsampling, by learning the location of sampling points, it is possible to more accurately restore the feature information of low-resolution images, thereby improving the ability to detect small targets and enhancing the model's feature parsing ability in complex backgrounds. The core idea is to use the pixel-level dynamic offset mechanism to make the upsampling process more flexible, adapt to targets of different scales and shapes, and improve the stability and robustness of detection. S3, lightweight dynamic detection head network The lightweight dynamic detection head optimizes the original YOLOv8 detection head and combines it with lightweight dynamic convolution technology to effectively reduce computational costs while enhancing feature perception capabilities. Compared with the traditional dynamic detection head, the lightweight dynamic detection head cuts redundant channels and computational complexity to improve the inference speed while maintaining high detection accuracy. It is suitable for scenarios with limited resources. This module adopts a task, scale, and space perception interaction mechanism, where: π L (Scale-aware) Adjust the information transfer between feature layers through average pooling and convolution operations, making the model more adaptable to objects of different scales; π S (Spatial perception) A dynamic offset mechanism is used to adjust the spatial distribution of feature maps so that the detection head can capture the target area more accurately; π T Extract high-level task features through global pooling and fully connected networks, and generate adaptive channel weights to improve the synergy between different tasks; π S , π L and π T It can be described by the following formula: Z(F)=π T (p S (p L (F)·F)·F)·F In formula (6), the feature tensor F∈R L×S×T , π S (·),π L (·) and π T (·) The three perception blocks are scale enhancement perception block, space enhancement perception block and task enhancement perception block; Where G is the number of sparse sampling positions, p g +Δp g is the position moved by the self-learned spatial offset, Δp g Learned by deformable convolution, used to focus on ambiguous areas, Δm g is the self-learning position p g Importance measure of p T (F)·F=max(ε 1 (F)·F t +b 1 (F),e 2 (F)·F t +b 2 (F)) where F t is the feature of the t-th segmentation channel, ε 1 ,ε 2 ,β 1 ,β 2 is a hyperparameter that is learned to control the activation threshold.
2. The method for detecting ship targets in SAR images based on YOLOv8 according to claim 1, characterized in that: In step S1, the specific implementation process is as follows: S1.1 Double Dilated Residual Module The double-expanded residual module replaces the original residual structure by combining the expanded expanded residual and the expanded reparameterized convolution, improves the multi-scale feature extraction capability, and adopts the expanded reparameterized convolution parallel structure with different expansion rates and convolution kernel sizes to optimize the residual connection and effectively enhance the detection capability of sparse patterns. The expanded expanded residual reparameterized convolution module further combines the efficient feature extraction of the expanded expanded residual convolution with the flexible multi-scale perception capability of the expanded reparameterized convolution, and can dynamically adjust the receptive field, improve the adaptability to targets of different scales and shapes, thereby enhancing the robustness and accuracy of SAR ship target detection. S1.2 Deformable Attention Mechanism Module The deformable attention mechanism module introduces a deformable attention mechanism based on fast spatial pyramid pooling, uses dynamic sampling points to focus on key areas, and improves the model's ability to detect small targets in complex backgrounds.
Citation Information
Cited By
Cross-modal smoke shielding human body identification method and device based on double-engine cooperation
CN121170846A