SAR image ship target detection method based on improved YOLOv8
By improving the YOLOv8 algorithm and adopting the Soft-NMS and CBAM attention mechanisms, the detection problem of small and dense targets in SAR images is solved, the detection accuracy and real-time performance are improved, and the robustness and generalization ability of the model are enhanced.
Patent Information
- Application Number
- CN202510784416.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-19
AI Technical Summary
Existing SAR image ship target detection methods have low detection accuracy and false detection and missed detection when dealing with small targets and dense targets. In addition, the model has weak generalization ability and is difficult to adapt to complex marine environments.
The improved YOLOv8 algorithm is adopted, and the standard NMS is replaced by Soft-NMS. The CBAM attention mechanism and SENet are integrated to suppress sea clutter, optimize detection box screening and feature extraction, and combine CIoU Loss and BCE Loss for target detection.
It improves the detection accuracy and real-time performance in dense scenes, enhances the robustness and generalization ability of the model, and can effectively identify small targets and ship targets in complex backgrounds.
Smart Images

Figure CN120673229A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and remote sensing image processing, and in particular to a method for detecting ship targets in SAR images based on an improved YOLOv8. Background Art
[0002] Synthetic Aperture Radar (SAR) has attracted considerable attention in the field of ship detection due to its all-day, all-weather operation and surface penetration capabilities. Traditional methods for detecting ship targets in SAR images rely primarily on manually designed features. For example, these methods utilize gray-level co-occurrence matrices and texture analysis, combined with classic machine learning algorithms, to perform image preprocessing such as filtering, denoising, and segmentation, followed by ship target detection through pattern recognition methods. However, these methods have significant limitations: manually designed features struggle to comprehensively and accurately describe the complex characteristics of ship targets, resulting in insufficient feature expression capabilities; when processing densely populated scenes such as ports and waterways, they are unable to effectively distinguish between closely spaced targets, resulting in poor processing capabilities; and, due to the high complexity of the algorithms, their real-time performance struggles to meet practical application requirements and makes them difficult to adapt to rapidly changing ocean environments.
[0003] In recent years, deep learning has been widely used in the field of ship detection using SAR images due to its powerful feature extraction and pattern recognition capabilities. It automatically extracts detailed textures and features from images for target prediction, resulting in improvements in detection efficiency and accuracy compared to traditional methods. However, it still faces many challenges: for smaller ship targets that occupy a small proportion of the image, the network struggles to extract sufficient effective features, resulting in poor detection performance; in densely populated scenes, false detections and missed detections are prone to occur; and due to the complex and changing imaging conditions of SAR images (such as varying sea conditions, imaging angles, and noise interference), the model's generalization ability is weak, resulting in unstable detection results in new scenarios.
[0004] Therefore, developing a method that can effectively solve the problems of missed detection of small targets, false detection of dense targets and sea clutter interference is a technical problem that technicians in this field urgently need to solve.
[0005] The information disclosed in this background technology section is only intended to enhance understanding of the overall background of the invention and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art already known to a person skilled in the art. Summary of the Invention
[0006] In response to the above technical problems, an embodiment of the present invention provides a SAR image ship target detection method based on improved YOLOv8 to solve the problems raised in the above background technology.
[0007] The present invention provides the following technical solution: a method for detecting ship targets in SAR images based on improved YOLOv8, comprising the following steps:
[0008] Select the target dataset and divide it into training set and test set according to the ratio of 7:3;
[0009] Preprocess the SAR images in the target dataset;
[0010] Construct an improved YOLOv8 algorithm model, which is based on the YOLOv8 network architecture, adopts Soft-NMS (flexible non-maximum suppression) instead of standard NMS (non-maximum suppression), and incorporates the CBAM (convolutional block attention module) attention mechanism;
[0011] Input the target data set into the improved YOLOv8 algorithm model for training until the training is completed to obtain the optimized YOLOv8 algorithm model;
[0012] The optimized YOLOv8 algorithm model is used to identify target ships in SAR images.
[0013] Preferably, the target dataset is obtained from the SAR-Ship-Dataset and SSDD public datasets; SAR-Ship-Dataset and SSDD are both open source datasets that contain a large number of SAR images.
[0014] Preferably, the SAR image preprocessing method includes: normalization, multi-scale enhancement and sea clutter suppression.
[0015] Preferably, sea clutter suppression uses SENet (channel attention mechanism) to generate channel descriptors through global average pooling, and dynamically allocates feature channel weights through a fully connected layer and a Sigmoid function to suppress sea clutter interference.
[0016] Preferably, the specific formula of SENet is:
[0017]
[0018] Among them, F sq (·) refers to the compression stage of SENet, namely global average pooling, which compresses the H×W×C feature map U containing global information into a 1×1×C statistic z c ;
[0019] z c is the output channel statistics vector, H and W are the height and width of the feature map respectively, u c(i, j) represents the activation value of the cth channel at the spatial position (i, j), reflecting the characteristic response strength of the position.
[0020] The above operation compresses the spatial dimension H×W to 1×1, forming a global context representation for each channel. In the excitation phase, by constructing a multi-layer perceptron with a bottleneck structure, the dimension of z is first compressed (dimensionality reduction coefficient r), and then the channel attention weight s is generated through the ReLU activation function and the Sigmoid function:
[0021] s=σ(W2δ(W1z));
[0022] Where W1 and W2 are learnable parameters, δ(·) represents the ReLU activation function, and σ is the Sigmoid function. Finally, the attention weight s is multiplied by the original feature map at the channel level to complete the feature recalibration:
[0023]
[0024] Preferably, the confidence decay formula of Soft-NMS is:
[0025]
[0026] Among them, σ is the attenuation coefficient, which is dynamically adjusted according to the dense scene. When the IoU exceeds the threshold, the attenuation strength is controlled by adjusting the σ value to retain potential targets in the dense area.
[0027] Preferably, the operation process of the CBAM attention mechanism realizes fine adjustment of the input feature map through the sequential processing of the channel attention submodule and the spatial attention submodule; the specific implementation is as follows:
[0028] Channel attention: on the input feature map F∈R C×H×W , perform global average pooling and maximum pooling respectively to generate two channel descriptors z avg and z max , generate the weight vector w by sharing the fully connected layer c , the formula is:
[0029] w c =σ(w1·z avg +w2·z max );
[0030] Among them, w c is the generated channel attention map, σ is the sigmoid function, w1 is the weight matrix of the first fully connected layer, z avg is the channel description vector obtained by global average pooling, w1 is the weight matrix of the second fully connected layer, z max is the channel description vector obtained by global maximum pooling;
[0031] Spatial attention: concatenate the channel attention output feature maps along the channel dimension and generate a spatial mask w through 7×7 convolution s ∈R 1×H×W , the formula is:
[0032] w s =σ(f 7×7 ([F avg ; F max ]));
[0033] Among them, w s is the generated spatial attention map, σ is the sigmoid function, f 7×7 is a 7×7 convolution operation, F avg yes
[0034] The spatial description vector obtained by channel average pooling, F max is the spatial description vector obtained by channel maximum pooling, [;] represents the splicing operation in the channel dimension;
[0035] The final output features are:
[0036] F out =w s ·(w c F);
[0037] Among them, F out is the final weighted feature map, w s is the generated spatial attention map, w c is the generated channel attention map, F is the input feature map, and [·] is the element-wise multiplication.
[0038] Preferably, the core formula of the YOLOv8 algorithm includes:
[0039] Bounding box prediction: CIoU Loss is used to optimize the detection box regression. The formula is:
[0040]
[0041] Let the detection box be A, the real box be B, and IoU represents the intersection-over-union ratio of A and B. IoU is the complete IoU (CIoU) loss, ρ is the Euclidean distance between the detection box and the center point of the real box, b is the center point of the detection box, and b gt is the center point of the real box, c is the minimum diagonal length of the bounding box, α is the weight coefficient, and v is the aspect ratio difference between the predicted box and the real box;
[0042] Classification loss: binary cross entropy loss (BCE Loss) is used, the formula is:
[0043] L cls =-∑(ylogp+(1-y)log(1-p);
[0044] Among them, L cls is the cross entropy loss function, y is the target label, represents the intersection over union (IoU) of the detection box and the true box for positive samples, and is 0 for negative samples; p is the probability value of the model predicting the existence of the target.
[0045] Preferably, the training method in the YOLOv8 algorithm model adopts alternating iterative optimization, specifically:
[0046] Fix YOLOv8 parameters and train the model until convergence;
[0047] Repeat the above steps until the total loss function is stable.
[0048] The embodiment of the present invention provides a method for detecting ship targets in SAR images based on an improved YOLOv8, which has the following beneficial effects:
[0049] (1) This paper uses Soft-NMS instead of standard NMS for the basic network YOLOv8, optimizes the detection box screening in dense scenes, and effectively improves the detection accuracy of target overlapping areas;
[0050] (2) At the same time, the CBAM attention mechanism is integrated to enhance the robustness and generalization ability of the algorithm;
[0051] (3) The improved algorithm shows high detection accuracy and real-time performance on both SAR-Ship and SSDD datasets, and also demonstrates excellent results in dense target scenarios, verifying the effectiveness of the coordinated optimization of the attention mechanism and the dynamic suppression strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 The detection accuracy of the unimproved YOLOv8 model on the SAR-Ship-Dataset dataset;
[0053] Figure 2 The detection accuracy of the improved YOLOv8 model on the SAR-Ship-Dataset dataset;
[0054] Figure 3 is the detection accuracy of the unimproved YOLOv8 model on the SSDD dataset;
[0055] Figure 4 The detection accuracy of the improved YOLOv8 model on the SSDD dataset;
[0056] Figure 5The improved YOLOv8 model of the present invention performs well in dense scenes.
[0057] Figure 6 This is the detection effect of the improved YOLOv8 model in the present invention on small targets. DETAILED DESCRIPTION
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0059] In response to the problems mentioned in the above background technology, the embodiments of the present invention provide a method for detecting ship targets in SAR images based on an improved YOLOv8 to solve the above technical problems. The technical solution is as follows:
[0060] The following is combined with Figure 1-6 , and specific implementation methods are further described to illustrate the present invention.
[0061] This example uses the SAR-Ship-Dataset and SSDD public datasets as examples to test the optimized model on these two datasets. First, install Python programming software on your terminal. This requires Python 3.10, tensoR, flow 2.16.1, and keras 3.0.5.
[0062] tensoR, flow, and keras are open source software libraries for deep learning. tensoR and flow provide a flexible platform for building, training, and deploying various complex neural network models, supporting multiple programming languages. Keras is an advanced neural network API that provides a simpler and more user-friendly interface based on tensoR and flow, allowing users to quickly build and run common neural network models.
[0063] 1. A method for detecting ship targets in SAR images based on improved YOLOv8, comprising the following steps:
[0064] The target dataset is selected from the SAR-Ship-Dataset and SSDD public datasets, and divided into training and test sets according to a 7:3 ratio;
[0065] Preprocess the SAR images in the target dataset. SAR image preprocessing methods include normalization, multi-scale enhancement, and sea clutter suppression. SENet is used to suppress sea clutter by generating channel descriptors through global average pooling, and dynamically assigning feature channel weights through a fully connected layer and a sigmoid function to suppress sea clutter interference.
[0066] Build an improved YOLOv8 algorithm model, which is based on the YOLOv8 network architecture, uses Soft-NMS instead of standard NMS, and incorporates the CBAM attention mechanism;
[0067] The target dataset is input into the improved YOLOv8 algorithm model for training. The training method in the YOLOv8 algorithm model adopts alternating iterative optimization, specifically:
[0068] Fix the YOLOv8 parameters and train the model until convergence; repeat the above steps until the total loss function stabilizes;
[0069] After training, the optimized YOLOv8 algorithm model is obtained;
[0070] The optimized model was tested on the public datasets SAR-Ship-Dataset and SSDD to identify ship targets in SAR images.
[0071] 2. Test Results
[0072] Figure 1-4 As shown in the figure, the Soft-NMS and CBAM modules improve the mAP of the two datasets by 2.0% and 3.5% respectively, proving that channel attention can effectively enhance the discriminative feature extraction of ship targets in SAR images, especially maintaining robust performance under speckle noise interference; the accuracy is significantly improved.
[0073] Figure 5 As shown in the figure, the improved model demonstrates the detection effect in dense scenes. The detection frame overlap conflicts are significantly reduced, and adjacent ships can be accurately distinguished. Ship targets are accurately identified, avoiding misjudgments and missed detections caused by dense targets.
[0074] Figure 6 As shown in the figure, the improved model more accurately locates the detection box for small targets and can effectively identify small ship targets that are easily missed by traditional algorithms even in complex sea conditions. The improved model performs well on datasets such as SSDD, which are dominated by small targets, because the CBAM module enhances high-frequency details in the shallow network through feature recalibration.
[0075] Although the introduction of the CBAM module resulted in a slight decrease in inference speed, the model size only increased by 2.9%, and it still maintained a real-time processing capability of 55FPS on the 3050 graphics card, meeting the needs of engineering deployment.
[0076] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A method for ship target detection in SAR images based on improved YOLOv8, characterized in that: The following steps are involved: Select the target dataset and divide it into training set and test set according to the ratio of 7:3; Preprocess the SAR images in the target dataset; Build an improved YOLOv8 algorithm model, which is based on the YOLOv8 network architecture, uses Soft-NMS instead of standard NMS, and incorporates the CBAM attention mechanism; Input the target data set into the improved YOLOv8 algorithm model for training until the training is completed to obtain the optimized YOLOv8 algorithm model; The optimized YOLOv8 algorithm model is used to identify target ships in SAR images.
2. The SAR image ship target detection method based on improved YOLOv8 according to claim 1 is characterized in that The target dataset is obtained from the SAR-Ship-Dataset and SSDD public datasets.
3. The SAR image ship target detection method based on improved YOLOv8 according to claim 1 is characterized in that SAR image preprocessing methods include normalization, multi-scale enhancement and sea clutter suppression.
4. The SAR image ship target detection method based on improved YOLOv8 according to claim 3 is characterized in that Sea clutter suppression uses SENet to generate channel descriptors through global average pooling, and dynamically assigns feature channel weights through a fully connected layer and a Sigmoid function to suppress sea clutter interference.
5. The SAR image ship target detection method based on improved YOLOv8 according to claim 4 is characterized in that: The specific formula of SENet is: Among them, F sq (·) refers to the compression stage of SENet, namely global average pooling, which compresses the H×W×C feature map U containing global information into a 1×1×C statistic z c ; z c is the output channel statistics vector, H and W are the height and width of the feature map respectively, u c (i, j) represents the activation value of the cth channel at the spatial position (i, j), reflecting the characteristic response strength of the position; The above operation compresses the spatial dimension H×W to 1×1, forming a global context representation for each channel. In the excitation phase, by constructing a multi-layer perceptron with a bottleneck structure, the dimension of z is first compressed (dimensionality reduction coefficient r), and then the channel attention weight s is generated through the ReLU activation function and the Sigmoid function: s=σ(W2δ(W1z)); Where W1 and W2 are learnable parameters, δ(·) represents the ReLU activation function, and σ is the Sigmoid function. Finally, the attention weight s is multiplied by the original feature map at the channel level to complete the feature recalibration:
6. The SAR image ship target detection method based on improved YOLOv8 according to claim 1 is characterized in that The confidence decay formula of Soft-NMS is: Among them, σ is the attenuation coefficient, which is dynamically adjusted according to the dense scene. When the IoU exceeds the threshold, the attenuation strength is controlled by adjusting the σ value to retain potential targets in the dense area.
7. The SAR image ship target detection method based on improved YOLOv8 according to claim 1 is characterized in that The CBAM attention mechanism operates through the sequential processing of the channel attention submodule and the spatial attention submodule to achieve fine-tuning of the input feature map. The specific implementation is as follows: Channel attention: on the input feature map F∈R C×H×W , perform global average pooling and maximum pooling respectively to generate two channel descriptors z avg and z max , generate the weight vector w by sharing the fully connected layer c , the formula is: In c =σ(w1 z avg +w2·z max ); Among them, w c is the generated channel attention map, σ is the sigmoid function, w1 is the weight matrix of the first fully connected layer, z avg is the channel description vector obtained by global average pooling, w1 is the weight matrix of the second fully connected layer, z max is the channel description vector obtained by global maximum pooling; Spatial attention: concatenate the channel attention output feature maps along the channel dimension and generate a spatial mask w through 7×7 convolution s ∈R 1×H×W , the formula is: w s =σ(f 7×7 ([F avg ;F max ])); Among them, w s is the generated spatial attention map, σ is the sigmoid function, f 7×7 is a 7×7 convolution operation, F avg yes The spatial description vector obtained by channel average pooling, F max is the spatial description vector obtained by channel maximum pooling, [;] represents the splicing operation in the channel dimension; The final output features are: F out =w s ·(w c ·F); Among them, F out is the final weighted feature map, w s is the generated spatial attention map, w c is the generated channel attention map, F is the input feature map, and [·] is the element-wise multiplication.
8. The SAR image ship target detection method based on improved YOLOv8 according to claim 1 is characterized in that: The core formula of the YOLOv8 algorithm includes: Bounding box prediction: CIoU Loss is used to optimize the detection box regression. The formula is: Let the detection box be A, the real box be B, and IoU represents the intersection-over-union ratio of A and B; Among them, L IoU is the complete IoU (CIoU) loss, ρ is the Euclidean distance between the center of the detection box and the true box, b is the center of the detection box, bgt is the center of the true box, c is the minimum diagonal length of the bounding box, α is the weight coefficient, and v is the aspect ratio difference between the predicted box and the true box; Classification loss: binary cross entropy loss (BCE Loss) is used, the formula is: L cls =-∑(ylogp+(1-y)log(1-p); Among them, L cls is the cross entropy loss function, y is the target label, represents the intersection over union (IoU) of the detection box and the true box for positive samples, and is 0 for negative samples; p is the probability value of the model predicting the existence of the target.
9. The SAR image ship target detection method based on improved YOLOv8 according to claim 1, characterized in that: The training method in the YOLOv8 algorithm model adopts alternating iterative optimization. Specifically, the YOLOv8 parameters are fixed and the model is trained until convergence; the above steps are repeated until the total loss function stabilizes.
Citation Information
Patent Citations
Improved YOLOv8-based SAR image ship target detection method, medium and device
CN117689870A
Shielded tomato positioning identification method based on improved YOLOv8
CN118781482A
Multi-scale ship detection method for complex sea surface scene
CN119313941A
Novel aerial image detection method
CN119863720A
Method for detecting infrared ship target based on improved yolov7
US20250078541A1