Improved YOLOv11-based SAR image ship small target robust detection method and system, storage medium and electronic equipment

Through the improved YOLOv11 model, combined with S2-MLPv2, SAFMN and SimAM modules, the TSIoU loss function is designed and slice-assisted reasoning strategy is applied, which solves the problems of noise interference, insufficient feature extraction and poor multi-scale adaptability of ship small target detection in SAR images, and achieves high-precision real-time detection.

CN120374955APending Publication Date: 2025-07-25HENAN UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510502993.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In existing SAR images, small ship target detection faces problems such as strong noise interference, insufficient feature extraction, poor multi-scale adaptability and low model convergence efficiency, making it difficult to achieve high-precision real-time detection in complex marine environments.

Method used

The improved YOLOv11 model is adopted to suppress spot noise by embedding the S2-MLPv2 module at the end of the backbone network, and a multi-scale receptive field is constructed with the SAFMN module, and the SimAM module is embedded to enhance feature expression, and the TSIoU loss function is introduced to optimize bounding box regression, and combined with slice-assisted reasoning strategies to improve detection recall.

Benefits of technology

High-precision real-time detection is realized on SSDD and HRSID datasets, with an average accuracy of 98.1% and 93.8%, significantly improving the detection performance of small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374955A_ABST
    Figure CN120374955A_ABST
Patent Text Reader

Abstract

The invention discloses an SAR image ship small target robust detection method and system based on improved YOLOv11, a storage medium and electronic equipment, and the method comprises the following steps: embedding an S2-MLPv2 module at the tail end of a backbone network, and suppressing speckle noise through spatial displacement operation and dynamic filtering; the SAFMN module is used for replacing traditional up-sampling, and the multi-scale feature fusion capability is enhanced in combination with a deformable convolution and gating mechanism; a SimAM non-parameter attention module is introduced, and ship texture and geometric features are adaptively focused based on an energy function; designing a TSIoU loss function, fusing a central point diagonal distance measurement and an end point distance measurement, and optimizing bounding box regression precision; a slice auxiliary reasoning strategy is applied, and the small target recall rate is increased through multi-scale slicing and parallel detection. The method solves the problems of strong noise interference, insufficient small target feature extraction, poor multi-scale adaptability and low convergence efficiency in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of synthetic aperture radar (SAR) image processing, and particularly to a robust detection method, system, storage medium and electronic device for small ship targets in SAR images based on an improved YOLOv11, which is applicable to the recognition of small targets in low-resolution and high-noise interference scenarios in complex marine environments. Background Art

[0002] Synthetic aperture radar (SAR), as an all-weather and high-resolution active microwave remote sensing technology, has important application values in fields such as ship monitoring and marine traffic management. However, the detection of small ship targets in SAR images faces severe challenges: on the one hand, the SAR imaging mechanism is vulnerable to speckle noise interference, resulting in a decline in image quality. Especially in complex sea conditions or low signal-to-noise ratio environments, it is difficult to distinguish target features from background noise; on the other hand, small ship targets occupy few pixels and have weak texture information in the image. Traditional convolutional neural networks are prone to losing details during deep feature extraction, leading to missed detections or false detections. In addition, the size differences of multi-scale ship targets are significant. Existing detection models often have difficulty effectively capturing the spatial correlations of targets at different scales during the feature fusion process due to fixed receptive fields or single upsampling methods, further restricting the detection accuracy.

[0003] Although current object detection algorithms based on deep learning have made some progress, they still have limitations. Single-stage detection models represented by the YOLO series have the advantage of real-time performance, but their backbone networks have insufficient noise suppression capabilities, and gradient disappearance problems are prone to occur during the construction of the feature pyramid in the neck network, resulting in insufficient expression of small target features. Two-stage detection models (such as Faster R-CNN) improve the accuracy through the region proposal mechanism, but have high computational complexity and are difficult to meet the real-time requirements in practical applications. In addition, traditional loss functions (such as CIoU) have ambiguity in the measurement of bounding box regression, resulting in slow model convergence speed and limited positioning accuracy. Existing attention mechanisms (such as SE, CBAM) often introduce additional parameters, increasing the model complexity and affecting the deployment efficiency. Therefore, how to improve the detection robustness and multi-scale adaptability of small targets in SAR images while ensuring real-time performance remains a technical problem to be solved urgently. Summary of the Invention

[0004] The purpose of the present invention is to provide a robust detection method, system, storage medium and electronic device for small ship targets in SAR images based on an improved YOLOv11. Aiming at the problems of strong noise interference, insufficient extraction of small target features, poor multi-scale adaptability and low model convergence efficiency in the prior art, through the optimization of the multi-level network architecture and the innovation of the loss function, high-precision real-time detection of small ship targets in complex environments is realized.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A robust detection method for small targets of ships in SAR images based on the improved YOLOv11 includes the following steps:

[0007] Step S101, data preparation and preprocessing;

[0008] Step S102, improvement based on the YOLOv11 model: embed the S2-MLPv2 module at the end of the backbone network. The S2-MLPv2 cyclically translates the feature map along the channel dimension through spatial displacement operations, establishes long-range dependencies across image patches, and combines dynamic filtering units to suppress speckle noise, outputting a high-quality feature map with 256 output channels; design the TSIoU loss function, integrate the center point diagonal distance metric and the endpoint distance metric into the distance loss function and the shape loss function of SIoU respectively, optimize the bounding box regression accuracy and accelerate the model convergence;

[0009] Step S103, train and optimize the improved model using the dataset: adopt the AdamW optimizer and the cosine annealing strategy to train the improved model, optimize the parameters through the backpropagation of the TSIoU loss function, monitor and save the weights with the best performance on the validation set;

[0010] Step S104, send the test set into the network model with the best training results for testing: aiming at the problem of feature weakening in small target detection, introduce the slice-assisted inference SAHI strategy, perform multi-scale segmentation and fusion on the input image through the sliding window mechanism, and significantly improve the detection recall rate of dense small targets.

[0011] The specific steps of the above-mentioned step S101 include the following steps: convert the original annotation into a TXT file in YOLO format, including normalized bounding box coordinates and class indices; this process involves normalizing the coordinates to the [0,1] interval and filtering invalid or damaged labels; keep the original aspect ratio of the image and pad the edges to a fixed size; simulate multi-target scenarios through Mosaic enhancement and image geometric changes to improve the small target detection ability.

[0012] In the above-mentioned step S102, the following steps are further included. In the neck network, replace the traditional bilinear upsampling with the SAFMN module, the core of which is a deformable convolutional kernel with a size of 3×3 and a dilation rate of 2 and a gating unit: the deformable convolution dynamically adjusts the receptive field to capture multi-scale targets, and the gating unit weights the feature transfer path through the sigmoid function to reduce gradient disappearance.

[0013] The specific steps of step S102 further include the following steps. In the feature fusion stage, the SimAM parameter-free attention module is embedded, and the three-dimensional attention weight is calculated based on the energy function to focus on the ship texture and geometric features and suppress background interference.

[0014] The improvement of the loss function in step S102 specifically includes the following steps: replacing the CIoU loss function of the YOLOv11 network model with the TSIoU loss function optimized based on SIoU, and using the ratio of the diagonal of the center points of the ground truth box and the predicted box to the diagonal of the smallest enclosing rectangle as the measure of the center point diagonal distance:

[0015]

[0016] Added to the SIoU distance loss function Δ, where W and H are the width and height of the smallest enclosing rectangle, and (x, x gt ) and (y, y gt ) are the center points of the target box and the predicted box respectively. The relevant calculation formula is:

[0017]

[0018] λ x , λ y is the measurement method of the original SIoU distance loss function:

[0019]

[0020] ∧ is the angular loss function, θ is the angle between the diagonal of the two center points and the abscissa, and the calculation formula is:

[0021] Λ = sin2θ

[0022] The ratio of the distances between the four end points of the predicted bounding box and the ground truth bounding box to the diagonal of the smallest enclosing matrix:

[0023]

[0024] Added to the SIoU shape Ω loss function, where w and h are the width and height of the predicted box, and w gt and h gt are the width and height of the ground truth box; the relevant calculation formula is as follows:

[0025]

[0026] w w , w h is the measurement method of the original SIoU shape loss function:

[0027]

[0028] Ci , where \(i = 1, 2, 3, 4\) are the Euclidean distances between the four endpoints of the predicted bounding box and the ground truth bounding box, respectively, so as to achieve the entire shape using the two sides of length and width and the four endpoints; the TSIoU loss function is defined as:

[0029]

[0030] In step 3 during the training phase, the AdamW optimizer is used, which specifically includes the following steps: the initial learning rate is set to \(1e - 4\), and the learning rate is dynamically adjusted in combination with the cosine annealing strategy, with the minimum learning rate dropping to \(1e - 6\); the batch size is set to 16, and the input image resolution is uniformly adjusted to \(800\times800\); the loss function weight distribution is classification loss: confidence loss: regression loss = 1:1:2 to strengthen the bounding box regression optimization; the training cycle is 200 epochs, and the model performance is evaluated on the validation set every 10 epochs, and the weights with the highest mAP are saved; the data augmentation strategy includes randomly horizontal flipping with a probability of 0.5, rotating by \(45^{\circ}\) with a probability of 0.3, and brightness jitter with an amplitude of \(\pm20\%\) to improve the generalization ability of the model to complex scenes.

[0031] Step S104 specifically includes the following steps: The test set is sent to the optimally trained network model for testing. The Slice-assisted Inference (SAHI) strategy is applied to perform multi-scale segmentation and parallel detection on the test set images. After fusing the subgraph results, redundant bounding boxes are filtered through Non-Maximum Suppression (NMS) to output the final detection results.

[0032] The robust small target detection system for SAR images based on the improved YOLOv11 includes the following units:

[0033] Preprocessing unit: Convert the original annotation into a YOLO-format TXT file, which contains the normalized bounding box coordinates and class indices. Keep the original aspect ratio of the image and pad the edges to a fixed size; enhance the ability to detect small targets by simulating multi-target scenarios through Mosaic augmentation and geometric transformations of the image;

[0034] Model construction unit: It includes an improved CSPDarknet53 backbone network, with an S2-MLPv2 module embedded at the end. Establish long-range dependencies through hierarchical spatial displacement operations, and combine dynamic filtering to suppress speckle noise in SAR images; the neck network replaces the traditional upsampling operation with the SAFMN module, constructs a multi-scale receptive field using deformable convolutional kernels, and dynamically adjusts the feature transmission path through a gating mechanism; embed the SimAM parameter-free attention module, derive three-dimensional attention weights based on the energy function, and focus on the texture and geometric features of ship targets; the output layer integrates the TSIoU loss function, and optimizes the bounding box regression accuracy by fusing the center point diagonal distance metric and the endpoint distance metric;

[0035] Training unit: The AdamW optimizer is adopted, and the initial learning rate is set to 1e-4. The learning rate is dynamically adjusted in combination with the cosine annealing strategy. The network parameters are optimized by backpropagation through the TSIoU loss function, and the model weights with the highest mean average precision (mAP) on the validation set are monitored and saved.

[0036] Inference unit: The Slice-assisted Inference (SAHI) strategy is applied to divide the input image into 256×256 overlapping sub-images, and the results are fused after parallel detection. The Non-Maximum Suppression (NMS) algorithm is used to filter redundant bounding boxes, and the target location, category, and confidence information are output, supporting visual display and data export.

[0037] A storage device stores multiple programs, and the programs are loaded and executed by a processor to implement the robust detection method for small ship targets in SAR images based on the improved YOLOv11.

[0038] An electronic device includes: a memory, a processor, and a program stored in the memory and executable on the processor. The processor executes the robust detection method for small ship targets in SAR images based on the improved YOLOv11.

[0039] In the present invention, the S2-MLPv2 module is embedded at the end of the backbone network to establish long-range dependencies through hierarchical spatial displacement operations, and the speckle noise in SAR images is suppressed by combining the dynamic filtering mechanism to improve the quality of feature extraction. In the neck network, the Spatial Adaptive Feature Modulation Network (SAFMN) is used to replace the traditional upsampling operation, and a multi-scale receptive field is constructed through deformable convolution. Combining with the gating mechanism, the feature transfer path is dynamically adjusted to alleviate the gradient disappearance and information loss problems during multi-scale feature fusion. Further, the SimAM parameter-free attention module is embedded in the neck network to derive the three-dimensional attention weights of the feature map through the energy function, and the network can be guided to focus on the texture and geometric structure features of ship targets without introducing additional parameters, enhancing the feature expression ability of key regions. Aiming at the feature weakening problem in small target detection, the Slice-assisted Inference (SAHI) strategy is introduced, and the input image is multi-scale segmented and fused through the sliding window mechanism, significantly improving the detection recall rate of dense small targets. In addition, the TSIoU loss function is designed, and the center point diagonal distance metric and the endpoint distance metric are respectively incorporated into the distance loss function and the shape loss function of SIoU to optimize the bounding box regression accuracy and accelerate the model convergence. The technical effects of the present invention are remarkable: Experiments on the public datasets SSDD and HRSID show that the average precision (mAP) of the improved YOLOv11-SST model reaches 98.1% and 93.8% respectively, providing an efficient and reliable solution for ship detection in SAR images. Description of the Drawings

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0041] Figure 1 : Network structure diagram of the improved YOLOv11 of the present invention;

[0042] Figure 2 : Schematic diagram of the S2-MLPv2 module structure of the present invention;

[0043] Figure 3 : Schematic diagram of the SAFMN module structure of the present invention;

[0044] Figure 4 : Schematic diagram of the SimAM module structure of the present invention;

[0045] Figure 5 : Schematic diagram of the SAHI module structure of the present invention;

[0046] Figure 6 : Visualization effect of small target detection on the HRSID dataset of the present invention;

[0047] Figure 7 : Comparative experiment results of the model of the present invention with other methods on the SSDD dataset;

[0048] Figure 8 : Comparative experiment results of the model of the present invention with other methods on the HRSID dataset. Detailed implementation manners

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0050] Step S101, data preparation and preprocessing;

[0051] Specifically, convert the original annotations into YOLO-format TXT files, including normalized bounding box coordinates and class indices. This process involves normalizing the coordinates to the [0,1] interval and filtering invalid or damaged labels. Keep the original aspect ratio of the image and pad the edges to a fixed size. Simulate multi-object scenarios through Mosaic augmentation and image geometric transformations to improve the small target detection ability.

[0052] Step S102: Improve based on the YOLOv11 model;

[0053] Specifically, the overall architecture of the model improvement is as Figure 1 shown. The backbone network adopts an improved CSPDarknet53 architecture, and the S2-MLPv2 module is embedded at the end. S2-MLPv2 cyclically translates the feature map along the channel dimension through spatial displacement operations, establishes long-range dependencies across image patches, and combines a dynamic filtering unit (implemented by 1×1 convolution and a gating mechanism) to suppress speckle noise, outputting a high-quality feature map with 256 output channels. Its module structure is shown in the appendix Figure 2 . In the neck network, the SAFMN module replaces the traditional bilinear upsampling. Its core is a deformable convolution kernel (3×3, dilation rate 2) and a gating unit: the deformable convolution dynamically adjusts the receptive field to capture multi-scale targets, and the gating unit weights the feature transfer path through the sigmoid function to reduce gradient disappearance. Its module structure is shown in the appendix Figure 3 . In the feature fusion stage, the SimAM parameter-free attention module is embedded, which calculates three-dimensional attention weights based on an energy function, focuses on ship texture and geometric features, and suppresses background interference. Its module structure is shown in the appendix Figure 4 .

[0054] In this example, the loss function CIoU of the YOLOv11 network model is replaced with the loss function TSIoU optimized based on SIoU, and the ratio of the diagonal of the center points of the ground truth box and the predicted box to the diagonal of the minimum bounding rectangle is used as the measure of the center point diagonal distance:

[0055]

[0056] is added to the SIoU distance loss function Δ, where W and H are the width and height of the minimum bounding rectangle, (x, x gt ) and (y, y gt ) are the center points of the target box and the predicted box respectively. The relevant calculation formula is:

[0057]

[0058] λ x , λ y is the measurement method of the original SIoU distance loss function:

[0059]

[0060] ∧ is the angular loss function, θ is the angle between the diagonal of the two center points and the abscissa, and the calculation formula is:

[0061] Λ = sin2θ

[0062] The ratio of the four endpoint distances between the predicted bounding box and the ground truth bounding box to the diagonal of the minimum enclosing matrix:

[0063]

[0064] Add it to the shape Ω loss function of SIoU. w and h are the width and height of the predicted bounding box, w gt and h gt are the width and height of the ground truth bounding box. The relevant calculation formula is as follows:

[0065]

[0066] w w , w h is the measurement method of the original SIoU shape loss function:

[0067]

[0068] C i , i = 1, 2, 3, 4 are the Euclidean distances between the four endpoints of the predicted bounding box and the ground truth bounding box in turn, so as to achieve the whole shape by the two sides of length and width and the four endpoints. The TSIoU loss function is defined as:

[0069]

[0070] TSIoU takes into account the vector angle between regressions, and can effectively avoid the aspect ratio ambiguity problem of CIoU. Add the distance between the center point diagonal and the diagonal of the enclosing rectangle to the SIoU distance loss function Δ to achieve a dual distance measurement, improve the detection accuracy of the network model, and add the ratio of the four endpoint distances between the ground truth bounding box and the predicted bounding box to the diagonal of the minimum enclosing rectangle to the SIoU shape function to optimize the measurement standard and accelerate the convergence speed of the network model.

[0071] Step S103: Use the dataset to train the improved model;

[0072] Specifically, the AdamW optimizer is used in the training stage, the initial learning rate is set to 1e-4, the learning rate is dynamically adjusted in combination with the cosine annealing strategy, and the minimum learning rate is reduced to 1e-6. The batch size is set to 16, and the input image resolution is uniformly adjusted to 800×800. The loss function weight assignment is classification loss: confidence loss: regression loss = 1:1:2, strengthening the bounding box regression optimization. The training period is 200 epochs, and the model performance is evaluated on the validation set every 10 epochs, and the weight with the highest mAP is saved. The data augmentation strategy includes random horizontal flipping (probability 0.5), 45° rotation (probability 0.3) and brightness jitter (amplitude ±20%), to improve the generalization ability of the model to complex scenes.

[0073] Step S104: Feed the test set into the optimally trained network model for testing;

[0074] Specifically, as Figure 5 shown, in the inference stage, the Slice-assisted Inference (SAHI) strategy is applied. The input image is sliced into 256×256 overlapping sub-images and fed into the model in parallel for detection. After the detection results of the sub-images are mapped to the original image by coordinates, non-maximum suppression (NMS) is used to filter redundant bounding boxes. The intersection-over-union threshold is set to 0.5, and the confidence threshold is set to 0.25. As Figure 7 shown, the quantitative evaluation of the SSDD dataset shows that the proposed SST-YOLO method outperforms other state-of-the-art models. The accuracy is 94.26%, and the recall is 94.48%, which are 0.53% and 2.55% higher than the accuracy and recall of YOLOv8 respectively. Compared with other methods such as FasterR-CNN and Cascade R-CNN, the precision of SST-YOLO is improved by 7.25% and 0.13% respectively, and the recall is improved by 4.05% and 3.64% respectively. The AP50 metric also proves the superiority of SST-YOLO, reaching 98.17%, which is 4.56% higher than the next best-performing model, YOLOv8. This model has the advantage of being lightweight, with only 24.32M Params and 63.37G FLOP, achieving a balance between detection accuracy and computational efficiency. Further experiments were conducted on the HRSID dataset, and the quantitative evaluation results are as Figure 8 shown. The precision of this method is 92.76%, the recall is 87.31%, and the AP50 is 93.82%, which are significantly higher than YOLOv8 and other comparison methods, further verifying its better detection accuracy and performance in SAR small target detection.

[0075] Exemplary System

[0076] An embodiment of this application provides a robust detection system for small ship targets in SAR images based on the improved YOLOv11, which consists of a preprocessing unit, a model construction unit, a training unit, and an inference unit. Each unit works together to achieve the robust detection of small ship targets in complex environments. The preprocessing unit is responsible for normalizing the original SAR image.

[0077] The model construction unit is the core module. Its backbone network adopts an improved CSPDarknet53 architecture, and the S2-MLPv2 module is embedded at the end. Long-range dependencies are established through hierarchical spatial displacement operations, and combined with a dynamic filtering mechanism to effectively suppress the speckle noise of SAR images, and a high signal-to-noise ratio feature map is output. The neck network replaces the traditional upsampling operation with the SAFMN module, constructs a multi-scale receptive field using deformable convolutional kernels, and dynamically adjusts the feature transfer path through a gating mechanism, significantly alleviating the problems of gradient disappearance and information loss. In the feature fusion stage, the SimAM parameter-free attention module is embedded, and three-dimensional attention weights are derived based on the energy function to guide the network to focus on the texture features and geometric structures of ship targets and suppress background noise interference. The output layer integrates the TSIoU loss function, optimizes the bounding box regression accuracy by fusing the center point diagonal distance metric and the endpoint distance metric, and improves the accuracy of detection and positioning.

[0078] The training unit uses the AdamW optimizer, sets the initial learning rate to 1e-4, and combines the cosine annealing strategy to dynamically adjust the learning rate to accelerate convergence. The network parameters are optimized by backpropagation through the TSIoU loss function, and the mean average precision (mAP) and recall rate metrics during the training process are monitored in real time to ensure the stability of the model performance. After training is completed, the model weights with the best performance on the validation set are saved to provide a reliable basis for subsequent inference and deployment.

[0079] The inference unit applies the slice-assisted inference (SAHI) strategy, divides the high-resolution input image into overlapping 256×256 sub-images, and inputs them into the model in parallel for detection, making full use of multi-scale information to enhance the feature expression of small targets. After fusing the detection results of the sub-images, the non-maximum suppression (NMS) algorithm is used to filter redundant bounding boxes, and the intersection over union (IoU) threshold is set to 0.5 to ensure the simplicity and accuracy of the output results. Finally, the system outputs the detection results including the target location, confidence, and category information, supporting visualization display and data export functions to meet the real-time monitoring requirements in practical applications.

[0080] The SAR image ship small target robust detection system based on the improved YOLOv11 provided by the embodiments of this application can implement the SAR image ship small target detection steps and processes and achieve the same technical effects, which will not be elaborated here one by one.

[0081] Exemplary Device

[0082] This application provides an electronic device, including: a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the SAR image ship small target robust detection method based on the improved YOLOv11.

[0083] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

[0084] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the device where the computer-readable storage medium is located executes the method for robust detection of small targets of ships in SAR images based on the improved YOLOv11 as described above. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory and other memories, etc.

[0085] An electronic device, including: a memory and a processor. A program that can run on the processor is stored on the memory. When the processor executes the program, it implements the method for robust detection of small targets of ships in SAR images based on the improved YOLOv11 as described above.

[0086] If the modules / units integrated in the electronic device of the present application are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by a computer program instructing relevant hardware devices. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented.

[0087] Furthermore, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of blockchain nodes, etc.

[0088] Computer-readable instructions are stored in the computer-readable storage medium. The computer-readable instructions are executed by a processor in the electronic device to implement the method for robust detection of small targets of ships in SAR images based on the improved YOLOv11 described in any of the above embodiments. The present invention well solves the problems of strong noise interference, insufficient extraction of small target features, poor multi-scale adaptability and low convergence efficiency in the prior art.

[0089] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0090] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0091] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

[0092] It should be noted that the terms "comprising" and "having" in the specification and claims of the present application, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0093] Note that the above is only the preferred embodiment of the present invention and the application of the technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention is described in detail through the above embodiments, the present invention is not limited to the specific embodiments described herein. Without departing from the concept of the present invention, more other effective embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

[0094] It should be noted that the terms "comprising" and "having" in the specification and claims of the present application, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

Claims

1. A robust detection method for small ship targets in SAR images based on the improved YOLOv11, characterized in that: It includes the following steps: Step S101, data preparation and preprocessing; Step S102, improvement based on the YOLOv11 model: Embed the S2-MLPv2 module at the end of the backbone network. The S2-MLPv2 cyclically translates the feature map along the channel dimension through spatial displacement operations, establishes long-range dependencies across image patches, and combines dynamic filtering units to suppress speckle noise, outputting a high-quality feature map with 256 output channels; Design the TSIoU loss function, incorporate the center point diagonal distance metric and the endpoint distance metric into the distance loss function and shape loss function of SIoU respectively, optimize the bounding box regression accuracy and accelerate model convergence; Step S103, use the dataset to train and optimize the improved model: Adopt the AdamW optimizer and cosine annealing strategy to train the improved model, optimize the parameters through backpropagation of the TSIoU loss function, and monitor and save the weights with the best performance on the validation set; Step S104, send the test set into the network model with the best training results for testing: Aiming at the problem of feature weakening in small target detection, introduce the slice-assisted inference SAHI strategy, perform multi-scale segmentation and fusion on the input image through the sliding window mechanism, and significantly improve the detection recall rate of dense small targets.

2. The robust detection method for small targets of ships in SAR images based on the improved YOLOv11 according to claim 1, wherein, The specific steps of the above-mentioned step S101 include the following steps: Convert the original annotation into a TXT file in YOLO format, including normalized bounding box coordinates and class indices; This process involves normalizing the coordinates to the [0,1] interval and filtering invalid or damaged labels; Keep the original aspect ratio of the image and pad the edges to a fixed size; Simulate multi-target scenarios through Mosaic augmentation and image geometric transformation to improve the small target detection ability.

3. The robust detection method for small targets of ships in SAR images based on the improved YOLOv11 according to claim 1, characterized in that, In the above-mentioned step S102, it also includes the following steps. In the neck network, replace the traditional bilinear upsampling with the SAFMN module, whose core is a deformable convolution kernel with a size of 3×3 and a dilation rate of 2 and a gating unit: The deformable convolution dynamically adjusts the receptive field to capture multi-scale targets, and the gating unit weights the feature transfer path through the sigmoid function to reduce gradient disappearance.

4. The robust detection method for small ship targets in SAR images based on the improved YOLOv11 according to claim 1, characterized in that, The specific steps of the above-mentioned step S102 also include the following steps. In the feature fusion stage, embed the SimAM parameter-free attention module, calculate the three-dimensional attention weight based on the energy function, focus on the ship texture and geometric features, and suppress background interference.

5. The robust detection method for small ship targets in SAR images based on the improved YOLOv11 according to claim 1, wherein The improvement of the loss function in the above-mentioned step S102 specifically includes the following steps: Replace the CIoU loss function of the YOLOv11 network model with the TSIoU loss function optimized based on SIoU, and use the ratio of the diagonal of the center points of the true box and the predicted box to the diagonal of the minimum enclosing rectangle as the center point diagonal distance metric: Added to the SIoU distance loss function Δ, W and H are the width and height of the minimum bounding rectangle, and (x, x gt ) and (y, y gt ) are the center points of the target box and the predicted box respectively. The relevant calculation formula is as follows: λ x , λ y is the measurement method of the original SIoU distance loss function: ∧ is the angular loss function, θ is the angle between the diagonal of the two center points and the abscissa, and the calculation formula is: Λ = sin2θ Use the ratio of the four endpoint distances between the predicted bounding box and the true bounding box to the diagonal of the minimum enclosing matrix: Added to the shape Ω loss function of SIoU, w and h are the width and height of the predicted bounding box, w gt and h gt are the width and height of the ground truth bounding box; the relevant calculation formula is as follows: w w ,w h is the measurement method of the original SIoU shape loss function: C i , where \(i = 1, 2, 3, 4\) are the Euclidean distances between the four endpoints of the predicted box and the ground truth box, respectively, so as to achieve the whole shape by using the two sides of length and width and the four endpoints; the TSIoU loss function is defined as:

6. The robust detection method for small targets of ships in SAR images based on the improved YOLOv11 according to claim 1, wherein In the training stage of step 3, the AdamW optimizer is adopted, which specifically includes the following steps: the initial learning rate is set to 1e-4, and the learning rate is dynamically adjusted in combination with the cosine annealing strategy, with the minimum learning rate reduced to 1e-6; the batch size is set to 16, and the input image resolution is uniformly adjusted to 800×800; the loss function weight is allocated as classification loss: confidence loss: regression loss = 1:1:2 to strengthen the optimization of bounding box regression. The training period is 200 epochs, and the model performance is evaluated on the validation set every 10 epochs, and the weights with the highest mAP are saved; the data augmentation strategy includes random horizontal flipping with a probability of 0.5, 45° rotation with a probability of 0.3, and brightness jitter with an amplitude of ±20% to improve the generalization ability of the model to complex scenes.

7. The robust detection method for small ship targets in SAR images based on the improved YOLOv11 according to claim 1, characterized in that, Step S104 specifically includes the following steps: The test set is sent into the optimal trained network model for testing. The Slice-assisted Inference (SAHI) strategy is applied to perform multi-scale segmentation and parallel detection on the test set images. After fusing the subgraph results, redundant bounding boxes are filtered through Non-Maximum Suppression (NMS) to output the final detection results.

8. A robust detection system for small ship targets in SAR images based on the improved YOLOv11, characterized in that, It includes the following units: Preprocessing unit: Convert the original annotation into a TXT file in YOLO format, which contains normalized bounding box coordinates and class indices. Keep the original aspect ratio of the image and pad the edges to a fixed size; enhance the ability to detect small targets by simulating multi-object scenarios through Mosaic augmentation and image geometric transformation. Model construction unit: It includes an improved CSPDarknet53 backbone network, and the S2-MLPv2 module is embedded at the end to establish long-range dependencies through hierarchical spatial displacement operations, and combined with dynamic filtering to suppress SAR image speckle noise; the neck network replaces the traditional upsampling operation with the SAFMN module, constructs a multi-scale receptive field using deformable convolutional kernels, and dynamically adjusts the feature transfer path through a gating mechanism. Embed the SimAM parameter-free attention module, derive the three-dimensional attention weights based on the energy function, and focus on the texture and geometric features of ship targets; the output layer integrates the TSIoU loss function, and fuses the center point diagonal distance metric and the endpoint distance metric to optimize the bounding box regression accuracy. Training unit: Adopt the AdamW optimizer, set the initial learning rate to 1e-4, and dynamically adjust the learning rate in combination with the cosine annealing strategy; optimize the network parameters through the backpropagation of the TSIoU loss function, monitor and save the model weights with the highest mean average precision (mAP) on the validation set. Inference unit: Apply the Slice-assisted Inference (SAHI) strategy to divide the input image into 256×256 overlapping subgraphs, fuse the results after parallel detection; use the Non-Maximum Suppression (NMS) algorithm to filter redundant bounding boxes, and output the target position, class, and confidence information, supporting visual display and data export.

9. A storage device that stores multiple programs, characterized in that, The described program is loaded and executed by a processor to implement the robust detection method for small ship targets in SAR images according to any one of claims 1-7.

10. An electronic device, characterized in that, It includes: A memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method for robust detection of small targets of ships in SAR images based on the improved YOLOv11 as described in any one of claims 1-7.

Citation Information

Cited By

  • Infrared fusion target detection and identification method under complex low-light background condition

    CN121190736A

  • Track foreign matter identification method and device

    CN121640367A

  • Multi-frame weak and small target detection method based on prior guidance and space-time cooperative constraint

    CN122244428A