High-precision lightweight method for ship target detection of synthetic aperture radar
By building a lightweight backbone network and optimizing the detection head structure, the YOLOv8 algorithm solves the problem of large amount of calculation and low detection accuracy in radar ship images, and achieves high-precision and lightweight ship target detection.
Patent Information
- Application Number
- CN202510390143.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The existing YOLOv8 algorithm is computationally expensive when detecting small targets in radar ship images, and the randomness and noise of synthetic aperture radar images affect detection accuracy, making it difficult to deploy on equipment with limited resources, and complex models increase computational costs.
Build a lightweight backbone network, introduce the GCADown module and ERC2f module, reduce the calculation amount through average pooling, maximum pooling and Ghost convolution, optimize the detection head structure, retain small target and medium target detection heads, and use path aggregation-feature pyramid structure for multi-scale fusion.
It realizes high-precision and lightweight detection in synthetic aperture radar images, reduces the calculation complexity and parameter quantity, and maintains high detection accuracy, and is suitable for equipment with limited resources.
Smart Images

Figure CN120259882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ship detection, and more specifically, to a high-precision lightweight method for synthetic aperture radar ship target detection. Background Art
[0002] Ship target detection technology has become increasingly important in modern marine activities. With the increasing busyness of maritime traffic and the rising demand for safety management, the need for accurate detection and identification of ships is constantly growing. This technology is widely applied in fields such as marine resource development and maritime safety, covering scenarios such as maritime traffic monitoring, marine resource development, and maritime search and rescue. In maritime traffic monitoring, the navigation status and position of ships can be monitored in real time, providing accurate information for management departments; in marine resource development, such as offshore oil extraction and fishing, the operating ships can be monitored to ensure safety and compliance; in maritime search and rescue missions, the position of the distressed ship can be quickly located, improving the search and rescue efficiency. However, ship target detection faces many challenges. The marine environment is complex and variable, and factors such as sea waves, sea fog, and light changes will interfere with detection and identification, increasing the difficulty. Ship targets in long-distance or low-resolution images often appear as small targets with unclear features, which are easily masked by background noise, posing higher requirements for detection algorithms. In addition, scenarios such as maritime traffic monitoring and maritime search and rescue require real-time acquisition of ship information, posing challenges to the running speed and efficiency of detection algorithms, and requiring fast detection while ensuring accuracy.
[0003] To address these challenges, ship target detection technology based on deep learning has been widely applied, and among them, the YOLO algorithm has become the first choice due to its efficient real-time detection ability. YOLOv8 was proposed by the Ultralytics team in January 2023. This algorithm consists of a backbone network, a neck network, and a detection head. In the backbone network, YOLOv8 replaces the C3 module in YOLOv5 with the C2f module, achieving further lightweighting. The C2f module divides the feature map into two parts, one part undergoes convolutional operations, and the other part is directly passed. Finally, the two parts of the features are fused, reducing the computational amount and enhancing the feature extraction ability. In the neck network, YOLOv8 uses a path aggregation-feature pyramid structure, which transmits high-level semantic information to low-level features through a top-down path, and at the same time transmits low-level detail information to high-level features through a bottom-up path, enhancing the fusion ability of multi-scale features. In the detection head part, for targets of different sizes, there are a total of three detection heads, and each detection head separately processes classification and regression tasks through decoupling, improving the detection accuracy.
[0004] However, the following disadvantages exist in the prior art:
[0005] 1. YOLOv8 is a general object detection model. Its design fully considers the problem of the diversity of object sizes in the object detection task. By outputting feature maps of different resolutions at different stages of the model, in radar ship images, small objects usually predominate. Therefore, adding a large object detection head in YOLOv8 increases the computational cost and the number of parameters, which is not conducive to deployment on devices with limited resources.
[0006] 2. Due to the randomness of the scattering characteristics of ground targets, the echo signal of the synthetic aperture radar has a certain randomness. This randomness will cause speckle noise in the SAR image. In addition, there are also problems such as non-linear set deformation and multi-scale objects, which make the visual effect of the image poor and thus affect the accuracy of object detection.
[0007] 3. Complex models often bring higher redundancy and computational cost, increasing the computational cost and being not conducive to deployment on some terminal devices with limited resources. The present invention uses the GCADown module to replace the convolutional layer, reducing the computational cost. Summary of the Invention
[0008] The purpose of the present invention is to provide a high-precision lightweight method for synthetic aperture radar ship target detection to solve the problems raised in the above background technology.
[0009] To achieve the above purpose, the present invention aims to provide a high-precision lightweight method for synthetic aperture radar ship target detection, including the following steps:
[0010] S1. Construct a lightweight backbone network: Based on the YOLOv8 algorithm, introduce the GCADown module and the ERC2f module. The GCADown module reduces the computational cost through the combination of average pooling, max pooling and Ghost convolution, and the ERC2f module adopts the EDWR structure to enhance the multi-scale feature fusion ability;
[0011] S2. Optimize the detection head structure: Remove the large object detection head in YOLOv8, retain the small object and medium object detection heads, and reduce the number of model parameters and computational complexity;
[0012] S3. Feature extraction and fusion: Perform multi-scale fusion on the feature maps output by the backbone network through the path aggregation-feature pyramid structure, and input them into the detection head to complete object detection.
[0013] As a further improvement of this technical solution, the specific operation of the GCADown module is:
[0014] The input feature map is successively subjected to average pooling and max pooling operations to generate feature maps of two branches respectively;
[0015] Each branch is processed by Ghost convolution, and the size of the output feature map is halved;
[0016] The final output feature maps are merged through a concatenation operation.
[0017] As a further improvement of this technical solution, the EDWR structure includes the following branches:
[0018] The first branch: uses a 1×1 convolution kernel with a dilation rate of 1 to enhance the non-linear expression ability;
[0019] The second branch: uses a 3×3 convolution kernel with a dilation rate of 3, combined with average pooling to extract local detailed features;
[0020] The third branch: uses a 5×5 convolution kernel with a dilation rate of 5 to expand the receptive field to capture context information;
[0021] The output feature maps of each branch are merged after adjusting the number of channels through concatenation and 1×1 convolution.
[0022] As a further improvement of this technical solution, the ERC2f module realizes multi-branch feature fusion by replacing the Bottleneck in the C2f module with the EDWR structure.
[0023] As a further improvement of this technical solution, the detection head only retains the small-target and medium-target detection heads, and decouples the classification and regression tasks to reduce the number of parameters and the amount of computation.
[0024] As a further improvement of this technical solution, the performance metrics of the method applied in synthetic aperture radar (SAR) images are:
[0025] On the SSDD dataset, the mAP@0.5:0.95 reaches 75.2%, the number of parameters is 1.57M, and the FLOPs is 6.5G;
[0026] On the iVision-MRSSD dataset, the mAP@0.5:0.95 reaches 61.8%, the number of parameters is 1.57M, and the FLOPs is 6.5G.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] 1. By reducing the detection head, the present invention makes the existing YOLOv8 algorithm more suitable for the radar ship image target detection task. However, reducing the detection head will cause a loss of accuracy. The present invention proposes the EDWR module to make up for this loss and constructs a lightweight and high-precision model.
[0029] 2. By proposing the GCADown module to replace the convolutional layer in YOLOv8, the present invention further reduces the complexity of the algorithm while maintaining the high accuracy of the algorithm. Brief Description of the Drawings
[0030] Figure 1 It is a schematic diagram of the overall algorithm structure of the present invention.
[0031] Figure 2 It is a schematic diagram of the GCADown structure of the present invention.
[0032] Figure 3 It is a schematic diagram of the ERC2f structure of the present invention.
[0033] Figure 4 It is a schematic diagram of the EDWR structure of the present invention. Detailed Embodiment
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0035] In a specific embodiment, the present invention provides a high-precision lightweight method for synthetic aperture radar ship target detection, which specifically includes the following steps:
[0036] S1. Construct a lightweight backbone network: Based on the YOLOv8 algorithm, introduce the GCADown module and the ERC2f module. The GCADown module reduces the computational complexity by combining average pooling, max pooling and Ghost convolution. The ERC2f module uses the EDWR structure to enhance the multi-scale feature fusion ability;
[0037] S2. Optimize the detection head structure: Remove the large target detection head in YOLOv8, and retain the small target and medium target detection heads to reduce the model parameter quantity and computational complexity;
[0038] S3. Feature extraction and fusion: Perform multi-scale fusion on the feature maps output by the backbone network through the path aggregation-feature pyramid structure (PA-FPN), and input them into the detection head to complete target detection.
[0039] At the same time, as Figure 1 shown, the method consists of two parts, the backbone network and the detection network, when running. Figure 1 Among them, Conv represents the convolutional layer, C2f represents the bottleneck layer, GCADown and ERC2f are the modules proposed by the present invention, SPPF represents the fast pyramid pooling operation, Upsample represents the upsampling operation, Concat represents the splicing operation, and Detect represents the detection operation.
[0040] The backbone network is responsible for extracting the features of the image, mainly composed of multiple convolutional layers, C2f modules, GCADown modules, ERC2f modules, and SPPF modules. The input image first passes through two convolutional layers. The convolutional layer consists of a convolutional kernel, batch normalization, and SiLU activation function. Among them, the convolutional kernel is a small matrix used to slide on the input image for convolution operations. Batch normalization is used to accelerate the training process and improve the stability of the model. The SiLU activation function is used to introduce non-linearity, enabling the network to learn complex features. The batch normalization process can be expressed as formula (1).
[0041]
[0042] Among them, is the standardized sample, and γ and β are two learnable parameters. The specific implementation of the SiLU activation function is expressed as formulas (2)-(3).
[0043]
[0044] After passing through the second convolutional layer, the feature map is input into C2f. The input features first pass through a convolutional layer, which performs convolution operations on the input feature map to generate an intermediate feature map. Then the intermediate feature map is split into two parts. One part of the feature map is directly passed to the final Concat module without any processing, retaining the original feature information. The other part of the feature map undergoes deeper feature extraction and enhancement through a series of convolutional layers in the Bottleneck module and then is passed to the Concat module. The feature map output from the third layer of C2f is input into the fourth layer of GCADown proposed in the present invention. The specific structure of GCADown is as Figure 2 shown,[[]]END]] Figure 2 where AvgPool and MaxPool represent average pooling and max pooling respectively, and Concat represents the concatenation operation.[[]]END]]
[0045] Assume the size of the feature map input into GCADown is C×H×W. In GCADown, it first undergoes the average pooling operation shown in formula (4).
[0046]
[0047] Among them, f i,j represents the value after average pooling, and n represents the number of regions for average pooling.[[]]END]] It represents the value of the k-th region of the input data. Then, the channels of the feature map after average pooling are divided into two branches in a 1:1 ratio, that is, the number of channels on each branch is C / 2×H×W. Then, the feature map on the first branch passes through GhostConv. In GhostConv, first, a 1×1 convolutional kernel is used to aggregate the information features of the channels, then a 3×3 convolutional kernel is used to perform convolution operations on the aggregated information features, and finally, the output structure and the channels of the original feature map are merged to obtain the final result. The size of the feature map of the final result is half of that before input, which is C / 2×H / 2×W / 2. The specific operation is shown in Equation (5).
[0048]
[0049] Among them, I and I o respectively represent the input feature map and the output feature map, Conv 1×1 and Conv 3×3 respectively represent the convolution operations with convolutional kernels of size 1×1 and 3×3, represents feature merging along the channel dimension. On the other branch, the input feature image undergoes a max-pooling operation to reduce the size of the image to half of the original, which is C / 2×H / 2×W / 2. The operation of max-pooling is shown in Equation (6).
[0050]
[0051] Among them, f i,j represents the value after max-pooling, represents the value on the k-th region that needs to be max-pooled. Then, the output feature map passes through GhosrConv to output the result of the second branch. Finally, the results of the two branches are concatenated to obtain an output feature map of size C / 2×H / 2×W / 2. Then, this feature map passes through C2f of the fifth layer and GCADown of the sixth layer and is input into ERC2f of the seventh layer. The specific structure of ERC2f is as Figure 3 shown, Figure 3 in which, Conv represents the convolutional layer, Split represents the division operation, EDWR is the module proposed in the present invention, and Concat is the concatenation operation.
[0052] This structure replaces the Bottleneck structure in C2f with the EDWR structure proposed in the present invention. The feature map input into ERC2f will pass through the EDWR structure. The EDWR structure is as Figure 4 shown, Figure 4 in which, Conv represents the convolutional layer, AvgPool represents the average pooling operation, and D-1, D-3, and D-5 respectively represent the dilation rates of the convolutional kernels being 1, 3, and 5.
[0053] In EDWR, the input feature map first undergoes a convolutional layer with a size of 3 to preprocess the input feature map. The preprocessed feature map is then passed through three different branches. On the first branch, it directly passes through a convolutional layer with a kernel size of 1 and a dilation rate of 1 to increase the number of channels and enhance the non-linear expression ability of the model. On the second branch, it first undergoes an average pooling operation to retain important detail information, and then passes through a convolutional layer with a kernel size of 3 and a dilation rate of 3 to extract the feature information of the image. On the third branch, it passes through a convolutional layer with a kernel size of 5 and a dilation rate of 5 to obtain more extensive context information, thereby obtaining a larger receptive field and compensating for the detection effect of large targets. The specific operations can be expressed as equations (7)-(9).
[0054] O1 = Conv 1×1,d=1 (I) (7)
[0055] O2 = Avg(Conv 3×3,d=3 (I)) (8)
[0056] O3 = Conv 5×5,d=5 (I) (9)
[0057] Among them, I represents the input feature map, O1, O2, and O3 respectively represent the outputs of the three branches, Conv 1×1 , Conv 3×3 and Conv 5×5 respectively represent convolutional layers with kernel sizes of 1, 3, and 5, d represents the dilation rate. Avg represents the average pooling operation, and the specific calculation method is shown in equation (10).
[0058]
[0059] Among them, i and j represent the row and column indices of the feature map, k represents the number of pooling regions, m and n represent the row and column indices within the pooling region, c represents the number of channels of the feature map, I and O respectively represent the input and output feature maps. After merging the outputs of the three branches, the number of channels is adjusted through a convolutional kernel with a size of 1, and then merged. Finally, it is merged into the original feature map to obtain the final output. This series of operations helps to improve the object detection performance of the network. The specific operations can be expressed as equation (11).
[0060] O final = I + Conv 1×1 (Concat(O1, O2, O3)) (11)
[0061] Among them, I and O final respectively represent the input and the final output, Conv 1×1It represents a convolutional layer with a kernel size of 1. O1, O2, and O3 respectively represent the outputs of three branches, and Concat represents the operation of concatenating feature maps. The specific operations of the feature map passing through ERC2f are shown in equations (12)-(14).
[0062] x conv = Conv(x) (12)
[0063] x concat = Concat(EDWR(Split(x conv ))) (13)
[0064] x output = Conv(x concat ) (14)
[0065] Among them, Conv represents the convolutional layer, Split represents the splitting operation, and Concat represents the merging operation. After passing through ERC2f of the seventh layer, the output feature map passes through GCADown, ERC2f, and SPPF in sequence, and then the extracted features are input into the detection head. In the detection head, the image is upsampled and fused through the feature pyramid structure, and then the obtained result is input into the small-object and medium-object detection heads for detection and recognition.
[0066] The present invention verifies the performance of the algorithm proposed in the present invention mainly from aspects such as precision (P), recall (R), mean average precision (mAP), number of parameters (Params), and floating-point operations per second (FLOPs) through two public datasets: SSDD and iVision-MRSSD.
[0067] Table 1: Comparative experiments of different algorithms on the SSDD dataset.
[0068]
[0069]
[0070] Table 2: Comparative experiments of different algorithms on the iVision-MRSSD dataset.
[0071]
[0072] As can be seen from Table 1 and Table 2, for the improved algorithm, compared with the existing algorithms, the accuracy is the highest in terms of mAP0.5:0.95, and the number of parameters and the computational cost are decreased by 47.8% and 19.8% respectively compared with the baseline algorithm yolov8n. Given this result, it can be shown that the algorithm proposed in the present invention is effective.
[0073] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A high-precision lightweight method for synthetic aperture radar ship target detection, characterized in that, It includes the following steps: S1. Construct a lightweight backbone network: Based on the YOLOv8 algorithm, introduce the GCADown module and the ERC2f module. The GCADown module reduces the computational load by combining average pooling, max pooling, and Ghost convolution. The ERC2f module adopts the EDWR structure to enhance the multi-scale feature fusion ability; S2. Optimize the detection head structure: Remove the large-object detection head in YOLOv8, and retain the small-object and medium-object detection heads to reduce the number of model parameters and computational complexity; S3. Feature extraction and fusion: Perform multi-scale fusion on the feature map output by the backbone network through the path aggregation-feature pyramid structure, and input it to the detection head to complete object detection.
2. The high-precision lightweight method for synthetic aperture radar ship target detection according to claim 1, characterized in that The specific operation of the GCADown module is as follows: The input feature map is respectively subjected to average pooling and max pooling operations to generate feature maps of two branches; Each branch is processed by Ghost convolution, and the output feature map size is halved; The final output feature map is merged through a concatenation operation.
3. The high-precision lightweight method for synthetic aperture radar ship target detection according to claim 1, characterized in that The EDWR structure includes the following branches: The first branch: Use a 1×1 convolutional kernel with a dilation rate of 1 to enhance the non-linear expression ability; The second branch: Use a 3×3 convolutional kernel with a dilation rate of 3, combined with average pooling to extract local detail features; The third branch: Use a 5×5 convolutional kernel with a dilation rate of 5 to expand the receptive field to capture context information; The output feature maps of each branch are merged after adjusting the number of channels through concatenation and 1×1 convolution.
4. The high-precision lightweight method for synthetic aperture radar ship target detection according to claim 1, characterized in that The ERC2f module realizes multi-branch feature fusion by replacing the Bottleneck in the C2f module with the EDWR structure.
5. The high-precision lightweight method for synthetic aperture radar ship target detection according to claim 1, characterized in that The detection head only retains the small-object and medium-object detection heads, and decouples the classification and regression tasks to reduce the number of parameters and computational load.
6. The high-precision lightweight method for synthetic aperture radar ship target detection according to claim 1, characterized in that The performance metrics of the method applied in synthetic aperture radar (SAR) images are as follows: On the SSDD dataset, mAP@0.5:0.95 reaches 75.2%, the number of parameters is 1.57M, and FLOPs is 6.5G; On the iVision-MRSSD dataset, mAP@0.5:0.95 reaches 61.8%, the number of parameters is 1.57M, and FLOPs is 6.5G.
Citation Information
Cited By
SAR (Synthetic Aperture Radar) image aircraft target detection method and system and medium
CN120853063A