A Method for Ship Target Detection in Synthetic Aperture Radar
By introducing ADown and ISPPF structures into the YOLOv8 backbone network, the problems of complex structure and loss of detailed information are solved, more efficient ship target detection is achieved, and detection accuracy and flexibility in model deployment are improved.
Patent Information
- Application Number
- CN202510197058.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing YOLOv8 ship object detection method uses convolutional layers in the backbone network to lead to high structural complexity, single feature learning, and SPPF structure leads to loss of image detail information, affecting detection accuracy.
The ADown structure and ISPPF structure are introduced into the backbone network of YOLOv8. Through the combination of average pooling and convolution operations, the amount of model parameters and calculations are reduced, while reducing the problem of loss of details.
The detection accuracy is improved, the number of parameters and calculations of the model is reduced, so that the model can be deployed on resource-constrained devices, and the accuracy of ship target detection is improved.
Smart Images

Figure CN119693806B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ship detection, and in particular to a method for detecting ship targets using a synthetic aperture radar. Background Art
[0002] Ship target detection is a key part of maritime monitoring. The most commonly used technical means for ship target detection is synthetic aperture radar (SAR). SAR generates high-resolution images by emitting electromagnetic waves and receiving returned echo signals. SAR has strong penetration and can obtain information in complex weather scenes such as cloud cover, rain and fog. It is not affected by complex weather, lighting and other factors. SAR can work all day and all weather. Therefore, SAR images can be used to realize ship target recognition.
[0003] Traditional SAR ship detection usually uses binarization, threshold segmentation and other methods to separate the target to be detected from the background image, and then detects the target through the geometric and texture features of the target. In simple scenes, the above methods can achieve good results. However, in bad weather or complex environments, due to the presence of more background interference in the image, texture features need to be added, but texture features will be affected by resolution. In addition, traditional recognition methods require manual feature design. When the image background noise is too large, manual features are difficult to establish, which makes ship recognition more difficult.
[0004] With the continuous development of deep learning, ship target detection algorithms have gradually shifted from traditional methods to target detection methods based on deep learning. Deep learning-based methods no longer require manual design of features, and their deep models can automatically learn effective feature representations from data. This shift has greatly promoted the development of ship target detection technology.
[0005] In the ship detection method based on deep learning, the YOLO series of algorithms are widely used. YOLOv8 is the eighth version of the YOLO series, which mainly consists of three parts: the backbone network, the neck network and the detection network. The backbone network uses the improved CSPDarkNet structure, which consists of convolutional layers, C2f structure and fast pyramid pooling (Spatial Pyramid PoolingFast, SPPF); the neck network part adopts the path aggregation network-feature pyramid network to achieve more effective feature fusion and enhance the model's detection ability for targets of different sizes; the detection network adopts the idea of no anchor box to separate the regression branch and the classification branch, making the network training and reasoning more efficient.
[0006] However, YOLOv8 uses convolutional layers for downsampling in the backbone network. This structure consists of a batch normalization layer, a convolutional operation, and an activation function, with a relatively high structural complexity. Moreover, the features learned by this structure are relatively single, which is not conducive to object detection. At the same time, the last structure of the backbone network uses the SPPF structure, which is composed of three consecutive max pooling operations and a skip connection. By performing a series of pooling operations on the input feature map, it can effectively capture receptive field information at different scales. Subsequently, by fusing features at different scales, it can effectively improve the detection ability of target feature information at different scales. However, when using the max pooling operation to process the feature map, during the process, due to the reduction in resolution, some detailed information of the image will be lost, reducing the detection accuracy. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide a method for synthetic aperture radar ship target detection. By introducing the ADown structure and the ISPPF structure into the backbone network, this method can reduce the number of parameters and computational complexity of the model while alleviating the problem of detailed information loss caused by the SPPF.
[0008] To achieve the above object, the present invention provides the following technical solution: A method for synthetic aperture radar ship target detection. The ship target detection method uses the collected synthetic aperture radar data as the detection data source and a target detection neural network as the detection means to detect ships. Among them, the target detection neural network is a neural network improved based on YOLOv8, including: a backbone network, a neck network, and a detection network. Among them, the backbone network includes: a convolutional layer, a C2f structure, an ADown structure, and an ISPPF structure.
[0009] Specifically, the algorithm process of the ADown structure is as follows: The input data is divided into two branches after average pooling. One branch performs a convolutional operation based on the pooled data, and the other branch performs a max pooling operation on the pooled data and then a convolutional operation. Finally, the data processed by the two branches is fused by Concat and then output.
[0010] Specifically, the algorithm process of the ISPPF structure is as follows:
[0011] (1) The input data X is output as X after average pooling, and the calculation formula is: avg , where, , represents the th number in the area to be average pooled in the th row and th column, and
[0012] (2) The output data X after average pooling avg is divided into two branches, X1 and X2, in a 1:1 ratio, where Split represents the splitting function;
[0013] (3) X1 is output as X after two convolutions 1b , and the calculation formula is:
[0014] ,
[0015] ,
[0016] where X 1a represents the output after the first convolution, k represents the convolution kernel size, s represents the stride, and d represents the dilation rate;
[0017] (4) X2 is output as the calculation result X3 of the second branch after pooling and skip connection feature extraction, and the formula is:
[0018] ,
[0019] ,
[0020] ,
[0021] where X 2a represents the output after X2 pooling, and X 2b represents the output after X 2a pooling, and Skipconnection represents the skip connection;
[0022] (5) The output result X 1b of the first branch and the output result X3 of the second branch are fused by Concat and then output as Y after the ISPPF operation, .
[0023] Preferably, when the object detection neural network is used as a detection means to detect ships, the training method of the model includes the following steps:
[0024] S1. Data collection, collecting images of ships through synthetic aperture radar;
[0025] S2. Data annotation, using the LabelImg tool to annotate the collected images, and dividing the annotated images into a training set and a validation set in a ratio of 8:2;
[0026] S3. Inputting the divided image data set into the improved network for training;
[0027] S4. Using the trained model to detect images with ship targets, and the specific calculation method is as follows:
[0028] ,
[0029] ,
[0030] ,
[0031] ,
[0032] Among them, TP represents the number that the model can correctly identify, FP represents the number that is misidentified by the model, FN represents false detection, and Q represents the number of categories.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] 1. By introducing the ADown structure into the backbone network of YOLOv8, the present invention can reduce the number of parameters and computational complexity of the model, which is beneficial for the model to be deployed on resource-constrained devices.
[0035] 2. The present invention proposes a new network structure ISPPF, which can alleviate the problem of loss of detailed information caused by SPPF in YOLOv8. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic diagram of the improved network framework of the present invention;
[0037] Figure 2 It is a schematic diagram of the ADown structure of the present invention;
[0038] Figure 3 It is a schematic diagram of the ISPPF structure of the present invention;
[0039] Figure 4 It is a flowchart of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0041] Based on YOLOv8, the present invention proposes a method for synthetic aperture radar ship target detection, which consists of a backbone network, a neck network, and a detection network. Among them, the backbone network consists of a convolutional layer, C2f, ADown, and ISPPF; the neck network consists of an upsampling layer, Concat, C2f, and a convolutional layer; the detection network consists of a convolutional layer, and the overall structure is asFigure 1 As shown in the figure, where Input represents the input, Conv represents the convolutional layer, C2f represents the bottleneck layer, ADown is the introduced downsampling structure, ISPPF is the enhanced pyramid pooling structure proposed in the present invention, Upsamples represents upsampling, Concat represents merging, and Detect represents detection.
[0042] In the backbone network, the input image first passes through two Conv layers for extracting the preliminary features of the image. The Conv layer consists of a convolutional operation, batch normalization, and the SiLU activation function. The feature map after passing through the Conv layer can be expressed as:
[0043] ,
[0044] where x represents the sample data input, y represents the output of the feature processing of the corresponding layer, conv represents the convolutional operation, bn represents batch normalization, and act represents the activation function. Batch normalization is expressed as:
[0045] ,
[0046] ,
[0047] ,
[0048] ,
[0049] where is the i-th sample in the batch, is the normalized sample, is the size of the batch, is the mean, is the variance, is a constant used to avoid division by zero, and are the learned parameters. The SiLU activation function is expressed as:
[0050] ,
[0051] where is the Sigmoid function. After passing through the Conv layer, it then passes through a C2f for further extracting and enhancing features. The C2f module enhances the performance of the model by optimizing the gradient flow. Then it passes through an ADown followed by a C2f module for extracting and enhancing features and at the same time aligning with the P3 layer of the feature pyramid. The structure of ADown is as Figure 2 shown in the figure, where Figure 2Avgpool in it represents average pooling, Conv represents the convolutional layer, Maxpool represents max pooling, and Concat represents concatenation; it simplifies the model structure by reducing the model's parameters, thereby improving the model's computational efficiency. Although ADown is used to reduce the size of the feature map, the design also fully considers retaining as much image information as possible to ensure that the model can accurately identify the target. After passing through ADown and then C2f, it is used to extract and enhance features, corresponding to the P4 layer of the feature pyramid.
[0052] Finally, after passing through ADown and C2f and then through ISPPF, the feature extraction of the input image is completed.
[0053] The structure of ISPPF is as Figure 3 shown, Avgpool represents average pooling, Split represents splitting, Conv represents the convolutional layer, Maxpool represents max pooling, and Concat represents concatenation; the input feature map first passes through average pooling, and the formula is as follows:
[0054] ,
[0055] where, represents the output after average pooling, represents the row column number in the area to be average pooled, number, represents the number of elements in the area to be average pooled.
[0056] Then the input feature map is divided into and two parts of feature maps in a 1:1 ratio:
[0057] ,
[0058] Feature map On the first branch, it passes through a convolutional layer with a kernel size of 3, a stride of 1, and a dilation rate of 2 and a convolutional layer with a kernel size of 3, a stride of 1, and a dilation rate of 3 respectively to extract the features of the image, represents the kernel size, represents the stride size, represents the dilation rate, represents the output after the first convolutional layer, represents the output after the second convolutional layer:
[0059] ,
[0060] ,
[0061] Feature map After passing through the second branch, more abstract features are further extracted through two max pooling operations and skip connections. Finally, the two branches are merged to obtain a more detailed feature map, alleviating the problem of partial detail loss in SPPF. Represents the output after the first max pooling operation. Represents the output after the second max pooling operation. Represents the final output of the second branch. Represents the output after passing through ISPPF.
[0062]
[0063] ,
[0064] ,
[0065] ,
[0066] The neck network consists of a Path Aggregation Network - Feature Pyramid Network, which is used to fuse features of different scales and sizes. The Path Aggregation Network upsamples the feature maps from different levels and then fuses them with features from other levels. This fusion can retain both low-level details and high-level semantic information. The Feature Pyramid Network constructs and fuses multiple feature representations at different feature levels, enhancing the model's understanding of images at multiple scales.
[0067] In the detection network, using the idea of decoupling, two parallel branches are used to independently extract class and location features, separating classification and detection to detect and identify ship targets.
[0068] The complete training and detection process requires several steps as shown in Figure 4 :
[0069] (1) Data collection, collecting images of ships through synthetic aperture radar.
[0070] (2) Data annotation, using the LabelImg tool to annotate the collected images, and dividing the annotated images into a training set and a validation set according to a ratio of 8:2.
[0071] (3) Input the divided image dataset into the improved network for training. During training, set the size of the input image to 640, the batch_size to 64, the initial learning rate to 0.01, and train for 200 epochs.
[0072] (4) Use the trained model to detect images with ship targets.
[0073] Finally, the present invention randomly selects 3,000 images from the publicly available dataset sar_ship_dataset to construct a small dataset, and verifies the performance in terms of precision (P), recall (R), mean average precision (mAP), number of parameters (Params), floating point operations per second (FLOPs), etc. The specific calculation method is as shown in the formula.
[0074] ,
[0075] ,
[0076] ,
[0077] ,
[0078] Among them, TP represents the number that the model can correctly identify, FP represents the number that is misidentified by the model, FN represents false detection, and Q represents the number of categories.
[0079] The performance of the ISPPF proposed by the present invention is compared with some existing improved pyramid pooling schemes, and the comparison results are shown in the following table:
[0080]
[0081] It can be seen from the table that the precision of the present invention is improved by 0.2% compared with the original YOLOv8, and mAP@0.5 is improved by 0.4% compared with the original YOLOv8, which is the highest among all comparison schemes. In addition, the present invention also conducts ablation experiments on the improved network to verify the effectiveness of the network.
[0082]
[0083] As shown in the table, for the improved network, mAP@0.5 is improved by 0.8% compared with the original YOLOv8, and the number of parameters, computational volume, and model size are all reduced by 5%.
[0084] In summary, the SAR ship image detection method based on the improvement of YOLOv8 proposed by the present invention can reduce the number of parameters and computational volume of the model compared with the traditional method, enable the model to be more conveniently deployed on resource-constrained devices, and at the same time, can alleviate the problem of loss of detailed information caused by SPPF and improve the detection accuracy.
[0085] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for synthetic aperture radar ship target detection, characterized in that, The ship target detection method uses the collected synthetic aperture radar data as the detection data source and the target detection neural network as the detection means to detect ships. Among them, the target detection neural network is a neural network improved based on YOLOv8, including: a backbone network, a neck network, and a detection network. Among them, the backbone network at least includes: a convolutional layer, a C2f structure, an ADown structure, and an ISPPF structure; The algorithm process of the ADown structure is: the input data is divided into two branches after average pooling. One branch performs a convolutional operation based on the pooled data, and the other branch performs a max pooling operation on the pooled data and then performs a convolutional operation. Finally, the data processed by the two branches is fused by Concat and then output; The algorithm process of the ISPPF structure is: (1)The input data X is output as X after average pooling avg , and the calculation formula is: , where represents the th number in the area to be average pooled in the th row and th column, represents the number of elements in the area to be average pooled; (2)The output data X after average pooling avg is divided into two branches, X1 and X2, in a 1:1 ratio, where Split represents the splitting function; (3) After two convolutions of X1, the output is X 1a and X 1b , and the calculation formula is: , , Among them, X 1a represents the output after the first convolution, X 1b represents the output after the second convolution, k represents the convolution kernel size, s represents the stride, and d represents the dilation rate; (4) X2 outputs the calculation result X3 of the second branch after pooling and skip connection feature extraction. The formula is: , , , Among them, X 2a represents the output after X2 pooling, and X 2b represents the output after X 2a pooling, and Skipconnection represents the skip connection; (5)The output result X of the first branch 1b After being fused with the output result X3 of the second branch through Concat, the output Y after the ISPPF operation is output. .
2. The method for synthetic aperture radar ship target detection according to claim 1, wherein The neck network is at least composed of an upsampling layer, Concat, C2f, and a convolutional layer.
3. The method for detecting ship targets using synthetic aperture radar according to claim 2, wherein The detection network is composed of convolutional layers.
4. The method for synthetic aperture radar ship target detection according to claim 3, characterized in that In the backbone network, the input image first passes through two Conv layers to extract the preliminary features of the image. The Conv layer consists of a convolutional operation, batch normalization, and a SiLU activation function. The feature map is represented as: , where x represents the input of sample data, y represents the output of the feature processing of this layer, conv represents the convolution operation, bn represents batch normalization, act represents the activation function, and batch normalization is expressed as: , , , , wherein, is the i-th sample in the batch, is the standardized sample, is the size of the batch, is the mean value, is the variance, is a constant used to avoid division by zero, and are the parameters for learning, and the SiLU activation function is expressed as: , Among them, is the Sigmoid function. After passing through the Conv layer, it goes through a C2f to extract and enhance features. The C2f module enhances the performance of the model by optimizing the gradient flow. Then, it goes through an ADown followed by a C2f module to extract and enhance features and align with the P3 layer of the feature pyramid.
5. The method for synthetic aperture radar ship target detection according to claim 4, wherein, The training method of the model when using the target detection neural network as the detection means to detect ships includes the following steps: S1. Data collection, collecting images of ships through synthetic aperture radar; S2. Data annotation, using the LabelImg tool to annotate the collected images, and dividing the annotated images into a training set and a validation set according to a ratio of 8:2; S3. Input the divided image data set into the improved network for training; S4. Use the trained model to detect images with ship targets. The specific calculation method is as follows: , , , , Among them, TP represents the number that the model can correctly identify, FP represents the number that is misidentified by the model, FN represents false detection, and Q represents the number of categories.