A method for effectively identifying ship targets in SAR images

By adopting the YOLOv8-BCN detection model in SAR image recognition, this model combines the Transformer-style backbone network and neck network to solve the problems of missed detection and missed detection in complex backgrounds, achieving higher detection accuracy and fewer error recognition.

CN119625555BActive Publication Date: 2025-05-27BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411681062.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-05-27
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

When identifying ship targets in SAR images, prior art is prone to problems of mis-checking and missed inspections in complex backgrounds, especially the similarity between nearshore ship targets and backgrounds is too high.

Method used

The YOLOv8-BCN detection model is adopted, which includes the backbone network backbone, the neck network neck, and the downsampling module CLG. The backbone network uses C3BoT, the neck network uses C3NeXt, and the downsampling module of the YOLOv8n model is replaced in the downsampling module, combining Transformer-style feature extraction.

Benefits of technology

The YOLOv8-BCN detection model can effectively pay attention to the target in a complex context, improves the detection accuracy of small target densely agglomerated areas, and reduces the occurrence of missed detection and missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625555B_ABST
    Figure CN119625555B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image recognition and relates to a method for effectively identifying ship targets in SAR images: collecting SAR images and images taken by high-resolution satellites and Sentinel satellites, and constructing a training set and a test set according to a ratio of 8:2; using the YOLOv8n model as the baseline model to construct the YOLOv8-BCN detection model, and the YOLOv8-BCN detection model includes a backbone network constructed by C3BoT, a neck network constructed by C3NeXt, and a downsampling module CLG; training the YOLOv8-BCN detection model according to the training images using the stochastic gradient descent algorithm, and saving the YOLOv8-BCN detection model with excellent training effects; inputting the images in the test set into the trained YOLOv8-BCN detection model for testing, and the YOLOv8-BCN detection model outputs the predicted detection results; the YOLOv8-BCN detection model of the present invention can achieve 12.2 ms / per image in image prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of computer application, and in particular relates to a method for effectively identifying ship targets in SAR images. Background Art

[0002] SAR image ship detection is one of the hot topics in military application research. Currently, the commonly used deep learning models include lightweight models with simple structures, including LiraNet, Tiny-YOLOv3, lightweight v5-MNE model based on YOLOv5, LSSD based on SSD structure, YOLOv8 and models based on vision transformer. Although the existing models use various methods to achieve lightweight models, the detection accuracy has also been improved to a certain extent, and the multi-scale features are used to pay attention to the missed targets. Combined with Tansformer to solve the uncertainty of the target scale, the prediction potential of the transformer has been improved to a certain extent. However, it is easy to have false detection and missed detection problems under complex backgrounds, especially when the similarity between the nearshore ship target and the background is too high. For example, small targets with small pixel occupancy are densely gathered and appear as bright spots, which are prone to missed detection. Edge targets and false alarms in the image cause a ship to be easily misidentified as multiple ships. Summary of the invention

[0003] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method for effectively identifying ship targets in SAR images.

[0004] In order to achieve the purpose of the present invention, the present invention will be implemented by adopting the following technical solutions.

[0005] A method for effectively identifying ship targets in SAR images includes the following steps:

[0006] S1. For SAR images collected from the SSDD dataset and images taken by high-resolution satellites and Sentinel satellites collected from the SAR-ship-dataset dataset, the training set and test set were constructed according to the training-test ratio of 8:2.

[0007] S2. Use the YOLOv8n model as the baseline model to build a YOLOv8-BCN detection model; wherein: the YOLOv8-BCN detection model includes a backbone network backbone, a neck network neck, and a downsampling module CLG; wherein:

[0008] The last stage of the backbone network uses C3BoT, and C3BoT has a transform style. The left branch of C3BoT passes the input image through a ConvBNSilu backbone After the convolution operation, it is input to the tail ConvLN of the right branch backbone ; The right branch is to pass the input image from top to bottom through the head ConvLN backbone 、BoT backbone and Concat backbone After operation, input to the tail ConvLN backbone ; wherein the BoT backbone The structure is a standard Bottleneck structure, and the 3*3 convolution in the Bottleneck structure is replaced by multi-head self-attention MHSA;

[0009] The downsampling module CLG is used to replace all Conv layers used for downsampling in the YOLOv8n model except the first two layers of Conv layers used for downsampling in the backbone network. The downsampling module CLG is composed of a convolutional layer with kernel_size=2 and stride=2, a layer norm normalization layer, and a GELU activation function set from top to bottom in the encoder structure of Transformer;

[0010] The neck network uses C3NeXt, and C3NeXt has a transformer style. The left branch of C3NeXt is a feature map input from the backbone network backbone through a ConvBNSilu ne ck After the convolution operation, it is input to the tail ConvLN of the right branch neck ; The right branch is to pass the feature map input from the backbone network backbone one from top to bottom through the head ConvLN neck 、ConvNext Block neck and Concat n eck After operation, input to the tail ConvLN neck ; Wherein, the ConvNext Block nec It is composed of Inverted bot tleneck;

[0011] S3, input the images in the training set into the YOLOv8-BCN detection model, train the YOLOv8-BCN detection model from scratch using the stochastic gradient descent algorithm, and save the YOLOv8-BCN detection model with the best training effect;

[0012] S4. Input the images in the test set into the trained YOLOv8-BCN detection model for testing. The YOLOv8-BCN detection model outputs the predicted detection results.

[0013] As a preferred solution of the present invention, the batch size adopted by the YOLOv8-BCN detection model is 16, and the training rounds are set to 300 rounds.

[0014] As a preferred solution of the present invention, the YOLOv8-BCN detection model can achieve 12.2ms / perimage in image prediction.

[0015] As a preferred solution of the present invention, the advancement of the YOLOv8-BCN detection model is verified by burning experiments on the SSDD dataset and the SAR-ship-dataset dataset.

[0016] As a preferred solution of the present invention, the layer norm standardization layer standardizes all channels in a batch.

[0017] As a preferred solution of the present invention, the down sampling module CLG can be connected with the BoT backbone and ConvNextBlock neck Use in combination.

[0018] As a preferred solution of the present invention, the SAR images and the images taken by the high-resolution satellite and the sentinel satellite all have ship targets.

[0019] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0020] The YOLOv8-BCN detection model has advantages in terms of parameter quantity and computational complexity. After experiments, the YOLOv8-BCN detection model can achieve 12.2ms / perimage in image prediction and can also focus on targets in some complex backgrounds. The Transformer style is effective in SAR image ship detection. The YOLOv8-BCN detection model has achieved relatively good results in the research pain points of complex backgrounds, image edges, and dense clusters of small targets, proving that self-attention is a superior mechanism for SAR image recognition. Changes in the backbone network, neck, and downsampling modes can improve the performance of the YOLOv8-BCN detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic diagram of the binding of BoT and C3 according to the present invention;

[0022] Figure 2 is a comparative schematic diagram of the down sampling module CLG and the down sampling module CBS of the present invention;

[0023] Figure 3It is a schematic diagram of the combination of ConvNext Block and C3 of the present invention;

[0024] Figure 4 It is a target image that is easily missed in the SSDD data set of the present invention;

[0025] Figure 5 It is a target image that is easily missed in the SAR-Ship-Dataset data set of the present invention;

[0026] Figure 6 This is a graph showing the detection results of the improved model in the SSDD dataset of the present invention;

[0027] Figure 7 This is a graph of the detection results of the improved model in the SAR-Ship-Dataset data set described in the present invention. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] As an embodiment of the present invention, Figures 1 to 3 As shown, a method for effectively identifying ship targets in SAR images includes the following steps:

[0030] S1. For SAR images collected from the SSDD dataset and images taken by high-resolution satellites and Sentinel satellites collected from the SAR-ship-dataset dataset, the training set and test set were constructed according to the training-test ratio of 8:2.

[0031] S2. Use the YOLOv8n model as the baseline model to build a YOLOv8-BCN detection model; wherein: the YOLOv8-BCN detection model includes a backbone network backbone, a neck network neck, and a downsampling module CLG; wherein:

[0032] The last stage of the backbone network uses C3BoT, and C3BoT has a transform style. The left branch of C3BoT passes the input image through a ConvBNSilu backbo ne After the convolution operation, it is input to the tail ConvLN of the right branch backbone ; The right branch is to pass the input image from top to bottom through the head ConvLN backbone、BoT backbone and Concat backbone After operation, input to the tail ConvLN backbone ; wherein the BoT backbone The structure of is a standard Bottleneck structure, and the 3*3 convolution in the Bottleneck structure is replaced by multi-head self-attention MHSA; Figure 1 As shown; the improvement of the backbone network backbone structure aims to solve the problem of weak target perception ability in complex backgrounds;

[0033] The downsampling module CLG is used to replace all Conv layers used for downsampling in the YOLOv8n model except the first two layers of Conv layers used for downsampling in the backbone network. The downsampling module CLG is composed of a convolution layer with kernel_size=2 and stride=2, a layer norm normalization layer, and a GELU activation function set from top to bottom in the encoder structure of Transformer; Figure 2 As shown; Improving the downsampling module of the YOLOv8n model into a downsampling module CLG can reduce the number of parameters to a certain extent, and can be well combined with the C3BoT and C3NeXt modules;

[0034] The neck network uses C3NeXt, and C3NeXt has a transformer style. The left branch of C3NeXt is a feature map input from the backbone network backbone through a ConvBNSilu ne ck After the convolution operation, it is input to the tail ConvLN of the right branch neck ; The right branch is to pass the feature map input from the backbone network backbone one from top to bottom through the head ConvLN neck 、ConvNext Block neck and Concat n eck After operation, input to the tail ConvLN neck ; Wherein, the ConvNext Block nec It is composed of Inverted bot tleneck; its purpose is to improve the recall of complex background targets, ensure that difficult-to-detect and missed targets can be focused on and accurately return to the target;

[0035] S3, input the images in the training set into the YOLOv8-BCN detection model, train the YOLOv8-BCN detection model from scratch using the stochastic gradient descent algorithm, and save the YOLOv8-BCN detection model with the best training effect;

[0036] S4. Input the images in the test set into the trained YOLOv8-BCN detection model for testing. The YOLOv8-BCN detection model outputs the predicted detection results.

[0037] As a preferred solution of the present invention, the batch size adopted by the YOLOv8-BCN detection model is 16, and the training rounds are set to 300 rounds.

[0038] As a further solution of the present invention, the prediction of the YOLOv8-BCN detection model can reach 12.2ms / perimage.

[0039] As a further solution of the present invention, the YOLOv8-BCN detection model is verified to be advanced through burning experiments on the SSDD dataset and the SAR-ship-dataset dataset.

[0040] As an embodiment of the present invention, the layer norm standardization layer standardizes all channels in one batch; Batchnorm standardizes all batches on one channel.

[0041] As an embodiment of the present invention, the down-sampling module CLG can be used in combination with the BoT and ConvNextBlock.

[0042] As an embodiment of the present invention, Figure 1 As shown, the present invention has made many attempts to design the module, placing it in different positions and making different changes to the details of the module. In the process of using C3BoT, the study found that the most suitable position for the module is to be placed in the fourth stage of the backbone network. During the experiment, it was tried to place it in each stage of the backbone network, and it was found that it would cause excessive computational complexity and affect training. The main impact is that placing it in the third stage will directly lead to difficulties in training and bring excessive parameters. Therefore, we gave up using C3BoT in every stage and followed the original structure of the BoT model, using it only in the last stage of the network.

[0043] Why use the previous generation YOLOv5's C3 module as the basis instead of the latest C2f module of YOLOv8? The C2f module adopts a diversion strategy after each standard bottleneck in the process of processing feature maps. Although it allows subsequent feature maps to have more gradient flow information, this structure is not suitable for adding BoT. The core of the BoT structure is the multi-head attention mechanism in the middle of the bottleneck. Compared with ordinary convolution, the multi-head attention mechanism will bring a certain amount of calculation, so it is not suitable for multiple cycles in a C2f. Therefore, the C3 module is used, and the standard bottleneck in the middle is directly replaced by BOT, such as Figure 1 shown.

[0044] Before using BOT, the present invention tried to use the Swin-Transformer stage in combination with the C3 module, but the computational complexity and model parameter quantity of Swin-Transformer were too large. At the same time, it was tried to directly use Swin-Transformer to replace the backbone network for feature extraction, and then processed by the neck network of YOLOv8 and finally detected by the head network, but the performance improvement was not significant. An attempt was made to imitate the BoT model and place the SwinTransformer stage at the last stage of the yolov8 backbone network. Although the parameters were significantly reduced, the performance improvement was still not as good as BOT, so the use of SwinTransformer was abandoned.

[0045] As an embodiment of the present invention, the present invention refers to the experimental details in the article about ConvNeXt and experiments on the number of heads of multi-head attention. The original number of heads was set to 16. Referring to the head setting of SwinTransformer, the number of heads was 3, 6, 12, and 24 in the four stages. Because multi-head attention is only used in the fourth stage of the model, the number of heads is increased to 24. From the experimental results of SAR-ship-dataset, the performance has even regressed, and the accuracy has dropped by 0.04 compared to the baseline model. The performance of other numbers of heads is not as good as 16 heads, so 16 heads is the most suitable in this model. The C3BoT model tried slight changes in the experiment. Because the purpose of the present invention is to build a transformer-style model, the present invention replaces the batchnorm in the three standard convolutions of the C3 module with layernorm, and the activation function is replaced from Silu to GELU, as shown in the Vit and other models. Figure 2As shown in the figure. From the perspective of mAP50, such a small change can have an impact on both SSDD and SARship datasets. On the SSDD dataset, only C3BoT is added to the backbone network, and the first convolution of the BoT module is changed as described above. The mAP50 will drop from 0.915 to 0.91, and the recall rate will drop while the precision will hardly change. However, if the activation function is cancelled, the mAP50 will increase to 0.922.

[0046] As an embodiment of the present invention, the present invention believes that the unification of styles is very important. According to the structural setting of the ConvNeXt model, the dependence on the activation function is reduced. In C3BoT, the model has three Convs, the batch norm of Conv is replaced by layernorm, the activation function is replaced by GELU, and the activation function is only used in the Conv of the tributary, and the activation function is not enabled in the other two Convs. From the results, only adding C3BoT, mAP50 increased by 0.04, the recall rate decreased by 0.23, and the precision increased by 0.18. In this way, C3BoT of this model is determined to be the most suitable version.

[0047] As an embodiment of the present invention, the present invention initially experimented with a new downsampling module using SwinTransformer's Patch merging. This module divides the feature map of the previous layer into four smaller feature maps, and the number of channels is increased by 4 times. The operation is similar to the convolution operation with a stride of 2 before each stage of CNN. Layernorm is adopted and the activation function is cancelled. Patch merging does not lose image information during the downsampling process, because in essence, this operation only slices adjacent pixels of the image. After the experiment, the use of patch merging alone does not seem to show a huge advantage in performance, and then the new down sample ~ CLG is used. In the initial experiment, the frequent use of activation functions was reduced, and layernorm was set before convolution, but the results of the experiment showed that the performance regressed and was not compatible with the C3BoT module. Therefore, in subsequent experiments, the order of layer norm and convolution was changed, with convolution first and layernorm later. After the experiment, there were still problems in the combination of the module and C3BoT, which led to the instability of the training process. Later, it was decided to re-add the Gelu activation function. After the experiment, whether used alone in SSDD or SARship dataset, it can bring performance improvement. And this final version of CLG is more suitable for the other two modules. This model tried to replace all YOLOv8 downsampling with CLG, only replace the downsampling of the backbone network, and only use it in the last stage with C3BoT. The results of the experiment show that mAP50 on the SARship dataset shows that only replacing the downsampling of the backbone network will exceed the baseline model by 0.004. Replacing all downsampling will cause mAP50 to drop by 0.009, but if the activation function of the downsampling of the neck network is cancelled, the performance will be restored. Finally, this model tried to use the new downsampling in stage4 of the backbone network, but the performance was not significantly outstanding. Finally, it was decided to use full replacement, but cancel the activation function in the neck. It is worth noting that this module cannot replace the first two layers of the YOLOv8 backbone network for downsampling Conv (directly downsampled from size 640 to 160). From the experimental results, it will cause too much fluctuation in training and difficult to return.

[0048] As an embodiment of the present invention, Figure 3As shown in the figure, this model has also made many attempts to use ConvNeXt. First, it is used as the backbone network. The experimental results show that adding it to the backbone network can well balance the precision and recall rate, both of which are around 0.89, but the mAP50 is slightly reduced. It is also compared to swap the positions of C3BoT and C3NeXt, and set C3BoT in the neck network. This combination can hardly improve the mAP50 of the model. In addition, all c2f modules of the backbone network except the fourth stage are replaced with C3NeXt. From the experimental results, several indicators are not as good as the improvement brought by placing them on the neck network. Finally, in the stylized test, all Convs of C3NeXt except the branches use Layernorm and cancel the activation function, and retain the original Batch norm on the branches. Using C3NeXt module in the neck network can improve the recall rate. This model makes a simple change to ConvNeXtblock. The original module adds layer scale and droppath after invertedbottleneck. This model cancels layerscale and droppath. Layerscale is a learnable parameter, and droppath randomly deletes the multi-branch structure of the module. Finally, the model only leaves Invertedbottleneck.

[0049] The datasets introduced are SSDD and SAR-ship-dataset. SSDD contains 1,160 SAR images, and SAR-ship-dataset contains more than 40,000 images from Gaofen and Sentinel. Both include ship targets of various scales and various backgrounds. Both datasets are used for experiments with a training-test ratio of 8:2.

[0050] The present invention follows the default strategy of training the data set from scratch in YOLOv8—the stochastic gradient descent algorithm SGD for training. The batch size used by the network is 16, and the number of training rounds is set to 300 rounds. Other unmentioned hyperparameters remain the same as in YOLOv8. In addition, when the present invention is compared with other methods, the parameters set are basically the same to ensure that the difference is caused by the change of the model rather than the change in performance caused by the use of certain tricks.

[0051] In order to demonstrate the advancedness of the model, the present invention conducted burning experiments on two datasets, SSDD and SAR ship dataset. The YOLOv8n baseline model was first trained according to the parameters set in the above experimental details. Then the experimental strategy of the present invention was to add modules C3BoT, C3NeXt, and downsampling module CLG to the structure of YOLOv8, and then test the model performance through different combinations.

[0052] For each result, we tried changing some details within the module to achieve the best training results.

[0053] Table 1 shows the burning experiment of the module in the SAR-Ship-Dataset dataset

[0054]

[0055] Table 2: Burning experiment results of different modules in SSDD dataset

[0056]

[0057] The results are compared with models of different styles. The data in Table 3 are all obtained from the SSDD and SAR shipdataset experiments.

[0058] Table 3:

[0059]

[0060]

[0061] From the results, the impact of adding the Transformer-style module is positive. Objects on the edges of some images and in some complex backgrounds are easily missed. Figure 4 and Figure 6 shown.

[0062] The addition of multi-head attention and the use of a more modern bottleneck structure allow the model to focus on some difficult-to-detect ship targets, such as Figure 5 and Figure 7 shown.

[0063] The Transformer structure brings a more modern feature extraction structure, which to some extent makes up for the information lost in the convolutional network.

[0064] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for effectively identifying ship targets in SAR images, characterized by: The steps include: S1. For SAR images collected from the SSDD dataset and images taken by high-resolution satellites and Sentinel satellites collected from the SAR-ship-dataset dataset, the training set and test set were constructed according to the training-test ratio of 8:

2. S2. Use the YOLOv8n model as the baseline model to build a YOLOv8-BCN detection model; wherein: the YOLOv8-BCN detection model includes a backbone network backbone, a neck network neck, and a downsampling module CLG; wherein: The last stage of the backbone network uses C3BoT, and C3BoT has a transformer style. The left branch of C3BoT passes the input image through a ConvBNSilu backbone After the convolution operation, it is input to the tail ConvLN of the right branch backbone ; The right branch is to pass the input image from top to bottom through the head ConvLN backbone 、BoT backbone and Concat backbone After operation, input to the tail ConvLN backbone ; wherein the BoT backbone The structure is a standard Bottleneck structure, and the 3*3 convolution in the Bottleneck structure is replaced by multi-head self-attention MHSA; The downsampling module CLG is used to replace all Conv layers used for downsampling in the YOLOv8n model except the first two layers of Conv layers used for downsampling in the backbone network. The downsampling module CLG is composed of a convolution layer with kernel_size=2 and stride=2, a layer norm normalization layer, and a GELU activation function set from top to bottom in the encoder structure of the Transformer; The neck network uses C3NeXt, and C3NeXt has a transformer style. The left branch of C3NeXt is a feature map input from the backbone network backbone through a ConvBNSilu neck After the convolution operation, it is input to the tail ConvLN of the right branch neck ; The right branch is to pass the feature map input from the backbone network backbone one from top to bottom through the head ConvLN neck 、ConvNext Block neck and Concat neck After operation, input to the tail ConvLN neck ; Wherein, the ConvNext Block nec It is composed of Inverted bot tleneck; S3, input the images in the training set into the YOLOv8-BCN detection model, train the YOLOv8-BCN detection model from scratch using the stochastic gradient descent algorithm, and save the YOLOv8-BCN detection model with the best training effect; S4. Input the images in the test set into the trained YOLOv8-BCN detection model for testing. The YOLOv8-BCN detection model outputs the predicted detection results.

2. The method for effectively identifying ship targets in SAR images according to claim 1, characterized in that: The batch size used by the YOLOv8-BCN detection model is 16, and the number of training rounds is set to 300.

3. The method for effectively identifying ship targets in SAR images according to claim 1, characterized in that: The YOLOv8-BCN detection model can achieve 12.2ms / perimage in image prediction.

4. The method for effectively identifying ship targets in SAR images according to claim 1, characterized in that: The advancement of the YOLOv8-BCN detection model is verified by burning experiments on the SSDD dataset and SAR-ship-dataset dataset.

5. The method for effectively identifying ship targets in SAR images according to claim 1, characterized in that: The layer norm standardization layer standardizes all channels in a batch.

6. The method for effectively identifying ship targets in SAR images according to claim 1, characterized in that: The down sampling module CLG can be connected with the BoT backbone and ConvNextBlock neck Use in combination.

7. The method for effectively identifying ship targets in SAR images according to claim 1, characterized in that: The SAR images and the images taken by the high-resolution satellite and the sentinel satellite all have ship targets.

Citation Information

Patent Citations

  • Marine ship detection method based on synthetic aperture radar data

    CN116665148A

  • SAR (Synthetic Aperture Radar) ship detection method and system based on space-ground synchronous neural network

    CN118053080A