SAR image ship detection method and detection terminal

By adopting the SAR ship detection model in the SAR image, including the backbone feature extraction network and feature aggregation and diffusion pyramid network, the problem of low target detection accuracy in the SAR image is solved, and higher detection accuracy and accuracy are achieved.

CN119992041APending Publication Date: 2025-05-13HEBEI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411704418.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the object detector designed based on visible light images has a problem of low accuracy in SAR images, mainly due to the complex background interference in SAR images and the diversity of ship reflection characteristics.

Method used

A SAR ship detection model is adopted, which includes a backbone feature extraction network, feature aggregation and diffusion pyramid network and detection head. The backbone feature extraction network is used to extract the initial feature map of multi-scale. The feature aggregation and diffusion pyramid network spreads features with rich context information to each detection scale through a diffusion mechanism. The detection head is used to detect the multi-scale target feature map.

Benefits of technology

It effectively improves the accuracy and accuracy of object detection of SAR images, reduces the impact of background interference, and improves the accuracy of ship detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992041A_ABST
    Figure CN119992041A_ABST
Patent Text Reader

Abstract

The invention provides an SAR (Synthetic Aperture Radar) image ship detection method and a detection terminal. The method comprises the following steps: acquiring a to-be-detected SAR image; inputting a to-be-detected SAR image into the SAR ship detection model to obtain a target detection result; wherein the SAR ship detection model comprises a trunk feature extraction network, a feature aggregation and diffusion pyramid network and a detection head; the trunk feature extraction network is used for performing feature extraction on a to-be-detected SAR image to obtain a multi-scale initial feature map; the feature aggregation and diffusion pyramid network is used for performing fusion diffusion on the extracted multi-scale initial feature map to obtain a multi-scale target feature map; according to the SAR image target detection method, the backbone network is adopted to extract multi-scale information, the feature aggregation and diffusion pyramid network is adopted to diffuse the features with rich context information to each detection scale by using a diffusion mechanism, and the accuracy and precision of SAR image target detection are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a SAR image ship detection method and a detection terminal. Background Art

[0002] Synthetic aperture radar imaging is a high-resolution imaging technology that uses radar technology to obtain surface information. Unlike traditional optical remote sensing, SAR imaging is not limited by factors such as weather, light and clouds, and can work effectively under various meteorological conditions. At the same time, SAR imaging is a high-resolution imaging that can perform detailed detection and analysis of specific targets on the surface. These advantages of SAR imaging have made it widely used in the field of ship detection.

[0003] In the prior art, most target detectors are designed based on visible light images and are usually used directly for SAR ship detection without considering the characteristics of SAR images. There are complex background interferences in SAR images, which are easily confused with target ships, reducing the accuracy of detection. The size and angle of the ship will also change with the distance, attitude and motion state, which means that under different conditions, the SAR reflection characteristics of the ship are diverse, resulting in false positives or false negatives, and low accuracy. Summary of the invention

[0004] The embodiment of the present invention provides a SAR image ship detection method and a detection terminal to solve the problem of low accuracy when the visible light image detection method in the prior art is applied to SAR image detection.

[0005] In a first aspect, an embodiment of the present invention provides a SAR image ship detection method, comprising:

[0006] Acquire the SAR image to be detected;

[0007] Input the SAR image to be detected into the SAR ship detection model to obtain the target detection result;

[0008] Among them, the SAR ship detection model includes: backbone feature extraction network, feature aggregation and diffusion pyramid network and detection head;

[0009] The backbone feature extraction network is used to extract features from the SAR image to be detected and obtain a multi-scale initial feature map;

[0010] The feature aggregation and diffusion pyramid network is used to fuse and diffuse the extracted multi-scale initial feature maps to obtain multi-scale target feature maps;

[0011] The detection head is used to detect multi-scale target feature maps and obtain target detection results.

[0012] In a second aspect, an embodiment of the present invention provides a detection terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the SAR image ship detection method provided in the first aspect or any possible implementation method of the first aspect are implemented.

[0013] The embodiment of the present invention provides a SAR image ship detection method and a detection terminal. The above-mentioned SAR image ship detection method includes: obtaining a SAR image to be detected; inputting the SAR image to be detected into a SAR ship detection model to obtain a target detection result; wherein the SAR ship detection model includes: a backbone feature extraction network, a feature aggregation and diffusion pyramid network and a detection head; the backbone feature extraction network is used to extract features of the SAR image to be detected to obtain a multi-scale initial feature map; the feature aggregation and diffusion pyramid network is used to fuse and diffuse the extracted multi-scale initial feature map to obtain a multi-scale target feature map; in the embodiment of the present invention, the backbone network is used to extract the multi-scale information of the SRA image, and the feature aggregation and diffusion pyramid network combines the diffusion mechanism with the pyramid network, and uses the diffusion mechanism to diffuse the features with rich context information to each detection scale, so that the extracted features are more comprehensive and not easily affected by the background, which effectively improves the accuracy and precision of the target detection of the SAR image. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0015] Figure 1 This is a flow chart of an implementation of a SAR image ship detection method provided by an embodiment of the present invention;

[0016] Figure 2 It is an architecture diagram of a backbone feature extraction network provided by an embodiment of the present invention;

[0017] Figure 3 is an architecture diagram of an edge and space information fusion unit provided by an embodiment of the present invention;

[0018] Figure 4 is an architecture diagram of an edge feature information extraction module provided by an embodiment of the present invention;

[0019] Figure 5 is an architecture diagram of a SobelConv module provided in an embodiment of the present invention;

[0020] Figure 6 is an architecture diagram of a fusion feature module provided in an embodiment of the present invention;

[0021] Figure 7 is a schematic diagram of the structure of a SAR image ship detection device provided by an embodiment of the present invention;

[0022] Figure 8 is a schematic diagram of a detection terminal provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention. However, it should be clear to those skilled in the art that the present invention may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present invention.

[0024] In order to make the purpose, technical solutions and advantages of the present invention more clear, specific embodiments will be described below in conjunction with the accompanying drawings.

[0025] See also Figure 1 , which shows a flow chart of an implementation of a SAR image ship detection method provided by an embodiment of the present invention, which is described in detail as follows:

[0026] The above SAR image ship detection method comprises:

[0027] S101: Acquire a SAR image to be detected;

[0028] S102: inputting the SAR image to be detected into the SAR ship detection model to obtain a target detection result;

[0029] Among them, the SAR ship detection model includes: backbone feature extraction network, feature aggregation and diffusion pyramid network and detection head;

[0030] The backbone feature extraction network is used to extract features from the SAR image to be detected and obtain a multi-scale initial feature map;

[0031] The feature aggregation and diffusion pyramid network is used to fuse and diffuse the extracted multi-scale initial feature maps to obtain multi-scale target feature maps;

[0032] The detection head is used to detect multi-scale target feature maps and obtain target detection results.

[0033] In the embodiment of the present invention, a backbone network is used to extract multi-scale information of SRA images, and feature aggregation and diffusion pyramid networks combine the diffusion mechanism with the pyramid network. The diffusion mechanism is used to diffuse features with rich contextual information to each detection scale, so that features at each scale have detailed contextual information, which is conducive to the detection of subsequent targets and is not easily affected by background interference, thereby effectively improving the accuracy and precision of target detection in SAR images.

[0034] Before S101, the above method may further include:

[0035] S103: Establishing an initial network model;

[0036] S104: Divide the SSDD dataset into a training set and a test set at a ratio of 8:2, and divide the ISDD dataset into a training set, a test set, and a validation set at a ratio of 7:1:2, train the initial network model, and obtain a SAR ship detection model.

[0037] In one possible implementation, reference Figure 2 , the backbone feature extraction network may include: a first CBS module, a second CBS module, a first edge and spatial information fusion (ESIF) module, a third CBS module, a second edge and spatial information fusion (ESIF) module, a fourth CBS module, a third edge and spatial information fusion (ESIF) module, a fifth CBS module, a fourth edge and spatial information fusion (ESIF) module and an SPPF module connected sequentially from top to bottom;

[0038] The second edge and spatial information fusion (ESIF) module also outputs the first feature map (P1);

[0039] The third edge and spatial information fusion (ESIF) module also outputs the second feature map (P2);

[0040] The SPPF module outputs the third feature map (P3);

[0041] The edge and spatial information fusion (ESIF) module is used for edge feature extraction and fusion;

[0042] Among them, the first feature map (P1), the second feature map (P2) and the third feature map (P3) form a multi-scale initial feature map.

[0043] The CBS module is used for feature extraction. In the CBS module, C stands for "convolution", which is one of the most commonly used operations in deep learning and is used to extract local features from input data. B stands for "batch normalization", a technique used to accelerate training and improve model performance. It reduces internal covariate shift by normalizing the output of each layer. S stands for the "SiLU" (Sigmoid Linear Unit) activation function, which is used to introduce nonlinearity to the network, enabling the network to learn complex patterns.

[0044] The edge and space information fusion (ESIF) module is used for edge feature extraction and fusion.

[0045] In the embodiment of the present invention, the ESIF module and the CBS module are fused and extracted to obtain multi-scale features.

[0046] In one possible implementation, reference Figure 3 , the edge and space information fusion (ESIF) module may include: a plurality of edge and space information fusion units connected sequentially from top to bottom; the edge and space information fusion unit may include: a sixth CBS module, a seventh CBS module, a Split module and three edge feature information extraction (EFIEM) modules;

[0047] The input end of the sixth CBS module forms the input end of the edge and spatial information fusion unit, and the output of the sixth CBS module is input into the Split module;

[0048] The output of the first output end of the Split module is input into the first edge feature information extraction (EFIEM) module, the output of the first edge feature information extraction (EFIEM) module is input into the second edge feature information extraction (EFIEM) module, and the output of the second edge feature information extraction (EFIEM) module is input into the third edge feature information extraction (EFIEM) module;

[0049] The output of the third edge feature information extraction (EFIEM) module, the output of the first edge feature information extraction (EFIEM) module and the output of the second output end of the Split module are spliced ​​and input into the seventh CBS module. The output end of the seventh CBS module forms the output end of the edge and spatial information fusion unit.

[0050] In the embodiment of the present invention, with reference to the C2f module, an edge and spatial information fusion unit is used to repeatedly extract edge features and fuse them, so that richer image features can be extracted.

[0051] In one possible implementation, the first edge and spatial information fusion (ESIF) module, the second edge and spatial information fusion (ESIF) module, the third edge and spatial information fusion (ESIF) module and the fourth edge and spatial information fusion (ESIF) module include 3, 6, 6 and 3 edge and spatial information fusion units, respectively.

[0052] The number of edge and spatial information fusion units in each edge and spatial information fusion (ESIF) module can be set according to actual application requirements and is not specifically limited here.

[0053] In one possible implementation, reference Figure 4 , the edge feature information extraction (EFIEM) module may include: a SobelConv module, an eighth CBS module, a ninth CBS module and a tenth CBS module;

[0054] The input end of the SobelConv module is connected to the input end of the eighth CBS module to form the input end of the edge feature information extraction (EFIEM) module, and the output of the SobelConv module is spliced ​​with the output of the eighth CBS module and input into the ninth CBS module;

[0055] The output of the ninth CBS module is added to the input of the input end of the edge feature information extraction (EFIEM) module and then input into the tenth CBS module. The output end of the tenth CBS module forms the output end of the edge feature information extraction (EFIEM) module.

[0056] In an embodiment of the present invention, the edge feature information extraction (EFIEM) module includes two branches, the SobelConv module is used to extract edge features, and the eighth CBS module is a convolution branch, which is used to extract spatial information, so that the edge feature information extraction (EFIEM) module can learn a richer image feature representation.

[0057] The formula of the edge feature information extraction (EFIEM) module is expressed as:

[0058] y=CBS(CBS(Concat(Sobel(x),CBS(x)))+x)

[0059] Among them, x is input, y is output, CBS is convolution, Concat is concatenation, and Sobel is the operation of the SobelConv module.

[0060] In one possible implementation, reference Figure 5 , the SobelConv module may include: a Sobel-X block and a Sobel-Y block;

[0061] The input of the Sobel-X block is connected to the input of the Sobel-Y block to form the input of the SobelConv module;

[0062] The output of the Sobel-X block is added to the output of the Sobel-Y block as the output of the SobelConv module.

[0063] The SobelConv module also includes two branches. The Sobel-X block represents the convolution kernel of the Sobel operator applied to the horizontal direction of the input feature map, and the Sobel-Y block represents the convolution kernel of the Sobel operator applied to the vertical direction of the input feature map, thereby extracting complete edge information.

[0064] In one possible implementation, reference Figure 2 , the feature aggregation and diffusion pyramid network may include: two fusion feature (FusionFeature) modules, the eleventh CBS module, the twelfth CBS module and two upsampling (Upsample) modules;

[0065] The first feature map (P1), the second feature map (P2) and the third feature map (P3) are respectively input into a first fusion feature (FusionFeature) module through a first input terminal, a second input terminal and a third input terminal, the first fusion feature (FusionFeature) module outputs a fifth feature map (P5), and the fifth feature map (P5) is respectively input into a first upsample (Upsample) module, an eleventh CBS module and a second input terminal of a second fusion feature (FusionFeature) module;

[0066] The output of the first upsample module is concatenated with the first feature map (P1) to obtain the fourth feature map (P4);

[0067] The output of the eleventh CBS module is concatenated with the third feature map (P3) to obtain the sixth feature map (P6);

[0068] The fourth feature map (P4) and the sixth feature map (P6) are also respectively input into a second fusion feature (FusionFeature) module through the first input terminal and the third input terminal, and the second fusion feature (FusionFeature) module outputs a second target feature map, and the second target feature map is respectively input into a second upsample (Upsample) module and a twelfth CBS module;

[0069] The output of the second upsample module is concatenated with the fourth feature map (P4) to obtain the first target feature map;

[0070] The output of the twelfth CBS module is concatenated with the sixth feature map (P6) to obtain the third target feature map;

[0071] The first target feature map, the second target feature map and the third target feature map form a multi-scale target feature map.

[0072] Feature aggregation and diffusion pyramid network applies diffusion mechanism in the pyramid network. The fusion feature module receives features of three different scales, fuses them and then diffuses them, thereby diffusing features with rich contextual information to each detection scale, so that the features of each detection scale have detailed contextual information, which is conducive to the detection of subsequent targets.

[0073] In one possible implementation, reference Figure 6 The fusion feature module may include: a thirteenth CBS module, a fourteenth CBS module, a third upsample module, a downsample module and at least two DwConv modules;

[0074] The first feature map (P1) is input into the downsample module, the second feature map (P2) is input into the thirteenth CBS module, and the third feature map (P3) is input into the third upsample module;

[0075] The output of the downsample module is concatenated with the output of the thirteenth CBS module and the output of the third upsample module to form a target concatenation graph, which is then input into each DwConv module.

[0076] The output of each DwConv module and the target splicing graph are added and input into the fourteenth CBS module. The output end of the fourteenth CBS module forms the output end of the fusion feature module;

[0077] Among them, the convolution kernels of each DwConv module are different.

[0078] In an embodiment of the present invention, the fusion feature (FusionFeature) module contains at least two convolution kernels of different sizes (each DwConv module) to capture rich information across multiple scales, which is then applied to feature aggregation and diffusion pyramid network to extract rich multi-scale information for target detection with high detection accuracy, effectively avoiding the influence of background and improving the accuracy of target detection.

[0079] In a possible implementation, the number of DwConv modules may be 4; the sizes of the convolution kernels of the DwConv modules may be 5×5, 7×7, 9×9, and 11×11, respectively.

[0080] The more DwConv modules there are, the richer the features extracted, but the amount of calculation will also be greater; therefore, it is necessary to reasonably set the number of DwConv modules to avoid excessive calculations affecting the detection speed while ensuring the quality of the features. The number of DwConv modules includes but is not limited to 4, which can be set according to actual application requirements. At the same time, the convolution kernels of each DwConv module are different, and the size of the convolution kernel can also be set according to actual application requirements, which is not specifically limited here.

[0081] In a possible implementation, the detection head may adopt an Anchor Free mechanism to detect feature maps of three different scales.

[0082] Corresponding to the above-mentioned SAR image ship detection method, it is described in detail below in conjunction with specific embodiments.

[0083] The Windows 11 operating system is used, the CPU is Intel(R) Core(TM) i9-10900K CPU@3.70GHz, the graphics card is NVIDIA RTX A6000, the video memory is 48GB, the PyTorch 1.12.1 framework is used to run the code, and CUDA11.3. The evaluation indicators are precision, recall, average precision at IOU of 0.5 (AP50), average precision at IOU of 0.75 (AP75), small target precision (APS) indicator, and medium target precision (APM) indicator.

[0084] The embodiment of the present invention uses the SSDD dataset and the ISDD dataset for comparative experiments. The SSDD dataset is a dataset for target detection, which contains 1160 images, including ship targets near the coast, as well as land and sea scenes. The ISDD dataset is also a dataset for small target detection, which contains 1284 infrared remote sensing images and 3061 ship target instances.

[0085] A comparative experiment was conducted between the SAR image ship detection provided by the embodiment of the present invention and YOLOv8.

[0086] When the SSDD data set is used, the comparison results between the SAR ship detection model provided by the embodiment of the present invention and YOLOv8 are shown in Table 1.

[0087] Table 1 Performance comparison of SAR ship detection model and YOLOv8 on SSDD dataset

[0088]

[0089] When the ISDD dataset is used, the comparison results between the AR ship detection model provided by the embodiment of the present invention and YOLOv8 are shown in Table 2.

[0090] Table 2 Performance comparison of SAR ship detection model and YOLOv8 on ISDD dataset

[0091]

[0092] It can be seen from Table 1 that, compared with YOLOv8, on the SSDD dataset, the precision of the SAR ship detection model provided by the embodiment of the present invention is improved by 1.1%, the recall rate is improved by 1.3%, and on the more convincing AP75, it is improved by 2.2%, the APS reaches 71.3%, and the APM reaches 78.2%.

[0093] As can be seen from Table 2, since the background in the ISDD dataset is more complex and there are more small targets, the overall detection accuracy is lower than that of the SSDD. However, the SAR ship detection model provided by the embodiment of the present invention performs better than YOLOv8, especially for the small target detection task, with an accuracy improvement of 3.5%, a recall rate improvement of 3.9%, an AP50 improvement of 2.8%, an AP75 improvement of 5.9%, an APS improvement of 3.5%, and an APM improvement of 45.5%.

[0094] The SAR image ship detection method provided by the implementation of the present invention uses the edge and spatial information fusion (ESIF) module to construct a backbone feature extraction network, and combines the SobelConv branch for extracting edge information and the convolution branch (CBS module) for extracting spatial information, so as to learn a richer image feature representation; at the same time, the embodiment of the present invention also proposes a feature aggregation and diffusion pyramid network, which accepts inputs of three scales through the proposed fusion feature (FusionFeature) module, and contains convolution kernels of different sizes to capture rich information across multiple scales, and then diffuses the features with rich contextual information to each detection scale through the diffusion mechanism, so that the features of each scale have detailed contextual information, which is conducive to the detection of subsequent targets. Compared with the YOLOv8 target detection method, when the embodiment of the present invention is applied to SAR image detection, more small ship targets can be detected when detecting small ship targets in SAR images, and the confidence is higher and the missed detection rate is lower, which effectively improves the accuracy of ship detection.

[0095] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0096] The following is an embodiment of the device of the present invention. For details not described in detail therein, reference may be made to the corresponding method embodiment described above.

[0097] Figure 7 The schematic diagram of the structure of the SAR image ship detection device provided by the embodiment of the present invention is shown. For the convenience of description, only the part related to the embodiment of the present invention is shown, which is described in detail as follows:

[0098] like Figure 7 As shown, the SAR image ship detection device includes:

[0099] An image acquisition module 21 is used to acquire a SAR image to be detected;

[0100] The target detection module 22 is used to input the SAR image to be detected into the SAR ship detection model to obtain the target detection result;

[0101] Among them, the SAR ship detection model includes: backbone feature extraction network, feature aggregation and diffusion pyramid network and detection head;

[0102] The backbone feature extraction network is used to extract features from the SAR image to be detected and obtain a multi-scale initial feature map;

[0103] The feature aggregation and diffusion pyramid network is used to fuse and diffuse the extracted multi-scale initial feature maps to obtain multi-scale target feature maps;

[0104] The detection head is used to detect multi-scale target feature maps and obtain target detection results.

[0105] In a possible implementation, the backbone feature extraction network may include: a first CBS module, a second CBS module, a first edge and spatial information fusion module, a third CBS module, a second edge and spatial information fusion module, a fourth CBS module, a third edge and spatial information fusion module, a fifth CBS module, a fourth edge and spatial information fusion module and an SPPF module connected sequentially from top to bottom;

[0106] The second edge and spatial information fusion module also outputs the first feature map;

[0107] The third edge and spatial information fusion module also outputs a second feature map;

[0108] The SPPF module outputs the third feature map;

[0109] The edge and spatial information fusion module is used for edge feature extraction and fusion;

[0110] Among them, the first feature map, the second feature map and the third feature map form a multi-scale initial feature map.

[0111] In a possible implementation, the edge and space information fusion module may include: a plurality of edge and space information fusion units connected sequentially from top to bottom; the edge and space information fusion unit includes: a sixth CBS module, a seventh CBS module, a Split module and three edge feature information extraction modules;

[0112] The input end of the sixth CBS module forms the input end of the edge and spatial information fusion unit, and the output of the sixth CBS module is input into the Split module;

[0113] The output of the first output terminal of the Split module is input into the first edge feature information extraction module, the output of the first edge feature information extraction module is input into the second edge feature information extraction module, and the output of the second edge feature information extraction module is input into the third edge feature information extraction module;

[0114] The output of the third edge feature information extraction module, the output of the first edge feature information extraction module and the output of the second output end of the Split module are spliced ​​and input into the seventh CBS module. The output end of the seventh CBS module forms the output end of the edge and spatial information fusion unit.

[0115] In a possible implementation, the first edge and spatial information fusion module, the second edge and spatial information fusion module, the third edge and spatial information fusion module and the fourth edge and spatial information fusion module may include 3, 6, 6 and 3 edge and spatial information fusion units, respectively.

[0116] In a possible implementation, the edge feature information extraction module may include: a SobelConv module, an eighth CBS module, a ninth CBS module, and a tenth CBS module;

[0117] The input end of the SobelConv module is connected to the input end of the eighth CBS module to form the input end of the edge feature information extraction module. The output of the SobelConv module is spliced ​​with the output of the eighth CBS module and then input into the ninth CBS module.

[0118] The output of the ninth CBS module is added to the input of the input end of the edge feature information extraction module and then input into the tenth CBS module. The output end of the tenth CBS module forms the output end of the edge feature information extraction module.

[0119] In a possible implementation, the SobelConv module includes: a Sobel-X block and a Sobel-Y block;

[0120] The input of the Sobel-X block is connected to the input of the Sobel-Y block to form the input of the SobelConv module;

[0121] The output of the Sobel-X block is added to the output of the Sobel-Y block as the output of the SobelConv module.

[0122] In a possible implementation, the feature aggregation and diffusion pyramid network may include: two fusion feature modules, an eleventh CBS module, a twelfth CBS module, and two upsampling modules;

[0123] The first feature map, the second feature map and the third feature map are respectively input into the first fusion feature module through the first input terminal, the second input terminal and the third input terminal, the first fusion feature module outputs the fifth feature map, and the fifth feature map is respectively input into the first upsampling module, the eleventh CBS module and the second input terminal of the second fusion feature module;

[0124] The output of the first upsampling module is concatenated with the first feature map to obtain the fourth feature map;

[0125] The output of the eleventh CBS module is concatenated with the third feature map to obtain the sixth feature map;

[0126] The fourth feature map and the sixth feature map are also respectively input into the second fusion feature module through the first input terminal and the third input terminal, and the second fusion feature module outputs a second target feature map, and the second target feature map is respectively input into the second upsampling module and the twelfth CBS module;

[0127] The output of the second upsampling module is concatenated with the fourth feature map to obtain the first target feature map;

[0128] The output of the twelfth CBS module is concatenated with the sixth feature map to obtain the third target feature map;

[0129] The first target feature map, the second target feature map and the third target feature map form a multi-scale target feature map.

[0130] In a possible implementation, the fusion feature module may include: a thirteenth CBS module, a fourteenth CBS module, a third upsampling module, a downsampling module, and at least two DwConv modules;

[0131] The first feature map is input into the downsampling module, the second feature map is input into the thirteenth CBS module, and the third feature map is input into the third upsampling module;

[0132] The output of the downsampling module is concatenated with the output of the thirteenth CBS module and the output of the third upsampling module to form a target concatenation graph, which is then input into each DwConv module.

[0133] The output of each DwConv module and the target splicing graph are added and input into the fourteenth CBS module. The output end of the fourteenth CBS module forms the output end of the fusion feature module;

[0134] Among them, the convolution kernels of each DwConv module are different.

[0135] In a possible implementation, the number of DwConv modules may be 4; the sizes of the convolution kernels of the DwConv modules are 5×5, 7×7, 9×9, and 11×11, respectively.

[0136] Figure 8 Schematic diagram of the detection terminal 3 provided in the embodiment of the present invention. Figure 8 As shown, the detection terminal 3 of this embodiment includes: a processor 30 and a memory 31. The memory 31 is used to store a computer program 32, and the processor 30 is used to call and run the computer program 32 stored in the memory 31 to perform the steps in the above-mentioned SAR image ship detection method embodiments, such as Figure 1 Alternatively, the processor 30 is used to call and run the computer program 32 stored in the memory 31 to implement the functions of each module / unit in the above-mentioned device embodiments, such as Figure 7 The functions of modules 21 to 22 are shown.

[0137] Exemplarily, the computer program 32 may be divided into one or more modules / units, one or more modules / units are stored in the memory 31 and executed by the processor 30 to implement the present invention. One or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the detection terminal 3. For example, the computer program 32 may be divided into Figure 7 Modules / units 21 to 22 are shown.

[0138] The detection terminal 3 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The detection terminal 3 may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art will appreciate that Figure 8 It is only an example of the detection terminal 3 and does not constitute a limitation on the detection terminal 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal may also include input and output devices, network access devices, buses, etc.

[0139] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0140] The memory 31 may be an internal storage unit of the detection terminal 3, such as a hard disk or memory of the detection terminal 3. The memory 31 may also be an external storage device of the detection terminal 3, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the detection terminal 3. Further, the memory 31 may also include both an internal storage unit and an external storage device of the detection terminal 3. The memory 31 is used to store computer programs and other programs and data required by the terminal. The memory 31 may also be used to temporarily store data that has been output or is to be output.

[0141] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0142] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0143] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0144] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0145] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0147] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0148] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.

Claims

1. A method for detecting ships in SAR images, characterized in that: include: Acquire the SAR image to be detected; Inputting the SAR image to be detected into a SAR ship detection model to obtain a target detection result; The SAR ship detection model includes: a backbone feature extraction network, a feature aggregation and diffusion pyramid network, and a detection head; The backbone feature extraction network is used to extract features from the SAR image to be detected to obtain a multi-scale initial feature map; The feature aggregation and diffusion pyramid network is used to fuse and diffuse the extracted multi-scale initial feature maps to obtain a multi-scale target feature map; The detection head is used to detect the multi-scale target feature map to obtain the target detection result.

2. The SAR image ship detection method according to claim 1, characterized in that: The backbone feature extraction network includes: a first CBS module, a second CBS module, a first edge and spatial information fusion module, a third CBS module, a second edge and spatial information fusion module, a fourth CBS module, a third edge and spatial information fusion module, a fifth CBS module, a fourth edge and spatial information fusion module and an SPPF module, which are sequentially connected from top to bottom; The second edge and spatial information fusion module further outputs a first feature map; The third edge and spatial information fusion module also outputs a second feature map; The SPPF module outputs a third feature map; The edge and spatial information fusion module is used to extract and fuse edge features; The first feature map, the second feature map, and the third feature map form the multi-scale initial feature map.

3. The SAR image ship detection method according to claim 2, characterized in that: The edge and spatial information fusion module includes: a plurality of edge and spatial information fusion units connected sequentially from top to bottom; the edge and spatial information fusion unit includes: a sixth CBS module, a seventh CBS module, a Split module and three edge feature information extraction modules; The input end of the sixth CBS module forms the input end of the edge and spatial information fusion unit, and the output of the sixth CBS module is input into the Split module; The output of the first output end of the Split module is input into the first edge feature information extraction module, the output of the first edge feature information extraction module is input into the second edge feature information extraction module, and the output of the second edge feature information extraction module is input into the third edge feature information extraction module; The output of the third edge feature information extraction module, the output of the first edge feature information extraction module, and the output of the second output end of the Split module are spliced ​​and input into the seventh CBS module. The output end of the seventh CBS module forms the output end of the edge and spatial information fusion unit.

4. The SAR image ship detection method according to claim 2, characterized in that: The first edge and spatial information fusion module, the second edge and spatial information fusion module, the third edge and spatial information fusion module and the fourth edge and spatial information fusion module include 3, 6, 6 and 3 edge and spatial information fusion units respectively.

5. The SAR image ship detection method according to claim 3, characterized in that: The edge feature information extraction module includes: a SobelConv module, an eighth CBS module, a ninth CBS module and a tenth CBS module; The input end of the SobelConv module is connected to the input end of the eighth CBS module to form the input end of the edge feature information extraction module, and the output of the SobelConv module is spliced ​​with the output of the eighth CBS module and then input into the ninth CBS module; The output of the ninth CBS module is added to the input of the input end of the edge feature information extraction module and then input into the tenth CBS module. The output end of the tenth CBS module forms the output end of the edge feature information extraction module.

6. The SAR image ship detection method according to claim 5, characterized in that: The SobelConv module includes: a Sobel-X block and a Sobel-Y block; The input of the Sobel-X block is connected to the input of the Sobel-Y block to form the input of the SobelConv module; The output of the Sobel-X block is added to the output of the Sobel-Y block as the output of the SobelConv module.

7. The SAR image ship detection method according to any one of claims 2 to 6, characterized in that: The feature aggregation and diffusion pyramid network includes: two fusion feature modules, the eleventh CBS module, the twelfth CBS module and two upsampling modules; The first feature map, the second feature map, and the third feature map are respectively input into a first fusion feature module through a first input terminal, a second input terminal, and a third input terminal, the first fusion feature module outputs a fifth feature map, and the fifth feature map is respectively input into a first upsampling module, the eleventh CBS module, and a second input terminal of the second fusion feature module; The output of the first upsampling module is concatenated with the first feature map to obtain a fourth feature map; The output of the eleventh CBS module is concatenated with the third feature map to obtain a sixth feature map; The fourth feature map and the sixth feature map are further input into the second fusion feature module through the first input terminal and the third input terminal respectively, and the second fusion feature module outputs a second target feature map, and the second target feature map is respectively input into the second upsampling module and the twelfth CBS module; The output of the second upsampling module is spliced ​​with the fourth feature map to obtain a first target feature map; The output of the twelfth CBS module is concatenated with the sixth feature map to obtain a third target feature map; The first target feature map, the second target feature map, and the third target feature map form the multi-scale target feature map.

8. The SAR image ship detection method according to claim 7, characterized in that: The fusion feature module includes: a thirteenth CBS module, a fourteenth CBS module, a third upsampling module, a downsampling module and at least two DwConv modules; The first feature map is input into the downsampling module, the second feature map is input into the thirteenth CBS module, and the third feature map is input into the third upsampling module; The output of the downsampling module is spliced ​​with the output of the thirteenth CBS module and the output of the third upsampling module to form a target splicing graph, and the target splicing graph is input into each DwConv module respectively; The output of each DwConv module and the target spliced ​​graph are added together and input into the fourteenth CBS module, and the output end of the fourteenth CBS module forms the output end of the fusion feature module; Among them, the convolution kernels of each DwConv module are different.

9. The SAR image ship detection method according to claim 8, characterized in that: The number of the DwConv modules is 4; the sizes of the convolution kernels of the DwConv modules are 5×5, 7×7, 9×9, and 11×11, respectively.

10. A detection terminal, characterized in that: The system comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the steps of the SAR image ship detection method according to any one of claims 1 to 9.