SAR image ship detection method and device for complex scene
By building a SAR image ship detection network and utilizing modules such as FCSI, VCCB, and WPCM, the problem of low accuracy of SAR image ship detection in complex scenarios is solved, more efficient feature extraction and noise resistance are achieved, and detection performance is improved.
Patent Information
- Application Number
- CN202510804231.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-10-21
AI Technical Summary
Existing SAR image ship detection methods have low detection rates in complex scenes, are difficult to effectively extract small target features, and are affected by coherent speckle noise and complex background interference, resulting in low detection accuracy.
A SAR image ship detection network is constructed, which adopts the full-channel spatial interaction module FCSI, the longitudinal compression backbone network VCCB, the wavelet preprocessing downsampling module WPCM and the short-distance aggregation dispersion neck network SADNet. Through feature extraction and fusion, the detection accuracy and anti-noise ability are improved.
The detection accuracy of ship targets in SAR images is significantly improved in complex scenarios, the multi-scale target detection performance is improved, and the model's anti-noise interference ability and feature fusion effect are enhanced.
Smart Images

Figure CN120823362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of SAR image ship detection, and in particular to a SAR image ship detection method and device for complex scenes. Background Art
[0002] Synthetic Aperture Radar (SAR) is an active microwave remote sensing technology that uses virtual arrays and pulse compression techniques to acquire high-resolution two-dimensional images of land objects. Due to its unique structure and rich polarimetric information, it can observe the ground in all weather conditions and at all times of the day, and can even acquire high-resolution images under adverse conditions such as nighttime or cloud cover. It is widely used in various civil and military fields, such as urban planning, environmental and natural disaster monitoring, and ship target detection. Ship target detection is a key means of achieving automatic ship identification and positioning, and is of great significance for coastal and waterway security monitoring. However, due to the very small size of various ships, they only appear as a few pixels in wide-width SAR images, and ship features are often submerged by complex backgrounds, resulting in low detection rates. Therefore, designing a ship target detection method for SAR images in complex scenes is of great significance for improving the accuracy of remote sensing ship detection.
[0003] Due to the unique imaging mechanism of SAR, which acquires ground target information through the backscattering of radar waves, SAR images, while rich in structural and physical characteristics, as well as intensity and phase information, are inferior to RGB images in terms of detail, texture, and color. Furthermore, the fading of radar target echo signals can lead to the generation of coherent speckle noise, which reduces the image's signal-to-noise ratio, blurring details, and disappearing features. Furthermore, various challenges arise in different complex scenarios. For example, ship detection using SAR images in seaport scenarios is affected by large variations in target scale and variable ground and sea clutter intensity. In port scenarios, SAR image targets have diverse structures and differing polarization modes, making target feature extraction difficult. In inland scenarios, the close spatial proximity of ships and ground clutter, coupled with high grayscale texture similarity, can also lead to low target detection accuracy.
[0004] Existing detection methods, such as CSP and ResNet, have poor detection performance for small objects and complex scenes with coherent speckle noise. This is primarily due to excessive downsampling and incomplete feature extraction. Furthermore, while the Feature Fusion Pyramid Network (FPN) can achieve cross-scale feature fusion to a certain extent, its excessively long fusion path can lead to small object loss and inaccurate feature localization, resulting in poor detection results. Summary of the Invention
[0005] In order to improve the detection accuracy of ship targets in SAR images under complex scenes, the present invention proposes a SAR image ship detection method and device for complex scenes.
[0006] In a first aspect, the present invention provides a SAR image ship detection method for complex scenes, comprising: constructing and training a SAR image ship detection network, and using the trained SAR image ship detection network to detect an input SAR image, characterized in that the SAR image ship detection network comprises a backbone network and a neck network; the backbone network extracts multi-scale features of the input SAR image using stacked feature extraction layers, the multi-scale features comprising P2 feature maps, P3 feature maps, and P4 feature maps output by the second, third, and fourth layers of the backbone network; the neck network is used to fuse and detect the multi-scale features output by the backbone network to obtain the position of the ship in the SAR image; wherein the feature extraction layer comprises a full-channel spatial interaction module FCSI, and the feature extraction process of the FCSI comprises:
[0007] For the input feature map F in Perform convolution with a convolution kernel of 1, normalization, and SiLU activation function activation processing; then perform channel segmentation on the processed feature map to form the initial feature maps f1 and f2 in the channel dimension;
[0008] Based on the initial feature maps f1 and f2, m batches of channel space interaction operations are continuously performed to obtain m intermediate feature maps f1 (i) and i=1,2,…,m, m is a positive integer not less than 2; wherein, t is set to be a positive integer, t∈[2,m], then the channel space interaction operation of the t-th batch specifically includes: the intermediate feature map f1 obtained in the (t-1)th batch (t-1) and F1 (t-1) Input the bottleneck module for feature extraction to form feature maps f1′ and f2′; then use CBS to extract the feature maps f1′ and f2′ respectively. 1×1 After adjusting the number of channels of feature maps f1′ and f2′ to 1, the Softmax activation function is activated to obtain the spatial weights a1 and b1 of feature maps f1′ and f2′ respectively; finally, the feature map f1′ and the weight b1 are multiplied element by element to obtain the intermediate feature map f1 (t) , perform element-by-element multiplication of the feature map f2′ and the weight a1 to obtain the intermediate feature map Among them, for the channel space interaction operation of the first batch, the initial feature maps f1 and f2 are used as the input of the bottleneck module respectively;
[0009] The obtained m intermediate feature maps f1 (i) and Perform splicing and fusion operations to obtain the output feature map F of the FCSI out .
[0010] Furthermore, the backbone network also includes a wavelet preprocessing downsampling module WPCM located before the first feature extraction layer; the WPCM is used to expand the receptive field of the input feature map and reduce noise interference; the corresponding feature map processing process includes: first using two branches to process the input feature map respectively, and then fusing the outputs of the two branches to obtain the output feature map of the WPCM: wherein, one branch includes: performing convolution processing on the input feature map to obtain the corresponding convolution result; the other branch includes: first performing wavelet transform on the input feature map to obtain a low-frequency component and high-frequency components in three directions of horizontal, vertical and diagonal directions; then performing convolution processing on the four components respectively to obtain four convolution results accordingly; finally, applying inverse wavelet transform to the four convolution results to obtain a reconstructed feature map.
[0011] Furthermore, the neck network includes a multi-scale information adaptive aggregation module, and the feature fusion process of the multi-scale information adaptive aggregation module includes:
[0012] P3″=CBR(Cat(U b (Conv 1×1 (P4))×w1,P3×w2,Conv 3×3 (P2)×w3))
[0013] Among them, P2, P3 and P4 are the P2 feature map, P3 feature map and P4 feature map output by the backbone network respectively; Conv 1×1 It is a convolution operation with a convolution kernel of 1; U b is an upsampling operation; Conv 3×3 is a convolution operation with a kernel size of 3 and a stride of 2; w1, w2, and w3 are three independent adaptive weights; CBR is a point-by-point convolution operation, batch normalization operation, and ReLU activation function; Cat is a splicing operation; and P3″ is the output of the multi-scale information adaptive aggregation module.
[0014] Furthermore, the neck network further includes a residual module, and the feature fusion process of the residual module includes:
[0015] F4=Cat(Conv 3×3 (P3″),Conv 1×1 (P4))
[0016] F3=Cat(P3″,P3)
[0017] F2=Cat(U a (P3″),P2)
[0018] Among them, F2, F3 and F4 are the output feature maps of the residual module.
[0019] Furthermore, the neck network also includes a grouped convolution residual feature fusion module GCRF, and the feature fusion process of the GCRF includes:
[0020] F=Conv 1×1 (F in )
[0021] F out =(n-1)Conv 3×3 ((n-1)GConv 5×5 (Conv 1×1 (GConv 5×5 (F))))+F
[0022] Among them, F in and F out It is the input feature map and output feature map of GCRF, + represents the residual connection, GConv 5×5 It is a grouped large kernel convolution operation, where n is a positive integer not less than 2.
[0023] In a second aspect, the present invention provides a SAR image ship detection device for complex scenes, comprising:
[0024] Network building unit, used to build and train SAR image ship detection network;
[0025] A detection unit, configured to detect an input SAR image using a trained SAR image ship detection network;
[0026] The SAR image ship detection network includes a backbone network and a neck network; the backbone network uses stacked feature extraction layers to extract multi-scale features of the input SAR image, and the multi-scale features include the P2 feature map, P3 feature map, and P4 feature map output by the second, third, and fourth layers of the backbone network; the neck network is used to fuse and detect the multi-scale features output by the backbone network to obtain the position of the ship in the SAR image; the feature extraction layer includes a full-channel spatial interaction module FCSI, and the feature extraction process of the FCSI includes:
[0027] For the input feature map F in Perform convolution with a convolution kernel of 1, normalization, and SiLU activation function activation processing; then perform channel segmentation on the processed feature map to form the initial feature maps f1 and f2 in the channel dimension;
[0028] Based on the initial feature maps f1 and f2, m batches of channel space interaction operations are continuously performed to obtain m intermediate feature maps f1 (i)and i=1,2,…,m, m is a positive integer not less than 2; wherein, t is set to be a positive integer, t∈[2,m], then the channel space interaction operation of the t-th batch specifically includes: the intermediate feature map f1 obtained in the (t-1)th batch (t-1) and F1 (t-1) Input the bottleneck module for feature extraction to form feature maps f1′ and f2′; then use CBS to extract the feature maps f1′ and f2′ respectively. 1×1 After adjusting the number of channels of feature maps f1′ and f2′ to 1, the Softmax activation function is activated to obtain the spatial weights a1 and b1 of feature maps f1′ and f2′ respectively; finally, the feature map f1′ and the weight b1 are multiplied element by element to obtain the intermediate feature map f1 (t) , perform element-by-element multiplication of the feature map f2′ and the weight a1 to obtain the intermediate feature map Among them, for the channel space interaction operation of the first batch, the initial feature maps f1 and f2 are used as the input of the bottleneck module respectively;
[0029] The obtained m intermediate feature maps f1 (i) and Perform splicing and fusion operations to obtain the output feature map F of the FCSI out .
[0030] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0031] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.
[0032] The beneficial effects of the present invention are:
[0033] (1) A full-channel spatial interaction module (FCSI) is designed, and based on it, a longitudinal compression backbone network (VCCB) is designed to reduce the model depth, thereby reducing the loss of detail features during downsampling. In addition, it also integrates more shallow feature maps in the detail features, improving the ability to extract detail features of small targets, which helps to improve the effect of small target detection.
[0034] (2) A preprocessing downsampling module (WPCM) combined with wavelet convolution is proposed. This module uses wavelet convolution to enhance the receptive field of the backbone network feature extraction and decomposes the feature map, thereby separately processing the high-frequency components where noise is concentrated, thereby increasing the model's ability to resist noise interference. This has a very significant improvement effect on the detection of ship targets in SAR images, where noise interference and complex background interference cannot be avoided.
[0035] (3) A short-distance aggregation and dispersion structured neck network (SADNet) is designed. This network shortens the fusion path based on the third-layer features, thereby reducing feature loss during fusion and improving target positioning accuracy. In addition, a grouped convolutional residual fusion module (GCRF) is introduced into this neck network to further improve the feature fusion effect.
[0036] Experiments validated the proposed method on various public SAR image datasets. The results show that the proposed method outperforms other advanced methods overall, enabling more accurate detection of multi-scale vessel targets in SAR images in complex scenarios. For example, the proposed method achieved a 3% improvement in mAP0.5 on the HRSID dataset, 0.3% on the SSDD dataset, and 3.1% on the LS-SSDD-v1.0 dataset. Furthermore, a series of complex scene experiments confirmed the superior detection and generalization performance of the proposed model in these scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 The overall structure of the CSS-YOLO SAR image ship detection network provided by the embodiment of the present invention;
[0038] Figure 2 The FCSI module structure provided by the embodiment of the present invention;
[0039] Figure 3 The WPCM module structure provided by the embodiment of the present invention;
[0040] Figure 4 The SADNet structure provided by the embodiment of the present invention;
[0041] Figure 5 Visual comparison results on the HRSID dataset provided by the embodiment of the present invention;
[0042] Figure 6Comparison results of the inland river scene of the HRSID dataset provided by the embodiment of the present invention: green boxes indicate correct predictions; red boxes indicate missed predictions; blue boxes indicate incorrect predictions; (a) actual results; (b) YOLOv5s; (c) YOLOv8s; (d) YOLOv9s; (e) YOLO11s; (f) CSS-YOLO;
[0043] Figure 7 Comparison results of the seaport scene of the HRSID dataset provided by the embodiment of the present invention: green boxes indicate correct predictions; red boxes indicate missed predictions; blue boxes indicate incorrect predictions; (a) actual situation; (b) YOLOv5s; (c) YOLOv8s; (d) YOLOv9s; (e) YOLO11s; (f) CSS-YOLO;
[0044] Figure 8 Comparison results of the port scene of the HRSID dataset provided by the embodiment of the present invention; green boxes indicate correct predictions; red boxes indicate missed predictions; blue boxes indicate incorrect predictions; (a) actual situation; (b) YOLOv5s; (c) YOLOv8s; (d) YOLOv9s; (e) YOLO11s; (f) CSS-YOLO;
[0045] Figure 9 The dataset specifications provided in the embodiments of the present invention, where the first row is the SSDD dataset and the second row is the LS-SSDD-v1.0 dataset; (a) target size distribution diagram; (b) target position distribution diagram;
[0046] Figure 10 This is a visualization comparison result on the SSDD dataset provided by an embodiment of the present invention; the pink arrow indicates that the pointed area is magnified, and the magnified content is the content in the dotted box; the red arrow indicates that the red box it points to is a missed target, the green arrow indicates that the green box it points to is a correctly detected target, and the blue box is an incorrectly predicted target;
[0047] Figure 11 Visual comparison results on the LS-SSDD-v1.0 dataset provided by an embodiment of the present invention; yellow and pink arrows indicate zooming in on the areas they point to, with the zoomed-in content being the content within the corresponding dashed boxes. Red arrows indicate missed targets within the red boxes they point to, while green arrows indicate correctly detected targets within the green boxes they point to, and blue boxes indicate incorrectly predicted targets.
[0048] Figure 12 A schematic diagram of the structure of a SAR image ship detection device for complex scenes provided by an embodiment of the present invention;
[0049] Figure 13This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0050] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0051] The present invention aims to improve the performance of SAR image ship detection in complex environments and solve the problem that targets are easily disturbed by complex backgrounds. The core of the present invention is to build a SAR image ship detection network CSS-YOLO, Figure 1 The overall architecture of CSS-YOLO consists of the VCCB backbone network and the SADNet network. The VCCB backbone network uses stacked feature extraction layers to extract multi-scale features from the input SAR image. These multi-scale features include the P2, P3, and P4 feature maps output by the second, third, and fourth layers of the backbone network. The SADNet network fuses and detects the multi-scale features output by the backbone network to determine the ship's position in the SAR image.
[0052] Example 1
[0053] In the field of ship target detection using SAR images in complex scenarios, the unique imaging mechanism of SAR images, particularly in ports, estuaries, and inland areas, makes the detection of small ships extremely susceptible to interference from sea and ground clutter. This places higher demands on target feature extraction and detection methods. As network depth increases, detailed target features are easily lost. However, shallow networks extract relatively little global semantic information, making it difficult to address the problem of similar backgrounds and targets.
[0054] Traditional target detection methods often enhance feature extraction capabilities by deepening convolutional layers. While this can increase the richness of semantic information, it can also blur details, reducing the accuracy of small target detection. However, small and medium-sized targets are the majority of targets in SAR image ship detection. Some researchers have opted to add a P2 layer detection head to improve small target detection performance, but this also significantly increases the number of parameters and computational complexity. Therefore, fully extracting features without losing small targets and detailed features is a challenge facing SAR image ship target detection in complex scenarios.
[0055] In order to strike a balance between detection performance and computational complexity, this paper proposes a longitudinal compression backbone network VCCB suitable for ship target detection in SAR images of complex scenes. Figure 1 As shown in the figure. In the backbone network, a P2 detection layer is added to improve the network's ability to extract detailed features. In addition, considering the insensitivity of high-level semantic information to small targets and detailed features, the P5 detection layer is removed, and SPPF is added to the P4 detection layer to reduce the blurring of details caused by the model being too deep. In other words, the backbone network in this embodiment will eventually extract feature maps of three scales from the input SAR image, namely the P2 feature map, P3 feature map, and P4 feature map output by the second, third, and fourth layers of the backbone network.
[0056] In addition, the network not only pays more attention to the detailed feature extraction of the target, but also uses fewer model parameters, making the model more lightweight. However, removing the P5 detection layer will weaken the overall feature extraction ability of the model. Therefore, the embodiment of the present invention designs a full-channel spatial interaction module FCSI and introduces it into the feature extraction layer of the backbone network to enhance the feature extraction ability. The structure of FCSI is as follows Figure 2 As shown, it mainly includes three stages:
[0057] The first stage: input feature map F in The convolution kernel is set to 1, followed by normalization and activation using the SiLU activation function. The processed feature map is then segmented along the channel dimension to form the initial feature maps f1 and f2. This feature processing can reduce the number of model parameters and computational complexity.
[0058] The second stage: Based on the initial feature maps f1 and f2, m batches of channel space interaction operations are continuously performed to obtain m intermediate feature maps f1. (i) and , i=1,2,…,m, m is a positive integer not less than 2;
[0059] Specifically, set t to a positive integer, t∈[2,m], then the channel space interaction operation of the t-th batch specifically includes: the intermediate feature map f1 obtained in the (t-1)th batch (t-1) and F1 (t-1) Input the bottleneck module Bottleneck for feature extraction to form feature maps f1′ and f2′; then use CBS 1×1 After adjusting the number of channels of feature maps f1′ and f2′ to 1, the Softmax activation function is activated to obtain the spatial weights a1 and b1 of feature maps f1′ and f2′ respectively; finally, the feature map f1′ and the weight b1 are multiplied element by element to obtain the intermediate feature map f1 (t), perform element-by-element multiplication of the feature map f2′ and the weight a1 to obtain the intermediate feature map Among them, for the channel space interaction operation of the first batch, the initial feature maps f1 and f2 are respectively used as the input of the bottleneck module, and the rest of the process is the same as the operation of the subsequent batches, which will not be repeated here. In this embodiment, the bottleneck module Bottleneck includes 4 serial 3×3 convolutions. After the above operation, while reducing the number of model parameters and computational complexity, the intermediate feature map f1 obtained in this stage is obtained through the interaction of spatial weights. (i) and All channel information is included.
[0060] The third stage: the obtained m intermediate feature maps f1 (i) and Perform splicing and fusion operations to obtain the output feature map F of the FCSI out .
[0061] In this embodiment, in order to improve the performance of feature extraction, spatial weights are introduced. On the basis of introducing spatial weights, the design of segmented channels is retained to reduce complexity. The feature map is divided into two parts on the channel for feature extraction. After each extraction, the two parts of features are interacted through spatial weights. This will enhance the spatial information expression of the two parts, make it more sensitive in the model backtracking process, and improve the effect and convergence speed of feature extraction.
[0062] Algorithm 1 below shows the specific steps of feature extraction in the VCCB backbone network. Step 1 shows the internal implementation of FCSI, Step 2 shows the details of each layer of VCCB feature extraction, and Step 3 shows the output after removing the fifth layer.
[0063]
[0064]
[0065] In this embodiment, the design purpose of FCSI is to solve the problem of incomplete feature extraction caused by VCCB removing the P5 detection layer to reduce the model depth. Compared with C2f, FCSI performs spatial weight interaction in a dual-branch parallel manner, which uses a very small amount of parameters in exchange for more effective feature extraction capabilities, which meets the needs of complex scene target detection for detail feature extraction. Not only that, the purpose of VCCB is to extract as much larger receptive field and more feature information as possible while reducing the extraction path and downsampling. The longitudinal compression removes the last downsampling feature extraction operation and adopts the FCSI with stronger extraction capability. The model not only has a larger receptive field to distinguish low clutter from targets, but also improves the ability to extract detailed features and reduces the loss of detailed features, thereby significantly improving the effect of SAR image ship target detection in complex scenes.
[0066] Example 2
[0067] During the SAR imaging process, noise interference is inevitably introduced by factors such as atmospheric reflection, sea surface reflection, and signal fading. The most common type of noise is speckle noise, which appears as dense white spots in the image. This type of noise severely degrades the signal-to-noise ratio (SNR) of the image and significantly impacts the accuracy of target detection. Furthermore, due to the unique imaging mechanism of SAR images, complex ground backgrounds can reflect features similar to those of the target, thus affecting target detection. Since ships are located on the coastal waters, in order to eliminate the effects of background reflections, the receptive field can generally be expanded to obtain global information such as the target's location and neighborhood for differentiation. However, existing methods for expanding the receptive field have numerous drawbacks. For example, using large kernel convolutions results in a large number of parameters, using dilated convolutions results in the loss of detailed features, and using adaptive convolution kernels increases model depth.
[0068] The embodiment of the present invention takes into account that the coherent speckle noise is mainly concentrated in the high-frequency components of the image. The wavelet transform (WT) can separate these high-frequency noises from other image features to form multiple wavelet domains, and perform separate processing in the wavelet domain to reduce the impact of noise interference. In addition, the inverse wavelet transform (IWT) can fuse information of different frequencies to make the image smoother. Therefore, in order to expand the receptive field and reduce noise interference, the embodiment of the present invention introduces wavelet convolution, that is, using the dimensional characteristics of the wavelet domain, convolution is performed in the wavelet domain, which can expand the receptive field with a very small increase in the number of parameters; and a preprocessing module based on wavelet convolution is designed to replace the first downsampling operation, such as Figure 3 As shown, it is named as wavelet preprocessing downsampling module WPCM, and its main steps mainly include four stages:
[0069] The first stage: First, the initial feature map with 3 input channels is subjected to the first wavelet transform. The specific implementation method is to first use the four filters shown in formula (1) to perform a depth convolution operation with a step size of 2 to form four different frequency components f in different directions. LL , f LH , f LH and f HH , and then using formula (2), the input feature map can be divided into a low-frequency component and three high-frequency components.
[0070]
[0071] [X LL ,X LH ,X HL ,X HH ]=WT([f LL , f LH , f LH , f HH ],x)(2)
[0072] Among them, f LL is a low-pass filter, f LH , f LH , f HH is a set of high-pass filters. WT represents the process of wavelet transform. X LL represents the low-frequency component, and X LH ,X HL ,X HH They represent the horizontal, vertical and diagonal high frequency components respectively. The dimensions of these four components are the input features Figure X half.
[0073] The second stage: Convolution is performed on the four components separately, and the high-frequency components where the noise is concentrated can be processed separately to reduce noise interference. This process can be expressed by formula (3):
[0074] [X′ LL ,X′ LH ,X′ HL ,X′ HH ]=Conv 3×3 ([X LL ,X LH ,X HL ,X HH ])(3)
[0075] The third stage: Apply the inverse wavelet transform (IWT) operation to the result to restore it to a feature map with 3 channels and the same dimension as the original. This process can be expressed by formula (4):
[0076] X′=Conv_transposed([f LL ,f LH ,f LH ,f HH ],[X′ LL ,X′ LH ,X′ HL ,X′ HH ]) (4)
[0077] The fourth stage: The result of convolution of the original feature map is fused with the result of inverse wavelet transform, retaining the original feature information, and finally performing downsampling operation through convolution. This process can be expressed by formula (5):
[0078] X″=CBS(X′+Conv 3×3 (X))(5)
[0079] Among them, Conv 3×3 is a convolution operation with a convolution kernel of 3. Conv_transposed is an inverse wavelet transform that directly uses the transposed convolution kernel to perform convolution due to the orthogonal basis formed by formula (4). + is an element-by-element addition operation. CBS is a convolution, normalization, and SiLU activation function activation operation.
[0080] In this embodiment, WPCM decomposes the input feature map into low-frequency and high-frequency components through wavelet transform and performs independent convolution on each. This not only effectively reduces the interference of coherent speckle noise but also expands the model's receptive field, enabling it to better capture global information and improving the model's detection performance for ship targets in SAR imagery under complex scenarios. Furthermore, the WPCM module design achieves significant performance improvements without significantly increasing the number of parameters.
[0081] Example 3
[0082] Because ship targets in SAR images are small and feature-rich, confidence levels are low. Integrating high-level semantics can further reduce confidence, leading to missed detections. Furthermore, due to insufficient semantic information when determining targets reflected from complex backgrounds, false detections can occur. This places higher demands on feature fusion for ship target detection in SAR images in complex scenarios, especially for detecting small targets.
[0083] In the study of neck feature fusion networks for ship detection in SAR images, some researchers are still studying and improving the classic FPN network, which utilizes a pyramid fusion approach for multi-scale fusion and further integrates it with corresponding convolutional modules. Although FPN networks have certain advantages, their performance varies across different detection tasks. In SAR images, some small ships are only a few pixels, and the top-down approach of FPN can cause these target and position features to be lost. Furthermore, performing the same fusion operation on feature maps of different scales fails to distinguish important information from noise, failing to highlight and utilize the advantages of feature maps of different scales.
[0084] In order to solve the problem that the fusion path is too long and the detail features are lost and the multi-scale features cannot be adaptively selected, the embodiment of the present invention proposes a feature fusion neck network, called the short-distance aggregation dispersion neck network SADNet, whose structure is as follows: Figure 4 The neck network provided in this embodiment performs feature fusion from three aspects.
[0085] First, a multi-scale information adaptive aggregation module based on the P3 feature map is proposed. The main idea of this module is that since the P3 layer is in the middle and not only has rich detail features but also contains rich semantic information, the P3 feature map is used as the benchmark. On this basis, the P2 feature map is downsampled and fused with the P3 feature map to enrich the small target features, and the P4 feature map is upsampled and fused with the P3 feature map to enrich the semantic information. Secondly, the present invention adds adaptive weights during fusion to adjust the advantages of different feature maps and balance the resulting negative effects.
[0086] Specifically, the P4 feature map and P2 feature map input from the backbone network are first upsampled and downsampled, and then fused after adding adaptive weights. The process can be expressed by formula (6):
[0087] P3″=CBR(Cat(U b (Conv 1×1 (P4))×w1,P3×w2,Conv 3×3 (P2)×w3))(6)
[0088] Among them, P2, P3, and P4 are the second, third, and fourth layer feature maps input from the backbone network respectively. 1×1 It is a convolution operation with a convolution kernel of 1, used to adjust the number of P4 channels. b It is a bilinear interpolation upsampling operation, Conv 3×3is a downsampling convolution operation with a kernel size of 3 and a stride of 2. w1, w2, and w3 are three independent adaptive weights. CBR is a point-by-point convolution operation, batch normalization, and ReLU activation function. P3″ is the output of the multi-scale information adaptive aggregation structure.
[0089] Secondly, in order to further improve the effect of feature fusion, a residual module is added. Specifically, the aggregated information is first divided into upsampling and downsampling, and then fused with the backbone network features of the corresponding layer. The process can be expressed by formulas (7)-(9):
[0090] F4=Cat(Conv 3×3 (P3″),Conv 1×1 (P4)) (7)
[0091] F3=Cat(P3″,P3)(8)
[0092] F2=Cat(U a (P3″),P2)(9)
[0093] Among them, the upsampling operation U a It is the nearest neighbor interpolation method, downsampling operation Conv 3×3 It is implemented using convolution. F4, F3, and F2 are the output feature maps of the residual module.
[0094] In the third aspect, based on the second aspect, a grouped convolution residual feature fusion module GCRF is further proposed to increase the extraction of objects of different scales in the aggregated features and ensure good traceback effect. The implementation of the GCRF module is as follows Figure 4 Specifically, the feature fusion process of this module includes: first, fusing the residual structure through convolution operation, and then further fusion and extraction of features using large kernel group convolution and ordinary convolution. The process can be expressed by formulas (10)-(11):
[0095] F=Conv 1×1 (Fin) (10)
[0096] F out =(n-1)Conv 3×3 ((n-1)GConv 5×5 (Conv 1×1 (GConv 5×5 (F))))+F(11)
[0097] Among them, F in and F out is the input and output of GCRF, the symbol + represents the residual connection, where GConv 5×5It is a grouped large-kernel convolution operation, (n-1) is the batch effect, and after entering the loop for the second time, the convolution operation with a convolution kernel of 1 is replaced by a convolution operation with a convolution kernel of 3. This is because the increase in the convolution depth requires more information for the connection between the upper and lower channels.
[0098] The SADNet proposed in this embodiment of the present invention, through its multi-scale adaptive information aggregation structure, residual structure, and GCRF module, not only retains rich detailed feature information but also integrates rich semantic information. In the task of ship target detection in SAR images of complex scenes, it enhances the ability to detect multi-scale targets and effectively improves the model's feature fusion performance.
[0099] In order to verify the effectiveness of the technical solution of the present invention, the present invention also provides the following experimental data.
[0100] The experimental environment is as follows: Ubuntu 22.04, using computing resources including an AMD EPYC 7402 CPU, an NVIDIA RTX4090 (24GB) GPU, Python 3.8, PyTorch 2.2.0, and CUDA 12.4. The image size was always set to 640×640 pixels. To ensure comparability, all compared methods were trained with a batch size of 8 and a maximum of 300 epochs. Training began with an initial learning rate of 0.001. The SGD optimizer was used during training.
[0101] In order to intuitively demonstrate the effectiveness of the improved method of the present invention, each improvement is discussed in detail through ablation experiments. The results are shown in Table 1 and analyzed as follows:
[0102] 1) VCCB: Experimental results show that after integrating the VCCB module into the baseline model, mAP0.5 increased by 1.7%, indicating that the VCCB backbone network reduced target loss by reducing the number of downsampling times. mAP0.5:0.95 increased by 1.5%, indicating that FCSI provides the model with stronger feature extraction capabilities, making the model more accurate in target recognition. The precision decreased by 0.1%, but the recall rate increased by 3%, indicating that the short-range design of the VCCB backbone network makes the model more sensitive to ship targets in SAR images in complex scenes. In particular, VCCB reduces the number of model parameters by 6.99M, which is only 37.2% of the baseline model, making it more lightweight. Therefore, using the VCCB backbone network for SAR image ship target detection in complex scenes can make the model more accurate and efficient.
[0103] 2) WPCM: Recall improved by 4% compared to the baseline model. This strongly demonstrates the effectiveness of WPCM's denoising. Precision also improved by 1.8%, demonstrating the importance of a large receptive field for ship detection in SAR images under complex conditions. The WPCM module's denoising capabilities and ability to expand the receptive field significantly improved the model's ability to discern targets, resulting in increases of 2.6% and 4.1% in mAP0.5 and mAP0.5:0.95, respectively. In particular, the WPCM module's method of expanding the receptive field through convolution in the wavelet domain does not significantly increase the number of parameters. Table 1 shows that the parameter increase does not exceed 0.005M.
[0104] 3) SADNet: This short-distance aggregation and dispersion feature fusion network aims to shorten the feature fusion path and reduce the loss of detailed features. The experimental results in Table 1 show that compared with the baseline model, replacing the SADNet neck alone reduces precision and recall by 0.2% and 0.5%, but improves mAP0.5 by 0.1%. Furthermore, based on the VCCB backbone, precision improves by 0.6%, recall decreases by 1.3%, and mAP0.5 improves by 0.2%. This result demonstrates the effectiveness of combining the SADNet neck with the VCCB backbone, addressing the negative effects of removing the fifth layer of VCCB. The results confirm that the short-distance aggregation and dispersion design of the SADNet neck is superior to the FPN structure.
[0105] Table 1 Ablation test results of each module of CSS-YOLO on HRSID dataset
[0106]
[0107] The present invention proposes a spatial information interactive feature extraction module FCSI, which aims to replace the C2f module whose information extraction is not thorough. The FCSI module can extract target features more efficiently and thoroughly, and make up for the loss of global information caused by the reduction of model depth by VCCB. The present invention conducted a comparative experiment on the VCCB backbone network, combining it with C2f and FCSI for comparison. The results detailed in Table 2 confirm that FCSI can improve the detection performance by replacing the C2f module. In particular, the detection accuracy of large targets was improved by 2.9%, confirming that FCSI is better than the C2f module in extracting global information. Although the small target accuracy and overall accuracy did not reach the highest level, the overall recall rate and average precision were improved.
[0108] Table 2 Comparative experiment on the effectiveness of FCSI in VCCB on the HRSID dataset
[0109]
[0110] In order to determine the optimal wavelet transform depth of WPCM in the task of ship target detection in complex scene SAR images, the present invention conducted a series of comparative experiments. Since the depth is proportional to the amount of computation, the present invention only conducted comparative experiments with a depth of 1 and a depth of 2. The results are shown in Table 3. The results show that when the WPCM depth is 1, all index results are better than those with a depth of 2, and the multi-scale detection results are also excellent. It was finally determined that the design with a depth of 1 not only ensures excellent detection performance, but also controls the computational complexity to the lowest range. Compared with the baseline model, the WPCM with a depth of 1 achieved significant performance improvement with a very small number of parameters.
[0111] Table 3 Comparative experiment of WPCM wavelet transform depth on HRSID dataset
[0112]
[0113] To better demonstrate the superiority of the proposed CSS-YOLO, we compared various SAR image ship target detection models on the HRSID dataset. The experimental results are shown in Table 4, where the results marked with an "*" are from this paper's experiments, while the others are referenced from recently improved models. As can be seen from Table 4, CSS-YOLO achieved a precision of 93.8% and a recall of 87.8%, surpassing all corresponding versions of YOLO and various recently proposed SAR image ship target detection methods. This resulted in a mean average precision (mAP) of 94.8%. A comparative analysis of APs, APm, and AP1 demonstrates that CSS-YOLO's multi-scale detection capabilities offer significant advantages. Compared to other detection methods, CSS-YOLO achieves the best detection results with a very small number of parameters, making it a significant advantage in SAR image ship target detection.
[0114] Table 4 Performance comparison of the proposed method and the most advanced detection method on the HRSID dataset
[0115]
[0116]
[0117] To test CSS-YOLO's performance in complex scenarios, we differentiated the datasets for different scenarios, including independent tests for complex scenes such as inland rivers, seaports, and ports. The results are shown in Table 5. Compared with the baseline model, our method improved its precision by 3.6%, its recall by 9%, and its mAP0.5 and mAP0.5:0.95 by 6.9% and 11.5% respectively. Multi-scale detection results show that the accuracy of small target detection increased by 14%, the accuracy of medium target detection increased by 1.9%, and the accuracy of large target detection increased by 12.6%. This demonstrates the powerful performance of CSS-YOLO in detecting ship targets in SAR images in complex scenarios, especially for small and large targets. The detection of small ships in complex backgrounds is a difficult problem that affects detection performance. The 14% improvement also shows that CSS-YOLO has very excellent sensitivity and detection performance for small ships in complex backgrounds.
[0118] Table 5. Comparison results of complex scenes on the HRSID dataset
[0119]
[0120] To further verify the efficiency of the proposed method in detecting HRSID dataset, we conducted a Grad-CAM visualization experiment. This visual test involves comparing the results between the baseline and CSS-YOLO methods, such as Figure 5 As shown. In order to provide a more comprehensive detection performance, a group of SAR images were selected in different scenes for comparison in our experiments. It can be clearly seen from the figure that YOLOv8s has difficulty in accurately detecting ships in the scene and cannot focus on the target, resulting in missed detections and false detections. This defect can be attributed to the complex influence of the complex background and the small size of the ship target, which has an adverse effect on the performance of the model. In contrast, the CSS-YOLO of the present invention shows excellent ability in capturing small ships in complex backgrounds close to the coast. This is mainly due to VCCB, WPCM and SADNet, which complement each other to greatly reduce the interference of the complex background and improve the model's sensitivity to small targets, so that CSS-YOLO still has excellent performance in complex scenes.
[0121] In order to verify the advantages of CSS-YOLO's detection performance in complex backgrounds, this paper makes a visual comparison of three scenes: inland river, sea outlet and port on the HRSID dataset. Figure 6 、 Figure 7 and Figure 8 shown.
[0122] In order to further verify the generalization of CSS-YOLO, the present invention conducted further comparative experiments on the SSDD dataset and the LS-SSDD-v1.0 dataset. The experimental results are shown in Table 6. The target size specifications of the two datasets are as follows Figure 9 As shown, objects on the SSDD dataset are smaller than 0.8 in size and are mostly distributed in the lower left-center region. On the LS-SSDD-v1.0 dataset, objects are smaller, smaller than 0.1 in size, and more evenly distributed. Table 6 shows that, while CSS-YOLO's multi-scale object detection results on the SSDD dataset are somewhat lower, it achieves an mAP0.5 of 98.6% with 44% of the parameters of the original model, surpassing all corresponding YOLO versions. Both precision and recall are 0.5% higher than the baseline model, YOLOv8s. Table 6 shows that on the LS-SSDD-v1.0 dataset, CSS-YOLO achieves precision of 86.1%, recall of 66.8%, mAP0.5 of 78.4%, and mAP0.5 / 0.95 of 31.3%, surpassing all corresponding YOLO versions and placing it at the top of the list in multi-scale detection capabilities. The comparative test results on the two datasets in Table 6 demonstrate that CSS-YOLO demonstrates excellent robustness.
[0123] Table 6 Performance comparison on SSDD dataset and LS-SSDD-v1.0 dataset
[0124]
[0125] To better examine CSS-YOLO's detection performance and generalization in complex scenes, we divided the SSDD and LS-SSDD-v1.0 datasets into different scenes and conducted experiments on complex scenes with background interference, such as offshore and harbor scenes. The results are shown in Table 7. Comparison results on the SSDD dataset show that CSS-YOLO achieves a 0.4% decrease in accuracy compared to the baseline model, but improves its recall by 0.4%. This demonstrates that the model is more sensitive to objects in complex scenes, resulting in more accurate detection in these scenes, improving mAP0.5 by 0.2% and mAP0.5:0.95 by 1.1%. Comparison results on the LS-SSDD-v1.0 dataset show that CSS-YOLO reduces accuracy by 0.2% compared to the baseline model, but improves recall by 4.2%, mAP0.5 by 6.4%, and mAP0.5:0.95 by 3.2%, exceeding the improvements seen on the SSDD dataset. Furthermore, the LS-SSDD-v1.0 dataset contains a high concentration of small objects, demonstrating that CSS-YOLO is highly sensitive to small objects and can significantly improve their detection performance. Furthermore, multi-scale object detection results show improvements of 3.3% and 3.9% in small object detection accuracy on both datasets, demonstrating CSS-YOLO's superior performance in detecting small objects in complex backgrounds. The comparison results on both datasets demonstrate CSS-YOLO's improved detection performance, and the complex scene comparison results on both datasets demonstrate CSS-YOLO's superior detection performance in complex scenes. CSS-YOLO improves its sensitivity to complex objects, particularly small objects. Furthermore, comparison results across multiple datasets demonstrate CSS-YOLO's excellent generalization.
[0126] Table 7 Performance comparison of complex scenes on the SSDD dataset and LS-SSDD-v1.0 dataset
[0127]
[0128] In the SSDD dataset, this paper selects four scenes of small targets for comparison, such as Figure 10 As shown in the figure, the first and second rows respectively depict scenes with a small number of targets and densely packed targets in the open sea, while the third and fourth rows depict scenes with overlapping targets and embedded targets in the near sea. The visual comparison results shown in the figure demonstrate that CSS-YOLO is superior in detecting dense targets. Furthermore, CSS-YOLO also demonstrates superior performance in the third and fourth scenes with overlapping and embedded targets.
[0129] Since the LS-SSDD-v1.0 dataset is a wide-width SAR image, the ships in it are all small and medium-sized targets, such as Figure 11 As shown in the figure, this paper selected four complex scenes: the first row shows a dense offshore scene with small objects, the second row shows an inland river scene with small objects, the third row shows an even denser inland river scene with small objects, and the fourth row shows a small object scene with a rocky background and significant noise interference. The visualization directly demonstrates the strong performance of CSS-YOLO in small object detection. In the dense scenes in the first and third rows, CSS-YOLO also performs excellent results. In the scene with strong noise interference in the fourth row, despite the clutter of noise and rocky background, CSS-YOLO accurately detects the ship in the image. The objects in the LS-SSDD-v1.0 dataset are smaller and more challenging to detect than in other datasets. In this experiment, the objects are smaller than those in the HRSID and SSDD datasets, which leads to the poor performance of various YOLO versions. In contrast, CSS-YOLO performs very well in small object detection, confirming its strong detection performance in complex scenes.
[0130] Based on the same inventive concept, Figure 12 As shown, an embodiment of the present invention provides a SAR image ship detection device for complex scenes, including a network building unit and a detection unit.
[0131] Specifically, the network building unit is used to build and train the SAR image ship detection network; the detection unit is used to detect the input SAR image using the trained SAR image ship detection network;
[0132] The SAR image ship detection network includes a backbone network and a neck network; the backbone network uses stacked feature extraction layers to extract multi-scale features of the input SAR image, and the multi-scale features include the P2 feature map, P3 feature map and P4 feature map output by the second, third and fourth layers of the backbone network; the neck network is used to fuse and detect the multi-scale features output by the backbone network to obtain the position of the ship in the SAR image; wherein the feature extraction layer includes a full-channel spatial interaction module FCSI, and the feature extraction process of the FCSI includes: in Perform convolution with a convolution kernel of 1, normalization, and SiLU activation function activation processing; then perform channel segmentation on the processed feature map to form initial feature maps f1 and f2 in the channel dimension; based on the initial feature maps f1 and f2, continuously perform m batches of channel space interaction operations to obtain m intermediate feature maps f1 (i) and , i=1,2,…,m, m is a positive integer not less than 2; wherein, setting t to be a positive integer, t∈[2,m], the channel space interaction operation of the t-th batch specifically includes: the intermediate feature map f1 obtained in the (t-1)th batch (t-1)and F1 (t-1) Input the bottleneck module for feature extraction to form feature maps f1′ and f2′; then use CBS to extract the feature maps f1′ and f2′ respectively. 1×1 After adjusting the number of channels of feature maps f1′ and f2′ to 1, the Softmax activation function is activated to obtain the spatial weights a1 and b1 of feature maps f1′ and f2′ respectively; finally, the feature map f1′ and the weight b1 are multiplied element by element to obtain the intermediate feature map f1 (t) , perform element-by-element multiplication of the feature map f2′ and the weight a1 to obtain the intermediate feature map Among them, for the channel space interaction operation of the first batch, the initial feature maps f1 and f2 are respectively used as the input of the bottleneck module; the obtained m intermediate feature maps f1 (i) and Perform splicing and fusion operations to obtain the output feature map F of the FCSI out .
[0133] It should be noted that the ship detection device provided in the embodiment of the present invention is for implementing the above method. Its specific functions can be referred to the above method embodiments and will not be described in detail here.
[0134] Figure 13 An example of a physical structure diagram of an electronic device is shown below. Figure 13 As shown, the electronic device may include: a processor (processor) 1301, a communication interface (Communications Interface) 1302, a memory (memory) 1303 and a communication bus 1304, wherein the processor 1301, the communication interface 1302, and the memory 1303 communicate with each other through the communication bus 1304. The processor 1301 can call the logic instructions in the memory 1303 to execute a SAR image ship detection method for complex scenes, the method comprising: constructing and training a SAR image ship detection network, and using the trained SAR image ship detection network to detect the input SAR image, characterized in that the SAR image ship detection network comprises a backbone network and a neck network; the backbone network uses stacked feature extraction layers to extract multi-scale features of the input SAR image, the multi-scale features comprising P2 feature maps, P3 feature maps, and P4 feature maps output by the second, third, and fourth layers of the backbone network; the neck network is used to fuse and detect the multi-scale features output by the backbone network to obtain the position of the ship in the SAR image; wherein the feature extraction layer comprises a full-channel spatial interaction module FCSI, and the feature extraction process of the FCSI comprises: performing a fusion on the input feature map F inPerform convolution with a convolution kernel of 1, normalization, and SiLU activation function activation processing; then perform channel segmentation on the processed feature map to form initial feature maps f1 and f2 in the channel dimension; based on the initial feature maps f1 and f2, continuously perform m batches of channel space interaction operations to obtain m intermediate feature maps f1 (i) and , i=1,2,…,m, m is a positive integer not less than 2; wherein, setting t to be a positive integer, t∈[2,m], the channel space interaction operation of the t-th batch specifically includes: the intermediate feature map f1 obtained in the (t-1)th batch (t-1) and F1 (t-1) Input the bottleneck module for feature extraction to form feature maps f1′ and f2′; then use CBS to extract the feature maps f1′ and f2′ respectively. 1×1 After adjusting the number of channels of feature maps f1′ and f2′ to 1, the Softmax activation function is activated to obtain the spatial weights a1 and b1 of feature maps f1′ and f2′ respectively; finally, the feature map f1′ and the weight b1 are multiplied element by element to obtain the intermediate feature map f1 (t) , perform element-by-element multiplication of the feature map f2′ and the weight a1 to obtain the intermediate feature map Among them, for the channel space interaction operation of the first batch, the initial feature maps f1 and f2 are respectively used as the input of the bottleneck module; the obtained m intermediate feature maps f1 (i) and Perform splicing and fusion operations to obtain the output feature map F of the FCSI out .
[0135] In addition, when the logic instructions in the above-mentioned memory 1303 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0136] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the SAR image ship detection method for complex scenes provided by the above-mentioned method embodiments.
[0137] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the ship detection method for SAR images of complex scenes provided by the above-mentioned method embodiments is implemented.
[0138] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A SAR image ship detection method for complex scenes, characterized by: include: A SAR image ship detection network is constructed and trained, and the trained SAR image ship detection network is used to detect the input SAR image. The SAR image ship detection network comprises a backbone network and a neck network; the backbone network extracts multi-scale features of the input SAR image using stacked feature extraction layers, and the multi-scale features include P2 feature maps, P3 feature maps, and P4 feature maps output by the second, third, and fourth layers of the backbone network; the neck network is used to fuse and detect the multi-scale features output by the backbone network to obtain the position of the ship in the SAR image; wherein the feature extraction layer includes a full-channel spatial interaction module FCSI, and the feature extraction process of the FCSI includes: For the input feature map F in Perform convolution with a convolution kernel of 1, normalization, and SiLU activation function activation processing; then perform channel segmentation on the processed feature map to form the initial feature maps f1 and f2 in the channel dimension; Based on the initial feature maps f1 and f2, m batches of channel space interaction operations are continuously performed to obtain m intermediate feature maps f1 (i) and i=1,2,…,m, m is a positive integer not less than 2; wherein, t is set to be a positive integer, t∈[2,m], then the channel space interaction operation of the t-th batch specifically includes: the intermediate feature map f1 obtained in the (t-1)th batch (t-1) and F1 (t-1) Input the bottleneck module for feature extraction to form feature maps f′1 and f′2; then use CBS to 1×1 After adjusting the number of channels of feature maps f′1 and f′2 to 1, the Softmax activation function is activated to obtain the spatial weights a1 and b1 of feature maps f′1 and f′2 respectively; finally, the feature map f′1 and the weight b1 are multiplied element by element to obtain the intermediate feature map f1 (t) , perform element-by-element multiplication of the feature map f′2 and the weight a1 to obtain the intermediate feature map Among them, for the channel space interaction operation of the first batch, the initial feature maps f1 and f2 are used as the input of the bottleneck module respectively; The obtained m intermediate feature maps f1 (i) and Perform splicing and fusion operations to obtain the output feature map F of the FCSI out .
2. The method for ship detection in SAR images for complex scenes according to claim 1, characterized in that: The backbone network also includes a wavelet preprocessing downsampling module WPCM located before the first feature extraction layer; The WPCM is used to expand the receptive field of the input feature map and reduce noise interference; The corresponding feature map processing process includes: first using two branches to process the input feature map respectively, and then fusing the outputs of the two branches to obtain the output feature map of the WPCM: wherein, one branch includes: performing convolution processing on the input feature map to obtain the corresponding convolution result; the other branch includes: first performing wavelet transform on the input feature map to obtain a low-frequency component and high-frequency components in three directions of horizontal, vertical and diagonal directions; then performing convolution processing on the four components respectively to obtain four corresponding convolution results; finally, applying inverse wavelet transform to the four convolution results to obtain the reconstructed feature map.
3. A SAR image ship detection method for complex scenes according to claim 1 or 2, characterized in that: The neck network includes a multi-scale information adaptive aggregation module, and the feature fusion process of the multi-scale information adaptive aggregation module includes: P3″=CBR(Cat(U b (Conv 1×1 (P4))×w1,P3×w2,Conv 3×3 (P2)×w3)) Among them, P2, P3 and P4 are the P2 feature map, P3 feature map and P4 feature map output by the backbone network respectively; Conv 1×1 It is a convolution operation with a convolution kernel of 1; U b is an upsampling operation; Conv 3×3 is a convolution operation with a kernel size of 3 and a stride of 2; w1, w2, and w3 are three independent adaptive weights; CBR is a point-by-point convolution operation, batch normalization operation, and ReLU activation function; Cat is a splicing operation; and P3″ is the output of the multi-scale information adaptive aggregation module.
4. The method for ship detection in SAR images for complex scenes according to claim 3, characterized in that: The neck network also includes a residual module, and the feature fusion process of the residual module includes: F4=Cat(Conv 3×3 (P3″),Conv 1×1 (P4)) F3=Cat(P3″,P3) F2=Cat(U a (P3″),P2) Among them, F2, F3 and F4 are the output feature maps of the residual module.
5. The method for ship detection in SAR images for complex scenes according to claim 4, characterized in that: The neck network also includes a grouped convolution residual feature fusion module GCRF, and the feature fusion process of the GCRF includes: F=Conv 1×1 (F in ) F out =(n-1)Conv 3×3 ((n-1)GConv 5×5 (Conv 1×1 (GConv 5×5 (F))))+F Among them, F in and F out It is the input feature map and output feature map of GCRF, + represents the residual connection, GConv 5×5 It is a grouped large kernel convolution operation, where n is a positive integer not less than 2.
6. A SAR image ship detection device for complex scenes, characterized by: include: Network building unit, used to build and train SAR image ship detection network; A detection unit, configured to detect an input SAR image using a trained SAR image ship detection network; The SAR image ship detection network includes a backbone network and a neck network; the backbone network uses stacked feature extraction layers to extract multi-scale features of the input SAR image, and the multi-scale features include the P2 feature map, P3 feature map, and P4 feature map output by the second, third, and fourth layers of the backbone network; the neck network is used to fuse and detect the multi-scale features output by the backbone network to obtain the position of the ship in the SAR image; the feature extraction layer includes a full-channel spatial interaction module FCSI, and the feature extraction process of the FCSI includes: For the input feature map F in Perform convolution with a convolution kernel of 1, normalization, and SiLU activation function activation processing; then perform channel segmentation on the processed feature map to form the initial feature maps f1 and f2 in the channel dimension; Based on the initial feature maps f1 and f2, m batches of channel space interaction operations are continuously performed to obtain m intermediate feature maps f1 (i) and i=1,2,…,m, m is a positive integer not less than 2; wherein, t is set to be a positive integer, t∈[2,m], then the channel space interaction operation of the t-th batch specifically includes: the intermediate feature map f1 obtained in the (t-1)th batch (t-1) and F1 (t-1) Input the bottleneck module for feature extraction to form feature maps f′1 and f′2; then use CBS to 1×1 After adjusting the number of channels of feature maps f′1 and f′2 to 1, the Softmax activation function is activated to obtain the spatial weights a1 and b1 of feature maps f′1 and f′2 respectively; finally, the feature map f′1 and the weight b1 are multiplied element by element to obtain the intermediate feature map f1 (t) , perform element-by-element multiplication of the feature map f′2 and the weight a1 to obtain the intermediate feature map Among them, for the channel space interaction operation of the first batch, the initial feature maps f1 and f2 are used as the input of the bottleneck module respectively; The obtained m intermediate feature maps f1 (i) and Perform splicing and fusion operations to obtain the output feature map F of the FCSI out .
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.