SAR Ship Detection Method Based on Multi-Scale Feature Extraction and FRFT Convolution
By using multi-scale feature extraction and fractional Fourier transform convolution method in SAR ship detection, a detection network model is constructed, which solves the problem of unsatisfactory detection effect in complex backgrounds, and achieves detection effects with high accuracy and low error detection rate.
Patent Information
- Application Number
- CN202411375836.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-09-30
AI Technical Summary
The existing SAR ship detection technology has poor detection results in complex backgrounds, high error detection rate, and inaccurate detection.
Using a detection method based on multi-scale feature extraction and fractional Fourier transform convolution, a detection network model including backbone network, neck network and probe head part is constructed, and multi-scale features are extracted using multi-level residual module and aggregation multi-scale channel attention module, and detection efficiency is enhanced through FRFT convolution.
It significantly improves the accuracy and accuracy of SAR ship detection, reduces the false detection rate, and enhances the model's adaptability to complex backgrounds.
Smart Images

Figure CN118884439B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a SAR ship detection method based on multi-scale feature extraction and FRFT convolution, and belongs to the technical field of target detection. Background Art
[0002] Synthetic Aperture Radar (SAR) is a high-resolution microwave imaging radar system with the ability to work all-weather and all-day, and is widely used in fields such as homeland security, environmental monitoring, and disaster assessment. SAR image target detection refers to automatically identifying and locating targets of interest in SAR images, such as ships, vehicles, or specific building structures.
[0003] The key issues in maritime ship target detection are various complex backgrounds, uncertain factors such as various irregular arrangements of ships and misdetection of similar targets, and the confusion caused by the inherent high gray level characteristics of the SAR image background. With the wide application of deep learning in the field of target detection, Du Yanling et al. improved the fully convolutional neural network and used different hierarchical convolutional layer fusion strategies to improve the SAR detection accuracy. Rostami et al. proposed to enable electro-optical domain images to be used for training SAR ship detection models by learning a shared invariant cross-domain embedding space. Raj et al. proposed a deep learning classification model with an additional preprocessing input stage to reduce misclassification. However, although the existing technologies have achieved certain effects in SAR detection, they have strong limitations for backgrounds such as complex backgrounds, and the detection effect is not ideal, with a high misdetection rate and inaccurate detection. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a SAR ship detection method based on multi-scale feature extraction and FRFT convolution, which can better obtain ship features, reduce the misdetection rate of detection, improve the detection accuracy, and the adaptation effect in various scenarios.
[0005] The present invention adopts the following technical solutions to solve the above technical problems:
[0006] A SAR ship detection method based on multi-scale feature extraction and FRFT convolution includes the following steps:
[0007] Step 1, obtain a SAR ship image dataset, and after preprocessing the dataset, divide it into a training set, a validation set, and a test set;
[0008] Step 2, construct a SAR ship detection network model based on multi-scale feature extraction and fractional Fourier transform convolution, and the SAR ship detection network model includes an input end, a backbone network, a neck network, and a detection head part;
[0009] In step 2, the backbone network includes a first CBS module, a second CBS module, a third CBS module, a fourth CBS module, a first multi-level residual module, a first downsampling module, a second multi-level residual module, a second downsampling module, a third multi-level residual module, a third downsampling module, and a fourth multi-level residual module connected in sequence; among them,
[0010] The structures of the first to fourth multi-level residual modules are the same, and they are all stepped residual structures composed of a conventional convolution, a partial convolution, a depthwise separable convolution, a batch normalization, and a ReLU activation function. The convolution kernel size of the conventional convolution is 1, and the convolution kernel size of the depthwise separable convolution is 3; the internal calculations of each multi-level residual module are as follows:
[0011] ,
[0012] where PConv represents the partial convolution, conv represents the conventional convolution, DWConv represents the depthwise separable convolution, BN represents the batch normalization, ReLU represents the ReLU activation function, Concat represents the concatenation operation, x represents the input of each multi-level residual module, and both represent the output of the node, and both represent the output of the node; After passing through a CBS module, the output of each multi-level residual module is obtained. The convolution kernel size of the convolution layer of the CBS module is 1;
[0013] Step 3: Use the training set to train the SAR ship detection network model constructed in step 2, calculate the loss function and perform backpropagation to update the network parameters, obtain the network model with the best parameters, that is, the trained SAR ship detection network model, and use the validation set for verification;
[0014] Step 4: Input the test set into the trained SAR ship detection network model and output the ship detection result.
[0015] As a preferred solution of the present invention, the specific process of step 1 is as follows:
[0016] Obtain the SAR ship image datasets SSDD and HRSID, and the corresponding ship position labels, unify the image sizes in the datasets to size, and divide the dataset with unified size into a training set, a validation set, and a test set according to the ratio.
[0017] As a preferred embodiment of the present invention, the first to fourth CBS modules are each composed of a convolutional layer, a batch normalization layer, and a SiLU activation function; the convolutional kernels of the convolutional layers in the first and third CBS modules have a size of 3 and a stride of 1; the convolutional kernels of the convolutional layers in the second and fourth CBS modules have a size of 3 and a stride of 2;
[0018] The first to third downsampling modules have the same structure and each includes a max pooling layer and fifth to seventh CBS modules; the input of each downsampling module passes through the max pooling layer and the fifth CBS module in sequence to obtain the output of the fifth CBS module, and the input of each downsampling module passes through the sixth CBS module and the seventh CBS module in sequence to obtain the output of the seventh CBS module. A connection operation is performed on the output of the fifth CBS module and the output of the seventh CBS module to obtain the output of each downsampling module; the convolutional kernels of the convolutional layers in the fifth and sixth CBS modules have a size of 1, and the convolutional kernel of the convolutional layer in the seventh CBS module has a size of 3.
[0019] As a preferred embodiment of the present invention, in step 2, the neck network includes an SPPCSPC module, first to fourth ACAM modules, eighth to eleventh CBS modules, and fourth to fifth downsampling modules; the output of the fourth multi-stage residual module passes through the SPPCSPC module and the eighth CBS module in sequence to obtain the output of the eighth CBS module; the output of the third multi-stage residual module passes through the ninth CBS module to obtain the output of the ninth CBS module; the output of the eighth CBS module is upsampled and fused with the output of the ninth CBS module to obtain a first fusion result; the first fusion result passes through the first ACAM module and the tenth CBS module in sequence to obtain the output of the tenth CBS module; the output of the second multi-stage residual module passes through the eleventh CBS module to obtain the output of the eleventh CBS module; the output of the tenth CBS module is upsampled and fused with the output of the eleventh CBS module to obtain a second fusion result; the second fusion result passes through the second ACAM module and the fourth downsampling module in sequence and is fused with the output of the first ACAM module to obtain a third fusion result; the third fusion result passes through the third ACAM module and the fifth downsampling module in sequence and is fused with the output of the SPPCSPC module to obtain a fourth fusion result; the output of the second ACAM module is used as the first output of the neck network, the output of the third ACAM module is used as the second output of the neck network, and the fourth fusion result passes through the fourth ACAM module and is used as the third output of the neck network.
[0020] As a preferred embodiment of the present invention, the structures of the first to fourth ACAM modules are the same, and each includes a channel multi-scale attention module and the twelfth to eighteenth CBS modules. The convolution kernel sizes of the convolution layers in the twelfth, seventeenth, and eighteenth CBS modules are all 1, and the convolution kernel sizes of the convolution layers in the thirteenth to sixteenth CBS modules are all 3. The input of each ACAM module passes through the twelfth CBS module, the thirteenth CBS module, the fourteenth CBS module, the fifteenth CBS module, the sixteenth CBS module, and the channel multi-scale attention module in sequence to obtain the output of the channel multi-scale attention module. The input of each ACAM module passes through the seventeenth CBS module to obtain the output of the seventeenth CBS module. The output of the seventeenth CBS module is fused with the output of the channel multi-scale attention module, the output of the thirteenth CBS module, the output of the fourteenth CBS module, and the output of the sixteenth CBS module. After the fusion result passes through the eighteenth CBS module, the output of each ACAM module is obtained.
[0021] The channel multi-scale attention module is divided into two stages. In the first stage, the input of the channel multi-scale attention module passes through adaptive max pooling and adaptive average pooling respectively. The outputs of the adaptive max pooling and adaptive average pooling enter a multi-layer perceptron. After fusing the output of the multi-layer perceptron and passing through the Sigmoid activation function, the output of the Sigmoid activation function is multiplied by the input of the channel multi-scale attention module to obtain the output of the first stage.
[0022] The output of the first stage passes through a first depth convolution. The output of the first depth convolution passes through the second and third depth convolutions in sequence to obtain the output of the third depth convolution. The output of the first depth convolution passes through the fourth and fifth depth convolutions in sequence to obtain the output of the fifth depth convolution. The output of the first depth convolution passes through the sixth and seventh depth convolutions in sequence to obtain the output of the seventh depth convolution. After fusing the output of the third depth convolution with the output of the fifth depth convolution and the output of the seventh depth convolution, and then passing through a convolution layer, the output of the convolution layer is multiplied by the output of the first stage to obtain the output of the second stage, that is, the output of the channel multi-scale attention module.
[0023] As a preferred embodiment of the present invention, the SPPCSPC module includes the nineteenth to twenty-fifth CBS modules and the first to third max pooling layers. The convolution kernel sizes of the convolution layers in the nineteenth, twenty-first, twenty-second, twenty-fourth, and twenty-fifth CBS modules are all 1, the convolution kernel sizes of the convolution layers in the twentieth and twenty-third CBS modules are all 3, and the convolution kernel sizes of the first to third max pooling layers are 5, 9, and 12 respectively.
[0024] The input of the SPPCSPC module sequentially passes through the nineteenth CBS module, the twentieth CBS module, and the twenty-first CBS module to obtain the output of the twenty-first CBS module; the output of the twenty-first CBS module passes through the first max pooling layer, the second max pooling layer, and the third max pooling layer respectively. After the output of the twenty-first CBS module is fused with the outputs of the first max pooling layer, the second max pooling layer, and the third max pooling layer, it sequentially passes through the twenty-second CBS module and the twenty-third CBS module to obtain the output of the twenty-third CBS module; after the input of the SPPCSPC module passes through the twenty-fourth CBS module, it is fused with the output of the twenty-third CBS module, and the fusion result passes through the twenty-fifth CBS module to obtain the output of the SPPCSPC module.
[0025] As a preferred embodiment of the present invention, the detection head part includes the first to third fractional Fourier transform convolution modules, the first to third RepConv modules, and the first to third detection heads; the first output of the neck network sequentially passes through the first fractional Fourier transform convolution module, the first RepConv module, and the first detection head for large-scale prediction; the second output of the neck network sequentially passes through the second fractional Fourier transform convolution module, the second RepConv module, and the second detection head for medium-scale prediction; the third output of the neck network sequentially passes through the third fractional Fourier transform convolution module, the third RepConv module, and the third detection head for small-scale prediction.
[0026] As a preferred embodiment of the present invention, the structures of the first to third fractional Fourier transform convolution modules are the same. The input of each fractional Fourier transform convolution module first passes through a convolution layer with a kernel size of 1 and is then evenly divided into four parts, each part is passed to different Fn modules, n = 1, 2, 3, 4. The four Fn modules are processed at different frequencies to obtain feature maps of different frequencies. The feature maps of different frequencies obtained by the four Fn modules are fused to obtain a fused feature map; the input of each fractional Fourier transform convolution module passes through a convolution layer with a kernel size of 1 again, and is fused with the fused feature map and the input of the fractional Fourier transform convolution module to obtain the output of each fractional Fourier transform convolution module;
[0027] In the Fn module, the learning tensor t is updated using the gradient descent algorithm. After passing through the Gaussian regularization operation, t is used as the input of the fractional Fourier filter, and the output of the fractional Fourier filter is as follows:
[0028] ,
[0029] where, represents the filter component, respectively represent the filter components in the X and Y directions,f is the frequency parameter, p is the fractional order, x , y are the elements in the 2D grid coordinate matrices in the X and Y directions respectively, generated by decomposing the input to the filter.
[0030] A computer device includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the steps of the SAR ship detection method based on multi-scale feature extraction and FRFT convolution are implemented.
[0031] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0032] 1. The present invention proposes a SAR ship detection method based on multi-scale feature extraction and fractional Fourier transform convolution network. The idea of gradual residual and different convolutions are used to construct a more effective and lightweight MLRM module as the feature extraction module of the backbone network, and a new structure ACAM module jointly constructed by an aggregation structure and an attention mechanism is used as the feature extraction module of the Neck network, enhancing the model's ability to capture input-end features. Finally, FRFT convolution is introduced to make the detection efficiency stronger.
[0033] 2. The network constructed by the present invention has significant advantages in improving the detection accuracy of synthetic aperture radar ships, reducing false detections and missed detections, effectively addressing the problems existing in synthetic aperture radar ship detection, and providing a high-precision detection method. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is the flowchart of the SAR ship detection method based on multi-scale feature extraction and FRFT convolution of the present invention;
[0035] Figure 2 is the overall network structure diagram of the MEFTNet constructed by the present invention;
[0036] Figure 3 is the structure diagram of the multi-level residual module (MLRM) designed by the present invention;
[0037] Figure 4 is the structure diagram of the aggregation multi-scale channel attention module (ACAM) designed by the present invention;
[0038] Figure 5 is the structure diagram of the channel multi-scale attention designed by the present invention;
[0039] Figure 6 is the structure diagram of the fractional Fourier transform (FRFT) convolution designed by the present invention;
[0040] Figure 7 This is a comparison experiment result graph of the detection method of the present invention and other mainstream methods. Among them, (a) is the true value, (b) is SSD detection, (c) is Faster-Rcnn detection, (d) is YOLOv5 detection, (e) is YOLOv8 detection, and (f) is the detection of the present invention. Specific implementation manners
[0041] The following details the implementation manners of the present invention, and the examples of the implementation manners are shown in the drawings. The implementation manners described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0042] Aiming at the problems of high false detection rate and inaccurate detection in complex backgrounds, the present invention proposes a SAR ship detection method based on multi-scale feature extraction and fractional Fourier transform convolutional network, as Figure 1 shown, which includes the following steps:
[0043] Step 1: Obtain the SSDD synthetic aperture radar image dataset and the HRSID dataset, and after preprocessing the dataset, divide it into a training set, a validation set, and a test set according to a set ratio. Specifically as follows:
[0044] Obtain the SSDD dataset and the HRSID dataset, as well as the corresponding ship position labels, in the YOLO format;
[0045] Unify the sizes of the images contained in the dataset to , and randomly divide the dataset of size into three parts, where the training set accounts for 70%, the validation set accounts for 20%, and the test set accounts for 10%.
[0046] Step 2: Build a multi-scale feature extraction and FRFT convolutional network (Multi scale feature extraction and fractional Fourier transform convolutional network, MEFTNet), as Figure 2 shown, including an input end, a backbone network, a Neck network, and a detection head part, where the backbone network and the Neck network cooperate to process features.
[0047] Step 2.1: Process the SAR image as the input image through the backbone network: The modules contained in the backbone network mainly consist of the initial 4 CBS modules, 4 efficient and lightweight multi-level residual modules (Multi level residual module, MLRM), and 3 downsampling (MP) modules. The CBS module is mainly composed of a convolutional layer, normalization, and a silu activation function, and its function is to initially extract image features and transform the image size. Among them, the convolutional kernel sizes of the first CBS and the third CBS modules are both 3, and the stride is 1. The convolutional kernel sizes of the second CBS and the fourth CBS modules are 3, and the stride is 2. At this time, the feature map size becomes , and then by the multi-level residual module, that is, the MLRM module: The MLRM module is a stepped residual structure composed of depthwise separable convolution, partial convolution, conventional convolution, Concat, normalization, and ReLU activation function. Among them, the convolutional kernel size of the conventional convolution is one, and the convolutional kernel size of the depthwise separable convolution is 3. The MP module consists of a max pooling layer and a parallel structure composed of two CBS modules with a convolutional kernel size of 1 and a CBS module with a convolutional kernel size of 3 for Concat operation. By dividing the input feature map, the maximum value is selected as the output for each divided region. In the backbone network, the original image undergoes the above operations, enabling each MLRM module in the backbone network to obtain multi-scale features as the input features for the Neck part.
[0048] The MLRM module is an efficient and lightweight multi-level residual module, a multi-level residual structure composed of multiple partial convolutions, depthwise separable convolutions, conventional convolutions, BN normalization, and ReLU activation functions. Different from the traditional feature extraction method that uses one backbone plus several residual structures, the present invention proposes to use a structure of gradually fusing residuals to achieve multi-scale processing of features by the backbone network. In the gradually residual structure proposed in the present invention, multi-level transmission is performed on each internal layer of features, and multiple residual structures are used for stepped hierarchical stacking to fully extract and fuse the feature outputs of each layer, so that the final output features have diversity. In order to avoid an increase in the number of parameters caused by multi-level stacking, partial convolution and depthwise separable convolution are used internally. Each layer uses a large-kernel depthwise separable convolution to facilitate full feature extraction. The partial convolution processes local features and the conventional convolution processes global features, and the local features and global features are integrated. Then, through the cascade structure and the residual structure, feature loss is avoided, and each level of features is stacked multiple times until final fusion, so as to fully extract features and fuse shallow features and deep features.
[0049] As Figure 3 shown, the internal calculation of MLRM is as follows:
[0050] ,
[0051] Step 2.2: At the end of the backbone network in Step 2.1, that is, the output feature map of the fourth MLRM module is used as the input of SPPCSPC in the Neck part: SPPCSPC (Spatial Pyramid Pooling and Cross-Stage Partial Network Module) is composed of 5 CBSs with a convolution kernel size of 1, two CBSs with a convolution kernel size of 3, a max pooling layer with a convolution kernel size of 5, a max pooling layer with a convolution kernel size of 9, and a max pooling layer with a convolution kernel size of 12. By using the pooling pyramid, feature maps of different scales can be obtained, which helps to detect objects at different scales and extract more expressive features. The Neck part is mainly composed of 1 SPPCSPC, 2 MP modules, 2 upsamplings, 4 Concat operations, 4 CBS modules, and 4 ACAM modules. The output feature maps of the last 3 layers of the MLRM in the backbone network are adjusted in terms of the corresponding number of channels through several CBS modules and upsamplings, and their sizes are gradually input into the Neck part for multi-scale feature fusion to reduce the loss during feature transmission. The upsampling operation is the UP sample, as Figure 4 shown. The ACAM module is composed of several CBS modules and a channel multi-scale attention module.
[0052] As Figure 5 shown, the channel multi-scale attention module consists of two computational stages. In the first stage, operations such as adaptive average pooling and adaptive max pooling are used to learn the importance of the input feature map in the channel dimension and adjust the weights of different channels, enabling the model to better capture the relationships and feature importance between different channels, thereby enhancing the model's feature expression ability.
[0053] Among them, adaptive average pooling and adaptive max pooling can simultaneously consider and integrate the local and global important features of the feature map as the output of a new feature map.
[0054] MLP refers to Multilayer Perceptron: The role of MLP in CBAM is to perform feature mapping for each channel. Through MLP, ACAM can model and learn the correlations between different channels, thereby extracting more useful features. MLP can be implemented by stacking multiple fully connected layers, and each fully connected layer uses a non-linear activation function.
[0055] In the second stage, multi-scale convolution operations are utilized to learn the importance of the input feature map in the spatial dimension. It extracts features of different scales through different convolutional layers and combines these features to make full use of the information at different scales in the image. Due to problems such as complex textures and spurious point backgrounds often existing in SAR images, multi-scale convolution can better capture these details. Without significantly increasing the model parameters and computational complexity, ACAM can improve the model's perception ability and expressive ability.
[0056] Among them, the multi-scale convolutions are respectively composed of convolutional kernels of sizes, covering the important features in all aspects of the first stage, which are integrated through conventional convolution and then output.
[0057] Step 2.3: After step 2.2, the output feature maps at the three output ends in the Neck part respectively correspond to the inputs of three detection heads of three different sizes. The detection head ends are respectively composed of fractional Fourier transform convolutional (FRFT) convolution, RepConv, and Head. When the original features pass through the backbone network and the Neck part and are processed, the output feature maps will be input to the detection head ends for prediction. RepConv is mainly composed of convolution and normalization. The Head part outputs feature maps with sizes of , , .
[0058] As Figure 6 shown, the fractional Fourier transform convolution uses a multi-branch structure and a residual structure to capture more feature information. In the main branch, the input feature map first passes through a convolutional layer with a kernel size of 1 and is evenly divided into four parts. Each part is passed to a different Fn (n = 1, 2, 3, 4) module for processing at different frequencies. Through this method, low-order (close to the time domain) and high-order (close to the frequency domain) features can be obtained simultaneously, and then the input features can be made to transition more smoothly from the time domain to the frequency domain, providing a more detailed perception of feature changes for the network model and enabling the network model to learn more abundant useful features. In the Fn module, the present invention proposes to update the learnable tensor t using the gradient descent algorithm. After Gaussian regularization operation, t is used as the input of the fractional Fourier filter, where the output weight k1 of the fractional Fourier filter has the same shape as the convolutional kernel. Convolution operation is performed on k1 to obtain the output feature, and the activation function is used to improve the expressive ability of the output feature.
[0059] In the FRFT filter, the input features are first decomposed to generate a 2D grid coordinate matrix in the X and Y directions. The filter components in the X and Y directions are calculated using the FRFT formula, and finally multiplied to obtain the final filter output. For each output channel, the corresponding filter components are calculated:
[0060] Among them, the calculation formula for the filter component in the X direction is as follows:
[0061] ,
[0062] The calculation formula for the filter component in the Y direction is as follows:
[0063] ,
[0064] The calculation formula for the final filter component is as follows:
[0065] ,
[0066] Among them, f is the frequency parameter, p is the fractional order.
[0067] Step 2.4: After steps 2.1, 2.2, and 2.3 are deployed, the original image is preprocessed and input into the backbone network. After passing through the first four CBS modules, the first MLRM is called the first stage, the second MLRM is called the second stage, the third MLRM is called the third stage, and the fourth MLRM is called the fourth stage. The output of the fourth stage is used as the input to the SPPCSPC of the Neck network. After passing through the pooling pyramid SPPCSPC, its output is transformed in terms of image size and feature extraction through the CBS module and upsampling operation. The output of the third stage passes through an output feature with a kernel size of 1 for fusion, which is called the first fusion. After passing through the first ACAM module in the Neck part, after passing through the CBS with a kernel size of 1 and upsampling operation, it is fused with the output of the CBS module with a kernel size of 1 from the output of the second stage, which is called the second fusion. The output passes through the second ACAM module and then undergoes large-scale prediction through FRFT convolution and RepConv. At the same time, the output of the second ACAM passes through the MP module and is fused with the output of the first ACAM. The output after passing through the third ACAM module and the MP module is fused with the SPPCSPC. The output after passing through the fourth ACAM module undergoes small-scale prediction through RepConv and FRFT convolution. The output of the third ACAM enters RepConv for medium-scale prediction. After going through all the above processes, the accuracy and loss are calculated for each round of training.
[0068] The experimental indicators are accuracy ( Precision ), recall rate ( Recall ), mAP 0.5 (average precision at IOU 0.5) represents the average accuracy when the IOU detection threshold is 0.5, and mAP 0.5:0.95 (average precision at IOU 0.95) represents the average accuracy when the IOU detection threshold ranges from 0.5 to 0.95. The formula is as follows:
[0069] ,
[0070] where represents the true positive samples, represents the false positive samples, represents the false negative samples that fail to be detected. IOU (Intersection over Union) is one of the commonly used metrics to evaluate the performance in object detection or image segmentation tasks.
[0071] In object detection tasks, IOU is used to measure the overlap degree between the predicted bounding box (or called detection box) and the true bounding box. It is measured by calculating the ratio of the intersection area of the two bounding boxes to their union area.
[0072] Step 3: Input the preprocessed SAR images of the training set and validation set in Step 1 into the MEFTNet network in Step 2 for training, calculate the loss function and perform backpropagation to update the network parameters and obtain the optimal parameter model.
[0073] Step 4: Input the preprocessed test set in Step 1 into the optimal parameter model trained in Step 3 to output the accurate recognition map of the SAR image.
[0074] To prove the effectiveness of the SAR ship detection method with multi-scale feature extraction and fractional Fourier transform convolution provided by the present invention, the SSDD dataset and the HRSID dataset are used to train, validate and test the model. From Figure 7 (a)-(f) and Table 1, it can be seen that all the evaluation indicators of the present invention are higher than those of the existing detection networks, and the detection effect is closest to the true value.
[0075] Table 1
[0076]
[0077] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the foregoing SAR ship detection method based on multi-scale feature extraction and FRFT convolution are implemented.
[0078] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0079] The present invention is described with reference to the flowcharts of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process in the flowchart and the combination of processes in the flowchart can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 process or multiple processes.
[0080] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 process or multiple processes.
[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or multiple processes.
[0082] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the present invention.
Claims
1. A SAR ship detection method based on multi-scale feature extraction and FRFT convolution is characterized by: The steps include: Step 1, obtain a SAR ship image dataset, pre-process the dataset, and divide it into a training set, a validation set, and a test set; Step 2, constructing a SAR ship detection network model based on multi-scale feature extraction and fractional Fourier transform convolution, wherein the SAR ship detection network model includes an input end, a backbone network, a neck network, and a detection head part; In step 2, the backbone network includes a first CBS module, a second CBS module, a third CBS module, a fourth CBS module, a first multi-level residual module, a first down-sampling module, a second multi-level residual module, a second down-sampling module, a third multi-level residual module, a third down-sampling module and a fourth multi-level residual module connected in sequence; wherein, The structures of the first to fourth multi-level residual modules are the same, which are all stepped residual structures composed of conventional convolution, partial convolution, depthwise separable convolution, batch normalization and ReLU activation function. The convolution kernel size of conventional convolution is 1, and the convolution kernel size of depthwise separable convolution is 3. The internal calculation of each multi-level residual module is as follows: , Among them, PConv means partial convolution, conv means regular convolution, DWConv means depth-separable convolution, BN Represents batch normalization, ReLU represents the ReLU activation function, and Concat represents the concatenation operation. x represents the input of each multi-level residual module, and They all represent the output of the node. and Both represent the output of the node; After a CBS module, the output of each multi-level residual module is obtained, and the convolution kernel size of the convolution layer of the CBS module is 1; The neck network includes the SPPCSPC module, the first to fourth ACAM modules, the eighth to eleventh CBS modules, and the fourth to fifth downsampling modules; the first to fourth ACAM modules have the same structure, including a channel multi-scale attention module and the twelfth to eighteenth CBS modules, the convolution kernel size of the convolution layer in the twelfth, seventeenth and eighteenth CBS modules is 1, and the convolution kernel size of the convolution layer in the thirteenth to sixteenth CBS modules is 3; Step 3, use the training set to train the SAR ship detection network model constructed in step 2, calculate the loss function and perform back propagation, update the network parameters, obtain the optimal parameter network model, that is, the trained SAR ship detection network model, and use the verification set for verification; Step 4: Input the test set into the trained SAR ship detection network model and output the ship detection results.
2. The SAR ship detection method based on multi-scale feature extraction and FRFT convolution according to claim 1 is characterized in that: The specific process of step 1 is as follows: The SAR ship image datasets SSDD and HRSID, as well as the corresponding ship position labels, are obtained. The image sizes in the datasets are unified to 640×640, and the unified datasets are divided into training set, validation set, and test set in a ratio of 7:2:
1.
3. The SAR ship detection method based on multi-scale feature extraction and FRFT convolution according to claim 1 is characterized in that: The first to fourth CBS modules are composed of convolutional layers, batch normalization layers and SiLU activation functions; the convolution kernel size of the convolutional layers in the first and third CBS modules is 3 and the step size is 1; the convolution kernel size of the convolutional layers in the second and fourth CBS modules is 3 and the step size is 2; The structures of the first to third downsampling modules are the same, and all include a maximum pooling layer and fifth to seventh CBS modules; the input of each downsampling module passes through the maximum pooling layer and the fifth CBS module in sequence to obtain the output of the fifth CBS module, and the input of each downsampling module passes through the sixth CBS module and the seventh CBS module in sequence to obtain the output of the seventh CBS module, and the output of the fifth CBS module and the output of the seventh CBS module are connected to obtain the output of each downsampling module; the convolution kernel size of the convolution layer in the fifth and sixth CBS modules is 1, and the convolution kernel size of the convolution layer in the seventh CBS module is 3.
4. The SAR ship detection method based on multi-scale feature extraction and FRFT convolution according to claim 1 is characterized in that: In the neck network described in step 2, the output of the fourth multi-level residual module is sequentially passed through the SPPCSPC module and the eighth CBS module to obtain the output of the eighth CBS module; the output of the third multi-level residual module is passed through the ninth CBS module to obtain the output of the ninth CBS module; The output of the eighth CBS module is fused with the output of the ninth CBS module after upsampling operation to obtain a first fusion result; the first fusion result is sequentially passed through the first ACAM module and the tenth CBS module to obtain the output of the tenth CBS module; The output of the second multi-level residual module is passed through the eleventh CBS module to obtain the output of the eleventh CBS module; The output of the tenth CBS module is fused with the output of the eleventh CBS module after an upsampling operation to obtain a second fusion result; The second fusion result is sequentially passed through the second ACAM module and the fourth down-sampling module, and then fused with the output of the first ACAM module to obtain a third fusion result; The third fusion result is sequentially passed through the third ACAM module and the fifth down-sampling module and fused with the output of the SPPCSPC module to obtain a fourth fusion result; The output of the second ACAM module is used as the first output of the neck network, the output of the third ACAM module is used as the second output of the neck network, and the fourth fusion result is used as the third output of the neck network after passing through the fourth ACAM module.
5. The SAR ship detection method based on multi-scale feature extraction and FRFT convolution according to claim 4 is characterized in that: The input of each ACAM module passes through the twelfth CBS module, the thirteenth CBS module, the fourteenth CBS module, the fifteenth CBS module, the sixteenth CBS module and the channel multi-scale attention module in sequence to obtain the output of the channel multi-scale attention module; the input of each ACAM module passes through the seventeenth CBS module to obtain the output of the seventeenth CBS module; the output of the seventeenth CBS module is fused with the output of the channel multi-scale attention module, the output of the thirteenth CBS module, the output of the fourteenth CBS module and the output of the sixteenth CBS module, and the fusion result passes through the eighteenth CBS module to obtain the output of each ACAM module; The channel multi-scale attention module is divided into two stages. In the first stage, the input of the channel multi-scale attention module is respectively subjected to adaptive maximum pooling and adaptive average pooling, and the outputs of the adaptive maximum pooling and adaptive average pooling enter the multi-layer perceptron. The output of the multi-layer perceptron is fused and then subjected to the Sigmoid activation function. The output of the Sigmoid activation function is multiplied by the input of the channel multi-scale attention module to obtain the output of the first stage; The output of the first stage is subjected to the first depth convolution, and the output of the first depth convolution is successively subjected to the second and third depth convolutions to obtain the output of the third depth convolution. The output of the first depth convolution is successively subjected to the fourth and fifth depth convolutions to obtain the output of the fifth depth convolution. The output of the first depth convolution is successively subjected to the sixth and seventh depth convolutions to obtain the output of the seventh depth convolution. The output of the third depth convolution is fused with the output of the fifth depth convolution and the output of the seventh depth convolution, and then passed through the convolution layer. The output of the convolution layer is multiplied with the output of the first stage to obtain the output of the second stage, that is, the output of the channel multi-scale attention module.
6. The SAR ship detection method based on multi-scale feature extraction and FRFT convolution according to claim 4 is characterized in that: The SPPCSPC module includes the 19th to 25th CBS modules and the 1st to 3rd maximum pooling layers, the convolution kernel size of the convolution layer in the 19th, 21st, 22nd, 24th and 25th CBS modules is 1, the convolution kernel size of the convolution layer in the 20th and 23rd CBS modules is 3, and the convolution kernel size of the 1st to 3rd maximum pooling layers is 5, 9 and 12 respectively; The input of the SPPCSPC module passes through the 19th CBS module, the 20th CBS module, and the 21st CBS module in sequence to obtain the output of the 21st CBS module; the output of the 21st CBS module passes through the first maximum pooling layer, the second maximum pooling layer, and the third maximum pooling layer respectively, and the output of the 21st CBS module is fused with the output of the first maximum pooling layer, the output of the second maximum pooling layer, and the output of the third maximum pooling layer, and then passes through the 22nd CBS module and the 23rd CBS module in sequence to obtain the output of the 23rd CBS module; The input of the SPPCSPC module passes through the twenty-fourth CBS module and then is fused with the output of the twenty-third CBS module. The fusion result passes through the twenty-fifth CBS module to obtain the output of the SPPCSPC module.
7. The SAR ship detection method based on multi-scale feature extraction and FRFT convolution according to claim 4 is characterized in that: The detection head part includes the first to third fractional-order Fourier transform convolution modules, the first to third RepConv modules and the first to third detection heads; the first output of the neck network is sequentially passed through the first fractional-order Fourier transform convolution module, the first RepConv module and the first detection head for large-scale prediction; the second output of the neck network is sequentially passed through the second fractional-order Fourier transform convolution module, the second RepConv module and the second detection head for medium-scale prediction; the third output of the neck network is sequentially passed through the third fractional-order Fourier transform convolution module, the third RepConv module and the third detection head for small-scale prediction.
8. The SAR ship detection method based on multi-scale feature extraction and FRFT convolution according to claim 7 is characterized in that: The structures of the first to third fractional-order Fourier transform convolution modules are the same. The input of each fractional-order Fourier transform convolution module first passes through a convolution layer with a convolution kernel size of 1, and is evenly divided into four parts. Each part is passed to a different Fn module, n=1, 2, 3, 4. The four Fn modules process at different frequencies to obtain feature maps of different frequencies. The feature maps of different frequencies obtained by the four Fn modules are fused to obtain a fused feature map. The input of each fractional-order Fourier transform convolution module passes through a convolution layer with a convolution kernel size of 1, and is fused with the fused feature map and the input of the fractional-order Fourier transform convolution module to obtain the output of each fractional-order Fourier transform convolution module. In the Fn module, the gradient descent algorithm is used to update the learning tensor t. After the Gaussian regular distribution operation, t is used as the input of the fractional Fourier filter. The output of the fractional Fourier filter is as follows: , in, represents the filter components, Represent the filter components in the X and Y directions respectively, f is the frequency parameter, p is the fractional order, x , y They are the elements of the 2D grid coordinate matrices in the X and Y directions generated by decomposing the filter input.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the SAR ship detection method based on multi-scale feature extraction and FRFT convolution as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
SAR vessel detection method based on multi-scale feature enhancement network
CN118196496A