Synthetic aperture radar ship image detection method and system
Through the feature extraction method of multi-branch preprocessing and multi-scale mapping modules, combined with the RetinaNet network, the accuracy and robustness of SAR ship detection in complex sea surface environments are solved, and efficient ship object detection is achieved.
Patent Information
- Application Number
- CN202510359462.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing synthetic aperture radar SAR ship detection method has limited detection accuracy and poor robustness when facing complex sea surface environments, making it difficult to adapt to the modeling characteristics of different sea areas. Conventional deep learning algorithms are not effective in SAR image processing.
The feature enhancement and extraction methods of multi-branch preprocessing module and multi-scale mapping module are adopted, and classification and regression processing is combined with the RetinaNet network, including convolution, normalization, maximum pooling, U-shaped structure and multi-scale mapping structure, suppress noise interference and extract feature information of different scales.
It significantly improves the accuracy and robustness of SAR ship detection, can effectively deal with target sparseness and noise interference, greatly reduces the computational complexity, and is suitable for a variety of sea surface environments and complex scenarios.
Smart Images

Figure CN120298983A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a synthetic aperture radar ship image detection method and system. Background Art
[0002] In a marine environment, real-time monitoring and identification of ships are crucial for maritime safety and the economy. Synthetic Aperture Radar (SAR), due to its microwave imaging ability, has become a key tool for ship detection. SAR is not restricted by light and weather and can monitor ships all-weather, thus providing strong support for maritime law enforcement and resource management.
[0003] SAR ship detection utilizes sea surface image data to automatically identify and detect ships through image processing and machine learning algorithms. However, due to problems such as sparse targets, large noise interference, and many small targets in SAR images, the detection difficulty is increased. Currently, SAR ship detection mainly uses Constant False Alarm Rate (CFAR) detection and conventional deep learning object detection algorithms. CFAR detection is mainly based on the differences in mathematical statistical characteristics between sea surface waves and ships for detection. However, due to the complexity and strong interference characteristics of the sea surface environment, the detection accuracy of CFAR is limited. At the same time, there are differences in the modeling characteristics of different sea areas, resulting in a CFAR algorithm being difficult to adapt to all sea surface conditions, so its robustness is poor. Conventional deep learning object detection algorithms have achieved remarkable results on ordinary optical images, but have poor support when dealing with SAR images.
[0004] Therefore, there is an urgent need for a prediction method with higher accuracy and stronger robustness to perform ship detection on synthetic aperture radar ship images. Summary of the Invention
[0005] The purpose of this application is to provide a synthetic aperture radar ship image detection method and system, which can improve the accuracy and robustness of SAR ship detection.
[0006] This application is implemented as follows:
[0007] In a first aspect, the present application provides a method for detecting synthetic aperture radar ship images, including the following steps: inputting a synthetic aperture radar ship image into a preprocessing module with multiple branches for feature enhancement to obtain a first feature map, where each branch sequentially performs convolution, normalization, and max pooling processing with different sizes. The outputs of each branch except the initial branch are superimposed by an adder to generate a weight matrix, which is multiplied by the output of the initial branch to obtain the first feature map; extracting features from the first feature map through a mapping module to obtain multiple sets of second feature maps, where the mapping module includes a U-shaped structure composed of multiple downsampling units and upsampling units with the same number as the branches, and a multi-scale mapping structure; inputting multiple sets of second feature maps into a RetinaNet network for classification and regression processing, and summarizing the detection results of each group to obtain the final ship detection result.
[0008] In some implementation manners, the preprocessing module includes three branches. The sizes of the max pooling layers in the three branches are 3×3, 1×3, and 3×1 respectively. The outputs of the two branches with the max pooling layer sizes of 1×3 and 3×1 are superimposed by an adder, and then a weight matrix is generated through a sigmoid activation function, which is multiplied by the output of the branch with the max pooling layer size of 3×3 to obtain the first feature map.
[0009] In some implementation manners, in the three branches of the preprocessing module, the convolution kernel sizes of the convolution layers are all 3×3, and the normalization layers all adopt BatchNorm combined with the ReLU function.
[0010] In some implementation manners, the U-shaped structure includes superimposing the downsampling unit and the upsampling unit in the following manner: the output of one downsampling unit is used as the input of the next downsampling unit until the downsampling unit is processed; the output of the last downsampling unit is used as the input of the first upsampling unit, and the output of the penultimate downsampling unit is used as the input of the second upsampling unit, and so on until all upsampling units are processed.
[0011] In some implementation manners, the downsampling unit includes: a first residual connection path, including a 1×1 convolution layer and a batch normalization layer connected in series, for outputting a first residual feature; a first feature extraction path, including a shape reorganization structure and a multi-head attention mechanism based on q, k, and v, for extracting a first global feature; a first feature fusion path, for splicing the first residual feature and the first global feature and then performing feature output through a 1×1 convolution layer and a normalization layer.
[0012] In some implementations, the upsampling unit includes: a second residual connection path, including a size scaling layer, a 1×1 convolutional layer, and a batch normalization layer connected in series, for performing 1×1 convolution processing after magnifying the size of the feature map by linear interpolation to output second residual features; a second feature extraction path, including a shape reorganization structure, a multi-head attention mechanism based on q, k, and v, and a size scaling layer at the tail, for extracting second global features; and a second feature fusion path, for directly adding the second residual features and the second global features and then outputting the features.
[0013] In some implementations, the multi-scale mapping structure is used to splice the feature maps output by the downsampling unit and the upsampling unit, and perform feature processing through a convolutional layer and a normalization layer to output multiple groups of second feature maps of different sizes.
[0014] In some implementations, the RetinaNet network includes a classification subnet and a bounding box regression subnet. The classification subnet is used to judge the category of the candidate target region, and the bounding box regression subnet is used to adjust the position of the candidate target region. Among them, the classification subnet and the bounding box regression subnet run simultaneously and share weights.
[0015] In some implementations, summarizing the detection results of each group to obtain the final ship detection result includes: passing the summarized detection results of each group through confidence screening and non-maximum suppression processing to obtain the final ship detection result.
[0016] In a second aspect, the present application provides a synthetic aperture radar ship image detection system, which includes:
[0017] A preprocessing component, for inputting the synthetic aperture radar ship image into a preprocessing module including multiple branches for feature enhancement to obtain a first feature map. Among them, each branch sequentially performs convolution, normalization, and maximum pooling processing of different sizes. The outputs of each branch except the initial branch are superimposed by an adder to generate a weight matrix, and multiplied by the output of the initial branch to obtain the first feature map; a feature extraction component, for extracting features from the first feature map through a mapping module to obtain multiple groups of second feature maps. Among them, the mapping module includes a U-shaped structure composed of multiple downsampling units and upsampling units with the same number as the number of branches, and a multi-scale mapping structure; a detection output component, for inputting the multiple groups of second feature maps into the RetinaNet network for classification and regression processing, and summarizing the detection results of each group to obtain the final ship detection result.
[0018] In a third aspect, the present application provides an electronic device, which includes a memory for storing one or more programs; a processor; when the above one or more programs are executed by the above processor, the method described in any item of the first aspect above is implemented.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above first aspects is implemented.
[0020] Compared with the prior art, the present application has at least the following advantages or beneficial effects:
[0021] The present application proposes a method for detecting synthetic aperture radar ship images. Through multiple branches of the preprocessing module, ship features are strengthened and noise interference is suppressed, improving the quality of the feature map. By the multi-scale mapping module, feature information at different scales is taken into account, avoiding the loss of small target information and improving the detection ability for small targets and complex backgrounds. At the same time, the design of the U-shaped structure composed of multiple downsampling units and upsampling units, as well as the multi-scale mapping structure, can adapt to the complex environments of different sea areas and reduce the dependence on single modeling characteristics. That is, through this innovative design of preprocessing, multi-scale feature extraction, and detection modules, the present application significantly improves the accuracy and robustness of SAR ship image detection, while reducing the computational complexity, thereby effectively dealing with problems such as sparse targets and large noise interference in SAR images and being applicable to various sea surface environments and complex scenarios. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a flowchart of an embodiment of a method for detecting synthetic aperture radar ship images according to the present application;
[0024] Figure 2 It is a framework diagram of an embodiment of a method for detecting synthetic aperture radar ship images according to the present application;
[0025] Figure 3 It is a schematic structural diagram of the preprocessing module in an embodiment of the present application;
[0026] Figure 4 It is a schematic structural diagram of the downsampling unit in an embodiment of the present application;
[0027] Figure 5 It is a schematic structural diagram of the upsampling unit in an embodiment of the present application;
[0028] Figure 6 It is a schematic structural diagram of the multi-scale mapping structure in an embodiment of the present application;
[0029] Figure 7 It is a structural block diagram of an embodiment of a synthetic aperture radar ship image detection system of the present application;
[0030] Figure 8 It is a structural block diagram of an electronic device provided by an embodiment of the present application. Specific Embodiments
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0032] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the various embodiments and the various features in the embodiments below can be combined with each other.
[0033] Embodiment 1
[0034] The embodiment of the present application provides a synthetic aperture radar ship image detection method, which realizes high-precision and robust detection of SAR ship images through feature enhancement and extraction of a multi-branch preprocessing module and a mapping module, combined with classification and regression processing of the RetinaNet network.
[0035] Please refer to Figure 1-2 , the synthetic aperture radar ship image detection method includes the following steps:
[0036] Step S101: Input the synthetic aperture radar ship image into a preprocessing module including multiple branches for feature enhancement to obtain a first feature map. Among them, each branch sequentially performs convolution, normalization, and max-pooling processing of different sizes. The outputs of each branch except the initial branch are superimposed by an adder to generate a weight matrix, which is multiplied by the output of the initial branch to obtain the first feature map;
[0037] Step S101 is the starting stage of the entire detection method. Its core lies in using a preprocessing module with multiple branches to enhance the features of the input synthetic aperture radar (SAR) ship image. Each branch performs a series of operations, including convolution, normalization, and max pooling with different sizes. These operations aim to extract feature information from the image from different angles and scales. Among them, the normalization layer and max pooling layer in the preprocessing module can effectively reduce the influence of noise and outliers, enabling stable detection even when facing complex synthetic aperture radar images, such as those with noise interference and sparse targets. In particular, except for the initial branch, the outputs of other branches are superimposed through an adder to generate a weight matrix. This weight matrix is then multiplied by the output of the initial branch to enhance key features and suppress secondary features, finally generating the first feature map.
[0038] Among them, through the processing of multiple branches, features of different scales and angles in the input synthetic aperture radar ship image can be extracted, increasing the diversity of features. At the same time, the introduction of the weight matrix strengthens the key features, which helps to improve the accuracy of feature extraction and classification in subsequent steps. That is, through the generation of a multi-branch structure and a weight matrix, ship features are strengthened, noise interference is suppressed, and the quality of the feature map is improved.
[0039] Step S102: Feature extraction is performed on the first feature map through a mapping module to obtain multiple groups of second feature maps. Among them, the mapping module includes a U-shaped structure composed of multiple downsampling units and upsampling units with the same number as the number of branches, and a multi-scale mapping structure.
[0040] In step S102, the first feature map is input into the mapping module for further feature extraction. The mapping module consists of multiple U-shaped structures with the same number as the number of branches and a multi-scale mapping structure. The U-shaped structure can extract features of different scales through the combination of downsampling units and upsampling units. The multi-scale mapping structure further enhances the feature extraction ability, extracts feature information from different scales through different mapping methods, and finally generates multiple groups of second feature maps. That is, the introduction of the U-shaped structure and the multi-scale mapping structure enables the method to handle ship targets of different scales, improves scale invariance, can take into account the detection requirements of large and small targets, and avoids information loss.
[0041] Step S103: Input multiple groups of second feature maps into the RetinaNet network for classification and regression processing, and summarize the detection results of each group to obtain the final ship detection result.
[0042] In step S103, multiple groups of second feature maps are respectively input into the RetinaNet network for classification and regression processing. The RetinaNet network is an advanced object detection network that combines the two tasks of classification and regression and can accurately identify and locate objects in images. For each group of second feature maps, the RetinaNet network outputs a group of detection results, including the category and location information of the objects. Finally, the final ship detection results are obtained by summarizing the detection results of each group.
[0043] In summary, in the above implementation, through multiple branches of the preprocessing module, ship features are strengthened and noise interference is suppressed, improving the quality of the feature maps; by the multi-scale mapping module, feature information of different scales is taken into account, avoiding the loss of small object information, and improving the detection ability for small objects and complex backgrounds. At the same time, the U-shaped structure composed of multiple downsampling units and upsampling units, as well as the design of the multi-scale mapping structure, can adapt to the complex environments of different sea areas and reduce the dependence on single modeling characteristics. Among them, data redundancy is reduced by the downsampling units, and the difficulty of model training is reduced by combining residual connections, improving the computing efficiency; the design of the multi-scale mapping structure avoids repeated calculations and optimizes the feature extraction process. It should be noted that through this innovative design of preprocessing, multi-scale feature extraction and detection modules, the accuracy and robustness of SAR ship image detection are significantly improved, while the computational complexity is reduced, so that problems such as sparse targets and large noise interference in SAR images can be effectively processed, and it is applicable to various sea surface environments and complex scenarios.
[0044] Please refer to Figure 3 , based on the foregoing solution, in some implementations of the present application, the preprocessing module includes three branches, and the sizes of the max-pooling layers in the three branches are 3×3, 1×3, and 3×1 respectively. After the outputs of the two branches with the max-pooling layer sizes of 1×3 and 3×1 are superimposed by an adder, a weight matrix is generated through a sigmoid activation function and then multiplied by the output of the branch with the max-pooling layer size of 3×3 to obtain the first feature map.
[0045] In the above implementation, the preprocessing module is composed of three branches, and each branch follows the same basic processing flow, that is, the input synthetic aperture radar ship image is sequentially subjected to convolution, normalization, and max-pooling processing. However, the sizes of the max-pooling layers in each branch are different, which are 3×3, 1×3, and 3×1 respectively.
[0046] Among them, the 3×3 max-pooling operation in the 3×3 max-pooling branch can reduce the dimension of the ship image as a whole. While reducing the data volume, it can remove some noise interference and retain the main feature information of the ship, which helps to grasp the overall shape of the ship. In the synthetic aperture radar ship image, the ship usually appears as a rectangular strip when observed from high altitude. The 1×3 strip pooling in the 1×3 max-pooling branch can better capture and strengthen the features of the horizontally arranged ships and enhance the horizontal feature information. Corresponding to the 1×3 pooling, the 3×1 pooling operation in the 3×1 max-pooling branch is mainly used to capture and strengthen the features of the vertically arranged ships and highlight the vertical feature information.
[0047] The outputs of the two branches with the max-pooling layer sizes of 1×3 and 3×1 will first be superimposed through an adder. The superimposed result is processed by the sigmoid activation function to generate a weight matrix. The sigmoid function maps the superimposed result to the interval (0, 1), making each element in the generated weight matrix represent the importance degree of the feature at the corresponding position. Finally, the generated weight matrix is multiplied by the output of the branch with the max-pooling layer size of 3×3. This multiplication operation actually weights the overall features extracted by the 3×3 branch and strengthens the ship features (such as horizontal or vertical features) highlighted in the 1×3 and 3×1 branches, thereby obtaining the first feature map.
[0048] In summary, the max-pooling branches with different sizes can extract ship features from multiple angles and directions. The 3×3 pooling focuses on the overall features, while the 1×3 and 3×1 poolings are respectively aimed at the horizontal and vertical features. Through this multi-branch design, the feature information of the ship can be captured more comprehensively and meticulously. The generation and application of the weight matrix enable the preprocessing module to enhance the overall features targeted according to the importance degree of the features in different directions. For ships with obvious horizontal or vertical features, the corresponding weights will increase, thereby highlighting these features and improving the recognition ability of the preprocessing module for ships with different shapes.
[0049] Among them, the 3×3 max-pooling can reduce the dimension of the feature map to a certain extent, remove some noise interference, and make the preprocessing module pay more attention to the main features of the ship. The role of the weight matrix further focuses on the meaningful ship features, reduces the influence of noise and irrelevant information on the subsequent detection, and improves the quality of the features. Thus, through the comprehensive extraction and targeted enhancement of ship features in the above manner, the preprocessing module can identify ship targets more accurately. Especially for the long strip-shaped ships that are easily overlooked in the synthetic aperture radar image, this technical solution can better capture their features, thereby improving the accuracy and reliability of the detection.
[0050] Based on the foregoing solution, in some implementation manners of the present application, among the three branches of the preprocessing module, the convolution kernel size of the convolutional layer is 3×3, and the normalization layer all adopts BatchNorm combined with the ReLU function.
[0051] Convolution operation is a basic operation for feature extraction in deep learning. By sliding the convolution kernel on the image and performing weighted summation on the local area, the features of the image are extracted. In the above implementation manner, in the three branches of the preprocessing module, the convolution kernel size of the convolutional layer is set to 3×3, so as to ensure the effective extraction of the local features of the image while maintaining a high computational efficiency. The accurate extraction of such local features helps the preprocessing module to accurately identify ship targets in complex backgrounds.
[0052] The use of BatchNorm (batch normalization) makes the input distribution of each layer in the network more stable during the training process, reducing the impact of internal covariate shift. This helps the network to converge faster, reduces the fluctuations during the training process, and improves the stability and efficiency of the training of the preprocessing module. When processing synthetic aperture radar ship images, due to the complexity and diversity of the images, the training process may face many challenges, and BatchNorm can effectively alleviate these problems, enabling the preprocessing module to better learn the features of the images.
[0053] The non-linear characteristic of the ReLU (rectified linear unit) function introduces more expressive power to the preprocessing module. It enables the network to learn more complex feature patterns and distinguish different types of ship features. By setting negative pixel values to 0, the ReLU function can effectively reduce redundant information and highlight important features, thereby improving the detection accuracy of the preprocessing module.
[0054] Please refer to Figure 2 , based on the foregoing solution, in some implementation manners of the present application, the U-shaped structure includes stacking the downsampling unit and the upsampling unit in the following manner: the output of one downsampling unit is used as the input of the next downsampling unit until the downsampling unit is processed; the output of the last downsampling unit is used as the input of the first upsampling unit, and the output of the penultimate downsampling unit is used as the input of the second upsampling unit, and so on until all upsampling units are processed.
[0055] In the above implementation, the core of the U-shaped structure lies in the stacking method of the downsampling unit and the upsampling unit. The downsampling unit is used to extract the features of the image and capture higher-level features by gradually reducing the resolution of the image (i.e., downsampling). The upsampling unit is used to gradually restore the resolution of the image in order to locate and segment the target area in the final output. Among them, the output of one downsampling unit is directly used as the input of the next downsampling unit. This cascading method helps to gradually extract the deep features in the image while reducing the dimension and computational amount of the data. Contrary to the downsampling unit, the connection of the upsampling unit is reversed. That is, the output of the last downsampling unit is used as the input of the first upsampling unit, the output of the penultimate downsampling unit is used as the input of the second upsampling unit, and so on until all upsampling units are processed. This connection method helps to combine the deep features with the detailed information of the original image, thereby improving the accuracy of segmentation.
[0056] Through the stacking of the downsampling unit and the upsampling unit in the above implementation, the U-shaped structure can effectively extract the deep features in the image and combine these features with the detailed information of the original image during the upsampling process. That is, the downsampling unit and the upsampling unit are symmetrically connected to form a U-shaped structure, ensuring the integrity and consistency of information during the process of dimensionality reduction and dimensionality increase. This way of feature extraction and fusion helps to improve the accuracy of image segmentation.
[0057] Please refer to Figure 4 , based on the foregoing solution, in some implementations of the present application, the downsampling unit includes: a first residual connection path, including a 1×1 convolutional layer and a batch normalization layer connected in series, for outputting first residual features; a first feature extraction path, including a shape reorganization structure and a multi-head attention mechanism based on q, k, v, for extracting first global features; a first feature fusion path, for splicing the first residual features and the first global features and then performing feature output through a 1×1 convolutional layer and a normalization layer.
[0058] As Figure 4 shown, the downsampling unit in the above implementation is an improved Transformer downsampling structure. The original Transformer structure has problems of huge data volume and difficult training. The improved Transformer downsampling structure realizes the reduction of data volume through shape reorganization. The implementation principle is to reduce the number of pixels in proportion and reduce data redundancy.
[0059] Figure 4The first residual connection path of the left part includes a 1×1 convolutional layer and a batch normalization layer (BatchNorm) in series, which can be used to perform dimensionality reduction and normalization processing on the input feature map, achieve the reduction of the size of the feature map and information connection, reduce the training difficulty of the downsampling unit, output the first residual feature, retain local feature information, and at the same time reduce the training difficulty of the downsampling unit.
[0060] Figure 4 The first feature extraction path of the right part includes a shape reorganization structure and a multi-head attention mechanism based on q (query), k (key), and v (value). The shape reorganization structure is mainly used to achieve the reduction of the data volume, reduce the pixel points in equal proportion, reduce data redundancy, and reduce the amount of calculation. The multi-head attention mechanism is the core component in the Transformer structure. It maps the input features to three different spaces of query (q), key (k), and value (v), calculates the similarity between the query and the key to obtain attention scores, and then weights and sums the values according to these scores to extract global features. In this path, the input feature map can be globally modeled through the multi-head attention mechanism, capturing the dependencies between different positions in the feature map, and outputting the first global feature.
[0061] Figure 4 The function of the first feature fusion path of the lower part is to splice the first residual feature and the first global feature. The splicing operation can fuse the local feature (the first residual feature) and the global feature (the first global feature), so that the output feature contains both local detailed information and global context information. The spliced feature is further processed by a 1×1 convolutional layer and a normalization layer to further adjust the dimension and distribution of the feature, and finally output the fused feature.
[0062] In summary, in the above implementation method, by combining the residual connection and the global attention mechanism, this technical solution can retain local detailed information and capture global context information at the same time, thereby enhancing the feature representation ability of the downsampling unit. Moreover, the use of the 1×1 convolutional layer effectively reduces the amount of calculation, while the batch normalization and feature fusion strategies help to accelerate the training process, improve the convergence speed and overall efficiency of the downsampling unit.
[0063] Please refer to Figure 5, Based on the foregoing solution, in some implementation manners of the present application, the upsampling unit includes: a second residual connection path, including a size scaling layer, a 1×1 convolutional layer, and a batch normalization layer connected in series, for performing 1×1 convolution processing after magnifying the size of the feature map by linear interpolation to output a second residual feature; a second feature extraction path, including a shape reorganization structure, a multi-head attention mechanism based on q, k, and v, and a size scaling layer at the tail, for extracting a second global feature; a second feature fusion path, for directly adding the second residual feature and the second global feature and then performing feature output.
[0064] Traditional upsampling structures are usually based on the values of pixels and do not have the ability to learn and adaptively adjust. The upsampling method based on transposed convolution has the ability to learn and adaptively adjust, but it focuses too much on local information and ignores the processing of global information. However, the upsampling unit in the above implementation manner can solve these problems. As Figure 5 shown, the upsampling unit in the above implementation manner is similar in structure to the Figure 4 downsampling unit in, but the difference is that in order to increase the dimension, the upsampling unit adds a size scaling layer, and the size scaling magnifies the size of the feature map by linear interpolation. In the second residual connection path, the size scaling layer is placed at the starting position, while in the second feature extraction path, it is placed at the ending position.
[0065] In addition, after directly adding the second residual feature and the second global feature and then performing feature output, the feature map at this time will have the same dimension as the feature map output by the downsampling unit, and the feature map at this time contains information about larger ship targets, which is beneficial for the upsampling unit to identify ship targets in image interference. In the above implementation manner, while ensuring the increase in the size of the feature map, by combining the residual connection and the global attention mechanism, it is possible to simultaneously retain the local details and global context information of the features, effectively avoiding the problem that transposed convolution upsampling only focuses on local information.
[0066] Please refer to Figure 6 , Based on the foregoing solution, in some implementation manners of the present application, the multi-scale mapping structure is used to splice the feature maps output by the downsampling unit and the upsampling unit, and perform feature processing through a convolutional layer and a normalization layer to output multiple groups of second feature maps with different sizes.
[0067] In the above implementation, after concatenating the feature maps output by the downsampling unit and the upsampling unit, they are processed through one or more convolutional layers. The convolutional layer slides a convolutional kernel over the feature map to extract local features and realizes the non-linear transformation of features through different combinations of convolutional kernels. This step helps to further fuse and extract the information in the feature map, while reducing the dimension of the feature map, providing convenience for subsequent processing. Then, the feature map after convolutional processing is normalized by a normalization layer. The normalization layer normalizes each channel of the feature map, making the distribution of the feature map more stable, which helps to accelerate the training of the multi-scale mapping structure and improve the performance.
[0068] In previous studies, shallow large-size feature maps support small targets well, and deep small-size feature maps support large targets well. Because small target information is prone to information loss during continuous reduction of feature maps. In the above implementation, the multi-scale mapping structure can output multiple groups of second feature maps of different sizes, thus ensuring the effective fusion and complementarity of shallow features and deep features. These feature maps not only contain feature information of different scales, but also have been optimized by convolutional and normalization processing, providing a more robust and effective feature representation for subsequent tasks.
[0069] Based on the foregoing solution, in some implementations of the present application, the RetinaNet network includes a classification subnet and a bounding box regression subnet. The classification subnet is used to judge the category of the candidate target region, and the bounding box regression subnet is used to adjust the position of the candidate target region. Among them, the classification subnet and the bounding box regression subnet run simultaneously and share weights.
[0070] In the above implementation, the classification subnet is used to judge the category (ship category) of the candidate target region, mainly including processing the input feature map through convolutional layers and fully connected layers, and outputting the category probability of each candidate region. The bounding box regression subnet is used to adjust the position (bounding box coordinates) of the candidate target region, mainly including processing the input feature map through convolutional layers and fully connected layers, and outputting the bounding box offset of each candidate region. Among them, the design of weight sharing reduces the number of parameters of the RetinaNet network, reduces the computational complexity, and at the same time improves the training efficiency and consistency of the RetinaNet network.
[0071] Based on the foregoing solution, in some implementations of the present application, summarizing the detection results of each group to obtain the final ship detection result includes: screening the summarized detection results of each group through confidence filtering and non-maximum suppression processing to obtain the final ship detection result.
[0072] When detecting synthetic aperture radar ship images, since there may be multiple ship targets in the image, and these targets may have different sizes, shapes, and occlusion situations, multiple candidate detection results will be generated. In order to screen out the most accurate and reliable detection results from these candidate results, the above implementation method proposes a method for summarizing the detection results of each group, which mainly includes two steps: confidence screening and non-maximum suppression (NMS) processing.
[0073] Through confidence screening and non-maximum suppression processing, inaccurate or redundant detection results can be effectively removed, and the most reliable and accurate detection results can be retained. This helps to improve the accuracy of ship detection and reduce the probabilities of false detection and missed detection. Among them, confidence screening can quickly reduce the number of candidate detection results and reduce the complexity of subsequent processing. At the same time, non-maximum suppression processing can also avoid repeated processing of overlapping detection frames, thereby further improving the detection efficiency. Non-maximum suppression processing can ensure that each ship target has only one optimal detection frame corresponding to it, avoiding the situation of multiple detection frames overlapping and being chaotic. This makes the detection results clearer, easier to understand and interpret.
[0074] Embodiment 2
[0075] Please refer to Figure 7 , this embodiment of the present application provides a synthetic aperture radar ship image detection system, which includes: a preprocessing component for inputting the synthetic aperture radar ship image into a preprocessing module with multiple branches for feature enhancement to obtain a first feature map. Among them, each branch sequentially performs convolution, normalization, and max-pooling processing of different sizes. The outputs of each branch except the initial branch are superimposed by an adder to generate a weight matrix, and are multiplied by the output of the initial branch to obtain the first feature map; a feature extraction component for extracting features from the first feature map through a mapping module to obtain multiple groups of second feature maps. Among them, the mapping module includes a U-shaped structure composed of multiple downsampling units and upsampling units with the same number as the number of branches, and a multi-scale mapping structure; a detection output component for inputting multiple groups of second feature maps into the RetinaNet network for classification and regression processing, and summarizing the detection results of each group to obtain the final ship detection result.
[0076] For the specific implementation process of the above system, please refer to a synthetic aperture radar ship image detection method provided in Embodiment 1, which will not be elaborated here.
[0077] Embodiment 3
[0078] Please refer to Figure 8, an embodiment of the present application provides an electronic device, which includes at least one processor 201 and at least one memory 202; wherein, the processor 201 is directly connected to the memory 202, or communicates with each other through a communication interface 203, or is electrically connected through one or more communication buses or signal lines to achieve data transmission or interaction; the memory 202 stores program instructions executable by the processor 201, and the processor 201 calls the program instructions to execute a synthetic aperture radar ship image detection method. For example, it can achieve:
[0079] Input the synthetic aperture radar ship image into a preprocessing module with multiple branches for feature enhancement to obtain a first feature map. Among them, each branch sequentially performs convolution, normalization, and max pooling with different sizes. The outputs of each branch except the initial branch are superimposed by an adder to generate a weight matrix, which is multiplied by the output of the initial branch to obtain the first feature map; extract features from the first feature map through a mapping module to obtain multiple groups of second feature maps. Among them, the mapping module includes a U-shaped structure composed of multiple downsampling units and upsampling units with the same number as the number of branches, and a multi-scale mapping structure; input multiple groups of second feature maps into the RetinaNet network for classification and regression processing, and summarize the detection results of each group to obtain the final ship detection result.
[0080] Among them, the memory 202 can be, but is not limited to, random access memory (Random Access Memory, RAM), read only memory (Read Only Memory, ROM), programmable read only memory (Programmable Read-Only Memory, PROM), erasable programmable read only memory (Erasable Programmable Read-Only Memory, EPROM), electrically erasable programmable read only memory (Electric Erasable Programmable Read-Only Memory, EEPROM), etc.
[0081] The processor 201 can be an integrated circuit chip with signal processing capabilities. The processor 201 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0082] It can be understood that Figure 8 The structure shown is only schematic, and the electronic device may also include more or fewer components than those shown Figure 8 in the figure, or have a different configuration from that shown Figure 8 in the figure. Figure 8 Each component shown in the figure can be implemented using hardware, software, or a combination thereof.
[0083] Embodiment 4
[0084] This application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor 201, it implements a synthetic aperture radar ship image detection method.
[0085] For example, it realizes:
[0086] Input the synthetic aperture radar ship image into a preprocessing module with multiple branches for feature enhancement to obtain a first feature map. Among them, each branch sequentially performs convolution, normalization, and max pooling with different sizes. The outputs of each branch except the initial branch are superimposed by an adder to generate a weight matrix, which is multiplied by the output of the initial branch to obtain the first feature map; perform feature extraction on the first feature map through a mapping module to obtain multiple groups of second feature maps. Among them, the mapping module includes a U-shaped structure composed of multiple downsampling units and upsampling units with the same number as the number of branches, and a multi-scale mapping structure; input multiple groups of second feature maps into the RetinaNet network for classification and regression processing, and summarize the detection results of each group to obtain the final ship detection result.
[0087] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0088] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above, and that the present application can be implemented in other specific forms without departing from the spirit or essential features of the present application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the present application is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims be included in the present application. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. A synthetic aperture radar ship image detection method, characterized in that It includes the following steps: Input the synthetic aperture radar ship image into a preprocessing module with multiple branches for feature enhancement to obtain a first feature map. Among them, each branch performs convolution, normalization, and max pooling processing with different sizes in sequence. The outputs of each branch except the initial branch are superimposed by an adder to generate a weight matrix, which is multiplied by the output of the initial branch to obtain the first feature map; Extract features from the first feature map through a mapping module to obtain multiple groups of second feature maps. Among them, the mapping module includes a U-shaped structure composed of multiple downsampling units and upsampling units with the same number as the number of branches, and a multi-scale mapping structure; Input multiple groups of second feature maps into the RetinaNet network for classification and regression processing, and summarize the detection results of each group to obtain the final ship detection result.
2. The method according to claim 1, characterized in that, The preprocessing module includes three branches. The sizes of the max pooling layers in the three branches are 3×3, 1×3, and 3×1 respectively. The outputs of the two branches with max pooling layer sizes of 1×3 and 3×1 are superimposed by an adder, and then a weight matrix is generated through a sigmoid activation function, which is multiplied by the output of the branch with a max pooling layer size of 3×3 to obtain the first feature map.
3. The method according to claim 2, wherein In the three branches of the preprocessing module, the convolution kernel sizes of the convolutional layers are all 3×3, and the normalization layers all adopt BatchNorm combined with the ReLU function.
4. The method according to claim 1, wherein The U-shaped structure includes the following method for superimposing the downsampling unit and the upsampling unit: The output of one downsampling unit is used as the input of the next downsampling unit until the downsampling unit is processed; the output of the last downsampling unit is used as the input of the first upsampling unit, and the output of the penultimate downsampling unit is used as the input of the second upsampling unit, and so on until all upsampling units are processed.
5. The method according to claim 1, characterized in that, The downsampling unit includes: The first residual connection path, including a 1×1 convolutional layer and a batch normalization layer connected in series, is used to output the first residual feature; The first feature extraction path, including a shape reorganization structure and a multi-head attention mechanism based on q, k, v, is used to extract the first global feature; The first feature fusion path is used to splice the first residual feature and the first global feature and then perform feature output through a 1×1 convolutional layer and a normalization layer.
6. The method according to claim 1, wherein The upsampling unit includes: The second residual connection path, including a size scaling layer, a 1×1 convolutional layer, and a batch normalization layer connected in series, is used to perform 1×1 convolution processing after magnifying the feature map size through linear interpolation to output the second residual feature; The second feature extraction path, including a shape reorganization structure, a multi-head attention mechanism based on q, k, v, and a size scaling layer at the end, extracts the second global feature; The second feature fusion path is used to directly add the second residual feature and the second global feature and then perform feature output.
7. The method according to claim 1, characterized in that The multi-scale mapping structure is used to splice the feature maps output by the downsampling unit and the upsampling unit, and perform feature processing through a convolutional layer and a normalization layer to output multiple groups of second feature maps with different sizes.
8. The method according to claim 1, characterized in that, The RetinaNet network includes a classification subnet and a bounding box regression subnet. The classification subnet is used to determine the category of the candidate target region, and the bounding box regression subnet is used to adjust the position of the candidate target region. Among them, the classification subnet and the bounding box regression subnet operate simultaneously and share weights.
9. The method according to claim 1, wherein Summarizing the detection results of each group to obtain the final ship detection result includes: screening the summarized detection results of each group through confidence filtering and non-maximum suppression processing to obtain the final ship detection result.
10. A synthetic aperture radar ship image detection system, characterized in that, Including: A preprocessing component, which is used to input the synthetic aperture radar ship image into a preprocessing module with multiple branches for feature enhancement to obtain a first feature map. Among them, each branch sequentially performs convolution, normalization, and max pooling processing of different sizes. The outputs of each branch except the initial branch are superimposed by an adder to generate a weight matrix, and then multiplied by the output of the initial branch to obtain the first feature map; A feature extraction component, which is used to extract features from the first feature map through a mapping module to obtain multiple groups of second feature maps. Among them, the mapping module includes a U-shaped structure composed of multiple downsampling units and upsampling units with the same number as the number of branches, and a multi-scale mapping structure; A detection output component, which is used to input multiple groups of second feature maps into the RetinaNet network for classification and regression processing, and summarize the detection results of each group to obtain the final ship detection result.