A Ship Target Recognition Method for SAR Images Based on Attention Mechanism

By adopting a small sample recognition method based on twin neural network in SAR image ship target recognition, combined with FSSD network and attention mechanism, the problems of low recognition rate and complex background influence of small ships are solved, and efficient multi-scale ship detection and accurate target recognition are achieved.

CN117218612BActive Publication Date: 2025-05-30NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311273815.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-05-30
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

The existing SAR image ship target recognition methods have low recognition rate in small ships, inability to effectively avoid the influence of noise and similar objects, inability to effectively identify nearshore targets, and excessive model parameters.

Method used

A small sample SAR target recognition method based on twin neural network is adopted, combined with FSSD network, a lightweight network structure and attention mechanism are introduced, and the performance of the backbone network is enhanced through the CA attention module and the RFB receptive field module, and a top-down feature pyramid design scheme and dense connection are adopted to fuse multi-scale features.

Benefits of technology

The recognition rate of multi-scale ships in SAR images is improved, the processing capability of complex backgrounds is enhanced, the amount of model parameters is reduced, and the detection accuracy and efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117218612B_ABST
    Figure CN117218612B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for SAR image ship target recognition based on an attention mechanism, including: using a CA attention module and an RFB receptive field module to enhance the performance of the MobileNetV3 backbone network; adopting a top-down feature pyramid design scheme to extract multi-scale significant features containing semantic and spatial information; combining and using dense connections and lateral stitching between feature maps of different scales to fuse non-blurred global features and significant local features, and adding a CA attention module to adaptively refine the feature maps to adapt to different scenarios in SAR images. The present invention is based on the FSSD neural network, and the present invention proposes an ASSD network to better meet the needs of the SAR ship target recognition task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to synthetic aperture radar technology, and specifically relates to a method for identifying ship targets in SAR images based on an attention mechanism. Background Art

[0002] Synthetic Aperture Radar (SAR for short) is a high-resolution imaging radar. As an active microwave imaging sensor, compared with traditional passive imaging sensors (such as infrared and optical sensors), the SAR imaging process is less affected by environmental factors (such as weather, light, and clouds). It can detect hidden objects and work under all-weather and all-day conditions. Due to its characteristics of being unaffected by many conditions such as light and weather, it is widely used in ship target recognition. With the rapid development of spaceborne and airborne SAR, SAR has been widely applied in civilian and military fields.

[0003] Traditional SAR image ship recognition methods need to go through three stages: preprocessing, feature extraction, and classification recognition. The difficulty lies in feature extraction and feature selection. On the one hand, manually designing and selecting features is a time-consuming and laborious method, and whether good features can be selected depends heavily on experience and luck. On the other hand, manually designed features have limitations and cannot fully reflect the high-level information contained in the image. For China, the research level of SAR image ship target recognition technology is relatively backward compared with foreign countries. Therefore, improving the SAR image translation ability in civilian and military fields and realizing more intelligent SAR image ship target recognition have become urgent problems to be solved.

[0004] The method of shifting attention to the most important regions of an image while ignoring irrelevant parts is called the attention mechanism; the human visual system uses the attention mechanism to help people analyze and understand complex scenes more efficiently, which in turn has prompted researchers to introduce the attention mechanism into the computer vision system to improve its performance. In visual tasks, the attention mechanism can be regarded as a dynamic selection process that adaptively weights features according to the importance of the input. The attention mechanism plays a very important role in many visual tasks, such as image classification, face recognition, object detection, semantic segmentation, medical image processing, 3D vision, etc. In the SAR image ship detection task, due to the influence of complex backgrounds and image noise, multi-scale ships, especially small ships, are easily overlooked. Introducing the attention mechanism into the SAR ship recognition task to improve the detection accuracy is a very promising research direction.

[0005] Traditional ship detection methods in SAR images mainly focus on the detection of ship targets and ship wakes. Since ship wakes do not always exist and their features are not as obvious as those of ship targets, the research on ship wake detection is not extensive. For ship target detection, there are mainly three methods, including detection algorithms based on statistical features, thresholds, and transforms. Among the above detection methods, the Constant False Alarm Rate (CFAR) detection algorithm is one of the most widely used ship target detection algorithms at present. It detects ship targets by statistically modeling the background clutter. However, these ship detection methods mainly focus on strong scattering targets. In large-scale SAR images with continuously improving resolution, the target area occupies fewer pixels and contains some weak scattering, which makes it difficult for traditional single-polarization SAR methods to achieve real-time and accurate ship detection.

[0006] Compared with single-polarization SAR, Polarimetric SAR (PolSAR) can provide more information from the scattering matrix, including target structure and motion regions, thus improving the detection accuracy of SAR images. For PolSAR images, a new decomposition method has been proposed and combined with a CFAR detector based on the generalized gamma distribution to improve the signal-to-clutter ratio. In addition, to detect ships in crowded waters, variational Bayesian inference has been proven to have good detection performance and conformal ability. It has been proposed to use three superpixel-level dissimilarity measures to automatically detect ships. Although the ship detection method of PolSAR may have better performance than single-polarization SAR, it is also of high research value to explore how to use the limited information of SAR images to achieve higher detection accuracy. The method proposed in this project is mainly used to study the ship detection problem in single-polarization SAR images.

[0007] In recent years, convolutional neural networks have achieved remarkable success in the field of computer vision with their powerful representation ability and automatic feature extraction ability. Object detectors based on CNN can be divided into two methods. One is the single-stage method, such as the YOLO network and the SSD network, which are fast and easy to train. The other is the two-stage method, such as R-CNN and fast R-CNN, which have high accuracy but large computational costs. Generally speaking, two-stage algorithms have higher accuracy because the computational cost is also large. On the contrary, single-stage algorithms have the characteristics of fast and simple training in the end-to-end mode, but the accuracy is lower than that of two-stage algorithms. However, these methods only use the final feature map to predict the target, and the features of the target are lost from the lower layers, so it is difficult to detect small targets. To make full use of multi-level features, the Feature Pyramid Network (FPN) extracts features from different layers for individual detection to obtain better detection performance.

[0008] In the field of ship recognition in SAR images, CNN also shows superior detection performance. In the early stage, CNN was used in various parts of ship detection, such as land masking, ROI extraction, and ship target recognition. At the present stage, most of the ship detection algorithms based on CNN adopt an end-to-end approach to send SAR images to CNN for detection, directly outputting the ship detection results, eliminating the preprocessing and sea-land segmentation operations, and improving the detection efficiency and accuracy. To solve the problem of insufficient SAR training sets, SSD combination transfer learning has been proposed. An improved faster R-CNN method has been proposed for ship detection and a SAR ship detection dataset (SSDD) has been provided to train and test the model. FPN and faster R-CNN have been combined to detect multi-scale ships, making the multi-scale fusion more effective. SSD has been improved by using a semantic aggregation module and an attention module to detect ships and directions simultaneously. Regarding the problem of nearshore ship target detection, relevant work shows that aggregating the ship semantic information in the SSD feature map can improve the accuracy of nearshore ship detection.

[0009] In multiple trainings and experiments, both the SSD network and the FSSD network have exposed very similar problems in the SAR ship target recognition task, and the specific analysis is as follows:

[0010] 1. The recognition rate of small ship targets is not high. Especially for a dense group of ships, it is very difficult for both networks to detect small ships with a pixel size of less than 20. Even though the FSSD network uses a lightweight fusion module to fuse features at different levels, it still cannot provide sufficient semantic information and context information for the detector that recognizes small targets;

[0011] 2. It is unable to effectively avoid the influence of nearshore noise and similar objects. Neither the SSD network nor the FSSD network can eliminate the influence of invalid pixels and noise in the SAR ship target recognition task, especially unable to effectively avoid the false alarm influence of approximate ship objects;

[0012] 3. It is unable to effectively recognize nearshore targets. Similarly, for ship targets close to the shore, neither the SSD network nor the FSSD network can distinguish them from other invalid targets, and it is easy to ignore nearshore ship targets and focus on invalid targets;

[0013] 4. The number of model parameters is too large. In the actual training process, both the SSD and FSSD networks are prone to overfitting problems. Since the SSDD dataset only has 1,160 photos and 2,456 training samples, the two single-stage detection models have exposed problems of too high complexity and too large number of parameters in the SAR ship target recognition task. Summary of the Invention

[0014] The main purpose of the present invention is to provide a small-sample SAR target recognition method based on Siamese neural network, which aims to implement the SAR image ship target detection task based on the FSSD network, and on this basis, introduce a more lightweight network structure for redesign, more effectively establish a feature pyramid to make full use of effective features, and introduce an attention mechanism to improve the network structure, so as to solve the problems of low recognition rate of multi-scale ships and the influence of complex background on the detection effect in the SAR image ship target detection task.

[0015] The technical solution adopted by the present invention is: a SAR image ship target recognition method based on an attention mechanism, including:

[0016] Using the CA attention module and the RFB receptive field module to enhance the performance of the MobileNetV3 backbone network;

[0017] Adopting a top-down feature pyramid design scheme to extract multi-scale significant features containing semantic and spatial information;

[0018] Combining dense connection and lateral stitching between feature maps of different scales to fuse non-blurred global features and significant local features, and adding a CA attention module to adaptively refine the feature maps to adapt to different scenarios in SAR images.

[0019] Further, the using the CA attention module and the RFB receptive field module to enhance the performance of the MobileNetV3 backbone network includes:

[0020] On the basis of the original three branches, a dilated convolution with a dilation rate of 7 is added. Through connecting the inputs of the four branches and residual connections, the features of multi-scale ships can be better extracted by the detection model; at the same time, referring to the improvement in the number of parameters in the original text, using 1×3 and 3×1 convolutions to replace the 3×3 convolution in the second branch, 1×5 and 5×1 convolutions to replace the 5×5 convolution in the third branch, and 1×7 and 7×1 convolutions to replace the 7×7 convolution in the fourth branch, a lightweight receptive field enhancement structure more suitable for multi-scale ship detection in SAR images is generated.

[0021] The specific formula is expressed as follows:

[0022] Adding a dilated convolution with a dilation rate of 7:

[0023] Y = Residule(Concat(X1, X2, X3, X4), Conv2d, factor = 6, K = 1, S = 1)

[0024] Replace the 3×3 convolution in the second branch with 1×3 and 3×1 convolutions, replace the 5×5 convolution in the third branch with 1×5 and 5×1 convolutions, and replace the 7×7 convolution in the fourth branch with 1×7 and 7×1 convolutions:

[0025] X 1 = Conv2d(X, factor = 6, K = 1, P = 0, S = 1)

[0026] X 2 = Conv2d(branch1, factor = 6, K = [1, 3], P = [0, 1], S = 1)

[0027] X 3 = Conv2d(branch2, factor = 6, K = [3, 1], P = [1, 0], S = 1)

[0028] X 4 = Conv2d(branch3, factor = 6, K = [3, 3], P = [1, 1], S = 1)

[0029] Y = Concat(X 1 , X 2 , X 3 , X 4 )

[0030] Among them, Residule represents the residual connection; Concat represents the feature map concatenation; Conv2d represents the two-dimensional convolution; X1, X2, X3, and X4 are the output feature maps of the first to fourth branches respectively; X is the input feature map, Y is the output feature map; K represents the convolution kernel size; S represents the stride; P represents the padding length; D represents the dilation rate; G represents the grouped convolution; factor represents the channel expansion magnification, that is, the multiple of the expansion of the feature map channel number in each convolution layer.

[0031] Furthermore, the adoption of the top-down feature pyramid design scheme to extract multi-scale significant features containing semantic and spatial information includes:

[0032] Feature fusion module: According to the different sizes of the output feature tensors, the backbone network is divided into block 1, block 2, block 3, and block 4; take the features output by the last bottleneck block of block 2, 3, and 4 as {B 2 , B 3 , B 4} and input them into the RFB receptive field module to enhance the receptive field of the effective area without changing the scale features and retain the small ship targets in the feature map to the greatest extent;

[0033] The features after receptive field enhancement will be input into the small feature fusion module in the FSSD network, and the fused features will be used to generate the feature pyramid. The fusion process is calculated by the formula:

[0034] F = [R(Upsample(B 4 ))), R(Upsample(B 3 ), R(B 2 ))

[0035] where F represents the fused features, R represents the operation of the RFB receptive field module, Upsample represents the upsampling operation, and [x 0 , x 2 ,..., x n represents the feature concatenation operation;

[0036] Top-down channel structure: A pyramid structure is adopted to extract multi-scale significant features containing semantic and spatial information to detect multi-scale ships in SAR images;

[0037] The feature pyramid consists of a top-down path, a CA attention module, and dense connections; only the four feature maps with the largest resolution are retained for multi-scale ship detection, which can not only ensure accuracy but also efficiency and reduce the number of parameters;

[0038] In the top-down feature pyramid structure, the network can upsample the high-level feature maps to obtain feature maps with rich semantic information and high resolution;

[0039] The CA attention module is placed between feature maps of different scales to highlight the significant features of a specific ratio, and fuses channel attention and spatial attention to improve the feature representation ability of the input feature maps;

[0040] Channel attention is applied globally, while spatial attention focuses on local information. In this way, at different levels, a large number of features containing more semantic and spatial information are refined to adaptively obtain important features of different scales for subsequent detection;

[0041] Dense connections: A pyramid structure based on dense connections is adopted to obtain more spatial and semantic information, and the feature maps with the same spatial size in the fused feature maps and the top-down path are merged through horizontal connections;

[0042] Integrates the multi-level unambiguous global fusion features in the backbone network and the local significant features of the top-down network. A 3x3 convolutional layer is added to each finally concatenated feature map to mitigate the aliasing effect of upsampling, and at the same time reduce the channel dimension to 256;

[0043] Define the feature map processed by the CA attention module as {A 2 ,A 3 ,A 4}, and define the finally fused feature map set as {D 1 ,D 2 ,D 3 ,D 4}. Then the calculation process of the feature pyramid is obtained by the following two formulas:

[0044]

[0045]

[0046] where F represents the fused feature, A represents the operation of the CA attention module, Upsample represents the upsampling operation, and [x 0 ,x 2 ,...,x n represents the feature concatenation operation;

[0047] Provide the fused feature map to the detection head network for subsequent detection.

[0048] Furthermore, the combination of dense connection and lateral concatenation is used between feature maps of different scales to fuse the non-blurred global features and significant local features, and the CA attention module is added to adaptively refine the feature map to adapt to different scenarios in SAR images, including:

[0049] MobileNetV3 uses a structural combination of lightweight attention modules to build the most effective layer structure, and introduces the h-swish activation function to replace the ReLU6 activation function to save computational costs; the h-wise activation function is calculated by the following formula:

[0050]

[0051] Based on MobileNetV3, some improvements are made, and the SE attention module is replaced by the CA module.

[0052] Advantages of the present invention:

[0053] Based on the FSSD neural network, the present invention proposes the ASSD network to better meet the needs of the SAR ship target recognition task. The main improvements are as follows:

[0054] (1) MobileNetV3 is introduced to build a lightweight backbone structure, and the CA attention module is adopted to improve the inverted residual structure in the backbone. Compared with the original backbone VGG16, the detection speed and detection accuracy of the backbone of the ASSD network are increased by 117% and 5.11% respectively, but the number of model parameters is only about 27.82% of the original.

[0055] (2) The RFB receptive field module is introduced to enhance the feature extraction ability of the backbone, and the accuracy is improved by about 2.52% while ensuring the lightweight of the backbone.

[0056] (3) The top-down feature pyramid design scheme is adopted to extract multi-scale significant features containing semantic and spatial information and effectively increase the spatial and semantic information of the entire network. Compared with the original bottom-up feature pyramid, the new scheme can improve the detection accuracy by about 3.09%.

[0057] Dense connections and lateral splicing are combined between feature maps of different scales to fuse the non-blurred global features and significant local features, and the CA attention module is added to adaptively refine the feature maps to adapt to different scenarios in SAR images. Combining the two connections can further increase the semantic information of the entire network and improve the detection accuracy by about 5.33% on the premise of slightly increasing the model parameters and detection speed.

[0058] In addition to the purposes, features and advantages described above, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the drawings.

[0059] BRIEF DESCRIPTION OF THE DRAWINGS The drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0060] It is the structure diagram of the CA attention module of the present invention;

[0061] Figure 1 It is the functional schematic diagram of the RFB module of the present invention;

[0062] Figure 2 It is two different RFB module architectures of the present invention;

[0063] Figure 3 It is the improved RFB module architecture diagram of the present invention;

[0064] Figure 4 It is the network architecture diagram of the present invention;

[0065] Figure 5 It is the network architecture diagram of the present invention;

[0066] Figure 6 is the scale feature map of the present invention using the attention mechanism map ((a) is the map of using the attention mechanism for the same scale feature map, (b) is the map of using the attention mechanism for different scale feature maps);

[0067] Figure 7 is the residual module and inverted residual module diagram of the present invention;

[0068] Figure 8 is the improved inverted residual structure diagram of the present invention;

[0069] Figure 9 is an example of the SSDD dataset of the present invention;

[0070] Figure 10 is the data augmentation effect diagram used in the present invention;

[0071] Figure 11 is the detection multi-scale target and approximate ship target effect diagram of the present invention respectively. Detailed implementation manners

[0072] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0073] Attention module:

[0074] The attention mechanism can help the object detection model shift its attention to the most important regions of the image and ignore the irrelevant pixel parts, which can effectively improve the performance of the object detection model in computer vision tasks. Aiming at the problem that the FSSD network cannot effectively distinguish ships from complex objects on the shore, the present invention plans to introduce Coordinate Attention to enhance the attention of the FSSD network to position information and channel information.

[0075] Different from most methods of computationally expensive non-local or self-attention modules, the CA (CoordinateAttention) module uses a more effective method to capture position information and channel relationships, and can more effectively enhance the feature representation of lightweight mobile networks. By decomposing the 2D global pooling into two one-dimensional encoding processes, the CA module can perform better than other attention methods with lightweight characteristics (such as SENet, CBAM, and TA).

[0076] A CA module can be regarded as a computational unit, aiming to enhance the expression ability of the network to learn features. It can take any intermediate feature tensor as input and output an enhanced feature transformation tensor Y = {y 1 ,y2 , ..., y C} and the tensor size is the same as X. The process of the CA module encoding channel relationships and long-term dependencies with precise position information can be divided into two steps: coordinated information embedding and coordinated information generation. The structural diagram of the CA module is as shown in Figure 1 as follows:

[0077] Coordinated information embedding: Global pooling is often used in channel attention to globally encode spatial information, but it compresses the global spatial information into a channel descriptor, so it is difficult to retain position information, which is crucial for capturing spatial structures in visual tasks. To enable the attention module to capture long-range interactions spatially with precise position information, the CA module decomposes global pooling into a one-to-one dimensional feature encoding process. Specifically, given the input X, we use two

[0078] kernel functions of the spatial range (H, 1) and (1, W) to encode each channel along the horizontal and vertical coordinates respectively. Therefore, the output of the c-th channel at height h can be expressed by the formula:

[0079]

[0080] Similarly, the output of the c-th channel at width w can be expressed by the formula:

[0081]

[0082] The above two transformations aggregate features along two spatial directions respectively, generating a stack of direction-aware feature maps, and allowing the attention block of the network to capture long-term dependencies along one spatial direction and retain precise position information along the other spatial direction, which helps the network to more accurately locate the object of interest.

[0083] Coordinated attention generation: Formulas (4-1) and (4-2) enable a global receptive field and encode precise position information. To make full use of the resulting expressive features, a second transformation called coordinated attention generation is proposed here. For the aggregated feature maps generated by the above two formulas, we first concatenate them and send them to a shared 1×1 convolutional transformation function F 1 , obtaining the formula:

[0084] f = δ(F 1 ([z h , z w ))

[0085] where [.,.] represents the concatenation operation along the spatial dimension, δ is a non-linear activation function, and is an intermediate feature map that encodes spatial information in the horizontal and vertical directions. Here, r is the reduction ratio used to control the block size, which is very similar to the similar design in the SE module. Then we split f into two independent tensors along the spatial dimension and Two other 1×1 convolutional transforms F h and F w are used to transform f h and f w into tensors with the same number of channels as the input X, which are represented by the following two formulas:

[0086] g h = σ(F h (f h ))

[0087] g w = σ(F w (f w ))

[0088] where σ is the sigmoid function. g h and g w are respectively expanded as attention weights, and the output Y of the CA module is calculated by the formula:

[0089]

[0090] It can be seen that the CA module applies attention in both the horizontal and vertical directions to the input tensor. Each element in the two attention maps reflects whether the object of interest exists in the corresponding row and column. This encoding process allows the CA module to more accurately locate the exact position of the object of interest, thus helping the entire model to better identify.

[0091] Receptive Field Module

[0092] An important reason why SSD and FSSD networks cannot effectively focus on small ship targets is that they accelerate the detection speed at the expense of some accuracy, which is also a drawback of most single-stage detection models. Even some recent works (such as DSSD network and RetinaNet) have significantly improved the detection accuracy of single-stage models, but their performance improvement also owes to very deep network models (such as ResNet-101), which greatly limits the efficiency of single-stage detection models. Therefore, a reasonable alternative is to enhance the feature representation of lightweight networks by introducing some handcrafted mechanisms instead of simply deepening the model.

[0093] Several findings in neuroscience suggest that in the human visual cortex, the size of the population receptive field (pRF) is a function of eccentricity in the retinotopic map. Although it varies across different maps, in each map, it increases with increasing eccentricity. This helps to highlight the importance of the central region and improve the insensitivity to small spatial displacements. Inspired by the structure of RFs in the human system, the RFB (Receptive Field Block) module can enhance the deep features learned from lightweight CNN models, thus contributing to fast and accurate detection. Specifically, the RFB module uses multi-branch pooling and different kernels corresponding to different sizes of RFs, applies dilated convolutional layers to control their eccentricity, and reshapes them to generate the final feature representation, as Figure 2 shown:

[0094] The RFB module is a multi-branch convolutional block, and its internal structure can be divided into two parts: a multi-branch convolutional layer with different kernels and a tail dilated pooling or convolutional layer. The first part is similar to the Inception module and is responsible for simulating pRFs of multiple sizes, while the second part reproduces the relationship between different pRF sizes and eccentricity in the human visual system. The results and functions of the two parts will be elaborated in detail below:

[0095] Multi-branch convolutional layer:

[0096] According to the definition of RF in the CNN network, using different convolutional kernels to implement multi-scale RFs is a simple and natural method, which is superior to sharing RFs of a fixed size. Here, the latest scheme, i.e., the multi-branch convolutional layer in InceptionV4, is used. Specifically, we first use a bottleneck structure (consisting of a 1×1 convolutional layer) in each branch to reduce the number of channels in the feature map and add an n×n convolutional layer. Secondly, we replace the 5×5 convolutional layer with two stacked 3×3 convolutional layers to reduce the parameters and have deeper non-linear layers. For the same reason, the RFB module also uses an n×1 convolutional layer to replace the original n×n convolutional layer. Finally, the shortcut design in ResNet is applied to the end of the RFB module.

[0097] Dilated pooling or convolutional layer: The basic intention of this structure is to generate higher-resolution feature maps, capture information in a larger area and more contexts, while maintaining the same number of parameters. In the RFB module, dilated convolution is used to simulate the effect of eccentricity of pRFs in the human visual cortex. Figure 3 Shows two combinations of the multi-branch convolutional layer and the dilated pooling or convolutional layer. On each branch, a convolutional layer with a specific kernel size is followed by a corresponding dilated pooling or convolutional layer, which has a similar positive correlation with the size and eccentricity of pRFs in the visual cortex. Finally, the feature maps of all branches are concatenated and merged into a spatial pooling or convolutional array.

[0098] In the SAR ship target recognition task, due to the significant differences in the sizes of ship targets in SAR images, small target ships are prone to being confused with speckle noise, resulting in the inability to effectively extract ship features during the feature extraction process. The dilated convolution and residual connection in the RFB structure can enhance the receptive field of the effective region without changing the scale features and maximize the retention of small ship targets in the feature map.

[0099] In order to better improve the detection accuracy of the network model for multi-scale ship targets in SAR images, the present invention makes some improvements to the original RFB module.

[0100] The present invention adds a dilated convolution with a dilation rate of 7 on the basis of the original three branches. By connecting the inputs of the four branches and the residual connection, the features of multi-scale ships can be better extracted by the detection model. At the same time, referring to the improvement in the number of parameters in the original text, the present invention uses 1×3 and 3×1 convolutions to replace the 3×3 convolution in the second branch, 1×5 and 5×1 convolutions to replace the 5×5 convolution in the third branch, and 1×7 and 7×1 convolutions to replace the 7×7 convolution in the fourth branch, generating a lightweight receptive field enhancement structure that is more suitable for multi-scale ship detection in SAR images.

[0101] The specific formula is expressed as follows:

[0102] Add a dilated convolution with a dilation rate of 7:

[0103] Y = Residule(Concat(X1, X2, X3, X4), Conv2d, factor = 6, K = 1, S = 1)

[0104] Use 1×3 and 3×1 convolutions to replace the 3×3 convolution in the second branch, 1×5 and 5×1 convolutions to replace the 5×5 convolution in the third branch, and 1×7 and 7×1 convolutions to replace the 7×7 convolution in the fourth branch:

[0105] X 1 = conv2d(X, factor = 6, K = 1, P = 0, S = 1)

[0106] X 2 = Conv2d(branch1, factor = 6, K = [1, 3], P = [0, 1], S = 1)

[0107] X 3 = Conv2d(branch2, factor = 6, K = [3, 1], P = [1, 0], S = 1)

[0108] X 4= Conv2d(branch3, factor = 6, K = [3, 3], P = [1, 1], S = 1)

[0109] Y = Concat(X 1 , X 2 , X 3 , X 4 )

[0110] Among them, Residule represents the residual connection; Concat represents the feature map concatenation; Conv2d represents the two-dimensional convolution; X1, X2, X3, X4 are the output feature maps of the first to fourth branches respectively; X is the input feature map, Y is the output feature map; K represents the convolution kernel size; S represents the stride; P represents the padding length; D represents the dilation rate; G represents the grouped convolution; factor represents the channel expansion magnification, that is, the multiple by which the number of channels of the feature map expands in each convolution layer.

[0111] The improved RFB module is as Figure 4 shown:

[0112] Design of a top-down feature pyramid based on dense connections.

[0113] The SSD network realizes the detection of multi-scale targets by detecting on multi-scale feature maps respectively. However, since the multi-scale feature maps calculate the feature hierarchy layer by layer and have the characteristics of a feature pyramid, this network feature hierarchy structure generates feature maps with different spatial resolutions, but due to different depths, a large semantic gap is introduced, resulting in insufficient semantic information in the high-resolution feature maps of the shallow layer and damaging its object recognition representation ability. In addition, the SSD network abandons reusing the already calculated layers and only uses the features of the high-level of the backbone network (such as conv4_3 of VGG16) to construct the feature pyramid, which misses the opportunity to reuse the high-resolution feature maps, and this is very important for detecting small-scale targets. The FSSD network fuses the features of multiple layers of the backbone network on the basis of the SSD network, strengthening the semantic information of the entire network to a certain extent and improving the performance in detecting small targets. However, since it still uses the feature pyramid structure of the SSD network, the high-resolution feature maps of the shallow layer still do not get enough semantic information to detect small objects, so the improvement is very limited.

[0114] Therefore, the present invention introduces the idea of the FPN network to construct a top-down feature pyramid and introduces dense connections similar to DenseNet to improve the semantic information of the entire feature pyramid, so as to enhance the recognition accuracy of small ship targets in SAR ship target recognition. The structure diagram of the entire network is as Figure 5 shown.

[0115] Feature Fusion Module: According to the different sizes of the output feature tensors, the backbone network of the present invention is divided into Block 1, Block 2, Block 3, and Block 4. Generally speaking, the output of the deepest layer of each block has the strongest features. Therefore, the present invention takes the features output by the last bottleneck blocks of Blocks 2, 3, and 4 as {B 2 , B 3 , B 4} and inputs them into the RFB receptive field module ( Figure 4-7 denoted by R in

[0116] ). This can enhance the receptive field of the effective area without changing the scale features and retain the small ship targets in the feature map to the greatest extent. The features after receptive field enhancement will be input into the small feature fusion module in the FSSD network, and the fused features are used to generate the feature pyramid. The fusion process is calculated by the formula: 4 ))), R(Upsample(B 3 ), R(B 2 ))

[0117] where F represents the fused features, R represents the operation of the RFB receptive field module, Upsample represents the upsampling operation, and [x 0 , x 2 , …, x n represents the feature concatenation operation.

[0118] Top-down Channel Structure: High-level low-resolution feature maps have a large receptive field, less semantic information, and contain more spatial information, which are suitable for detecting large ships. On the contrary, low-level low-resolution feature maps with a small receptive field, less spatial information but more semantic information are more suitable for detecting small-scale ships. Therefore, effectively fusing semantic information and spatial information is the key to multi-scale ship detection. In addition, the present invention needs to extract significant information at a specific scale from a large amount of features, highlight ship targets at different scales, and eliminate the influence of complex backgrounds, so as to achieve better performance in multi-scale ship detection in various scenarios of SAR images. Inspired by this, the present invention adopts a pyramid structure to extract multi-scale significant features containing semantic and spatial information to detect multi-scale ships in SAR images.

[0119] As shown in Figure 8As shown, the feature pyramid consists of a top-down path, a CA attention module, and dense connections. Compared with the FSSD network and the SSD network, the improved network proposed in the present invention deletes the three feature maps with the lowest resolution (sizes of 1×1, 2×2, and 4×4) because most of the ships in the SSDD dataset are medium and small ships, and the feature maps with low resolution cannot effectively identify these targets, while the feature map with a size of 8×8 is sufficient to identify large ships, and these small-resolution feature maps are slightly redundant. Therefore, the present invention only retains the four feature maps with the largest resolution for multi-scale ship detection, which can not only ensure the accuracy but also ensure the efficiency and reduce the number of parameters.

[0120] In the top-down feature pyramid structure, the network can upsample the high-level feature maps to obtain feature maps with rich semantic information and higher resolution. Since some redundant information needs to be suppressed among a large number of features, while emphasizing the target features, the CA attention module is placed between feature maps of different scales to highlight the significant features of a specific ratio, and fuses channel attention and spatial attention to improve the feature representation ability of the input feature maps. Channel attention is applied globally, and spatial attention focuses on local information. In this way, at different levels, a large number of features containing more semantic information and spatial information are refined to adaptively obtain important features of different scales for subsequent detection.

[0121] The attention module is used for feature maps at the same scale. As shown in (a) of Figure 6 , this method applies the attention module to each feature map separately, which results in the non-target features generated by the higher-level feature maps not being able to be eliminated, leading to the transmission of irrelevant features layer by layer from the higher-level feature maps downwards. Therefore, there may be more false detections and lower detection accuracy. In contrast, the method proposed in the present invention embeds the attention module into the FPN to construct a feature pyramid, as shown in (b) of Figure 6 . In this way, the attention module can refine the feature maps layer by layer, so as to reduce the irrelevant features in the high-level feature maps in the lower-level feature maps through the attention mechanism, which can enable the network to focus on more important features at a specific scale, thus greatly improving the accuracy of multi-scale ship detection.

[0122] Dense Connection: Making full use of features at different levels to obtain more spatial and semantic information is the key to multi-scale ship detection. Therefore, the present invention adopts a pyramid structure based on dense connection to obtain more spatial and semantic information. Compared with other related methods applying dense connection, the dense connection in the present invention adopts a splicing operation instead of simple element addition, which can retain the mapping of multi-scale fusion features to a greater extent. In addition, although related methods use dense connection to fuse feature maps of different scales for multi-scale ship detection, it lacks an adaptive feature selection process at a specific scale. Since some targets are easily overlooked at different scales, this will greatly interfere with the detection results, while the CA attention module is used to weight the feature maps of different scales to highlight the important features at a specific scale. In this way, the feature maps at a specific scale can be more suitable for detecting ships of a specific scale. Then, the fused feature maps and the feature maps with the same spatial size in the top-down path are merged through lateral connection. In this way, it integrates the multi-level and unambiguous global fusion features in the backbone network and the local significant features of the top-down network. Finally, a 3x3 convolutional layer is added to each finally spliced feature map to mitigate the aliasing effect of upsampling, and at the same time reduce the channel dimension to 256. As Figure 5 shown, the feature map processed by the CA attention module is defined as {A 2 , A 3 , A 4}, and the set of finally fused feature maps is defined as {D 1 , D 2 , D 3 , D 4}. Then, the calculation process of the feature pyramid is obtained by the following two formulas:

[0123]

[0124]

[0125] where F represents the fused feature, A represents the operation of the CA attention module, Upsample represents the upsampling operation, and [x 0 , x 2 ,..., x n represents the feature splicing operation.

[0126] The feature pyramid can extract multi-scale fusion feature maps containing rich spatial and semantic information for multi-scale ship detection, and can also extract the unambiguous global features and significant local features contained in the finally fused feature maps to adapt to different scenarios in SAR images. Then, the finally fused feature maps are provided to the detection head network for subsequent detection.

[0127] Lightweight Backbone Network Design

[0128] In the SAR ship target detection task, the original backbone network VGG16 in the FSSD network is too large in scale and number of parameters, and is not very suitable for the subsequent generation of the feature pyramid. The number of feature channels output by the three feature levels to be fused is too large, which makes the lightweight feature fusion module in the FSSD network need to reduce the dimension of the feature with 2048 channels to 512, which will result in a large amount of original features being lost and is not conducive to the backbone network to complete the task of feature extraction. In addition, for the SSDD dataset, the complexity of VGG16 is too high, the number of parameters is too large, and it is very easy to have overfitting problems.

[0129] To improve the network scale problem, the present invention introduces a new lightweight network MobileNetV3 with a new architecture as the new backbone network for feature extraction. The characteristics of MobileNetV3 are as follows:

[0130] Depthwise separable convolution: MobileNetV1 introduced depthwise separable convolution as an effective alternative to the traditional convolutional layer. Depthwise separable convolution effectively decomposes the traditional convolution by separating the spatial filtering and the feature generation mechanism. Depthwise separable convolution is defined by two independent layers. The first layer is called the depthwise convolution, which performs lightweight filtering by applying a single convolutional filter to each input channel. The second layer is a 1×1 convolution, also known as the pointwise convolution, which is responsible for constructing new features by calculating the linear combination of the input channels.

[0131] The input size of a standard convolution is a tensor of h i ×w i ×d i and applies a convolutional kernel to generate an output vector L of size h i ×w i ×d j . Therefore, the computational cost of the standard convolutional layer is h j . i ·w i ·d i ·d j ·k·k.

[0132] Depthwise separable convolution is a substitute for the standard convolutional layer. Empirically, their performance is almost as good as that of the conventional convolution, but the computational cost is only as shown in the formula:

[0133] h i ·w i ·d i (k 2 +d j )

[0134] ​This is the number of depthwise and 1×1 pointwise convolutions. Compared with traditional convolutions, depthwise separable convolutions effectively reduce the computational cost by almost k 2 times. MobileNetV3 uses a kernel coefficient k = 3 (3×3 depthwise separable convolution), so the computational cost is 8 to 9 times smaller than that of standard convolutions, but the accuracy is only slightly reduced.

[0135] Linear bottleneck layer: MobileNetV2 introduced linear bottleneck layers and inverted residual structures to utilize the low-order characteristics of the problem to form a more efficient layer structure. For the depthwise separable convolutions of MobileNetV1, the M-dimensional space compressed by the width multiplier passes through a non-linear activation function ReLU. According to the properties of ReLU, if the input feature is positive, the output feature after passing through the activation layer is the restored input value, which is equivalent to a linear transformation. However, if the input feature is negative, the feature of this channel will be zeroed, and the consequence is that the already compressed features will further lose feature information. But if there are many channels, this loss can be compensated by using the activation values of other channels. Therefore, to avoid losing too much information, the paper's authors adopted linear bottleneck layers, that is, non-linear activation functions are not used in bottleneck layers with relatively few channels, which can prevent non-linear activation functions from destroying too much information 1 .

[0136] Inverted residual structure: The bottleneck layer actually contains all the necessary information, while the expansion layer is actually only an implementation detail for the non-linear change of the adjoint tensor. Inspired by this idea, the paper's authors directly used a residual structure between the bottlenecks.

[0137] The traditional residual structure first uses a 1×1 convolution to reduce the number of channels, then performs a normal convolution operation, and finally uses a 1×1 convolution to restore the number of channels to the original size. The residual connection is established between two layers with relatively many channels, and each layer uses the ReLU activation function, as Figure 7 shown in the left figure;

[0138] The inverted residual structure first uses a 1×1 convolution to increase the number of channels, then performs a depthwise separable convolution operation, and finally uses a 1×1 convolution to reduce the number of channels to the original size. The residual connection is established between two bottleneck layers with relatively few channels. And because the number of channels in the middle two layers is relatively large, an activation function is used, and the information loss is small; while the number of channels at both ends is relatively small, and the loss will be large if an activation function is used, so the ReLU activation function is not used. This structure expands to a higher-dimensional feature space inside while remaining compact at the input and output ends to increase the expressive power of non-linear channel transformations, as Figure 7 shown in the right figure;

[0139] Lightweight Attention Module: MnasNet is based on the MobileNetV2 structure, and a lightweight attention module, the SE module, is introduced in the bottleneck structure. This module is placed after the depthwise separable convolution to apply attention to the maximum feature representation.

[0140] MobileNetV3 uses a combination of the above structures to build the most efficient layer structure, and introduces the h-swish activation function to replace the ReLU6 activation function to save computational costs. The h-wise activation function is calculated by formula (4-8):

[0141]

[0142] In addition, in order to make MobileNetV3 more suitable for the task of SAR ship target recognition, the present invention makes some improvements on the basis of MobileNetV3, and replaces the SE attention module with the CA module. Because the SE attention module only considers encoding the information between channels and ignores the importance of location information, while location information is crucial for capturing object structures in visual tasks. The improved inverted residual structure is as Figure 8 shown:

[0143] In the selection of the backbone network, the present invention selects the MobileNetV3-Large version, which can better meet the requirements of SAR ship target recognition in terms of the number of parameters and model complexity, and truncates the remaining layers after the last bottleneck layer. The specific backbone network structure is shown in Table 1:

[0144] Table 1. MobileNetV3 Network Structure Table

[0145]

[0146] Among them, CA indicates whether there is a CA attention module in this block. NL indicates the type of non-linear activation function used, HS indicates the h-swish activation function, RE indicates the ReLU activation function, and s represents the stride size.

[0147] In the actual backbone network design, the present invention divides the backbone network into Block 1, Block 2, Block 3, and Block 4 with the bottleneck blocks with output feature tensor sizes of 64×64, 32×32, and 16×16 as boundaries respectively.

[0148] In terms of data augmentation, considering that the SSDD dataset itself does not contain enough samples, Figure 9Some examples of the SSDD dataset are presented. This sample only contains a limited number of targets with different scales, including inshore, offshore, and samples with approximately confused targets. To enhance the diversity of training data and improve the robustness of model training, more diverse data augmentation methods, such as mosaic data augmentation, can be introduced during network training in the future. The data augmentation effects used in this project are as Figure 10 shown. For the same sample, we use methods such as rotation, scaling, and cropping to increase the number of typical samples so that the model can learn more robust features. In the design of the loss function, the loss function designed by the SSD network is still used in this invention. However, for the SAR ship target detection task, there are many difficult-to-distinguish samples in the dataset. Therefore, the focal loss function may achieve better training results because it focuses on adding weights to the losses corresponding to samples according to the difficulty of sample discrimination, and can better handle the problem of unbalanced sample classification.

[0149] This project has achieved good results in detecting multi-scale targets and approximately separated ship targets, as Figure 11 shown. The targets outlined by the square rectangles in the figure are all ship targets recognized by the model. It can be seen that the model used in this project can still maintain a very high accuracy when dealing with a large number of dense ship targets. For inshore targets, this model can also well distinguish similar-shaped confused targets and accurately identify inshore ship targets. The experimental results on the SAR ship detection dataset (SSDD) show that this method can achieve a detection speed of 171.86 FPS and a detection accuracy of 90.28% mAP, while the number of model parameters is only 8.89M, taking into account lightweight and detection speed while maintaining extremely high detection accuracy.

[0150] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for ship target recognition in SAR images based on the attention mechanism, characterized in that, it includes: using the CA attention module and the RFB receptive field module to enhance the performance of the MobileNetV3 backbone network; adopting a top-down feature pyramid design scheme to extract multi-scale salient features containing semantic and spatial information; combining dense connection and lateral splicing between feature maps of different scales to fuse non-blurred global features and salient local features, and adding a CA attention module to adaptively refine the feature maps to adapt to different scenarios in SAR images; On the basis of the original three branches, a dilated convolution with a dilation rate of 7 is added. By connecting the inputs of the four branches and using residual connections, the features of multi-scale ships can be better extracted by the detection model; respectively use and convolutions to replace the convolution in the second branch, and convolutions to replace the convolution in the third branch, and convolutions to replace the convolution in the fourth branch, generating a lightweight receptive field enhancement structure that is more suitable for multi-scale ship detection in SAR images.

2. The method for ship target recognition in SAR images based on the attention mechanism according to claim 1, characterized in that, the using the CA attention module and the RFB receptive field module to enhance the performance of the MobileNetV3 backbone network includes: adding a dilated convolution with a dilation rate of 7: , using 1×3 and 3×1 convolutions to replace the 3×3 convolution in the second branch, 1×5 and 5×1 convolutions to replace the 5×5 convolution in the third branch, and 1×7 and 7×1 convolutions to replace the 7×7 convolution in the fourth branch: , Among them, represents a residual connection; represents feature map concatenation; represents a two-dimensional convolution; , , , are the output feature maps of the first to fourth branches respectively; is the input feature map, is the output feature map; represents the convolution kernel size; represents the stride; represents the padding length; represents the dilation rate; represents grouped convolution; represents the channel expansion ratio, that is, the multiple by which the number of channels of the feature map expands at each convolutional layer.

3. The method for ship target recognition in SAR images based on the attention mechanism according to claim 1, characterized in that, the adopting a top-down feature pyramid design scheme to extract multi-scale salient features containing semantic and spatial information includes: Feature Fusion Module: According to the different sizes of the output feature tensors, the backbone network is divided into Block 1, Block 2, Block 3, and Block 4; the features output by the last bottleneck blocks of Blocks 2, 3, and 4 are taken as inputs to the RFB receptive field module, which enhances the receptive field of the effective region without changing the scale features and maximally retains small ship targets in the feature map; the features after receptive field enhancement will be input into the small feature fusion module in the FSSD network, and the fused features are used to generate the feature pyramid, and the fusion process is calculated by the formula: , Among them The features after fusion Represents the operation of the RFB receptive field module Represents the upsampling operation Represents the feature concatenation operation Top-down channel structure: adopting a pyramid structure to extract multi-scale salient features containing semantic and spatial information to detect multi-scale ships in SAR images; The feature pyramid is composed of a top-down path, a CA attention module, and dense connections; only the four feature maps with the largest resolution are retained for multi-scale ship detection, which can not only ensure accuracy but also ensure efficiency and reduce the number of parameters; In the top-down feature pyramid structure, the network can upsample the high-level feature map to obtain a feature map with rich semantic information and high resolution; The CA attention module is placed between feature maps of different scales to highlight salient features of a specific ratio, and fuses channel attention and spatial attention to improve the feature representation ability of the input feature map; Channel attention is applied globally, while spatial attention focuses on local information. In this way, at different levels, a large amount of features containing more semantic and spatial information are refined to adaptively obtain important features of different scales for subsequent detection; Dense connection: adopting a pyramid structure based on dense connection to obtain more spatial and semantic information, and horizontally connecting and merging the fused feature map and the feature map with the same spatial size in the top-down path; integrating the multi-level non-blurred global fusion features in the backbone network and the local salient features of the top-down network, adding a 3x3 convolutional layer to each finally stitched feature map to mitigate the aliasing effect of upsampling, and at the same time reducing the channel dimension to 256; Define the feature map processed by the CA attention module as , and define the set of finally fused feature maps as . Then the calculation process of the feature pyramid is obtained by the following two formulas: , Among them represents the fused feature, represents the operation of the CA attention module, represents the upsampling operation, represents the feature concatenation operation; Provide the fused feature map to the detection head network for subsequent detection.

4. The method for identifying ship targets in SAR images based on the attention mechanism according to claim 1, wherein, combining the use of dense connection and lateral stitching between feature maps of different scales to fuse the non-blurred global features and the significant local features, and adding a CA attention module to adaptively refine the feature map to adapt to different scenarios in SAR images includes: MobileNetV3 uses a structural combination of lightweight attention modules to construct the most effective layer structure, and introduces the h-swish activation function to replace the ReLU6 activation function to save computational costs; the h-wise activation function is calculated by the following formula: 。

Citation Information

Patent Citations

  • Attention pyramid network-based SAR image multi-scale ship detection method

    CN110084210A

  • SAR image small sample target identification method based on multi-scale network

    CN115995040A