A High-Aspect-Ratio-Aware SAR Ship Target Detection Method
By using a variety of high-period convolution kernels to extract features and using an adaptive feature fusion network to fusion these features, the difficulties of ship target detection in SAR images are solved, and high-precision ship target detection and low false alarm rate are achieved.
Patent Information
- Application Number
- CN202210666346.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-06-13
AI Technical Summary
Ship target detection faces many difficulties in SAR images, including coherent spot noise interference, unknown differences in ship size, resulting in few features, complex background, false alarms, and difficult to utilize the large aspect ratio characteristics of ships.
A variety of aspect ratio convolution kernels are used to extract the aspect ratio features, and the adaptive feature fusion network is fused with different aspect ratio features to construct a SAR ship target detection network model with aspect ratio perception.
It effectively improves the accuracy of SAR ship target detection, reduces the false alarm rate, and makes full use of the ship's large aspect ratio feature information.
Smart Images

Figure CN115131652B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of SAR signal processing, and particularly to a method for detecting SAR ship targets with aspect ratio perception. Background Art
[0002] The detection of ship targets on the sea surface can be used for tasks such as marine supervision, naval battle situation monitoring, intelligence reconnaissance, etc., and is of great significance for safeguarding national security and interests. China has a long coastline and vast sea areas, and there are many threatening objects in the sea areas. Therefore, it is the key to maintaining national security and interests to master the dynamics of hot sea areas in real time, quickly identify and obtain suspicious target information from complex backgrounds, and achieve early warning, early layout, and seizing the initiative.
[0003] Compared with visible light imaging, synthetic aperture radar (SAR) has the characteristics of all-day, all-weather, and large swath width, enabling it to image vast sea areas even in complex environments, and it is an important information source for marine monitoring and marine intelligence extraction. Therefore, it is of great significance to study the rapid and accurate detection of ships from wide-swath SAR images.
[0004] However, the detection of ship targets in SAR images faces many difficulties:
[0005] 1) The special imaging mechanism of the SAR system causes the SAR image to be interfered by special speckle noise. The cross stars caused by the strong sidelobes of ship targets may not only cause repeated detections, but also may cover the surrounding small ship targets, resulting in missed detections.
[0006] 2) The SAR image has a large swath width and may contain a large number of ships of different types. The biggest difference among these ships in the SAR image is their different sizes. Small ships are less than 20 meters in length and only occupy 4 pixel points in a 5-meter resolution SAR image, resulting in few features and being easily submerged by noise and interference. Large ships are prone to breakage in the SAR image, resulting in a single ship being detected as multiple ships.
[0007] 3) The background near the coastline and ports is complex, with many strong scattering objects, which are extremely prone to false alarms. SAR ships docked at the shore are easily misjudged as land, resulting in missed detections.
[0008] 4) SAR ship targets usually have a large aspect ratio and cannot directly apply existing detection models. It is necessary to design a targeted network structure model according to their characteristics. Summary of the Invention
[0009] Aiming at the problems existing in the prior art, the present invention provides a method for extracting various aspect ratio features using various aspect ratio convolutional kernels; adopting an adaptive feature fusion network to fuse different aspect ratio features, making full use of the information of each aspect ratio feature, improving the model detection accuracy, having a low false alarm rate, and being able to provide core technical support for SAR-based marine monitoring and intelligence acquisition, namely an aspect ratio-aware SAR ship target detection method.
[0010] The purpose of the present invention is achieved through the following technical solutions.
[0011] An aspect ratio-aware SAR ship target detection method includes the following steps:
[0012] 1) Construction of ship detection network model: Construct an aspect ratio-aware ship detection network model;
[0013] 2) Design of loss function: Design the model loss function to measure the quality of model prediction and formulate the model learning criterion;
[0014] 3) Model training: Use the gradient descent method to train the model and dynamically adjust the learning rate during the training process;
[0015] 4) SAR ship target detection: Divide the wide-swath SAR image into blocks, send the divided images into the trained ship detection network model, perform non-maximum suppression processing on the model output results to obtain the final ship detection results, and then fuse and splice the detection results to obtain the final wide-swath SAR image ship target detection results.
[0016] An aspect ratio-aware SAR ship target detection network model is constructed using the pyramid feature extraction network of YOLOv3, which outputs convolutional features at three scales to predict ship targets of three different sizes. The multi-scale convolutional features output by the pyramid network are input into the aspect ratio feature extraction module to extract target aspect ratio features of different aspect ratios, and the aspect ratio feature fusion module is used to adaptively weight the aspect ratio features.
[0017] The aspect ratio feature extraction module uses three types of convolutional kernels, namely 1×3, 1×5, and 1×7, to extract ship features with an aspect ratio less than or equal to 0.5, uses three types of convolutional kernels, namely 3×1, 5×1, and 7×1, to extract ship features with an aspect ratio greater than 2, and uses three types of convolutional kernels, namely 1×5, 5×1, and 1×1, to extract ship features with an aspect ratio greater than 0.5 and less than or equal to 2. The extracted features are fused through concatenation respectively. Behind all convolutional layers in the aspect ratio feature extraction module, there are immediately followed by a batch normalization layer and a RELU activation layer.
[0018] The aspect ratio feature fusion module performs weighted fusion on the extracted features with different aspect ratios, adaptively selects the optimal aspect ratio features for ship detection; each aspect ratio feature passes through a 1×1 convolutional layer, then the output features are concatenated, and then passed through a convolutional layer with three 1×1 convolutional kernels. After passing through the softmax layer, three weight maps are generated as the weights of the three aspect ratio features, and the final output features are obtained by weighting the three aspect ratio features.
[0019] In the design of the loss function, the anchor boxes adopt the anchor boxes set by YOLOv3 on the coco dataset. Five outputs are obtained for each anchor box, which are , , , and , corresponding to the x and y coordinates, width, height, and confidence of the regression box respectively. The five outputs are transformed as follows:
[0020]
[0021]
[0022]
[0023]
[0024] Among them, is the sigmoid function. Assume that is the coordinate variable corresponding to the ground truth box, and its calculation method is as follows:
[0025]
[0026]
[0027]
[0028] Among them, and are the width and height of the anchor box, and are the coordinates of the anchor box projected on the feature map x , y axis, w and h are the width and height of the target on the input image, s is the total number of stride steps of the current output feature map relative to the input image. The model loss function is:
[0029]
[0030] Among them, , , , are the coordinate variables corresponding to the truth box, i is the serial number of the grid, is the number of grids, j is the serial number of the anchor box. If the j-th anchor box in the i-th grid has the largest overlap with a certain truth box, then otherwise . If the overlap rate between the j-th anchor box in the i-th grid and all truth boxes is less than 0.7, then otherwise .
[0031] During the detection process, each anchor box obtains 5 outputs, which are respectively , , , and , and the converted center coordinates of the rectangular box are:
[0032]
[0033]
[0034]
[0035]
[0036] Among them, and are the width and height of the anchor box, and are the coordinates of the projection of the anchor box on the feature map x , y axis, s is the total number of stride steps of the current output feature map relative to the input image.
[0037] Compared with the prior art, the advantages of the present invention are as follows: The present invention realizes a high-precision SAR ship target detection method, providing core technical support for SAR-based marine monitoring and intelligence acquisition. Compared with the existing SAR ship target detection methods, its significant advantages are: designing a ship detection network model with aspect ratio perception, making full use of the large aspect ratio characteristics of ships, effectively improving the ship detection rate while reducing the false alarm rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Overall network model structure diagram of the method.
[0039] Figure 2 Network structure diagram of the aspect ratio feature extraction module.
[0040] Figure 3 Network structure diagram of the aspect ratio feature fusion module
[0041] Figure 4 Effect diagram of ship detection in SAR images based on this method Specific implementation manners
[0042] The present invention will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments
[0043] The aspect ratio-aware SAR ship target detection method includes the following steps
[0044] (10) Construction of the ship detection network model: Construct an aspect ratio-aware ship detection network model
[0045] (20) Design of the loss function: Design the model loss function to measure the quality of the model prediction and formulate the model learning criterion
[0046] (30) Model training: Use the gradient descent method to train the model and dynamically adjust the learning rate during the training process
[0047] (40) SAR ship target detection: Divide the wide-swath SAR image into blocks, send the divided images into the trained ship detection network model, perform non-maximum suppression processing on the model output results to obtain the final ship detection results, and then fuse and splice the detection results to obtain the final wide-swath SAR image ship target detection results
[0048] An aspect ratio-aware SAR ship target detection network model is constructed. The model uses the pyramid feature extraction network of YOLOv3 to output convolutional features at three scales and predict ship targets of three different sizes. In order to make full use of the ship aspect ratio information, the multi-scale convolutional features output by the pyramid network are input into the aspect ratio feature extraction module to extract the aspect ratio features of different targets. Then, the aspect ratio feature fusion module adaptively weights the aspect ratio features to realize the utilization of the target aspect ratio feature information. The overall network model structure is as Figure 1 shown
[0049] An aspect ratio feature extraction module is constructed. This module uses three types of convolutional kernels of 1×3, 1×5, and 1×7 to extract ship features with an aspect ratio less than or equal to 0.5, uses three types of convolutional kernels of 3×1, 5×1, and 7×1 to extract ship features with an aspect ratio greater than 2, and uses three types of convolutional kernels of 1×5, 5×1, and 1×1 to extract ship features with an aspect ratio greater than 0.5 and less than or equal to 2. The extracted features are fused through concatenation. After all convolutional layers in the aspect ratio feature extraction module, there are immediately batch normalization layers and RELU activation layers. The network structure of the aspect ratio feature extraction module is as Figure 2 shown
[0050] A height-width ratio feature fusion module is constructed. This module weights and fuses the extracted features with different height-width ratios, and adaptively selects the best aspect ratio features for ship detection. Each height-width ratio feature passes through a 1×1 convolutional layer (followed by a batch normalization layer and a RELU activation layer), then the output features are concatenated, and then passed through a convolutional layer with three 1×1 convolutional kernels. Immediately afterwards, a softmax layer is used to generate three weight maps as the weights of the three height-width ratio features, and the final output feature is obtained by weighting the three height-width ratio features. The network structure of the height-width ratio feature fusion module is as Figure 3 shown. Embodiment
[0051] Airborne / spaceborne SAR ship targets have the following characteristics: 1) Ships are presented in a top-down view in SAR images; 2) Ships usually have a large aspect ratio and will present a larger or smaller height-width ratio shape in the image. Existing SAR ship target detection methods do not make full use of the large aspect ratio information of ships, resulting in the inability to further improve the detection accuracy. In order to make full use of the large aspect ratio information of ships, this method proposes to extract ship features with multiple height-width ratios using multiple height-width ratio convolutional kernels, and adopt an adaptive feature fusion network to fuse features with different height-width ratios. Finally, ships with different height-width ratios are output through different height-width ratio feature fusion modules. The specific implementation steps are as follows:
[0052] (10) Construction of the ship detection network model: This method uses the pyramid feature extraction network of YOLOv3 to output convolutional features at three scales and predict ship targets of three different sizes. In order to make full use of the ship aspect ratio information, the multi-scale convolutional features output by the YOLOv3 network are input into the height-width ratio feature extraction module, as Figure 1 shown. The height-width ratio feature extraction module uses three convolutional kernels of 1×3, 1×5, and 1×7 to extract ship features with a height-width ratio less than or equal to 0.5, uses three convolutional kernels of 3×1, 5×1, and 7×1 to extract ship features with a height-width ratio greater than 2, and uses three convolutional kernels of 1×5, 5×1, and 1×1 to extract ship features with an aspect ratio greater than 0.5 and less than or equal to 2. The extracted features are fused by concatenation respectively, as Figure 2 shown, where the content (m×n, f) in the brackets represents the convolution sum size of m×n and the number of convolutional kernels is f. All convolutional layers in the height-width ratio feature extraction module are immediately followed by a batch normalization layer and a RELU activation layer. The height-width ratio feature fusion module weights and fuses the extracted features with different height-width ratios, and adaptively selects the best aspect ratio features for ship detection. Its network structure is as Figure 3As shown. Each aspect ratio feature passes through a 1×1 convolution layer (the convolution layer is followed by a batch normalization layer and a Relu activation layer), and then the output features are concatenated, and then passed through a convolution layer with 3 1×1 convolution kernels, followed by a softmax layer to generate 3 weight maps as the weights of the three aspect ratio features. The final output feature is obtained by weighting the three aspect ratio features. The prediction part uses the loss function and post-processing method of YOLOv3, and uses different prediction branches to predict the three different aspect ratio ships with aspect ratios (AR) less than or equal to 0.5, greater than 0.5 and less than or equal to 2, and greater than 2, as shown below. Figure 1 shown.
[0053] (20) Loss function design: Since we only detect ships and do not classify them, there is no category-related loss. The anchor box uses the anchor box set by YOLOv3 on the coco dataset. The model obtains 5 outputs for each anchor box, which are , , , and , corresponding to the x, y coordinates, width, height and confidence of the regression box. The 5 outputs are transformed as follows:
[0054] (1)
[0055] (2)
[0056] (3)
[0057] (4)
[0058] in, is the sigmoid function. Assume is the coordinate variable corresponding to the true value frame, which is calculated as follows:
[0059] (5)
[0060] (6)
[0061] (7)
[0062] (8)
[0063] in, and is the width and height of the anchor box, and Projection of anchor boxes on feature maps x, y The coordinates on the axis, w and h are the width and height of the target on the input image, s and is the total number of strides (stride) of the current output feature map relative to the input image. The model loss function is:
[0064] (9)
[0065] Among them, , , , are the coordinate variables corresponding to the ground truth box, i is the serial number of the grid (cell), is the number of grids, and j is the serial number of the anchor box. If the j-th anchor box in the i-th grid has the largest overlap with a certain ground truth box, then , otherwise . If the overlap rate of the j-th anchor box in the i-th grid with all ground truth boxes is less than 0.7, then , otherwise .
[0066] (30) Model training: Use a self-made dataset and train the model using the Stochastic Gradient Descent (SGD) method. The batch size is set to 16, and the input image size is 416×416. The model is trained for a total of 300 epochs. The initial learning rate is set to 0.001, and the learning rate is multiplied by 0.5 every 50 epochs. During all training processes, the weight decay is set to 0.0005, and the momentum is set to 0.9.
[0067] (40) SAR ship target detection: Divide the wide SAR into blocks, send them into the trained convolutional neural network for ship detection, then splice the detection results, and perform non-maximum suppression to obtain the final result. The ship detection results are as Figure 4 shown. During the detection process, each anchor box gets 5 outputs, which are respectively , , , and , and the converted rectangular box center coordinates are:
[0068] (10)
[0069] (11)
[0070] (12)
[0071] (13)
[0072] Among them, and are the width and height of the anchor box, and are the coordinates of the projection of the anchor box on the feature map x , y on the s axis, and is the total number of strides (stride) of the current output feature map relative to the input image.
Claims
1. A high aspect ratio perception-based SAR ship target detection method, characterized in that It includes the following steps: 1) Construction of a ship detection network model: Construct a ship detection network model with aspect ratio perception; 2) Design of the loss function: Design the model loss function to measure the quality of the model prediction and formulate the model learning criterion; 3) Model training: Use the gradient descent method to train the model and dynamically adjust the learning rate during the training process; 4) SAR ship target detection: The wide-swath SAR image is segmented, and the segmented images are fed into the trained ship detection network model. The output results of the model are processed by non-maximum suppression to obtain the final ship detection results. Then, the detection results are fused and stitched to obtain the final wide-swath SAR image ship target detection results. In the design of the loss function, the anchor boxes are the anchor boxes set by YOLOv3 on the coco dataset. Five outputs are obtained for each anchor box, which are respectively , , , and , corresponding to the x and y coordinates, width, height, and confidence of the regression box respectively. The five outputs are converted as follows: Among them, is the sigmoid function. Assume that is the coordinate variable corresponding to the ground truth box, and its calculation method is as follows: wherein, and are the width and height of the anchor box, and are the coordinates of the projection of the anchor box on the feature map x , y on the w axis, h and s are the width and height of the target on the input image, s is the total number of stride steps of the current output feature map relative to the input image, and the model loss function is: Among them, , , , are the coordinate variables corresponding to the ground truth boxes, i is the serial number of the grid, is the number of grids, j is the serial number of the anchor box. If the j-th anchor box in the i-th grid has the maximum overlap with a certain ground truth box, then , otherwise, if the overlap rate of the j-th anchor box in the i-th grid with all ground truth boxes is less than 0.7, then , otherwise .
2. The high - aspect - ratio perception - based SAR ship target detection method according to claim 1, wherein, A ship target detection network model with aspect ratio perception is constructed. It adopts the pyramid feature extraction network of YOLOv3, outputs convolutional features at three scales, predicts ship targets of three different sizes, inputs the multi-scale convolutional features output by the pyramid network into the aspect ratio feature extraction module to extract the aspect ratio features of different targets, and uses the aspect ratio feature fusion module to adaptively weight the aspect ratio features.
3. The high-aspect-ratio perception-based SAR ship target detection method according to claim 2, wherein The aspect ratio feature extraction module uses three convolution kernels of 1×3, 1×5, and 1×7 to extract ship features with an aspect ratio less than or equal to 0.5, uses three convolution kernels of 3×1, 5×1, and 7×1 to extract ship features with an aspect ratio greater than 2, and uses three convolution kernels of 1×5, 5×1, and 1×1 to extract ship features with an aspect ratio greater than 0.5 and less than or equal to 2. The extracted features are fused through concatenation respectively. Behind all convolutional layers in the aspect ratio feature extraction module, there are immediately followed by a batch normalization layer and a RELU activation layer.
4. The aspect ratio-aware SAR ship target detection method according to claim 2, wherein, The aspect ratio feature fusion module weights and fuses the extracted aspect ratio features of different aspect ratios, and adaptively selects the best aspect ratio feature for ship detection; Each aspect ratio feature passes through a 1×1 convolutional layer, then the output features are concatenated, and then through a convolutional layer with three 1×1 convolutional kernels. After passing through the softmax layer, three weight maps are generated as the weights of the three aspect ratio features, and the final output feature is obtained by weighting the three aspect ratio features.
5. The high-aspect-ratio perception-based SAR ship target detection method according to claim 1, wherein During the detection process, each anchor box gets 5 outputs, which are respectively , , , and , and the center coordinates of the rectangular box are converted to: Among them, and are the width and height of the anchor box, and are the coordinates of the projection of the anchor box on the feature map x , y axis, s is the total number of strides of the current output feature map relative to the input image.