SAR Ship Target Detection Method Based on Mask Network Fusion of Image Features
By constructing mask network branching and adaptive semantic segmentation label generation strategy, combined with weight allocation loss function, the false alarm and missing alarm problems of ship target detection in SAR images are solved, and the detection accuracy and network deployment efficiency are improved.
Patent Information
- Application Number
- CN202211567684.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-07
AI Technical Summary
The existing SAR image object detection methods have problems such as many false alarms and missing alarms and low detection accuracy in complex scenarios, and are highly dependent on semantic segmentation mask labels, making it difficult to deploy.
A mask network branch is constructed, and an adaptive semantic segmentation mask label is generated using the luminance gradient difference. Combining the mask feature fusion sub-network and the target detection network, a cross-entropy loss function for adaptive weight allocation is designed, and a mask feature fusion branch is added to highlight the ship's target characteristics and suppress background interference.
It improves the detection accuracy of ship targets in SAR images, reduces false alarms, improves detection accuracy, and effectively combines target detection and segmentation tasks in the absence of semantic segmentation labels.
Smart Images

Figure CN115965862B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to a method for detecting ship targets by synthetic aperture radar (SAR) based on mask network fusion in the technical field of SAR image target detection. The present invention can be applied to detect ship targets in SAR images. Background Art
[0002] Synthetic Aperture Radar (SAR) is an imaging radar with high resolution. Compared with optical detection means such as visible light imaging, infrared detection, and laser detection, SAR is not restricted by natural conditions such as clouds and fog and can detect stationary time-sensitive targets such as ships at a long distance. As the primary stage of SAR automatic target recognition, SAR image target detection has received extensive attention. Constant False Alarm Rate (CFAR) is the most widely used and deeply studied traditional SAR target detection method. This type of method uses background information by modeling the statistical distribution of background clutter, and then obtains an adaptive threshold. Then, the gray value of the pixel is compared with the adaptive threshold through a sliding window to obtain the detection result. Therefore, determining a suitable clutter statistical model is very important for ensuring the detection performance of CFAR. However, it is difficult to select a suitable clutter statistical model for measured SAR images because they contain a large amount of complex background clutter, resulting in a decline in detection performance. With the development of deep learning, many methods based on convolutional neural networks have been proposed. Since a large amount of labeled training data is learned by the network, these methods have made significant progress in target detection. However, due to the complex scene of measured SAR images, the SAR target detection method based on CNN still has many problems of false alarms and missed alarms, and the detection accuracy urgently needs to be improved.
[0003] The Fourteenth Research Institute of China Electronics Technology Group Corporation proposed a SAR ship target detection method based on background and scale perception in its patent document "A SAR Ship Target Detection Method Based on Background and Scale Perception" (Patent Application No.: CN 202111613298.X, Publication No.: CN114219997 A). This method effectively reduces the false alarm rate by designing a ship detection network model with background perception. By designing a scale perception loss function, the model pays more attention to small targets during the training process, improving the detection rate of small targets. At the same time, a multi-scale training strategy is adopted, where the size of the image input to the network is fixed each time while ensuring multi-scale training, so that the time required for each iteration and the required hardware resources remain consistent. However, the deficiencies of this method are still as follows: The multi-scale training strategy introduces multiple downsampling operations into the network, and small target features are lost during each downsampling process. Although the network designs a scale perception loss function to avoid the loss of small target features as much as possible, the small target information lost during the sampling process is still difficult to fully recover, resulting in a decrease in the detection accuracy of the network. In addition, while paying attention to scale changes, it also introduces more noise into the network, which is prone to false alarms and ultimately leads to a decline in the network monitoring effect.
[0004] Xidian University proposed a SAR ship target detection method based on balanced sample regression loss in its patent document "SAR Ship Target Detection Method Based on Balanced Sample Regression Loss" (Patent Application No.: CN 202011544100.2, Publication No.: CN 112668440A). This method selects the Faster-RCNN network as the training network model, improves the original loss function of the training network, constructs a new total loss function to train the network, and obtains the finally trained network model, improving the detection accuracy of the network for ship targets. However, the deficiencies of this method are still as follows: When extracting features from the input SAR image, this method processes the entire scene without discrimination. When the target and the background are not significantly different and the scene contains a large amount of complex background clutter, it will cause a large number of false alarms and missed alarms, resulting in a decrease in detection accuracy. In addition, this method uses a two-stage network, which has a large number of parameters and occupies a large amount of storage space, making it difficult to deploy the network and the model.
[0005] Zhang proposed a SAR ship target detection method based on the Mask Attention Interaction and Scale Enhancement Network (MAI-SE-Net) in his published paper "A Mask Attention Interaction and Scale Enhancement Network for SAR Ship Instance Segmentation" (IEEE Geoscience and Remote Sensing Letters (Volume: 19) 2021). This method models long-range spatial dependencies with non-local blocks, uses content-aware feature recombination blocks to generate additional pyramid bottoms to improve the performance of small ships, uses feature balance operations to improve scale feature description, and uses global context blocks to refine features. Through the above operations, MAI-SE-Net completed the semantic segmentation and target detection tasks and achieved good results. However, the deficiencies of this method are as follows: This method requires a dataset with both target detection bounding boxes and semantic segmentation mask labels, and has high requirements for the dataset. In addition, the network used in this method is a two-stage network, and the number of parameters and the model storage space are much larger than those of a one-stage network, making it difficult to deploy the network and the model. Summary of the Invention
[0006] The purpose of the present invention is to propose a SAR target detection method based on mask network fusion in view of the above deficiencies of the prior art. The aim is to solve the problems in the prior art that the network processes targets and backgrounds without discrimination, resulting in many false alarms and missed alarms and low detection accuracy when detecting complex scene SAR images, and the network highly depends on semantic segmentation mask labels.
[0007] The technical idea for achieving the object of the present invention is as follows: The present invention constructs a mask network branch, which is used to highlight the features of ship targets while suppressing background feature information. The features extracted by the trained mask network branch are fused with the features extracted by the backbone network to perform target detection on SAR images, avoiding the problems of high false alarms and low detection accuracy caused by the undifferentiated processing of the entire scene in the prior art. The present invention proposes an adaptive semantic segmentation mask label making strategy. For different SAR data sets, the present invention can adaptively complete the mask label making, combining the target detection and target segmentation tasks even in the case of lack of semantic segmentation labels, and solving the problem of the high dependence of the network on semantic segmentation mask labels in the prior art. The present invention designs a cross-entropy loss function with weight coefficients for the mask branch network, focusing the attention of the branch on ship targets and solving the problem of foreground-background imbalance. The present invention adds an adaptive weight allocation mechanism to the loss function of the target detection network, enabling the network to adaptively and reasonably allocate weights for different tasks during the training process and solving the problem of unstable network training process.
[0008] The specific steps of the present invention are as follows:
[0009] Step 1, generating a sample set:
[0010] Step 1.1, collecting at least 900 SAR ship image samples with a scale larger than 320×320 pixels, and each sample contains at least one ship target;
[0011] Step 1.2, annotating a target detection label file for each sample, and each target detection label file contains the bounding box coordinates of all ship targets in the corresponding sample;
[0012] Step 1.3, forming a sample set by all SAR ship image samples and their corresponding target detection label files;
[0013] Step 2, generating an adaptive SAR image ship semantic segmentation label by using the brightness gradient difference between ships and the background:
[0014] Step 2.1, magnifying the coordinates of each bounding box in the target detection label of each sample in the sample set by 1.5 times to obtain the magnified coordinates of each bounding box;
[0015] Step 2.2, according to the magnified coordinate positions, cropping out all ship targets included in each sample from each sample to form the target slice map of the sample;
[0016] Step 2.3, successively apply bilateral filtering and adaptive median filtering to each target slice image; the size of the filtering kernel selected for bilateral filtering is 5, and the filtering range is 10; dynamically adjust the size of the filter window according to the gray values in the area covered by the filter window.
[0017] Step 2.4, use the OTSU method of maximum inter-class variance to perform adaptive processing on each filtered target slice image to obtain the binary target slice image corresponding to each target slice image.
[0018] Step 2.5, successively perform erosion-dilation and dilation-erosion operations on each binary target slice image to obtain the processed binary target slice image.
[0019] Step 2.6, crop each processed binary target slice image according to the size coordinates of the bounding box in Step 1.1, and paste the cropped binary target slice image onto a completely black image with the same size as the sample scale and all pixel values being 0 to obtain a semantic segmentation mask label with the pixel value of the ship target being 255 and the pixel value of the background being 0.
[0020] Step 3, generate a training set:
[0021] Step 3.1, perform the operation of generating the semantic segmentation label for ships in SAR images adaptively on each sample in the training sample set to obtain the semantic segmentation label of the training sample set.
[0022] Step 3.2, preprocess each sample in the training sample set and its corresponding semantic segmentation mask label to obtain the preprocessed training samples, semantic segmentation labels, and object detection labels.
[0023] Step 3.3, form a training set with all the preprocessed training samples, semantic segmentation labels, and object detection labels.
[0024] Step 4, construct a mask feature fusion object detection network composed of a mask feature fusion sub-network, a feature extraction sub-network, a multi-scale feature fusion sub-network, and its detection head.
[0025] Step 4.1: Build an 8-layer masked feature fusion sub-network, which includes: a first GateC module, a first CSP module, a second GateC module, a second CSP module, a convolutional layer, a sigmoid layer, a first convolutional sampling layer, and a second convolutional sampling layer; among them, the first GateC module, the first CSP module, the second GateC module, the second CSP module, the convolutional layer, and the sigmoid layer are connected in series; the first convolutional sampling layer is connected in series with the first GateC module; the second convolutional sampling layer is connected in series with the second GateC module; set the input channels of the first and second CSP modules to 64 and 32 respectively, the output channels to 64 and 32 respectively, the number of Bottleneck modules to 1 each, and turn off the residual structure; the network structures of the first and second convolutional sampling layers are the same, and each is composed of a convolutional layer and an upsampling layer connected in series; set the input channels of the convolutional layers in the first and second convolutional sampling layers to 128 and 256 respectively, the output channels to 1 each, the kernel size of the convolutional kernels to 1×1, and the stride to 1; the upsampling layers in the first and second convolutional sampling layers are both set to 640×640 pixels;
[0026] The first and second CSP modules have the same structure. Each CSP module includes: a first CBS module, a Bottleneck module, a second CBS module, a splicing layer, and a third CBS module; the network structure of the CSP module is: the first CBS module, the Bottleneck module, the splicing layer, and the third CBS module are connected in series, and the second CBS module is connected in parallel with the first CBS module and the Bottleneck module at the splicing layer;
[0027] The first and second GateC modules have the same structure. The structure of each GateC module is in turn: a first splicing layer, a first batch normalization layer, a first convolutional layer, a ReLU activation layer, a second convolutional layer, a second batch normalization layer, a sigmoid layer, a multiplication layer, an addition layer, and a third convolutional layer; set the input channels of the first to third convolutional layers in the first GateC module to 65, 65, and 64 respectively, and the output channels to 65, 1, and 64 respectively; set the input channels of the first to third convolutional layers in the second GateC module to 33, 33, and 32 respectively, and the output channels to 33, 1, and 32 respectively; the kernel size of all convolutional layers in the two GateC modules is 1×1, and the stride is 1;
[0028] Step 4.2: Build a feature extraction sub-network, whose structure is in sequence: the first CBS module, the second CBS module, the first CSP module, the third CBS module, the second CSP module, the fourth CBS module, the third CSP module, the fifth CBS module, the ASPPF module, the fourth CSP module; Set the input channel numbers of the first to fifth CBS modules to 1, 32, 64, 128, 256 respectively, the output channel numbers to 32, 64, 128, 256, 512 respectively, the strides to 1, 2, 2, 2, 2 respectively, and the convolutional kernel sizes to 3×3; Set the input channel numbers of the first to fourth CSP modules to 64, 128, 256, 512 respectively, the output channel numbers to 64, 128, 256, 512 respectively, and the numbers of Bottleneck modules to 1, 2, 2, 2, 2 respectively; The structures of the first to third CSP modules are with the residual structure enabled, and the structure of the fourth CSP module is with the residual structure disabled;
[0029] The structures of the first to third CSP modules are the same as the CSP module in Step 4.1; The structures of the first to fifth CBS modules are the same, and the network of each CBS module is composed of a convolutional layer, a batch normalization layer, and a SiLU activation layer connected in series;
[0030] The ASPPF module includes: the first CBM module, the first max-pooling layer, the second max-pooling layer, the third max-pooling layer, a concatenation layer, and the second CBM module; Among them, the first CBM module, the first max-pooling layer, the second max-pooling layer, and the third max-pooling layer are connected in series; The concatenation layer is connected in series with the second CBM module; The first CBM module, the first max-pooling layer, the second max-pooling layer, and the third max-pooling layer are connected in parallel at the concatenation layer; Set the convolutional kernel sizes of the first to third max-pooling layers to 5×5; The structures of the first and second CBM modules are the same, and each CBM module is composed of a convolutional layer, a batch normalization layer, and a MetaAcon adaptive activation layer connected in series; Set the input channel numbers of the first and second CBM modules to 512, 1024 respectively, and the output channel numbers to 256, 512 respectively; The convolutional kernel sizes of the convolutional layers are 1×1, and the strides are 1;
[0031] Step 4.3, construct a multi-scale feature fusion sub-network, the structure of which includes: the first CBS module, the first upsampling layer, the first splicing layer, the first CSP module, the second CBS module, the second upsampling layer, the second splicing layer, the second CSP module, the third CBS module, the third splicing layer, the third CSP module, the fourth CBS module, the fourth splicing layer, the fourth CSP module, and the Fusion module; among them, the first CBS module, the first upsampling layer, the first splicing layer, the first CSP module, the second CBS module, the second upsampling layer, the second splicing layer, the second CSP module, the third CBS module, the third splicing layer, the third CSP module, the fourth CBS module, the fourth splicing layer, and the fourth CSP module are connected in series in turn; the Fusion module is connected in series with the second splicing layer; the first CBS module is connected in parallel with the fourth splicing layer; the first CSP module is connected in parallel with the third splicing layer; the first to second upsampling layers are set to 80×80 pixels and 160×160 pixels respectively; the input channel numbers of the first to fourth CBS modules are set to 512, 256, 128, and 256 respectively, the output channels are set to 256, 128, 128, and 256 respectively, the convolution kernel sizes are set to 1×1, 1×1, 3×3, and 3×3 respectively, and the strides are set to 1, 1, 2, and 2 respectively; the input channel numbers of the first to fourth CSP modules are set to 512, 256, 256, and 256 respectively, the output channels are set to 256, 128, 512, and 512 respectively, the number of BottleNeck modules is set to 1, and the residual structure is turned off;
[0032] The structure of the Fusion module is, in turn, a downsampling layer, a splicing layer, a first convolutional layer, a first batch normalization layer, a first ReLU activation layer, a dropout layer, a second convolutional layer, a second batch normalization layer, and a second ReLU activation layer; the Fusion module has two inputs, namely a mask feature input and a backbone feature input; the downsampling layer is set to 160×160 pixels; the input channels of the first and second convolutional layers are set to 288 and 320 respectively, the output channels are set to 320 and 256 respectively, the convolution kernel sizes are both set to 3, and the strides are both set to 1; the dropout rate of the dropout layer is set to 0.1;
[0033] Step 4.4, connect the first CSP module in the feature extraction sub-network to the first GateC module in the mask feature fusion sub-network as the gated input of the first GateC module; connect the first convolutional sampling layer and the second convolutional sampling layer of the mask feature fusion sub-network to the second CSP module and the third CSP module in the feature extraction sub-network respectively; connect the fourth CBS module in the feature extraction sub-network to the Fusion module in the multi-scale feature fusion sub-network as the backbone feature input of the Fusion module; connect the fifth CBS module and the fourth CSP module in the feature extraction sub-network to the first splicing layer and the first CBS module in the multi-scale feature fusion sub-network respectively; connect the outputs of the second, third, and fourth CSP modules of the multi-scale feature fusion sub-network to the detection head 1, detection head 2, and detection head 3 respectively to obtain the mask feature fusion object detection network;
[0034] Step 5, generate the object detection loss function Loss all as follows:
[0035]
[0036] where Loss target represents the loss value of the object detection task, and Loss target represents the loss function of the traditional Yolov5 network; Loss bce represents the weighted cross-entropy loss function value, and σ1 and σ2 respectively represent the weight coefficients automatically assigned to Loss target and Loss target during the network training process, and log represents the logarithmic operation with the natural constant e as the base;
[0037] The weighted cross-entropy loss function Loss bce is as follows:
[0038]
[0039] ω p = n neg / (n pos + n neg )
[0040] ω n = n pos / (n pos + n neg )
[0041] where N represents the total number of pictures of the semantic segmentation mask labels in the current network iterative training, y i represents the pixel value of the i-th semantic segmentation mask label, p i represents the output of the i-th semantic segmentation mask label through the network model, ωp and ω n respectively represent the weight values assigned to the ship target and the background, and n neg and n pos respectively represent the total number of ship target pixels and the total number of background pixels in the semantic segmentation mask label;
[0042] Step 6. Train the mask feature fusion object detection network:
[0043] Input the images in the training set into the mask feature fusion object detection network, use the SGD optimization algorithm, and iteratively update the network weight values until the loss function converges, obtaining a trained mask feature fusion object detection network;
[0044] Step 7. Detect the positions of ship targets in the SAR image:
[0045] Sample the SAR ship target image to be detected to 640×640 pixels and then input it into the trained mask network feature fusion object detection network, and output the target box coordinates and confidence levels of each ship target in the image to be detected.
[0046] The present invention has the following advantages compared with the existing technologies:
[0047] First, the present invention adds a mask network feature fusion branch in the Yolov5 network. By extracting image features from the Yolov5 network and combining the gated convolution and feature fusion operations, the attention of the network is concentrated on the ship target, and the interference brought by the background information is weakened, enabling the object detection network to better distinguish the target from the background, avoiding the problems of high false alarms and low detection accuracy caused by the network's indiscriminate processing of the target and the background, and improving the detection accuracy of ship targets in the SAR image.
[0048] Second, the present invention designs an adaptive semantic segmentation mask label making strategy. The present invention separates the target from the background by using the brightness gradient difference between the target and the background, and adaptively extracts the semantic segmentation mask label, enabling the present invention to combine the semantic segmentation work with the object detection work even without semantic segmentation labels, further improving the detection accuracy of ship targets in the SAR image.
[0049] Third, the present invention designs a cross-entropy loss function with weight coefficients for the mask branch, improves the constraint weight of the branch for the ship target, solves the problem of foreground-background imbalance, and further optimizes the object detection effect. Description of the Drawings
[0050] Figure 1 is the flowchart of the present invention;
[0051] Figure 2It is a schematic diagram of the overall network structure of the present invention;
[0052] Figure 3 It is a schematic diagram of the residual block CSP module of the present invention;
[0053] Figure 4 It is a schematic diagram of the gated convolution GateC module of the present invention;
[0054] Figure 5 It is a schematic diagram of the adaptive fully connected layer ASPPF module of the present invention;
[0055] Figure 6 It is a schematic diagram of the Fusion feature fusion module of the present invention;
[0056] Figure 7 It is a schematic diagram of the visual object detection result of the simulation experiment of the present invention on the SSDD dataset.
[0057] Figure 8 It is a schematic diagram of the visual object detection result of the simulation experiment of the present invention on the HRSID dataset. Detailed implementation manners
[0058] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0059] Refer to Figure 1 to further describe in detail the implementation steps of the embodiments of the present invention.
[0060] Step 1, generate a training sample set and a test sample set.
[0061] The SAR image samples used in the embodiments of the present invention are derived from the SSDD and HRSID datasets. The SSDD dataset is a multi - resolution, multi - size, and multi - sensor SAR ship dataset published by Li et al. (J. Li, C. Qu, and J. Shao, “Ship detection in sar images based on an improved faster r - cnn,” in 2017 SAR in Big Data Era: Models, Methods and Applications (BIGSARDATA), 2017, pp. 1–6.) and is widely used for SAR ship detection. The SSDD dataset contains 1160 images of different scales, with a total of 2456 ships. The scale range of the ship targets is from 7×7 to 211×298 pixels. The HRSID dataset is a SAR ship dataset publicly available by Wei et al. (S. Wei, X. Zeng, Q. Qu, M. Wang, H. Su and J. Shi, "HRSID: A high - resolution SAR images dataset for ship detection and instance segmentation", IEEE Access, vol. 8, pp. 120234 - 120254, 2020.). The HRSID dataset contains 5604 images with a scale of 800×800 pixels each, with a total of 16951 ships, and the scale of the ship targets varies.
[0062] Each sample in the SSDD dataset and the HRSID dataset contains at least one ship target. For each sample, both datasets annotate the bounding box coordinates of each ship target in each sample in the labeling method of PASCAL VOC. Users can obtain the four vertex coordinates of the bounding boxes of all ship targets in each sample from the object detection labels.
[0063] The present invention forms a SSDD test sample set with 232 images named ending with 1 and 9 in the SSDD dataset, and the remaining 928 images form the SSDD training sample set. The 3642 images divided according to the division method in the author's paper in the HRSID dataset form the HRSID training sample set, and the remaining 1962 images form the HRSID test sample set.
[0064] Step 2, generate adaptive SAR image ship semantic segmentation labels.
[0065] Due to the characteristics of SAR images where ship targets are bright and the ocean background is dark, there is a brightness gradient difference between the background and the targets at the boundaries of ship targets. By utilizing this brightness gradient difference, the targets can be separated from the background, and semantic segmentation mask labels can be extracted to assist the mask feature fusion object detection network in completing training.
[0066] Embodiments of the present invention separate the targets and the background in each sample of the SSDD dataset and the HRSID dataset, and extract the semantic segmentation mask labels in this sample. The steps are as follows:
[0067] Step 2.1: Enlarge the coordinates of each bounding box in the object detection label of each sample by 1.5 times to obtain the enlarged coordinates of each bounding box.
[0068] Step 2.2: According to the enlarged coordinate positions, crop out all the ship targets included in each sample from each sample to form the target slice map of this sample.
[0069] Step 2.3: To reduce noise interference, bilateral filtering and adaptive median filtering adopted in the embodiments of the present invention are used to filter each target slice map in turn to obtain a clean target slice map. The size of the filtering kernel selected for bilateral filtering is 5, and the filtering range is 10. Adaptive median filtering adds a step of adaptively changing the sliding window size on the basis of median filtering, and dynamically adjusts the filter window size according to the gray values of the area covered by the filter window. In the embodiments of the present invention, the change range of the filter window is from 3 to 7.
[0070] Step 2.4: Using the OTSU method of maximum inter-class variance, according to the gray values of each filtered target slice map, obtain the binary segmentation threshold of this slice map. Since the slice map is rectangular and the ship target is fusiform, there is still some background information in the slice map. According to the calculated threshold, the part greater than the threshold in each target slice map is determined as the target, and the part less than the threshold is determined as the background. This slice map is divided into two parts: the background and the target. The pixel values of the background part are set to 0, and the pixel values of the target part are set to 255 to obtain the binary target slice map after adaptive processing.
[0071] Step 2.5, in order to reduce the interference of holes and snowflakes in the binary image, the erosion-dilation and dilation-erosion operations are sequentially performed on each binary target slice image. Since the target scale is variable, in order to obtain a more accurate mask label closer to the real target, embodiments of the present invention select different sizes of filtering kernels in the two operations of erosion-dilation and dilation-erosion. Based on the scale size of the target slice, different operations are performed on the binary target slice images divided into different operation intervals. Specifically, for binary target slice images with a width and height less than 20 pixels, no erosion-dilation and dilation-erosion operations are performed. For binary target slice images with a width and height in the range of 20 - 150 pixels, 150 - 210 pixels, and greater than 210 pixels, the sizes of the filtering kernels selected for the erosion-dilation and dilation-erosion operations are 3 and 2, 4 and 2, and 5 and 4 respectively. Through the erosion-dilation and dilation-erosion operations, the processed binary target slice images are obtained.
[0072] Step 2.6, in order to avoid introducing interference in the target detection task, according to the scale of the original bounding box corresponding to the sample label, a cropping operation is performed on each processed binary target slice image.
[0073] Step 2.7, according to the vertex coordinates of the original bounding box corresponding to the sample label, the cropped binary target slice image is pasted on a completely black image with the same scale as the sample and all pixel values being 0, to obtain a semantic segmentation mask label with the pixel value of the ship target being 255 and the pixel value of the background being 0.
[0074] Step 3, generate the training set and the test set.
[0075] Step 3.1, since the present invention only uses semantic segmentation labels during the training process of the network, semantic segmentation labels are not generated for test samples. For each sample in the SSDD training sample set and the HRSID training sample set respectively, the adaptive SAR image ship semantic segmentation label generation operation is performed to obtain the semantic segmentation labels of the SSDD training sample set and the semantic segmentation labels of the HRSID dataset training sample set.
[0076] Step 3.2, in order to ensure the alignment of the semantic segmentation mask, the target detection label, and the training samples after data augmentation, for each sample in the SSDD and HRSID training sample sets respectively, operations of sampling, random flipping, random cropping, and mosaic splicing data augmentation are sequentially performed to obtain training samples with an input scale size of 640×640 pixels. The semantic segmentation mask and the training samples are subjected to the same operations of sampling, random flipping, random cropping, and mosaic splicing data augmentation, and the coordinates of the bounding box in the target detection label are adjusted accordingly according to the position of the ship target in the preprocessed training samples. The preprocessed training samples, semantic segmentation labels, and target detection labels are obtained.
[0077] Step 3.3, generate the training set and the test set. All the samples, semantic segmentation labels, and object detection labels after preprocessing the SSDD training sample set are combined to form the SSDD training set. The SSDD test sample set and the object detection labels are combined to form the SSDD test set. All the preprocessed samples, semantic segmentation labels, and object detection labels in the HRSID training samples are combined to form the HRSID training set, and the HRSID test sample set and the object detection labels are combined to form the HRSID test set.
[0078] Step 4, construct a mask feature fusion object detection network.
[0079] Refer to Figure 2 , and make a further description of the constructed mask feature fusion object detection network structure.
[0080] The mask feature fusion object detection network is composed of four interconnected networks. The four networks are: the mask feature fusion subnet, the feature extraction subnet, the multi-scale feature fusion subnet, and the detection head.
[0081] Step 4.1, build an 8-layer mask feature fusion subnet. As shown in the mask feature fusion subnet part of the overall network structure in Figure 2 .
[0082] The mask feature fusion subnet includes: the first GateC module, the first CSP module, the second GateC module, the second CSP module, a convolutional layer, a sigmoid layer, the first convolutional sampling layer, and the second convolutional sampling layer. Among them, the first GateC module, the first CSP module, the second GateC module, the second CSP module, the convolutional layer, and the sigmoid layer are connected in series. The first convolutional sampling layer is connected in series with the first GateC module. The second convolutional sampling layer is connected in series with the second GateC module. The input channels of the first and second CSP modules are set to 64 and 32 respectively, the output channels are set to 64 and 32 respectively, the number of Bottleneck modules is set to 1 for both, and the residual structure is turned off for both. The network structures of the first and second convolutional sampling layers are the same, and both are composed of a convolutional layer and an upsampling layer connected in series. The input channels of the convolutional layers in the first and second convolutional sampling layers are set to 128 and 256 respectively, the output channels are both set to 1, the kernel size is both set to 1×1, and the stride is both set to 1. The upsampling layers in the first and second convolutional sampling layers are both set to 640×640 pixels.
[0083] The structures of the first to second CSP modules are the same, as shown in Figure 4As shown in the figure. The CSP module includes: a first CBS module, a Bottleneck module, a second CBS module, a splicing layer, and a third CBS module. The network structure of the CSP module is: the first CBS module, the Bottleneck module, the splicing layer, and the third CBS module are connected in series, and the second CBS module is connected in parallel with the first CBS module, and the Bottleneck module is connected in parallel at the splicing layer.
[0084] The structures of the first and second GateC modules are the same, as Figure 3 shown. The structure of each GateC module is in turn: a first splicing layer, a first batch normalization layer, a first convolutional layer, a ReLU activation layer, a second convolutional layer, a second batch normalization layer, a sigmoid layer, a multiplication layer, an addition layer, and a third convolutional layer. The input channels of the first to third convolutional layers in the first GateC module are set to 65, 65, and 64 respectively, and the output channels are set to 65, 1, and 64 respectively. The input channels of the first to third convolutional layers in the second GateC module are set to 33, 33, and 32 respectively, and the output channels are set to 33, 1, and 32 respectively. The kernel sizes of all convolutional layers in the two GateC modules are set to 1×1, and the strides are set to 1.
[0085] The first and second GateC modules have two inputs, namely a gating input and a localization feature input. The gating input and the localization feature input are connected in parallel at the first splicing layer. The localization feature input and the output of the sigmoid layer are multiplied at the first multiplication layer and then added to the output of the multiplication layer at the addition layer. The outputs of the first convolutional sampling layer and the second convolutional sampling layer are the localization inputs of the first GateC module and the second GateC module respectively.
[0086] Step 4.2, build a feature extraction sub-network, as Figure 2 shown in the feature extraction sub-network part of the overall network structure.
[0087] The structure of the feature extraction sub-network is successively: the first CBS module, the second CBS module, the first CSP module, the third CBS module, the second CSP module, the fourth CBS module, the third CSP module, the fifth CBS module, the ASPPF module, and the fourth CSP module. Among them, the structures of the first to third CSP modules are the same as the CSP module in step 4.1. The input channel numbers of the first to fifth CBS modules are set to 1, 32, 64, 128, and 256 respectively, the output channel numbers are set to 32, 64, 128, 256, and 512 respectively, the strides are set to 1, 2, 2, 2, and 2 respectively, and the convolution kernel sizes are all set to 3×3. The input channel numbers of the first to fourth CSP modules are set to 64, 128, 256, and 512 respectively, the output channel numbers are set to 64, 128, 256, and 512 respectively, and the numbers of Bottleneck modules are set to 1, 2, 2, and 2 respectively. The structures of the first to third CSP modules are open residual structures, and the structure of the fourth CSP module is a closed residual structure.
[0088] The structures of the first to fifth CBS modules are the same. The network structure of each CBS module is: a convolutional layer, a batch normalization layer, and a SiLU activation layer in series.
[0089] Refer to Figure 5 A further description of the structure of the ASPPF module is made. The ASPPF module includes: the first CBM module, the first max pooling layer, the second max pooling layer, the third max pooling layer, a concatenation layer, and the second CBM module. Among them, the first CBM module, the first max pooling layer, the second max pooling layer, and the third max pooling layer are in series. The concatenation layer is in series with the second CBM module. The first CBM module, the first max pooling layer, the second max pooling layer, and the third max pooling layer are in parallel at the concatenation layer. The convolution kernel sizes of the first to third max pooling layers are all set to 5×5. The structures of the first and second CBM modules are the same. The network structure of the CBM module is: a convolutional layer, a batch normalization layer, and a MetaAcon adaptive activation layer in series. The input channel numbers of the first and second CBM modules are set to 512 and 1024 respectively, and the output channel numbers are set to 256 and 512 respectively. The convolution kernel sizes of the convolutional layers are all set to 1×1, and the strides are all set to 1.
[0090] Step 4.3, construct a multi-scale feature fusion sub-network, the structure of which is as Figure 2 shown in the multi-scale feature fusion sub-network part of the overall network structure.
[0091] The multi-scale feature fusion sub-network includes: the first CBS module, the first upsampling layer, the first splicing layer, the first CSP module, the second CBS module, the second upsampling layer, the second splicing layer, the second CSP module, the third CBS module, the third splicing layer, the third CSP module, the fourth CBS module, the fourth splicing layer, the fourth CSP module, and the Fusion module. Among them, the first CBS module, the first upsampling layer, the first splicing layer, the first CSP module, the second CBS module, the second upsampling layer, the second splicing layer, the second CSP module, the third CBS module, the third splicing layer, the third CSP module, the fourth CBS module, the fourth splicing layer, and the fourth CSP module are connected in series. The Fusion module is connected in series with the second splicing layer, the first CBS module is connected in parallel with the fourth splicing layer, and the first CSP module is connected in parallel with the third splicing layer;. The first to second upsampling layers are set to 80×80 pixels and 160×160 pixels respectively. The input channel numbers of the first to fourth CBS modules are set to 512, 256, 128, 256 respectively, the output channels are set to 256, 128, 128, 256 respectively, the convolutional kernel sizes are set to 1×1, 1×1, 3×3, 3×3 respectively, and the strides are set to 1, 1, 2, 2 respectively. The input channel numbers of the first to fourth CSP modules are set to 512, 256, 256, 256 respectively, the output channels are set to 256, 128, 512, 512 respectively, the number of BottleNeck modules is set to 1, and the residual structure is turned off for all of them.
[0092] The structure of the said Fusion module is as Figure 6 shown, and its structure is successively the downsampling layer, the splicing layer, the first convolutional layer, the first batch normalization layer, the first ReLU activation layer, the dropout layer, the second convolutional layer, the second batch normalization layer, and the second ReLU activation layer. The Fusion module has two inputs, namely the mask feature input and the backbone feature input. The downsampling layer is set to 160×160 pixels. The input channels of the first and second convolutional layers are set to 288 and 320 respectively, the output channels are set to 320 and 256 respectively, and the convolutional kernel sizes are both set to 3 and the strides are both set to 1. The dropout rate of the dropout layer is set to 0.1.
[0093] Step 4.4, Combine the four-part network. Connect the first CSP module in the feature extraction sub-network to the first GateC module in the mask feature fusion sub-network as the gated input of the first GateC module; connect the first convolutional sampling layer and the second convolutional sampling layer of the mask feature fusion sub-network to the second CSP module and the third CSP module in the feature extraction sub-network respectively; connect the fourth CBS module in the feature extraction sub-network to the Fusion module in the multi-scale feature fusion sub-network as the backbone feature input of the Fusion module; connect the fifth CBS module and the fourth CSP module in the feature extraction sub-network to the first splicing layer and the first CBS module in the multi-scale feature fusion sub-network respectively; connect the outputs of the second, third, and fourth CSP modules in the multi-scale feature fusion sub-network to Detection Head 1, Detection Head 2, and Detection Head 3 respectively to obtain the mask feature fusion object detection network.
[0094] The CSP module mentioned in the mask feature fusion object detection network refers to the same module as the CSP module in the Yolov5 network, and the user can independently select the number of Bottleneck modules connected in the CSP module; the network structure of the Bottleneck module is divided into two types: residual structure and non-residual structure, and the user can independently choose whether to enable the residual structure. In the case of turning off the residual structure, the network structure of the Bottleneck module is two CBS modules connected in series; in the case of enabling the residual structure, the network structure of the Bottleneck module is two CBS modules connected in series, and its output is added to the input of the Bottleneck module to form a residual connection.
[0095] Step 6, Construct the network loss function.
[0096] The loss function of the network consists of two parts: the object detection loss function and the weighted cross-entropy loss of the mask fusion sub-network.
[0097] Step 6.1, Construct the object detection loss function. The object detection loss function Loss target is consistent with Yolov5 and is composed of the localization loss l box , the confidence loss l obj and the classification loss l cls added together. That is:
[0098] Loss target = l box + l obj + l cls
[0099] Step 6.2, Construct the weighted cross-entropy loss Loss bce. To enhance the network's attention to ship targets, the present invention improves on the basis of the cross-entropy loss and introduces a weight coefficient ω when calculating the cross-entropy loss. p and ω n . The calculation method of Loss bce is as follows:
[0100]
[0101] ω p = n neg / (n pos + n neg )
[0102] ω n = n pos / (n pos + n neg )
[0103] where p i represents the output result of the network model, ω p and ω n respectively represent the weights assigned to the ship target and the background, n neg and n pos respectively represent the total number of pixel values of the ship target and the background in the adaptive mask label, log represents the logarithmic function with the natural number e as the base, and the weights ω p and ω n and the weighted cross-entropy loss Loss bce can be calculated from the above formula.
[0104] Step 6.3, automatic weight assignment for the loss function. The mask feature fusion sub-network and the target detection network can be regarded as two different tasks of semantic segmentation and target detection respectively. Therefore, the training process of the network is unstable and prone to the situation where the loss function does not converge, resulting in the failure of network training. Therefore, the present invention adds an automatic weight assignment method to the loss function.
[0105] The total loss function Loss all is as follows:
[0106]
[0107] The present invention assigns weight coefficients σ1 and σ2 to the target detection loss Loss target and the weighted cross-entropy loss Loss bce . σ1 and σ2 are automatically updated as the network is trained to adapt to the optimization process of the network.
[0108] Step 7, train the mask feature fusion target detection network.
[0109] Input the preprocessed training set images into the network for training. The present invention uses the SGD optimization algorithm to iteratively update the network weight values until the loss function converges, obtaining a trained object detection network.
[0110] Step 8: Detect the positions of ship targets in the SAR image.
[0111] Sample each test sample in the test set to 640×640 pixels and then input it into the trained mask network feature fusion object detection network to output the bounding box coordinates and confidence levels of the ship targets in each test sample.
[0112] The effect of the present invention can be further demonstrated by the following simulation.
[0113] 1. Simulation experiment conditions:
[0114] The hardware platform for the simulation experiment of the present invention is: NVIDIA A100. The software platform for the simulation experiment is: Ubuntu18.04.6 operating system, based on the PyTorch1.7.1 deep learning framework, and the programming language is Python3.7.
[0115] 2. Simulation experiment content and result analysis:
[0116] The simulation experiment of the present invention uses the method of the present invention and an existing technology Yolov5 to perform object detection on the test samples in the SSDD and HRSID datasets respectively, and the detection results are as Figure 7 and Figure 8 shown.
[0117] In the simulation experiment, the existing technology Yolov5 refers to the object detection model proposed by Ultralytics in “ultralytics / yolov5:v5.0” (2021. [Online]. Available: https: / / github.com / ultralytics / yolov5).
[0118] The following combines Figure 7 and Figure 8 's simulation diagrams to further describe the effect of the present invention.
[0119] Figure 7 For the comparison diagram of the detection effects between the Yolov5 network and the present invention on the SSDD dataset, Figure 7 a total of 5 pictures are selected for comparison. Figure 7 The first row in
[0120] represents the detection results of the Yolov5 network, the second row represents the detection results of the present invention, and the third row represents the correct results labeled by the object detection labels.Figure 7 As can be seen from column (a) in [reference], the bounding box coordinates predicted by the present invention are more accurate than the prediction results of the Yolov5 network. From Figure 7 columns (b) and (c) in [reference], it can be seen that in the SSDD dataset, when the noise interference is large, the Yolov5 network is prone to misidentifying interference targets such as offshore reefs as ship targets, resulting in false alarms, while the present invention can accurately distinguish reefs from ships. From Figure 7 columns (d) and (e) in [reference], it can be seen that in the nearshore background where ships are densely arranged, due to the background interference and the lack of obvious details of ship targets, it is difficult for the Yolov5 network to correctly identify closely arranged ships, while the present invention can better distinguish closely arranged ships. Figure 7 The detection results shown indicate that the network designed by the present invention has better detection effects in the presence of noise interference and complex backgrounds.
[0121] Figure 8 Fig. [figure number] is a comparison chart of the detection effects of the Yolov5 network and the present invention on the HRSID dataset. Figure 8 A total of 5 pictures were selected for comparison. The first row represents the detection results of the Yolov5 network, the second row represents the detection results of the present invention, and the third row represents the correct results marked by the target detection labels.
[0122] From Figure 8 columns (a) and (e) in [reference], it can be seen that in the HRSID dataset, for images with complex backgrounds, it is difficult for the Yolov5 network to distinguish targets from the background, and the detection results are prone to missed detections and false alarms, while the present invention effectively avoids these phenomena. Figure 8 From columns (b), (c), and (d) in [reference], it can be seen that the bounding box coordinates predicted by the present invention are more accurate than the prediction results of the Yolov5 network. Compared with Yolov5, the network designed by the present invention effectively reduces the problems of missed detections and false alarms and has better detection effects.
[0123] To verify the simulation effect of the present invention, the evaluation index mAP was used to evaluate the Yolov5 and the method of the present invention, and the integral method was used to calculate the area enclosed by the precision-recall curve and the coordinate axes for all categories. The specific calculation formula is as follows:
[0124]
[0125]
[0126]
[0127] Among them, TP, FP, and FN refer to the number of correctly detected ships, false alarms, and missing ships respectively. N represents the number of categories, AP is the area under the precision and recall curve, and mAP is the average of APs for each category. Since the present invention only targets one type of object, mAP is equal to AP.
[0128] Table 1. Comparison table of mAP of different methods on the SSDD dataset in the simulation experiment
[0129] Method <![CDATA[mAP 50 (%)]]> <![CDATA[mAP 75 (%)]]> Yolov5 97.3 64.6 The present invention 97.9 64.7
[0130] Table 2. Comparison table of mAP of different methods on the HRSID dataset in the simulation experiment
[0131] Method <![CDATA[mAP 50 (%)]]> <![CDATA[mAP 75 (%)]]> Yolov5 91.2 65.2 The present invention 93.4 69.5
[0132] Combining Table 1 and Table 2, it can be seen that the average detection accuracy mAP of the SAR image ship target detection method of the present invention on the SSDD and HRSID datasets 50 is 97.9% and 93.4% respectively, which are 0.6% and 2.2% higher than the Yolov5 network respectively; mAP 75 is 64.7% and 69.5% respectively, which are 0.1% and 3.7% higher than the Yolov5 network respectively, proving that the present invention can obtain higher SAR ship target detection accuracy.
[0133] The above simulation experiments show that: the present invention proposes a SAR ship target detection method based on mask network fusion of image features. The present invention designs and constructs a mask network feature fusion branch, which highlights the target features and reduces the interference of background information, avoiding the problem of the decrease in detection accuracy caused by the network's indiscriminate processing of targets and backgrounds. The adaptive semantic segmentation mask label making strategy proposed by the present invention effectively separates the target and the background using the brightness gradient difference between the target and the background, and completes the production of semantic segmentation labels for different SAR datasets, avoiding the high dependence of the task on semantic segmentation labels. In addition, the present invention carefully designs a weighted cross-entropy loss and an adaptive loss weight allocation strategy for the network, enabling the network to better separate the target and the background and maintain the stability of the training process. The present invention can reduce false alarms and missed alarms when detecting targets in SAR images of complex scenes, thereby improving the target detection accuracy and having important practical application value.
Claims
1. A SAR ship target detection method based on mask network fusion of image features, characterized in that Generate an adaptive SAR image ship semantic segmentation label using the brightness gradient difference between the ship and the background, construct a mask feature fusion object detection network, and construct a weighted cross-entropy loss function; the steps of this object detection method are as follows: Step 1, generate a sample set: Step 1.1, collect at least 900 SAR ship image samples with a scale larger than 320×320 pixels, and each sample contains at least one ship target; Step 1.2, annotate an object detection label file for each sample, and each object detection label file contains the bounding box coordinates of all ship targets in the corresponding sample; Step 1.3, form a sample set by combining all SAR ship image samples and their corresponding object detection label files; Step 2, generate an adaptive SAR image ship semantic segmentation label using the brightness gradient difference between the ship and the background: Step 2.1, magnify the coordinates of each bounding box in the object detection label of each sample in the sample set by 1.5 times to obtain the magnified coordinates of each bounding box; Step 2.2, according to the magnified coordinate positions, crop all ship targets contained in each sample to form the target slice map of this sample; Step 2.3, sequentially perform bilateral filtering and adaptive median filtering on each target slice map; the filter kernel size selected for bilateral filtering is 5, and the filtering range is 10; dynamically adjust the filter window size according to the gray values in the area covered by the filter window; Step 2.4, use the OTSU method of maximum inter-class variance to perform adaptive processing on each filtered target slice map to obtain the binary target slice map corresponding to each target slice map; Step 2.5, sequentially perform erosion-dilation and dilation-erosion operations on each binary target slice map to obtain the processed binary target slice map; Step 2.6, crop each processed binary target slice map according to the size coordinates of the bounding box in Step 1.1, and paste the cropped binary target slice map on a completely black image with the same scale as the sample and all pixel values being 0 to obtain a semantic segmentation mask label with the ship target pixel value being 255 and the background pixel value being 0; Step 3, generate a training set: Step 3.1, perform the adaptive SAR image ship semantic segmentation label generation operation on each sample in the training sample set to obtain the semantic segmentation labels of the training sample set; Step 3.2, preprocess each sample in the training sample set and its corresponding semantic segmentation mask label to obtain the preprocessed training samples, semantic segmentation labels, and object detection labels; Step 3.3, form a training set by combining all the preprocessed training samples, semantic segmentation labels, and object detection labels; Step 4, construct a mask feature fusion object detection network composed of a mask feature fusion sub-network, a feature extraction sub-network, a multi-scale feature fusion sub-network, and its detection head: Step 4.1, build an 8-layer masked feature fusion sub-network, which includes: a first GateC module, a first CSP module, a second GateC module, a second CSP module, a convolutional layer, a sigmoid layer, a first convolutional sampling layer, and a second convolutional sampling layer; among them, the first GateC module, the first CSP module, the second GateC module, the second CSP module, the convolutional layer, and the sigmoid layer are connected in series; the first convolutional sampling layer is connected in series with the first GateC module; the second convolutional sampling layer is connected in series with the second GateC module; set the input channels of the first and second CSP modules to 64 and 32 respectively, the output channels to 64 and 32 respectively, and the number of Bottleneck modules to 1 for both, and turn off the residual structure; the network structures of the first and second convolutional sampling layers are the same, and both are composed of a convolutional layer and an upsampling layer connected in series; set the input channels of the convolutional layers in the first and second convolutional sampling layers to 128 and 256 respectively, the output channels to 1 for both, the kernel size to 1×1, and the stride to 1; the upsampling layers in the first and second convolutional sampling layers are both set to 640×640 pixels; The first and second CSP modules have the same structure, and each CSP module includes: a first CBS module, a Bottleneck module, a second CBS module, a splicing layer, and a third CBS module; the network structure of the CSP module is: the first CBS module, the Bottleneck module, the splicing layer, and the third CBS module are connected in series, and the second CBS module is connected in parallel with the first CBS module and the Bottleneck module at the splicing layer; The first and second GateC modules have the same structure, and the structure of each GateC module is in turn: a first splicing layer, a first batch normalization layer, a first convolutional layer, a ReLU activation layer, a second convolutional layer, a second batch normalization layer, a sigmoid layer, a multiplication layer, an addition layer, and a third convolutional layer; set the input channels of the first to third convolutional layers in the first GateC module to 65, 65, and 64 respectively, and the output channels to 65, 1, and 64 respectively; set the input channels of the first to third convolutional layers in the second GateC module to 33, 33, and 32 respectively, and the output channels to 33, 1, and 32 respectively; the kernel sizes of all convolutional layers in the two GateC modules are 1×1, and the strides are 1; Step 4.2: Build a feature extraction sub-network, whose structure is successively: the first CBS module, the second CBS module, the first CSP module, the third CBS module, the second CSP module, the fourth CBS module, the third CSP module, the fifth CBS module, the ASPPF module, the fourth CSP module; set the input channel numbers of the first to fifth CBS modules to 1, 32, 64, 128, 256 respectively, the output channel numbers to 32, 64, 128, 256, 512 respectively, the strides to 1, 2, 2, 2, 2 respectively, and the convolution kernel sizes to 3×3; set the input channel numbers of the first to fourth CSP modules to 64, 128, 256, 512 respectively, the output channel numbers to 64, 128, 256, 512 respectively, and the numbers of BottleNeck modules to 1, 2, 2, 2, 2 respectively; the structures of the first to third CSP modules are the open residual structure, and the structure of the fourth CSP module is the closed residual structure; The structures of the first to third CSP modules are the same as the CSP module in Step 4.1; the structures of the first to fifth CBS modules are the same, and the network of each CBS module is composed of a convolutional layer, a batch normalization layer, and a SiLU activation layer in series; The ASPPF module includes: the first CBM module, the first max-pooling layer, the second max-pooling layer, the third max-pooling layer, a concatenation layer, and the second CBM module; among them, the first CBM module, the first max-pooling layer, the second max-pooling layer, and the third max-pooling layer are in series; the concatenation layer is in series with the second CBM module; the first CBM module, the first max-pooling layer, the second max-pooling layer, and the third max-pooling layer are in parallel at the concatenation layer; set the convolution kernel sizes of the first to third max-pooling layers to 5×5; the structures of the first and second CBM modules are the same, and each CBM module is composed of a convolutional layer, a batch normalization layer, and a MetaAcon adaptive activation layer in series; set the input channel numbers of the first and second CBM modules to 512, 1024 respectively, and the output channel numbers to 256, 512 respectively; the convolution kernel sizes of the convolutional layers are 1×1, and the strides are 1; Step 4.3, construct a multi-scale feature fusion sub-network, the structure of which includes: the first CBS module, the first upsampling layer, the first splicing layer, the first CSP module, the second CBS module, the second upsampling layer, the second splicing layer, the second CSP module, the third CBS module, the third splicing layer, the third CSP module, the fourth CBS module, the fourth splicing layer, the fourth CSP module, and the Fusion module; among them, the first CBS module, the first upsampling layer, the first splicing layer, the first CSP module, the second CBS module, the second upsampling layer, the second splicing layer, the second CSP module, the third CBS module, the third splicing layer, the third CSP module, the fourth CBS module, the fourth splicing layer, and the fourth CSP module are connected in series in sequence; the Fusion module is connected in series with the second splicing layer; the first CBS module is connected in parallel with the fourth splicing layer; the first CSP module is connected in parallel with the third splicing layer; the first to second upsampling layers are respectively set to 80×80 pixels and 160×160 pixels; the input channel numbers of the first to fourth CBS modules are respectively set to 512, 256, 128, 256, and the output channels are respectively set to 256, 128, 128, 256, the convolution kernel sizes are respectively set to 1×1, 1×1, 3×3, 3×3, and the strides are respectively set to 1, 1, 2, 2; the input channel numbers of the first to fourth CSP modules are respectively set to 512, 256, 256, 256, and the output channels are respectively set to 256, 128, 512, 512, and the number of BottleNeck modules is set to 1, and the residual structure is turned off; The structure of the Fusion module is in sequence the downsampling layer, the splicing layer, the first convolutional layer, the first batch normalization layer, the first ReLU activation layer, the dropout layer, the second convolutional layer, the second batch normalization layer, the second ReLU activation layer; the Fusion module has two inputs, namely the mask feature input and the backbone feature input; the downsampling layer is set to 160×160 pixels; the input channels of the first and second convolutional layers are respectively set to 288, 320, and the output channels are respectively set to 320, 256, and the convolution kernel sizes are both set to 3, and the strides are both set to 1; the dropout rate of the dropout layer is set to 0.1; Step 4.4: Connect the first CSP module in the feature extraction sub-network to the first GateC module in the mask feature fusion sub-network as the gated input of the first GateC module; connect the first convolutional sampling layer and the second convolutional sampling layer of the mask feature fusion sub-network to the second CSP module and the third CSP module in the feature extraction sub-network respectively; connect the fourth CBS module in the feature extraction sub-network to the Fusion module in the multi-scale feature fusion sub-network as the backbone feature input of the Fusion module; connect the fifth CBS module and the fourth CSP module in the feature extraction sub-network to the first concatenation layer and the first CBS module in the multi-scale feature fusion sub-network respectively; connect the outputs of the second, third, and fourth CSP modules of the multi-scale feature fusion sub-network to Detection Head 1, Detection Head 2, and Detection Head 3 respectively to obtain the mask feature fusion object detection network; Step 5, generate the target detection loss function Loss all as follows: Among them, Loss targer represents the loss value of the object detection task, and Loss targer represents the loss function of the traditional Yolov5 network; Loss bce represents the weighted cross-entropy loss function value. σ1 and σ2 respectively represent the weight coefficients automatically assigned to Loss target and Loss target during the network training process, and log represents the logarithmic operation with the natural constant e as the base; The weighted cross-entropy loss function Loss bce is as follows: ω p = n neg / (n pos + n neg ) ω n = n pos / (n pos + n neg ) Among them, N represents the total number of pictures of semantic segmentation mask labels in the current network iterative training, and y i represents the pixel value of the i-th semantic segmentation mask label, and p i represents the output of the i-th semantic segmentation mask label through the network model, ω p and ω n respectively represent the weight values assigned to the ship target and the background, n neg and n pos respectively represent the total number of ship target pixels and the total number of background pixels in the semantic segmentation mask label; Step 6. Train the mask feature fusion object detection network: Input the images in the training set into the mask feature fusion object detection network, use the SGD optimization algorithm, and iteratively update the network weight values until the loss function converges to obtain the trained mask feature fusion object detection network; Step 7. Detect the positions of ship targets in SAR images: Sample the SAR ship target image to be detected to 640×640 pixels and then input it into the trained mask network feature fusion object detection network to output the target box coordinates and confidence levels of each ship target in the image to be detected.
2. The SAR ship target detection method based on masked network fusion of image features according to claim 1, wherein, The maximum inter-class variance method OTSU described in Step 2.4 means that, based on the gray values of each filtered target slice image, the binary segmentation threshold of the slice image is obtained; the part in each target slice image that is greater than the binary segmentation threshold of the slice image is determined as the target, and the part less than the threshold is determined as the background; Set the pixel values of the background part to 0 and the pixel values of the target part to 255 to obtain the adaptively processed binary target slice image.
3. The SAR ship target detection method based on masked network fusion of image features according to claim 1, wherein The erosion dilation and dilation erosion operations described in Step 2.5 mean that different-sized filter kernels are selected in the two operations of erosion dilation and dilation erosion. Based on the scale size of the target slice, different operations are performed on the binary target slice images divided into different operation intervals.
4. The SAR ship target detection method based on mask network fusion of image features according to claim 1, characterized in that The preprocessing described in Step 3.2 means that each sample in the training sample set is sequentially subjected to data augmentation operations such as sampling, random flipping, random cropping, and mosaic splicing to obtain training samples with an input scale size of 640×640 pixels; Perform the same operations of sampling, random flipping, random cropping, and mosaic splicing data augmentation on the semantic segmentation mask labels as on the training samples. The coordinates of the bounding boxes in the object detection labels are adjusted accordingly according to the positions of the ship targets in the preprocessed training samples to obtain the preprocessed training samples, semantic segmentation labels, and object detection labels.
Citation Information
Patent Citations
SAR ship target detection method based on balanced sample regression loss
CN112668440A
SAR Ship Target Detection Method Based on Balanced Sample Regression Loss
CN112668440B
SAR ship target detection method based on background and scale perception
CN114219997A
Rapid ship target detection method, storage medium and computing equipment
CN111914924A
Image processing method and system
US11341758B1