Semi-Supervised SAR Ship Detection Method Based on Scene Feature Learning
By adopting a semi-supervised detection method based on scene characteristic learning in SAR ship detection, using scene-level annotation to train the network and design a hierarchical testing process, the problem of severe dependence on target-level annotation and poor detection performance in the existing technology is solved, and more efficient detection accuracy is achieved.
Patent Information
- Application Number
- CN202310015820.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-05
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-01-05
AI Technical Summary
The prior art relies on a large number of target-level annotations in SAR ship detection, and it is easy to generate false alarms and missed alarms in complex scenarios, resulting in poor detection performance.
A semi-supervised SAR ship detection method based on scene characteristic learning is adopted. By constructing a semi-supervised SAR ship detection network, a scene-level annotation training network is used, and a hierarchical testing process from scene to target is designed during testing, and different detection strategies are set according to scene characteristics.
It reduces the dependence of network training on target-level annotations, reduces the occurrence of false alarms and missed alarms, and improves the accuracy and performance of SAR ship detection.
Smart Images

Figure CN115953695B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radar image, and further relates to a semi-supervised SAR ship detection method, which can be used to detect ship targets of interest from SAR images. Background Art
[0002] Synthetic Aperture Radar (SAR) has the advantage of providing remote sensing images under all-day and all-weather conditions and is widely used in military and civilian fields. With the rapid development of radar imaging technology, the field of SAR automatic target recognition has developed rapidly, and more and more high-resolution SAR images can be obtained. As the primary stage of SAR automatic target recognition, SAR automatic target detection has received extensive attention. In the field of SAR automatic target detection, an important branch is SAR ship detection, which is of great significance for maritime vessel surveillance and military intelligence acquisition. Constant False Alarm Rate (CFAR) is the most widely used and deeply studied traditional SAR ship detection method. This type of method uses background clutter statistical distribution modeling to utilize background information, thereby obtaining an adaptive threshold, and then comparing the gray value of the pixel with the adaptive threshold through a sliding window to obtain the detection result. Therefore, determining a suitable clutter statistical model is very important for ensuring the detection performance of CFAR. However, it is difficult to select a suitable clutter statistical model for measured SAR images due to the large amount of complex background clutter, resulting in a decline in detection performance. With the development of deep learning, many methods based on convolutional neural networks have been proposed and have achieved performance superior to traditional CFAR methods. These methods have made significant progress in SAR ship detection. However, these methods require that all SAR images have very fine target-level annotations. However, obtaining target-level annotations for SAR images requires a large amount of manpower and material resources, which is difficult to obtain in practice.
[0003] The patent document with the patent number CN201610561587.2 discloses a SAR image target detection method based on a convolutional neural network. It designs a SAR target detection network based on a convolutional neural network, and then uses the labeled training SAR images to train the target detection network. After the training converges, the trained model is used to test the test SAR images to obtain the detection results of the test SAR images. This method utilizes the feature extraction ability and non-linear mapping ability of the convolutional neural network and has good performance. However, this method requires a large number of SAR images with target-level labels as training data and has a high dependence on SAR images with target-level labels. In some cases where it is difficult to obtain SAR images with target-level labels, the detection performance of this method is poor.
[0004] The patent document with the application number CN201910016413.1 discloses a "SAR image target detection system and method based on semi - supervised CNN", which designs a self - learning algorithm to train the SAR target detection network. First, it uses the training SAR images with target - level labels to train the model; then uses the trained model to predict the training SAR images without target - level labels, and takes the prediction results with high confidence as pseudo - target - level labels; finally, uses all the training SAR images with target - level labels to train the model again; the last two steps of the whole process are repeated multiple times until convergence. Although this method reduces the dependence of network training on target - level labeled samples, it still has two deficiencies: one is that the pseudo - target - level labels generated based on self - learning may be incorrect, thus affecting network training and leading to a decline in detection performance; the other is that when the scene of the SAR image contains a large amount of complex background clutter, such as in complex inland and near - shore scenes, a large number of false alarms will be generated, resulting in a decline in detection accuracy. Summary of the Invention
[0005] The purpose of the present invention is to address the above - mentioned deficiencies of the prior art and propose a semi - supervised SAR ship detection method based on scene feature learning to improve the detection performance of the SAR target detection network under the condition of having a small number of target - level annotations, reduce the false alarms of SAR images in inland and near - shore scenes, and improve the detection accuracy.
[0006] The technical idea of the present invention is: by constructing a semi - supervised SAR ship detection network based on scene feature learning, and by designing a scene feature learning sub - network parallel to the detection sub - network during network training to make full use of the scene - level annotations of SAR images, so as to solve the problem that the prior art relies too much on target - level annotations in the training stage; the present invention constructs a hierarchical testing process from scene to target. During testing, it first identifies the scene type of the SAR image, and designs different ship detection strategies for SAR images identified as different scenes, so as to solve the problem that the prior art has too many false alarms and low detection accuracy when detecting targets in complex scene SAR images such as inland and near - shore. Its implementation steps are as follows:
[0007] (1) Generate a training set:
[0008] Collect at least 21 large - scale SAR images, and crop each large - scale SAR image into multiple sub - images of size 512×512; randomly select 30% of the sub - images containing ship targets for target - level labeling and scene - level labeling, and the remaining sub - images are only labeled with scene - level labels. All the labeled sub - images are combined to form a training set;
[0009] (2) Construct a semi - supervised SAR ship detection network:
[0010] (2a) Build a feature extraction sub-network composed of eight convolutional blocks connected in series;
[0011] (2b) Build a detection sub-network composed of four convolutional blocks and four detection heads, where the four convolutional blocks are first connected in series in sequence, and then each convolutional block is respectively connected to its corresponding detection head;
[0012] (2c) Build a scene feature learning sub-network composed of a scene recognition module and a scene aggregation module connected in parallel;
[0013] (2d) Connect the scene feature learning sub-network and the detection sub-network in parallel, and then connect them in series with the feature extraction sub-network to form a semi-supervised SAR ship detection network;
[0014] (3) Input the training set into the semi-supervised SAR ship detection network, and use the stochastic gradient descent algorithm to iteratively update the weight values of the network, optimize the total loss function of the network until it converges, and obtain the trained semi-supervised SAR ship detection network;
[0015] (4) Detect the position of the target box in the SAR image to be tested:
[0016] (4a) Slide and crop the large-scale SAR image to be tested into multiple sub-images of size 512×512;
[0017] (4b) Input each test sub-image into the trained feature extraction sub-network and the scene recognition module in sequence to obtain the scene recognition result of the test sub-image;
[0018] (4c) According to the scene recognition result, obtain the position and category of the target box:
[0019] For the test sub-image with the scene recognition result of the inland scene, output the detection result as no target;
[0020] For the test sub-image with the scene recognition result of the nearshore scene, input the test sub-image into the detection sub-network and set the nearshore detection threshold th in , and obtain the position and category of the target box of the test sub-image;
[0021] For the test sub-image with the scene recognition result of the open sea scene, input the test sub-image into the detection sub-network and set the open sea detection threshold th off , and obtain the position and category of the target box of the test sub-image;
[0022] (5) According to the order of the sliding window, map the position of the target box of each test sub-image to the corresponding position of each large-scale SAR image to be tested, and obtain the ship detection result of the large-scale SAR image.
[0023] The present invention has the following advantages compared with the existing technologies:
[0024] First, by simultaneously using the object-level annotation and scene-level annotation of SAR images to train the network, the present invention can learn features beneficial to SAR ship detection from the scene-level annotation of SAR images, thereby reducing the dependence of network training on object-level annotation and avoiding the problem of pseudo-object-level marking errors caused by self-learning, and improving the SAR ship detection performance under the condition of a small number of object-level markings.
[0025] Second, since a hierarchical testing process from scene to object is designed during testing, by fully considering the scene characteristics of SAR images and setting different detection strategies for SAR images of different scenes, the present invention significantly reduces inland false alarms, near-shore false alarms and open-sea missed alarms, thereby improving the SAR ship detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is the implementation flowchart of the present invention;
[0027] Figure 2 is the overall structural schematic diagram of the semi-supervised SAR ship detection network constructed in the present invention;
[0028] Figure 3 is the simulation diagram of the detection results of two large test original images of the AIR-SARShip-1.0 dataset using the present invention and the existing technology RefineDet. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The embodiments and effects of the present invention will be further described in detail below with reference to the drawings.
[0030] Refer to Figure 1 , the implementation steps of this example are as follows:
[0031] Step 1, generate a training set and a test set.
[0032] 1.1) In this embodiment, 21 large synthetic aperture radar (SAR) images are first collected from the measured AIR-SARShip-1.0 dataset; then, using a sliding window operation with a window size of 512×512, each large SAR image is cropped into sub-images with each pixel being 512×512.
[0033] 1.2) Randomly select 30% of the sub-images containing ship targets for object-level marking and scene-level marking, and only perform scene-level marking on the remaining all sub-images, and form a training set with all the marked sub-images;
[0034] 1.3) Select the remaining 10 large synthetic aperture radar (SAR) images in the measured AIR-SARShip-1.0 dataset. Using a sliding window operation with a window size of 512×512, each large SAR image is cropped into sub-images with a pixel size of 512×512, and these cropped sub-images are combined to form a test set.
[0035] Step 2: Construct a semi-supervised SAR ship detection network.
[0036] Refer to Figure 2 , and the specific implementation of this step is as follows:
[0037] 2.1) Construct a feature extraction sub-network:
[0038] Build a feature extraction sub-network composed of eight convolutional blocks, and its structure is in sequence: the first convolutional block, the second convolutional block, the third convolutional block, the fourth convolutional block, the fifth convolutional block, the sixth convolutional block, the seventh convolutional block, and the eighth convolutional block are cascaded;
[0039] The first convolutional block includes a convolutional layer and a pooling layer, and the convolutional kernel size of this convolutional layer is set to 7×7;
[0040] The number of convolutional layers in the second convolutional block is set to 4, and the convolutional kernel sizes of these 4 convolutional layers are respectively set to 1×1, 3×3, 1×1, 1×1;
[0041] The number of convolutional layers in the third convolutional block is set to 3, and the convolutional kernel sizes of these 3 convolutional layers are respectively set to 1×1, 3×3, 1×1;
[0042] The number of convolutional layers in the fourth convolutional block is set to 3, and the convolutional kernel sizes of these 3 convolutional layers are respectively set to 1×1, 3×3, 1×1;
[0043] The number of convolutional layers in the fifth convolutional block is set to 4, and the convolutional kernel sizes of these 4 convolutional layers are respectively set to 1×1, 3×3, 1×1, 1×1;
[0044] The number of convolutional layers in the sixth convolutional block is set to 3, and the convolutional kernel sizes of these 3 convolutional layers are respectively set to 1×1, 3×3, 1×1;
[0045] The number of convolutional layers in the seventh convolutional block is set to 3, and the convolutional kernel sizes of these 3 convolutional layers are respectively set to 1×1, 3×3, 1×1;
[0046] The number of convolutional layers in the eighth convolutional block is set to 3, and the convolutional kernel sizes of these 3 convolutional layers are respectively set to 1×1, 3×3, 1×1;
[0047] Take the output of the eighth convolutional block as the first convolutional feature map. The input of the feature extraction sub-network is the SAR image in the training set, and the output is the deep feature of the SAR image, that is, the first convolutional feature map;
[0048] 2.2) Construct the detection sub-network:
[0049] After the feature extraction sub-network, build a detection sub-network composed of four convolutional blocks and four detection heads. Its structure is as follows: the first convolutional block, the second convolutional block, the third convolutional block, and the fourth convolutional block are first connected in series in sequence, and then each convolutional block is respectively connected to its corresponding detection head;
[0050] The number of convolutional layers in the first convolutional block is set to 10, and the kernel sizes of these 10 convolutional layers are respectively set to 1×1, 3×3, 1×1, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1;
[0051] The number of convolutional layers in the second convolutional block is set to 9, and the kernel sizes of these 9 convolutional layers are respectively set to 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1;
[0052] The number of convolutional layers in the third convolutional block is set to 10, and the kernel sizes of these 10 convolutional layers are respectively set to 1×1, 3×3, 1×1, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1;
[0053] The number of convolutional layers in the fourth convolutional block is set to 4, and the kernel sizes of these 4 convolutional layers are respectively set to 1×1, 3×3, 1×1, 1×1;
[0054] Among the four detection heads, each detection head contains two parallel convolutional layers, and the kernel size of each convolutional layer is set to 3×3.
[0055] The input of this detection sub-network is the output of the feature extraction sub-network, and the output is the detection result of the SAR image; during the training process, add the location regression loss and the class loss to constrain the detection sub-network. These two loss functions are expressed as follows:
[0056]
[0057]
[0058]
[0059] Among them, I represents the total number of target boxes output by the network, i represents the serial number of the target box output by the network, J represents the total number of ground truth boxes manually marked, and j represents the serial number of the ground truth box manually marked; x ij represents the matching state between the i-th target box output by the network and the j-th ground truth box. If x ij takes the value of 0, it means non-matching, and if it is 1, it means matching; l i represents the i-th target box output by the network, represents the j-th ground truth box manually marked, z i represents the true class label corresponding to the i-th target box output by the network, s i represents the predicted class probability of the network for the i-th target box;
[0060] 2.3) Construct a scene feature learning sub-network:
[0061] Build a scene feature learning sub-network composed of a parallel connection of a scene recognition module and a scene aggregation module;
[0062] The scene recognition module is composed of a pooling layer and two fully connected layers, and its structure is in turn: the first pooling layer, the first fully connected layer, and the second fully connected layer; the pooling area size of the first pooling layer is set to 2×2, and the number of nodes of the first and second fully connected layers are set to 512 and 3 respectively;
[0063] The scene aggregation module is composed of a pooling layer and two fully connected layers, and its structure is in turn: the second pooling layer, the third fully connected layer, and the fourth fully connected layer; the pooling area size of the second pooling layer is set to 2×2, and the number of nodes of the third and fourth fully connected layers are set to 512 and 128 respectively.
[0064] The inputs of the scene recognition module and the scene aggregation module are both the outputs of the feature extraction sub-network. The output of the scene recognition module is the scene recognition result of the SAR image, and the output of the scene aggregation module is the embedded feature of the SAR image. During the training process, add scene classification loss and scene aggregation loss to constrain the scene recognition module and the scene aggregation module respectively. These two loss functions are expressed as follows:
[0065]
[0066]
[0067] Among them, N represents the total number of SAR images in the same minibatch, p, q, and n represent three serial numbers of SAR images in the same minibatch, sr pDenote the scene recognition result of the scene recognition module in the semi - supervised SAR ship detection network for the \(p\) - th SAR image. and respectively denote the scene - level annotations of the \(p\) - th and \(q\) - th SAR images, \(z\) p , \(z\) q and \(z\) n respectively denote the embedding features of the \(p\) - th, \(q\) - th and \(n\) - th SAR images output by the scene aggregation module in the semi - supervised SAR ship detection network. exp represents the exponential operation, and \(\tau\) is the temperature coefficient, whose value is taken as 0.07; \(I\) [q≠p] \(\in\{0,1\}\), \(I\) [n≠p] \(\in\{0,1\}\) is an indicator function, and its expression is as follows:
[0068]
[0069]
[0070]
[0071] where denotes the cosine similarity between \(z\) p and \(z\) q , and \(T\) represents the transpose of the vector;
[0072] 2.4) Connect the scene feature learning sub - network and the detection sub - network in parallel, and then connect them in series with the feature extraction sub - network to form a semi - supervised SAR ship detection network.
[0073] Step 3: Train the semi - supervised SAR ship detection network.
[0074] Input the training set into the semi - supervised SAR ship detection network, and use the stochastic gradient descent algorithm to iteratively update the weight values of the network, optimize the total loss function of the network until it converges, and obtain the trained semi - supervised SAR ship detection network. The specific implementation of this step is as follows:
[0075] 3.1) Randomly initialize the weight values \(\omega\) of the semi - supervised SAR ship detection network;
[0076] 3.2) Input the training set into the semi - supervised SAR ship detection network to obtain the predicted output of the network;
[0077] 3.3) Calculate the total loss function according to the predicted output of the network and the true annotation
[0078]
[0079] where Represents the position loss between the target bounding boxes output by the detection subnet in the semi-supervised SAR ship detection network and the marked ground truth boxes. Represents the class loss of the target bounding boxes output by the detection subnet in the semi-supervised SAR ship detection network. Represents the scene classification loss between the scene category of the SAR image output by the scene recognition module in the semi-supervised SAR ship detection network and the scene-level annotation. Represents the scene aggregation loss calculated from all SAR images belonging to the same minibatch in the scene aggregation module of the semi-supervised SAR ship detection network. α represents the weight of the scene classification loss function, and β represents the weight of the scene aggregation loss function. These two weight values are both taken in the range of α ∈ [0,1], β ∈ [0,1] according to the dimension of each loss function, and α ≠ β.
[0080] 3.4) According to the total loss function Take the partial derivative of the weight value ω of the network
[0081] 3.5) According to the formula Update the network weight value ω, where μ is the learning rate parameter.
[0082] 3.6) Repeat steps 3.2) to 3.5) a total of 120,000 times or end after the total loss function Converges to obtain the trained semi-supervised SAR ship detection network.
[0083] Step 4, detect the position of the target bounding box in the SAR image to be tested.
[0084] 4.1) Use a sliding window operation with a window size of 512×512 to crop each large-scale SAR image to be tested into multiple sub-images with a pixel size of 512×512.
[0085] 4.2) Input each test sub-image into the trained feature extraction subnet and scene recognition module in turn to obtain the scene recognition result of the test sub-image.
[0086] 4.3) According to the scene recognition result, obtain the target bounding box position and target bounding box class:
[0087] For the test sub-image with the scene recognition result of the inland scene, the detection result is output as no target.
[0088] For the test sub-image with the scene recognition result of the nearshore scene, input the test sub-image into the detection subnet and set the nearshore detection threshold th in And compare the confidence score of the target bounding box output by the detection subnet with the nearshore detection threshold th inMake a comparison: the confidence score is greater than the inshore detection threshold th in The target bounding boxes are retained, and those less than the inshore detection threshold th in The target bounding boxes are discarded, so as to obtain the final target bounding box positions and target bounding box categories in this test sub-image;
[0089] For a test sub-image with a scene recognition result of an open sea scene, input this test sub-image into the detection sub-network and set the open sea detection threshold th off , and compare the confidence score of the target bounding box output by the detection sub-network with this open sea detection threshold th off Make a comparison: the confidence score is greater than the open sea detection threshold th off The target bounding boxes are retained, and those less than the open sea detection threshold th off The target bounding boxes are discarded, so as to obtain the final target bounding box positions and target bounding box categories in this test sub-image.
[0090] Step 5, obtain the large-scale SAR image after target detection.
[0091] In accordance with the order of the sliding window, map the target bounding box positions of each test sub-image to the corresponding positions of each large-scale SAR image to be tested, and mark out the target bounding boxes in the large-scale SAR image according to the mapped target bounding box positions, so as to obtain the large-scale SAR image after target detection.
[0092] Next, the effect of the present invention will be further described in combination with simulation experiments.
[0093] 1. Simulation experiment conditions:
[0094] The hardware platform for the simulation experiment of the present invention is: the processor is an Intel Xeon Silver 4114 CPU, the processor main frequency is 2.20 GHz, the memory is 128 GB, and the graphics card is an NVIDIA GTX 2080Ti.
[0095] The software platform for the simulation experiment of the present invention is: ubuntu 16.04 LTS operating system, Pytorch, Python 3.6.
[0096] The dataset used in the simulation experiment of the present invention is the AIR-SARShip-1.0 measured dataset. There are ship targets in the AIR-SARShip-1.0 data images, and it also contains complex backgrounds, such as: sea surface, port facilities, buildings, grasslands, trees. In this experiment, the ship targets therein are used as the detection targets.
[0097] The dataset contains 31 original large images, each with a size of 3000×3000 pixels and in the tiff image format. In this experiment, 21 images were selected from the original 31 SAR images as training images, and the remaining 10 images were selected as test images. The original training and test SAR images were respectively cropped to obtain sub-images with a size of 512×512.
[0098] 2. Contents and result analysis of the simulation experiment:
[0099] An existing technology RefineDet used in the simulation experiment refers to the object detection model proposed by S. Zhang et al. in "Single-Shot Refinement Neural Network for Object Detection". The detection results are as Figure 3 shown.
[0100] Contents of the simulation:
[0101] Object detection was respectively performed on two test original large images in the AIR-SARShip-1.0 dataset using the present invention and the existing technology RefineDet. The detection results are as Figure 3 shown. Among them:
[0102] Figure 3 (a) shows the detection result of the existing technology RefineDet for the first test original large image in the AIR-SARShip-1.0 dataset;
[0103] Figure 3 (b) shows the detection result of the existing technology RefineDet for the second test original large image in the AIR-SARShip-1.0 dataset;
[0104] Figure 3 (c) shows the detection result of the present invention for the first test original large image in the AIR-SARShip-1.0 dataset,
[0105] Figure 3 (d) shows the detection result of the present invention for the second test original large image in the AIR-SARShip-1.0 dataset.
[0106] Figure 3 The solid rectangular frames indicate the correct detection results, the solid circular frames indicate the incorrect detection results, and the dashed rectangular frames indicate the missed ship targets.
[0107] From Figure 3 (a) and Figure 3(b) It can be seen that there are a large number of solid circular boxes in the detection result diagram of the prior art RefineDet, that is, false detection results, namely false alarms, and a large number of dashed rectangular boxes, that is, missed detection of ship targets, missed alarms. Compared with this, Figure 3 (c) and Figure 3 the solid circular boxes and dashed rectangular boxes in (d) are greatly reduced, that is, false alarms and missed alarms are greatly reduced.
[0108] From Figure 3 (c) and Figure 3 it can also be found in (d) that the present invention does not generate any false alarms in inland areas, while there are only a small number of false alarms in some nearshore and open sea areas. This is because the scattering results of some port facilities in the nearshore area and the sea clutter structure in the sea area are extremely similar to those of ship targets, so false alarms are very likely to occur. The present invention has a small number of missed alarms in the nearshore area and the open sea area. This is because the arrangement of some ship targets in these areas is relatively dense and the scattering intensity of some ship targets is low, which brings certain difficulties to detection.
[0109] Comparing Figure 3 (a) with Figure 3 (c), Figure 3 (b) with Figure 3 (d) in the detection result diagrams, it can be found that the method of the present invention can effectively reduce the number of false alarms and missed alarms in target detection and improve the accuracy of SAR ship detection.
[0110] In order to further verify the simulation effect of the present invention, the F1-score formula is used to evaluate the detection results of the two methods respectively:
[0111]
[0112] Among them, represents the detection precision rate, and the higher its value, the fewer false alarms in the detection results;
[0113] represents the detection recall rate, and the higher its value, the fewer missed alarms in the detection results.
[0114] The statistical results of all the calculation results of the above two methods are shown in Table 1:
[0115] Table 1. Quantitative analysis table of the detection results of the present invention and the prior art in the simulation experiment
[0116]
[0117] It can be seen from Table 1 that the F1-score of the present invention has increased by 6.73% compared with the prior art RefineDet, indicating that the present invention has better detection performance than the prior art.
[0118] The above simulation experiments show that: for a semi-supervised SAR ship detection method based on scene feature learning proposed by the present invention, since it can simultaneously use the object-level annotation and scene-level annotation of SAR images to train the network, it alleviates the problem of excessive dependence of network training on object-level annotation, and improves the SAR ship detection performance under the condition of a small number of object-level labels. At the same time, the present invention proposes a hierarchical testing process from scene to object. Since it fully considers the scene features of SAR images and sets different detection strategies for SAR images in different scenes, it significantly reduces the inland false alarms, near-shore false alarms and open sea missed alarms, thereby improving the SAR ship detection performance.
Claims
1. A semi-supervised SAR ship detection method based on scene feature learning, characterized in that, It includes the following steps: (1) Generate a training set: Collect at least 21 large-scale SAR images, and crop each large-scale SAR image into multiple sub-images of size 512×512; randomly select 30% of the sub-images containing ship targets for object-level marking and scene-level marking, and only perform scene-level marking on the remaining sub-images, and form a training set with all the marked sub-images; (2) Construct a semi-supervised SAR ship detection network: (2a) Build a feature extraction sub-network composed of eight cascaded convolutional blocks; (2b) Build a detection sub-network composed of four convolutional blocks and four detection heads, where the four convolutional blocks are first connected in series in sequence, and then each convolutional block is respectively connected to its corresponding detection head; (2c) Build a scene feature learning sub-network composed of a scene recognition module and a scene aggregation module in parallel; (2d) Connect the scene feature learning sub-network and the detection sub-network in parallel, and then connect them in series with the feature extraction sub-network to form a semi-supervised SAR ship detection network; (3) Input the training set into the semi-supervised SAR ship detection network, use the stochastic gradient descent algorithm to iteratively update the weight values of the network, and optimize the total loss function of the network until it converges to obtain a trained semi-supervised SAR ship detection network; (4) Detect the position of the target box in the SAR image to be tested: (4a) Slide and crop the large-scale SAR image to be tested into multiple sub-images of size 512×512; (4b) Input each test sub-image into the trained feature extraction sub-network and the scene recognition module in sequence to obtain the scene recognition result of the test sub-image; (4c) According to the scene recognition result, obtain the position and category of the target box: For the test sub-image with the scene recognition result being an inland scene, output the detection result as no target; For the test sub-image whose scene recognition result is the inshore scene, input the test sub-image into the detection sub-network and set the inshore detection threshold th in , and obtain the target box position and target box category of the test sub-image; For the test sub-image with the scene recognition result being the open sea scene, input the test sub-image into the detection sub-network and set the open sea detection threshold th off , and obtain the target box position and target box category of the test sub-image; (5) Map the position of the target box of each test sub-image to the corresponding position of each large-scale SAR image to be tested according to the order of the sliding window, and obtain the ship detection result of the large-scale SAR image.
2. The method according to claim 1, wherein: In the feature extraction sub-network built in step (2a), the structural parameters of each convolutional block are as follows: The first convolutional block contains a convolutional layer and a pooling layer, where the number of convolutional layers is set to 1, the convolutional kernel size of this convolutional layer is set to 7×7, the number of pooling layers is set to 1, and the pooling area size is set to 2×2; The number of convolutional layers in the second convolutional block is set to 4, and the convolutional kernel sizes of these 4 convolutional layers are 1×1, 3×3, 1×1, and 1×1 respectively; The number of convolutional layers in the third convolutional block is set to 3, and the convolutional kernel sizes of these 3 convolutional layers are 1×1, 3×3, and 1×1 respectively; The number of convolutional layers and the convolutional kernel sizes of the fourth convolutional block, the sixth convolutional block, the seventh convolutional block, and the eighth convolutional block are the same as those of the third convolutional block; The number of convolutional layers and the convolutional kernel sizes of the fifth convolutional block are the same as those of the second convolutional block.
3. The method according to claim 1, characterized in that: In the detection sub-network built in step (2b), the structural parameters of each convolutional block are as follows: The number of convolutional layers in the first convolutional block is set to 10, and the kernel sizes of these 10 convolutional layers are set to 1×1, 3×3, 1×1, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1 respectively; The number of convolutional layers in the second convolutional block is set to 9, and the kernel sizes of these 9 convolutional layers are set to 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1 respectively; The number of convolutional layers in the third convolutional block is set to 10, and the kernel sizes of these 10 convolutional layers are set to 1×1, 3×3, 1×1, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1 respectively; The number of convolutional layers in the fourth convolutional block is set to 4, and the kernel sizes of these 4 convolutional layers are set to 1×1, 3×3, 1×1, 1×1 respectively; In the four detection heads, each detection head contains two convolutional layers in parallel, and the size of each convolutional kernel is 3×3.
4. The method according to claim 1, wherein: In the scene feature learning sub-network constructed in step (2c), the structural parameters of each module are as follows: The scene recognition module consists of a pooling layer and two fully connected layers, and its structure is in turn: the first pooling layer, the first fully connected layer, the second fully connected layer; the pooling area size of the first pooling layer is set to 2×2, and the number of nodes in the first and second fully connected layers are set to 512 and 3 respectively; The scene aggregation module consists of a pooling layer and two fully connected layers, and its structure is in turn: the second pooling layer, the third fully connected layer, the fourth fully connected layer; the pooling area size of the second pooling layer is set to 2×2, and the number of nodes in the third and fourth fully connected layers are set to 512 and 128 respectively.
5. The method according to claim 1, characterized in that: The total loss function of the network in step (3) is expressed as follows: Among them, represents the total loss function of the semi-supervised SAR ship detection network, represents the location regression loss between the output target boxes of the detection subnet in the semi-supervised SAR ship detection network and the marked ground truth boxes, represents the classification loss of the output target boxes of the detection subnet in the semi-supervised SAR ship detection network, represents the scene classification loss between the scene categories of the SAR images output by the scene recognition module in the semi-supervised SAR ship detection network and the scene-level annotations, represents the scene aggregation loss calculated from all SAR images belonging to the same minibatch in the scene aggregation module of the semi-supervised SAR ship detection network. α represents the weight of the scene classification loss function, and β represents the weight of the scene aggregation loss function. These two weight values are both taken in α ∈ [0, 1], β ∈ [0, 1] according to the dimensions between the loss functions, and α ≠ β.
6. The method according to claim 5, wherein The position regression loss is calculated by the following formula: Among them, I represents the total number of target boxes output by the network, i represents the serial number of the target box output by the network, J represents the total number of ground truth boxes marked manually, and j represents the serial number of the ground truth box marked manually; x ij represents the matching status between the i-th target box output by the network and the j-th ground truth box. If x ij takes the value of 0, it means non-matching, and if it is 1, it means matching; l i represents the i-th target box output by the network, and represents the j-th ground truth box marked manually.
7. The method according to claim 5, wherein The category loss is expressed as follows: Among them, I represents the total number of target boxes output by the network, i represents the serial number of the target box output by the network, and z i represents the true class label corresponding to the i-th target box output by the network, and s i represents the predicted class probability of the network for the i-th target box.
8. The method according to claim 5, characterized in that, The scene classification loss is expressed as follows: Among them, N represents the total number of SAR images in the same minibatch, p represents the serial number of the SAR image in the same minibatch, and sr p represents the scene recognition result of the scene recognition module in the semi-supervised SAR ship detection network for the p-th SAR image, represents the scene-level annotation of the p-th SAR image.
9. The method according to claim 5, wherein The scene aggregation loss is expressed as follows: Among them, N represents the total number of SAR images in the same minibatch, and p, q, and n represent three serial numbers of SAR images in the same minibatch, z p , z q and z n respectively represent the embedded features of the p-th, q-th, and n-th SAR images output by the scene aggregation module in the semi-supervised SAR ship detection network. exp represents the exponential operation, τ is the temperature coefficient, and its value is taken as 0.07; I [q≠p] ∈ {0, 1}, I [n≠p] ∈ {0, 1} is an indicator function, and its expression is as follows: Among them, and respectively represent the scene-level annotations of the p-th and q-th SAR images; represents z p and z q 's cosine similarity, and T represents the transpose of the vector.
10. The method according to claim 1, characterized in that: In step (3), the random gradient descent algorithm is used to iteratively update the weight values of the network. The steps are as follows: (3a) Randomly initialize the weight values ω of the semi-supervised SAR ship detection network; (3b) Input the training set into the semi-supervised SAR ship detection network to obtain the predicted output of the network; (3c) Calculate the total loss function based on the predicted output and true annotation of the network (3d)Derive the partial derivative of the weight value ω of the network with respect to the total loss function (3e) Update the network weight value ω according to the formula where μ is the learning rate parameter; (3f) Repeat steps (3b) to (3e) 120,000 times or until the total loss function converges and then end.
Citation Information
Patent Citations
SAR Image Target Detection Method Based on Convolutional Neural Network
CN106228124B
An SAR image target detection system and method based on a semi-supervised CNN
CN109740549A
Scene perception data enhancement method for SAR ship detection
CN113902975A
Ship image detection method and device and storage medium
CN115082781A