CFAR-Guided YOLO Network SAR Image Target Detection Method
By introducing the weight guidance map of CFAR detection in the YOLO network, the network's target detection accuracy problem in complex environments is solved, and higher detection accuracy and higher detection efficiency are achieved.
Patent Information
- Application Number
- CN202310146807.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-02-21
AI Technical Summary
The accuracy of object detection in the YOLO network in the context of complex environments is affected, mainly due to the deformity of the loss function caused by the serious imbalance of positive and negative samples.
Using a weight guidance diagram based on CFAR detection, by correcting the confidence loss and loss functions, we ensure that positive and difficult-to-separate negative samples receive appropriate attention during the training process.
It effectively improves the target detection performance of YOLO network in complex environments, improves the detection accuracy, simplifies the detection process, and improves the detection efficiency.
Smart Images

Figure CN116258974B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning and radar, and particularly relates to a method for detecting SAR image targets based on a YOLO network guided by CFAR. Background Art
[0002] Constant False Alarm Rate (CFAR) is a relatively mature traditional detection method in the target detection task of Synthetic Aperture Radar (SAR) images. During the CFAR detection process, the size of the working window is first determined according to the prior information of the size of the target to be detected. The working window includes a protection window and a background window. A fixed false alarm probability is set according to requirements or experiments. During detection, the pixel positions in the image are traversed. Based on the pixel information within the working window, the data around the current position to be detected is modeled, and its statistical characteristics are calculated. According to the set false alarm rate, an adaptive threshold based on the environmental information around the position to be detected can be obtained. The data at the position to be detected is compared with the obtained adaptive threshold to obtain a determination result on whether the pixel is a target pixel. By detecting all pixel positions on the SAR image, a binary result map can be obtained, and this binary map represents the determination of whether each pixel position belongs to a target pixel by the CFAR algorithm. The binary map obtained by CFAR processing usually also undergoes morphological filtering operations, such as dilation and erosion operations, to eliminate detection regions of certain shapes and areas.
[0003] As a pixel-level target detection method, the detection performance of CFAR is greatly affected by noise and the contrast between the target and clutter. If the target is under a relatively uniform background clutter, such as a plain or the sea surface, the algorithm has good detection performance; if the target is in a more complex environmental background, such as complexly arranged buildings, etc., the accuracy of the detection result will be affected to a certain extent.
[0004] The YOLO network is a single-stage image target detection deep convolutional neural network based on anchor boxes and can be used for target detection tasks on SAR images. Generally, YOLO contains multiple output layers of different sizes and scales. For an output layer with a size of n×n and the number of anchor boxes set to m in multiple output layers, the meaning of its detection and output data can be explained as follows: this output layer represents dividing the image into n×n regions, each region is called a cell unit (cell), and each cell unit is responsible for predicting the target whose center point coordinates are located inside the unit. During detection, each cell unit will generate m basic detection boxes (anchor boxes). The network classifies whether there is a target in the anchor box. If it is determined that there is a target, the position and size of the anchor box are further corrected, and the target inside the box is classified, and finally the detection result is output.
[0005] During the training process of YOLO, first, for each image, a matching ground-truth label matrix for calculating the loss of the network output is established according to its real target information. After the network completes the forward propagation, the confidence loss, classification loss, and regression loss of the network output are calculated respectively. Among them, the confidence loss is used to evaluate the accuracy of the network's judgment on whether there are target objects of any category in the anchor boxes. Since there are usually only a few targets in an image, and the number of anchor boxes is much larger than the number of targets, when the network performs the classification discrimination of whether there are targets, the number of positive samples (with targets) and negative samples (without targets) is seriously unbalanced. The serious imbalance in the number of samples of different categories affects the performance surface of the network's loss function, ultimately affecting the performance of the network after convergence, resulting in low target detection efficiency and affecting detection accuracy. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a CFAR-guided YOLO network SAR image target detection method. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0007] An embodiment of the present invention provides a CFAR-guided YOLO network SAR image target detection method, including the steps of:
[0008] S1. Perform CFAR target detection on the SAR input images in the training set that are the same size as the input images of the YOLO network, and add corrections of real targets on the CFAR target detection map according to the real target annotation information to obtain a corrected CFAR annotation map;
[0009] S2. Downsample the corrected CFAR annotation map according to the size of each output layer in the YOLO network to obtain a weight guidance map for each output layer;
[0010] S3. Use the training set to train the YOLO network. When calculating the confidence loss of all detection boxes after each forward propagation, different weight coefficients are applied to the detection boxes output by the output layer using the weight guidance map of each output layer to obtain a corrected confidence loss, and a corrected loss function is calculated using the corrected confidence loss;
[0011] S4. Use the corrected loss function to optimize the parameters of the YOLO network belonging to it, and repeat step S3 in each iteration process to obtain the network training weights, thereby obtaining a trained YOLO network. Then, use the trained YOLO network to perform target detection on the SAR image to be detected to obtain the detection targets.
[0012] In an embodiment of the present invention, step S1 includes:
[0013] S11. Slice or scale the original SAR image to form a SAR input image with the same size as the input image of the YOLO network, and form a training set with each SAR input image and the corresponding ground truth annotation information of the SAR input image;
[0014] S12. According to the ground truth annotation information, statistically analyze the size distribution of the targets in the training set, and select the CFAR window length parameter according to the size statistical information of the targets;
[0015] S13. Using the CFAR window length parameter, use the CFAR detection method to detect the SAR input images in the training set, and obtain a CFAR detection binary map with the same size as the SAR input images;
[0016] S14. Draw the ground truth boxes on the CFAR detection binary map according to the ground truth annotation information, and fill the ground truth boxes with the value 1 to form a binary image with the same size and corresponding position as the SAR input image, and obtain the corrected CFAR annotation map.
[0017] In an embodiment of the present invention, step S2 includes:
[0018] Divide the corrected CFAR annotation map of the first size into a number of cells of the second size with the first side length, each cell corresponding to a pixel position in the weight guidance map, and the number of cells being the same as the size of each output layer in the YOLO network;
[0019] Statistically analyze the number of pixel points with the value 1 in each cell. If the proportion of pixel points with the value 1 is greater than or equal to 0.5, the pixel at the corresponding pixel position in the weight guidance map of the cell is 1. If the proportion of pixel points with the value 1 is less than 0.5, the pixel at the corresponding pixel position in the weight guidance map of the cell is 0, thereby obtaining the weight guidance map.
[0020] In an embodiment of the present invention, before obtaining the weight guidance map, it further includes:
[0021] For each training image in the training set in the weight guidance map of each output layer, the pixels corresponding to the positive samples and the difficult negative samples screened by the CFAR detection method are 1, and the pixels corresponding to the simple negative samples in the corrected CFAR annotation map are 0.
[0022] In an embodiment of the present invention, in step S3, when training the YOLO network using the training set and calculating the confidence loss of all detection frames after each forward propagation, different weight coefficients are applied to the detection frames output by the output layer using the weight guidance map of each output layer to obtain the corrected confidence loss, including:
[0023] Train the YOLO network using the training set, and obtain a number of detection frames after each forward propagation;
[0024] Obtain the weight guidance map corresponding to each output layer. When calculating the confidence loss of the detection frames, the confidence loss caused by the detection frames generated at the positions where the value in the weight guidance map is 1 is the original weight, and the confidence loss caused by the detection frames generated at the positions where the value in the weight guidance map is 0 is the product of the original weight and the weight coefficient, so as to obtain the corrected confidence loss.
[0025] In an embodiment of the present invention, the value range of the weight coefficient is 0 to 1.
[0026] In an embodiment of the present invention, in step S3, the corrected loss function is calculated using the corrected confidence loss, including:
[0027] Add the corrected confidence loss, classification loss, and bounding box regression loss to obtain the corrected loss function:
[0028]
[0029] Where is the corrected confidence loss, L cls is the classification loss, and L box is the bounding box regression loss.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] 1. The method of the present invention uses the CFAR detection method combined with the professional knowledge in the radar field to form a weight guidance map for the training process of the YOLO network. During the training process, while retaining all samples for training, different samples have different weights in the loss calculation. The training guidance makes the training process pay more attention to the distinction between positive samples and difficult-to-separate negative samples on the original basis, effectively improving the problem of the loss surface distortion caused by the serious imbalance in the number of positive and negative samples during the training process of the target detection network, thereby improving the detection performance of the network and the accuracy of the YOLO network for image target detection.
[0032] 2. The method of the present invention is based on modifying the calculation process of the loss function during the network training process. The direct aim is to train better weights under the same network structure as the original one, without the need to modify the structure and data interface of the original detection network. For the newly obtained training weights after improvement, only the new weights need to be used to replace the original network during the testing and application phases. Compared with other network improvement measures that may require network structure adjustment and data interface adjustment, this method is more concise and has higher detection efficiency when performing image target detection.
[0033] 3. The method of the present invention starts from the interpretive meaning of the data in the output layer of the YOLO network, that is, the data position distribution in the network output layer represents the position information of the detection box samples. It proposes to use the CFAR detection method combining professional knowledge in the radar field to guide the network training process. Each intermediate process of the implementation steps has clear meanings and purposes, and has better interpretability compared with traditional network improvement strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic flowchart of a SAR image target detection method for a YOLO network guided by CFAR provided by an embodiment of the present invention;
[0035] Figure 2 It is a schematic diagram of the steps of a YOLO network training method for SAR image target detection guided by CFAR provided by an embodiment of the present invention;
[0036] Figure 3 It is a schematic diagram of the results of step S1 and step S2 of an embodiment of the present invention;
[0037] Figure 4 It is a schematic diagram of the downsampling process for generating the final weight guidance map from the CFAR results after adding GT during the generation process of the weight guidance map provided by an embodiment of the present invention;
[0038] Figure 5 It is a schematic diagram of the CFAR annotation map guiding the calculation of the loss function of the output layer provided by an embodiment of the present invention;
[0039] Figures 6a - 6b It is a schematic diagram of the experimental process and results of the verification experiment provided by an embodiment of the present invention;
[0040] Figures 7a - 7b It is a performance comparison diagram of detection on the same test set after training with the original training method of YOLO and the method of the present invention on the same dataset;
[0041] Figures 8a - 8b It is a visual comparison diagram of the test results of detection on the same test set after training with the original training method of YOLO and the method of the present invention on the same dataset. Specific Embodiments
[0042] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0043] Embodiment 1
[0044] For the specific task of using the YOLO network for target detection in SAR images, the present invention makes the following improvements. Since there is a corresponding relationship between the data of each output layer of the YOLO network and the position of the cell unit, relevant professional technical knowledge in the radar field, CFAR detection, can be used to provide prior knowledge guidance for the training of YOLO. Different weights are used for samples in different regions when calculating the loss, reducing the interference of a large number of simple negative samples on network training, and making the network training process pay more attention to the discrimination of positive samples and difficult-to-separate negative samples. For the training set SAR input image I ori perform CFAR target detection, and each training set image will obtain a corresponding CFAR detection binary map I CFAR , and supplement and correct the true target information for this map to obtain According to the network structure, the size of the i-th output layer of YOLO is n×n. Downsample to the size of n×n to obtain the weight guidance map G of the network output layer i . When calculating the confidence loss, different weights are applied to each sample according to G i to correct the performance surface of the network loss function to obtain a better training convergence result.
[0045] Please refer to Figure 1 and Figure 2 , Figure 1 which is a schematic flow chart of a method for SAR image target detection based on a CFAR-guided YOLO network provided by an embodiment of the present invention,[[]]END]] Figure 2 which is a schematic diagram of the steps of a method for training a CFAR-guided SAR image target detection YOLO network provided by an embodiment of the present invention. Among them, Dataloader is a data reader used to provide data and related information during training. Images is the input image, GT represents the label, and Images Paths represents the image storage path, which is used to confirm the images used in this round of training according to this information to read the weight indication map of the corresponding training images. Figure 2 The main improvement work on the YOLO network in this embodiment is within the dashed box. In the step of generating the CFAR annotation map, the CFAR annotation map is constructed according to the following steps S1 and S2. In the step of computing loss, the loss of the YOLO network is calculated by guiding the CFAR annotation map according to the following step S3.
[0046] The object detection method for SAR images based on the CFAR-guided YOLO network in this embodiment includes the following steps:
[0047] S1. Perform CFAR object detection on the SAR input image I with the same size as the input image of the YOLO network in the training set ori and add corrections for real objects on the CFAR object detection map I CFAR according to the real annotation information of the objects to obtain a corrected CFAR annotation map Specifically, it includes the following steps:
[0048] S11. Slice or scale the original SAR image to form a SAR input image with the same size as the input image of the YOLO network, and form a training set from each SAR input image and the corresponding real annotation information of the objects in the SAR input image.
[0049] Specifically, slice or scale the original SAR image used for training so that the size of the image used for training is the same as the size of the input image of the YOLO network used, denoted as SAR input image I ori . Each SAR input image I ori has corresponding real annotation information of the objects. The training set used for training is composed of the SAR input image I ori and the real annotation information (label file) of the objects in the SAR input image I ori .
[0050] S12. Statistically analyze the size distribution of the objects in the training set according to the real annotation information of the objects, and select the CFAR window length parameter according to the size statistical information of the objects.
[0051] Specifically, according to the real annotation information of the objects, statistically analyze the size distribution of the objects in the training set. According to the size statistical information of the objects, select an appropriate CFAR window length parameter. According to the requirements of the CFAR algorithm, when detecting each target pixel point, the remaining target pixel points should fall within the protection window, and the background window is the background pixels around the object. In practical applications, it is appropriate to select a detection window length slightly larger than twice the side length of the vast majority of object frames in the training dataset.
[0052] S13. Use the CFAR window length parameter and the CFAR detection method to detect the SAR input images in the training set to obtain a CFAR detection binary map with the same size as the SAR input image.
[0053] Specifically, use the CFAR window length parameter and the CFAR detection method to detect the SAR input image I ori in the training set to obtain a result with the same size as the SAR input image Iori CFAR detection binary image I with the same size CFAR 。CFAR detection binary image I CFAR is a binary image with the same size as the SAR input image I ori and corresponding pixel positions. In the CFAR detection binary image I CFAR the pixel positions determined as target pixel points by the CFAR detection algorithm are 1 (white), and the pixel positions determined as background pixels by the detection algorithm are 0 (black).
[0054] S14. Draw a true target box on the CFAR detection binary image according to the true target annotation information, and fill the true target box with the value 1 to form a binary image with the same size and corresponding position as the SAR input image, obtaining a corrected CFAR annotation map.
[0055] Specifically, to prevent the loss of positive samples for network training in subsequent steps due to unsatisfactory CFAR detection results, draw a true target box on the CFAR detection binary image according to the true target annotation information of the image, and fill the true target box with the value 1 (white) to obtain a corrected CFAR annotation map The corrected CFAR annotation map is still a binary image with the same size and corresponding position as the SAR input image I ori and is a correction of the CFAR detection binary image I CFAR to ensure that in the subsequent generated indication map, the true target positive sample positions are marked as positive to ensure the use of the expected normal loss weights, and at the same time, there is a certain suppression of the data that deviates too much from the true target generated by the YOLO data augmentation strategy.
[0056] The correction step of this embodiment ensures that effective positive samples can be fully utilized, and at the same time, it can also avoid the influence of inappropriate positive samples generated by the YOLO data augmentation method, ensuring that the positive samples and the false alarms detected by CFAR participate in the loss calculation as difficult negative samples with the expected larger weights.
[0057] S2. Downsample the corrected CFAR annotation map according to the size of each output layer in the YOLO network to obtain a weight guidance map for each output layer.
[0058] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the results of step S1 and step S2 of the embodiment of the present invention. Specifically, for the corrected CFAR annotation map Downsample according to the size of the YOLO output layer. Based on the structure of the actually used YOLO detection network, obtain the size of each YOLO detection output layer. YOLO networks with different structures generally have 2 to 3 detection output layers to adapt to the detection of targets of different sizes. For each SAR input image in the training set For each output layer, downsampling of the corresponding size is required. For the i-th detection output layer, there is a corresponding downsampling result, that is, the weight guidance map G i . The weight guidance map G i is a binary map with the same size as the i-th output layer of the YOLO detection network, and is called the weight guidance map of the i-th output layer. The description of the downsampling operation for the corrected CFAR annotation map is as follows.
[0059] The SAR input image I of the network ori has a size of m×m, and the size of the corrected CFAR annotation map is also m×m. The size of the i-th output layer of the YOLO network is n×n. The output layer only considers the size of the image, and different channels are not considered as the representation of the detection box. Thus, the downsampling ratio of the i-th output layer of the network, that is, the downsampling ratio that the corrected CFAR annotation map corresponding to this output layer needs to adopt is required to downsample the corrected CFAR annotation map of size m×m to a binary map G of size n×n i . Please refer to Figure 4 . Figure 4 is a schematic diagram of the downsampling process for generating the final weight guidance map from the CFAR result after adding GT in the process of generating the weight guidance map provided by the embodiment of the present invention. Further, first divide the corrected CFAR annotation map of the first size into several cells of the second size with the first side length. Each cell corresponds to a pixel position in the weight guidance map. The number of cells is the same as the size of each output layer in the YOLO network, that is, divide the corrected CFAR annotation map of the first size of m×m into n×n cells of size k×k of the second size with k as the first side length. At this time, each cell corresponds to a pixel position in the weight guidance map G i ; then, count the number of pixel points with a value of 1 in each cell. If the ratio of the number of pixel points with a value of 1 is greater than or equal to 0.5, the pixel at the corresponding pixel position in the weight guidance map G i is 1. If the ratio of the number of pixel points with a value of 1 is less than 0.5, the pixel at the corresponding pixel position in the weight guidance map is 0. After counting and mapping all cells, the downsampled image, that is, the weight guidance map G, can be obtained.i 。
[0060] The YOLO network often has more than one detection output layer with different scales, and it is necessary to perform corresponding downsampling of different sizes on the image. The conventional YOLO detection network has 3 detection output layers. For the nth training image in the training set of the corrected CFAR annotation map needs to be downsampled 3 times respectively, each time downsampled to the size corresponding to the output layer, to obtain the weight guidance map of the SAR input image in the training set for different detection output layers and
[0061] In addition, when forming the weight guidance map, in the weight guidance map of each training image in the training set for each output layer, the pixels corresponding to the positive samples and the difficult negative samples screened by the CFAR detection method are 1, and the pixels corresponding to the simple negative samples in the corrected CFAR annotation map are 0. Specifically, because the weight guidance maps are the downsampling results of the corrected CFAR annotation map and the corrected CFAR annotation map has been corrected at the real target position, then in the corresponding weight guidance map the positions with targets (positive samples) are 1 after downsampling. The false alarms of CFAR are used as the difficult negative samples screened by the CFAR algorithm, and the positions with difficult negative samples in the weight guidance map are also 1, while the remaining positions that are 0 in the corrected CFAR annotation map are judged as simple negative samples and are also 0 at the corresponding positions in the weight guidance map G i (n) 。
[0062] S3. Use the training set to train the YOLO network. When calculating the confidence loss of all detection frames after each forward propagation, use the weight guidance map of each output layer to apply different weight coefficients to the detection frames output by the output layer to obtain the corrected confidence loss, and use the corrected confidence loss to calculate the corrected loss function.
[0063] Specifically, according to the weight guidance map G i of the output layer, respectively guide the loss calculation of the training samples corresponding to the output layer after forward propagation, and its working principle is as Figure 5 shown Figure 5Schematic diagram for calculating the loss function of the output layer guided by the CFAR annotation map provided by the embodiments of the present invention. The size of the CFAR annotation map is the same as that of the output layer. When calculating the loss value, it plays a masking and filtering role for the loss values at different positions on the output layer, so as to apply different weights to the detection outputs of different regions, achieving the purpose of making the network pay more attention to the fewer positive samples and difficult-to-separate negative samples.
[0064] Specifically, it includes the following steps:
[0065] S31. Use the training set to train the YOLO network. After each forward propagation of the YOLO network, a series of detection box outputs are obtained.
[0066] S32. Obtain the weight guidance map corresponding to each output layer. When calculating the confidence loss of the detection box, the confidence loss caused by the detection box generated at the position where the value in the weight guidance map is 1 is the original weight, and the confidence loss caused by the detection box generated at the position where the value in the weight guidance map is 0 is the product of the original weight and the weight coefficient, so as to obtain the corrected confidence loss.
[0067] Specifically, the loss during the training process of the YOLO network is expressed as follows:
[0068] L = L obj + L cls + L box
[0069] In the formula, L obj is the confidence loss, which needs to be calculated for all samples. L cls is the classification loss, and L box is the bounding box regression loss. These two losses are only calculated for positive samples.
[0070] When the network performs forward propagation and calculates the loss for the i-th output layer, the classification loss L cls and the regression loss L box are calculated as usual. When calculating the confidence loss L obj of the i-th output layer with a size of n×n, read the corresponding weight guidance map with a size of n×n At this time, the size of the output layer being calculated is the same as that of the weight guidance map When calculating the confidence loss of the detection box output at each position of the output layer, check the weight guidance map Whether the corresponding position is 1. If it is 1, calculate the confidence loss normally. If it is 0, after calculating the confidence loss, multiply the confidence loss by a weight coefficient c less than 1. The value range of the weight coefficient c is 0 to 1. For example, the weight coefficient c can be 0.2, but the value of the weight coefficient c is not limited to 0.2. In this embodiment, the confidence loss calculated normally and the confidence loss multiplied by the weight coefficient are collectively referred to as the corrected confidence loss as
[0071] S33. Calculate the corrected loss function by using the corrected confidence loss.
[0072] Specifically, add the corrected confidence loss, the classification loss, and the bounding box regression loss to sum, and obtain the corrected loss function:
[0073]
[0074] Among them, is the corrected confidence loss, L cls is the classification loss, L box is the bounding box regression loss.
[0075] S4. Optimize the parameters of the YOLO network using the corrected loss function, and repeat step S3 in each iteration process to obtain the network training weight w, so as to obtain the trained YOLO network. Then use the trained YOLO network to perform target detection on the SAR image to be detected to obtain the detection target.
[0076] Specifically, the corrected loss function L (c) Compared with the original loss function L, the loss surface changes after being adjusted according to positive and negative samples, which is beneficial to iteratively converge to model parameters with better generalization robustness, thereby improving the detection accuracy of SAR image targets.
[0077] In this embodiment, the loss function of the YOLO target detection network is optimized to achieve the purpose of optimizing and training the problem of unbalanced positive and negative samples in the SAR image detection task, and the problem of unbalanced positive and negative samples is common to all anchor box-based target detection networks. The same idea can be applied to other anchor box-based target detection networks, such as Fast-RCNN and its series of detection networks, EfficientDet target detection networks, etc.
[0078] The method of this embodiment uses the CFAR detection method combined with professional knowledge in the radar field to form a weight guidance map for the training process of the YOLO network. During the training process, all samples are retained for training, while different samples have different weights in loss calculation. The training guidance makes the training process pay more attention to the distinction between positive samples and difficult negative samples on the original basis, effectively improving the problem of the loss surface deformation caused by the serious imbalance in the number of positive and negative samples during the training process of the target detection network, thereby improving the detection performance of the network and the accuracy of the YOLO network for image target detection.
[0079] The method of this embodiment is based on the modification of the loss function calculation process during the network training process. The direct purpose is to train better weights under the same network structure as the original. There is no need to modify the structure and data interface of the original detection network. For the new training weights obtained after improvement, only the new weights need to be used to replace the original network during the test and application phases. Compared with other network improvement measures that may require network structure adjustment and data interface adjustment, this method is more concise and has higher detection efficiency when performing image target detection.
[0080] The method of this embodiment starts from the interpretive meaning of the data in the output layer of the YOLO network, that is, the data position distribution in the network output layer represents the position information of the samples forming the detection box, and proposes to use the CFAR detection method combined with professional knowledge in the radar field to form guidance for the network training process. Each intermediate process of the implementation steps has clear meanings and purposes, and has better interpretability compared with traditional network improvement strategies.
[0081] The effects of the present invention are further verified and illustrated through the following simulation experiments.
[0082] (I) Experimental conditions
[0083] The data used in the experiment comes from the ground airport area images taken by the Gaofen-3 radar remote sensing satellite. The data resolution used is 1 meter, including a total of 2000 SAR images with sizes of 600*600 and 1024*1024, divided into six types of clearly defined model airplanes: A220(1344), A320 / 321(747), A330(104), ARJ21(524), Boeing737-800(692), Boeing787(839) and one type of other(2306) other types of airplanes, a total of 6556 airplanes. The training set is randomly divided to contain 1598 images, including 5228 targets, and the test set contains 402 images, including 1328 targets.
[0084] The model used in the experiment is the YOLOv3-tiny model, which contains two detection output layers. The downsampling rates of the two detection output layers relative to the input image are 16 and 32 respectively.
[0085] Before network training, the data is first preprocessed. The data images are sliced using a window of the network input size with an appropriate stride. When slicing the images in the training set, only the slices containing the targets are retained for subsequent training. The CFAR detection and the real label correction of the CFAR detection results are performed on all the training set slices in the manner of step S1. Then, according to step S2, based on the output layer parameters of the network used, the corrected CFAR images are downsampled 16 times and 32 times respectively according to the binary map downsampling method used in this embodiment to obtain the loss CFAR guidance maps G1 and G2 of the two output layers.
[0086] (II) Experimental content and results:
[0087] Experiment 1: Train the YOLOv3-tiny object detection network used with the traditional YOLO network training method. Input the divided test set into the trained model, and calculate the AP metrics of each category of the model obtained by the original network training method on the training set, which are A220(0.94), A320 / 321(0.91), A330(0.94), ARJ21(0.93), Boeing737-800(0.92), Boeing787(0.98), other(0.94), and calculate the overall mAP metric as 93.59%.
[0088] Experiment 2: Under the same data set division and hyperparameter settings as in Experiment 1, train the detection network used with the CFAR guidance training method of the present invention. Test the obtained model under the same conditions, and the AP metrics of each category on the test set are A220(0.94), A320 / 321(0.92), A330(0.94), ARJ21(0.91), Boeing737-800(0.95), Boeing787(0.98), other(0.95), and calculate the overall mAP metric as 94.22%.
[0089] Experiment 3: Please refer to Figures 6a - 6b , Figures 6a - 6b which is the schematic diagram of the experimental process and results of the verification experiment provided by the embodiment of the present invention, Figure 6a is the flowchart for generating the superimposed image of the prediction confidence error and the corrected CFAR map, Figure 6b and is the schematic diagram of the superimposed image of the prediction confidence error and the corrected CFAR map.
[0090] Use asFigure 6a For the process of Figure 6a , use the training weights obtained by the original method to test the test set images, draw the heat map of the confidence prediction error, and overlay this map with the CFAR image after adding the GT target box correction to obtain Figure 6b the result. By observing the overlay result of the corrected CFAR result map and the heat map of the network confidence error, it can be concluded that: in the detection results of the weights obtained by the original training method on the test images, the regions where the confidence prediction values have a large gap with the true labels basically fall in the white regions that appear after CFAR detection. The confidence prediction errors are mainly distributed in the positive part of the CFAR map marked with GT information, verifying the rationality of the YOLO network SAR image target detection method based on CFAR guidance in this embodiment.
[0091] Experiment 4: Statistically count the average number of positive and negative samples per image on the training data set before and after adding the CFAR guidance map. The training set has a total of 11,562 training slices. Each image slice generates 2,535 detection box samples in total at two output layers after passing through the network. On average, each slice contains 19.69 positive samples and 2,515.31 negative samples, and the ratio of positive to negative samples is approximately 1:128. After screening by the CFAR indication map, an average of 392.69 samples are selected per image, including 19.26 positive samples and 373.42 negative samples, and the ratio of positive to negative samples is approximately 1:19. The remaining samples that do not pass the CFAR indication map screening participate in network training with a lower weight to ensure that sample information is not lost.
[0092] Please refer to Figures 7a - 7b , Figures 7a - 7b FIG. Figures 7a - 7b is a performance comparison chart of detection on the same test set after training using the original YOLO training method and the method of the present invention on the same data set. Among them, Figure 7a is the original training method, Figure 7b is the training method of the present invention. mAP is a performance metric for object detection. The higher this metric, the better the performance of the detection model. By comparing Figure 7a and Figure 7b it can be seen that the detection method of the present invention has better performance.
[0093] Please refer to Figures 8a - 8b , Figures 8a - 8b FIG. Figures 8a - 8b is a visual comparison chart of the test results of detection on the same test set after training using the original YOLO training method and the method of the present invention on the same data set. From Figure 8a and Figure 8b it can be seen that the method of the present invention has an obvious inhibitory effect on difficult-to-distinguish false alarms during the detection process, and at the same time has an improvement effect on the detection scores of positive samples.
[0094] Comparing the results of Experiment 1, Experiment 2, Experiment 3 and Experiment 4, it can be concluded that CFAR can detect difficult-to-distinguish samples for the network to a certain extent. It is reasonable to use CFAR to guide training. The CFAR indication diagram can effectively optimize the imbalance problem of the number of positive and negative samples of the detection box samples. At the same time, the detection model obtained by the SAR image target detection YOLO network training method guided by CFAR in the present invention is significantly better than the model obtained by the training method of the original YOLO network.
[0095] In summary, the experiments verify the correctness, effectiveness and reliability of the present invention.
[0096] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A SAR image target detection method for YOLO network based on CFAR guidance, characterized in that Including the steps: S1. Perform CFAR target detection on SAR input images in the training set that are the same size as the input images of the YOLO network, and add corrections for real targets on the CFAR target detection map based on the real annotation information of the targets to obtain a corrected CFAR annotation map; S2. Downsample the corrected CFAR annotation map according to the size of each output layer in the YOLO network to obtain a weight guidance map for each output layer; S3. Use the training set to train the YOLO network. When calculating the confidence loss of all detection boxes after each forward propagation, different weight coefficients are applied to the detection boxes output by the output layer using the weight guidance map for each output layer to obtain a corrected confidence loss, and a corrected loss function is calculated using the corrected confidence loss; S4. Use the corrected loss function to optimize the parameters of the YOLO network belonging to it, and repeat step S3 during each iteration to obtain network training weights, thereby obtaining a trained YOLO network. Then, use the trained YOLO network to perform target detection on the SAR image to be detected to obtain detection targets.
2. The CFAR-guided YOLO network SAR image target detection method according to claim 1, wherein Step S1 includes: S11. Slice or scale the original SAR image to form SAR input images that are the same size as the input images of the YOLO network, and form a training set from each SAR input image and the corresponding real annotation information of the targets in the SAR input image; S12. Statistically analyze the size distribution of the targets in the training set according to the real annotation information of the targets, and select the CFAR window length parameter according to the size statistical information of the targets; S13. Use the CFAR window length parameter and the CFAR detection method to detect the SAR input images in the training set to obtain a CFAR detection binary map that is the same size as the SAR input image; S14. Draw real target boxes on the CFAR detection binary map according to the real annotation information of the targets, and fill the real target boxes with the value 1 to form a binary image that is the same size and in the same position as the SAR input image, obtaining a corrected CFAR annotation map.
3. The CFAR-guided YOLO network SAR image target detection method according to claim 1, characterized in that, Step S2 includes: Divide the corrected CFAR annotation map of the first size into several cells of the second size with the first side length. Each cell corresponds to a pixel position in the weight guidance map, and the number of cells is the same as the size of each output layer in the YOLO network; Count the number of pixel points with the value 1 in each cell. If the proportion of pixel points with the value 1 is greater than or equal to 0.5, the pixel at the corresponding pixel position in the weight guidance map for the cell is 1. If the proportion of pixel points with the value 1 is less than 0.5, the pixel at the corresponding pixel position in the weight guidance map for the cell is 0, thereby obtaining the weight guidance map.
4. The CFAR-guided YOLO network-based SAR image target detection method according to claim 3, wherein Before obtaining the weight guidance map, it also includes: In the weight guidance map of each training image in the training set for each output layer, the pixels corresponding to the positive samples and the difficult negative samples screened by the CFAR detection method are 1, and the pixels corresponding to the simple negative samples in the corrected CFAR annotation map are 0.
5. The CFAR-guided YOLO network SAR image target detection method according to claim 1, wherein In step S3, when training the YOLO network using the training set and calculating the confidence loss of all detection boxes after each forward propagation, different weight coefficients are applied to the detection boxes output by the output layer using the weight guidance map of each output layer to obtain the corrected confidence loss, including: Training the YOLO network using the training set, and obtaining a number of detection boxes after each forward propagation; Obtaining the weight guidance map corresponding to each output layer. When calculating the confidence loss of the detection boxes, the confidence loss caused by the detection boxes generated at the positions where the value in the weight guidance map is 1 is the original weight, and the confidence loss caused by the detection boxes generated at the positions where the value in the weight guidance map is 0 is the product of the original weight and the weight coefficient, so as to obtain the corrected confidence loss.
6. The CFAR-guided YOLO network-based SAR image target detection method according to claim 5, wherein The value range of the weight coefficient is 0 to 1.
7. The CFAR-guided YOLO network-based SAR image target detection method according to claim 1, characterized in that, In step S3, the corrected loss function is calculated using the corrected confidence loss, including: Adding the corrected confidence loss, the classification loss, and the bounding box regression loss to sum, to obtain the corrected loss function: Among them, is the corrected confidence loss, L cls is the classification loss, L box is the bounding box regression loss.
Citation Information
Patent Citations
On-line SAR target detection method based on deep learning
CN107563411A
A ship target detection method based on the fusion of CFAR and Fast-RCNN in SAR image
CN109145872A