Low-altitude small flying object detection method based on convolutional neural network
By applying a low-altitude small flying object detection method based on convolutional neural network in low-altitude airspace, the problems of low detection accuracy and large error in candidate box positioning in the prior art are solved, and higher detection accuracy and position accuracy are achieved.
Patent Information
- Application Number
- CN202210157359.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-02-21
AI Technical Summary
The prior art has low detection accuracy for small drones in low-altitude airspace, and there are large errors in positioning of candidate boxes, making it difficult to apply to the current situation.
The low-altitude small flying object detection method based on convolutional neural network is used to adjust the image size, use the Resnet50 feature extraction network, the K-means clustering algorithm to generate anchor boxes, the soft-nms algorithm to optimize candidate boxes, and the ROI Align layer to map candidate areas, and finally predict the results of low-altitude small flying object through the fully connected layer.
The sensitivity and accuracy of the model to small targets is improved, the detection ability of the model when the objects to be detected are dense, and the position accuracy of low-altitude small target detection is improved.
Smart Images

Figure CN114529805B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to a method for detecting low-altitude small flying objects based on a convolutional neural network. Background Art
[0002] A drone is an unmanned aircraft controlled by radio remote control. With the gradual opening of the low-altitude airspace, various types of aircraft have gradually come into the public view. With its advantages such as low cost and perfect functions, drones can meet the operation requirements of most fields, such as military reconnaissance, outdoor shooting and performance, post-disaster search and rescue, and agricultural irrigation, etc. At the same time, with the breakthroughs in various technical means such as the field of artificial intelligence, Micro Electro Mechanical System (MEMS) sensors, miniature Global Positioning System (GPS), and microelectronics technology, drone technology has gradually developed towards miniaturization and high precision. With the improvement of the utilization rate of the low-altitude airspace, many problems have followed. The situations of "unauthorized flight" and "random flight" are repeatedly prohibited. More seriously, some drone users ignore air traffic control and take off wantonly near airports, even posing a threat to aircraft. If the flight order of drones cannot be managed well, it will inevitably lead to a series of serious consequences such as air accidents. Therefore, small targets will become the difficulties and key points of future low-altitude airspace surveillance.
[0003] Before the rise of the deep learning field, traditional target detection was all realized based on digital image processing and computer vision technology. Its process was to select some candidate regions in the input image, then use feature descriptors to extract corresponding features, and finally send them into a classifier. However, the target size of the flying object image in the low-altitude airspace is small and the flight speed is fast, resulting in blurred edges. The artificially designed feature extractor has poor generalization ability, long running time, is easily affected by light, and there are also large errors in the positioning of the candidate boxes. The effect of traditional target detection methods on the detection of small targets is not ideal and it is difficult to apply to the current situation.
[0004] With the continuous development of the deep learning field, various deep models based on convolutional networks have been continuously developed. The features in the image dataset are no longer based on artificial design, but use convolutional layers and pooling layers to extract the semantic information of the target, and adopt upsampling and downsampling techniques to achieve the dimensional conversion of image features, so as to learn strong semantic information on the high-resolution feature map. Moreover, deep learning technology also has the advantages of high information reusability and shared weights. It not only improves the generalization ability of the model, but also can reduce the data dimension, reduce redundant information, thereby reducing the computational cost and greatly improving the detection speed and accuracy. Since 2012, the deep convolutional neural network was first proposed and applied to the field of object detection to complete intelligent detection tasks. Since then, object detection has made great breakthroughs, and convolutional neural networks have officially been used to complete tasks such as object recognition.
[0005] However, small and medium-sized unmanned aerial vehicles in the low-altitude airspace have characteristics such as small scale and large quantity, which makes it difficult to extract their texture features, and the detection accuracy of existing classical deep learning detection algorithms for such targets is relatively low. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a detection method for low-altitude small flying objects based on a convolutional neural network in view of the above-mentioned deficiencies of the prior art. Optimization and improvement are carried out from two aspects: improving the model's ability to obtain semantic information of small targets and the accurate positioning of the model for candidate boxes, which improves the sensitivity of the model to small targets, the accuracy of the model in the case of dense appearance of objects to be detected, and can improve the position accuracy of the model for detecting low-altitude small targets.
[0007] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0008] A detection method for low-altitude small flying objects based on a convolutional neural network, comprising the following steps:
[0009] Step 1: First, readjust the size of all data sample pictures, scale the images to a fixed size, that is, 800×600, and the model sends the picture dataset into the feature extraction network Resnet50. After a series of operations of convolutional layers and pooling layers, a feature map is output.
[0010] Step 2: Transfer the feature map into the region proposal network, and use the k-means clustering algorithm in the region proposal network to cluster the sizes of the ground truth boxes in the data sample set, replacing the sizes preset based on artificial experience in the original model.
[0011] Step 3: Perform linear scaling on the clustered anchor box sizes, then divide positive and negative samples and correct the positions of the anchor boxes in the Region Proposal Network (RPN). The RPN passes the positive and negative sample scores and the coordinate parameters of the anchor boxes through two fully connected layers respectively. The first fully connected layer is used to generate the target probabilities of these anchor boxes, producing the scores of positive samples and negative samples. The second fully connected layer is used to encode the four coordinate values of each anchor box, that is, calculate the offset of the anchor box to make the candidate box closer to the true box, and finally calculate the candidate regions.
[0012] Step 4: After calculating the candidate regions, optimize the number of anchor boxes again. Use the soft-nms algorithm to reduce the weights of overlapping anchor boxes, which can improve the accuracy of model prediction and reduce missed detection cases in the case of multiple small object overlaps, and then pass the optimized candidate regions into the ROI Align layer.
[0013] Step 5: The ROI Align layer fixes the candidate regions of different sizes, maps the candidate regions back to the original image, and finally predicts the results of low-altitude small flying objects through a fully connected layer.
[0014] The beneficial effects of adopting the above technical solutions are as follows: The low-altitude small flying object detection method based on convolutional neural network provided by the present invention is optimized and improved from two aspects: improving the model's ability to obtain semantic information of small objects and the accurate positioning of the model for candidate boxes. By improving the structure of the feature extraction network and adopting the residual structure with skip connections to improve the sensitivity of the model to small objects; in the post-processing part, using soft-nms to optimize redundant anchor boxes to improve the accuracy of the model in the case of dense appearance of objects to be detected; in the generation of anchor boxes, adopting ROI Align and replacing the original two rounding operations with linear interpolation algorithm to improve the position accuracy of the model for detecting low-altitude small objects, so that the convolutional neural network model can accurately detect low-altitude small flying objects. Brief Description of the Drawings
[0015] Figure 1 It is the algorithm structure diagram of the low-altitude small flying object detection method based on convolutional neural network provided by the embodiment of the present invention;
[0016] Figure 2 It is the structure diagram of the feature extraction network Resnet50 provided by the embodiment of the present invention;
[0017] Figure 3 It is the structure diagram of the ResNet residual network provided by the embodiment of the present invention;
[0018] Figure 4 It is the process of extracting feature maps of the ResNet residual network structure provided by the embodiment of the present invention;
[0019] Figure 5 Flow chart of generating anchor boxes for the regional candidate network provided by the embodiment of the present invention;
[0020] Figure 6 Flow chart of the regional candidate network provided by the embodiment of the present invention;
[0021] Figure 7 Schematic diagram of the IoU calculation process provided by the embodiment of the present invention;
[0022] Figure 8 Effect diagram of flexible non-maximum suppression provided by the embodiment of the present invention;
[0023] Figure 9 Schematic diagram of the structure of the ROI network Align layer provided by the embodiment of the present invention;
[0024] Figure 10 Schematic diagram of the mapping process of ROI Align candidate boxes provided by the embodiment of the present invention;
[0025] Figure 11 Schematic diagram of the ROI Align calculation provided by the embodiment of the present invention;
[0026] Figure 12 Schematic diagram of the prediction of the optimized model for a single small target provided by the embodiment of the present invention;
[0027] Figure 13 Schematic diagram of the prediction of the optimized model for dense small targets provided by the embodiment of the present invention. Detailed implementation manners
[0028] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0029] As Figure 1 shown, the method of this embodiment is described as follows.
[0030] Step 1: First, readjust the sizes of all data sample pictures, scale the images to a fixed size, that is, 800×600, and send the scaled images into the feature extraction network Resnet50. After a series of operations of convolutional layers and pooling layers, feature maps are output. The shallow network of the feature extraction network of the Faster R-CNN model is mainly used to extract geometric information such as the shape and size of the object to be detected. Deepening the number of network layers helps to better extract the semantic information of the object to be detected, but problems such as "network degradation" and "gradient explosion" may occur. Therefore, the ResNet50 network is selected as the feature extraction network of Faster R-CNN, and its overall structure is as Figure 2As shown, in addition to the forward propagation of the traditional convolutional network, it adds a unique residual structure to increase the information transmission between adjacent network layers. Specifically, as Figure 3 shown, this structure of skip links can improve the configurable depth of the network, enabling it to learn more complex and accurate semantic information. Figure 4 The ResNet50 residual network structure uses convolutional layers to perform forward convolution to obtain a highest-layer feature map.
[0031] Step 2: Use the K-Means clustering algorithm in the candidate region network to generate anchor boxes with appropriate sizes. The K-means algorithm can replace the artificial setting of anchor box sizes in traditional object detection models, helping the deep learning detection model to better design the sizes of anchor boxes, reducing redundant anchor boxes, saving computing power, and improving the operation speed of the model while increasing the accuracy of the model. Moreover, the detection task in this embodiment mainly targets small objects, and there are relatively more small-sized drones and birds in the dataset. In the traditional deep learning detection model, the number of generated anchor boxes of large, medium, and small sizes is the same, which is obviously unreasonable in the detection requirements of this embodiment. The K-means algorithm can solve this problem. The K-Means algorithm is relatively simple to implement. For a given dataset of data samples, it is divided into K clusters. Let the divided clusters be: (C1, C2,... C k ), then the minimum squared error E is shown in formula (1):
[0032]
[0033] In the formula: K represents the number of clusters; C i represents the divided cluster; μ i represents the mean vector of cluster C i . Among them, μ i is also called the centroid of the cluster, and the expression is shown in formula (2):
[0034]
[0035] The specific solution process of the K-means algorithm is as follows: (1) For an initial dataset of data samples, randomly select two points as the initial cluster centers; (2) Calculate the distances from all points in the data samples to the initial cluster centers, and mark the category of each point, that is, the category of the initial cluster center closest to this point. After one round of calculation, the categories of the first iteration of all sample data can be obtained; (3) Repeat the above process, and calculate the new centroids for the sample dataset roughly divided for the first time; the new centroids are the final clustering results of the algorithm. In actual applications, the above process is often repeated multiple times to ensure the accuracy of the results.
[0036] In the above traditional clustering algorithm, the Euclidean distance is used to calculate the difference. However, this calculation method is not applicable to all cases. Through a large number of experimental results, it is found that when the sizes of the model anchor boxes are relatively concentrated, the results of the K-means clustering algorithm are often inaccurate. To address such problems, the IoU value is introduced in the algorithm calculation, and a distance d is defined to represent the error. The calculation formula of d is shown in (3):
[0037] d = 1 - IoU (3)
[0038] Among them, the collective calculation process of IoU (Intersection over Union) is as Figure 7 shown. It measures the overlap degree between two anchor boxes and is obtained by calculating the intersection-over-union ratio of the anchor boxes. In the clustering algorithm, IoU refers to the overlap degree between each clustering center and other boxes. The larger the IoU, the higher the overlap degree. In the clustering algorithm, it is expected that the higher the overlap degree, the shorter the distance, and the better the clustering effect. Therefore, it is represented by formula (3). After running the K-means algorithm in the cases of N = 5 and N = 9, the correct rates are 69.12% and 75.27% respectively. So, 9 kinds of anchor boxes are selected. After clustering operation, the sizes [width, length] of these 9 kinds of anchor boxes are: {[9.36, 9.24], [14, 17.81], [16.9, 27.82], [21.97, 38.39], [26.61, 22.68], [30.16, 54.87], [34.53, 33.73], [46.17, 48.92], [63.41, 83.76]}.
[0039] Step 3: Since the targets detected in this embodiment mainly focus on low-altitude small flying objects, the self-collected data set is clustered. The sizes of the anchor boxes are relatively concentrated and small. Considering that there are still large-scale targets in the actual detected data set, in this embodiment, the sizes of the clustered anchor boxes are linearly scaled and stretched to avoid affecting the correct rate of other scale targets.
[0040] The specific calculation process is shown in formulas (4) to (7):
[0041] x′1 = αx1 (4)
[0042] x′9 = βx9 (5)
[0043]
[0044]
[0045] Among them, x a and y a are the width and length of the a-th anchor box, and x′ a and y′a are the width and length after scaling for the a-th anchor box, where a = 1, 2, …, 9; α and β are constants, taking 0.5 and 1.5 respectively.
[0046] In the region proposal network, positive and negative samples are divided and the positions of the anchor boxes are corrected. As Figure 5 shown, the region proposal network will pass the 2k positive and negative sample scores and 4k coordinate parameters (k is the 9 types of anchor box sizes in step 2, k = 9) through two fully connected layers respectively. The coordinate parameters include the length, width and center coordinates of the anchor box. As Figure 6 shown, the first fully connected layer is used to generate the target probabilities of these anchor boxes, and will generate 2k scores, that is, the scores of positive samples and negative samples; the second fully connected layer is used to encode the four coordinate values of each anchor box, that is, to calculate the offset of the anchor box, so that the candidate box is closer to the true box, and finally the candidate region is calculated.
[0047] Step 4: Optimize the confidence of the anchor box using the soft-nms algorithm, which can improve the accuracy of model prediction and reduce the occurrence of missed detections in the case of overlapping of multiple small targets.
[0048] Candidate regions are obtained through the region proposal network, and dense prediction is performed on this region. In two-stage detection algorithms, the bounding box regression operation regards the targets in the generated adjacent candidate boxes as the same object. Small targets in the low-altitude airspace, such as drones and birds, always appear densely. That is to say, there will be a large number of redundant candidate rectangular boxes in the result map pointing to the same result when the basic model detects these small targets, resulting in a decrease in the accuracy of the model. Therefore, the soft non-maximum suppression (soft-nms) algorithm is adopted to filter unnecessary candidate boxes.
[0049] Considering the classification confidence and localization confidence of the anchor box, a Gaussian weight function is used to screen redundant anchor boxes, as shown in formula (8) specifically:
[0050]
[0051] In the formula: M is the box with the highest current score, b i is the box to be processed, s i is the score of b i and σ is a hyperparameter, generally taking 0.5;
[0052] The non - maximum suppression algorithm is used twice in the Faster RCNN deep detection model. The first time is in the training stage, when a large number of anchor boxes are generated by the region proposal network, the non - maximum suppression algorithm is used to roughly filter them. The second time is in the post - processing stage, where the non - maximum suppression algorithm is used to accurately screen the anchor boxes. Experimental verification is carried out for the above - mentioned situations, and the results are shown in Table 1.
[0053] Table 1 Experimental verification results
[0054]
[0055] From the actual detection situation and combined with the data in the above table, modifying the non - maximum suppression algorithm in the training stage not only increases the running time of the model and wastes computing power, but also a large number of redundant anchor boxes appear during the prediction process. When using the non - maximum suppression algorithm to process the anchor boxes for the first time, a large number of useless anchor boxes should be deleted, rather than increasing the weight function to measure the importance of the anchor boxes and reducing the anchor box weights. When modifying the non - maximum suppression algorithm in both places, it not only causes the model to run for too long. During this process, due to the lack of an initial process of massive screening and deletion, more anchor boxes also appear during the final prediction; when modifying the non - maximum suppression algorithm in the post - processing stage, the average accuracy rate in the initial prediction results of the model is 1% higher than the accuracy rate when modifying the non - maximum suppression algorithm for the second time. The significance of modifying the non - maximum suppression algorithm lies in that the predicted anchor boxes are more accurate when modified for the second time. Compared with the screenshot of the initial prediction of the model, 1 / 3 of the volume of the drone on the right is outside the anchor box. When modifying the non - maximum suppression algorithm for the second time, although some accuracy is lost, the positioning of the anchor box is more accurate, and the confidence level of the lost anchor box is relatively high at 93%, which is within the acceptable range. Therefore, only the non - maximum suppression algorithm for the second time needs to be optimized. The effect of flexible non - maximum suppression is as Figure 8 shown. There are redundant anchor boxes on the object to be detected in the image dataset. By comparing the IOU values of the objects to be detected, the anchor box with the maximum confidence level is retained.
[0056] Step 5: Input the feature map calculated in Step 1 and the optimized candidate regions obtained in Step 4 into the ROI Align layer. The specific process is as Figure 9 shown. This layer can fix candidate regions of different sizes, map the candidate regions back to the original image, and perform the final result prediction on low - altitude small flying objects through the fully - connected layer.
[0057] When the target detection model first inputs data, it fixes the size of the original image once. Since the number of network layers in the feature extraction network is fixed, the size of the generated feature map is also fixed. Ignoring the dimensions, the ratio between the two is also fixed, that is, M = m / 16 and N = n / 16. In most cases, evenly dividing the size of the feature map into x parts will result in floating-point numbers. At this time, the model will default to rounding down, which will finally cause the bias of the candidate box. For some targets with a large enough size, these effects are negligible and can be ignored. However, for small targets, the effects can even lead to incorrect model classification. Therefore, in this embodiment, the data in the mapping process is not operated on, but new values are regenerated, and new values are calculated through an interpolation algorithm commonly used in image affine transformation to avoid the precision loss caused by quantization and not perform rounding. This algorithm has excellent performance both in terms of the accuracy of the results and the running speed. The specific calculation process of this algorithm is as Figure 10 shown. Figure 10 There is a candidate region of 124×124, corresponding to a scale of 16 times. After retaining the floating-point number, its size is 7.75×7.75. Similarly, it is evenly divided into four small regions, as Figure 11 shown. The size of each small region is 3.875×3.875. Then, pixel values are extracted using the interpolation algorithm in each small region. This linear interpolation algorithm uses four points near the target pixel coordinates for calculation, so the small region is divided into four parts again. The ROI Align layer has improved both in terms of the selection of sampling points and the accuracy of the results. The backpropagation formula of ROI Align is shown in formula (9):
[0058]
[0059] In the formula, x i represents the i-th pixel point on the feature map before pooling; y rj represents the j-th point of the r-th candidate region after pooling, d(.) is the pixel point on the feature map before pooling; Δh is the difference in the abscissa between x i and x i*(r,j) ; Δw is the difference in the ordinate between x i and x i*(r,j) , and x i*(r,j) is the coordinate position of a floating-point number (the sampling point calculated during forward propagation). During the backpropagation process, ROI Align no longer only targets the pixel value of a certain point. It targets the pixel values within a certain range. Because the quantization operation is cancelled during the calculation process and the corresponding point cannot be found, the floating-point coordinates mapped to the feature map are used as the center of the circle, and all feature points within a circle with a radius of 1 are backpropagated, that is, each feature point related to x i*(r,j)Points where both the horizontal and vertical coordinates are less than 1. Finally, a fully connected layer is used to classify and predict low-altitude small flying objects and output the results.
[0060] Figure 12 is a single small target data sample, Figure 13 is a multiple small target data sample. The model Figure 12 for prediction, detecting the UAV (drone) of the sample type, with a confidence level of 93%, and the prediction effect is good; for Figure 13 for prediction, detecting that the sample type is bird (bird), and for the case where multiple samples overlap, the model has a good prediction effect and the confidence level is generally greater than 90%.
[0061] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.
Claims
1. A method for detecting low-altitude small flying objects based on a convolutional neural network, characterized in that: The method includes the following steps: Step 1: First, readjust the sizes of all data sample pictures, scale the images to a fixed size, i.e., 800×600. The model feeds the picture dataset into the feature extraction network Resnet50. After a series of operations in the convolutional layer and pooling layer, a feature map is output. Step 2: Feed the feature map into the region proposal network. In the region proposal network, use the k-means clustering algorithm to cluster the sizes of the ground truth boxes in the data sample set. The IoU value is introduced in the calculation of the k-means clustering algorithm, and the distance d is defined to represent the error. The calculation formula of d is shown in (3): d = 1 - IoU (3) where IoU (Intersection over Union) is a measure of the overlap between two anchor boxes, obtained by calculating the intersection over union of the anchor boxes. In the clustering algorithm, IoU refers to the degree of overlap between each clustering center and other boxes. After the clustering operation, the [width, length] sizes of 9 anchor boxes are obtained as: {[9.36, 9.24], [14, 17.81], [16.9, 27.82], [21.97, 38.39], [26.61, 22.68], [30.16, 54.87], [34.53, 33.73], [46.17, 48.92], [63.41, 83.76]}. Step 3: Perform linear scaling processing on the sizes of the clustered anchor boxes, and then perform positive and negative sample division and correction of the anchor box positions in the region proposal network. The region proposal network passes the positive and negative sample scores and the coordinate parameters of the anchor boxes through two fully connected layers respectively. The first fully connected layer is used to generate the target probabilities of these anchor boxes, generating positive sample scores and negative sample scores. The second fully connected layer is used to encode the four coordinate values of each anchor box, that is, calculate the offset of the anchor box, so that the candidate box is closer to the ground truth box, and finally calculate the candidate region. The calculation process of performing linear scaling processing on the sizes of the clustered anchor boxes is shown in formulas (4) to (7): x'1 = αx1 (4) x'9 = βx9 (5) Among them, x a , y a are the width and length of the a-th anchor box, x' a , y' a are the width and length of the a-th anchor box after scaling, where a = 1, 2, …, 9; α and β are constants, taking 0.5 and 1.5 respectively; Step 4: After calculating the candidate regions, use the soft-nms algorithm to optimize the anchor box confidence, optimize the number of anchor boxes, reduce the weights of overlapping anchor boxes, and feed the optimized candidate regions into the ROI Align layer. Step 5: The ROI Align layer fixes the candidate regions of different sizes, maps the candidate regions back to the original image, and finally predicts the results of low-altitude small flying objects through a fully connected layer. When the ROI Align layer fixes the candidate regions of different sizes, for each candidate region, after corresponding scaling, keep the floating point number, and then evenly divide it into four small regions. Pixel values are extracted using linear interpolation algorithm in each small region. The backpropagation formula of ROI Align is shown in formula (9): Where x i represents the i-th pixel point on the feature map before pooling; y rj represents the j-th point of the r-th candidate region after pooling, d(.) is the pixel point on the feature map before pooling; Δh is the difference between the abscissa of x i and x i*(r,j) ; Δw is the difference between the ordinate of x i and x i*(r,j) ; x i*(r,j) is the coordinate position of a floating point number, that is, the sampling point calculated during forward propagation. During the backpropagation process, ROI Align no longer targets the pixel value of a single point. Instead, it targets the pixel values within a certain range. Using the floating-point coordinates mapped to the feature map as the center, all the feature points within a circle with a radius of 1 are backpropagated, that is, every point where both the i*(r,j) horizontal and vertical coordinates are less than 1.
2. The method for detecting low-altitude small flying objects based on a convolutional neural network according to claim 1, characterized in that: The soft-nms algorithm uses a Gaussian weight function to filter redundant anchor boxes, specifically shown in formula (8): Among them, M is the currently highest-scoring box, b i is the box to be processed, s i is the score of b i and σ is a hyperparameter; By comparing the IOU values of the objects to be detected, retain the anchor box with the highest confidence.
Citation Information
Patent Citations
Image semantic feature constrained remote sensing target detection method
CN112101277A
Unmanned aerial vehicle platform multi-scale target detection method and device
CN112819100A