An Unmanned Aerial Vehicle (UAV) Attitude Detection Method and System with Confidence Estimation
Skyline detection is carried out through a fully convolutional neural network and confidence estimation is carried out in combination with Gaussian discriminant analysis model. The problems of single scenes, poor anti-interference ability and low generalization ability in the existing technology are solved, and efficient and accurate drone attitude angle estimation under complex conditions are achieved, which enhances the autonomy and concealment of the drone.
Patent Information
- Application Number
- CN202111277574.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-10-29
AI Technical Summary
The existing drone attitude angle estimation technology has the problems of single scenarios, poor anti-interference ability and low generalization ability. Especially in complex backgrounds, there are large errors in the detection results, which cannot effectively avoid risks.
The skyline detection is performed by using a full convolutional neural network, and confidence estimation is performed in combination with the Gaussian discriminant analysis model. The detection results are filtered through confidence, and error results are effectively filtered out and risks are avoided.
It realizes high adaptability and high precision skyline detection under different environments and complex terrain conditions, provides reliability assessment of the drone's attitude angle estimation results, and enhances the autonomy and concealment of the drone.
Smart Images

Figure CN113888630B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image information processing, and particularly relates to a method and system for detecting the attitude of an unmanned aerial vehicle with confidence estimation. Background Art
[0002] Navigation is an important research field for the autonomous flight of unmanned aerial vehicles. Among them, the attitude angle is essential navigation information required for autonomous flight. Due to the limitations of many factors such as weight, volume, and power consumption, computer vision technology with a camera as the main sensor has become the main development trend. The present invention will realize the real-time estimation of the attitude angle of the unmanned aerial vehicle by detecting the position of the skyline.
[0003] In recent years, researchers at home and abroad have achieved certain results in the field of skyline detection and application research. The existing skyline detection algorithms can be divided into four categories: 1) Model methods based on linear boundaries. This method aims to perform Gaussian distribution modeling or Hough transformation on image information based on the assumption that the skyline is a straight line. However, the assumption of a straight horizon is only valid in specific scenarios. When the altitude is too low, obstacles and hills will produce a horizon that is not a straight line. Therefore, due to the limitations of this assumption, this model method cannot meet the needs of more actual scenarios. 2) Edge detection-based methods. This method realizes the contour extraction of the skyline by identifying the edge information of the boundary between the sky and the ground. However, edge detection highly depends on the setting of parameters, and the algorithm has poor generalization ability. Secondly, the contours of clouds and mountains will also interfere with the skyline edge detection, reducing the detection accuracy. 3) Classifier-based methods through machine learning. This method uses image color and texture features, such as average intensity, entropy, smoothness, uniformity, etc., to train a classifier, and then applies the classifier to the sky and non-sky regions for skyline extraction. Commonly used classifiers include: SVM, J48, and Naive Bayes classifiers. However, this method has an unsatisfactory detection effect for skylines with low color differentiation in the skyline area. 4) Deep learning-based methods. This method applies a convolutional neural network to skyline detection, which is a faster and more robust skyline detection method. Existing methods mainly use CNN to train the sky and non-sky regions and the skyline in the flight video, and then verify the proposed method using a large dataset, and the detection accuracy is better than SVM and random forests. However, the research in this field is not yet fully mature, and the applied convolutional neural network frameworks are relatively simple, with room for subsequent improvement.
[0004] More importantly, when there are complex backgrounds such as clouds, rain, fog, mountains, etc., which lead to large errors in the skyline detection results, the estimated UAV attitude angle information cannot be used. Therefore, in response to this situation, it is necessary to estimate the confidence of the detection results as a corresponding reliability reference value to avoid risks caused by incorrect results. However, there is little research on the confidence estimation of UAV attitude angle detection results at home and abroad. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to overcome the defects of the existing UAV attitude angle estimation technology, and provide a UAV attitude detection method and system with confidence estimation, and a skyline detection method based on a fully convolutional neural network, which solves the problems of single use scenario, poor anti-interference ability, low generalization ability, etc. of the original method. The confidence estimation algorithm based on Gaussian discriminant analysis effectively filters out incorrect results when there are large or serious errors in the detection results, helping the UAV avoid risks.
[0006] The present invention is realized through the following technical solutions:
[0007] A UAV attitude detection method with confidence estimation includes the following steps:
[0008] Step A: Read the input image of the current frame, perform pixel-level segmentation of the sky area and non-sky area on the input image through a fully convolutional neural network, extract the skyline coordinates from the image of the sky area, and fit the optimal straight line equation according to the skyline coordinates to obtain the skyline fitting straight line.
[0009] Step B: Estimate the confidence of the skyline fitting straight line through a trained Gaussian discriminant analysis model. If the confidence of the skyline fitting straight line is higher than the preset optimal classification threshold, then go to Step C; otherwise, return to Step A to read the input image of the next frame.
[0010] Step C: Real-time estimate the UAV attitude angle information based on the skyline fitting straight line.
[0011] Preferably, in Step A, the fully convolutional neural network includes an encoding network, a decoding network, a class calibration module, and an optimal straight line extraction module;
[0012] Step A is specifically: Use the encoding network to extract image features and encode them into corresponding heatmaps; the decoding network uses upsampling to enlarge the heatmap to the size of the input image and decodes it into the classification probability of each pixel, and outputs a probability map; the class calibration module performs class calibration on the probability map pixel by pixel to generate a segmentation binary map to obtain the image of the sky area; the optimal straight line extraction module extracts the skyline coordinates from the image of the sky area and fits the optimal straight line equation according to the skyline coordinates to obtain the skyline fitting straight line.
[0013] Further, the decoding network uses upsampling to enlarge the heatmap to the size of the input image, decodes it into the classification probabilities of each pixel, and outputs a probability map, expressed as:
[0014] M = F de (H)
[0015] M ij0 = P(p ij = sky)
[0016] M ij1 = P(p ij = nonsky)
[0017] where, F de represents the decoding network, represents upsampling; H represents the heatmap, which is the input of the decoding network; M represents the probability map, which is the output of the decoding network; M ijk represents the value of the probability map M at the coordinate (i, j) in channel k, and k takes the value of 0 or 1; p ij represents the pixel at the coordinate (i, j) in the input image I.
[0018] Further, step B specifically includes the following steps:
[0019] 1) According to the probability map and the segmentation binary map output by the fully convolutional neural network, quantify to obtain the segmentation quality Q and the curvature T of the skyline;
[0020] 2) Use the trained Gaussian discriminant analysis model to perform multivariate Gaussian modeling on the segmentation quality Q and the curvature T; according to the learned sample distribution, the Gaussian discriminant analysis model calculates the confidence P of the skyline fitting straight line.
[0021] Further, in step B, the trained Gaussian discriminant analysis model in step 2) is obtained through the following training method:
[0022] Use m training samples (x (1) , y (1) ), (x (2) , y (2) ), (x (3) , y (3) ), …, (x (m) , y (m) ) to perform offline training on the Gaussian discriminant analysis model, where, y (i) ∈0, 1; x represents multivariate sample data, which is the quantization value of the segmentation quality Q and the curvature T; y represents the category of the sample data, y (i) = 1 represents that the skyline fitting straight line is reliable; y (i) = 0 represents that the skyline fitting straight line is unreliable;
[0023] Assume that the class y of the sample data follows a Bernoulli distribution under given conditions, and the sample data x in different classes y follows a multivariate Gaussian distribution respectively:
[0024] y ∼ Bernoulli(φ)
[0025] x|y = 0 ∼ N(μ0, Σ)
[0026] x|y = 1 ∼ N(μ1, Σ)
[0027] Where Bernoulli(φ) represents the Bernoulli distribution, and μ and Σ represent the mean and covariance of the multivariate Gaussian distribution respectively. Then we have:
[0028]
[0029]
[0030] The values of the three parameters μ0, μ1, and Σ are obtained through the maximum likelihood estimation function:
[0031]
[0032]
[0033]
[0034] According to Bayes' formula, the probability values of the class y of the sample data being positive and negative samples are obtained under the condition of known sample data x:
[0035]
[0036]
[0037] Among them, p(y = 0|x) is considered the confidence of the skyline fitting line, and its value range is [0, 1].
[0038] Preferably, in step A, the skyline coordinates are extracted from the image of the sky region, and the optimal straight-line equation is fitted according to the skyline coordinates to obtain the skyline fitting line. Specifically: the lower boundary coordinates of the largest contour in the sky region are extracted as the skyline coordinates, and the filtering point algorithm is used to fit a straight line from the skyline coordinates to obtain the skyline fitting line.
[0039] Preferably, in step B, the method for setting the optimal classification threshold is: a large number of samples are used to offline train the Gaussian discriminant analysis model, and the optimal classification threshold of the confidence of the skyline fitting line is obtained through the training results.
[0040] Preferably, step C is specifically:
[0041] By obtaining the skyline to fit the straight line equation y = kx + b, through geometric calculation, the calculation formulas for the roll angle φ and the pitch angle θ are respectively:
[0042]
[0043]
[0044] Among them, f x and f y are the camera internal parameters, and (u0, v o ) are the coordinates of the principal point of the image.
[0045] An unmanned aerial vehicle (UAV) attitude detection system with confidence estimation includes:
[0046] A fully convolutional neural network, which is used to read the input image of the current frame, perform pixel-level segmentation of the sky and non-sky regions on the input image, extract the skyline coordinates from the image of the sky region, fit the optimal straight line equation according to the skyline coordinates, and obtain the skyline fitting straight line;
[0047] A confidence estimation module, which is used to estimate the confidence of the skyline fitting straight line through a Gaussian discriminant analysis model. If the confidence of the skyline fitting straight line is higher than the preset optimal classification threshold, the UAV attitude angle estimation module works; otherwise, the fully convolutional neural network reads the input image of the next frame;
[0048] A UAV attitude angle estimation module, which is used to estimate the UAV attitude angle information in real time through geometric calculation and the equation of the skyline fitting straight line.
[0049] Preferably, the fully convolutional neural network includes an encoding network, a decoding network, and a class calibration and optimal straight line extraction module;
[0050] The encoding network is used to extract image features and encode them into corresponding heatmaps;
[0051] The decoding network is used to enlarge the heatmap to the size of the input image by using upsampling and decode it into the classification probabilities of each pixel, and output a probability map;
[0052] The class calibration module is used to perform class calibration on each pixel of the probability map, generate a segmentation binary map, and obtain the image of the sky region;
[0053] The optimal straight line extraction module is used to extract the skyline coordinates from the image of the sky region and fit the optimal straight line equation according to the skyline coordinates to obtain the skyline fitting straight line.
[0054] Compared with the prior art, the present invention has the following beneficial technical effects:
[0055] The detection method of the present invention estimates the attitude angle by detecting the skyline, which has the advantage of strong autonomy and can overcome the dependence on external navigation methods. The confidence estimation function designed by the method of the present invention can provide corresponding reliability reference values for the detection results in real time, especially when there are large or serious errors in the detection results, effectively avoiding risks. At the same time, the method of visual navigation does not need to emit signals outward, providing strong concealment for the unmanned aerial vehicle. In addition, the present invention uses a fully convolutional neural network to perform pixel-level classification on the image, retaining the spatial information in the original input image, and finally performing per-pixel classification on the upsampled feature map. Therefore, the present invention can realize the prediction and classification of sky and non-sky pixels by pixel, enabling it to have high adaptability and high precision detection capabilities in different environments, different terrains, and complex meteorological conditions, and can be better applied to the attitude angle estimation of unmanned aerial vehicles. The present invention has the advantages of good autonomy, strong concealment, small size, and light weight. It solves the problems of single use scenario, poor anti-interference ability, and low generalization ability of the original method. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a schematic diagram of the skyline detection method based on a fully convolutional neural network of the present invention;
[0057] Figure 2 It is a schematic diagram of the image segmentation process of the embodiment of the present invention.
[0058] Figure 3 (a) is the ROC curve graph of the confidence estimation module in the embodiment of the present invention; Figure 3 (b) is the distribution graph of the confidence values of the actually negative samples; Figure 3 (c) is the distribution graph of the confidence values of the actually positive samples. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The following further elaborates on the present invention in detail with specific embodiments, which is an explanation rather than a limitation of the present invention.
[0060] An unmanned aerial vehicle attitude detection method with confidence estimation according to the present invention, as Figure 1 shown, includes the following steps:
[0061] Step A: Read the input image of the current frame, and through a fully convolutional neural network suitable for skyline detection, perform pixel-level segmentation of the sky area and non-sky area of the input image, extract the corresponding skyline position coordinates from the sky area, and extract the optimal straight line equation through the RANSAC filtering point algorithm, that is, the skyline fitting straight line, which is used for subsequent unmanned aerial vehicle pose calculation;
[0062] Step B: During the actual flight of the drone, the effectiveness of the skyline fitting line is tested online through the trained Gaussian discriminant analysis model, and the output result is the confidence level of the skyline fitting line. The effectiveness of the skyline fitting line is judged according to the confidence level. If the confidence level is higher than the threshold, it is considered that the skyline fitting line is effective, and thus Step C is carried out, and the attitude angles calculated in Step C are adopted by the navigation device. On the contrary, if the confidence level is lower than the threshold, it is considered that the skyline fitting line is ineffective, and Step C is directly skipped, and the system returns to Step A to read the next frame of input image.
[0063] Step C: Through geometric calculation, the roll angle and pitch angle of the drone are estimated in real time according to the skyline fitting line.
[0064] The specific design of the fully convolutional neural network structure in the above Step A is as follows:
[0065] This fully convolutional neural network N full is mainly composed of an encoding network F en and a decoding network F de and class calibration. The size of the input image I is not limited by its network, and the output segmentation binary map O is always the same size as the input image I, realizing end-to-end pixel-level segmentation. Among them,
[0066] O = N full (I) = argmax(F de (F en (I)))
[0067] 1) Encoding network design
[0068] The encoding network is mainly used to extract image features and encode them into corresponding heatmaps. Each point of the heatmap represents the target detection result of the receptive field area. The specific function expression is:
[0069] H = F en (I)
[0070] Among them, I is the input image, with a size of h×w×c; F en represents the encoding network; H is the output heatmap, with a size of h H ×w H ×c. Due to the processing of the convolutional layer and the pooling layer, the size of the heatmap is smaller than that of the input image, but the number of channels of the two remains the same.
[0071] The front-end structure of the encoding network is basically the same as that of the classification network, and feature extraction is achieved through consecutive convolutional layers and pooling layers. Different from the classification network, this encoding network replaces the fully connected layer with a convolutional layer, thus removing the constraint on the size of the input image. Among them, the convolutional layer performs a convolution operation on the input layer or the output of the previous layer through a number of convolutional kernels, and combines the convolution results into a feature image through an activation function. The convolutional layer function is expressed as:
[0072] s = f(x × w + b)
[0073] Among them, s represents the output data of the convolutional layer, x represents the input data of the convolutional layer, w represents the weight of the convolutional kernel, b represents the bias, and f represents the activation function.
[0074] 2) Decoding network design
[0075] The role of the decoding network is to use the upsampling method to enlarge the heat map to the size of the original image, so as to decode the image feature information into the classification probability of pixels. After decoding, the decoding network will output a probability map with a size of h × w × 2. The size of the probability map is exactly the same as that of the original input image, achieving pixel-level correspondence; the number of channels in the 2 layers represents two types of targets: sky and non-sky, representing the probability that the corresponding pixel becomes sky or non-sky. This decoding network can be expressed as:
[0076] M = F de (H)
[0077] M ij0 = P(p ij = sky)
[0078] M ij1 = P(p ij = nonsky)
[0079] Among them, F de represents the decoding network, represents upsampling; the heat map H is used as the input of the decoding network; the probability map M is the output of the decoding network, and M ijk represents the value of the probability map M at the coordinate (i, j) in the channel k, and p ij represents the pixel at the coordinate (i, j) in the input image I.
[0080] 3) Class calibration
[0081] The probability map generated by the decoding network needs to generate the final segmentation binary map O through pixel-by-pixel class calibration. Class calibration obtains the channel sequence value where the maximum probability value of a certain pixel is located by comparing different channels of the probability map, that is, obtains the classification result corresponding to the pixel. This method can generate a sky or non-sky prediction for each pixel, retain the spatial information in the original input image, and achieve pixel-level coherent segmentation between sky regions. The specific expression is as follows:
[0082] O = argmax(M ij0 , M ij1 ), O ij ∈ {0, 1}
[0083] Where O represents the segmented binary map generated after class calibration, with a size of h×w×1. Since the probability map has only two channels, the values of the segmented binary map after class calibration are 0 or 1.
[0084] 4) Loss function
[0085] In this method, each pixel in the fully convolutional neural network is a classification task, and each image has the same number of samples as the corresponding number of pixels. When calculating the loss function, the softmax loss function is calculated for each pixel in the segmented binary map O, and after all are accumulated, a gradient update is performed:
[0086] m = max(O ij ), O ij ∈ {0, 1}
[0087]
[0088] Where O ij is the predicted label (sky, non-sky) corresponding to the pixel with coordinates (x, y) in the segmented binary map O, is the actual classification label of this pixel, and f is the softmax function.
[0089] 5) Optimal straight line equation of the skyline
[0090] (1) Obtain the maximum contour coordinates of the sky area from the segmented binary map O output by the fully convolutional neural network;
[0091] (2) Remove the upper, left, and right boundary coordinates of the maximum contour of the sky area, and extract the lower boundary coordinates as the detected skyline coordinates;
[0092] (3) Extract the optimal straight line equation through the RANSAC filtering algorithm, and this straight line is used for subsequent UAV pose calculation.
[0093] The basic principle in step B is as follows:
[0094] Without the correct reference conditions, the accuracy of the skyline detection results is not available. However, two relevant factors that have a decisive impact on the accuracy of skyline detection: the segmentation quality Q of the fully convolutional neural network and the curvature T of the predicted skyline, can be quantified and data can be obtained. In the present invention, the Gaussian discriminant analysis algorithm indirectly measures the "reliability" degree of the skyline detection results by performing a multivariate Gaussian modeling on the segmentation quality Q and the curvature T. The skyline detection results refer to the skyline fitting line.
[0095] For m samples (x (1) , y (1) ), (x (2) , y (2) ), (x (3) , y (3) ), …, (x (m) , y (m) ), y (i) ∈ 0, 1. x represents the multivariate sample data, which is the quantified value of the segmentation quality Q and the curvature T in the present invention; y represents the category of the sample data, and y (i) = 1 represents that the skyline fitting line is reliable and has a higher accuracy; y (i) = 0 represents that the skyline fitting line is unreliable and has a lower accuracy. In this confidence estimation algorithm, there are two prior assumptions: one is that the category y of the sample data follows a Bernoulli distribution given, and the other is that the sample data x in different categories respectively follows a multivariate Gaussian distribution:
[0096] y ~ Bernoulli(φ)
[0097] x|y = 0 ~ N(μ0, Σ)
[0098] x|y = 1 ~ N(μ1, Σ)
[0099] Among them, Bernoulli(φ) represents the Bernoulli distribution, that is, the 0-1 distribution or the binomial distribution. μ and Σ respectively represent the expectation and covariance of the multivariate Gaussian distribution. Then there are:
[0100]
[0101]
[0102] The values of the three parameters μ0, μ1, and Σ can be obtained through the maximum likelihood estimation function:
[0103]
[0104]
[0105]
[0106] According to Bayes' formula, the probability values of the class y of the sample data being positive and negative samples can be obtained given the known sample x:
[0107]
[0108]
[0109] Among them, p(y = 0|x) is regarded as the confidence level P of the skyline detection result. The higher this confidence level P, the higher the probability that the detection result becomes a reliable result; the lower this confidence level P, the higher the probability that the detection result is incorrect.
[0110] Step B specifically includes the following steps:
[0111] (1) Use a large number of samples to offline train a Gaussian discriminant analysis model, and obtain the optimal classification threshold of the confidence level through the training result.
[0112] (1.1) Prepare training samples, and quantify to obtain the segmentation quality Q and the curvature T of the skyline fitting line according to the probability map and the segmented binary map output by the fully convolutional neural network;
[0113] (1.2) Use the segmentation quality Q and the curvature T to perform online training on the Gaussian discriminant analysis model; the training result can obtain the confidence level P of the skyline fitting line according to the learned sample distribution.
[0114] (1.3) The confidence level obtained by Gaussian determination analysis has a value range of [0, 1], and the ROC curve is used to further determine the optimal classification threshold.
[0115] (2) When the drone is actually flying, online test the effectiveness of the skyline fitting line through the trained Gaussian discriminant analysis model, and the output result is the confidence level.
[0116] (2.1) According to the probability map M and the segmented binary map O output by the fully convolutional neural network, quantify the segmentation quality Q and the curvature T of the skyline fitting line;
[0117] (2.2) Use the Gaussian discriminant analysis model to perform multivariate Gaussian modeling on the segmentation quality Q and the curvature of the skyline fitting line T;
[0118] (2.3) According to the learned sample distribution, the Gaussian discriminant analysis model obtains the confidence level P of a newly detected skyline fitting line.
[0119] (3) If the confidence level is higher than the above-mentioned optimal classification threshold, the skyline fitting line is considered valid, and thus step C is performed, and the attitude angles calculated in step C are adopted by the navigation device. Conversely, if the confidence level is lower than the above-mentioned optimal classification threshold, the skyline fitting line is considered invalid, and step C is directly skipped, and the system reads the next frame.
[0120] Step C specifically includes the following steps:
[0121] From the obtained skyline fitting line equation y = kx + b, through geometric calculation, the calculation formulas for the roll angle φ and the pitch angle θ are respectively:
[0122]
[0123]
[0124] where f x and f y are the camera internal parameters, and there are (u0, v o ) are the coordinates of the principal point of the image.
[0125] The present invention proposes an unmanned aerial vehicle attitude detection system with confidence estimation, including:
[0126] A skyline extraction module, which is used to perform pixel-level segmentation of the sky area and the non-sky area on the input image through a fully convolutional neural network, extract the skyline coordinates from the image of the sky area, and fit the optimal straight line equation according to the skyline coordinates to obtain the skyline fitting line.
[0127] Among them, the fully convolutional neural network includes an encoding network, a decoding network and a class calibration module;
[0128] The encoding network is used to extract image features and encode them into corresponding heat maps;
[0129] The decoding network is used to enlarge the heat map to the size of the input image by means of upsampling and decode it into the classification probability of each pixel, and output a probability map;
[0130] The class calibration module is used to perform class calibration on each pixel of the probability map to generate a segmentation binary map and obtain the image of the sky area.
[0131] A confidence estimation module, which is used to estimate the confidence of the skyline fitting line through a Gaussian discriminant analysis model.
[0132] An unmanned aerial vehicle attitude angle estimation module, which solves the roll angle and pitch angle of the unmanned aerial vehicle through geometric calculation and the skyline fitting line equation.
[0133] Embodiment
[0134] The application scenario of this embodiment is as follows: By detecting the skyline in the images captured by the forward-looking camera of the drone, the roll and pitch angles of the drone can be calculated in real time. The present invention realizes the ability of the drone to detect the skyline with high adaptability and high precision under different environments, terrains, and complex meteorological conditions through a skyline detection method based on a fully convolutional neural network, and thus calculates the accurate roll angle and pitch angle of the drone. The detection algorithm of the present invention can achieve accurate pixel-by-pixel segmentation of the sky and non-sky regions without relying on any assumptions. At the same time, the present invention proposes a confidence estimation algorithm based on Gaussian discriminant analysis to provide a reliable value for the detection result for reference.
[0135] The framework of the drone attitude angle estimation algorithm with confidence in this embodiment is carried out according to the following steps:
[0136] Step A: Take the image captured in real time by the forward-looking camera of the drone as the input image I, and use a fully convolutional neural network structure N with VGG16 as the decoding network and deconvolution as the upsampling method full to achieve pixel-level segmentation of the sky and non-sky.
[0137] Step B: Obtain the segmentation quality Q through the probability map M, obtain the curvature T through the skyline position coordinates, and use the segmentation quality Q and the curvature T to perform offline training on the Gaussian discriminant algorithm. The training results are used to obtain the optimal classification threshold. During the flight of the drone, the confidence of the fitted straight line of the detected skyline is estimated online in real time and judged against the optimal classification threshold: When the confidence is higher than the classification threshold, proceed to Step C; otherwise, read the next frame.
[0138] Step C: Use this straight line to calculate the roll angle φ and pitch angle θ of the drone at this moment.
[0139] The specific design of the fully convolutional neural network structure in Step A is as Figure 2 shown, specifically:
[0140] 1) An encoding network mainly based on VGG16
[0141] In this embodiment, VGG16 is used as the encoding network to extract features through continuous convolutional layers and pooling layers to generate corresponding heat maps. Different from this, three fully connected layers of VGG16 in this encoding network are changed to convolutional layers, and the rest are retained. Secondly, the output channels are adjusted to 2, corresponding to 2 categories of "sky" and "non-sky" respectively. As Figure 2As shown below, the specific modification method is as follows: In this embodiment, the input image size of the encoding network is set to 256×256×3. After a series of convolutional layers and pooling layers, the image size is reduced to a data volume of 15×15×512. The first fully connected layer of VGG16 is adjusted to a convolutional layer with a convolutional kernel size of k = 7, and the output feature map size is 9×9×4096; the second fully connected layer of VGG16 is adjusted to a convolutional layer with a convolutional kernel size of k = 1 and a depth of c = 2, and the output feature map size is 9×9×4096; the third fully connected layer of VGG16 is adjusted to a convolutional layer with a convolutional kernel size of k = 1 and a depth of c = 2, and the output feature map size is 9×9×2. This feature map is the heat map H output by the encoding network.
[0142] 2) Use a deconvolution-based decoding network
[0143] This fully convolutional neural network uses deconvolution for upsampling to enlarge the image. However, it often cannot be enlarged exactly to the original image size and requires further size trimming, and finally decodes into a probability map M of the same size as the original image. After upsampling, the image size is enlarged to 320×320×2, and then cropped into a probability map M of 256×256×2. In this way, the heat map is restored to the same size as the input image and decoded into pixel-to-pixel classification probability values.
[0144] 3) Model training
[0145] The probability map M is class-labeled to generate a segmentation binary map O. The value of each point on the segmentation binary map represents the predicted classification of the pixel at the corresponding position. The gradient is updated and the model is trained by calculating the softmax loss function between the predicted classification and the actual classification.
[0146] 4) Optimal straight line equation extraction
[0147] The image or real-time video information is input into the trained fully convolutional neural network, and a segmentation binary map of the same size as the original image is output. According to the segmentation binary map, the maximum contour coordinates of the sky region predicted to be classified are obtained, and the lower boundary coordinates are extracted as the detected skyline coordinate set U sky ;
[0148] According to the detected skyline coordinates, the RANSAC algorithm is used to synthesize the optimal fitting straight line L p , and the fitting equation is:
[0149] L p = ax + b
[0150] The specific steps in step B are as follows:
[0151] 1) Calculate the segmentation quality Q and the predicted skyline curvature T of the fully convolutional neural network.
[0152] In the present invention, the Gaussian discriminant analysis algorithm needs to perform multivariate Gaussian modeling on the segmentation quality Q and the curvature T to evaluate the reliability of the detection result. In this embodiment, it is proposed that the segmentation quality Q is the absolute value of the average probability difference between the sky and non-sky regions, which characterizes the accuracy of the detection result of the fully convolutional neural network; the skyline curvature T is the predicted skyline S p and the fitted straight line L obtained by the least squares method p The average pixel distance between them measures the degree of curvature of the detected skyline. The calculation formulas for the segmentation quality Q and the skyline curvature T are as follows:
[0153]
[0154]
[0155] Q = |μ0 - μ1|
[0156]
[0157] where μ0 represents the average probability value that a pixel predicted as the sky becomes the sky, and μ1 represents the average probability value that a pixel predicted as non-sky becomes the sky. The larger the segmentation quality Q, the greater the probability difference between the sky and non-sky pixels becoming the sky, indicating a better segmentation effect. and respectively represent the predicted skyline S p and the fitted straight line L p The row numbers in the j-th column of the image, and N is the total number of columns in the test image or video.
[0158] 2) Prepare multivariate sample data for offline training.
[0159] The multivariate sample data can be expressed as (x (1) , y (1) ), (x (2) , y (2) ), (x (3) , y (3) ), …, (x (m) , y (m) ), y (i) ∈0, 1. x represents the multivariate sample data, which in this embodiment are the values of the segmentation quality Q and the curvature T; y represents the sample category, y (i) = 1 represents that the skyline detection result is reliable and has a high accuracy; y (i) = 0 represents that the skyline detection result is unreliable and has a low accuracy. For the value of the true sample category y, in this embodiment, according to the predicted skyline S p the actual fitted straight line L r and the predicted fitted straight line Lp It is determined by the average pixel error between them. The specific label setting rules are as follows:
[0160]
[0161] 3) Determine the optimal classification threshold.
[0162] The probability value obtained from Gaussian discriminant analysis ranges from [0, 1]. It is necessary to use the ROC curve to further determine the optimal classification threshold to obtain the best classification result. The abscissa and ordinate in the ROC curve are the true positive rate TPR and the false positive rate FPR respectively. The ROC curve is drawn by traversing all thresholds. When the TPR reaches the highest and the FPR reaches the lowest (i.e., the ROC curve is the steepest) at a certain threshold or threshold interval, the classification accuracy of the model is the highest. At this time, the threshold or threshold interval is set as the optimal threshold. The detection result is determined whether it can be defined as "reliable" according to the optimal threshold. Figure 3 In the ROC curve shown in (a), the AUC of the Gaussian discriminant model is 0.99, indicating that the classification model has excellent quality. The range of the optimal threshold selected from the ROC curve is [0.67, 0.72]. Therefore, the threshold is determined to be 0.70. Among them, the calculation formulas of the true positive rate TPR and the false positive rate FPR are respectively:
[0163]
[0164]
[0165] 4) Real-time confidence estimation.
[0166] After the multi-sample data preparation is completed, the trained Gaussian discriminant analysis model is used to perform real-time estimation of the confidence of the skyline fitting line. The specific process is as follows: First, estimate the prior probability and the mean and covariance matrix of the multivariate Gaussian distribution, and then use the Bayes formula to calculate the probabilities of a new sample belonging to two categories respectively. Among them, the probability p(y = 0|x) of belonging to the sky category is the confidence P of the skyline fitting line required.
[0167] 5) Validity judgment.
[0168] If the confidence is higher than the optimal classification threshold, it is considered that the skyline fitting line is valid, and then step C is performed, and the attitude angle calculated in step C is adopted by the navigation device. On the contrary, if the confidence is lower than the optimal classification threshold, it is considered that the skyline fitting line is invalid, and step C is directly skipped, and the system reads the next frame.
[0169] Figure 3 (b) and (c) are the confidence distribution plots of the skyline fitting lines for actual negative and positive samples respectively. With a threshold of 0.7 as the optimal classification threshold, 207 out of 210 test samples are predicted correctly, indicating that Gaussian discriminant analysis can perform confidence judgment with relatively high accuracy.
[0170] Step C is specifically as follows:
[0171] According to the fitting line L p , inversely deduce the roll angle φ and pitch angle θ of the UAV at the current moment:
[0172]
[0173]
Claims
1. A method for detecting the attitude of an unmanned aerial vehicle with confidence estimation, characterized in that, Including the following steps: Step A: Read the input image of the current frame, extract image features through a fully convolutional neural network and encode them into corresponding heatmaps, enlarge the heatmaps to the size of the input image, decode them into classification probabilities for each pixel, output a probability map, perform pixel-by-pixel class calibration on the probability map to generate a segmentation binary map, obtain the image of the sky region, extract the skyline coordinates from the image of the sky region, and fit the optimal straight-line equation according to the skyline coordinates to obtain the skyline fitting straight line; Step B: Quantify the segmentation quality based on the probability map and the segmented binary map and the curvature of the skyline ; Use the trained Gaussian discriminant analysis model to perform multivariate Gaussian modeling on the segmentation quality and the curvature ; According to the learned sample distribution, the Gaussian discriminant analysis model calculates the confidence of the skyline fitting line. If the confidence of the skyline fitting line is higher than the preset optimal classification threshold, then proceed to Step C; otherwise, return to Step A to read the input image of the next frame Segmentation quality and degree of curvature are calculated as follows: Among them, represents the average probability value that the pixels predicted as the sky become the sky, represents the average probability value that the pixels predicted as non-sky become the sky, and respectively represent the predicted skyline and the skyline fitting straight line at the row number of the column in the input image, and N is the total number of columns in the input image; represents the value at the coordinate in the channel of the probability map , and the channel represents the sky; is the predicted label corresponding to the pixel at the coordinate in the segmentation binary map , is the input image; Step C: Based on the skyline fitting straight line, estimate the attitude angle information of the UAV in real time.
2. The method for detecting the attitude of a drone with confidence estimation according to claim 1, wherein, In Step A, the fully convolutional neural network includes an encoding network, a decoding network, a class calibration module, and an optimal straight-line extraction module; Specifically, Step A is as follows: Use the encoding network to extract image features and encode them into corresponding heatmaps; the decoding network uses an upsampling method to enlarge the heatmaps to the size of the input image and decode them into classification probabilities for each pixel, outputting a probability map; the class calibration module performs pixel-by-pixel class calibration on the probability map to generate a segmentation binary map and obtain the image of the sky region; the optimal straight-line extraction module extracts the skyline coordinates from the image of the sky region and fits the optimal straight-line equation according to the skyline coordinates to obtain the skyline fitting straight line.
3. The method for detecting the attitude of an unmanned aerial vehicle with confidence estimation according to claim 2, wherein The decoding network uses an upsampling method to enlarge the heatmaps to the size of the input image and decode them into classification probabilities for each pixel, outputting a probability map, which is expressed as: Among them, represents the decoding network and represents upsampling; serves as the input to the decoding network; represents the probability map, which is the output of the decoding network; indicates the probability map at the coordinate in the channel with the value of taking values of 0 or 1; represents the pixel at the coordinate in the input image with the coordinate of 4. The method for detecting the attitude of a drone with confidence estimation according to claim 1, wherein In Step B, the trained Gaussian discriminant analysis model in Step 2) is obtained through the following training method: Apply training samples to perform offline training on the Gaussian discriminant analysis model, where ; represents multivariate sample data and is the quantization value of the segmentation quality and the degree of curvature ; represents the category of the sample data, represents that the skyline fitting line is reliable; represents that the skyline fitting line is unreliable; Suppose the class of the sample data follows a Bernoulli distribution under the given conditions, and the sample data in different classes respectively follow a multivariate Gaussian distribution: Among them, represents the Bernoulli distribution, and represent the mean and covariance of the multivariate Gaussian distribution respectively. Then we have: )) )) Obtained through the maximum likelihood estimation function , and The values of the three parameters: Given the known sample data obtained according to Bayes' formula the probability values of the sample data being positive or negative samples for the class of the sample data are as follows: Among them, is considered the confidence level of the skyline fitting line, and the value range is , and the ROC curve is used to further determine the optimal classification threshold.
5. The method for detecting the attitude of a drone with confidence estimation according to claim 1, wherein, In Step A, the skyline coordinates are extracted from the image of the sky region, and the optimal straight-line equation is fitted according to the skyline coordinates to obtain the skyline fitting straight line. Specifically: The lower boundary coordinates of the largest contour in the sky region are extracted as the skyline coordinates, and a straight line is fitted from the skyline coordinates using a filtering point algorithm to obtain the skyline fitting straight line.
6. The method for detecting the attitude of a drone with confidence estimation according to claim 1, wherein In Step B, the method for setting the optimal classification threshold is as follows: Use a large number of samples to offline train a Gaussian discriminant analysis model, and obtain the optimal classification threshold of the confidence level of the skyline fitting straight line through the obtained training results.
7. The method for detecting the attitude of a drone with confidence estimation according to claim 1, characterized in that, Specifically, Step C is: Fitting a straight line equation based on the obtained skyline , through geometric calculation, it can be known that the roll angle and the pitch angle The calculation formulas are respectively: Among them, and are the camera internal parameters, is the principal point coordinate.
8. An unmanned aerial vehicle attitude detection system with confidence estimation, characterized in that, Including: A fully convolutional neural network, which is used to read the input image of the current frame, extract image features through the fully convolutional neural network and encode them into corresponding heatmaps, enlarge the heatmaps to the size of the input image, decode them into classification probabilities for each pixel, output a probability map, perform pixel-by-pixel class calibration on the probability map to generate a segmentation binary map, obtain the image of the sky region, extract the skyline coordinates from the image of the sky region, and fit the optimal straight-line equation according to the skyline coordinates to obtain the skyline fitting straight line; A confidence estimation module, which is used to quantify the segmentation quality according to the probability map and the segmentation binary map and the curvature of the skyline ; perform multivariate Gaussian modeling on the segmentation quality and the curvature by using the trained Gaussian discriminant analysis model; according to the learned sample distribution, the Gaussian discriminant analysis model obtains the confidence of the skyline fitting line. If the confidence of the skyline fitting line is higher than the preset optimal classification threshold, the UAV attitude angle estimation module works. Otherwise, the fully convolutional neural network reads the input image of the next frame; the segmentation quality and the curvature are calculated as follows: Among them, represents the average probability value that the pixels predicted as the sky become the sky, represents the average probability value that the pixels predicted as non-sky become the sky, and respectively represent the predicted skyline and the skyline fitting straight line at the row number of the th column of the input image, and N is the total number of columns in the input image; represents the value of the probability map at the coordinate in the channel , and the channel represents the sky; is the predicted label corresponding to the pixel at the coordinate in the segmentation binary map , and is the input image; A UAV attitude angle estimation module, which is used to estimate the UAV attitude angle information in real time through geometric calculations and the equation of the skyline fitting straight line.
9. The UAV attitude detection system with confidence estimation according to claim 8, characterized in that The fully convolutional neural network includes an encoding network, a decoding network, and a class calibration and optimal straight-line extraction module; The encoding network is used to extract image features and encode them into corresponding heatmaps; The decoding network is used to use an upsampling method to enlarge the heatmaps to the size of the input image and decode them into classification probabilities for each pixel, outputting a probability map; The class calibration module is used to perform pixel-by-pixel class calibration on the probability map to generate a segmentation binary map and obtain the image of the sky region; The optimal straight line extraction module is used to extract the skyline coordinates from the image of the sky area, and fit the optimal straight line equation according to the skyline coordinates to obtain the skyline fitting straight line.
Citation Information
Patent Citations
Unmanned aerial vehicle landing method based on optical flow method and horizon line detection
CN105644785A
Sea-sky-line online detection method based on full convolutional neural network
CN110705623A