Unmanned Aerial Vehicle Forest Exploration System Based on Monocular Depth Prediction and Deep Reinforcement Learning

By adopting monocular depth of field prediction and deep enhancement learning methods on drones, the problems of drones' autonomous obstacle avoidance and target detection in forest environments are solved, and the efficient exploration and search and rescue capabilities of drones in complex environments are achieved.

CN114943757BActive Publication Date: 2025-06-20ZHEJIANG LAB +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210622959.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-06-20
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

The prior art is difficult to achieve autonomous obstacle avoidance and target detection of drones in forest environments, especially when obstacles are dense and the environment is unknown.

Method used

Using a method based on monocular depth of field prediction and deep enhancement learning, obstacle distance information is obtained through the depth of field prediction recognition optimization module, obstacle avoidance strategy is constructed in combination with the PPO deep enhancement learning algorithm, and a lightweight object detection model is deployed on the onboard computer for real-time detection.

Benefits of technology

It realizes autonomous obstacle avoidance and real-time target detection of drones in dense forest scenes, and improves the exploration and search and rescue capabilities of drones in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943757B_ABST
    Figure CN114943757B_ABST
Patent Text Reader

Abstract

The present invention provides a drone forest exploration system based on monocular depth prediction and deep reinforcement learning. Logically, the present invention mainly includes four parts: depth prediction recognition optimization, reinforcement learning rule-based obstacle avoidance, target detection, and a scalable perception, computing, and control drone link formed in the form of a RESTful service. In terms of the process, first, a monocular depth prediction model for forest scenes is established, and according to a whole-image uncertainty estimation method, the uncertainty of the depth prediction model result is estimated by estimating the variance, solving the problem of depth prediction failure in drone flight samples; then, an obstacle avoidance mode with the avoidance of small obstacles in forest scenes as the planning target is constructed under the hardware constraints of a monocular forward drone; finally, the PPO deep reinforcement learning algorithm is used to achieve drone obstacle avoidance and navigation, solving the problem of possible entrapment in local optimal regions after policy training, and personnel search is realized by deploying a lightweight target detection network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a drone forest exploration technology based on monocular depth prediction and deep reinforcement learning. Background Art

[0002] Due to the complexity of the forest environment, with its variable internal terrain and dense trees, it is very difficult for ground robots to operate smoothly. Therefore, forest exploration and search and rescue are difficult tasks. Thanks to the unique dexterity of multi-rotor drones, they are suitable for replacing manual work in scenarios such as forests with dense obstacles and complex terrains, which can greatly speed up the search and rescue process and save a large amount of manpower and material resources. In the forest exploration task, the drone needs to be able to fly autonomously while avoiding obstacles. Relying solely on manual flight control is difficult to complete flight tasks in complex obstacle scenarios such as forests. Therefore, the drone needs to have the ability to avoid obstacles autonomously. In addition, in the exploration and search and rescue tasks, the drone needs to determine the search and rescue target autonomously. When operating the flight control manually, it is difficult to judge the search area through the limited photos taken from the drone's perspective. Therefore, the autonomous exploration of drones for forest environments is a problem with great application value.

[0003] There are many methods to achieve autonomous flight of drones in the prior art, but most of them consider scenarios with sparse obstacles or known global environments, and are not applicable to scenarios such as forest search and rescue with unknown environments and dense obstacles. To achieve the exploration and search and rescue of multi-rotor drones in forest environments, support needs to be provided in terms of hardware and algorithms. In terms of hardware, although lidar is a method with good effects and widely used in commercial drones, lidar has disadvantages such as high cost and large volume. In terms of algorithms, obstacle avoidance is the key to enabling the drone to fly smoothly in a dense forest. Most commercial drones, including DJI, adopt the method of immediately stopping flying after identifying an obstacle. Although this approach can achieve obstacle avoidance, it cannot meet the requirements of flight tasks. Planning a path based on obtaining global environment information through SLAM is also a common method, and its biggest disadvantage is that map building is time-consuming. Moreover, in forest scenarios, most of the obstacles are small obstacles such as trees, which generally will not be marked on the map and are prone to change over time. A better method should be to perform real-time avoidance based on the currently perceived information. In addition, in exploration and rescue, the drone also needs to be able to autonomously determine whether it has searched for other targets to be detected such as trapped people. Using object detection methods based on deep learning can obtain a very robust detection model. Thanks to the rapid development of the Internet of Things and GPUs in recent years, it has become possible to carry low-power CPUs and GPUs on drones to support real-time deep learning calculations. Summary of the Invention

[0004] The present invention first uses a depth of field prediction and recognition optimization module to obtain the obstacle distance information and its predicted uncertainty in the forest environment where the UAV is currently located. Then, it uses the PPO deep reinforcement learning algorithm and a stochastic simulation forest environment to construct a training UAV obstacle avoidance strategy to improve the generalization ability and learning ability of the model, and constructs a forward monocular UAV forest scene obstacle avoidance mode. At the same time, a lightweight object detection model trained for personnel detection in dense forest scenes is deployed on the on-board computer to detect in real time whether other search and rescue targets such as people appear in the UAV's field of view. Finally, an extensible perception, computing, and control UAV link is formed in the form of a RESTful service, realizing UAV exploration in forest scenes.

[0005] To solve the above problems, the present invention discloses a UAV forest exploration system based on monocular depth prediction and deep reinforcement learning for dense forest scenes, including the following steps:

[0006] Step 1: Define a monocular depth prediction model based on self-supervised learning for dense forest scenes using the following steps:

[0007] Step 1-1: Construct a training specific dataset, which is constructed by frame-by-frame extraction of the video recorded by the experimental UAV when traveling in a dense forest scene;

[0008] Step 1-2: Process the video taken by the experimental UAV to obtain an image sequence of length N. Take the current frame in the image sequence as frame t, and the next frame of the current frame as the next frame t' to be reconstructed;

[0009] Step 1-3: Input the current frame t obtained in Step 1-2 into the constructed depth network Dispnet. After passing through the fully convolutional network, the convolutional layer feature information of the same size is sequentially stacked and connected, and after upsampling, the depth information is finally output. Input the t frame and t' frame obtained in Step 1-2 into the constructed camera pose network posenet to obtain six scalar values of the translation and rotation along the X, Y, and Z axes between the two frames of the estimated camera. It is directly constructed using convolutional layers and fully connected layers, and then based on these six scalar values, a pose change estimation matrix between the two frames is generated. Then reconstruct frame t';

[0010] Step 1-4: Calculate the reconstruction loss of a specific pixel location and the depth smoothness loss through the two frames reconstructed in Step 1-3. and the depth smoothness loss

[0011] The reconstruction loss of a specific pixel location is constructed as:

[0012]

[0013] In the above formula, SSIM represents the structural similarity index. The depth smoothing loss is constructed as follows:

[0014]

[0015] Finally, the complete loss function of the entire model is:

[0016]

[0017] In the above formula, is the weight of the sliding loss;

[0018] Steps 1-5: According to the relative depth, obtain the absolute depth of the actual obstacle from the camera according to the following formula, The specific steps are to take points on multiple potential obstacles in 20 unfamiliar scene pictures in the experiment that are not in the test set, and measure the actual depth distance with a rangefinder. At the same time, obtain the pixel depth estimation value of the corresponding picture in this area, and calculate the ratio based on this value, and restore the absolute depth based on this ratio and calculate the error;

[0019] Steps 1-6: According to a method for estimating the uncertainty of the whole image, estimate the uncertainty by estimating the variance, which can effectively be used as an index for quickly identifying the uncertainty of the whole image depth estimation; calculate θ according to the combination ratio of SSIM and L1 loss, and judge the threshold β. During the flight of the drone, if |motion perception information S| > 0 or |current control signal C| > 0, according to I t and I t-1 , calculate the average difference between the two images

[0020]

[0021] Estimate the depth t of I and the pose change based on the network model Calculate the average difference between the reconstructed area and the original image location

[0022]

[0023] Thus, the calculated uncertainty

[0024] Step 2: Construct an obstacle avoidance mode with the avoidance of small obstacles in forest scenes such as trees as the planning goal under the hardware constraints of a single forward-facing drone, analyze the characteristics of depth of field prediction in Step 1, and design a basic but effective obstacle avoidance algorithm for forest scenes;​

[0025] Step 2-1: Obtain the current perception image and the motion perception information S from the drone in Step 1, determine the link width w of the drone obstacle avoidance system, fix the forward speed s, the feasible vector W, initialize the sliding window size o according to the drone size, the proximity threshold Unconfid, the timeout threshold Outtime, and the redirection threshold Redirect;

[0026] Step 2-2: Determine the distance between the current position Pnow and the target position Ptar. If it is less than the proximity threshold Near, the loop ends; otherwise, execute the following steps sequentially;

[0027] Step 2-3: Pass the current perception image I t through Step 1 to obtain the monocular depth prediction map and the pose matrix Intercept the estimated map The middle area Area and calculate the uncertainty R unconfid , if the uncertainty R unconfid is greater than the uncertainty threshold R unconfid , then the current waiting time +1, and determine the waiting time t. If t is greater than the waiting timeout threshold Outtime, alarm for manual processing or landing, and return to Step 2-2;

[0028] Step 2-4: Compress Area into a one-dimensional vector, mark the ones below the threshold Danger in the W vector, and perform smoothing processing;

[0029] Step 2-5: Traverse bias w times, slide the o window in W to find the feasible Yaw window and mark it. If there is no feasible region, rotate the Yaw angle clockwise by one FOV; otherwise, use min(Yawtarget, Yawnext) to weight-select Yaw in the feasible Yaw and rotate to this position;

[0030] Step 2-6: Move forward at a fixed speed s;

[0031] Step 2-7: Determine if the redirection threshold Redirect == 0 and cos(direction of travel, target direction) is less than zero, then assign Yawtarget to Yaw, and assign 30 to Redirect;

[0032] Step 2-8: Reset Redirect -= 1, the waiting timeout value e = 0, reset all elements of the W vector, and execute Step 2-2;

[0033] Step 3: Use the deep reinforcement learning strategy to solve the problem of drone obstacle avoidance and navigation based on self-supervised monocular depth prediction;

[0034] Step 3-1: Obtain the environmental state s each time t and the reward r are input into the reinforcement learning decision model, initialize the policy model variable θ, and the value model variable Initialize the traversal parameter k, whose initial value is 1, and set the reward r as the following distribution function:

[0035]

[0036] Step 3-2: Run the policy Collect the trajectory D that changes in the environment T times k , and obtain and calculate A based on the current value function π (s t ,a t ), and update the policy parameter θ based on stochastic gradient ascent to increase the target value with a clip function:

[0037]

[0038] Optimize the value estimation model parameters through mean squared error loss

[0039]

[0040] Step 3-3: Increase the value of k by 1, return to execute Step 3-2 until the value of k reaches the set termination parameter;

[0041] Step 3-4: The advantage value A π can be directly approximated based on the previous policy output and the actual reward obtained from the trajectory after masking. The depth perception observed by the 64*3 drone for the current and the previous two frames, and the remaining different state variables include the following:

[0042] (alphaCos*3,alphaSin*3,lastaction*3,inTempTarget,stopTime)

[0043] alphaCos and alphaSin are the cosine and sine values of the angle between the three-frame target vector and the current direction vector respectively, lastaction is the previous action value, inTempTarget indicates whether in the temporary target search, and stopTime is the time count for the drone to stay in place;

[0044] Thus, obtain π θ ​The output of (·|s) is the (yaw, pitch, tempTarget) vector of specific behaviors, which respectively represent the rotation degree of the current Yaw angle, with the range of [-20.45, 20.45] degrees, the forward distance amplitude, and the target direction angle if using the temporary target. And the model only outputs the average overall value in the current state;

[0045] Step 4: Define a lightweight object detection network model based on self-supervised learning using the following steps:

[0046] Step 4-1: Construct a human detection dataset by using an open-source human object detection dataset and pictures taken by the experimental drone marked by ourselves in the dense forest scene. Randomly flip, scale, and change the color gamut of the pictures in the dataset to enhance the dataset, and preprocess the pictures in the dataset to meet the input of the neural network. At the same time, divide the training set, test set, and validation set according to a certain ratio;

[0047] Step 4-2: Construct the YOLO V4 object detection network model structure, including a backbone feature extraction network for preliminary feature extraction, an enhanced feature extraction network for enhanced feature extraction, and a prediction network for obtaining the final prediction result. The backbone feature extraction network can obtain three preliminary effective feature layers, and the enhanced feature extraction network performs feature fusion on the three preliminary effective feature layers to extract better features and obtain three more effective effective feature layers. The final prediction network uses the more effective effective feature layers to obtain the prediction result. To reduce the number of network parameters and the amount of computation, replace the backbone feature extraction network with the lightweight MobileNet V3 and use depthwise separable convolution instead of the ordinary convolution used in YOLOV4;

[0048] Step 4-3: Decode the prediction result output by the prediction network. Add its corresponding x_offset and y_offset to each grid point to obtain the center of each prediction box, and then use the prior box, h, and w to calculate the length and width of the prediction box. Then use the position and score of the box for non-maximum suppression to obtain the position of the entire prediction box;

[0049] Step 4-4: Define the loss function during training. The loss function consists of three parts: the regression loss function, the confidence loss function, and the loss function for the predicted category. Among them, the regression loss function is defined as 1 - CIoU. IoU is a ratio concept and is insensitive to the scale of the target object. CIOU takes into account the distance between the target and the anchor, the overlap rate, the scale, and the penalty term, making the target box regression more stable and avoiding problems such as divergence during training like IoU. The penalty factor takes into account the aspect ratio of the predicted box fitting the aspect ratio of the target box. The specific formula for CIoU is as follows:

[0050]

[0051] where ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the predicted box and the ground truth box. c represents the distance of the diagonal of the smallest closed region that can contain both the predicted box and the ground truth box. The formulas for α and v are as follows:

[0052]

[0053]

[0054] Step 4-5: Conduct multiple rounds of training on the training set to obtain the final object detection model. Export the parameter file of the model and deploy it to the on-board computer for actual testing. Preprocess each frame of the image in the drone camera and put it into the model to obtain the prediction result.

[0055] Advantages of the present invention:

[0056] (1) By measuring the depth of the obstacle point group within a certain range, obtain the conversion ratio between the relative depth estimate and the absolute depth and give the estimation error, providing depth of field information for subsequent forest exploration of the drone. Analyze the situation where the monocular depth prediction method has estimation errors in the drone scenario, and propose the method of using the original image as a reference for the average reconstruction loss difference to evaluate the estimation uncertainty, providing a solution for early judgment of whether the obstacle depth estimation fails in drone forest exploration and navigation.

[0057] (2) Aiming at the deficiency that the traditional method of constructing a global map based on SLAM and then performing path planning is not suitable for the dense forest environment, introduce the depth reinforcement learning algorithm. Model the obstacle avoidance mode based on the depth reinforcement learning method, construct a training strategy network for the dense forest simulation model based on the obstacle avoidance mode, discuss the limitations of the reinforcement learning method, and give a method of temporary random target based on the strategy with reference to the simulated annealing algorithm. Finally, integrate it into a complete reinforcement learning obstacle avoidance strategy through the state machine mode.

[0058] (3) Multiple flight experiments were conducted in a real forest scenario, simultaneously verifying the effectiveness of the UAV obstacle avoidance method based on monocular depth prediction and the lightweight target detection network deployed on the on-board computer. It has the actual avoidance ability in a dense forest scenario, can detect people in real time, and return the detected positions through the established UAV communication network. Description of the Drawings

[0059] Figure 1 is the architecture diagram of the depth estimation network;

[0060] Figure 2 is an example of depth estimation for test set images;

[0061] Figure 3 is an example of depth estimation for images in an unfamiliar scenario;

[0062] Figure 4 is an example of depth estimation failure during flight;

[0063] Figure 5 is an example of depth estimation failure in abnormal situations;

[0064] Figure 6 is an example of depth expectation and variance estimation;

[0065] Figure 7 is an example of image depth estimation in different normal and failure scenarios;

[0066] Figure 8 is a simplified schematic diagram of the operation of the UAV obstacle avoidance method;

[0067] Figure 9 is the architecture of the policy and value network;

[0068] Figure 10 is the state transition of the combined obstacle avoidance method;

[0069] Figure 11 is the comparison of training rewards with and without policy-based temporary targets;

[0070] Figure 12 is the comparison of training with different reinforcement learning algorithms;

[0071] Figure 13 is the service providing structure of the UAV module;

[0072] Figure 14 is the UAV forest experiment environment. Detailed Implementation Manner

[0073] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. It should be noted that the terms "front", "rear", "left", "right", "up" and "down" used in the following description refer to the directions in the accompanying drawings, and the terms "inner" and "outer" respectively refer to the directions towards or away from the geometric center of a specific component.

[0074] This embodiment provides a drone forest exploration system based on monocular depth prediction and depth reinforcement learning, including the following steps;

[0075] Step 1: Define a monocular depth prediction model based on self-supervised learning for dense forest scenes by the following steps:

[0076] Step 1-1: Construct a training specific dataset, which is constructed by frame-by-frame intercepting of the video recorded by the experimental drone when traveling in the dense forest scene;

[0077] Step 1-2: Process the video taken by the experimental drone to obtain an image sequence of length N. Take the current frame in the image sequence as frame t, and the next frame of the current frame as the next frame t' to be reconstructed;

[0078] Step 1-3: Input the current frame t obtained in Step 1-2 into the constructed depth network Dispnet, the network structure of which is as Figure 1 shown. After passing through the fully convolutional network, the convolutional layer feature information of the same size is sequentially stacked and connected, and after upsampling, the depth information is finally output The test results are as Figure 2 、 Figure 3 shown. Input the t frame and t' frame obtained in Step 1-2 into the constructed camera pose network posenet to obtain 6 scalar values of the translation and rotation along the X, Y, and Z axes between the two frames of the estimated camera. It is directly constructed using convolutional layers and fully connected layers, and then based on these six scalar values, a pose change estimation matrix between the two frames is generated Subsequently, reconstruct frame t';

[0079] Step 1-4: Calculate the reconstruction loss at a specific pixel location and the depth smoothness loss through the two frames reconstructed in Step 1-3 and

[0080] The reconstruction loss at a specific pixel location is constructed as:

[0081]

[0082] In the formula, SSIM represents the structural similarity index.

[0083] The depth smoothness loss is constructed as:

[0084]

[0085] Finally, the complete loss function of the entire model is as follows:

[0086]

[0087] Among them, is the weight of the sliding loss.

[0088] Steps 1-5: According to the relative depth, obtain the absolute depth of the actual obstacle from the camera according to the following formula The specific steps are to take points on multiple potential obstacles in 20 unfamiliar scene pictures in the experiment that are not in the test set, measure the actual depth distance with a rangefinder, and at the same time obtain the pixel depth estimation value of the corresponding picture in this area. Based on this, calculate the ratio value, and restore the absolute depth based on this ratio and calculate the error;

[0089] Steps 1-6: There are situations where the depth estimation fails as shown in Figure 4 and Figure 5 . Therefore, consider calculating the uncertainty of the depth estimation result. According to a whole-image uncertainty estimation method, estimate the uncertainty by estimating the variance Figure 6 is the result graph of the depth expectation and variance estimation, which can effectively be used as an index for quickly identifying the uncertainty of the whole-image depth estimation; Calculate θ according to the combination ratio of SSIM and L1 loss, and judge the threshold β. During the flight of the drone, if |motion perception information S|>0 or |current control signal C|>0, according to I t and I t-1 , calculate the average difference between the two pictures

[0090]

[0091] Estimate the depth of I t based on the network model and the pose change reconstruct Calculate the average difference between the reconstructed area and the original picture location average difference

[0092]

[0093] Thus, obtain the calculated uncertainty Figure 7 are the image depth estimation results in different normal and failure scenarios.

[0094] Step 2: Construct an obstacle avoidance mode with the avoidance of small obstacles in forest scenarios such as trees as the planning goal under the hardware constraints of a single forward drone, Figure 8 which is a simplified flowchart of the operation process of this mode. Analyze the depth of field prediction characteristics in Step 1 and design a basic but effective obstacle avoidance algorithm for forest scenarios;

[0095] Step 2-1: Obtain the current perception image and the motion perception information S from the drone in Step 1. Determine the link width w of the drone obstacle avoidance system, the fixed forward speed s, the feasible vector W, initialize the sliding window size o according to the size of the drone, the proximity threshold Unconfid, the timeout threshold Outtime, and the redirection threshold Redirect;

[0096] Step 2-2: Determine the distance between the current position Pnow and the target position Ptar. If it is less than the proximity threshold Near, the loop ends; otherwise, execute the following steps in sequence;

[0097] Step 2-3: Pass the current perception image I t through Step 1 to obtain the monocular depth of field prediction map and the pose matrix to intercept the estimated map in the middle area Area and calculate the uncertainty R unconfid . If the uncertainty R unconfid is greater than the uncertainty threshold R unconfid , the current waiting time is incremented by 1, and the waiting time t is determined. If t is greater than the waiting timeout threshold Outtime, an alarm is issued for manual handling or landing, and return to Step 2-2;

[0098] Step 2-4: Compress Area into a one-dimensional vector, mark those below the threshold Danger in the W vector, and perform smoothing processing

[0099] Step 2-5: Traverse bias w times, slide the o window in W to find the feasible Yaw window and mark it. If there is no feasible region, rotate the Yaw angle clockwise by one FOV; otherwise, use min(Yawtarget, Yawnext) to weighted-select Yaw in the feasible Yaw and rotate to this position

[0100] Step 2-6: Move forward at a fixed speed s

[0101] Step 2-7: Determine if the redirection threshold Redirect == 0 and cos(direction of travel, target direction) is less than zero. If so, assign Yawtarget to Yaw and assign 30 to Redirect,

[0102] Step 2-8: Reset Redirect-=1, wait for the timeout value e=0, reset all elements of the W vector, and execute Step 2-2;

[0103] Step 3: Use the deep reinforcement learning strategy to solve the problem of UAV obstacle avoidance and navigation based on self-supervised monocular depth prediction. The architecture of its policy value network is as Figure 9 shown.

[0104] Step 3-1: Obtain the environmental state s t and the reward r at each time and input them into the reinforcement learning decision model. Initialize the policy model variable θ and the value model variable Initialize the traversal parameter k, whose initial value is 1, and set the reward r as the following distribution function:

[0105]

[0106] Step 3-2: Run the policy Collect the trajectory D that changes in the environment for T times k , and obtain and calculate A based on the current value function π (s t , a t ). Update the policy parameter θ based on stochastic gradient ascent to increase the target value with the clip function:

[0107]

[0108] Optimize the value estimation model parameters through the mean squared error loss

[0109]

[0110] Step 3-3: Increase the value of k by 1, return to execute Step 3-2 until the value of k reaches the set termination parameter;

[0111] Step 3-4: The advantage value A π can be directly approximated based on the output of the previous policy and the actual reward obtained from the trajectory after masking. The depth perception observed by the 64*3 UAV for the current and the previous two frames, and the remaining different state variables include the following:

[0112] (alphaCos*3, alphaSin*3, lastaction*3, inTempTarget, stopTime)

[0113] alphaCos and alphaSin are the cosine and sine values of the angles between the three-frame target vectors and the current direction vector respectively, lastaction is the previous action value, inTempTarget indicates whether in the temporary target search, and stopTime is the time count for the UAV to stay in place. The state transition of the combined obstacle avoidance method for the entire process is as Figure 10 shown.

[0114] Thus, π is obtained θ (·|s) outputs the (yaw, pitch, tempTarget) vector of the specific behavior, whose meanings are the rotation degrees of the current Yaw angle, in the range of [-20.45, 20.45] degrees, the distance amplitude for moving forward, and the target direction angle if using the temporary target. And the model only outputs the average overall value in the current state. The network training results of the entire model are as Figure 11 、 12 shown.

[0115] Step 4: Define a lightweight object detection network model based on self-supervised learning using the following steps:

[0116] Step 4-1: Construct a human body detection dataset by using an open-source human body target detection dataset and the pictures taken by the experimental UAV marked by oneself in the dense forest scene. Randomly flip, scale, and change the color gamut of the pictures in the dataset to enhance the dataset, and preprocess the pictures in the dataset to meet the input of the neural network. At the same time, divide the training set, test set, and validation set according to a certain ratio;

[0117] Step 4-2: Construct the YOLO V4 object detection network model structure, including a backbone feature extraction network for preliminary feature extraction, an enhanced feature extraction network for enhanced feature extraction, and a prediction network for obtaining the final prediction result. The backbone feature extraction network can obtain three preliminary effective feature layers, and the enhanced feature extraction network performs feature fusion on the three preliminary effective feature layers to extract better features and obtain three more effective effective feature layers. The final prediction network uses the more effective effective feature layers to obtain the prediction result. In order to reduce the number of network parameters and the amount of computation, replace the backbone feature extraction network with the lightweight MobileNet V3 and use depthwise separable convolution instead of the ordinary convolution used in YOLOV4.

[0118] Step 4-3: Decode the prediction results output by the prediction network. Add the corresponding x_offset and y_offset to each grid point to obtain the center of each prediction box. Then, calculate the length and width of the prediction box by combining the prior box with h and w. Finally, perform non-maximum suppression using the position and score of the box to obtain the position of the entire prediction box.

[0119] Step 4-4: Define the loss function during training. The loss function consists of three parts: the regression loss function, the confidence loss function, and the loss function for the predicted classes. The regression loss function is defined as 1-CIoU. IoU is a ratio concept and is insensitive to the scale of the target object. CIOU takes into account the distance between the target and the anchor, the overlap rate, the scale, and the penalty term, making the target box regression more stable and avoiding problems such as divergence during training like IoU. The penalty factor takes into account the aspect ratio of the predicted box fitting the aspect ratio of the target box. The specific formula for CIoU is as follows:

[0120]

[0121] where ρ 2 (b,b gt ) represents the Euclidean distance between the centers of the prediction box and the ground truth box. c represents the distance of the diagonal of the smallest closed region that can simultaneously contain the prediction box and the ground truth box. The formulas for α and v are as follows:

[0122]

[0123]

[0124] Step 4-5: Perform multiple rounds of training on the training set to obtain the final object detection model. Export the parameter file of the model and deploy it to the on-board computer for actual testing. Preprocess each frame of the image in the drone camera and put it into the model to obtain the prediction results. Figure 13 Shows the service provision structure of the entire system server module. The final experimental results of the forest scene are as Figure 14 shown.

[0125] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features.

Claims

1. A drone forest exploration system based on monocular depth prediction and deep reinforcement learning, characterized in that, including the following steps; Step 1: Define a monocular depth network posen prediction model based on self-supervised learning for dense forest scenarios using the following steps: The specific steps of Step 1 include the following steps; Step 1-1: Construct a training specific dataset, which is constructed by frame-by-frame intercepting of the video recorded by the experimental drone when traveling in the dense forest scenario; Step 1-2: Process the video captured by the experimental drone to obtain an image sequence of length N. Take the current frame in the image sequence as the t frame, and the next frame of the current frame as the next frame t' to be reconstructed; Step 1-3: Input the current frame t obtained in Step 1-2 into the constructed depth network Dispnet. After passing through the fully convolutional network, successively stack and connect the convolutional layer feature information of the same size, and perform upsampling, and finally output the depth information Input the t-frame and t'-frame obtained in Step 1-2 into the constructed camera pose network posenet to obtain six scalar values of the translation and rotation along the X, Y, and Z axes between the two frames of the estimated camera. It is directly constructed using convolutional layers and fully connected layers, and then based on these six scalar values, a pose change estimation matrix between the two frames is generated Subsequently, reconstruct the t'-frame; Step 1-4: Calculate the reconstruction loss of specific pixel locations and the depth smoothness loss based on the two frames reconstructed in Steps 1-3 and the depth smoothness loss The reconstruction loss of a specific pixel location is constructed as: In the above formula, SSIM represents the structural similarity index; the depth smoothness loss is constructed as: The final complete loss function of the entire model is: In the above formula, is the weight of the slip loss; Step 1-5: Obtain the absolute depth of the actual obstacle from the camera according to the relative depth using the following formula, The specific steps are as follows: Take points on multiple potential obstacles in 20 unfamiliar scene pictures in the experiment that are not in the test set, measure the actual depth distance with a rangefinder, and at the same time obtain the pixel depth estimation value of the corresponding picture in this area. Based on this, calculate the ratio value, and based on this ratio, restore the absolute depth and calculate the error; Steps 1-6: According to a method for estimating the uncertainty of the entire image, the uncertainty is represented by estimating the variance, which can effectively serve as an index for quickly identifying the uncertainty of the entire image depth estimation; calculate θ according to the combined ratio of SSIM and L1 loss, and judge the threshold β. During the flight of the drone, if |motion perception information S|>0 or |current control signal C|>0, according to I t and I t-1 , calculate the average difference between the two images Estimate I based on the network model t depth and pose change reconstruction Calculate the average difference between the reconstructed area and the original map location average difference Thus, the calculation uncertainty is obtained Step 2: Construct an obstacle avoidance mode with the avoidance of small obstacles in the forest scenario as the planning goal under the hardware constraints of the monocular forward drone. Analyze the depth prediction characteristics in Step 1 and design a basic but effective obstacle avoidance algorithm for the forest scenario; Step 3: Use a deep reinforcement learning strategy to solve the problem of obstacle avoidance and navigation of drones based on self-supervised monocular depth prediction; Step 4: Adopt a YOLO V4 lightweight object detection network model based on self-supervised learning to achieve personnel detection in the dense forest scenario.

2. The drone forest exploration system based on monocular depth prediction and deep reinforcement learning according to claim 1, characterized in that, The specific steps of Step 2 include the following: Step 2-1: Obtain the current perception image and the motion perception information S from the drone in Step 1. Determine the link width w of the drone obstacle avoidance system, the fixed forward speed s, the feasible vector W, initialize the sliding window size o according to the size of the drone, the proximity threshold Unconfid, the timeout threshold Outtime, and the redirection threshold Redirect; Step 2-2: Determine the distance between the current position Pnow and the target position Ptar. If it is less than the proximity threshold Near, the loop ends, otherwise execute the following steps in sequence; Step 2-3: Take the current perceived image I t After Step 1, obtain the monocular depth prediction map and the pose matrix Intercept the estimated map The middle area Area and calculate the uncertainty R unconfid , if the uncertainty R unconfid is greater than the uncertainty threshold R unconfid , then the current waiting time +1, and determine the waiting time t. If t is greater than the waiting timeout threshold Outtime, then alarm for manual processing or landing, and return to Step 2-2; Step 2-4: Compress Area into a one-dimensional vector, mark the ones below the threshold Danger in the W vector, and perform smoothing processing; Step 2-5: Traverse bias w times, slide the o window in W to find the feasible Yaw window and mark it. If there is no feasible region, rotate the Yaw angle clockwise by a FOV. Otherwise, use min(Yawtarget, Yawnext) to weighted select Yaw in the feasible Yaw and rotate to this; Step 2-6: Move forward at a fixed speed s; Step 2-7: Determine if the redirection threshold Redirect == 0 and cos(direction of travel, target direction) is less than zero, then assign Yawtarget to Yaw, and assign 30 to Redirect; Step 2-8: Reset Redirect -= 1, wait for the timeout value e = 0, reset all elements of the W vector, and execute Step 2-2.

3. The drone forest exploration system based on monocular depth prediction and deep reinforcement learning according to claim 1, characterized in that, The specific steps of Step 3 include the following: Step 3-1: Obtain the environmental state s each time t and the reward r are input into the reinforcement learning decision model, initialize the policy model variable θ, and the value model variable Initialize the traversal parameter k, whose initial value is 1, and set the reward r as the following distribution function: Step 3-2: Running Policy Collect the trajectories D that change in the environment T times k , and obtain and calculate A based on the current value function Calculate A π (s t , a t ), update the policy parameter θ based on stochastic gradient ascent to increase the objective value with the clip function: Optimize the value estimation model parameters by mean squared error loss Step 3-3: Increase the value of k by 1, and return to execute Step 3-2 until the value of k is the set termination parameter; Step 3-4: Advantage value A π It can be directly approximated based on the previous policy output and the reward mask actually obtained from the trajectory. The current and the previous two frames of depth perception observed by the 64*3 drone, and the remaining different state variables include the following: (alphaCos*3, alphaSin*3, lastaction*3, inTempTarget, stopTime) Here, alphaCos and alphaSin are respectively the cosine and sine values of the angle between the target vector of three frames and the current direction vector, lastaction is the previous action value, inTempTarget indicates whether in the process of searching for a temporary target, and stopTime is the time count for the UAV to stay in place; Thus, π is obtained. θ (·|s) outputs the (yaw, pitch, tempTarget) vector of the specific behavior, whose respective meanings are the rotation degrees of the current Yaw angle, with the range being [-20.45, 20.45] degrees, the distance amplitude moving forward, and the target direction angle if using the temporary target. The model only outputs the average overall value in the current state.

4. The unmanned aerial vehicle forest exploration system based on monocular depth prediction and depth reinforcement learning according to claim 1, characterized in that, The specific steps of step 4 are as follows: Step 4-1: Construct a human detection dataset by using an open-source human target detection dataset and the pictures taken by the experimental UAV marked by oneself in a dense forest scene. Randomly flip, scale, and change the color gamut of the pictures in the dataset to enhance the dataset, and preprocess the pictures in the dataset to meet the input requirements of the neural network. At the same time, divide the training set, test set, and validation set according to a certain ratio; Step 4-2: Construct the YOLO V4 target detection network model structure, including a backbone feature extraction network for preliminary feature extraction, an enhanced feature extraction network for enhanced feature extraction, and a prediction network for obtaining the final prediction result. The backbone feature extraction network can obtain three preliminary effective feature layers, and the enhanced feature extraction network performs feature fusion on the three preliminary effective feature layers to extract better features and obtain three more effective effective feature layers. The final prediction network uses the more effective effective feature layers to obtain the prediction result. In order to reduce the number of network parameters and the amount of computation, replace the backbone feature extraction network with the lightweight MobileNet V3, and use depthwise separable convolution instead of the ordinary convolution used in YOLO V4; Step 4-3: Decode the prediction result output by the prediction network. Add its corresponding x_offset and y_offset to each grid point to obtain the center of each prediction box, and then use the prior box, h, and w to calculate the length and width of the prediction box. Then, perform non-maximum suppression using the position and score of the box to obtain the position of the entire prediction box; Step 4-4: Define the loss function in the training process. The loss function consists of three parts, the regression loss function, the confidence loss function, and the loss function of the predicted class. Among them, the regression loss function is defined as 1-CIoU; IoU is a concept of ratio and is insensitive to the scale of the target object; CIOU takes into account the distance, overlap rate, scale, and penalty term between the target and the anchor, making the target box regression more stable; and the penalty factor takes into account the aspect ratio of the predicted box fitting the aspect ratio of the target box. The specific formula of CIoU is: Among them, ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the predicted box and the ground truth box; c represents the distance of the diagonal of the smallest closed region that can contain both the predicted box and the ground truth box. The formulas for α and v are as follows: Step 4-5: Perform multiple rounds of training on the training set to obtain the final target detection model, export the parameter file of the model and deploy it to the on-board computer for actual testing. Preprocess each frame of the picture in the UAV camera and put it into the model to obtain the prediction result.

Citation Information

Patent Citations

  • Self-supervised depth estimation method based on multi-frame attention

    CN113240722A

  • Pseudo RGB-d for self-improving monocular slam and depth prediction

    US20210065391A1