A UAV navigation method, device, computer equipment and storage medium
By constructing a simulated environment and generating reference trajectories in drone navigation, the problems of large computing volume and high latency in the prior art are solved, and a higher speed and safe drone flight is achieved.
Patent Information
- Application Number
- CN202310385515.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-04-11
AI Technical Summary
The existing end-to-end navigation algorithms for drones have large calculation volume and high latency, which cannot support drones to fly at higher speeds.
By constructing a simulation environment, a reference trajectory is generated and a trajectory probability distribution is constructed, a collision-free trajectory is obtained, and the neural network is trained to generate real-time trajectories to support the high-speed flight of the drone.
It reduces the amount of calculations during navigation, eliminates delays, increases the success rate of obstacle avoidance of drones, and supports drones to fly at higher speeds.
Smart Images

Figure CN116358543B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of autonomous navigation of unmanned aerial vehicles, and in particular to a navigation method, device, computer equipment and storage medium for unmanned aerial vehicles. Background Art
[0002] The vision-based autonomous navigation algorithms for drones are mainly divided into traditional methods and end-to-end methods. In the end-to-end method, the observations (depth images, attitude, speed, etc.) obtained by the drone's onboard sensors are directly mapped into control instructions. The flight control board receives the instructions to achieve autonomous navigation, and the entire process does not require the construction of a map. However, most end-to-end algorithms impose certain constraints on drones to achieve autonomous navigation, such as limiting the movement of drones within a plane, or discretizing their movements at the cost of losing the flexibility of the drone. Recently, some intelligent methods based on deep neural networks have not imposed constraints on drones. However, the network efficiency they use is not high enough, and there is still a possibility of collision when flying at high speeds. Therefore, the end-to-end algorithms used in existing drone intelligent navigation have large computational complexity and high latency, and cannot support drones to fly at higher speeds. Summary of the invention
[0003] The present application provides a drone navigation method, device, computer equipment and storage medium, which can reduce the amount of calculation during the navigation process, eliminate delays to a certain extent, and support the drone to fly at a higher speed.
[0004] In a first aspect, an embodiment of the present application provides a drone navigation method, which is applied to a drone navigation device, and the method includes:
[0005] Build a simulation environment and obtain the simulation environment information and drone status information of the simulation environment;
[0006] Generate a reference trajectory based on the simulated environment information and the drone state information, and construct a trajectory probability distribution based on the reference trajectory and the distance between the drone and obstacles in the simulated environment;
[0007] Based on the trajectory probability distribution, several collision-free trajectories are obtained, and the trajectory cost calculation results corresponding to each collision-free trajectory are obtained by calculating the several collision-free trajectories according to the trajectory cost function;
[0008] Sort several collision-free trajectories according to the trajectory cost calculation results, and select the three collision-free trajectories with the smallest trajectory cost calculation results as the three optimal trajectories;
[0009] The neural network is trained according to the three optimal trajectories to obtain a trained trajectory neural network;
[0010] The drone sensor observations are input into the trained trajectory neural network, and three real-time trajectories and trajectory cost values corresponding to the three real-time trajectories are output;
[0011] The input cost of each real-time trajectory is calculated based on the cost value of each trajectory, and the real-time trajectory with the lowest input cost is selected for tracking.
[0012] Furthermore, the simulation environment includes: a forest simulation environment, a geometric body simulation environment and a narrow slit simulation environment;
[0013] Building a simulation environment, including: building a simulation environment on an end-to-end simulation platform.
[0014] Furthermore, the above-mentioned generation of a reference trajectory based on the simulated environment information and the drone state information includes:
[0015] According to the simulated environment information and UAV status information, a reference trajectory is generated based on the search trajectory planning algorithm.
[0016] Furthermore, several collision-free trajectories are obtained based on the trajectory probability distribution, including:
[0017] Sampling from the trajectory probability distribution to obtain several trajectories, and detecting whether the distances between the points on the several trajectories and the obstacles in the simulation environment are greater than the radius value of the UAV model;
[0018] If the distance is greater than the radius value of the drone model, the trajectory is taken as a collision-free trajectory, and several collision-free trajectories are obtained.
[0019] Furthermore, the above sampling from the trajectory probability distribution obtains several trajectories, including:
[0020] Several trajectories are obtained by sampling from the trajectory probability distribution through the Metropolis-Hastings algorithm.
[0021] Furthermore, the trajectories in the trajectory probability distribution are all represented by a cubic B-spline curve with three control points and a uniform knot vector.
[0022] Furthermore, drone sensor observations include: depth image data, drone speed and attitude information, and desired flight direction.
[0023] Furthermore, the trained trajectory neural network includes a visual branch, a state branch and three three-layer perceptrons, the visual branch includes a lightweight neural network, and the state branch includes a four-layer perceptron.
[0024] Furthermore, the above inputs the drone sensor observations into the trained trajectory neural network, and outputs three real-time trajectories and trajectory cost values corresponding to the three real-time trajectories, including:
[0025] Inputting the depth image data into a lightweight neural network in a visual branch to obtain a first visual feature vector, a second visual feature vector, and a third visual feature vector;
[0026] The speed and attitude information of the drone and the desired flight direction are input into the four-layer perceptron in the state branch to obtain the first state feature vector, the second state feature vector and the third state feature vector;
[0027] The first visual feature vector and the first state feature vector are input into the first three-layer perceptron to obtain a first real-time trajectory and a first trajectory cost value, the second visual feature vector and the second state feature vector are input into the second three-layer perceptron to obtain a second real-time trajectory and a second trajectory cost value, and the third visual feature vector and the third state feature vector are input into the third three-layer perceptron to obtain a third real-time trajectory and a third trajectory cost value.
[0028] Furthermore, the above-mentioned calculation of the input cost of each real-time trajectory based on the cost value of each trajectory and selecting the real-time trajectory with the lowest input cost for tracking includes:
[0029] Obtain the minimum trajectory cost value among the three trajectory cost values;
[0030] Calculate the ratios of the minimum trajectory cost value to the first trajectory cost value, the second trajectory cost value, and the third trajectory cost value, calculate the input costs of the real-time trajectories whose ratios are greater than or equal to 0.95, and select the real-time trajectory with the lowest input cost for tracking.
[0031] In a second aspect, an embodiment of the present application provides a drone navigation device, which is applied to a drone navigation device, and the device includes:
[0032] The environment construction module is used to construct the simulation environment and obtain the simulation environment information and UAV status information of the simulation environment;
[0033] A probability distribution generation module is used to generate a reference trajectory based on the simulated environment information and the UAV state information, and to construct a trajectory probability distribution based on the reference trajectory and the distance between the UAV and obstacles in the simulated environment;
[0034] A sampling module is used to obtain a number of collision-free trajectories based on trajectory probability distribution, calculate a number of collision-free trajectories according to a trajectory cost function, and obtain trajectory cost calculation results corresponding to each collision-free trajectory;
[0035] An optimal trajectory generation module is used to sort a number of collision-free trajectories according to trajectory cost calculation results, and select three collision-free trajectories with the smallest trajectory cost calculation results as three optimal trajectories;
[0036] A training module is used to train the neural network according to the optimal trajectory to obtain a trained trajectory neural network;
[0037] The real-time trajectory generation module is used to input the drone sensor observations into the trained trajectory neural network and output three real-time trajectories and the trajectory cost values corresponding to the three real-time trajectories;
[0038] The trajectory tracking module is used to calculate the input cost of each real-time trajectory based on the cost value of each trajectory, and select the real-time trajectory with the lowest input cost for tracking.
[0039] Furthermore, the probability distribution generation module is used to generate a reference trajectory based on the search-based trajectory planning algorithm according to the simulated environment information and the UAV state information.
[0040] Furthermore, the sampling module is used to sample from the trajectory probability distribution to obtain a number of trajectories, and respectively detect whether the distance between the points on the several trajectories and the obstacles in the simulation environment is greater than the radius value of the UAV model; if the distance is greater than the radius value of the UAV model, the trajectory is taken as a collision-free trajectory to obtain a number of collision-free trajectories.
[0041] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of a drone navigation method as described in any of the above embodiments are performed.
[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a drone navigation method as in any of the above embodiments.
[0043] In summary, compared with the prior art, the technical solution provided in the embodiment of the present application has at least the following beneficial effects:
[0044] The embodiment of the present application provides a method for unmanned aerial vehicle navigation. By first constructing a simulation environment, a neural network is trained using the optimal trajectory generated by calculation in the simulation environment. When the trained neural network is applied to the navigation of an actual unmanned aerial vehicle, it can be calculated based on the observations of the sensors on the unmanned aerial vehicle flying in the real scene, and the real-time trajectory can be quickly output to enable the unmanned aerial vehicle to track the flight. The above method can reduce the amount of calculation in the existing end-to-end algorithm and eliminate the delay in the existing algorithm to a certain extent; at the same time, the neural network is trained with a collision-free trajectory, so that the trained trajectory neural network can improve the obstacle avoidance success rate of the unmanned aerial vehicle in actual application, thereby supporting the unmanned aerial vehicle to fly at a higher speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1A flowchart of a drone navigation method provided as an exemplary embodiment of the present application.
[0046] Figure 2 A flowchart of the steps of obtaining a collision-free trajectory is provided for an exemplary embodiment of the present application.
[0047] Figure 3 A flowchart of the steps of obtaining real-time trajectory and trajectory cost value provided for an exemplary embodiment of the present application.
[0048] Figure 4 A schematic diagram of the structure of a trajectory neural network provided for yet another exemplary embodiment of the present application.
[0049] Figure 5 A structural diagram of a drone navigation device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0051] See also Figure 1 The embodiment of the present application provides a drone navigation method, which is applied to a drone navigation device. The method may include:
[0052] Step S1, construct a simulation environment, and obtain simulation environment information and drone status information of the simulation environment.
[0053] The simulated environment information may specifically be a 3D point cloud map of the environment, and the drone status information may include the flight speed information and flight attitude information of the drone.
[0054] Specifically, the simulation environment can be built through the end-to-end simulation platform Project AirSim simulator, which is a new AirSim simulation platform for developing, training and testing drones. The Project AirSim platform is cloud-based, thanks to Microsoft Azure and Bing Maps data to create large maps containing millions of data points. The simulation platform can also extract data from a huge external map library and quickly generate realistic models based on 3D models and map data.
[0055] Step S2, generating a reference trajectory according to the simulated environment information and the drone state information, and constructing a trajectory probability distribution according to the reference trajectory and the distance between the drone and obstacles in the simulated environment.
[0056] Among them, the distance between the drone and obstacles in the simulated environment is also obtained based on the simulated environment information and the drone status information.
[0057] Specifically, the trajectory probability distribution P(τ|τ ref ,) depends on the reference trajectory τ ref and the surrounding 3D point cloud C∈R n×3 If the UAV can stay away from obstacles in the simulated environment and approach the reference trajectory τ when flying according to the trajectory τ ref , then the probability of trajectory τ in the trajectory probability distribution is larger. The trajectory probability distribution P is defined as:
[0058]
[0059] where Z = ∫ τ P(τ|τ ref ,C) is the normalization factor, c(τ,τ ref ,C)∈R + is the trajectory cost function, which represents the difference between the actual trajectory of the UAV flight and the reference trajectory τ ref The proximity of the drone and the distance between the drone and obstacles in the simulated environment.
[0060] Step S3, obtaining a plurality of collision-free trajectories based on the trajectory probability distribution, calculating the plurality of collision-free trajectories according to the trajectory cost function, and obtaining trajectory cost calculation results corresponding to each collision-free trajectory.
[0061] Among them, the trajectory cost function is defined as:
[0062]
[0063] where λ c =1000, Q is the semi-positive definite state cost matrix, C collision is the collision cost, which is a measure of the distance from the drone to the point cloud C, i.e., the point in the simulated environment. The drone is modeled as a q = 0.2m sphere, and define the collision cost as d c The truncated quadratic function, d c That is, the distance between the drone and the nearest point in the 3D point cloud of the environment:
[0064]
[0065] Specifically, not all trajectories sampled from the trajectory probability distribution are collision-free trajectories at the beginning, after all, they are randomly sampled. The method of removing trajectories that may collide from the sampled trajectories and obtaining several collision-free trajectories is as follows: detect whether the distance between the points on the trajectory obtained after sampling and the obstacles in the simulation environment is less than the radius of the drone model. If d c <r q , then it is judged that the trajectory will collide with the obstacle, so it is deleted from the sampled trajectory, thus obtaining several collision-free trajectories.
[0066] In the specific implementation process, because the trajectory probability distribution P is very complex, it is often multimodal in the presence of obstacles and in highly dynamic environments. Since obstacles can be avoided through multiple feasible trajectories. Therefore, the analytical calculation of P is usually very difficult. The trajectory probability distribution P is approximated by random sampling. The sampled trajectories will asymptotically cover all different modes of P. In order to estimate P, the sampling algorithm requires a target score function s(τ)∝P(τ|τ ref ,C). Define s(τ)=exp(-c(τ,τ ref ,C)), where c(τ,τ ref ,C) is the trajectory cost calculation result of trajectory τ. The number of trajectories sampled from the trajectory probability distribution can be 50,000, and then detection calculation is performed from these 50,000 sampled trajectories to obtain several collision-free trajectories.
[0067] Step S4, sorting a number of collision-free trajectories according to the trajectory cost calculation results, and selecting three collision-free trajectories with the smallest trajectory cost calculation results as three optimal trajectories.
[0068] Specifically, after sorting the trajectory cost calculation results of several collision-free trajectories from large to small, the last three collision-free trajectories are selected as the three optimal trajectories.
[0069] Step S5, training the neural network according to the three optimal trajectories to obtain a trained trajectory neural network.
[0070] Among them, the neural network includes a visual branch, a state branch and three 3-layer perceptrons.
[0071] Specifically, when training the neural network, the collision cost in the optimal trajectory will also be used for the training of the neural network, so that the trained trajectory neural network can not only generate real-time trajectories, but also calculate the trajectory cost value corresponding to each real-time trajectory.
[0072] Step S6, inputting the drone sensor observations into the trained trajectory neural network, and outputting three real-time trajectories and trajectory cost values corresponding to the three real-time trajectories.
[0073] Specifically, the UAV sensor observations are received as input to the trained trajectory neural network. Based on this input, it predicts three real-time trajectories T n (n=1,2,3) and its corresponding trajectory cost c k ,Right now where c k ∈R + , M is the number of real-time trajectories, here it is 3. Here the real-time trajectory T predicted by the trajectory neural network n It is not represented by the state evolution of the UAV, but by the position of the UAV, i.e. Specifically, the real-time trajectory generated by the trained trajectory neural network Described as:
[0074]
[0075] Where p(t i )∈R 3 is the UAV at time t = t i The position relative to the current state x.
[0076] To account for multiple hypothesis predictions, the Relaxed-Winner-Takes-All (R-WTA) loss is minimized for each sample:
[0077]
[0078] Among them, T e is the set of three optimal trajectories used to train the neural network in step S5, and remains in the trained trajectory neural network after the training is completed; T n is the set of three real-time trajectories predicted by the trajectory neural network based on the input UAV sensor observations, τ e,p Represents τ e The position component of Defined as:
[0079]
[0080] Among them, the weight 1-ε=0.95, ε / (M-1)=0.025. In addition, the trajectory neural network can predict the trajectory cost c of the real-time trajectory k The collision cost C in step S3 is used collision Calculated true value of collision cost For training, the final training loss function for each sample is:
[0081]
[0082] where λ 1=10,λ 2 = 0.1, the training batch size is 8, and the Adam optimizer is used with a learning rate of 1×10 -3 .
[0083] Step S7, calculating the input cost of each real-time trajectory based on the cost value of each trajectory, and selecting the real-time trajectory with the lowest input cost for tracking.
[0084] Specifically, the components of the three real-time trajectories on each coordinate axis are mapped to the polynomial space to obtain three mapped real-time trajectories, and c is selected. * / k ≥0.95(c * =inc k )’s real-time trajectory corresponding to the mapped real-time trajectory calculates its input cost.
[0085] The polynomial space may be a fifth-order polynomial space.
[0086] Specifically, the three real-time trajectories predicted by the trajectory neural network are mapped to the 5th-order polynomial space on each coordinate axis. For example, in the x direction, the polynomial projection is defined as in T(t) T =[1,t,…,t 5 ]. Projection term a x By solving the following optimization problem:
[0087]
[0088]
[0089]
[0090]
[0091] in yes The x-component of the ith element of s x (0), is the x-position of the drone and its derivatives (i.e., velocity and acceleration) obtained from the current state estimate, which in Project AirSim corresponds to the real state provided by the simulator and on the physical platform corresponds to the calculated state estimate. In order to reduce sudden changes in velocity and strong pitching motions during flight, the average velocity of the polynomial is therefore limited to the desired value v des To do this, we can transform the polynomial μ x (t) is scaled by a factor of time t That is, t′=t, where
[0092] Once all three predicted real-time trajectories are mapped, the best one is selected for tracking. To this end, the input cost of the mapped real-time trajectory corresponding to the real-time trajectory is calculated based on the trajectory cost value of the real-time trajectory, and the drone selects the real-time trajectory corresponding to the mapped real-time trajectory with the lowest input cost to track. From then on, intelligent navigation is completed.
[0093] The above method for calculating the input cost may specifically be:
[0094]
[0095] Among them, μ r and μ ψ is a constant that makes the integrand dimensionless, and r is the three-dimensional coordinate, that is, r = [x, y, z] T , ψ is the yaw angle of the UAV. σ T =[x T ,y T ,z T ,ψ T ,] T , σ i =[x i ,y i ,z i ,ψ i ,] T , k r =4,k ψ =2.
[0096] A method for unmanned aerial vehicle navigation provided in the above embodiment first constructs a simulation environment, and uses the optimal trajectory calculated and generated in the simulation environment to train the neural network. When the trained neural network is applied to the navigation of the unmanned aerial vehicle in actual flight, it can be calculated based on the observations of the sensors on the unmanned aerial vehicle flying in the real scene, and the real-time trajectory can be quickly output to enable the unmanned aerial vehicle to track the flight. The above method can reduce the amount of calculation in the existing end-to-end algorithm and eliminate the delay existing in the existing algorithm to a certain extent; at the same time, the neural network is trained with a collision-free trajectory, so that the trained trajectory neural network can improve the obstacle avoidance success rate of the unmanned aerial vehicle in actual application, thereby supporting the unmanned aerial vehicle to fly at a higher speed.
[0097] In some embodiments, the simulation environment includes: a forest simulation environment, a geometric body simulation environment, and a narrow gap simulation environment;
[0098] Step S1 may specifically further include: constructing a simulation environment on an end-to-end simulation platform.
[0099] The end-to-end simulation platform may be a Project AirSim simulator, the geometric body simulation environment may specifically be a cuboid or a cylinder, and the narrow slit in the narrow slit simulation environment may be a narrow slit perpendicular to the ground.
[0100] The above implementation method uses a simulation platform to simulate the three most common drone flight environments in reality, so that the optimal trajectory obtained by training is more suitable for the real environment, thereby making the obstacle avoidance success rate of the drone real-time trajectory output by the trajectory neural network trained by the optimal trajectory higher.
[0101] In some embodiments, generating a reference trajectory according to the simulated environment information and the drone status information in step S2 may include: generating a reference trajectory according to the simulated environment information and the drone status information using a search-based trajectory planning algorithm.
[0102] Among them, search-based trajectory planning algorithms include blind algorithms and heuristic search algorithms; Dijkstra's algorithm for solving directed positive weighted graph planning, Ford's algorithm that can calculate negative weights, and breadth-first algorithms that systematically check all nodes in the graph are all blind algorithms. The characteristics of this type of algorithm are that it can find the optimal path or the shortest path, although the search process is too blind, which means it takes a long time and the search range is wide. Heuristic search algorithms use heuristic information for guided search and heuristic functions for navigation, effectively avoiding the drawbacks of blind search in traditional methods. Typical heuristic search algorithms include the best-first algorithm and the greedy algorithm, but most heuristic path planning algorithms can only find approximate solutions to problems rather than optimal solutions.
[0103] Therefore, the LPA (Lifelong Planning A) algorithm, which combines the advantages of heuristic search and Dijkstra algorithm, is more commonly used. This algorithm is usually used to deal with the shortest path problem from a given starting point to a given target point in a dynamic environment. However, during the walking process, the nodes in the environment change. If you want to obtain a new path, you have to use the current node of the drone as the starting point and re-solve it with the LPA algorithm, which reduces the algorithm's computational efficiency and increases its time complexity.
[0104] The above-mentioned embodiment can make the calculated reference trajectory a global collision-free optimal trajectory from the starting point to the target point, so that most of the trajectories in the trajectory probability distribution established according to the reference trajectory can be biased towards the obstacle-free area. When the optimal trajectory used for neural network training is biased towards the obstacle-free area, the probability of collision of the real-time trajectory generated by the trained neural network is reduced.
[0105] In some embodiments, see Figure 2 The step S3 of obtaining a plurality of collision-free trajectories based on the trajectory probability distribution may specifically include the following steps:
[0106] Step S31, sampling from the trajectory probability distribution to obtain a number of trajectories.
[0107] Step S32, respectively detecting whether the distances between points on a plurality of trajectories and obstacles in the simulation environment are greater than the radius value of the drone model.
[0108] Step S33: if the distance is greater than the radius value of the drone model, the trajectory is taken as a collision-free trajectory to obtain a plurality of collision-free trajectories.
[0109] In the specific implementation process, the UAV is modeled as a q = 0.2m sphere, d c That is, the distance between the drone and the nearest point in the 3D point cloud of the environment. If d c <r q , it is judged that the trajectory will collide with the obstacle in the simulation environment, so it is deleted from the sampled trajectory to obtain several collision-free trajectories.
[0110] The above embodiment detects the sampled trajectory samples and eliminates the trajectories that may collide due to being too close to obstacles, thereby obtaining a number of collision-free trajectories, ensuring that the optimal trajectory for neural network training is a trajectory that does not collide, thereby improving the obstacle avoidance success rate of the real-time trajectory generated by the trajectory neural network.
[0111] In some embodiments, step S31 may further specifically include: sampling from the trajectory probability distribution by using the Metropolis-Hastings algorithm to obtain a plurality of trajectories.
[0112] Among them, the Metropolis–Hastings algorithm (MH algorithm for short) is a Markov Monte Carlo (MCMC) method in statistics and statistical physics, which is used to extract a sequence of random samples from a probability distribution when direct sampling is difficult. The obtained sequence can be used to estimate the probability distribution or calculate integrals (such as expected values). Metropolis–Hastings or other MCMC algorithms are generally used to sample from multivariate (especially high-dimensional) distributions. For univariate distributions, other methods such as adaptive rejection sampling that can extract independent samples are often used without the problem of sample autocorrelation in MCMC.
[0113] Specifically, considering the amount of computation, the present application can discretize the trajectory probability distribution at a time interval of 0.1s. Specifically, the present application uses Gaussian distribution with variances of 2, 5, and 10 as the recommended distribution of the MH algorithm to sample a total of 50,000 trajectories for the trajectory probability distribution P, that is, approximately 16,000 sampled trajectory samples are obtained under each Gaussian distribution with each variance.
[0114] The above embodiment solves the problem that the analytical calculation of the trajectory probability distribution P is very difficult and there is no way to obtain its analytical solution by using the MH algorithm to sample in the trajectory probability distribution P. The trajectory samples after random sampling can infinitely approach the trajectory probability distribution P, so that the optimal trajectory is more biased towards the obstacle-free area biased by the trajectory probability distribution P, thereby making the real-time trajectory obstacle avoidance success rate generated by the trajectory neural network trained by the optimal trajectory higher.
[0115] In some embodiments, the trajectories in the trajectory probability distribution are each represented by a cubic B-spline curve having three control points and a uniform knot vector.
[0116] Among them, the B-spline curve is a generalization of the Bezier curve (also known as the Bezier curve), which can be further extended to the non-uniform rational B-spline (NURBS), so that accurate models can be built for more general geometric bodies.
[0117] The surface of B-spline curve has many excellent properties such as geometric invariance, convex hull, convexity preservation, variation reduction, local support, etc. It is a commonly used geometric representation method in CAD system. Therefore, parameterization based on measurement data and B-spline surface reconstruction are one of the research hotspots and key technologies in reverse engineering. The definition of B-spline curve is: given n+1 control points P i (i=0,1,2,...,n), then the expression of k-order B-spline curve can be defined as Among them, N i,k () is the k-th B-spline basis function, also called the harmonic function, or the k-th canonical B-spline basis function.
[0118] In the above embodiment, the trajectory in the trajectory probability distribution is represented by a cubic B-spline curve with three control points and a uniform node vector, which greatly reduces the dimension of the sampling space, thereby making the sampling process more computationally efficient.
[0119] In some embodiments, drone sensor observations may specifically include: depth image data, drone speed and attitude information, and desired flight direction.
[0120] The desired flight direction is represented by a normalized vector of a reference point after 1 second, and the reference point after 1 second specifically refers to the reference point where the current state of the drone will fly towards after 1 second in the future.
[0121] Specifically, the drone flying in the real environment is different from the drone simulating the flight state in the simulated environment. When the drone flies in the real environment, the sensor on the drone collects the surrounding environment image data in real time and converts the depth image data d∈R 640×480 , UAV speed v∈R 3 , UAV attitude q∈R 9 and the desired flight direction w∈R 3 As the input of the trained trajectory neural network, three real-time trajectories are output through the calculation of the trajectory neural network.
[0122] The above embodiment uses the information collected by the drone sensor when the drone is flying in a real environment as the input of the neural network, so that the generated real-time trajectory completely matches the flight status of the drone at the current moment. At the same time, the real-time trajectory includes the flight reference direction of the drone within the next 1 second, so that the drone can fly at high speed with low latency while successfully avoiding obstacles.
[0123] In some embodiments, the trained trajectory neural network includes a visual branch, a state branch and three three-layer perceptrons, the visual branch includes a lightweight neural network, and the state branch includes a four-layer perceptron.
[0124] Among them, the lightweight neural network can be a Mobile-Former network. The Mobile-Former network is the latest and most efficient lightweight network. It connects the MobileNet and Transformer structures in parallel through a double-wire bridge. This method combines the advantages of MobileNet's local expression ability and Transformer's global expression ability. This bridge can bidirectionally integrate locality and globality. Compared with other lightweight networks, Mobile-Former has smaller computational complexity and higher efficiency. A Mobile-Former network consists of multiple Mobile-Former modules. After pre-training on the ImageNet dataset, the last two fully connected layers are removed and it can be used for drone data processing applications.
[0125] The above embodiment adopts a lightweight neural network in the visual branch, so that the trajectory neural network has a smaller calculation amount for the real-time trajectory and lower latency, ensuring good real-time performance, so that the drone can avoid obstacles in time when flying at high speed.
[0126] In some embodiments, see Figure 3 and Figure 4 In step S6, the drone sensor observations are input into the trained trajectory neural network, and three real-time trajectories and trajectory cost values corresponding to the three real-time trajectories are output, which may specifically include the following steps:
[0127] Step S61, input the depth image data into the lightweight neural network in the visual branch to obtain a first visual feature vector, a second visual feature vector and a third visual feature vector.
[0128] Step S62, input the drone speed and attitude information and the desired flight direction into the four-layer perceptron in the state branch to obtain a first state feature vector, a second state feature vector and a third state feature vector.
[0129] Step S63, input the first visual feature vector and the first state feature vector into the first three-layer perceptron to obtain the first real-time trajectory and the first trajectory cost value, input the second visual feature vector and the second state feature vector into the second three-layer perceptron to obtain the second real-time trajectory and the second trajectory cost value, input the third visual feature vector and the third state feature vector into the third three-layer perceptron to obtain the third real-time trajectory and the third trajectory cost value.
[0130] Specifically, the visual branch uses a pre-trained Mobile-Former lightweight neural network to efficiently extract features from deep image data, which are then processed by one-dimensional convolution to generate 3 visual feature vectors of size 32.
[0131] The state branch connects the current drone speed and attitude information with the desired flight direction and is processed by a four-layer perceptron with [64, 32, 32, 32] neurons, using LeakyReLU as the activation function and one-dimensional convolution to obtain three 32-dimensional state feature vectors.
[0132] The three visual feature vectors and the three state feature vectors are processed by three three-layer perceptrons with [64, 128, 128] neurons and LeakyReLU activation function respectively. The three-layer perceptron generates a real-time trajectory and a trajectory cost value corresponding to the real-time trajectory based on the visual feature vector and the state feature vector.
[0133] The above embodiment processes the drone sensor observations input into the neural network through the visual branch, state branch and three three-layer perceptrons in the neural network to obtain three real-time trajectories and corresponding trajectory cost values, with fast calculation speed and low latency, thereby supporting the drone to fly at a higher speed.
[0134] In some embodiments, the step S7 of calculating the input cost of each real-time trajectory based on the cost value of each trajectory and selecting the real-time trajectory with the lowest input cost for tracking includes:
[0135] Obtain the minimum trajectory cost value among the three trajectory cost values;
[0136] Calculate the ratios of the minimum trajectory cost value to the first trajectory cost value, the second trajectory cost value, and the third trajectory cost value, calculate the input costs of the real-time trajectories whose ratios are greater than or equal to 0.95, and select the real-time trajectory with the lowest input cost for tracking.
[0137] Specifically, the components of the three real-time trajectories on each coordinate axis are mapped to the polynomial space to obtain three mapped real-time trajectories, and c is selected. * / k ≥0.95(c * =inc k )’s real-time trajectory corresponding to the mapped real-time trajectory calculates its input cost.
[0138] The polynomial space may be a 5th-order polynomial space. For more detailed calculation steps, please refer to step S7.
[0139] In the specific implementation process, the number of trajectories for calculating the input cost is likely to be less than 3. When the ratio of the minimum trajectory cost value to the trajectory cost value of the real-time trajectory among the three real-time trajectories is less than 0.95, the input cost of the mapped real-time trajectory corresponding to the real-time trajectory will no longer be calculated.
[0140] The above embodiment screens the real-time trajectory based on the trajectory cost value output by the trained trajectory neural network, ensuring that the probability of collision of the real-time trajectory is extremely small, thereby greatly improving the obstacle avoidance success rate of the drone.
[0141] See also Figure 5 Another embodiment of the present application provides a drone navigation device, which is applied to a drone navigation device, and includes:
[0142] The environment construction module 101 is used to construct a simulation environment and obtain simulation environment information and drone status information of the simulation environment;
[0143] The probability distribution generation module 102 is used to generate a reference trajectory according to the simulated environment information and the drone state information, and to construct a trajectory probability distribution according to the reference trajectory and the distance between the drone and the obstacle in the simulated environment;
[0144] A sampling module 103 is used to obtain a plurality of collision-free trajectories based on the trajectory probability distribution, and calculate the plurality of collision-free trajectories according to the trajectory cost function to obtain a trajectory cost calculation result corresponding to each collision-free trajectory;
[0145] The optimal trajectory generation module 104 is used to sort a plurality of collision-free trajectories according to the trajectory cost calculation results, and select three collision-free trajectories with the smallest trajectory cost calculation results as three optimal trajectories;
[0146] A training module 105 is used to train the neural network according to the optimal trajectory to obtain a trained trajectory neural network;
[0147] A real-time trajectory generation module 106 is used to input the drone sensor observations into the trained trajectory neural network, and output three real-time trajectories and trajectory cost values corresponding to the three real-time trajectories;
[0148] The trajectory tracking module 107 is used to calculate the input cost of each real-time trajectory based on the cost value of each trajectory, and select the real-time trajectory with the lowest input cost for tracking.
[0149] The above embodiment provides a drone navigation device, which constructs a simulation environment through the environment construction module 101, and uses the optimal trajectory generated in the optimal trajectory generation module 104 to train the neural network. When the trained neural network is applied to the actual flight drone navigation, the real-time trajectory generation module 106 will calculate according to the observed quantity of the sensor on the drone flying in the real scene, and quickly output the real-time trajectory to enable the trajectory tracking module 107 to track the flight. The above method can reduce the amount of calculation in the existing end-to-end algorithm, and eliminate the delay existing in the existing algorithm to a certain extent; at the same time, the neural network is trained with a collision-free trajectory, so that the trained trajectory neural network can improve the obstacle avoidance success rate of the drone in actual application, thereby supporting the drone to fly at a higher speed.
[0150] In some embodiments, the probability distribution generation module 102 is used to generate a reference trajectory based on a search-based trajectory planning algorithm according to the simulated environment information and the drone state information.
[0151] The above-mentioned embodiment can make the calculated reference trajectory a global collision-free optimal trajectory from the starting point to the target point, so that most of the trajectories in the trajectory probability distribution established according to the reference trajectory can be biased towards the obstacle-free area. When the optimal trajectory used for neural network training is biased towards the obstacle-free area, the probability of collision of the real-time trajectory generated by the trained neural network is reduced.
[0152] In some embodiments, the sampling module 103 is used to sample from the trajectory probability distribution to obtain a number of trajectories, and respectively detect whether the distance between the points on the several trajectories and the obstacles in the simulated environment is greater than the radius value of the drone model; if the distance is greater than the radius value of the drone model, the trajectory is taken as a collision-free trajectory to obtain a number of collision-free trajectories.
[0153] The above embodiment detects the sampled trajectory samples through the sampling module 103, eliminates the trajectories that may collide due to being too close to the obstacle, thereby obtaining a number of collision-free trajectories, ensuring that the optimal trajectory for neural network training is a trajectory that will not collide, thereby improving the obstacle avoidance success rate of the real-time trajectory generated by the trajectory neural network.
[0154] The specific limitations of a drone navigation device provided in this embodiment can be found in the above embodiment of a drone navigation device method, which will not be repeated here. Each module in the above drone navigation device can be implemented in whole or in part by software, hardware, and a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0155] An embodiment of the present application provides a computer device, which may include a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the processor executes the steps of a method for a drone navigation device as in any of the above embodiments.
[0156] The working process, working details and technical effects of the computer equipment provided in this embodiment can be found in the above embodiment of a method for a drone navigation device, and will not be described in detail here.
[0157] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of a method for a drone navigation device as in any of the above embodiments are implemented. The computer-readable storage medium refers to a carrier for storing data, which may include but is not limited to a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash drive and / or a memory stick, etc., and the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0158] The working process, working details and technical effects of the computer-readable storage medium provided in this embodiment can be found in the above embodiment of a method for a drone navigation device, and will not be described in detail here.
[0159] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0160] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0161] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A drone navigation method, It is characterized in that The method comprises: Constructing a simulation environment, and obtaining simulation environment information and drone status information of the simulation environment; Generate a reference trajectory according to the simulated environment information and the drone state information, and construct a trajectory probability distribution according to the reference trajectory and the distance between the drone and obstacles in the simulated environment; Based on the trajectory probability distribution, a plurality of collision-free trajectories are obtained, and the plurality of collision-free trajectories are calculated according to a trajectory cost function to obtain trajectory cost calculation results corresponding to each of the collision-free trajectories; Sorting a plurality of the collision-free trajectories according to the trajectory cost calculation result, and selecting three collision-free trajectories with the smallest trajectory cost calculation results as three optimal trajectories; The neural network is trained according to the three optimal trajectories to obtain a trained trajectory neural network; Inputting the drone sensor observations into the trained trajectory neural network, and outputting three real-time trajectories and trajectory cost values corresponding to the three real-time trajectories; The input cost of each of the real-time trajectories is calculated based on the cost value of each of the trajectories, and the real-time trajectory with the lowest input cost is selected for tracking.
2. The method according to claim 1, It is characterized in that The simulation environment includes: a forest simulation environment, a geometric body simulation environment and a narrow gap simulation environment; The constructing of the simulation environment includes: constructing the simulation environment on an end-to-end simulation platform.
3. The method according to claim 1, It is characterized in that The generating a reference trajectory according to the simulated environment information and the drone state information includes: The reference trajectory is generated based on a search trajectory planning algorithm according to the simulated environment information and the drone state information.
4. The method according to claim 1, It is characterized in that The obtaining of a plurality of collision-free trajectories based on the trajectory probability distribution includes: Sampling from the trajectory probability distribution to obtain a plurality of trajectories, and respectively detecting whether the distances between the points on the plurality of trajectories and obstacles in the simulation environment are greater than the radius value of the drone model; If the distance is greater than the radius value of the UAV model, the trajectory is used as the collision-free trajectory to obtain a plurality of collision-free trajectories.
5. The method according to claim 4, It is characterized in that The sampling from the trajectory probability distribution to obtain a plurality of trajectories includes: The plurality of trajectories are obtained by sampling from the trajectory probability distribution through the Metropolis-Hastings algorithm.
6. The method according to claim 1, It is characterized in that The trajectories in the trajectory probability distribution are all represented by a cubic B-spline curve with three control points and a uniform node vector.
7. The method according to claim 1, It is characterized in that The drone sensor observations include: depth image data, drone speed and attitude information, and desired flight direction.
8. The method according to claim 7, It is characterized in that The trajectory neural network includes a visual branch, a state branch and three three-layer perceptrons, the visual branch includes a lightweight neural network, and the state branch includes a four-layer perceptron.
9. The method according to claim 8, It is characterized in that The method of inputting the drone sensor observations into the trained trajectory neural network and outputting three real-time trajectories and trajectory cost values corresponding to the three real-time trajectories includes: Inputting the depth image data into the lightweight neural network in the visual branch to obtain a first visual feature vector, a second visual feature vector, and a third visual feature vector; Inputting the speed and attitude information of the drone and the desired flight direction into the four-layer perceptron in the state branch to obtain a first state feature vector, a second state feature vector and a third state feature vector; The first visual feature vector and the first state feature vector are input into a first three-layer perceptron to obtain a first real-time trajectory and a first trajectory cost value, the second visual feature vector and the second state feature vector are input into a second three-layer perceptron to obtain a second real-time trajectory and a second trajectory cost value, and the third visual feature vector and the third state feature vector are input into a third three-layer perceptron to obtain a third real-time trajectory and a third trajectory cost value.
10. The method according to claim 9, It is characterized in that The calculating the input cost of each of the real-time trajectories based on the cost value of each of the trajectories, and selecting the real-time trajectory with the lowest input cost for tracking, includes: Obtain a minimum trajectory cost value among the three trajectory cost values; calculate the ratios of the minimum trajectory cost value to the first trajectory cost value, the second trajectory cost value, and the third trajectory cost value, respectively, calculate the input cost of the real-time trajectory for which the ratio is greater than or equal to 0.95, and select the real-time trajectory with the lowest input cost for tracking.
11. A drone navigation device, It is characterized in that Applied to unmanned aerial vehicle navigation equipment, the device comprises: An environment construction module is used to construct a simulation environment and obtain simulation environment information and drone status information of the simulation environment; A probability distribution generation module, used to generate a reference trajectory according to the simulated environment information and the drone state information, and to construct a trajectory probability distribution according to the reference trajectory and the distance between the drone and obstacles in the simulated environment; A sampling module, used for obtaining a plurality of collision-free trajectories based on the trajectory probability distribution, calculating the plurality of collision-free trajectories according to a trajectory cost function, and obtaining a trajectory cost calculation result corresponding to each of the collision-free trajectories; An optimal trajectory generation module, used for sorting a plurality of the collision-free trajectories according to the trajectory cost calculation result, and selecting three collision-free trajectories with the smallest trajectory cost calculation results as three optimal trajectories; A training module, used to train the neural network according to the optimal trajectory to obtain a trained trajectory neural network; A real-time trajectory generation module is used to input the drone sensor observations into the trained trajectory neural network, and output three real-time trajectories and trajectory cost values corresponding to the three real-time trajectories; The trajectory tracking module is used to calculate the input cost of each of the real-time trajectories based on the trajectory cost value, and select the real-time trajectory with the lowest input cost for tracking.
12. The device according to claim 11, It is characterized in that The probability distribution generation module is used to generate the reference trajectory based on the search trajectory planning algorithm according to the simulation environment information and the drone state information.
13. The device according to claim 11, It is characterized in that The sampling module is used to sample from the trajectory probability distribution to obtain a plurality of trajectories, and respectively detect whether the distance between the points on the plurality of trajectories and the obstacles in the simulation environment is greater than the radius value of the drone model; if the distance is greater than the radius value of the drone model, the trajectory is used as the collision-free trajectory to obtain a plurality of collision-free trajectories.
14. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.
15. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Training method and device of vehicle track evaluation network model and storage medium
CN113239986A
Model-enhanced unmanned aerial vehicle flight path reinforcement learning optimization method
CN114879738A