Unmanned Aerial Vehicle Route Planning Method and System Based on Deep Reinforcement Learning

By applying deep reinforcement learning technology on drones, the problems of low efficiency and insolable strategies of drones in complex three-dimensional environments are solved, and more efficient independent exploration and path planning are achieved.

CN118896610BActive Publication Date: 2025-06-17YUNNAN MINZU UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410940842.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-06-17
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

When exploring in complex three-dimensional unstructured environments, existing drones lack the ability to explore previous experiences or strategies, have high computational complexity, are prone to dimensional disasters, and their independent exploration strategies are not robust, making it difficult to cope with complex three-dimensional environments.

Method used

The method based on deep reinforcement learning is adopted to separate the time-consuming training process from the real-time decision-making process. Through end-to-end learning, high-level task representation and control strategies are learned from sensor data, a route planning model based on deep learning is built, and the PPO algorithm is used to optimize the model.

Benefits of technology

It improves the efficiency of autonomous exploration of drones in complex three-dimensional environments and the robustness of strategies, reduces the computational complexity, avoids dimensional disasters, and achieves more efficient exploration and path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118896610B_ABST
    Figure CN118896610B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for unmanned aerial vehicle (UAV) route planning based on deep reinforcement learning. The method includes: constructing a three-dimensional unstructured map using gradient-based Berlin noise and a digital elevation model; constructing a partially observable Markov decision process model with constraints for autonomous exploration of a fixed-wing UAV to constrain the flight route of the fixed-wing UAV; constructing a route planning model based on deep learning according to the partially observable Markov decision process model, and optimizing the route planning model using the Proximal Policy Optimization (PPO) algorithm. The present invention uses a deep neural network to fit components such as the action value function, policy, and model of reinforcement learning, constructs an approximate mapping of a deep neural network from local observations to value functions and policy functions, establishes a partially observable Markov decision process model, completes the construction of the reinforcement learning framework, and improves the robustness and generalization ability of the model when dealing with large-scale state spaces.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep reinforcement learning, and particularly relates to a method and system for unmanned aerial vehicle route planning based on deep reinforcement learning. Background Art

[0002] When performing tasks such as natural resource exploration, disaster monitoring and assessment, and border security patrol and monitoring, unmanned aerial vehicles (UAVs) usually aim at spatial exploration, rather than the ordinary sense of trajectory planning from a clear starting point to an ending point. When performing such autonomous exploration tasks, UAVs generally first establish an environmental map through pure exploration, and then use the generated map for navigation and path planning. Current exploration methods usually use artificial maps with prior experience or heuristic methods, such as frontier-based exploration, which require clear environmental models and rules. For complex and unknown environments, the performance decreases and it is not applicable to large-scale state spaces, and generalization is difficult. Other methods using autonomous learning often only focus on learning strategies for specific tasks, use sample-inefficient random exploration or make unrealistic assumptions about the availability of global maps, and there is an urgent need for more efficient methods to improve the exploration effect.

[0003] Traditional boundary-driven autonomous exploration algorithms do not interact adaptively with the current task environment, so they lack the ability to explore experiences or strategies that did not exist previously. During the policy execution process, some solution links rely more on direct search proportional to the spatial dimension and the number of spatial discretizations, with high computational complexity and prone to the curse of dimensionality. For example, when calculating the gain corresponding to the next step of a boundary point, all discrete boundary points need to be traversed, and the computational amount for three-dimensional problems is often unacceptable, consuming a large amount of computational resources, thus resulting in a reduction in the overall performance of the UAV. In addition, some studies limit the exploration scenario to a two-dimensional slice grid environment in order to simplify the model, perform binary processing on the plane where the UAV is located, set the flyable area to 0, and the non-flyable area to 1. Although this processing reduces the computational amount, it causes the loss of elevation information and weakens the influence of the three-dimensional space on the learning strategy. Eventually, the obtained autonomous exploration strategy is not robust enough to cope with complex three-dimensional environments.

[0004] Current exploration methods usually use artificial maps with prior experience or heuristic methods, such as frontier-based exploration, which not only requires explicit environmental models and rules, but also does not interact adaptively with the current task environment. Therefore, it lacks the ability to explore previously non-existent experiences or strategies. During the strategy execution process, some solution steps rely more on direct search proportional to the spatial dimension and the number of spatial discretizations, with high computational complexity and prone to the curse of dimensionality. For example, when calculating the gain corresponding to the next step of a boundary point, all discrete boundary points need to be traversed, and the computational amount for three-dimensional problems is often unacceptable, consuming a large amount of computing resources of the drone's on-board computer, thus reducing the overall performance of the drone. For complex and unknown environments, the performance of the algorithm degrades and it is not applicable to large-scale state spaces, making generalization difficult.

[0005] Other methods using autonomous learning often only focus on learning strategies for specific tasks, using sample-inefficient random exploration or making unrealistic assumptions about the availability of the global map. Some studies also limit the exploration scenario to a two-dimensional slice grid environment in order to simplify the model. The plane where the drone is located is binarized, with the flyable area set to 0 and the non-flyable area set to 1. Although this reduces the computational amount, it causes the loss of elevation information and weakens the influence of the three-dimensional space on the learning strategy. Eventually, the obtained autonomous exploration strategy is not robust enough to handle complex three-dimensional environments. Summary of the Invention

[0006] The object of the present invention is to propose a method and system for UAV route planning based on deep reinforcement learning. Aiming at the shortcomings of the prior art in the face of complex three-dimensional unstructured environments, such as the need for prior experience, poor interactivity, large computational burden, poor generalization, low exploration efficiency, and one-sided strategies, the present invention adopts a method based on deep reinforcement learning, separating the time-consuming training process from the real-time decision-making process. After training a model with excellent performance, it is then deployed to the UAV, which can not only ensure the response speed of the UAV, but also provide a better way to solve the curse of dimensionality. Deep reinforcement learning allows the UAV to directly learn high-level task representations and control strategies from sensor data through an end-to-end learning method, and can automatically learn high-level feature representations without manually designing features or controllers, which helps the agent better adapt to complex and changing environments. Through experience replay and parameter sharing of deep neural networks, the deep reinforcement learning method can better utilize limited data during training and improve learning efficiency.

[0007] To achieve the above object, in the first aspect of the present invention, there is provided a method for UAV route planning based on deep reinforcement learning, the method comprising:

[0008] S1. Construct a three-dimensional unstructured map using gradient-based Berlin noise and a digital elevation map;

[0009] S2. Construct a partially observable Markov decision process model with constraints for the autonomous exploration of fixed-wing UAVs to constrain the flight route of the fixed-wing UAV;

[0010] S3. Construct a route planning model based on deep learning according to the partially observable Markov decision process model, and optimize the route planning model using the PPO algorithm.

[0011] Further, the gradient-based Perlin noise generates a lattice rectangular grid covering the entire map, and a gradient vector is randomly initialized at each lattice point, specifically as follows:

[0012] Assume that a point P to be calculated within the lattice, and the four lattice points to which the calculation point P belongs are denoted as p0, p1, p2, p3 respectively, then its gradient vector is <grad0, grad1, grad2, grad3>;

[0013] Calculate the offsets <delta0, delta1, delta2, delta3> of point P from the four lattice points respectively, and then sum the dot products of the gradient vectors and the distance vectors at each lattice point to obtain the random noise value of point P. The Perlin calculation formula is expressed as follows:

[0014]

[0015] Among them, f(·) represents a high-order curve function used to enhance the smoothness of Perlin noise, and the calculation is as follows:

[0016] f(u) = 6u 5 -15u 4 +10u 3

[0017] Among them, u represents the independent variable of the high-order curve function;

[0018] Generate noise using a single frequency, then calculate by superimposing noises of multiple different frequencies, and at the same time use persistence to represent the amplitude of each frequency component. The calculation formula is expressed as follows:

[0019] frequency = 2 k

[0020] amplitude = persistence k

[0021]

[0022] Among them, "frequency" represents frequency, "amplitude" represents amplitude, "Perlin" represents the Perlin noise operation performed on the frequency and amplitude with different values for each point, that is, the k-th noise function to be superimposed, "N" represents the octave number, "persistence" represents persistence, "Noise(N)" represents the Perlin noise at the N-th octave, and "p" represents the point to be calculated.

[0023] Further, the digital elevation map is constructed as follows:

[0024] Normalize the random map and map it into the range [0, 1], which is expressed as follows:

[0025]

[0026] Among them, "Y" represents the value obtained after normalization, "X" represents the original input value, "X" max represents the maximum value in the map data set, and "X" min represents the minimum value in the map data set;

[0027] Record the elevation data at each point using 8-bit binary integer variables, and uniformly quantize 256 levels of elevation values.

[0028] Further, perform 3D simulation interaction on the 3D unstructured map, specifically including:

[0029] Intercept a boolean slice map of the 3D environment at the same horizontal plane as the starting point of the UAV takeoff;

[0030] Set the initial generation radius "R" init , and in the boolean slice map, take the candidate point as the center and intercept a square matrix with a side length of 2R init and denote it as "A" init , and perform a dot product with the matrix "A" of the same dimension. Among them, the elements in "A" unit take values that satisfy the following formula: unit

[0031]

[0032] Among them, "a" and "b" respectively represent the horizontal and vertical coordinates of the point, and "x" a,b , "y" a,b respectively represent the horizontal and vertical distances of the point from the central candidate point;

[0033] According to "A" init ·"A" unitDetermine whether the candidate point is a feasible initial point by judging the maximum value in the result; if the maximum value is not greater than 0, it indicates that the area near the candidate point is relatively open and suitable as the starting flight point, and add it to the candidate point queue for subsequent random sampling in the research; otherwise, it indicates that there are relatively close obstacles near this point and it is not suitable as the starting flight point.

[0034] Further, the partially observable Markov decision process model for the autonomous exploration of the drone includes a flight dynamics constraint model and a drone action space model.

[0035] Further, the flight dynamics constraint model is expressed as follows:

[0036] Assume that the drone flies at a fixed speed V UAV Then its minimum turning radius R min , is expressed as follows:

[0037]

[0038] where g represents the acceleration due to gravity, and n y represents the maximum allowable normal overload coefficient of the drone;

[0039] Calculate the maximum course half-angle ψ, which is expressed as follows:

[0040]

[0041] where λ represents the movement step length of the drone, and dt represents the time interval of path planning;

[0042] Obtain the position and direction of the next step of the autonomous exploration plan of the drone, which is expressed as follows:

[0043]

[0044] where (x u , y u ), (x' u , y' u ) represent the position of the drone at the current moment and the position at the next moment respectively, respectively represent the direction of the drone at the current moment and the direction at the next moment.

[0045] Further, the drone action space model establishes a discrete action space and a continuous action space according to the relative course angle of the drone; among them, the discrete action space discretizes the maximum course angle range of the drone to obtain a finite countable action set; the continuous action space maps the maximum course angle range of the drone to [-1, 1] through the tanh activation function to make the action space continuous, and finally obtains an infinite real number action set.

[0046] Furthermore, the steps for constructing the route planning model include constructing local observations and a deep neural network from the local observations to the value function and the policy function;

[0047] Among them, the local observations use the ray casting method to simulate the fan-shaped field of view obtained by the UAV, and the state input during network training is the original local map restored from the local map observed by the UAV;

[0048] The deep neural network adopts the Actor-Critic framework based on policy and value. Among them, Actor represents the policy function π θ (a|s); Critic represents the value function V π (s);

[0049] The complex mapping relationship from local observations to the value function and the policy function can be fitted by a deep neural network. A cascaded three-layer convolutional neural network is used to extract feature information, and then two fully connected layers are used to output actions and values;

[0050] Construct a reward function, and the reward value R is expressed as follows:

[0051]

[0052] Among them, r e represents the reward.

[0053] Furthermore, the route planning model uses the PPO algorithm for gradient optimization. Among them, the objective function of PPO-Clip is designed as:

[0054]

[0055] Among them, θ represents the weight parameter of the route planning model, τ represents the trajectory explored according to the current policy π θ to obtain, represents the previous policy, represents the state s t when taking the action a t of the advantage function, ρ t represents importance sampling; clip(ρ t (θ),1-ε,1+ε) represents the clipping function, which means truncating ρ t (θ) within [1-ε,1+ε]. That is, when the amplitude representing the relative change of the new policy to the old policy in the function exceeds 1+ε, 1+ε is output, and when it is less than 1-ε, 1-ε is output; E represents the mathematical expectation, represents the mathematical expectation of the trajectory τ obtained according to the current exploration policy π θ ; Among them, the importance sampling ρ based on the weight parameter of the route planning modelt (θ) is calculated as follows:

[0056]

[0057] Wherein, represents the old advantage function, and π θ (a t |s t ) represents the updated advantage function; wherein, when the advantage function value is greater than 0 and is the largest, it means that the state s t - action a t has the best adaptability in the environment; when ρ t (θ) is the smallest when it exceeds 1 + ε and is the largest.

[0058] In the second aspect of the present invention, a drone route planning system based on deep reinforcement learning is provided, and the system includes:

[0059] A map construction module for constructing a three-dimensional unstructured map by using gradient-based Perlin noise and a digital elevation model;

[0060] A constraint model construction module for constructing a partially observable Markov decision process model for autonomous exploration of a fixed-wing drone with constraints to constrain the flight route of the fixed-wing drone;

[0061] A route planning optimization module for constructing a route planning model based on deep learning according to the partially observable Markov decision process model, and optimizing the route planning model by using the PPO algorithm.

[0062] The beneficial technical effects of the present invention are at least as follows:

[0063] (1) The present invention uses Perlin noise to autonomously construct a three-dimensional unstructured simulation environment and restore the missing elevation information of the Boolean map as the environment for subsequent experiments and tests. Reinforcement learning emphasizes interaction with the environment, so the rationality of the environment design will affect the efficiency of the agent's strategy learning to a certain extent. In the construction of a three-dimensional uncertain environment, the lack of an accurate environment model for interaction may lead to the lack of practicality of the obtained drone autonomous exploration strategy, and accurately modeling applicable terrains and obstacles is a complex problem. Therefore, the present invention uses Perlin noise with natural fractal properties to generate a continuous three-dimensional unknown environment and stores it as a digital elevation model as the map for the drone to explore the unknown environment, which can not only better simulate the continuous natural mountain landscape, but also achieve various different terrains by changing the random seed.

[0064] (2) In order to make the learned autonomous exploration strategy more robust, the present invention uses a fixed-wing UAV with more complex flight dynamics constraints as a reference object for flight dynamics constraint modeling, and completes action space modeling on this basis to constrain the behavior of the UAV during exploration and prevent the UAV from making physically unfeasible actions, thus establishing a foundation for the subsequent reinforcement learning framework. The establishment of the UAV flight dynamics constraint model will directly affect the state action space of the intelligent agent. This process not only requires the use of a fixed-wing UAV with more complex flight dynamics constraints as a reference object, but also requires the mapping of the flight dynamics constraint model to the reinforcement learning action space.

[0065] (3) The present invention uses the powerful feature representation ability of deep neural networks to fit the action value function, strategy, model and other components of reinforcement learning, constructs a deep neural network approximate mapping from local observations to value functions and strategy functions, and establishes a partially observable Markov decision process model in combination with the application scenario characteristics of autonomous exploration. It adopts the PPO algorithm suitable for fine control tasks and high-dimensional action spaces, and constructs a reasonable reward function to complete the reinforcement learning framework. It provides a better way to solve the curse of dimensionality, improves the robustness and generalization ability of the model when dealing with large-scale state spaces, makes reinforcement learning technology practical, and solves complex problems in real scenarios.

[0066] (4) The present invention adopts a method based on deep reinforcement learning to separate the time-consuming training process from the real-time decision-making process. After training a model with excellent performance, it is deployed on the drone. This not only ensures the response speed of the drone, but also provides a better way to solve the curse of dimensionality. Deep reinforcement learning allows drones to learn high-level task representations and control strategies directly from sensor data through end-to-end learning. It can automatically learn high-level feature representations without manually designing features or controllers, which helps intelligent agents better adapt to complex and changing environments. Through experience replay and parameter sharing of deep neural networks, deep reinforcement learning methods can better utilize limited data during training and improve learning efficiency. The present invention uses the method of deep reinforcement learning to solve the problem of autonomous exploration of UAVs in a three-dimensional uncertain environment, and explores how to use deep reinforcement learning to cope with the challenges of unstructured terrain. It can further improve the performance of the autonomous exploration system of UAVs, not only providing an important application scenario for the field of deep reinforcement learning, but also providing useful guidance for the application of UAVs in similar terrain areas. It is expected to fill the knowledge gap in the field of autonomous exploration, provide new theories and methods for autonomous exploration of UAVs, and promote the development of the field of deep reinforcement learning. At the same time, it can also provide valuable contributions to future research and practice, thereby improving work efficiency, reducing personnel costs, and reducing potential risks in multiple fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The present invention will be further described with reference to the accompanying drawings. However, the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the following drawings without creative efforts.

[0068] Figure 1 This is a flowchart of the UAV route planning method based on deep reinforcement learning according to the present invention.

[0069] Figure 2 This is a schematic diagram of the generation process of the base Berlin noise in an embodiment of the present invention.

[0070] Figure 3 This is a schematic diagram of the random map and its digital elevation map in an embodiment of the present invention.

[0071] Figure 4 This is a schematic diagram of the generation process of the sliced map in an embodiment of the present invention.

[0072] Figure 5 This is a schematic diagram of the feasible initial point and the infeasible initial point in an embodiment of the present invention.

[0073] Figure 6 This is a schematic diagram of the discrete and continuous action space modeling in an embodiment of the present invention.

[0074] Figure 7 This is a schematic diagram of the original local map and the UAV-observed local map in an embodiment of the present invention.

[0075] Figure 8 This is a schematic diagram of the neural network structure in an embodiment of the present invention.

[0076] Figure 9 This is a schematic diagram of the PPO algorithm process in an embodiment of the present invention.

[0077] Figure 10 This is a schematic diagram of the UAV autonomous environment exploration in an embodiment of the present invention.

[0078] Figure 11 This is a framework diagram of the UAV route planning system based on deep reinforcement learning according to the present invention. Detailed implementation manners

[0079] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0080] As Figure 1 shown, the UAV route planning method based on deep reinforcement learning provided by the embodiment of the present invention includes:

[0081] S1. Construct a three-dimensional unstructured map using gradient-based Perlin noise and a digital elevation model.

[0082] Specifically, Perlin noise is a continuous random texture generation algorithm proposed by Ken Perlin and is commonly used to simulate irregular patterns in nature, such as clouds and terrain. Traditional value-based noise divides the entire spatial region into several lattices, sets random values at the vertices of each lattice, i.e., lattice points, and calculates and processes the regions outside the lattice points using an interpolation function. Affected by the limitations of the generation method, value-based noise exhibits artificial regularity in some aspects and is not close enough to nature. In contrast, gradient-based noise generates gradient vectors instead of values at lattice points. Due to considering gradient information, gradient noise is usually smoother than value noise and has more advantages in simulating natural terrain and generating detailed patterns and textures. It is also the method used in the present invention to generate random terrain. The main generation process of gradient-based Perlin noise is as Figure 2 shown.

[0083] Furthermore, the gradient-based Perlin noise first generates a lattice rectangular grid covering the entire map and randomly initializes a gradient vector at each lattice point, specifically as follows:

[0084] Assume that a point P to be calculated within a lattice has four lattice points p0, p1, p2, and p3 to which it belongs, and its gradient vector is <grad0, grad1, grad2, grad3>. Then, by summing the dot products of the gradient vectors and the distance vectors at each lattice point according to formula (1), the random noise value of point P can be obtained. Calculate the offsets <delta0, delta1, delta2, delta3> of point P from the four lattice points respectively, and then sum the dot products of the gradient vectors and the distance vectors at each lattice point to obtain the random noise value of point P. The Perlin calculation formula is expressed as follows:

[0085]

[0086] where f(·) represents a high-order curve function used to enhance the smoothness of Perlin noise. f(u) is a high-order curve function to ensure the second-order continuous differentiability of the interpolation function and is used to enhance the smoothness of Perlin noise. Specifically, it is expanded as formula (2) and calculated as follows:

[0087] f(u) = 6u 5 - 15u 4 + 10u 3 (2)

[0088] where u represents the independent variable of the high-order curve function;

[0089] If only a single frequency is used to generate noise, although it is smooth overall but lacks local details, this can be solved by superimposing noises of multiple different frequencies. Such superimposed noise is also called fractal noise. In Perlin noise, Persistence is used to represent the amplitude of each frequency component. The definitions of its frequency and amplitude are shown in Formulas (3) and (4), where i represents the i-th superimposed noise function. The frequency of each noise function is twice that of the previous one, and each superimposed layer of noise is an octave. The expression of the final obtained noise is shown in Formula (5). N is the number of octaves. Perlin represents the Perlin noise operation on the frequencies and amplitudes with different values at each point, that is, the process of Formulas (1) and (2). The calculation formula is as follows:

[0090] frequency=2 k (3)

[0091] amplitude=persistence k (4)

[0092]

[0093] Among them, frequency represents the frequency, amplitude represents the amplitude, Perlin represents the Perlin noise operation on the frequencies and amplitudes with different values at each point, that is, the k-th superimposed noise function. N represents the number of octaves, persistence represents the Persistence, Noise(N) represents the Perlin noise at the N-th octave, and p represents the point to be calculated.

[0094] Specifically, a digital elevation model is a digital map that stores surface elevation information in the form of raster data. This elevation information can be used for many geographical and environmental science applications such as drawing topographic maps, performing geographic information system analysis, terrain modeling, and hydrological modeling. The map generated using Perlin noise is also stored in the form of raster data, but its values are in the interval [-x, x] (0 < x < 1). To facilitate better extraction of data information in subsequent research, before converting the random map into a digital elevation model, the random map needs to be normalized according to Formula (6) and mapped into [0, 1]. In the formula, Y is the value obtained after normalization, X is the original input value, X max is the maximum value in the map data set, and X min is the minimum value in the map data set.

[0095]

[0096] Meanwhile, in order to reduce the storage burden of the UAV, when storing the elevation data at each point, an 8-bit binary integer variable is used for recording, and a total of 256 levels of elevation values can be evenly quantized. Due to the limitations of its own power supply and communication range, the flight altitude of the UAV is limited. In the present invention, the vertical resolution is 1 meter, and an elevation map with an altitude of 0 - 255 meters can be recorded. Finally, the normalized random map and its corresponding digital elevation map are as Figure 3 shown.

[0097] Furthermore, after establishing the three-dimensional simulation interaction environment, it is necessary to randomly set the initial flight point of the UAV at a suitable position to avoid the occurrence of random initial flight points that violate physical conditions. In the present invention, the flight altitude of the starting point of the UAV is set to 128 meters, and a three-dimensional environment on the same horizontal plane as the starting point is intercepted for making a Boolean slice map. The specific process is as Figure 4 shown.

[0098] Map the slice map to the normalized map, set the height threshold to 0.508, perform threshold segmentation calculation, set the data higher than the threshold to 1, and the data not greater than the threshold to 0. The points with a value of 0 in the threshold-segmented Boolean slice map are potential initial flight points. However, some points with a value of 0 are too close to obstacles. Even if the UAV is initialized at these points, there is a high probability that the UAV will crash in the initial stage of exploration. Therefore, it is necessary to detect obstacles around the candidate initial flight points before initialization to ensure that the UAV starts exploration in a relatively stable and safe environment. The specific detection method is to set the initial generation radius R init , and in the slice map, with the candidate point as the center, intercept a square matrix with a side length of 2R init and denote it as A init , and perform a dot product with the matrix A unit of the same dimension. Among them, the element values in A unit satisfy formula (7), where a and b are the horizontal and vertical coordinates of the point respectively, and x a,b , y a,b are the horizontal and vertical distances of the point from the central candidate point respectively, and are expressed as follows:

[0099]

[0100] Among them, a and b respectively represent the horizontal and vertical coordinates of the point, and x a,b , y a,b respectively represent the horizontal and vertical distances of the point from the central candidate point.

[0101] According to A init ·A unitJudge whether the candidate point is a feasible initial point based on the maximum value in the result. If the maximum value is not greater than 0, it indicates that the area near the candidate point is relatively open and suitable as a starting flight point, and it will be added to the candidate point queue for subsequent random sampling in the research; otherwise, it indicates that there are relatively close obstacles near this point and it is not suitable as a starting flight point. The schematic diagrams of feasible initial points and infeasible initial points are as shown in Figure 5 shown. The initial direction of the UAV is randomly generated. The inner circle of the ring is the field of view range that the UAV may observe, and the outer circle is the detection range of the initial generation radius.

[0102] S2. Construct a partially observable Markov decision process model for the autonomous exploration of UAVs with constraints based on fixed-wing UAVs to constrain the flight route of fixed-wing UAVs.

[0103] Specifically, the partially observable Markov decision process model for the autonomous exploration of UAVs includes a flight dynamics constraint model and a UAV action space model.

[0104] Among them, for the exploration of outdoor unstructured environments, the flight range is large, the altitude changes greatly, and the flight dynamics model and flight dynamics constraints of fixed-wing UAVs are more complex than those of multi-rotor UAVs. Therefore, for the sake of generality, a fixed-wing UAV is proposed as the reference object for flight dynamics constraint modeling.

[0105] Preferably, based on the fixed-wing UAV flight dynamics model, flight dynamics constraints represented by parameters such as the minimum turning radius and the maximum heading half-angle of the UAV in the Boolean slice map are derived. Assume that the UAV flies at a fixed speed V UAV The minimum turning radius R of it min is shown in formula (8):

[0106]

[0107] In the formula, g is the acceleration due to gravity, taking 9.82 m / s 2 , n y is the maximum allowable normal overload coefficient of the UAV, taking 20. In the present invention, V UAV is set to 20 m / s. By calculation, the minimum turning radius R of the UAV model min is approximately 2.04 m. Substituting it into formula (9) can obtain the maximum heading half-angle ψ, where the UAV movement step length λ is shown in formula (10). dt is the time interval for path planning, taking 0.1 s. Therefore, it is calculated that λ is 2 m, and then the maximum heading half-angle of the UAV in the present invention is approximately 29.3°, and the maximum heading angle range of the UAV is [-29.3°, 29.3°].

[0108]

[0109] λ = V UAV × dt (10)

[0110] The update of the next position and direction in the autonomous exploration and planning of the UAV is shown in Equation (11):

[0111]

[0112] In the formula, (x u , y u ) and (x' u , y' u ) are the positions of the UAV at the current moment and the next moment respectively, are the directions of the UAV at the current moment and the next moment respectively.

[0113] Furthermore, an action is a means for an agent to interact with the environment and is a core part of reinforcement learning. It is crucial for the performance of the system and the learning of the policy, and directly affects the behavior of the agent in the environment and the acquisition of reward results. The action space can be divided into a discrete action space, a continuous action space, and a mixed action space according to its nature. Different action space modellings will also affect the construction of the subsequent reinforcement learning framework.

[0114] The system in the present invention adopts the same type of action internally and does not involve mixed actions. After determining the maximum heading angle range of the UAV, the set of its action space is also determined, and the values of all actions are within the heading angle range. The design of the action space considers the relative heading angle of the UAV, and a discrete action space and a continuous action space are established for it, as Figure 6 shown. The discrete action evenly discretizes the maximum heading angle range of the UAV, and finally obtains a finite countable action set, which can reduce the dimension of the action space; the continuous action maps the maximum heading angle range of the UAV to [-1, 1] through the tanh activation function, making the action space continuous, and finally obtains an infinite set of real number actions, which has higher complexity but stronger adaptability to the environment.

[0115] S3. Construct a route planning model based on deep learning according to the partially observable Markov decision process model, and use the PPO algorithm to optimize the gradient of the route planning model.

[0116] Specifically, reinforcement learning requires the agent to continuously take actions to interact with the environment to obtain experience. Based on this experience, it improves the policy through learning and then takes actions according to the policy, repeating this process until the optimal converged policy is obtained. Therefore, the research work of reinforcement learning mainly consists of two parts: the reinforcement learning method framework and the corresponding training method. This part proposes approximate expression forms of the value function and the policy function, and the internally motivated reward function, and constructs a learning framework adapted to the characteristics of the problem based on the POMDP mathematical model (Partially Observable Markov Decision Process model) considering various constraints.

[0117] Among them, in the reinforcement learning of partially observable problems, mapping from the observation information to the value function and the policy function is the basis for the entire reinforcement learning to perform value estimation and policy improvement. At the same time, using a deep network for approximate mapping is also an important means to solve the curse of dimensionality. This part mainly consists of two aspects: on the one hand, constructing local observations; on the other hand, building a deep neural network from local observations to the value function and the policy function.

[0118] For the convenience of expression, the partially observable Markov decision process model can be represented as a tuple <S, A, T, R, Ω, O, γ>, where each element represents a different part of the partially observable Markov decision process model: S is a finite set of states s, which represents all situations that the environment can present; A is the set of action spaces; T is the state transition probability function; R is the immediate reward function; Ω is the set of observation spaces; O is the observation probability function; γ is the discount factor for future rewards obtained.

[0119] Constructing local observations: The sum of the unknown areas where the drone needs to perform exploration tasks is the global map. Under this condition, the drone cannot fully observe all the state information. During the exploration process, the drone can obtain the environmental digital elevation information within the sensing range through various cameras and radars as the local map, but the original objective state information will be distorted after passing through the drone's sensors. The present invention uses the ray casting method to simulate the fan-shaped field of view obtained by the drone. The state input during network training is the original local map restored from the local map observed by the drone. The schematic diagrams of the two are as Figure 7 shown.

[0120] Building a deep neural network from local observations to the value function and the policy function: In order to enable the agent to learn the optimal policy, the present invention intends to adopt the Actor-Critic framework based on policy and value. Actor-Critic, also translated as actor-critic, is a reinforcement learning method that combines policy gradient and temporal difference learning. Among them, the actor refers to the policy function π θ (a|s), that is, learning a policy to obtain the highest possible return, which determines how the agent takes actions according to the current state; the critic refers to the value function Vπ (s) estimates the value of the current policy, that is, evaluates the quality of the actor by the magnitude of the estimated value in the current state. With the help of the value function, the actor-critic algorithm can perform single-step parameter updates without waiting until the end of the episode.

[0121] The complex mapping relationship from local observations to the value function and the policy function can be approximated by a deep neural network. However, based on only a single input local observation map, it is difficult for the neural network to fully grasp the motion trend of the UAV. Therefore, four consecutive 40×40 local observation map images are input as the current state s. Next, a cascaded three-layer convolutional neural network is used to extract feature information, and then two fully connected layers are used to output actions and values. To avoid excessive changes in the UAV's heading angle, which may affect the safe flight of the UAV, in the present invention, the action space A in the discrete state is equally divided into 21 parts, i.e., a0, a1,..., a 20} and each action of the UAV originates from this state space. The structure of the deep neural network is as Figure 8 shown. The dimensions of the last fully connected layer of the policy network and the value network are not the same because the final output of the policy network is one of 21 actions, and the output of the value network is a single estimated value.

[0122] Furthermore, the steps of constructing the route planning model include constructing local observations and a deep neural network from local observations to the value function and the policy function;

[0123] Among them, the local observations use the ray casting method to simulate the fan-shaped field of view obtained by the UAV, and the state input during network training is the original local map restored from the local map observed by the UAV;

[0124] The deep neural network adopts the Actor-Critic framework based on policy and value, where Actor represents the policy function π θ (a|s); Critic represents the value function V π (s);

[0125] The complex mapping relationship from local observations to the value function and the policy function can be approximated by a deep neural network. A cascaded three-layer convolutional neural network is used to extract feature information, and then two fully connected layers are used to output actions and values;

[0126] In reinforcement learning, the reward function plays a pivotal role and is the main guide for the learning and decision-making of the agent. The reward function defines the goal of the problem. By giving the agent an immediate reward after taking a certain action at each time step, it guides the agent to find a strategy that can maximize the cumulative reward. In this invention, in order to ensure that the drone can explore as many unknown areas as possible, the first thing is to let the drone learn the autonomous obstacle avoidance strategy to prevent crashes, and construct a reward function. The reward value R is expressed as follows:

[0127]

[0128] Among them, r e Represents reward. When the drone collides with an obstacle, the reward value is a large negative value of -1000, and the exploration of this round will end directly. This state is called done in reinforcement learning. When the drone is in a normal exploration state, its reward is r e The area of ​​the unknown area newly explored in a single step multiplied by a fixed coefficient. Considering the physical limitations of the drone's endurance, in addition to ending the exploration when it hits an obstacle, it is also set that the exploration will end when the drone's exploration steps per round reach the maximum number of steps.

[0129] Furthermore, the route planning model is optimized using the PPO algorithm.

[0130] Among them, the PPO algorithm is an on-policy algorithm based on the policy gradient optimization method proposed by OpenAI. It can effectively improve the stability of policy optimization and sample utilization efficiency. Due to its relatively simple implementation and good performance, it has become one of the preferred algorithms for many reinforcement learning tasks. The PPO algorithm has two main variants: Proximal Policy Optimization Penalty (PPO-Penalty) and Proximal Policy Optimization Clipping (PPO-Clip). PPO-Clip is better than PPO-Penalty and is the default variant when PPO is mentioned. The present invention intends to use this algorithm to build a reinforcement learning framework. Its overall process is as follows Figure 9 As shown in the figure, it can be divided into three parts: reinforcement learning environment interaction, Critic network update and Actor network update.

[0131] The environmental interaction part is a classic framework for reinforcement learning. The agent takes actions in the environment according to its current state and its own policy. The environment feedbacks an immediate reward to the agent, and at the same time updates the agent to the next state, forming a partially observable Markov decision process. In the present invention, this part is handed over to the Actor with the latest policy to collect the trajectory before the drone reaches the end state, and the maximum step length is not exceeded. Before entering the network update part, multiple states collected within each episode need to be packed into states for the network to perform multiple rounds of training and update. The update of the Critic network needs to calculate the advantage function based on the target value and the predicted value, and use the least mean square error as the loss function to perform backpropagation update on the network weight parameters. The Actor network needs the respective probability distributions of the new and old network outputs for the current state, and then updates the network according to the objective function of PPO-Clip.

[0132] Among them, the objective function of PPO-Clip is designed as:

[0133]

[0134] Among them, θ represents the weight parameter of the route planning model, τ represents the trajectory explored according to the current policy π θ obtained by exploration, represents the previous policy, represents the state s t and taking action a t under the state s t of the advantage function, ρ t represents importance sampling; clip(ρ t (θ), 1 - ε, 1 + ε) represents the clipping function, which means truncating ρ (θ) within [1 - ε, 1 + ε], that is, when the amplitude representing the relative change of the new policy with respect to the old policy in the function exceeds 1 + ε, it outputs 1 + ε, and when it is less than 1 - ε, it outputs 1 - ε; E represents the mathematical expectation, represents the mathematical expectation of the trajectory τ obtained according to the current exploration policy π θ ; among them, the importance sampling ρ t (θ) based on the weight parameter of the route planning model is calculated as follows:

[0135]

[0136] Among them, represents the old advantage function, and π θ (a t | s t ) represents the updated advantage function; among them, when the advantage function value is greater than 0 and is the largest, it means that the state s t - action at is the best adapted to the environment; when ρ t (θ) exceeds 1 + ε and is at its maximum, the value of the updated advantage function is minimized.

[0137] Specifically, the clipping function clip(ρ t (θ), 1 - ε, 1 + ε) truncates ρ t (θ) within [1 - ε, 1 + ε], that is, when the amplitude representing the relative change of the new policy with respect to the old policy in the function exceeds 1 + ε, it outputs 1 + ε, and when it is less than 1 - ε, it outputs 1 - ε, thus ensuring the similarity between the new and old policies. When the value of the advantage function is greater than 0, it means that this state-action pair is better adapted to the environment and helps to obtain a greater reward, that is, π θ (a t |s t ) is better, and a greater objective function value can be obtained. However, when ρ t (θ) exceeds 1 + ε, it does not bring an improvement in the training effect for the model and is not conducive to the learning of the optimal policy, and vice versa. Finally, the smaller of the truncated objective function and the untruncated objective function is taken as the final objective function. It is such a conservative update method that greatly improves the stability and convergence of model training.

[0138] Preferably, the present invention preliminarily constructs an experimental environment, covering parts such as three-dimensional unstructured environment generation, drone exploration environment, learning strategy update, etc., and the example is Figure 10 as shown.

[0139] Among them, Figure 10 is an example of the preliminarily constructed experimental environment, covering parts such as three-dimensional unstructured environment generation, drone exploration environment, learning strategy update, etc. Figure 10 (a) is a top-down three-dimensional environment generated using Perlin noise, and a map with a size of 1000m × 1000m is framed as the exploration area. Since this global map is unknown to the drone, a mask is covered on it, and the exploration map is updated whenever the drone explores a new area. When training, after the drone is initialized at the starting point, it randomly selects a heading angle and starts to explore at a constant speed until done or reaches the maximum number of steps to stop the exploration of the current episode. Figure 10 (b) is the final exploration map. The lines with color changes in the figure are the exploration trajectories obtained by the drone over time, and the starting point and ending point positions are also marked with dots. Figure 10 (c) is the change of the reward curve. As the number of training episodes increases, the total reward value that the obtained drone can get in the environment is also continuously rising.

[0140] Such as Figure 11As shown in the figure, the embodiment of the present invention also provides a UAV route planning system based on deep reinforcement learning, and the system includes:

[0141] A map construction module 101, configured to construct a three-dimensional unstructured map by using gradient-based Perlin noise and a digital elevation map;

[0142] A constraint model construction module 102, configured to construct a partially observable Markov decision process model for autonomous exploration of a fixed-wing UAV with constraints to constrain the flight route of the fixed-wing UAV;

[0143] A route planning optimization module 103, configured to construct a route planning model based on deep learning according to the partially observable Markov decision process model, and use the PPO algorithm to perform gradient optimization on the route planning model.

[0144] In summary, the technical solution of the present invention adopts a method based on deep reinforcement learning, which can automatically learn high-level features from the original input data, does not require prior knowledge or an accurate environmental model, and does not require manual feature design. The PPO algorithm, which is suitable for fine control tasks and high-dimensional action spaces, can adapt to high-dimensional and complex environments, including continuous action and state spaces, and is suitable for dealing with nonlinear and unstructured problems. The finally obtained model can transfer the experience learned from one task to related tasks, has generalization ability, and continuously learns and improves in the interaction with the environment to adapt to changing environmental conditions. Deploying the trained model directly on the UAV ensures the real-time nature of decision-making and improves task efficiency. For the method of binarizing the map area, the present invention retains the elevation information of the three-dimensional map, and the more comprehensive information can help the network train an intelligent agent model with a more sound autonomous exploration strategy, improving the adaptability of the UAV in different situations.

[0145] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0146] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0147] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0148] Although the embodiments of the present invention have been shown and described, those skilled in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A UAV route planning method based on deep reinforcement learning, characterized in that: The method comprises: S1. Constructing 3D unstructured maps using gradient-based Perlin noise and digital elevation maps; S2. Construct a partially observable Markov decision process model for autonomous exploration of fixed-wing UAVs with constraints to constrain the flight path of fixed-wing UAVs. S3. Constructing a route planning model based on deep learning according to the partially observable Markov decision process model, and optimizing the route planning model using the PPO algorithm; The gradient-based Perlin noise first generates a lattice rectangular grid covering the entire image, and randomly initializes a gradient vector at each lattice point, as follows: Assume that a point to be calculated in the lattice is , the point to be calculated The four lattice points are denoted as , , , , then its gradient vector is < , , , > Calculate The offset of each point from the four lattice points is < , , , >, and then the dot product of the gradient vector and the distance vector at each lattice point is summed to obtain The random noise value of the point, The calculation formula is as follows: ; in, Represents a high-order curve function, which is used to improve the smoothness of Perlin noise. The calculation is as follows: ; in, Represents the independent variable of the higher-order curve function; Use a single frequency to generate noise, then calculate by superimposing multiple noises of different frequencies, and use persistence to represent the amplitude of each frequency component. The calculation formula is as follows: ; ; ; in, Indicates frequency, represents the amplitude, represents the Perlin noise operation with different values ​​of frequency and amplitude at each point, that is, the kth superimposed noise function, N represents the number of multiples, Indicates the duration, It represents Perlin noise at N times the frequency. Indicates the point to be calculated; The digital elevation map is constructed as follows: Normalize the random map and map it to [0, 1], as follows: ; Among them, Y represents the normalized value, and X represents the original input value. Indicates the maximum value in the map data set. Represents the minimum value in the map data set; The elevation data at each point is recorded using an 8-bit binary integer variable, and 256 levels of elevation values ​​are evenly quantized; The three-dimensional unstructured map is used for three-dimensional simulation interaction, including: Intercept the three-dimensional environment at the same level as the starting point of the drone takeoff to create a Boolean slice map; Set the initial generation radius , in the Boolean slice map, the candidate point is taken as the center and the length of the interception is The square matrix is ​​denoted as , and a matrix of the same dimension Perform dot multiplication, where The value of the inner element satisfies the following formula: ; in, Respectively represent the horizontal and vertical coordinates of the candidate points, , Respectively represent the horizontal distance and vertical distance between the candidate point and the central candidate point; according to The maximum value of the result is used to determine whether the candidate point is a feasible starting flight point; if the maximum value is not greater than 0, it means that the area near the candidate point is relatively empty and suitable as a starting flight point, and it is added to the candidate point queue for subsequent random sampling; otherwise, it means that there are relatively close obstacles near the point and it is not suitable as a starting flight point.

2. The UAV route planning method based on deep reinforcement learning according to claim 1, characterized in that: The partially observable Markov decision process model for autonomous exploration of the UAV includes a flight dynamics constraint model and a UAV action space model.

3. The UAV route planning method based on deep reinforcement learning according to claim 2 is characterized in that: The flight dynamics constraint model is expressed as follows: Assume that the drone is flying at a fixed speed Flying, the minimum turning radius , which is expressed as follows: ; Where g represents the acceleration due to gravity, Indicates the maximum allowable normal overload factor of the drone; Calculate the maximum heading half angle , which is expressed as follows: ; ; in, represents the motion step length of the drone, Indicates the time interval for path planning; Get the next position and direction of the UAV's autonomous exploration plan, as follows: ; in, Respectively represent the position of the drone at the current moment and the position at the next moment. , They represent the direction of the drone at the current moment and the direction at the next moment respectively.

4. The UAV route planning method based on deep reinforcement learning according to claim 2 is characterized in that: The UAV action space model establishes a discrete action space and a continuous action space according to the relative heading angle of the UAV; wherein the discrete action space separates and discretizes the maximum heading angle range of the UAV to obtain a finite countable action set; the continuous action space maps the maximum heading angle range of the UAV to [-1, 1] through a tanh activation function, so that the action space is continuous, and finally an infinite real number action set is obtained.

5. The method for UAV route planning based on deep reinforcement learning according to claim 1, characterized in that: The route planning model building step includes constructing local observations and a deep neural network from local observations to value functions and policy functions; The local observation uses a ray casting method to simulate the fan-shaped field of view obtained by the drone, and the input state during network training is the original local map restored by the drone observing the local map; The deep neural network adopts the Actor-Critic framework based on strategy and value, where Actor represents the strategy function ; Critic represents the value function ; The complex mapping relationship from local observation to value function and policy function can be fitted through a deep neural network. A cascaded three-layer convolutional neural network is used to extract feature information, and two fully connected layers are used to output actions and values. Construct reward function, reward value It is expressed as follows: ; in, Indicates reward.

6. The method for UAV route planning based on deep reinforcement learning according to claim 5, characterized in that: The route planning model uses the PPO algorithm for gradient optimization, where the objective function of PPO-Clip is Designed for: ; in, represents the weight parameter of the route planning model, Indicates that according to the current strategy The trajectory obtained, represents the previous strategy, Indicates status Take action The advantage function of represents importance sampling; Represents the clipping function, which means Truncated at [ ], that is, when the function represents the relative change of the new strategy to the old strategy exceeds Output , less than It will output ; E represents mathematical expectation, According to the current exploration strategy The obtained trajectory The mathematical expectation value of ; Among them, the importance sampling of the weight parameters based on the route planning model The calculation is as follows: ; in, represents the old advantage function, Represents the updated advantage function; when the advantage function value is greater than 0 and is the maximum, it represents the state -action The best adaptability to the environment; when When more than And when is maximum, the value of the updated advantage function reaches the minimum.

7. The UAV route planning system based on deep reinforcement learning is characterized by: The system comprises: A map construction module for constructing 3D unstructured maps using gradient-based Perlin noise and digital elevation maps; The constraint model building module is used to build a partially observable Markov decision process model for autonomous exploration of fixed-wing UAVs with constraints to constrain the flight path of fixed-wing UAVs. A route planning optimization module, used to construct a deep learning-based route planning model according to a partially observable Markov decision process model, and optimize the route planning model using a PPO algorithm; The gradient-based Perlin noise first generates a lattice rectangular grid covering the entire image, and randomly initializes a gradient vector at each lattice point, as follows: Assume that a point to be calculated in the lattice is , the point to be calculated The four lattice points are denoted as , , , , then its gradient vector is < , , , > Calculate The offset of each point from the four lattice points is < , , , >, and then the dot product of the gradient vector and the distance vector at each lattice point is summed to obtain The random noise value of the point, The calculation formula is as follows: ; in, Represents a high-order curve function, which is used to improve the smoothness of Perlin noise. The calculation is as follows: ; in, Represents the independent variable of the higher-order curve function; Use a single frequency to generate noise, then calculate by superimposing multiple noises of different frequencies, and use persistence to represent the amplitude of each frequency component. The calculation formula is as follows: ; ; ; in, Indicates frequency, represents the amplitude, represents the Perlin noise operation with different values ​​of frequency and amplitude at each point, that is, the kth superimposed noise function, N represents the number of multiples, Indicates the duration, It represents Perlin noise at N times the frequency. Indicates the point to be calculated; The digital elevation map is constructed as follows: Normalize the random map and map it to [0, 1], as follows: ; Among them, Y represents the normalized value, and X represents the original input value. Indicates the maximum value in the map data set. Represents the minimum value in the map data set; The elevation data at each point is recorded using an 8-bit binary integer variable, and 256 levels of elevation values ​​are evenly quantized; The three-dimensional unstructured map is used for three-dimensional simulation interaction, including: Intercept the three-dimensional environment at the same level as the starting point of the drone takeoff to create a Boolean slice map; Set the initial generation radius , in the Boolean slice map, the candidate point is taken as the center and the length of the interception is The square matrix is ​​denoted as , and a matrix of the same dimension Perform dot multiplication, where The value of the inner element satisfies the following formula: ; in, Respectively represent the horizontal and vertical coordinates of the candidate points, , Respectively represent the horizontal distance and vertical distance between the candidate point and the central candidate point; according to The maximum value of the result is used to determine whether the candidate point is a feasible starting flight point; if the maximum value is not greater than 0, it means that the area near the candidate point is relatively empty and suitable as a starting flight point, and it is added to the candidate point queue for subsequent random sampling; otherwise, it means that there are relatively close obstacles near the point and it is not suitable as a starting flight point.

Citation Information

Patent Citations

  • Data processing method and device

    CN114201569A

  • Evolution-based multi-objective reinforcement learning vehicle route planning method

    CN115907254A

  • Unmanned aerial vehicle visual obstacle avoidance and autonomous navigation method based on improved PPO

    CN117705113A

  • Robot path planning method and device based on reinforcement learning

    CN118163101A