Unmanned aerial vehicle intelligent dynamic path guidance method and system based on deep reinforcement learning
Through technologies such as deep reinforcement learning and graph convolutional networks, a dynamic path planning system for drones in complex environments is built, which solves the problems of insufficient path security and multi-index conflicts in traditional methods, and achieves safe and efficient drone flight.
Patent Information
- Application Number
- CN202510255328.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional UAV dynamic path planning methods are difficult to effectively deal with three-dimensional terrain changes and dynamic obstacle characteristics, resulting in insufficient path safety and difficult to balance conflicts of multiple indicators such as energy consumption, time and safety, and are prone to falling into local optimality.
Using a method based on deep reinforcement learning, an environmental topology diagram is constructed by obtaining three-dimensional terrain point cloud data and obstacle geometric information, a graph convolution network is used to extract environmental features, and a multi-objective reward function is constructed, and a path scheme is optimized by combining dynamic programming and gradient descent method.
It has achieved safe flight of drones, optimized path planning, enhanced model adaptability and improved independent decision-making capabilities in complex environments, and can effectively balance energy consumption, time and safety.
Smart Images

Figure CN120103856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for intelligent dynamic path guidance of an unmanned aerial vehicle based on deep reinforcement learning. Background Art
[0002] With the widespread application of drones in complex environments such as disaster relief, logistics transportation, and military reconnaissance, dynamic path planning technology faces many challenges. First, traditional methods are mostly based on grid maps or two-dimensional topological modeling, which makes it difficult to effectively handle three-dimensional terrain changes and dynamic obstacle characteristics, resulting in insufficient path safety. For example, in urban environments, drones need to consider the height and shape changes of buildings, as well as the real-time positions of dynamic obstacles (such as moving vehicles), which puts higher requirements on traditional models. At the same time, existing reinforcement learning methods mostly use a single reward function, which makes it difficult to balance the conflicts of multiple indicators such as energy consumption, time, and safety, and are prone to fall into local optimality. For example, in logistics distribution scenarios, drones need to reach their destinations in the shortest time while ensuring minimum energy consumption and flight safety. This multi-objective optimization problem is difficult to effectively solve in existing methods. Moreover, static planning methods cannot cope with sudden obstacle movements or changes in environmental parameters, and re-planning efficiency is low. For example, in a dynamic environment, drones need to adjust their paths in real time to avoid sudden obstacles, and traditional methods often cannot respond quickly in this case.
[0003] Although the existing technology has made some progress, it still has some shortcomings. For example, the path planning method based on convolutional neural network proposed in patent CN112230675A has improved the efficiency of path planning to a certain extent, but it has not solved the problems of three-dimensional feature extraction and dynamic adjustment of multi-target weights, and it is difficult to meet the wide demand of drones. Summary of the invention
[0004] The main purpose of the present invention is to provide a method and system for intelligent dynamic path guidance of unmanned aerial vehicles based on deep reinforcement learning, so as to achieve safe flight of unmanned aerial vehicles in complex environments, optimize path planning, enhance model adaptability and improve autonomous decision-making capabilities.
[0005] To achieve the above object, the present invention provides a method for intelligent dynamic path guidance of a UAV based on deep reinforcement learning, comprising the following steps:
[0006] Obtain the 3D terrain point cloud data and obstacle geometry information of the UAV flight area, construct the environment topology structure map, extract the local and global features of the environment topology structure map through a graph convolutional network containing 3 layers of graph convolutional layers, and generate a high-dimensional environment feature representation;
[0007] Based on the high-dimensional environmental feature vector, a state space of deep reinforcement learning is constructed. The state space includes the position coordinates of the drone, the environmental feature vector, and the remaining energy state. A multi-objective reward function is designed. The reward function is obtained by weighted summation of energy consumption index, time index, and safety index.
[0008] The deep reinforcement learning network in the path planning model is initialized with random parameters uniformly distributed in [-0.1, 0.1], and an initial path plan is generated according to the input state space information. The path planning model performs path planning based on the deep reinforcement learning state space including the drone position coordinates, the environmental feature vector and the remaining energy state;
[0009] A dynamic programming algorithm is used to optimize the path node sequence in the initial path plan. When the path length exceeds the dynamic threshold corresponding to the task type, the node sequence is readjusted and each objective function value is calculated. The objective function value includes energy consumption, time and safety indicators.
[0010] Acquire the path length of the initial path plan, and when the path length exceeds the dynamic threshold corresponding to the current flight mission, adjust the path node sequence through a dynamic programming algorithm to generate a current path plan;
[0011] Obtaining the values of each objective function of the current path plan, and judging whether the optimal balance point is reached, wherein the optimal balance point is a solution in which the energy consumption, time and safety indicators are all within the preset range of the task; if the optimal balance point is not reached, adjusting the weight coefficient of the reward function by the gradient descent method, inputting the updated weight coefficient into the path planning model to regenerate the path plan, and repeating the optimization and judgment steps until the final path plan is output;
[0012] The final path plan is converted into waypoint coordinates and speed instructions, and the environmental changes are monitored in real time during the flight. When the obstacle movement exceeds the safety threshold, the path replanning mechanism is triggered. Furthermore, the three-dimensional terrain point cloud data and obstacle geometry information of the UAV flight area are obtained, the environment topology structure map is constructed, and the local and global features of the environment topology structure map are extracted through a graph convolution network containing 3 layers of graph convolution layers to generate a high-dimensional environment feature representation step, including:
[0013] The three-dimensional terrain point cloud data of the flight area is obtained by laser radar scanning, and the geometric features and location information of obstacles are captured by visual sensors. The terrain elevation data and the obstacle spatial coordinates are integrated into an environmental topological structure diagram, in which the nodes of the topological structure diagram represent the terrain or obstacle feature points, and the edges represent the spatial connection relationship between the nodes;
[0014] Based on the node information and edge connection relationship of the environment topology graph, a network with three layers of graph convolution layers is constructed. Multi-layer graph convolution layers are used to extract features of topological relationships to obtain low-dimensional local feature representation.
[0015] Skip connections are used to concatenate the low-dimensional local features output by the first layer of graph convolution with the global features output by the third layer to generate a high-dimensional environment feature vector.
[0016] Furthermore, the step of constructing a state space of deep reinforcement learning based on a high-dimensional environment feature vector and designing a multi-objective reward function includes:
[0017] Get the current position coordinates of the drone, combine the environmental characteristics and energy state, and construct the state space;
[0018] A deep reinforcement learning algorithm is used to define a multi-objective reward function in the state space, wherein the reward function includes three indicators: energy consumption, time, and safety.
[0019] Among them, the energy consumption index is calculated based on the power consumption per unit distance, the time index is calculated based on the expected flight time, and the safety index is calculated based on the obstacle distance.
[0020] Furthermore, the step of using a random parameter uniformly distributed in [-0.1, 0.1] to initialize the deep reinforcement learning network in the path planning model and generating an initial path plan according to the input state space information includes:
[0021] The weight parameters of the deep reinforcement learning network are set to be evenly distributed in the range of [-0.1, 0.1], and the curvature of the network output path is limited to not exceed the maximum steering angle of the drone;
[0022] The constructed state space information including the drone position coordinates, high-dimensional environmental feature vectors, and residual energy state is input into the initialized deep reinforcement learning network to generate an initial path plan in the form of a path node sequence;
[0023] Calculate the total path length L and energy consumption E of the initial path plan, and evaluate whether the steering between each path node in the initial path plan is within the achievable range of the drone according to the maximum steering angle and energy constraint conditions preset by the drone, and determine whether L≤L max And E≤E max , wherein L max is the preset maximum path length, the E max is the preset maximum energy consumption;
[0024] If the initial path plan satisfies the preset maximum steering angle, L≤L max And E≤E maxIf the constraints are met, the initial path plan is accepted.
[0025] Furthermore, the path planning model adopts a deep deterministic policy gradient algorithm, with a combination of policy gradient and deep neural network as the basic architecture, and takes the deep reinforcement learning state space including the drone position, environmental characteristics and residual energy as input. After processing through the fully connected layer, it is input into the policy network composed of multiple hidden layers to generate flight control instructions for direction and speed. The path planning model also includes a value network for evaluating the value of actions, and introduces an experience replay mechanism.
[0026] Furthermore, the step of obtaining the path length of the initial path plan and, when the path length exceeds a dynamic threshold corresponding to the current flight mission, adjusting the path node sequence by a dynamic programming algorithm to generate a current path plan includes:
[0027] Calculate the corresponding dynamic threshold according to the current flight mission, and obtain the path length from the flight start point to the end point in the initial path plan;
[0028] Calculate the objective function values of the initial path plan, the energy consumption, time and safety indicators of the objective function values, and determine whether the path length exceeds the dynamic threshold of the current flight mission. If the path length exceeds the dynamic threshold, compress the path node sequence through a dynamic programming algorithm, delete redundant nodes, and generate an optimized current path plan.
[0029] Furthermore, the step of optimizing the path node sequence in the initial path plan by using a dynamic programming algorithm, and re-adjusting the node sequence and calculating each objective function value when the path length exceeds the dynamic threshold corresponding to the task type, further includes:
[0030] Using GPU parallel computing framework, it can simultaneously perform forward reasoning and back propagation tasks of graph neural network and deep reinforcement learning network;
[0031] If the single iteration calculation time exceeds the threshold, the network structure is adjusted, and the network structure adjustment includes reducing the number of graph convolutional network layers to 2 layers, reducing the number of hidden layer neurons in the deep reinforcement learning network, and recalculating the adjusted network performance.
[0032] Further, the step of obtaining the objective function values of the current path plan, judging whether the optimal balance point is reached, and if the optimal balance point is not reached, adjusting the reward function weight coefficient by the gradient descent method, inputting the updated weight coefficient into the path planning model to regenerate the path plan, and repeating the optimization and judgment steps until the final path plan is output includes:
[0033] Obtain the values of each objective function under the current path plan. If all objective function values are within the preset range, define the current point as the optimal balance point. The optimal balance point is used to determine whether the path plan meets the flight mission requirements and serves as a basis for whether to terminate the path optimization.
[0034] Calculate the deviation between each objective function value and the preset equilibrium point;
[0035] Based on the corresponding deviation value, the weight coefficient in the reward function is adjusted by the gradient descent method, and the reward function is updated;
[0036] The updated reward function is input into the path planning model for optimization calculation to generate a new path plan;
[0037] Obtain the objective function value set of the new path plan and make another balance point judgment;
[0038] If the new objective function value still does not reach the equilibrium point, the weight value correction and reward function update process are repeated until the optimal equilibrium point is reached.
[0039] When the objective function value of the path plan reaches the optimal balance point, the calculation is terminated and the final path planning plan is determined.
[0040] Furthermore, the steps of converting the final path plan into waypoint coordinates and speed instructions, and monitoring environmental changes in real time during the flight execution, and triggering the path replanning mechanism when it is detected that the obstacle movement exceeds the safety threshold, include:
[0041] Obtaining a final path plan, extracting waypoint coordinates and flight speed parameters therefrom, and generating a flight control instruction, wherein the flight control instruction includes the waypoint coordinates and the flight speed;
[0042] Collect environmental data in real time through sensors to obtain obstacle location information in the current flight environment;
[0043] According to the preset safety range threshold, it is determined whether the obstacle movement exceeds the safety range. If the obstacle movement exceeds the threshold, the path replanning mechanism is triggered to recalculate the flight path based on the current environmental data.
[0044] Determine new flight path parameters based on the re-planning results, and dynamically adjust flight parameters to adapt to the new flight path by continuously monitoring environmental changes;
[0045] The adjusted flight path parameters are converted into flight control instructions, and updated control instructions are generated to execute the flight mission.
[0046] The present invention also provides a UAV intelligent dynamic path guidance system based on deep reinforcement learning, comprising:
[0047] The feature extraction unit is used to obtain the three-dimensional terrain point cloud data and obstacle geometry information of the UAV flight area, construct the environment topology structure map, extract the local and global features of the environment topology structure map through a graph convolution network containing three layers of graph convolution layers, and generate a high-dimensional environment feature representation;
[0048] Function design unit, used to construct the state space of deep reinforcement learning based on high-dimensional environment feature vectors and design multi-objective reward functions;
[0049] A model initialization unit is used to initialize the deep reinforcement learning network in the path planning model using random parameters uniformly distributed in [-0.1, 0.1], and generate an initial path plan based on the input state space information;
[0050] A node optimization unit, used to optimize the path node sequence in the initial path plan by using a dynamic programming algorithm, and when the path length exceeds a dynamic threshold corresponding to the task type, readjust the node sequence and calculate each objective function value;
[0051] A solution adjustment unit, used to obtain the path length of the initial path solution, and when the path length exceeds the dynamic threshold corresponding to the current flight mission, adjust the path node sequence through a dynamic programming algorithm to generate a current path solution;
[0052] A judgment unit, used to obtain the values of each objective function of the current path plan, and judge whether an optimal balance point is reached, wherein the optimal balance point is a solution in which the energy consumption, time and safety indicators are all within the preset range of the task. If the optimal balance point is not reached, the weight coefficient of the reward function is adjusted by the gradient descent method, and the updated weight coefficient is input into the path planning model to regenerate the path plan, and the optimization and judgment steps are repeated until the final path plan is output;
[0053] The execution and monitoring unit is used to convert the final path plan into waypoint coordinates and speed instructions, and monitor environmental changes in real time during the flight execution. When it is detected that the obstacle movement exceeds the safety threshold, the path replanning mechanism is triggered.
[0054] The UAV intelligent dynamic path guidance method and system based on deep reinforcement learning provided by the present invention have the following beneficial effects:
[0055] (1) Use 3D terrain point cloud data and obstacle information to build a topological map, and combine it with a graph convolutional network to extract local and global features, providing accurate and comprehensive environmental information for path planning, ensuring that a feasible path is planned in a complex geographical environment;
[0056] (2) By constructing a deep reinforcement learning state space and multi-objective reward function containing multiple factors, energy consumption, time and safety are comprehensively considered to achieve a balance between various objectives during path planning, thereby improving the efficiency and economy of task execution;
[0057] (3) Using a dynamic programming algorithm, the path node sequence is dynamically adjusted according to the path length and task type, combined with an iterative optimization mechanism to continuously approach the optimal solution and adapt to different task requirements and environmental changes;
[0058] (4) Randomly initialize the parameters of the deep reinforcement learning network to broaden the search range of the solution space, avoid local optimality, and improve the model's optimization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a flow chart of a method for intelligent dynamic path guidance of a UAV based on deep reinforcement learning in one embodiment of the present invention;
[0060] Figure 2 It is a structural block diagram of a UAV intelligent dynamic path guidance system based on deep reinforcement learning in one embodiment of the present invention;
[0061] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0063] Reference Figure 1 , which is a flow chart of a method for intelligent dynamic path guidance of a UAV based on deep reinforcement learning proposed by the present invention, comprising the following steps:
[0064] S1, obtains the 3D terrain point cloud data and obstacle geometry information of the UAV flight area, constructs the environment topology structure map, extracts the local and global features of the environment topology structure map through a graph convolutional network containing 3 layers of graph convolutional layers, and generates a high-dimensional environment feature representation;
[0065] S2, constructing a state space of deep reinforcement learning based on a high-dimensional environmental feature vector, the state space includes the position coordinates of the drone, the environmental feature vector and the remaining energy state, and designing a multi-objective reward function, which is obtained by weighted summation of energy consumption index, time index and safety index;
[0066] S3, using random parameters uniformly distributed in [-0.1, 0.1] to initialize the deep reinforcement learning network in the path planning model, and generating an initial path plan according to the input state space information, wherein the path planning model performs path planning based on the deep reinforcement learning state space including the position coordinates of the drone, the environmental feature vector, and the remaining energy state;
[0067] S4, using a dynamic programming algorithm to optimize the path node sequence in the initial path plan, when the path length exceeds the dynamic threshold corresponding to the task type, readjust the node sequence and calculate each objective function value, the objective function value includes energy consumption, time and safety indicators
[0068] S5, obtaining the path length of the initial path plan, and when the path length exceeds the dynamic threshold corresponding to the current flight mission, adjusting the path node sequence through a dynamic programming algorithm to generate a current path plan;
[0069] S6, obtaining the values of each objective function of the current path plan, and judging whether the optimal balance point is reached, wherein the optimal balance point is a solution in which the energy consumption, time and safety indicators are all within the preset range of the task. If the optimal balance point is not reached, the weight coefficient of the reward function is adjusted by the gradient descent method, and the updated weight coefficient is input into the path planning model to regenerate the path plan, and the optimization and judgment steps are repeated until the final path plan is output;
[0070] S7 converts the final path plan into waypoint coordinates and speed instructions, and monitors environmental changes in real time during the flight. When it is detected that the obstacle movement exceeds the safety threshold, the path replanning mechanism is triggered.
[0071] As described in step S1 above, the flight area of the drone is scanned with a laser radar to obtain three-dimensional terrain point cloud data, and the geometric features and location information of obstacles are captured by visual sensors. The terrain elevation data and the spatial coordinates of obstacles are integrated to form an environmental topological structure diagram. In this topological structure diagram, nodes represent terrain or obstacle feature points, and edges represent the spatial connection relationship between nodes. Such a topological diagram can clearly show the terrain and obstacle distribution in the flight area. According to the node information and edge connection relationship of the environmental topological structure diagram, a network containing 3 layers of graph convolution layers is constructed, and the feature dimensions of each layer output are 256, 128, and 64 respectively. After each layer of graph convolution operation, the ReLU activation function and batch normalization layer are processed; the topological relationship is extracted through multiple layers of graph convolution layers to obtain low-dimensional local feature representation. The graph convolution layer can effectively process graph structure data and mine information between nodes and edges. The low-dimensional local features output by the first layer of graph convolution and the global features output by the third layer are channel-joined by using a jump connection method, and finally a high-dimensional environmental feature vector is generated. Skip connections help retain feature information at different levels, so that the generated high-dimensional environment feature vector can more comprehensively describe the flight environment.
[0072] As described in step S2 above, the current position coordinates of the drone are obtained, and then combined with the environmental features generated in step S1 and the energy state of the drone to construct the state space of deep reinforcement learning. This state space contains the location of the drone, the surrounding environment information and the remaining energy situation, which can provide rich information for path planning. Using the deep reinforcement learning algorithm, a multi-objective reward function is defined in the state space. The reward function includes three indicators: energy consumption, time and safety. The energy consumption index is calculated based on the power consumption per unit distance, the time index is calculated based on the expected flight time, and the safety index is calculated based on the obstacle distance. Finally, the three indicators are weighted and summed to obtain the final reward function. By setting the weights reasonably, the drone can be guided to comprehensively consider energy consumption, time and safety when planning the path. In the process of reinforcement learning, the flight strategy of the drone is optimized through the reward function to balance energy consumption, time and safety. If the energy consumption index exceeds the preset threshold, the flight speed is adjusted to reduce energy consumption; if the time index exceeds the preset threshold, the flight path is optimized to shorten the flight time; if the safety index is lower than the preset threshold, the flight path is replanned to avoid obstacles.
[0073] Defining energy consumption indicators Where E current is the cumulative power consumption of the current path, E total is the total power of the drone, k 1 is the energy consumption weight coefficient; define the time index Where T current is the estimated flight time of the current path, T maxis the maximum allowed time for the task, k 2 is the time weight coefficient; define the safety index where d i is the Euclidean distance between the path node and the nearest obstacle, k 3 is the safety weight coefficient; the weighted sum of the above indicators is used to obtain the total reward function R total =R energy +R time +R safety .
[0074] As described in step S3 above, the weight parameters of the deep reinforcement learning network in the path planning model are set to be uniformly distributed in the range of [-0.1, 0.1]. This random initialization method can give the network different parameter starting points at the beginning of training, which helps to explore a wider solution space. At the same time, the curvature of the network output path is limited to not exceed the maximum steering angle θ of the drone. max . To ensure that the generated path conforms to the flight physical characteristics of the drone, the drone cannot achieve excessive turning curvature during actual flight, so this restriction is used to ensure the feasibility of the path. The constructed state space information containing the drone position coordinates, high-dimensional environmental feature vectors, and residual energy state is input into the initialized deep reinforcement learning network. The network generates an initial path plan based on the input information, which is presented in the form of a path node sequence P = {P1, P2, ..., Pn}. Each node P i It represents a location point of the UAV on the path. These nodes are connected in sequence to form the flight path of the UAV.
[0075] Calculate the total length L of the generated path and the energy consumption E. The total length L of the path reflects the distance the drone flies from the starting point to the end point, and the energy consumption E represents the energy required for the drone to fly along the path. According to the preset maximum steering angle and energy constraints of the drone, the initial path plan is comprehensively evaluated. On the one hand, check whether the steering between each path node is within the range that the drone can achieve, and on the other hand, determine whether L≤L max And E≤E max , where L max is the preset maximum path length, E max is the preset maximum energy consumption. These two conditions constrain the path plan from the perspective of path length and energy consumption, respectively, to ensure that the UAV can complete the flight mission within an acceptable distance and energy range. If the initial path plan satisfies the preset maximum steering angle and path length L≤L max and energy consumption E≤E maxIf the constraints are not met, the initial path plan is accepted. This means that the path plan meets the mission requirements in terms of flight feasibility, distance and energy, and can be used as the basis for subsequent path optimization. If the above constraints are not met, the network parameters are reinitialized, that is, the weight parameters of the deep reinforcement learning network are set to be uniformly distributed in the range of [-0.1, 0.1] again, and a new path node sequence is regenerated. This process is repeated until an initial path plan that meets all constraints is obtained.
[0076] As described in step S4 above, the path node sequence in the initial path plan is optimized using a dynamic programming algorithm. When the path length exceeds the dynamic threshold corresponding to the task type, the node sequence is readjusted. This dynamic threshold is calculated based on the task type. threshold : If it is a reconnaissance mission, L threshold =1.2*L shortest ; If it is a material transportation task, L threshold =1.5*L shortest ; where L shortest is the theoretical shortest path length from the starting point to the end point; when the path length L>L threshold When , the dynamic programming algorithm is used to compress the path node sequence and delete redundant nodes; the optimal balance point is defined as satisfying R energy ≥R energy-min , R time ≥R time-min , R salety ≥R safety-min The path plan, where R energy-min , R time-min , R safety-min Preset according to task type.
[0077] After adjusting the node sequence, calculate the values of each objective function, including energy consumption, time and safety indicators, so as to evaluate the quality of the path later. Use the GPU parallel computing framework to simultaneously perform the forward reasoning and back propagation tasks of the graph neural network and deep reinforcement learning network to improve computing efficiency. If the single iteration calculation time exceeds the threshold T compute >T threshold (Tt hreshold If the network performance is set to 100ms, the network structure is adjusted, such as reducing the number of graph convolutional network layers to 2, reducing the number of hidden layer neurons in the deep reinforcement learning network, and recalculating the adjusted network performance. compute ≤T threshold , the current network structure is locked to ensure a balance between computing efficiency and network performance.
[0078] As described in step S5 above, the corresponding dynamic threshold is calculated according to the current flight mission, and the path length from the flight start point to the end point in the initial path plan is obtained at the same time. The objective function values (energy consumption, time and safety index) of the initial path plan are calculated to determine whether the path length exceeds the dynamic threshold of the current flight mission. If the path length exceeds the dynamic threshold, the path node sequence is compressed by the dynamic programming algorithm, the redundant nodes are deleted, and the optimized current path plan is generated.
[0079] As described in step S6 above, the objective function values under the current path plan are obtained. If all objective function values are within the preset range, the current point is defined as the optimal balance point. This optimal balance point is the basis for judging whether the path plan meets the flight mission requirements and is also a sign for terminating the path optimization. Calculate the deviation between the current objective function value and the preset balance point △ R=|R current -R target |, where R current is the objective function value of the current path plan, R target is the objective function value corresponding to the preset balance point. The deviation reflects the gap between the current path plan and the optimal balance point. threshold If △R>△R threshold , then the weight coefficient of the reward function needs to be adjusted. Update the weight coefficient, where α is the learning rate, and i∈{1,2,3} corresponds to the weights of energy consumption, time, and safety respectively. The learning rate α controls the step size of each weight adjustment. By continuously adjusting the weight coefficient, the objective function value gradually approaches the preset equilibrium point. The updated reward function is input into the path planning model for optimization calculation to generate a new path plan.
[0080] Get the objective function value set of the new path plan, calculate the deviation △R again, and determine whether △R≤△R threshold Or whether the maximum number of iterations, 100, has been reached. If the △R corresponding to the new objective function value is still greater than △R threshold If the maximum number of iterations is not reached, the weight value correction and reward function update process are repeated until △R≤△R threshold Or the maximum number of iterations reaches 100. When the termination condition is met, the calculation is terminated and the final path planning solution is determined.
[0081] As described in step S7 above, the final path plan is obtained, the waypoint coordinates and flight speed parameters are extracted, and a flight control instruction is generated. The instruction includes the waypoint coordinates and the flight speed, which is used to control the flight of the drone. Environmental data is collected in real time through sensors to obtain obstacle location information in the current flight environment. According to the preset safety range threshold, it is determined whether the obstacle movement exceeds the safety range. If the obstacle movement exceeds the threshold, the path replanning mechanism is triggered to recalculate the flight path based on the current environmental data. According to the replanning results, the new flight path parameters are determined, and the flight parameters are dynamically adjusted to adapt to the new flight path by continuously monitoring environmental changes. The adjusted flight path parameters are converted into flight control instructions, updated control instructions are generated, and flight missions are executed to ensure that the drone can adapt to environmental changes during flight and complete the mission safely.
[0082] Among them, the path planning model adopts the deep deterministic policy gradient (DDPG) algorithm in deep reinforcement learning as the basic architecture. The DDPG algorithm combines deterministic policy gradient with deep neural network, which is particularly suitable for processing path planning problems in continuous action space. In the UAV path planning scenario, flight direction and speed are continuous control quantities. The DDPG algorithm can effectively make decisions on these continuous control quantities and generate reasonable flight instructions for the UAV. The model takes the deep reinforcement learning state space constructed based on the high-dimensional environmental feature vector as input. This state space contains the UAV position coordinates, environmental feature vectors and residual energy state. These rich input information can comprehensively describe the current flight environment and its own state of the UAV. The input information is transmitted to the model through the fully connected layer. The number of neurons in the fully connected layer will be reasonably set according to the dimension of the input feature, in order to ensure that the model can accurately process the input information and provide an accurate data basis for subsequent decision-making.
[0083] The policy network consists of multiple hidden layers, each of which uses a rectified linear unit (ReLU) as an activation function. The ReLU activation function can introduce nonlinear features, allowing the model to learn more complex mapping relationships, thereby improving the model's expressiveness. The output layer of the policy network is a continuous action space, which outputs control instructions such as the flight direction and speed of the drone. Through the policy network, the model can generate corresponding action decisions based on the current state information to guide the flight of the drone. The value network also consists of multiple hidden layers, and its input includes the state space and the actions output by the policy network. The main function of the value network is to evaluate the value of performing a specific action in the current state. During the training process, the mean square error loss function is used to optimize the model's estimate of the action value by minimizing the mean square error between the predicted value and the true value, so that the model can more accurately judge the pros and cons of different actions.
[0084] In order to improve the training efficiency and stability of the model, the experience replay mechanism is introduced. During the flight of the drone, information such as state, action, reward, and next state will be generated, and this information will be stored in the experience replay pool. During training, a batch of data is randomly sampled from the experience replay pool for learning. This can avoid the correlation between data from having a negative impact on the training effect, allowing the model to make more full use of historical data for learning and improve the generalization ability of training.
[0085] When the deep reinforcement learning network is initialized to generate the initial path plan, the environmental feature representation is input into the policy network of the path planning model. The policy network outputs a series of action decisions based on the current state information. These action decisions correspond to the control instructions such as the flight direction and speed of the drone, thereby generating an initial path plan. This initial path plan provides the basis for subsequent path optimization. After adjusting the weight coefficient of the reward function through the gradient descent method, the updated reward function is input into the path planning model. The model will regenerate action decisions through the policy network based on the new reward function and the current state information. These new action decisions will optimize the path and generate a new path plan. In this process, the model will continuously adjust the path according to the new reward function and state information to seek a better flight path. After obtaining the set of objective function values of the new path plan, it is determined whether the optimal balance point is reached. The optimal balance point means that the energy consumption, time and safety indicators are all within the preset range of the task. If the optimal balance point is not reached, the above weight value correction and reward function update process is repeated, and the updated reward function is continuously input into the path planning model for optimization calculation. Through this iterative process, the model will gradually find a path plan that meets the optimal balance point conditions and ultimately determine the final path planning plan.
[0086] In one embodiment, path planning is performed through a forest fire reconnaissance mission. The laser radar carried by the drone is used to scan the terrain of the 100×100m forest fire area, and a point cloud map with elevation data (accuracy ±0.4m) is generated to accurately present the terrain undulations. With the help of binocular vision to identify the fire area and the smoke diffusion area (detection accuracy>93%), a 3D environmental model with 300 topological map nodes is constructed to fully reflect the fire scene environment. A 3-layer graph convolutional network (GCN) (output dimension of each layer 64→128→256) is used to extract features, and a 512-dimensional environmental feature vector is generated through jump connections to provide rich environmental information for subsequent path planning.
[0087] Define the state space, including the drone GPS coordinates (WGS84 format), remaining power (percentage), and environmental feature vectors, to fully describe the current state of the drone. Set the initial value of the reward function weight, energy consumption coefficient α = 0.3, time coefficient β = 0.3, safety factor γ = 0.3, and balance the importance of each goal. Construct the DDPG network structure, the policy network is 3×128 fully connected layers, the value network is 2×256 fully connected layers, and the experience pool capacity is 1*10 5 , ensuring the network's learning and decision-making capabilities. The initial path length is 120m (dynamic threshold 100m), triggering dynamic planning node compression, deleting 5 redundant turning nodes, shortening the path to 90m, and improving path efficiency. The objective function deviation shows that the energy consumption exceeds the standard by 15%. The weights α=0.35, β=0.25, γ=0.4 are adjusted by gradient descent to regenerate the path and optimize energy consumption performance. The fire spread speed in the fire area is monitored in real time>3m / s (safety threshold 2.5m / s), and the local path replanning is completed quickly within 45ms to ensure the safety of drone reconnaissance.
[0088] In another embodiment, dynamic obstacle avoidance is performed for port logistics drones. The port building and facility model is constructed through VSLAM (accuracy 0.05m), and the ship and vehicle location data pushed by the port dispatching system are received in real time (update frequency 20Hz) to obtain accurate dynamic environmental information. The topological map nodes cover building contour points (static nodes) and mobile ships and vehicle predicted trajectory points (dynamic nodes), accurately reflecting the complex dynamic environment of the port. It is detected that a large cargo ship suddenly changes its berthing position (lateral displacement>4m / s), and the safety mechanism is triggered quickly. GPU parallel computing (NVIDIAA100) is used to efficiently complete within 10ms, environmental feature re-extraction (GCN reasoning takes 5ms), quickly update environmental information, DRL strategy network generates new waypoints (taking 4ms), and timely adjust the flight path. After the update, the path waypoint spacing is adjusted from 8m to 5m, which significantly improves the obstacle avoidance accuracy and adapts to the complex environment of the port. The energy consumption of the original path was 2000J (distance 2km). After optimization, the speed command was adjusted from 18m / s to 16m / s, the flight speed was reasonably controlled, and 4 sharp turns were reduced through dynamic planning (the curvature was reduced from 0.4 to 0.25), which reduced the flight energy consumption. The final energy consumption was 1700J (reduced by 15%) and the time delay was controlled within 4%, achieving a good balance between energy consumption and time.
[0089] Reference Figure 2 , is a structural block diagram of a UAV intelligent dynamic path guidance system based on deep reinforcement learning in one embodiment of the present invention, including:
[0090] The feature extraction unit is used to obtain the three-dimensional terrain point cloud data and obstacle geometry information of the UAV flight area, construct the environment topology structure map, extract the local and global features of the environment topology structure map through a graph convolution network containing three layers of graph convolution layers, and generate a high-dimensional environment feature representation;
[0091] Function design unit, used to construct the state space of deep reinforcement learning based on high-dimensional environment feature vectors and design multi-objective reward functions;
[0092] A model initialization unit is used to initialize the deep reinforcement learning network in the path planning model using random parameters uniformly distributed in [-0.1, 0.1], and generate an initial path plan based on the input state space information;
[0093] A node optimization unit, used to optimize the path node sequence in the initial path plan by using a dynamic programming algorithm, and when the path length exceeds a dynamic threshold corresponding to the task type, readjust the node sequence and calculate each objective function value;
[0094] A solution adjustment unit, used to obtain the path length of the initial path solution, and when the path length exceeds the dynamic threshold corresponding to the current flight mission, adjust the path node sequence through a dynamic programming algorithm to generate a current path solution;
[0095] A judgment unit, used to obtain the values of each objective function of the current path plan, and judge whether an optimal balance point is reached, wherein the optimal balance point is a solution in which the energy consumption, time and safety indicators are all within the preset range of the task. If the optimal balance point is not reached, the weight coefficient of the reward function is adjusted by the gradient descent method, and the updated weight coefficient is input into the path planning model to regenerate the path plan, and the optimization and judgment steps are repeated until the final path plan is output;
[0096] The execution and monitoring unit is used to convert the final path plan into waypoint coordinates and speed instructions, and monitor environmental changes in real time during the flight execution. When it is detected that the obstacle movement exceeds the safety threshold, the path replanning mechanism is triggered.
[0097] For the specific implementation of each unit in the above device example, please refer to the above method embodiment, which will not be repeated here.
[0098] In summary, the intelligent dynamic path guidance method for UAV based on deep reinforcement learning includes obtaining three-dimensional terrain point cloud data and obstacle geometry information of the UAV flight area, constructing an environmental topological structure map, extracting local and global features of the environmental topological structure map through a graph convolution network containing three layers of graph convolution layers, and generating a high-dimensional environmental feature representation; constructing a state space of deep reinforcement learning based on the high-dimensional environmental feature vector, and designing a multi-objective reward function; using random parameters uniformly distributed in [-0.1, 0.1] to initialize the deep reinforcement learning network in the path planning model, and generating an initial path plan according to the input state space information; using a dynamic programming algorithm to optimize the path node sequence in the initial path plan, and when the path length exceeds the dynamic threshold corresponding to the task type, readjust the node sequence and calculate the value of each objective function; obtaining the The path length of the initial path plan is obtained. When the path length exceeds the dynamic threshold corresponding to the current flight mission, the path node sequence is adjusted by the dynamic programming algorithm to generate the current path plan; the objective function values of the current path plan are obtained to determine whether the optimal balance point is reached; if the optimal balance point is not reached, the reward function weight coefficient is adjusted by the gradient descent method, the updated weight coefficient is input into the path planning model to regenerate the path plan, and the optimization and judgment steps are repeated until the final path plan is output; the final path plan is converted into waypoint coordinates and speed instructions, and the environmental changes are monitored in real time during the flight execution. When it is detected that the obstacle movement exceeds the safety threshold, the path re-planning mechanism is triggered to achieve the purpose of safe flight of the UAV in complex environments, optimize path planning, enhance model adaptability and improve autonomous decision-making capabilities.
[0099] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided by the present invention and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM.
[0100] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0101] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for intelligent dynamic path guidance of unmanned aerial vehicles based on deep reinforcement learning, characterized in that: The following steps are involved: Obtain the 3D terrain point cloud data and obstacle geometry information of the UAV flight area, construct the environment topology structure map, extract the local and global features of the environment topology structure map through a graph convolutional network containing 3 layers of graph convolutional layers, and generate a high-dimensional environment feature representation; Based on the high-dimensional environmental feature vector, a state space of deep reinforcement learning is constructed. The state space includes the position coordinates of the drone, the environmental feature vector, and the remaining energy state. A multi-objective reward function is designed. The reward function is obtained by weighted summation of energy consumption index, time index, and safety index. The deep reinforcement learning network in the path planning model is initialized with random parameters uniformly distributed in [-0.1, 0.1], and an initial path plan is generated according to the input state space information. The path planning model performs path planning based on the deep reinforcement learning state space including the drone position coordinates, the environmental feature vector and the remaining energy state; A dynamic programming algorithm is used to optimize the path node sequence in the initial path plan. When the path length exceeds the dynamic threshold corresponding to the task type, the node sequence is readjusted and each objective function value is calculated. The objective function value includes energy consumption, time and safety indicators. Acquire the path length of the initial path plan, and when the path length exceeds the dynamic threshold corresponding to the current flight mission, adjust the path node sequence through a dynamic programming algorithm to generate a current path plan; Obtaining the values of each objective function of the current path plan, and judging whether the optimal balance point is reached, wherein the optimal balance point is a solution in which the energy consumption, time and safety indicators are all within the preset range of the task; if the optimal balance point is not reached, adjusting the weight coefficient of the reward function by the gradient descent method, inputting the updated weight coefficient into the path planning model to regenerate the path plan, and repeating the optimization and judgment steps until the final path plan is output; The final path plan is converted into waypoint coordinates and speed instructions, and environmental changes are monitored in real time during the flight. When it is detected that the obstacle movement exceeds the safety threshold, the path replanning mechanism is triggered.
2. The method for intelligent dynamic path guidance of unmanned aerial vehicle based on deep reinforcement learning according to claim 1 is characterized in that: The steps of obtaining three-dimensional terrain point cloud data and obstacle geometry information of the UAV flight area, constructing an environment topology structure map, extracting local and global features of the environment topology structure map through a graph convolution network including three graph convolution layers, and generating a high-dimensional environment feature representation include: The three-dimensional terrain point cloud data of the flight area is obtained by laser radar scanning, and the geometric features and location information of obstacles are captured by visual sensors. The terrain elevation data and the obstacle spatial coordinates are integrated into an environmental topological structure diagram, in which the nodes of the topological structure diagram represent the terrain or obstacle feature points, and the edges represent the spatial connection relationship between the nodes; Based on the node information and edge connection relationship of the environment topology graph, a network with three layers of graph convolution layers is constructed. Multi-layer graph convolution layers are used to extract features of topological relationships to obtain low-dimensional local feature representation. Skip connections are used to concatenate the low-dimensional local features output by the first layer of graph convolution with the global features output by the third layer to generate a high-dimensional environment feature vector.
3. The method for intelligent dynamic path guidance of unmanned aerial vehicle based on deep reinforcement learning according to claim 1 is characterized in that: The steps of constructing a state space of deep reinforcement learning based on a high-dimensional environment feature vector and designing a multi-objective reward function include: Get the current position coordinates of the drone, combine the environmental characteristics and energy state, and construct the state space; A deep reinforcement learning algorithm is used to define a multi-objective reward function in the state space, wherein the reward function includes three indicators: energy consumption, time, and safety. Among them, the energy consumption index is calculated based on the power consumption per unit distance, the time index is calculated based on the expected flight time, and the safety index is calculated based on the obstacle distance.
4. The method for intelligent dynamic path guidance of unmanned aerial vehicle based on deep reinforcement learning according to claim 1 is characterized in that: The step of using random parameters uniformly distributed in [-0.1, 0.1] to initialize the deep reinforcement learning network in the path planning model and generating an initial path plan according to the input state space information includes: The weight parameters of the deep reinforcement learning network are set to be evenly distributed in the range of [-0.1, 0.1], and the curvature of the network output path is limited to not exceed the maximum steering angle of the drone; The constructed state space information including the drone position coordinates, high-dimensional environmental feature vectors, and residual energy state is input into the initialized deep reinforcement learning network to generate an initial path plan in the form of a path node sequence; Calculate the total path length L and energy consumption E of the initial path plan, and evaluate whether the steering between each path node in the initial path plan is within the achievable range of the drone according to the maximum steering angle and energy constraint conditions preset by the drone, and determine whether L≤L max And E≤E max , wherein L max is the preset maximum path length, the E max is the preset maximum energy consumption; If the initial path plan satisfies the preset maximum steering angle, L≤L max And E≤E max If the constraints are met, the initial path plan is accepted.
5. The method for intelligent dynamic path guidance of unmanned aerial vehicle based on deep reinforcement learning according to claim 4 is characterized in that: The path planning model adopts a deep deterministic policy gradient algorithm, with a combination of policy gradient and deep neural network as the basic architecture. It takes the deep reinforcement learning state space including the drone position, environmental characteristics and residual energy as input, and after processing through the fully connected layer, it is input into the policy network composed of multiple hidden layers to generate flight control instructions for direction and speed. The path planning model also includes a value network for evaluating the value of actions, and introduces an experience replay mechanism.
6. The method for intelligent dynamic path guidance of unmanned aerial vehicle based on deep reinforcement learning according to claim 1, characterized in that: The step of obtaining the path length of the initial path plan and, when the path length exceeds a dynamic threshold corresponding to the current flight mission, adjusting the path node sequence by a dynamic programming algorithm to generate a current path plan comprises: Calculate the corresponding dynamic threshold according to the current flight mission, and obtain the path length from the flight start point to the end point in the initial path plan; Calculate the objective function values of the initial path plan, the energy consumption, time and safety indicators of the objective function values, and determine whether the path length exceeds the dynamic threshold of the current flight mission. If the path length exceeds the dynamic threshold, compress the path node sequence through a dynamic programming algorithm, delete redundant nodes, and generate an optimized current path plan.
7. The method for intelligent dynamic path guidance of unmanned aerial vehicle based on deep reinforcement learning according to claim 1, characterized in that: After the step of optimizing the path node sequence in the initial path plan by using a dynamic programming algorithm and re-adjusting the node sequence and calculating each objective function value when the path length exceeds the dynamic threshold corresponding to the task type, the method further includes: Using GPU parallel computing framework, it can simultaneously perform forward reasoning and back propagation tasks of graph neural network and deep reinforcement learning network; If the single iteration calculation time exceeds the threshold, the network structure is adjusted, and the network structure adjustment includes reducing the number of graph convolutional network layers to 2 layers, reducing the number of hidden layer neurons in the deep reinforcement learning network, and recalculating the adjusted network performance.
8. The method for intelligent dynamic path guidance of unmanned aerial vehicle based on deep reinforcement learning according to claim 1, characterized in that: The step of obtaining the objective function values of the current path plan, judging whether the optimal balance point is reached, and if the optimal balance point is not reached, adjusting the reward function weight coefficient by the gradient descent method, inputting the updated weight coefficient into the path planning model to regenerate the path plan, and repeating the optimization and judgment steps until the final path plan is output includes: Obtain the values of each objective function under the current path plan. If all objective function values are within the preset range, define the current point as the optimal balance point. The optimal balance point is used to determine whether the path plan meets the flight mission requirements and serves as a basis for whether to terminate the path optimization. Calculate the deviation between each objective function value and the preset equilibrium point; Based on the corresponding deviation value, the weight coefficient in the reward function is adjusted by the gradient descent method, and the reward function is updated; The updated reward function is input into the path planning model for optimization calculation to generate a new path plan; Obtain the objective function value set of the new path plan and make another balance point judgment; If the new objective function value still does not reach the equilibrium point, the weight value correction and reward function update process are repeated until the optimal equilibrium point is reached. When the objective function value of the path plan reaches the optimal balance point, the calculation is terminated and the final path planning plan is determined.
9. The method for intelligent dynamic path guidance of unmanned aerial vehicle based on deep reinforcement learning according to claim 1, characterized in that: The steps of converting the final path plan into waypoint coordinates and speed instructions, and monitoring environmental changes in real time during the flight execution, and triggering the path replanning mechanism when it is detected that the obstacle movement exceeds the safety threshold, include: Obtaining a final path plan, extracting waypoint coordinates and flight speed parameters therefrom, and generating a flight control instruction, wherein the flight control instruction includes the waypoint coordinates and the flight speed; Collect environmental data in real time through sensors to obtain obstacle location information in the current flight environment; According to the preset safety range threshold, it is determined whether the obstacle movement exceeds the safety range. If the obstacle movement exceeds the threshold, the path replanning mechanism is triggered to recalculate the flight path based on the current environmental data. Determine new flight path parameters based on the re-planning results, and dynamically adjust flight parameters to adapt to the new flight path by continuously monitoring environmental changes; The adjusted flight path parameters are converted into flight control instructions, and updated control instructions are generated to execute the flight mission.
10. An intelligent dynamic path guidance system for unmanned aerial vehicles based on deep reinforcement learning, characterized in that: include: The feature extraction unit is used to obtain the three-dimensional terrain point cloud data and obstacle geometry information of the UAV flight area, construct the environment topology structure map, extract the local and global features of the environment topology structure map through a graph convolution network containing three layers of graph convolution layers, and generate a high-dimensional environment feature representation; Function design unit, used to construct the state space of deep reinforcement learning based on high-dimensional environment feature vectors and design multi-objective reward functions; A model initialization unit is used to initialize the deep reinforcement learning network in the path planning model using random parameters uniformly distributed in [-0.1, 0.1], and generate an initial path plan based on the input state space information; A node optimization unit, used to optimize the path node sequence in the initial path plan by using a dynamic programming algorithm, and when the path length exceeds a dynamic threshold corresponding to the task type, readjust the node sequence and calculate each objective function value; A solution adjustment unit, used to obtain the path length of the initial path solution, and when the path length exceeds the dynamic threshold corresponding to the current flight mission, adjust the path node sequence through a dynamic programming algorithm to generate a current path solution; A judgment unit, used to obtain the values of each objective function of the current path plan, and judge whether an optimal balance point is reached, wherein the optimal balance point is a solution in which the energy consumption, time and safety indicators are all within the preset range of the task. If the optimal balance point is not reached, the weight coefficient of the reward function is adjusted by the gradient descent method, and the updated weight coefficient is input into the path planning model to regenerate the path plan, and the optimization and judgment steps are repeated until the final path plan is output; The execution and monitoring unit is used to convert the final path plan into waypoint coordinates and speed instructions, and monitor environmental changes in real time during the flight execution. When it is detected that the obstacle movement exceeds the safety threshold, the path replanning mechanism is triggered.
Citation Information
Patent Citations
Unmanned aerial vehicle task allocation method considering operation environment and performance in collaborative search and rescue
CN112230675A
Cited By
Communication ad hoc network method and system applied to capital construction non-signal site
CN120378922A
Unmanned aerial vehicle path planning method and system based on reinforcement learning
CN120403661A
Unmanned aerial vehicle group fire cooperative surrounding control method based on MADDPG algorithm
CN120469481A
Encasement path planning method and system based on deep learning
CN120645233A
Unmanned aerial vehicle flight path adjustment method based on obstacle avoidance in non-visual state
CN120704367A