A UAV route planning method and system based on deep reinforcement learning

Through the UAV route planning method based on deep reinforcement learning, SLAM technology and deep learning algorithms are used to build a three-dimensional map and dynamically adjust the UAV path, which solves the problem of high-precision position recognition of traditional UAVs in complex environments and realizes high-precision navigation and stable flight of UAVs in complex environments.

CN120178933BActive Publication Date: 2025-09-23SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510351020.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-09-23
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Traditional drone navigation systems have difficulty ensuring high-precision location recognition in complex and changing environments, resulting in location recognition errors and affecting flight safety.

Method used

A drone route planning method based on deep reinforcement learning is adopted. By collecting drone status parameters and comprehensive navigation situation data, SLAM technology is used to build a three-dimensional map. The obstacle return path is calculated by combining deep learning and quantum algorithms to dynamically adjust the flight route.

Benefits of technology

It realizes high-precision navigation and autonomous navigation capabilities of UAVs in complex environments, ensures accurate monitoring and dynamic adjustment of flight paths, and enhances the stable operation capabilities of UAVs in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178933B_ABST
    Figure CN120178933B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for UAV route planning based on deep reinforcement learning, which relates to the field of intelligent control technology. The method comprises collecting the current state parameters and comprehensive navigation situation data of the UAV; calculating the UAV's route planning deviation rate based on the collected current state parameters of the UAV; processing the comprehensive navigation situation data using a deep learning algorithm to identify the UAV's current precise position; constructing a three-dimensional map model using SLAM technology; and generating a three-dimensional map using a point cloud fusion method; constructing a deep reinforcement learning model based on the UAV's current precise position, current state parameters, and the three-dimensional map; and calculating the UAV's return path when facing obstacles using a quantum algorithm; and dynamically adjusting the UAV's return path based on the UAV's route planning deviation rate to generate the UAV's route plan. Through multi-source data fusion and advanced algorithm processing, the present invention enables UAVs to operate stably in various environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent control technology, and in particular to a method and system for unmanned aerial vehicle (UAV) route planning based on deep reinforcement learning. Background Art

[0002] With the rapid development of drone technology, its application in logistics distribution, agricultural monitoring and disaster relief is becoming increasingly extensive. Traditional drone navigation systems mainly rely on GPS and inertial navigation. Although they can provide reliable positioning services in most cases, in complex environments, existing technologies still face many challenges in dealing with efficient path planning in complex three-dimensional environments, especially in high-precision location recognition and rapid response.

[0003] In the field of intelligent control technology, traditional methods struggle to ensure high-precision drone position recognition in complex and ever-changing environments. Due to the diverse and rapidly changing nature of obstacles, traditional technologies may not be able to update map information in a timely manner, leading to position recognition errors during drone missions and compromising flight safety. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a UAV route planning method and system based on deep reinforcement learning to solve the problem that traditional technologies are difficult to ensure high-precision position recognition of UAVs in complex and changing environments.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a method and system for UAV route planning based on deep reinforcement learning, which includes collecting the current state parameters and comprehensive navigation situation data of the UAV;

[0008] Calculate the deviation rate of the drone's route planning based on the collected drone's current state parameters;

[0009] Use deep learning algorithms to process comprehensive navigation situation data, identify the current precise location of the drone, use SLAM technology to build a three-dimensional map model, and use point cloud fusion method to generate a three-dimensional map;

[0010] Based on the drone's current precise location, current state parameters, and 3D map, a deep reinforcement learning model is built, using quantum algorithms to calculate the drone's return path when facing obstacles.

[0011] Based on the deviation rate of the UAV’s route planning, the return path of the UAV is dynamically adjusted to generate the UAV route planning.

[0012] As a preferred solution of the UAV route planning method based on deep reinforcement learning described in the present invention, collecting the current state parameters and comprehensive navigation situation data of the UAV includes the following steps:

[0013] The current state parameters of the UAV include speed and attitude, and the comprehensive navigation situation data includes the current position information, target position information, flight status and environmental perception data of the UAV;

[0014] The environmental perception data includes inertial measurement unit data, image data and lidar data.

[0015] As a preferred solution of the UAV route planning method based on deep reinforcement learning described in the present invention, wherein: based on the collected current state parameters of the UAV, calculating the UAV route planning deviation rate includes the following steps:

[0016] The UAV's flight mission is set based on its current location information and target location information. The UAV's current state parameters are combined with the UAV's flight mission, and the ideal flight path of the UAV is obtained using a fast exploration random tree path optimization algorithm.

[0017] The current location information of the drone is recorded at fixed time intervals to form sampling points;

[0018] Determine the actual position of the drone using sensor fusion technology using inertial measurement unit data and lidar data;

[0019] The corresponding points on the ideal path and the actual position of the UAV are found through sampling points, and the ball tree algorithm is used to calculate the UAV's route planning deviation rate.

[0020] As a preferred solution of the UAV route planning method based on deep reinforcement learning described in the present invention, the process of using a deep learning algorithm to process the comprehensive navigation situation data and identify the current precise location of the UAV includes the following steps:

[0021] Use simultaneous localization and mapping algorithms to extract corner feature points from image data;

[0022] Use the fast library approximate nearest neighbor search method to match the current corner feature points with the previously extracted corner feature points, use the geometric matrix to calculate the relative displacement and rotation matrix, and obtain the matched feature point set and its corresponding transformation matrix;

[0023] Based on the transformation matrix corresponding to the matched feature point set, a posture network model is constructed and the image data is processed sequentially to obtain the initial position and posture of the UAV;

[0024] Based on the preliminary position and attitude of the UAV and combined with the inertial measurement unit data, the extended Kalman filter is used to fuse the deep learning model to calculate the precise position of the UAV.

[0025] As a preferred solution of the UAV route planning method based on deep reinforcement learning described in the present invention, wherein: using SLAM technology to build a three-dimensional map model, and using point cloud fusion method to generate a three-dimensional map includes the following steps:

[0026] The system uses LiDAR sensors to obtain 3D information about the surrounding environment, uses SLAM technology to build a 3D map model, and extracts different point cloud data from the surrounding environment.

[0027] Use the normal distribution transformation algorithm to fuse different point cloud data to generate a complete three-dimensional point cloud map;

[0028] Convert the complete 3D point cloud into a 3D mesh using the marching cubes algorithm;

[0029] Texture mapping is performed on the complete 3D mesh to obtain a preliminary 3D map;

[0030] The preliminary three-dimensional map is smoothed by point cloud fusion method to obtain a complete three-dimensional map.

[0031] As a preferred solution of the UAV route planning method based on deep reinforcement learning described in the present invention, wherein: based on the current precise position of the UAV, the current state parameters and the three-dimensional map, a deep reinforcement learning model is constructed, and a quantum algorithm is used to calculate the return path of the UAV when facing an obstacle, including the following steps:

[0032] Use point cloud processing technology in the complete 3D map to obtain the location of obstacles and the height of the drone above the ground;

[0033] The state space of the drone is defined by its speed, attitude and target position;

[0034] Define the drone’s action space through its flight state;

[0035] Build a deep Q network based on the drone's state space and action space and define the loss function;

[0036] Predict the drone’s action strategy in each state through the loss function;

[0037] Convert the complete 3D map data into a quantum-computing 3D map structure and apply a cost function to optimize the drone’s actions;

[0038] Based on the optimized UAV action strategy, a quantum approximate optimization algorithm is applied to obtain the return path of the UAV when facing obstacles.

[0039] As a preferred solution of the UAV route planning method based on deep reinforcement learning described in the present invention, wherein: based on the UAV route planning deviation rate, the return path of the UAV is dynamically adjusted, and the generation of the UAV route planning includes the following steps:

[0040] Based on the current position of the UAV and the position of the obstacle, the deviation rate between the current path and the actual flight path is obtained through the iterative closest point algorithm;

[0041] When the current position of the UAV and the position of the obstacle are less than the deviation rate of the UAV's route planning, the UAV continues to fly along the current path;

[0042] When the current position of the UAV and the position of the obstacle are greater than or equal to the deviation rate of the UAV's route planning, the greedy algorithm is used to optimize the deviation part to obtain the UAV route planning.

[0043] In a second aspect, the present invention provides a UAV route planning system based on deep reinforcement learning, comprising a data acquisition module for collecting comprehensive navigation situation data of the current state parameters of the UAV;

[0044] The deviation rate module calculates the deviation rate of the drone's route planning based on the collected drone's current state parameters;

[0045] The 3D map module uses deep learning algorithms to process comprehensive navigation situation data, identify the current precise location of the drone, use SLAM technology to build a 3D map model, and use point cloud fusion method to generate a 3D map;

[0046] The path return module builds a deep reinforcement learning model based on the drone's current precise position, current state parameters, and 3D map, and uses quantum algorithms to calculate the drone's return path when facing obstacles;

[0047] The path planning module dynamically adjusts the return path of the drone based on the drone's route planning deviation rate and generates the drone's route planning.

[0048] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the drone route planning method based on deep reinforcement learning as described in the first aspect of the present invention is implemented.

[0049] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the drone route planning method based on deep reinforcement learning as described in the first aspect of the present invention.

[0050] The beneficial effects of the present invention are as follows: by calculating the route planning deviation rate of the UAV, accurate monitoring and dynamic adjustment of the flight path are achieved, the flight mission is set based on the current position and target position of the UAV, and the RT path optimization algorithm is used to plan the ideal path, the actual position is recorded at fixed time intervals to form sampling points, the actual position is determined by combining inertial measurement unit data and lidar data, the ball tree algorithm is used to calculate the deviation rate to ensure high-precision navigation even in complex environments, corner feature points are extracted from visual data and matched through a fast library approximate nearest neighbor search method, a posture network model is constructed to obtain the preliminary position and posture, the extended Kalman filter is used to fuse the deep learning model, the precise position of the UAV is calculated, and the autonomous navigation capability of the UAV in dynamic and complex environments is enhanced. Through multi-source data fusion and advanced algorithm processing, the UAV can operate stably in various environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1 This is a flowchart of the drone route planning method based on deep reinforcement learning in Example 1.

[0053] Figure 2 This is a module diagram of the drone route planning system based on deep reinforcement learning in Example 1. DETAILED DESCRIPTION

[0054] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0055] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0056] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0057] Example 1, with reference to Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides a method and system for drone route planning based on deep reinforcement learning, including the following steps:

[0058] S1. Collect the current state parameters and comprehensive navigation situation data of the UAV.

[0059] S1.1. The current state parameters of the UAV include speed and attitude. The comprehensive navigation situation data includes the UAV's current position information, target position information, flight status and environmental perception data.

[0060] Furthermore, the speed (linear velocity and angular velocity) and attitude (pitch, roll and yaw angles) of the drone are integrated with the current position information and target position information of the drone to ensure the understanding of the specific geographical location and final destination of the drone, as well as the flight status (such as flight mode and battery power), so as to facilitate the real-time assessment of the health status and mission execution capability of the drone, and finally collect environmental perception data.

[0061] S1.2. Environmental perception data includes inertial measurement unit data, image data, and lidar data.

[0062] Furthermore, the accelerometer and gyroscope data provided by the inertial measurement unit are used to monitor the acceleration and rotation rate of the drone, thereby accurately calculating its attitude changes. The image data captured by the camera not only contains rich visual information, but can also be used for advanced functions such as target recognition and obstacle detection. In particular, by processing image data through algorithms such as simultaneous positioning and map construction, a high-precision estimation of the drone's position can be achieved. Finally, the three-dimensional point cloud data generated by the lidar provides detailed information about the surrounding environment structure, which helps to build accurate map models and avoid obstacles.

[0063] S2. Calculate the UAV's route planning deviation rate based on the collected UAV current state parameters.

[0064] S2.1. Set the flight mission of the UAV according to the current position information and target position information of the UAV, combine the current state parameters of the UAV with the flight mission of the UAV, and use the fast exploration random tree path optimization algorithm to obtain the ideal flight path of the UAV.

[0065] Furthermore, a specific flight mission is set based on the current position and target position information of the UAV. Then, combined with the current state parameters of the UAV's speed and attitude, the RRT algorithm is used to find a feasible path in the UAV's flight environment through random sampling and gradual expansion to cope with dynamic changes and dense obstacles. This method not only improves the flexibility of path planning, but also enhances the UAV's ability to respond to emergencies.

[0066] S2.2. Record the current location information of the drone at fixed time intervals to form sampling points.

[0067] The actual position of the drone is determined using sensor fusion technology using inertial measurement unit data and lidar data.

[0068] Furthermore, the current position information of the UAV is recorded at fixed time intervals (for example, once per second) to form a series of continuous sampling points. The sampling points are used to calculate the subsequent path deviation rate. The actual position of the UAV is determined by using inertial measurement unit data and lidar data through sensor fusion technology (such as particle filters). This multi-source data fusion method can effectively improve positioning accuracy, overcome the limitations of a single sensor, and ensure the reliability and accuracy of the UAV's position.

[0069] S2.3. Find the corresponding point on the ideal path and the actual position of the UAV through the sampling points, and use the Ball tree algorithm to calculate the UAV's route planning deviation rate. The expression is:

[0070]

[0071] in, The deviation rate of the drone’s route planning, is the number of sampling points, No. The distance difference between the actual position of each sampling point and the corresponding position on the ideal path, For the sampling point.

[0072] Furthermore, the sampling points are used to find the corresponding points on the ideal path that are closest to the actual position of the drone. Specifically, for each sampling point, the Ball tree algorithm is used to efficiently search for its nearest neighbor on the ideal path and calculate the distance difference between the two. , calculate the average distance difference of all sampling points to get the UAV route planning deviation rate The deviation rate reflects the degree to which the UAV deviates from the preset path during flight, providing a quantitative basis for subsequent adjustments to the flight path, ensuring that the UAV can correct deviations in a timely manner and maintain high-precision navigation performance.

[0073] S3. Use deep learning algorithms to process the comprehensive navigation situation data and identify the current precise location of the drone.

[0074] S3.1. Use the simultaneous localization and mapping algorithm to extract corner feature points from image data.

[0075] Furthermore, using a simultaneous localization and mapping algorithm to extract corner feature points from image data collected by the camera can provide stable visual features in complex environments.

[0076] S3.2. Use the fast library approximate nearest neighbor search method to match the current corner feature point with the previously extracted corner feature point, use the geometric matrix to calculate the relative displacement and rotation matrix, and obtain the matched feature point set and its corresponding transformation matrix.

[0077] Furthermore, a fast library approximate nearest neighbor search method is used to match the corner feature points in the current frame with the corner feature points extracted in the previous frame. The fast library approximate nearest neighbor search method can efficiently find each pair of closest feature points to establish a correspondence, and use geometric matrices (such as basic matrices or essential matrices) to calculate relative displacement and rotation matrices to obtain a reliable set of matching feature points and their corresponding transformation matrices.

[0078] S3.3. Based on the transformation matrix corresponding to the matched feature point set, a posture network model is constructed and the image data is processed sequentially to obtain the preliminary position and posture of the UAV.

[0079] Furthermore, a posture network model is constructed based on the transformation matrix corresponding to the matched feature point set. The posture network is a deep learning model specifically used to estimate the posture (i.e., position and orientation) of the camera from an image sequence. The image data is input into the posture network model, and high-level features are extracted through a convolutional neural network. The posture parameters corresponding to each frame of the image are predicted, and the preliminary position and posture information of the drone in consecutive frames can be obtained.

[0080] S3.4. Based on the preliminary position and attitude of the UAV and combined with the inertial measurement unit data, the extended Kalman filter is used to fuse the deep learning model to calculate the precise position of the UAV.

[0081] Furthermore, the preliminary position and attitude of the UAV are combined with the inertial measurement unit data, and the extended Kalman filter is used for multi-sensor data fusion. The inertial measurement unit data provides high-frequency but potentially drifting acceleration and angular velocity information, while the position and attitude information provided by the attitude network model has a lower frequency but higher accuracy. The EKF can fuse the advantages of these two data sources, dynamically adjust the state estimation, compensate for their respective errors, and ultimately obtain the precise position of the UAV.

[0082] S4. Use SLAM technology to build a three-dimensional map model and use point cloud fusion method to generate a three-dimensional map.

[0083] S4.1. Obtain three-dimensional information of the surrounding environment through the lidar sensor, use SLAM technology to build a three-dimensional map model, and extract different point cloud data from the surrounding environment.

[0084] Furthermore, lidar sensors are used to perform high-precision scans of the drone's surroundings, collecting large amounts of distance measurement data. This data represents the position of objects on the surface in the form of points, forming a point cloud dataset. SLAM algorithms are then applied to process this point cloud data in real time, simultaneously localizing the drone and building a map. The drone's position estimate is continuously updated, and the new point cloud data is matched and integrated with existing map data.

[0085] S4.2. Use the normal distribution transformation algorithm to fuse different point cloud data to generate a complete three-dimensional point cloud image.

[0086] Furthermore, the normal distribution transformation algorithm divides the point cloud data into multiple small cells and calculates a normal distribution model for each cell. Through iterative optimization, it minimizes the differences between different point clouds and finds the optimal transformation matrix to maximize the overlap between the two point cloud data sets. After multiple iterations and optimizations, the normal distribution transformation algorithm can generate a complete three-dimensional point cloud image.

[0087] S4.3. Use the marching cubes algorithm to convert the complete 3D point cloud into a 3D mesh.

[0088] Furthermore, after obtaining a complete three-dimensional point cloud map, the marching cube algorithm is applied to traverse each voxel in the point cloud data (i.e., a small cube in three-dimensional space) and generate a polygonal mesh based on these surfaces. The marching cube algorithm defines 8 vertices in each voxel and checks whether each vertex is above or below the surface. The position of the surface inside the voxel can be inferred and the corresponding triangular facets can be generated. The marching cube algorithm generates a three-dimensional mesh composed of triangular facets.

[0089] S4.4. Perform texture mapping on the complete three-dimensional mesh to obtain a preliminary three-dimensional map.

[0090] Furthermore, texture information is extracted from the image captured by the camera and mapped onto the three-dimensional mesh surface. Feature points in the image are identified and matched with corresponding vertices in the three-dimensional mesh. The pixel values ​​on the two-dimensional image are mapped to the three-dimensional mesh through projection transformation to ensure that each facet can correctly display its corresponding texture. The visual effect can be further enhanced through lighting model and shadow processing. Through texture mapping, a preliminary three-dimensional map with rich details and realism is generated.

[0091] S4.5. Smooth the preliminary three-dimensional map using the point cloud fusion method to obtain a complete three-dimensional map.

[0092] Furthermore, point cloud fusion technology is used to filter out noise in the preliminary map, and bilateral filters or Gaussian filters are applied to smooth the point cloud data to remove isolated points and outliers. The three-dimensional grid is optimized through surface reconstruction algorithms to ensure surface smoothness and continuity. Multi-scale analysis methods can be used to process the map at different resolution levels to further improve its fineness and consistency, and ultimately generate a complete three-dimensional map.

[0093] S5. Based on the drone’s current precise position, current state parameters, and three-dimensional map, a deep reinforcement learning model is constructed, and a quantum algorithm is used to calculate the drone’s return path when facing obstacles.

[0094] S5.1. Use point cloud processing technology in the complete 3D map to obtain the location of obstacles and the height information of the drone and the ground.

[0095] Furthermore, a voxel grid filter is applied to reduce the amount of point cloud data, and then a region growing algorithm or Euclidean clustering extraction algorithm is used to segment different objects. The location of obstacles is determined based on these segmentation results, and the height of the drone relative to the ground is calculated using height information.

[0096] S5.2. Define the state space of the UAV by its velocity, attitude and target position;

[0097] Furthermore, based on the state parameters of the drone, such as speed, attitude, current position and target position, its state space is defined. Each state consists of a series of variables that describe all relevant information of the drone at a certain moment.

[0098] S5.3. Define the UAV's action space based on the UAV's flight state;

[0099] Furthermore, the action space of the drone is defined according to its flight state, including the basic actions of forward, backward, left turn, right turn, ascent and descent. Each action corresponds to a set of specific control instructions, changing the speed of the motor or adjusting the angle of the servo. In order to facilitate implementation and optimization, the action is usually discretized into a finite number of options. The action can be expressed as ∈{forward, backward, turn left, turn right, up, down}, the well-defined action space is helpful for training deep Q-networks.

[0100] S5.4. Build a deep Q network based on the state space and action space of the drone and define the loss function, which is expressed as:

[0101] ;

[0102] in, is the loss function, is the current state, In the current state The following actions are taken, To perform an action The reward value, To perform an action Next state, is the discount factor, is the optimal action in the next state, is the weight of the main network, The weights of the target network, For the next state Take the best action Maximum goals obtained value, Take action for the current state Predicted value, For the buffer zone, In the buffer zone Randomly drawn samples The mean of .

[0103] Furthermore, by randomly sampling from the experience replay buffer By calculating the difference between the predicted Q value and the target Q value, the network weights can be continuously adjusted to minimize the loss function, thereby optimizing the drone's action strategy in different states, ensuring that it can make the best decision based on environmental changes, and effectively improving the accuracy and adaptability of the drone's path planning.

[0104] S5.5. Predict the UAV’s action strategy in each state through the loss function.

[0105] Furthermore, the weights of the deep Q network are updated by minimizing the loss function, thereby continuously optimizing the predicted Q value. During each iteration, samples are randomly drawn from the experience buffer, the difference between the predicted Q value and the target Q value is calculated, and the network parameters are adjusted based on the difference. After multiple iterations, the network can more accurately predict the value of taking different actions in each state, thereby guiding the drone to choose the optimal action strategy.

[0106] S5.6 converts the complete three-dimensional map data into a three-dimensional map structure for quantum computing and applies the cost function to optimize the UAV’s action strategy, which is expressed as:

[0107] ;

[0108] in, is the cost function, For the The Pauli-Z matrix of the 3D map nodes, For the The Pauli-Z matrix of the identified nodes, For the The cost coefficient of each 3D map node, is the 3D map node index, To identify the node index, is the sum of the three-dimensional map nodes of the three-dimensional map structure, It is a 3D map structure and a 3D map node collection.

[0109] Furthermore, each node in the three-dimensional map is mapped, and the three-dimensional map is decomposed into the sum of the three-dimensional map nodes and the three-dimensional map node set, where the three-dimensional map node represents a spatial point or area, It represents the connection relationship between the identification node and the map node. Each 3D map node is described by the Pauli-Z matrix, thereby simulating whether the drone passes through the 3D map node.

[0110] S5.7 uses the quantum approximate optimization algorithm based on the optimized UAV action strategy to obtain the return path of the UAV when facing obstacles.

[0111] Furthermore, based on the optimized action strategy, a quantum approximate optimization algorithm is used to generate the optimal return path for the drone when facing obstacles. This allows the drone to explore a large number of possible paths in a short period of time and find a path that is both safe and can quickly reach the target location. This improves the efficiency of path planning and enhances the drone's ability to cope with complex environmental changes. The drone returns safely to its destination along this optimized path.

[0112] S6. Based on the deviation rate of the UAV’s route planning, the return path of the UAV is dynamically adjusted to generate the UAV route planning.

[0113] S6.1 Based on the current position of the UAV and the position of the obstacle, the deviation rate between the current path and the actual flight path is obtained by iterative closest point algorithm. The expression is:

[0114] ;

[0115] in, is the deviation rate between the current path and the actual flight path, is the flight time, For drones in time Actual location, is the expected position on the planned path, is the time index.

[0116] Furthermore, the actual position data at a series of consecutive time points and the expected positions on the planned path are obtained from the drone sensor. The iterative closest point algorithm is used to match these actual positions with the expected positions, find each pair of closest points, and calculate the deviation rate between the current path and the actual flight path based on the distance difference between these matching point pairs.

[0117] S6.2 When the current position of the UAV and the position of the obstacle are less than the UAV's route planning deviation rate, continue to fly along the current path.

[0118] Furthermore, if the distance between the drone's current position and the nearest obstacle is greater than the deviation rate between the current path and the actual flight path, it means that the drone is currently within the safe range and the path deviation is small. The drone will continue to fly along the current path without the need for path adjustment, which can reduce unnecessary path adjustments and maintain flight stability and efficiency. The drone's status and environmental changes will be continuously monitored to ensure that any potential risks can be discovered and handled in a timely manner.

[0119] S6.3 When the current position of the UAV and the position of the obstacle are greater than or equal to the deviation rate of the UAV's route planning, the greedy algorithm is used to optimize the deviation part to obtain the UAV route planning.

[0120] Furthermore, if the distance between the drone's current position and the nearest obstacle is less than the deviation rate between the current path and the actual flight path, it indicates that the drone may face a high collision risk or has significantly deviated from the preset path. In this case, a greedy algorithm will be used to optimize the deviation part and select a path that minimizes the deviation and avoids obstacles in the shortest time. The greedy algorithm gradually constructs a global path through local optimal solutions to ensure that the drone can return to a safe path as soon as possible, and generates a new drone route plan so that the drone can safely bypass obstacles and continue to perform its mission.

[0121] This embodiment also provides a drone route planning system based on deep reinforcement learning, including:

[0122] Data acquisition module, collects the current state parameters and comprehensive navigation situation data of the UAV;

[0123] The deviation rate module calculates the deviation rate of the drone's route planning based on the collected drone's current state parameters;

[0124] The 3D map module uses deep learning algorithms to process comprehensive navigation situation data, identify the current precise location of the drone, use SLAM technology to build a 3D map model, and use point cloud fusion method to generate a 3D map;

[0125] The path return module builds a deep reinforcement learning model based on the drone's current precise position, current state parameters, and 3D map, and uses quantum algorithms to calculate the drone's return path when facing obstacles;

[0126] The path planning module dynamically adjusts the return path of the drone based on the drone's route planning deviation rate and generates the drone's route planning.

[0127] This embodiment also provides a computer device, which is suitable for the case of a drone route planning method based on deep reinforcement learning, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the drone route planning method based on deep reinforcement learning proposed in the above embodiment.

[0128] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad on the computer device housing, or an external keyboard, touchpad, or mouse.

[0129] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for implementing a drone route planning method based on deep reinforcement learning as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0130] In summary, the present invention realizes accurate monitoring and dynamic adjustment of the flight path by calculating the route planning deviation rate of the UAV, sets the flight mission based on the current position and target position of the UAV, and plans the ideal path using the RT path optimization algorithm. The actual position is recorded at fixed time intervals to form sampling points, and the actual position is determined by combining the inertial measurement unit data and the lidar data. The ball tree algorithm is used to calculate the deviation rate to ensure high-precision navigation even in complex environments. Corner feature points are extracted from visual data and matched through a fast library approximate nearest neighbor search method. A posture network model is constructed to obtain the preliminary position and posture. The extended Kalman filter is used to fuse the deep learning model to calculate the precise position of the UAV, thereby enhancing the autonomous navigation capability of the UAV in dynamic and complex environments. Through multi-source data fusion and advanced algorithm processing, the UAV can operate stably in various environments.

[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A UAV route planning method based on deep reinforcement learning, characterized by: include, Collect the current status parameters and comprehensive navigation situation data of the UAV; Calculate the deviation rate of the drone's route planning based on the collected drone's current state parameters; Use deep learning algorithms to process comprehensive navigation situation data, identify the current precise location of the drone, use SLAM technology to build a three-dimensional map model, and use point cloud fusion method to generate a three-dimensional map; Based on the drone's current precise location, current state parameters, and 3D map, a deep reinforcement learning model is built, using quantum algorithms to calculate the drone's return path when facing obstacles. Based on the deviation rate of the drone’s route planning, the return path of the drone is dynamically adjusted to generate the drone’s route planning; Collecting the current state parameters and comprehensive navigation situation data of the UAV includes the following steps: The current state parameters of the UAV include speed and attitude, and the comprehensive navigation situation data includes the current position information, target position information, flight status and environmental perception data of the UAV; The environmental perception data includes inertial measurement unit data, image data and lidar data; Based on the collected current state parameters of the UAV, the calculation of the UAV route planning deviation rate includes the following steps: The UAV's flight mission is set based on its current location information and target location information. The UAV's current state parameters are combined with the UAV's flight mission, and the ideal flight path of the UAV is obtained using a fast exploration random tree path optimization algorithm. The current location information of the drone is recorded at fixed time intervals to form sampling points; Determine the actual position of the drone using sensor fusion technology using inertial measurement unit data and lidar data; The corresponding points on the ideal path and the actual position of the UAV are found through sampling points, and the ball tree algorithm is used to calculate the UAV's route planning deviation rate.

2. The UAV route planning method based on deep reinforcement learning according to claim 1, characterized in that: Using deep learning algorithms to process comprehensive navigation situation data and identify the current precise location of the drone includes the following steps: Use simultaneous localization and mapping algorithms to extract corner feature points from image data; Use the fast library approximate nearest neighbor search method to match the current corner feature point with the previously extracted corner feature point, use the geometric matrix to calculate the relative displacement and rotation matrix, and obtain the transformation matrix corresponding to the matched feature point set; Based on the transformation matrix corresponding to the matched feature point set, a posture network model is constructed and the image data is processed sequentially to obtain the initial position and posture of the UAV; Based on the preliminary position and attitude of the UAV and combined with the inertial measurement unit data, the extended Kalman filter is used to fuse the deep learning model to calculate the precise position of the UAV.

3. The UAV route planning method based on deep reinforcement learning according to claim 1, characterized in that: Using SLAM technology to build a 3D map model and point cloud fusion method to generate a 3D map includes the following steps: The laser radar sensor is used to obtain 3D information of the surrounding environment, and the SLAM technology is used to build a 3D map model, and different point cloud data are extracted from the 3D information of the surrounding environment; Use the normal distribution transformation algorithm to fuse different point cloud data to generate a complete three-dimensional point cloud map; Convert the complete 3D point cloud into a 3D mesh using the marching cubes algorithm; Texture mapping is performed on the complete 3D mesh to obtain a preliminary 3D map; The preliminary three-dimensional map is smoothed by point cloud fusion method to obtain a complete three-dimensional map.

4. The UAV route planning method based on deep reinforcement learning according to claim 1, characterized in that: Based on the current precise position of the UAV, the current state parameters and the three-dimensional map, a deep reinforcement learning model is constructed. The quantum algorithm is used to calculate the return path of the UAV when facing obstacles. The following steps are included: Use point cloud processing technology in the complete 3D map to obtain the location of obstacles and the height of the drone above the ground; Define the state space of the drone by its velocity, attitude, and target position; Define the drone’s action space through its flight state; Build a deep Q network based on the drone's state space and action space and define the loss function; Predict the drone’s action strategy in each state through the loss function; Convert the complete 3D map data into a quantum-computing 3D map structure and apply the cost function to optimize the UAV’s action strategy; Based on the optimized UAV action strategy, a quantum approximate optimization algorithm is applied to obtain the return path of the UAV when facing obstacles.

5. The UAV route planning method based on deep reinforcement learning according to claim 1, characterized in that: Based on the deviation rate of the UAV's route planning, the return path of the UAV is dynamically adjusted to generate the UAV route planning The following steps are included: Based on the current position of the UAV and the position of the obstacle, the deviation rate between the current path and the actual flight path is obtained through the iterative closest point algorithm; When the current position of the UAV and the position of the obstacle are less than the deviation rate of the UAV's route planning, the UAV continues to fly along the current path; When the current position of the UAV and the position of the obstacle are greater than or equal to the deviation rate of the UAV's route planning, the greedy algorithm is used to optimize the deviation part to obtain the UAV route planning.

6. A UAV route planning system based on deep reinforcement learning, based on the UAV route planning method based on deep reinforcement learning according to any one of claims 1 to 5, characterized in that: include, Data acquisition module, collects the current state parameters and comprehensive navigation situation data of the UAV; The deviation rate module calculates the deviation rate of the drone's route planning based on the collected drone's current state parameters; The 3D map module uses deep learning algorithms to process comprehensive navigation situation data, identify the current precise location of the drone, use SLAM technology to build a 3D map model, and use point cloud fusion method to generate a 3D map; The path return module builds a deep reinforcement learning model based on the drone's current precise position, current state parameters, and 3D map, and uses quantum algorithms to calculate the drone's return path when facing obstacles; The path planning module dynamically adjusts the return path of the drone based on the drone's route planning deviation rate and generates the drone's route planning.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the drone route planning method based on deep reinforcement learning are implemented in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the drone route planning method based on deep reinforcement learning are implemented.

Citation Information

Patent Citations

  • Unmanned aerial vehicle autonomous navigation method and system based on deep reinforcement learning, and medium

    CN118225106A

  • Unmanned aerial vehicle navigation positioning method based on bridge detection

    CN119414875A