Unmanned aerial vehicle navigation method and system integrating reinforcement learning and model prediction
By combining deep reinforcement learning and model predictive control, and using lidar to build a 3D point cloud map and optimize trajectory planning, the problems of planning collisions and control stability of UAVs in complex indoor environments were solved, and safe and stable flight was achieved.
Patent Information
- Application Number
- CN202511721445.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-24
AI Technical Summary
In existing technologies, drones are prone to collisions when planning in complex indoor environments, and the system control stability is poor. Traditional methods struggle to balance computational complexity, real-time performance, and adaptability.
By combining deep reinforcement learning and model predictive control, environmental information is collected by LiDAR to build a 3D point cloud map, extract a global 2D plane map, search for the optimal global path, and optimize the trajectory through model predictive control to achieve smooth trajectory tracking.
It improves the safety and control stability of UAVs in complex obstacle environments, reduces computational complexity, enhances the robustness and smoothness of trajectory planning, and ensures optimal path and minimum energy consumption.
Smart Images

Figure CN121558024A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) navigation, and more specifically to a UAV navigation method and system that integrates reinforcement learning and model prediction. Background Technology
[0002] Currently, the low-altitude economy is developing rapidly, encompassing various economic forms such as drone logistics, aerial sightseeing, short-haul transportation, emergency rescue, agricultural plant protection, and aviation sports. Meanwhile, research on drone trajectory planning and control has become a hot topic. Numerous studies have emerged on drone trajectory and control, but autonomous navigation and trajectory planning in complex environments remain the core challenges for the practical application of this technology.
[0003] Drones have numerous applications in indoor environments. For example, in large e-commerce warehouses, drones need to quickly move between shelves to perform inventory checks or sort goods. However, adjustments to shelf layouts, forklift movements, and personnel movement can alter the environmental topology. For indoor drone navigation, traditional GPS-dependent navigation fails indoors. In the confined and relatively low-speed indoor environment, traditional drone navigation methods rely heavily on accurate environmental models and predefined path planning, performing poorly in environments with obstacles and exhibiting weak robustness.
[0004] To address these issues, academia and industry have recently begun exploring intelligent trajectory planning methods that combine deep learning, reinforcement learning, and multi-sensor fusion. However, existing solutions often struggle to balance computational complexity, real-time performance, and adaptability, especially on resource-constrained airborne computing platforms. Therefore, there is an urgent need for a novel, efficient, robust, and adaptable indoor navigation and trajectory planning technology for unmanned aerial vehicles (UAVs) to support their safe, stable, and autonomous operation in complex and dynamic environments. This patent aims to propose an innovative solution to overcome the shortcomings of existing technologies and promote the widespread adoption of UAVs in indoor applications.
[0005] Deep Reinforcement Learning (DRL), as an emerging artificial intelligence technology, has achieved significant results in fields such as games and robot control by learning optimal strategies through interaction with the environment. It does not require pre-established precise environmental models and can autonomously learn and adapt to complex environments, providing a new approach for UAV navigation. For example, the existing invention patent application document CN115345281A, entitled "A Deep Reinforcement Learning Accelerated Training Method for UAV Image Navigation," includes training an object detection model and training an action selection strategy. The action selection strategy is trained based on the simulated output values of the object detection model using a deep reinforcement learning method. In image-based deep reinforcement learning training, the simulated output values of the image detection model are used instead of the actual output values to accelerate training. However, this existing approach also suffers from low sample efficiency and training instability, making it difficult to directly apply to UAV navigation tasks with high safety and real-time requirements. For example, the existing invention patent application document CN120595853A, entitled "A Shape Adaptive Planning and Control Method for Deformable Unmanned Aerial Vehicles," describes a method that includes: searching for an initial path based on a scene-based occupancy grid map using a kinematic A* path planning algorithm with variable dimensions; using the initial path as initial conditions, performing shape adaptive trajectory optimization in continuous spatiotemporal space to obtain an optimized trajectory, wherein the UAV's centroid position and deformation parameters are simultaneously optimized during trajectory optimization; and based on the optimized trajectory, using a nonlinear model predictive controller (MPC) combined with an incremental nonlinear dynamic inverse algorithm, calculating and compensating for external force disturbances and external torque disturbances caused by UAV deformation and load. While MPC can provide accurate trajectory tracking and obstacle avoidance capabilities for UAVs, its performance heavily relies on the accuracy of the model and is difficult to handle complex and changing environments.
[0006] The existing invention patent application document CN119458337A, entitled "A Robot Assembly Optimization Method and System Based on Federated Learning," describes a method that uses a compensation algorithm employing a classic PID (Proportional-Integral-Derivative) controller or Model Predictive Control (MPC) to correct robot movements in real time. It also utilizes a Deep Reinforcement Learning (DRL) strategy to achieve motion control.
[0007] However, this motion compensation algorithm relies heavily on the fusion of features from both image data and point cloud data, and the inference process involving deep reinforcement learning will inevitably increase the computational overhead of the processor.
[0008] The existing invention patent application document CN120769324A, entitled "Polarization-Spacetime Coding Anti-jamming Communication System for UAV Swarm Based on Dynamic Topology Reconstruction," describes a method that includes: running a core multi-objective optimization decision algorithm (such as a policy network based on deep reinforcement learning (DRL)) to generate cooperative control commands and issue them to each execution subsystem, thereby achieving closed-loop intelligent control of the entire system. It employs an artificial potential field method or model predictive control (MPC) to plan a smooth and safe flight trajectory, while simultaneously coordinating the MAC layer for link establishment and teardown.
[0009] This patent uses artificial potential fields or model predictive control to plan the trajectory and directly sends the results as commands to the flight control system. While this method reduces the computational burden, it does not sufficiently smooth the trajectory, which may lead to abrupt changes in the trajectory, causing collisions or instability in the drone's control.
[0010] The existing invention patent application document CN120686872A, entitled "A Control Method, Device, Electronic Equipment, and Readable Storage Medium for an Unmanned Aerial Vehicle," describes a method that includes: real-time tracking and risk avoidance of the UAV path using Model Predictive Control (MPC). This is replaced by a Deep Reinforcement Learning (DRL) strategy: taking the risk field and UAV state as input, the action space includes velocity / altitude increments, and the reward function is directly mapped to the negative value of the cost function J. This patent uses either an A*+MPC method or a deep reinforcement learning method for planning.
[0011] In summary, existing technologies suffer from technical problems such as easy collisions during drone planning, insufficient trajectory smoothness, susceptibility to sudden changes in control, and poor system control stability. Summary of the Invention
[0012] The technical problem to be solved by this invention is: how to solve the technical problems of easy collisions and poor system control stability in the existing technology of drone planning.
[0013] This invention solves the above-mentioned technical problems by employing the following technical solution: a UAV navigation method integrating reinforcement learning and model prediction, comprising: S1. Utilize autonomous exploration and LiDAR to collect environmental information and establish a three-dimensional point cloud map of the environment; S2. At the preset height of the target point, extract the global two-dimensional planar map of the global map projection; S3. Search for the optimal global path on the global two-dimensional plane map; S4. Based on the optimal global path, a smooth trajectory is obtained. The information of the smooth trajectory includes: UAV position, UAV speed, and UAV attitude. S5. Based on the smooth trajectory, optimize the operation through model prediction control to perform trajectory tracking operation on the UAV.
[0014] This invention combines deep reinforcement learning and model predictive control for UAV navigation, enabling trajectory planning and control of UAVs in indoor environments with various obstacles. It solves the collision problem common in traditional UAV planning methods, while introducing model prediction to improve system control stability. This invention utilizes optimized methods for UAV tracking control, enhancing overall control accuracy and stability. It allows for safe and stable flight even in environments with numerous obstacles.
[0015] This invention provides a superior solution to the problems of existing technologies. The core of the invention is to replace complex deep reinforcement learning with a lightweight path point inference mechanism to improve computational efficiency. Furthermore, by constraining the rate of change of the control input in the MPC cost function, a smoother trajectory tracking that is more in line with UAV dynamics is achieved, thereby achieving a balance between computational complexity, control accuracy, and system robustness.
[0016] In a more specific technical solution, S1 includes: When entering an indoor environment, data is collected using the lidar onboard the drone; during the drone's movement, the drone's pose is estimated in real time and the lidar data is fused to generate a local point cloud map; a global map is constructed by using an incremental map construction method to fuse the new point cloud data with the current global map to obtain a global point cloud map. By processing the global point cloud map, removing dynamic objects and noise, and using point cloud processing tools to filter and optimize the global map, loop closure detection is performed to construct a globally consistent static map. By controlling the map density and accuracy through controlled parameters, an appropriate map density is obtained to produce an indoor environment point cloud map.
[0017] The inference of path points obtained by point cloud information and global search in this invention can ensure the accuracy of the target point. Combined with MPC control, it can ensure the precision of control, significantly reducing the computational complexity while ensuring the accuracy of navigation.
[0018] In a more specific technical solution, S2 includes: A two-dimensional rasterized map is obtained; let the map data be matrix M, with... This represents the value of the raster cell in the i-th row and j-th column; using a contour line extraction algorithm, contour lines are extracted from the map data to complete the information creation operation of the contour 2D map; Identify obstacles and target areas, analyze the 2D map of the elevation, record the location and size of closed areas to represent obstacles, and identify and extract obstacles based on area threshold conditions to obtain obstacle information data; Navigation of drones is based on contour-based 2D maps.
[0019] In a more specific technical solution, S3 includes: The global two-dimensional planar map is constructed into a vector map, and the optimal global path is found by using an improved global path planning algorithm. The goal is to find an optimal path without collisions and to perform micro-expansion processing on obstacles. By using global path planning, a feasible path at the expected height is obtained. The desired height is then extended to the feasible path to obtain a three-dimensional global path.
[0020] This invention uses the optimal global path obtained first through global planning to provide local target points for subsequent trajectory planning, which can accelerate the convergence speed of deep reinforcement learning algorithms, increase the reward value obtained during training, and thus generate a more reasonable trajectory.
[0021] This invention uses model predictive control to first calculate the cost of the planned path points, which can smooth the trajectory and generate a trajectory that is more in line with the kinematics and dynamics model of the UAV, thereby improving the stability and safety of the overall motion control.
[0022] In a more specific technical solution, S4 includes: Training yields suitable network parameters; By using applicable network parameters for model inference, the UAV inputs environmental information into the policy network to obtain actions, and updates the state through actions. Multi-state information trajectories are generated through deep reinforcement learning, and discrete path points are optimized to obtain smooth trajectories.
[0023] This invention generates trajectories with multi-state information through deep reinforcement learning, further optimizing the originally discrete path points into smooth trajectories and simultaneously reducing the probability of collisions with obstacles. Trajectories containing multiple state information also help improve the stability and accuracy of control.
[0024] This invention is designed for indoor environments. It achieves optimality and smoothness in the planned path, and uses global path planning to connect trajectory planning, ensuring optimal path and minimum energy consumption.
[0025] In a more specific technical solution, during the operation of training the applicable network parameters, the neural network parameters are randomly initialized; Initialize the drone's state values; The UAV observation information is acquired, fed into the reinforcement learning framework for training, and the action and reward are obtained. The state observations input into the deep reinforcement learning include the UAV state value, the external point cloud information value observed by the UAV, and the temporary target point relative to the global path. The state, action, reward, and updated value after each state transition are added to the experience pool; the following logic is used for updating the state of the drone agent: Mecha speed update ; Updates to yaw and pitch angles: ; Location update: ; Combine the updated state value with the previous one. Add them together to the experience pool; When the total step size reaches the update frequency, take one bacth of experience from the experience pool and use priority experience replay to update the network parameters; Jump to the operation of obtaining the action and reward, and execute it in a loop until the applicable network parameters are obtained.
[0026] In a more specific technical solution, the desired roll angle is calculated based on the desired yaw angle:
[0027] Status obtained: When the target point is reached, at least two time series consisting of states are obtained, which are used as the trajectory obtained from the planning.
[0028] The trained model can be mounted on the UAV's onboard host, requiring minimal computing power and making deployment relatively easy. It also exhibits strong robustness, capable of generating reasonable trajectories through continuous inference based on input observations in various environments, demonstrating strong adaptability.
[0029] In a more specific technical solution, S5 includes: Based on the dynamics and kinematics model of the UAV, a state transition equation is established, and a unified UAV controller is built to control the UAV's position and attitude. Specifically, the continuous-time dynamics model is discretized to obtain a discrete-time state-space model. Define the prediction time domain and the control time domain. The prediction time domain predicts the state for the next Np steps. The control time domain, Nc, optimizes the control input for the next Nc steps. Construct an objective function, where the penalty terms of the objective function include: a penalty term for the current state and tracking error, a smoothing penalty term for the UAV thrust and torque input, and a penalty term for the terminal error in the nth future time domain; Based on the objective function, constraints are added, including: UAV state constraints, control input constraints, and dynamic constraints; Under constraints, the OSQP solver library is called to simplify the UAV nonlinear model into a linear model.
[0030] In a more specific technical solution, the objective function is constructed using the following logic:
[0031] In the formula, Q, R, and P are weight matrices, corresponding to the weights of state error, control input, and terminal state, respectively.
[0032] This invention achieves stable control of UAVs by constructing an optimization function through model prediction. The control method with future prediction properties can achieve better and more stable control of UAVs. At the same time, the introduction of an input smoothing term penalty ensures that the control input does not change too much in adjacent time domains, thus achieving the smoothness and safety of UAV control. In more specific technical solutions, drone navigation systems that integrate reinforcement learning and model prediction include: The map building module is used to autonomously explore using LiDAR to collect environmental information and build a 3D point cloud map of the environment; The 2D map acquisition module is used to extract the global 2D planar map of the global map projection at a preset height of the target point. The 2D map acquisition module is connected to the map building module. The optimal path search module is used to search for the optimal global path on the global two-dimensional planar map. The optimal path search module is connected to the two-dimensional map acquisition module. The smoothing module is used to process the optimal global path to obtain a smooth trajectory. The smooth trajectory information includes: UAV position, UAV speed, and UAV attitude. The smoothing module is connected to the optimal path search module. The estimation and tracking module is used to perform trajectory tracking operations on the UAV by using model prediction control optimization based on the smooth trajectory. The estimation and tracking module is connected to the smoothing processing module.
[0033] The present invention has the following advantages over the prior art: This invention combines deep reinforcement learning and model predictive control for UAV navigation, enabling trajectory planning and control of UAVs in indoor environments with various obstacles. It solves the collision problem common in traditional UAV planning methods, while introducing model prediction to improve system control stability. This invention utilizes optimized methods for UAV tracking control, enhancing overall control accuracy and stability. It allows for safe and stable flight even in environments with numerous obstacles.
[0034] This invention uses the optimal global path obtained first through global planning to provide local target points for subsequent trajectory planning, which can accelerate the convergence speed of deep reinforcement learning algorithms, increase the reward value obtained during training, and thus generate a more reasonable trajectory.
[0035] This invention generates trajectories with multi-state information through deep reinforcement learning, further optimizing the originally discrete path points into smooth trajectories and simultaneously reducing the probability of collisions with obstacles. Trajectories containing multiple state information also help improve the stability and accuracy of control.
[0036] This invention is designed for indoor environments. It achieves optimality and smoothness in the planned path, and uses global path planning to connect trajectory planning, ensuring optimal path and minimum energy consumption.
[0037] This invention achieves stable control of UAVs by constructing an optimization function through model prediction. The control method with future prediction properties can achieve better and more stable control of UAVs. At the same time, the introduction of an input smoothing term penalty ensures that the control input does not change too much in adjacent time domains, thus achieving the smoothness and safety of UAV control.
[0038] This invention solves the technical problems of easy collisions during drone planning and poor system control stability in the prior art. Attached Figure Description
[0039] Figure 1 This is a schematic diagram of the basic steps of the UAV navigation method that integrates reinforcement learning and model prediction according to Embodiment 1 of the present invention; Figure 2 The map used for verification in Embodiment 1 of the present invention; Figure 3 This is a contour map extracted from a map according to Embodiment 1 of the present invention; Figure 4 This is a diagram showing the result obtained using global path planning in Embodiment 1 of the present invention; Figure 5 This is a trajectory diagram obtained from different starting points using deep reinforcement learning in Embodiment 1 of the present invention; Figure 6 The actual trajectory of the UAV after trajectory tracking is obtained using the model prediction algorithm in Embodiment 1 of the present invention; Figure 7 This is a flowchart of the trajectory planning algorithm in Embodiment 1 of the present invention; Figure 8 This is a schematic diagram of the data flow processing of the control optimization algorithm in Embodiment 1 of the present invention. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0041] Example 1 like Figure 1 As shown, the UAV navigation method integrating reinforcement learning and model prediction provided by this invention includes the following basic steps: S1. The drone autonomously explores the environment, collects environmental information using lidar, and builds a map.
[0042] like Figure 2 As shown, in this embodiment, upon entering an indoor environment, data is first collected autonomously by a drone. The drone, equipped with a LiDAR, moves randomly and irregularly within the target area to ensure that the LiDAR-detected area covers the entire region. During movement, the drone's pose is estimated in real time and fused with LiDAR data to generate a local point cloud map. Then, a global map is constructed using an incremental map building method, continuously fusing new point cloud data with the global map. Simultaneously, real-time point cloud data is published during movement, and this data accumulates to form a global point cloud map.
[0043] In this embodiment, the UAV radar sensor acquires point cloud information in space. Through the transformation relationship between the Earth coordinate system and the UAV coordinate system, a point cloud map under a global map is generated. The UAV obtains its own position and attitude information through its IMU sensor, which yields a relative rotation matrix between the UAV and the world coordinate system. Using the obtained point cloud coordinate information from the lidar, this is converted into map information in the world coordinate system. The transformation relationship between the UAV and the world coordinate system is expressed as follows:
[0044] The position transformation relationship is expressed as follows: Let the position of the UAV in the world coordinate system be... The position of the point cloud in the body coordinate system The position of a point cloud in world coordinates is represented as .
[0045] In this embodiment, after the map is built, it needs to be processed to remove dynamic objects or noise. Point cloud processing tools are used to filter and optimize the global map. Outliers and unstable points are removed, as these may represent dynamic objects or noise. After generation, loop closure detection is also performed. A globally consistent static map needs to be built, and the map density and accuracy are controlled by adjusting parameters to obtain an appropriate map density. Next, the map is saved. After the UAV autonomously explores and scans the entire area, the accumulated global point cloud map can be saved. The Robot Operating System (ROS) tools are used to save the point cloud data as a .pcd or .ply file.
[0046] After the above steps, a point cloud map of the indoor environment can be obtained, which consists of several point cloud points. Each point cloud contains the following information. in, The coordinates of the point cloud in the world coordinate system are represented by I, and the intensity value of the point cloud is represented by I. This is used to establish a three-dimensional point cloud map of the environment.
[0047] S2. After obtaining the global map, the communication transmission enables the UAV to obtain the target point location information and extract the contour line two-dimensional map at the height of the target point, including obstacle information and target information.
[0048] like Figure 3 As shown in this embodiment, after obtaining the relative position information between the target task point and the UAV, the target point is first converted to the global map, and then the environmental map is simplified to extract the two-dimensional map on the contour line of the target point, including all obstacles and target information.
[0049] From the 3D map data (.pcd) obtained using ROS, a 2D rasterized map is extracted and saved as a .yaml file. Map data is typically stored in raster format, where the value of each raster cell represents the height or occupancy status of the area. Assuming the map data is a matrix M, it can be... This represents the value of the raster cell in the i-th row and j-th column.
[0050] Contour line extraction can be performed using contour line extraction algorithms to extract contour lines from map data. The main implementation method involves first preprocessing the map data, converting the map data M into a grayscale image I, and then binarizing image I to generate a binary image B. The formula is expressed as follows:
[0051] Contour lines are composed of a series of points, represented as ,in These are the coordinates of a point on a contour line.
[0052] The information is then published through ROS. After the two-dimensional map information based on the contour lines is established, obstacles and target areas need to be identified. First, the contour map is analyzed to record the position and size of closed areas, representing several possible obstacles. Then, based on whether the area is greater than a threshold, it is determined whether it is an obstacle. Finally, all obstacle information data is extracted.
[0053]
[0054] The target location area is marked using special markers and colors, conveying information about the target's range. Map information including obstacles, the target itself, and its length and width is retained. The final result is a two-dimensional planar map showing the height of the target point.
[0055] Creating a two-dimensional map is used to plan paths at a fixed altitude, allowing the drone to navigate by first ascending to that altitude and then moving horizontally. This not only minimizes energy loss but also reduces the computational complexity of the planning process.
[0056] S3. Generate a reasonable collision-free path on the extracted 2D map using a global path planning algorithm.
[0057] The constructed 2D map is a vector map, and the goal is to find a collision-free optimal path. Obstacles on the map are present, and these obstacles are slightly magnified to become convex. An improved A* algorithm is used to find the optimal path.
[0058] In this embodiment, the A* algorithm, as a heuristic algorithm, finds the optimal path by minimizing the cost, which is achieved by evaluating the cost function of each node. To achieve this, the cost function for node n consists of two parts, where... This represents the actual cost from the starting point to node n, and the heuristic cost function from node n to the target point. Euclidean distance is chosen as the preferred method. The specific implementation steps are as follows; Define the positions of the target point and the initial point, which is represented as...
[0059] Create an open list and an explore list to store the nodes to be explored; add the starting point to the open list and initialize it.
[0060] The nodes are expanded by traversing their neighboring nodes. Since a vector map is used, the expansion area and range need to be manually set. Here, to select paths more efficiently and quickly, an expansion method of eight neighboring nodes is adopted, with each expansion distance in a minimum unit of 1. Therefore, the expanded nodes can be represented as (1,0), (0,1), (1,1), (0,-1), (-1,0), (-1,1), (1,-1), (-1,-1). Then, for each neighboring node, it is judged: if it is within the obstacle range, the node is skipped; if it is not within the obstacle range, the cost function is calculated. The actual cost is the original node cost plus the movement distance. In calculating the value of the total cost function
[0061] The node with the minimum cost function is selected and added to the exploration set. This process is repeated until the found node is within the range of the target point, or there are no nodes to explore, indicating that there is no feasible path.
[0062] This invention introduces the change in input as part of the loss calculation in the cost function of model predicting control input. This improvement can better constrain drastic fluctuations in control input, prevent the UAV from losing balance due to sudden input changes, and enhance the control robustness of the system.
[0063] After the search is complete, a feasible path is obtained by backtracking the parent node of each node.
[0064] like Figure 4 As shown, in this embodiment, a feasible path at the expected height is obtained through global path planning. The path includes a series of two-dimensional coordinate points of equal height. The expected height is extended to this path to obtain a three-dimensional global path. Using the optimal global path obtained first through global planning provides local target points for subsequent trajectory planning, which can accelerate the convergence speed of the deep reinforcement learning algorithm, increase the reward value obtained during training, and thus generate a more reasonable trajectory.
[0065] S4. Use deep reinforcement learning algorithms to obtain a smooth trajectory with the drone's position, velocity, and attitude information.
[0066] The global path obtained by the search only contains location information, and the resulting point set is not continuous and smooth enough. Therefore, by using deep reinforcement learning, the drone is regarded as an intelligent agent, and the most reasonable trajectory information with the drone's position, speed and attitude is obtained through its interaction with the environment. In this embodiment, to obtain better network parameters while minimizing training time, an improved SAC algorithm is used for training. This algorithm is based on the maximum entropy reinforcement learning framework, combining the Actor-Critic method and entropy regularization. This improves exploration capabilities and demonstrates excellent performance in continuous control tasks. Its core idea is to maximize not only the cumulative reward but also the policy's entropy when optimizing the policy. Entropy represents the randomness of the policy; high-entropy policies are more exploratory and help discover better policies. Simultaneously, a priority experience replay mechanism is introduced. By calculating the error of each experience in the experience pool, higher-value experiences are selected from the pool to update the network parameters, which accelerates the convergence speed of the trained neural network.
[0067] like Figure 7 As shown, in this embodiment, the trajectory planning algorithm includes the following specific steps: S41. Initialize neural network parameters; S42. Obtain the generated path point information and initialize the UAV agent information; S43. Obtain UAV status information and use lookahead to obtain local target points; S44. Receive an action based on the state observation value and be rewarded; S45. Add the state, state transition process, action, and reward to the experience pool; S46. Determine if the update step size has been reached; S47. If so, select higher-value experiences from the experience pool and update the network parameters; S48. If not, determine whether the maximum step size has been completed or reached. S49. If yes, then in this round of training, the aforementioned steps S42 to S48 are executed repeatedly; if no, then based on the state transition equation, step S43 is executed.
[0068] In this embodiment, the specific operations of the trajectory planning training method include: Randomly initialize the neural network parameters; Initialize the drone's state values; like Figure 5 and Figure 6As shown, in this embodiment, UAV observation information is acquired, fed into a reinforcement learning framework for training, and the action and reward are obtained. The state observations input into the deep reinforcement learning framework should include the UAV's own state values, the external point cloud information values observed by the UAV, and temporary target points relative to the global path. The selection of temporary target points employs a UAV guidance system method. First, the point closest to the UAV's current coordinates is found. Starting from this point, the search proceeds along the path. When the distance between a path point and the UAV agent's current position is greater than the minimum exploration value, that point is selected as the temporary target point. The observation information includes the UAV agent's own position, velocity, temporary target points, and the conversion of the observed forward point cloud information into distance and direction to obstacles.
[0069] The reward function includes sparse and dense rewards to encourage the UAV to approach the collision-free and optimal search path. Dense rewards include penalties for collisions, penalties for going out of bounds, and rewards for reaching temporary and target points. Sparse rewards include rewards for position and orientation approaching the target, rewards for height stability, and rewards for speed approaching the desired speed. The total reward includes the following components;
[0070] The output action includes the horizontal angular velocity, the vertical angular velocity, and the magnitude of a linear acceleration. This action will be used to update the state. The state, action, reward, and updated value after each state transition are added to the experience pool. The state update of the drone agent mainly includes updates to the overall speed of the drone. The update formulas for yaw and pitch angles are as follows: The velocity in each direction is calculated and expressed as:
[0071] The position update equation is expressed as Combine the updated state value with the previous one. Add them together to the experience pool.
[0072] When the total step size reaches the update frequency, an experience from the experience pool is taken to update the network parameters. Here, in order to improve utilization and speed up convergence, the experience selection adopts the method of prioritizing experience replay.
[0073] The method for updating network parameters is as follows: First, calculate the network Q-value for each tuple. Then, update the parameter values of the critic network by minimizing the loss function (MSE). Next, update the parameters and regularization coefficients of the actor network in the same way. Finally, update the target network using a soft update method. .
[0074] Returning to the step of training using the reinforcement learning framework based on drone observations, this process is repeated iteratively. If, in a single round, the agent's `done` value is `True`, or the step size has exceeded the maximum step size, then training for that round is stopped, and the process returns to the step of initializing the drone's state values to continue execution.
[0075] After all rounds are completed, the trained network parameter values are obtained and saved.
[0076] In this embodiment, when a trajectory is needed, this parameter is used for model inference. The UAV continuously obtains actions by inputting environmental information into the policy network, and updates the state through these actions. The desired roll angle can be roughly calculated based on the desired yaw angle. The final state is obtained. When the target point is reached, several time series consisting of states can be obtained as the trajectory obtained from the planning, which includes position, velocity, and Euler angle information.
[0077] S5. Achieve trajectory tracking control for quadcopter drones through optimized methods; Through the previous steps, a trajectory of a quadcopter drone in an indoor environment, including position, velocity, and Euler angles, can be obtained. An optimization function is then constructed using the principles of Model Predictive Control (MPC) to achieve trajectory tracking control of the drone. Specifically, Model Predictive Control (MPC) is a model-based optimal control method that effectively handles system constraints and uncertainties through online rolling optimization and feedback correction.
[0078] The system model predicts behavior over a future period and employs rolling optimization. In each control cycle, a finite-time optimization problem is solved and feedback correction is performed. After the first control variable is applied to the system, the optimization is restarted in the next cycle.
[0079] like Figure 8 As shown, in this embodiment, firstly, a state transition equation is established based on the dynamics and kinematics model of the UAV. Then, a unified UAV controller is established to achieve simultaneous control of the UAV's position and attitude. The continuous-time dynamics model is discretized to obtain a discrete-time state-space model. Then, the prediction time domain and the control time domain are defined. The prediction time domain predicts the state for the next Np steps. The control time domain Nc optimizes the control input for the next Nc steps.
[0080] In this embodiment, a reasonable objective function is constructed. Here, this patent designs an objective function that includes three penalty terms: a penalty term for the current state and tracking error, and a smoothing penalty term for the UAV thrust and torque input. And the terminal error penalty term for the nth time domain in the future, Q, R, and P are weight matrices, corresponding to the weights of state error, control input, and terminal state, respectively. These weights need continuous adjustment based on actual control results. In addition to the objective function, constraint conditions are added, including UAV state constraints, control input constraints, and dynamic constraints. Under these constraints, the OSQP library is used to simplify the UAV's nonlinear model into a linear model. This is because the changes in attitude angles are relatively small during trajectory tracking for a quadcopter UAV; simplifying the UAV motion to linear will not produce significant differences. Furthermore, using a linear model to control the UAV results in lower complexity and faster response time. This is suitable for the small-scale tracking and relatively low-disturbance environments of indoor UAV navigation in this patent.
[0081] Stable control of UAVs is achieved by constructing an optimization function through model prediction. The control method with future prediction properties can achieve better and more stable control of UAVs. At the same time, the introduction of an input smoothing term penalty ensures that the control input does not change too much in adjacent time domains, thus achieving the smoothness and safety of UAV control.
[0082] In summary, this invention combines deep reinforcement learning and model predictive control for UAV navigation, enabling trajectory planning and control of UAVs in indoor environments with various obstacles. It solves the collision problem common in traditional methods and improves system control stability by introducing model prediction. This invention utilizes optimized methods for UAV tracking control, enhancing overall control accuracy and stability. It allows for safe and stable flight even in environments with numerous obstacles.
[0083] This invention uses the optimal global path obtained first through global planning to provide local target points for subsequent trajectory planning, which can accelerate the convergence speed of deep reinforcement learning algorithms, increase the reward value obtained during training, and thus generate a more reasonable trajectory.
[0084] This invention generates trajectories with multi-state information through deep reinforcement learning, further optimizing the originally discrete path points into smooth trajectories and simultaneously reducing the probability of collisions with obstacles. Trajectories containing multiple state information also help improve the stability and accuracy of control.
[0085] This invention is designed for indoor environments. It achieves optimality and smoothness in the planned path, and uses global path planning to connect trajectory planning, ensuring optimal path and minimum energy consumption.
[0086] This invention achieves stable control of UAVs by constructing an optimization function through model prediction. The control method with future prediction properties can achieve better and more stable control of UAVs. At the same time, the introduction of an input smoothing term penalty ensures that the control input does not change too much in adjacent time domains, thus achieving the smoothness and safety of UAV control.
[0087] This invention solves the technical problems of easy collisions during drone planning and poor system control stability in the prior art.
[0088] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A UAV navigation method integrating reinforcement learning and model prediction, characterized in that, The method includes: S1. Utilize autonomous exploration and LiDAR to collect environmental information and establish a three-dimensional point cloud map of the environment; S2. At the preset height of the target point, extract the global two-dimensional planar map of the global map projection; S3. Search for the optimal global path on the global two-dimensional plane map; S4. Based on the optimal global path, a smooth trajectory is obtained, wherein the information of the smooth trajectory includes: UAV position, UAV speed, and UAV attitude. S5. Based on the smooth trajectory, perform trajectory tracking operation on the UAV through model prediction control optimization.
2. The UAV navigation method integrating reinforcement learning and model prediction according to claim 1, characterized in that, S1 includes: When entering an indoor environment, the drone uses its onboard LiDAR to collect data; during the drone's movement, the drone's pose is estimated in real time and the LiDAR data is fused to generate a local point cloud map; a global map is constructed by using an incremental map construction method to fuse the new point cloud data with the current global map to obtain a global point cloud map. By processing the global point cloud map, dynamic objects and noise are removed, and the global map is filtered and optimized using point cloud processing tools. Loop closure detection is performed to construct a globally consistent static map. The map density and accuracy are controlled by adjusting the parameters to obtain an appropriate map density, thus obtaining an indoor environment point cloud map.
3. The UAV navigation method integrating reinforcement learning and model prediction according to claim 1, characterized in that, S2 includes: A two-dimensional rasterized map is obtained; let the map data be matrix M, with... This represents the value of the raster cell in the i-th row and j-th column; using a contour line extraction algorithm, contour lines are extracted from the map data to complete the information creation operation of the contour 2D map; The system identifies obstacles and target areas, analyzes the elevated 2D map, records the location and size of closed areas to represent obstacles, and judges and extracts the obstacles based on area threshold conditions to obtain the information data of the obstacles. The drone is navigated based on the contour 2D map.
4. The UAV navigation method integrating reinforcement learning and model prediction according to claim 1, characterized in that, S3 includes: The global two-dimensional planar map is constructed into a vector map, and the optimal global path is found using an improved global path planning algorithm. The goal is to find an optimal path without collisions and to perform micro-layout processing on obstacles. By using global path planning, a feasible path at the expected height is obtained. The desired height is then extended to the feasible path to obtain a three-dimensional global path.
5. The UAV navigation method integrating reinforcement learning and model prediction according to claim 1, characterized in that, S4 includes: Training yields suitable network parameters; Using the applicable network parameters for model inference, the UAV inputs environmental information into the policy network to obtain an action, and updates the state through the action; Multi-state information trajectories are generated through deep reinforcement learning, and discrete path points are optimized to obtain the smooth trajectory.
6. The UAV navigation method integrating reinforcement learning and model prediction according to claim 5, characterized in that, In the process of training to obtain applicable network parameters, the neural network parameters are randomly initialized. Initialize the drone's state values; The UAV observation information is acquired, fed into the reinforcement learning framework for training, and the action and reward are obtained. The state observations input into the deep reinforcement learning framework include the UAV state value, the UAV's external point cloud information value, and the temporary target point relative to the global path. The state, action, reward, and updated value after each state transition are placed into the experience pool; the state update of the drone agent is performed using the following logic: Mecha speed update ; Updates to yaw and pitch angles: ; Location update: ; Combine the updated state value with the previous one. Add them together to the experience pool; When the total step size reaches the update frequency, a bacth of experience is taken from the experience pool, and the network parameters are updated using priority experience replay; Jump to the operation of obtaining the action and reward, and execute it in a loop until the applicable network parameters are obtained.
7. The UAV navigation method integrating reinforcement learning and model prediction according to claim 1, characterized in that, Calculate the desired roll angle based on the desired yaw angle: Status obtained: When the target point is reached, at least two time series consisting of states are obtained, which are used as the trajectory obtained from the planning.
8. The UAV navigation method integrating reinforcement learning and model prediction according to claim 1, characterized in that, S5 includes: Based on the dynamics and kinematics model of the UAV, a state transition equation is established, and a unified UAV controller is built to control the position and attitude of the UAV; wherein, the continuous-time dynamics model is discretized to obtain a discrete-time state-space model. Define the prediction time domain and the control time domain. The prediction time domain predicts the state for the next Np steps. The control time domain, Nc, optimizes the control input for the next Nc steps. Construct an objective function, wherein the penalty term of the objective function includes: a penalty term for the current state and tracking error, a smoothing penalty term for the UAV thrust and torque input, and a penalty term for the terminal error in the nth future time domain; Based on the objective function, constraints are added, including: UAV state constraints, control input constraints, and dynamic constraints; Under the constraints, the OSQP solver library is invoked to simplify the UAV nonlinear model into a linear model.
9. The UAV navigation method integrating reinforcement learning and model prediction according to claim 8, characterized in that, Construct the objective function using the following logic: In the formula, Q, R, and P are weight matrices, corresponding to the weights of state error, control input, and terminal state, respectively.
10. A drone navigation system integrating reinforcement learning and model prediction, characterized in that, The system includes: The map building module is used to autonomously explore using LiDAR to collect environmental information and build a 3D point cloud map of the environment; A two-dimensional map acquisition module is used to extract a global two-dimensional planar map of the global map projection at a preset height of the target point. The two-dimensional map acquisition module is connected to the map creation module. An optimal path search module is used to search for the optimal global path on the global two-dimensional plane map. The optimal path search module is connected to the two-dimensional map acquisition module. A smoothing module is used to process the optimal global path to obtain a smooth trajectory, wherein the information of the smooth trajectory includes: UAV position, UAV speed and UAV attitude, and the smoothing module is connected to the optimal path search module; An estimation tracking module is used to perform trajectory tracking operations on the UAV based on the smoothed trajectory and through model prediction control optimization. The estimation tracking module is connected to the smoothing processing module.
Citation Information
Patent Citations
Deep reinforcement learning accelerated training method for unmanned aerial vehicle image navigation
CN115345281A
Robot assembly optimization method and system based on federal learning
CN119458337A
Shape adaptive planning and control method for deformable unmanned aerial vehicle
CN120595853A
Unmanned aerial vehicle control method and device, electronic equipment and readable storage medium
CN120686872A
Unmanned aerial vehicle cluster polarization-space-time coding anti-interference communication system based on dynamic topology reconstruction
CN120769324A
Cited By
Underwater robot tracking control method and system fusing reinforcement learning navigation and model prediction control
CN122131811A