An Unmanned Aerial Vehicle (UAV) Autonomous Navigation Method and System Based on Deep Reinforcement Learning

By introducing deep reinforcement learning methods in the autonomous navigation of drones, a state space and dynamic reward function containing angle change information is constructed, which solves the problem of insufficient navigation capabilities of drones in complex environments, and achieves efficient and robust autonomous navigation performance.

CN118550320BActive Publication Date: 2025-06-13HARBIN INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410557434.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-06-13
Estimated Expiration
2044-05-07

AI Technical Summary

Technical Problem

The autonomous navigation capabilities of the drone in complex environments are insufficient, the training model converges slowly, and it is difficult to balance exploration and utilization behavior, and the dynamic change trend of the agent is not considered.

Method used

A drone autonomous navigation method based on deep reinforcement learning is proposed. By constructing a state space containing angular change information, designing a dynamic reward function, and using an experience replay pool for training, to improve the autonomous navigation performance of the drone in high-density and high-dynamic environments.

Benefits of technology

It realizes autonomous path planning for drones in high-density and high-dynamic environments, improves navigation success rate and flight efficiency, balances exploration and utilization behavior, and adapts to the dynamic trend of agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118550320B_ABST
    Figure CN118550320B_ABST
Patent Text Reader

Abstract

A method and system for autonomous navigation of unmanned aerial vehicles based on deep reinforcement learning, which relates to the field of autonomous navigation of unmanned aerial vehicles. It solves the problems of slow convergence speed of the training model of unmanned aerial vehicles, the balance between conservative and aggressive behaviors of unmanned aerial vehicles, and the influence of the dynamic change trend of agents on the navigation task not being considered. The method includes: obtaining the current state information and the state information at the previous moment; constructing the state space of the unmanned aerial vehicle at the current moment according to the state information at the current moment, the state information at the previous moment, and the information of the unmanned aerial vehicle itself at the current moment; obtaining the action space at the current moment; the unmanned aerial vehicle executes the desired action and obtains a reward, and at the same time updates the state information at the next moment; constructing the state space at the next moment; storing the state space, action space, reward, and state space at the next moment at the current moment as samples in the experience replay pool; training the autonomous navigation model of the unmanned aerial vehicle; and performing the autonomous navigation task of the unmanned aerial vehicle according to the trained model. It is applied to the industrial and agricultural fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous navigation of unmanned aerial vehicles, and particularly to an autonomous navigation method for unmanned aerial vehicles based on deep reinforcement learning. Background Art

[0002] In recent years, thanks to the rapid development of control algorithms and computing platforms, a series of progress has been made in the research of unmanned aerial vehicles. Due to their high efficiency, portability, and economy, unmanned aerial vehicles have been widely used in many fields, such as data collection, remote rescue, environmental detection, and pesticide spraying. In most tasks performed by unmanned aerial vehicles, their core role can be defined as flying from a starting point to a target point, and during this process, the unmanned aerial vehicle can achieve automatic navigation and obstacle avoidance. However, in real-world scenarios, high-density and high-dynamic working environments such as woods and pedestrians pose great challenges to the autonomous navigation and obstacle avoidance of unmanned aerial vehicles, which requires the unmanned aerial vehicle to be able to effectively identify obstacles and respond in a timely manner. Therefore, how to improve the autonomous navigation ability of unmanned aerial vehicles in complex environments and ensure the safe flight of unmanned aerial vehicles has become the core issue of concern for scholars and the industrial community.

[0003] The autonomous navigation scheme of unmanned aerial vehicles based on deep reinforcement learning has been proven to be effective in many studies, but there are still some problems that need to be further explored. These mainly include the following aspects:

[0004] (1) Sparse rewards: When training the model, the agent only receives a reward when it reaches the target point, which usually causes the agent to be difficult to effectively learn the correct behavior strategy during the exploration process, resulting in a slow convergence rate or even non-convergence of the trained model.

[0005] (2) Exploration and exploitation: The exploration of the agent means trying new behaviors to discover possible high-reward states, while exploitation is to select the current optimal behavior based on known information, which corresponds to the conservative and aggressive behaviors of the unmanned aerial vehicle. How to balance exploration and exploitation requires further research.

[0006] (3) State space: The selection of the state space is crucial for the deep reinforcement learning model. Most agents make decisions only based on the information at the current moment, while ignoring the dynamic change trend of the agent (such as the translation and rotation of the unmanned aerial vehicle). Summary of the Invention

[0007] In view of the problems of slow convergence rate of the unmanned aerial vehicle training model, the balance between the conservative and aggressive behaviors of the unmanned aerial vehicle, and the influence of the dynamic change trend of the agent on the navigation task not being considered, the present invention proposes an autonomous navigation method for unmanned aerial vehicles based on deep reinforcement learning, and the method includes:

[0008] S1: The unmanned aerial vehicle obtains the current state information and the state information of the previous moment based on the on-board sensor;

[0009] S2: Construct the state space of the UAV at the current moment based on the state information at the current moment, the state information at the previous moment, and the UAV's own information at the current moment;

[0010] S3: The UAV obtains the action space at the current moment based on the current state space;

[0011] S4: The UAV executes the desired action and obtains a reward, while updating the state information at the next moment;

[0012] S5: Repeat step S1 and step S2 to construct the state space at the next moment;

[0013] S6: Store the state space, action space, reward, and state space at the next moment at the current moment as a sample in the experience replay pool for training the UAV autonomous navigation model;

[0014] S7: Repeat the actions in step S1 to step S6 to train the UAV autonomous navigation model;

[0015] S8: According to the trained UAV autonomous navigation model, repeat the actions in step S1 to step S5 to execute the UAV's autonomous navigation task.

[0016] Furthermore, a preferred method is also proposed. The step S1 includes:

[0017]

[0018] where, represents the measurement value of the i-th laser beam in the lidar at time t, and θ i is the angle between the laser beam and the heading of the first view of the UAV, and α represents the yaw angle of the UAV.

[0019] Furthermore, a preferred method is also proposed. The on-board sensor is a lidar, and the field of view angle of the lidar is 4π / 3, and the angular resolution is π / 18.

[0020] Furthermore, a preferred method is also proposed. The state space of the UAV at the current moment in step S2 is:

[0021]

[0022] where, s t represents the state space of the UAV at time t, represents the UAV's own state information, which consists of the linear velocity v lin , angular velocity v yaw , linear acceleration a lin , angular acceleration a yaw, the distance d between the drone and the target ut and the angle β between the first view of the drone and the line connecting the drone and the target point, Δο t represents the difference between the state information of the drone at the current moment and the state information at the previous moment:

[0023]

[0024] Among them, represents the measurement value of the i-th laser beam in the lidar at time t-1, α t-1 represents the yaw angle of the drone at time t-1, α t represents the yaw angle of the drone at time t, Δα t represents the difference in the yaw angle of the drone between adjacent moments;

[0025] The state space of the drone includes static information and dynamic information for perceiving the environment, where the static information represents the observation data at the current moment, and the dynamic information represents the movement trend of the drone relative to the obstacle.

[0026] Furthermore, a preferred method is also proposed. The method further includes restricting the actual flight speed of the drone:

[0027] v lin ∈[0,0.5], v yaw ∈[-1,1].

[0028] Furthermore, a preferred method is also proposed. The reward function in step S4 is expressed as:

[0029] R = r dis + r arr + r cra + r laser + r step + r lin + r yaw

[0030] Among them, r dis represents the distance reward, r arr represents the reward for reaching the target, r cra represents the reward for colliding with an obstacle, r laser represents the free space reward, r step represents the step reward, r lin represents the linear velocity reward and r yaw represents the yaw angular velocity reward.

[0031] Furthermore, a preferred method is also proposed. The specific reward function is:

[0032]

[0033]

[0034]

[0035] and

[0036]

[0037] r laser =-∑ i (1 - d i / max(ο)) 4

[0038] r step =-ε 3 *N step ,and

[0039]

[0040] r lin =-ε 4 *|a lin |

[0041] r yaw =-ε 5 *|a yaw |

[0042] wherein, N step represents the number of steps executed by the UAV, ε 1 to ε 5 are positive coefficients, r 1 and r 2 represent constants set according to the task, E step represents the current training round, and E max represents the set maximum number of rounds.

[0043] Based on the same inventive concept, the present invention also provides an autonomous navigation system for UAVs based on deep reinforcement learning, and the system includes:

[0044] A state information acquisition unit, configured to enable the UAV to obtain the current state information and the state information at the previous moment based on on-board sensors;

[0045] A state space construction unit, configured to construct the state space of the UAV at the current moment according to the state information at the current moment, the state information at the previous moment, and the self-information of the UAV at the current moment;

[0046] An action space acquisition unit, configured to enable the UAV to obtain the action space at the current moment based on the current state space;

[0047] A reward unit, which is used to construct a reward function to encourage the drone to perform desired actions and update the state information at the next moment;

[0048] A loop unit, which is used to repeat the state information acquisition unit and the state space construction unit to construct the state space at the next moment;

[0049] A storage unit, which is used to store the state space, action space, reward, and state space at the next moment at the current moment as a sample in the experience replay pool for training the drone autonomous navigation model;

[0050] A training unit, which is used to repeat the actions from the state information acquisition unit to the storage unit to train the drone autonomous navigation model;

[0051] A navigation unit, which is used to repeat the actions from the state information acquisition unit to the state information acquisition unit according to the trained drone autonomous navigation model to perform the autonomous navigation task of the drone.

[0052] Based on the same inventive concept, the present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes a drone autonomous navigation method based on deep reinforcement learning as described in any one of the above.

[0053] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, which is used to store a computer program. The computer program executes a drone autonomous navigation method based on deep reinforcement learning as described in any one of the above.

[0054] The beneficial effects of the present invention are as follows:

[0055] The present invention solves the problems of slow convergence speed of the drone training model, the balance between the conservative and aggressive behaviors of the drone, and the problem of not considering the influence of the dynamic change trend of the agent on the navigation task.

[0056] A drone autonomous navigation method based on deep reinforcement learning proposed by the present invention realizes autonomous path planning of the drone in a high-density and high-dynamic environment. First, by analyzing the influence of the position change and angle change of the drone in a complex flight environment on the flight performance, a state space representation method including angle change information is proposed. Then, to balance the exploration and conservative behaviors of the agent, a dynamic reward function is constructed. Finally, a high-dynamic and high-density simulation environment is built to verify the proposed algorithm. The results show that the proposed algorithm can have the highest navigation success rate and can improve the navigation efficiency of the drone while improving the autonomous navigation performance.

[0057] The present invention takes into account the influence of the position change and angle change of the unmanned aerial vehicle (UAV) on the navigation performance in complex scenarios, introduces the angle change information as the input of the deep reinforcement learning model, and expands the thinking for the research of the navigation algorithm based on lidar.

[0058] Based on the non-sparse reward function, the present invention constructs a dynamic reward function to balance the conservative behavior and exploration behavior of the agent during the model training process, and improves the flight efficiency while enhancing the navigation success rate of the UAV.

[0059] The present invention is applied to the industrial and agricultural fields. Description of the Drawings

[0060] Figure 1 Flowchart of an unmanned aerial vehicle autonomous navigation method based on deep reinforcement learning according to Embodiment 1;

[0061] Figure 2 Network structure of the unmanned aerial vehicle autonomous navigation model according to Embodiment 1, where State represents the state space, Sample represents random sampling, and Action represents the action space;

[0062] Figure 3 Schematic diagram of the training scenario according to Embodiment 11;

[0063] Figure 4 Schematic diagram of the reward convergence curve obtained by training four algorithms according to Embodiment 11;

[0064] Figure 5 Schematic diagram of the high-density test scenario according to Embodiment 11, where Figure 5 (a) represents the training scenario containing 90 obstacles, Figure 5 (b) represents the training scenario containing 120 obstacles, Figure 5 (c) represents the training scenario containing 150 obstacles, Figure 5 (d) represents the training scenario containing 180 obstacles, Figure 5 (e) represents the training scenario containing 210 obstacles;

[0065] Figure 6 Average flight distance and average flight steps in the high-density scenario according to Embodiment 11, where Figure 6 (a) represents the average flight distance when the four algorithms complete the task in different test scenarios, Figure 6 (b) represents the average flight distance steps when the four algorithms complete the task in different test scenarios;

[0066] Figure 7 Schematic diagram of the high-dynamic test scenario according to Embodiment 11;

[0067] Figure 8 The average flight distance and average number of flight steps in a high-dynamic scenario described in Embodiment XI, where Figure 8 (a) represents the average flight distance when four algorithms complete tasks in different test scenarios, Figure 8 (b) represents the average number of flight steps when four algorithms complete tasks in different test scenarios. Specific Embodiment

[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0069] Embodiment 1. Refer to Figure 1 and Figure 2 to illustrate this embodiment. A method for autonomous navigation of an unmanned aerial vehicle based on deep reinforcement learning according to this embodiment includes:

[0070] S1: The unmanned aerial vehicle obtains the current state information and the state information of the previous moment based on on-board sensors;

[0071] S2: Construct the state space of the unmanned aerial vehicle at the current moment according to the state information at the current moment, the state information of the previous moment, and the self-information of the unmanned aerial vehicle at the current moment;

[0072] S3: The unmanned aerial vehicle obtains the action space at the current moment based on the current state space;

[0073] S4: The unmanned aerial vehicle executes the desired action and obtains a reward, and at the same time updates the state information of the next moment;

[0074] S5: Repeat steps S1 and S2 to construct the state space of the next moment;

[0075] S6: Store the state space, action space, reward, and state space of the next moment at the current moment as a sample in the experience replay pool for training the autonomous navigation model of the unmanned aerial vehicle;

[0076] S7: Repeat the actions in steps S1 to S6 to train the autonomous navigation model of the unmanned aerial vehicle;

[0077] S8: According to the trained autonomous navigation model of the unmanned aerial vehicle, repeat the actions in steps S1 to S5 to execute the autonomous navigation task of the unmanned aerial vehicle.

[0078] A method for autonomous navigation of unmanned aerial vehicles (UAVs) based on deep reinforcement learning proposed in this embodiment enables UAVs to learn from experience and continuously optimize the autonomous navigation model by using deep reinforcement learning, thereby accelerating the training convergence speed. By constructing a reward function, the UAV is encouraged to perform desired actions while avoiding overly conservative or aggressive behaviors, enabling the UAV to more effectively complete the navigation task while ensuring safety. In different environments and tasks, the dynamic change trends of the intelligent agent may vary. Through the experience replay pool and the continuously updated state space, the UAV can adapt to different environments and tasks and continuously optimize the autonomous navigation model to adapt to the dynamic change trends. The method proposed in this embodiment is based on the idea of deep reinforcement learning. The UAV continuously learns through interaction with the environment and optimizes the autonomous navigation model according to the guidance of the reward function. Through the experience replay pool, the UAV can reuse previous experiences, thereby improving the learning efficiency and the stability of the model. Finally, the UAV can efficiently and safely complete the navigation task in different environments according to the trained autonomous navigation model.

[0079] A method for autonomous navigation of unmanned aerial vehicles (UAVs) based on deep reinforcement learning proposed in this embodiment realizes autonomous path planning of UAVs in high-density and high-dynamic environments. First, by analyzing the influence of the position change and angle change of the UAV in a complex flight environment on flight performance, a state space representation method including angle change information is proposed. Then, to balance the exploration and conservative behaviors of the intelligent agent, a dynamic reward function is constructed. Finally, a high-dynamic and high-density simulation environment is built to verify the proposed algorithm. The results show that the proposed algorithm can have the highest navigation success rate and can improve the navigation efficiency of the UAV while enhancing the autonomous navigation performance.

[0080] Embodiment 2: This embodiment further limits a method for autonomous navigation of unmanned aerial vehicles (UAVs) based on deep reinforcement learning described in Embodiment 1. The step S1 includes:

[0081]

[0082] Among them, represents the measurement value of the i-th laser beam in the lidar at time t, and θ i is the angle between the laser beam and the heading of the first perspective of the UAV, and α represents the yaw angle of the UAV.

[0083] In this embodiment, the current state information and the state information of the previous moment are obtained through on-board sensors, including the measurement values of lidar and the yaw angle of the UAV, etc., which can provide more accurate and comprehensive state information and help the UAV perceive the surrounding environment more accurately. By using a variety of sensor information, such as lidar and yaw angle, etc., the state of the UAV is comprehensively considered. Such comprehensive information acquisition helps to improve the UAV's environmental perception ability, thereby improving the accuracy and reliability of autonomous navigation. The UAV can obtain and process the information of on-board sensors in real time, ensuring the timeliness and continuity of state information, enabling the UAV to quickly respond in a dynamic environment, thus enhancing the real-time performance and flexibility of autonomous navigation. Obtaining more accurate and comprehensive state information helps to establish a more accurate state space, thereby providing a more reliable data basis for the training of the deep reinforcement learning model and accelerating the convergence speed and performance improvement of the model.

[0084] Embodiment 3: This embodiment further limits a method for autonomous navigation of a UAV based on deep reinforcement learning described in Embodiment 2. The on-board sensor is a lidar, and the field of view angle of the lidar is 4π / 3, and the angular resolution is π / 18.

[0085] In this embodiment, a lidar is used as the on-board sensor, which can provide more accurate and high-resolution environmental perception data. The lidar can obtain obstacle information in the environment in real time at different distances and directions, thereby providing more comprehensive and accurate state information for the UAV, which helps to improve the accuracy and safety of autonomous navigation.

[0086] In this embodiment, the field of view angle of the lidar is 4π / 3, and the angular resolution is π / 18. This means that the lidar can cover a wider range and has a higher angular resolution. Such an optimized design enables the UAV to more comprehensively perceive the surrounding environment, accurately identify the position and shape of obstacles, and improve the safety and reliability of navigation. Further, the high resolution and wide field of view angle of the lidar enable the UAV to adapt to complex and changing environments, including different terrains and climate conditions such as urban areas, forests, and mountains. This enables the UAV to perform effective autonomous navigation and complete various tasks in various challenging environments.

[0087] Embodiment 4: This embodiment further limits a method for autonomous navigation of a UAV based on deep reinforcement learning described in Embodiment 2. The state space of the UAV at the current moment in step S2 is:

[0088]

[0089] where s t represents the state space of the UAV at time t, Indicates the state information of the UAV itself, including the linear velocity v of the UAV lin , angular velocity v yaw , linear acceleration a lin , angular acceleration a yaw , the distance d between the UAV and the target ut and the angle β between the first view of the UAV and the line connecting the UAV and the target point, Δο t Indicates the difference between the state information of the UAV at the current moment and the state information at the previous moment:

[0090]

[0091] Among them, Indicates the measurement value of the i-th laser beam in the lidar at time t - 1, α t-1 Indicates the yaw angle of the UAV at time t - 1, α t Indicates the yaw angle of the UAV at time t, Δα t Indicates the difference in the yaw angle of the UAV between adjacent moments;

[0092] The state space of the UAV includes static information and dynamic information for perceiving the environment. The static information represents the observed data at the current moment, and the dynamic information represents the movement trend of the UAV relative to the obstacles.

[0093] By combining the state information of the UAV with the measurement values of the lidar, a more comprehensive and complete state space is formed. This comprehensive design can better reflect the current environmental perception situation and the self-movement state of the UAV, providing richer and more accurate input data for the deep reinforcement learning model, thereby improving the accuracy and stability of navigation.

[0094] Taking the movement trend of the UAV relative to the obstacles as part of the state space enables the model to better understand the movement of the obstacles in the environment, thereby better planning the path and making responses. This design considering dynamic information enables the UAV to more flexibly respond to complex environmental changes, improving the robustness and adaptability of navigation.

[0095] By using the difference between the state information at the current moment and the state information at the previous moment to represent the state space of the UAV, the model can more easily capture the changes and trends between states, improving the efficiency and performance of the model.

[0096] Combining the static information and dynamic information for perceiving the environment enables the UAV to more comprehensively understand the environment and make corresponding navigation decisions. The static information provides the environmental observation data at the current moment, while the dynamic information provides the movement trend of the obstacles. The combination of the two can help the UAV better plan the path, avoid collisions and fly safely.

[0097] Embodiment 5. This embodiment further limits a method for autonomous navigation of an unmanned aerial vehicle based on deep reinforcement learning described in Embodiment 4. The method further includes restricting the actual flight speed of the unmanned aerial vehicle:

[0098] v lin ∈[0, 0.5], v yaw ∈[-1, 1].

[0099] By restricting the actual flight speed of the unmanned aerial vehicle, this embodiment can ensure that the unmanned aerial vehicle maintains a safe speed range during navigation. This restriction can effectively reduce the risks during flight, reduce the occurrence of accidental collisions or emergencies, thereby enhancing the safety of navigation.

[0100] Embodiment 6. This embodiment further limits a method for autonomous navigation of an unmanned aerial vehicle based on deep reinforcement learning described in Embodiment 1. The reward function in step S4 is expressed as:

[0101] R = r dis + r arr + r cra

[0102] + r laser + r step + r lin + r yaw

[0103] where r dis represents the distance reward, r arr represents the reward for reaching the target, r cra represents the reward for colliding with an obstacle, r laser represents the free space reward, r step represents the step reward, r lin represents the linear velocity reward and r yaw represents the yaw angular velocity reward.

[0104] In this embodiment, by designing a reward function, the drone can be effectively guided to learn good navigation strategies. The distance reward, the target-reached reward, and the free-space reward can encourage the drone to maintain an appropriate distance from the target, reach the target successfully, and preferably choose a safe and collision-free flight path, thereby improving the navigation efficiency and safety. The reward function includes a penalty term for colliding with obstacles, which helps the drone avoid collisions with obstacles and ensure flight safety. By penalizing collision behaviors, the drone will tend to choose paths that avoid obstacles, thus reducing the occurrence of accidental collisions. This reward function covers multiple aspects of rewards and penalties, including distance, speed, yaw angular velocity, etc., and can comprehensively guide the drone to learn effective navigation strategies. This multi-dimensional reward design enables the drone to learn more comprehensive and robust navigation strategies to adapt to different environmental and task requirements. The introduction of the step reward and the linear velocity reward helps improve the learning efficiency and the accuracy of speed control. The step reward can prompt the drone to reach the target as soon as possible, while the linear velocity reward helps control the flight speed of the drone to avoid flying too fast or too slow.

[0105] Embodiment 7. This embodiment further limits a drone autonomous navigation method based on deep reinforcement learning described in Embodiment 6, and the reward is specifically:

[0106] and

[0107]

[0108]

[0109] and

[0110]

[0111] r laser =-∑ i (1 - d i / max(ο)) 4

[0112] r step =-ε 3 *N step ,and

[0113]

[0114] r lin =-ε 4 *|a lin |

[0115] r yaw =-ε 5 *|ayaw |

[0116] Among them, N step represents the number of steps executed by the drone, and ε 1 to ε 5 are positive coefficients, r 1 and r 2 represent constants set according to the task, E step represents the current training round, and E max represents the set maximum number of rounds.

[0117] The main function of the reward function is to convey the goal to the agent, that is, to maximize the expected probability of the cumulative sum of the scalar reward signals received by the agent. The design of the reward function directly determines the performance evaluation of the agent in specific actions, and thus affects the performance and convergence speed of the reinforcement learning algorithm. In this embodiment, a dynamic reward function is constructed based on the non-sparse reward function to balance the conservative behavior and exploration behavior of the agent during the model training process, and improve the flight efficiency while increasing the success rate of the drone navigation.

[0118] Embodiment 8. A drone autonomous navigation system according to the present embodiment, the system includes:

[0119] A state information acquisition unit for the drone to obtain the current state information and the state information of the previous moment based on the on-board sensor;

[0120] A state space construction unit for constructing the state space of the drone at the current moment according to the state information of the current moment, the state information of the previous moment, and the self-information of the drone at the current moment;

[0121] An action space acquisition unit for the drone to obtain the action space at the current moment based on the current state space;

[0122] A reward unit for constructing a reward function to encourage the drone to execute the desired action and update the state information of the next moment;

[0123] A loop unit for repeating the state information acquisition unit and the state space construction unit to construct the state space of the next moment;

[0124] A storage unit for storing the state space, action space, reward, and state space of the next moment at the current moment as a sample in the experience replay pool for training the drone autonomous navigation model;

[0125] A training unit for repeating the actions from the state information acquisition unit to the storage unit to train the drone autonomous navigation model;

[0126] The navigation unit is used to repeat the actions from the state information acquisition unit to the state information acquisition unit according to the trained UAV autonomous navigation model, and execute the autonomous navigation task of the UAV.

[0127] Embodiment 9. A computer device described in this embodiment includes a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes a method for UAV autonomous navigation based on deep reinforcement learning as described in any one of Embodiments 1 to 7.

[0128] Embodiment 10. A computer-readable storage medium described in this embodiment is used to store a computer program, and the computer program executes a method for UAV autonomous navigation based on deep reinforcement learning as described in any one of Embodiments 1 to 7.

[0129] Embodiment 11. Refer to Figures 3 to 8 Describe this embodiment. This embodiment provides a specific example for a method for UAV autonomous navigation based on deep reinforcement learning described in Embodiment 1, and is also used to explain Embodiments 2 to 7. Specifically:

[0130] To verify the method proposed by the present invention, a simulation scenario with high-density and high-dynamic obstacles was constructed based on Gazebo. Its training scenario is as Figure 3 shown. There are 150 obstacles in a 20×20 rectangular field. In each episode of training, the coordinates of all obstacles are randomly generated, and the target point is randomly selected from a, b, c, and d. The flight task of the UAV is defined as flying from the origin to the target point within the specified number of steps. The criteria for the end of each episode of training include reaching the target point, encountering an obstacle, and the UAV neither reaching the target nor colliding within the maximum number of steps.

[0131] To verify the advantages of the method proposed by the present invention, the DDPG, TD3, and SAC algorithms were used for comparison respectively. In addition, to objectively evaluate the performance of the algorithms, the following five quantitative indicators were selected to measure the completion of the task:

[0132] (1) Success rate R s : The ratio of the number of times the UAV completes the navigation task to the total number of trials;

[0133] (2) Collision rate R c : The ratio of the number of times the UAV encounters an obstacle to the total number of trials;

[0134] (3) Loss rate R l : The ratio of the number of times the UAV neither reaches the target point nor collides within the specified number of steps to the total number of attempts;

[0135] (4) Average flight distance L dis : The average flight distance when the UAV successfully completes the mission;

[0136] (5) Average flight steps N step : The average number of steps required for the UAV to successfully complete the mission.

[0137] Among the above indicators, the larger R s , the better the autonomous navigation performance of the UAV. The smaller L dis and N step , the higher the flight efficiency of the UAV.

[0138] In the training scenario as shown in Figure 3 , the method proposed in the present invention and three comparison algorithms (DDPG, TD3, SAC) are trained. For objective comparison, the hyperparameters and training settings of the four algorithms are kept consistent. In the reward function, r 1 = 200, r 2 = 300, E max = 400. The results of the cumulative rewards of different algorithms varying with the number of training sets are as shown in Figure 4 . All the data in the figure have been smoothed.

[0139] According to Figure 4 shown, the method proposed in the present invention has a faster convergence speed and higher rewards than the other three algorithms. Among them, the method proposed in the present invention starts to converge around the 150th episode, SAC and DDPG start to converge around the 200th episode, while TD3 starts to converge around the 300th episode and has the lowest rewards. Therefore, the training results show that the method proposed in the present invention has better convergence.

[0140] To verify the effect of the proposed algorithm, the UAV autonomous navigation model proposed in the present invention is tested in a simulation scenario with different numbers of obstacles. The test scenario is as shown in Figure 5 , including five scenarios with different numbers of obstacles. The UAVs based on the four algorithms respectively execute 100 navigation tasks in five scenarios with different obstacle densities. Similar to the training process, the positions of the obstacles are randomly regenerated and the target points are reselected after each flight mission ends.

[0141] The results of the navigation success rate, collision rate, and loss rate of the four algorithms under different obstacle densities are shown in Table 1. As can be seen from the table, the method proposed in the present invention achieved the highest success rate and the lowest collision rate in all five scenarios, and the navigation performance of the SAC algorithm was better than that of DDPG and TD3. Due to the high obstacle density in the experimental scenarios, flight loss occurred only in DDPG. Comparing the experimental results of Scenario 1 and Scenario 5, it can be seen that as the obstacle density increased significantly, the success rate of our algorithm decreased by only 18%, while the success rates of the other three algorithms decreased by 28%, 48%, and 29% respectively, which further demonstrated the stability of the method proposed in the present invention.

[0142] Table 1 Experimental results under different density scenarios

[0143]

[0144] Based on the flight trajectory data, the average flight distance and average number of flight steps for all experiments that completed the navigation task were calculated as Figure 6 shown.

[0145] According to Figure 6 (a), the method proposed in the present invention had the shortest average flight distance in Scenarios 2 and 3, and had an average flight distance only greater than that of SAC in Scenarios 1 and 5. According to Figure 6 (b), the method proposed in the present invention required the fewest average number of flight steps in all five scenarios. Therefore, based on the above analysis, it can be seen that compared with the other three classic algorithms, the method proposed in the present invention not only had the highest navigation success rate and stability in a high-density flight environment but also had the optimal flight efficiency.

[0146] To further verify the performance of the trained algorithm, a high-dynamic test scenario was constructed as Figure 7 shown. As shown in the figure, the starting coordinates of the UAV were [0, -7], the target point coordinates were [0, 7], and the flight task of the UAV was defined as autonomously flying from the starting point to the target point. Each time a test was conducted, 30 obstacles were randomly generated in the middle area between the starting point and the target point. The movement directions of the obstacles were marked by the arrows in the figure. After reaching the edge of the site, the obstacles would continue to move in the opposite direction until the end of this test. The movement speeds of the obstacles were set to five cases: 0.01 m / s, 0.05 m / s, 0.1 m / s, 0.15 m / s, and 0.2 m / s. Each group of experiments was conducted 100 times, and the positions of the obstacles would be regenerated after each test.

[0147] The experimental results of all algorithms in 5 dynamic scenarios are shown in Table 2. It can be seen from Table 2 that the method proposed in the present invention has obtained the highest navigation success rate and the lowest collision rate in the experiments with different moving speeds of obstacles, and there is no situation of flight loss, which indicates the effectiveness of the method proposed in the present invention in high-dynamic scenarios. As a comparison, flight loss occurred in all the other three algorithms, and the loss rates of SAC and DDPG were both relatively high. It should be noted that as the moving speed of the obstacle increases, the difficulty for the UAV to avoid the obstacle increases sharply, resulting in a significant reduction in the success rate of all algorithms.

[0148] Table 2 Experimental Results in Different Dynamic Scenarios

[0149]

[0150] Based on the flight trajectory of the UAV, the average flight distance and average number of flight steps of the UAV in the dynamic scenario are calculated as Figure 8 shown.

[0151] From Figure 8 it can be seen that the method proposed in the present invention has the shortest average flight distance and the fewest average number of flight steps in all 5 dynamic flight scenarios. Figure 8 It shows that the proposed algorithm not only has a high navigation success rate but also has the highest navigation efficiency when facing a high-dynamic flight environment, enabling the UAV to achieve autonomous navigation quickly and accurately.

[0152] The present invention proposes a UAV autonomous navigation method based on deep reinforcement learning, which realizes the autonomous path planning of the UAV in a high-density and high-dynamic environment. First, by analyzing the influence of the position change and angle change of the UAV in a complex flight environment on the flight performance, a state space representation method including angle change information is proposed. Then, to balance the exploration and conservative behaviors of the intelligent agent, a dynamic reward function is constructed. Finally, a high-dynamic and high-density simulation environment is built to verify the proposed method, and the results show that the proposed algorithm can have the highest navigation success rate and can improve the navigation efficiency of the UAV while enhancing the autonomous navigation performance.

[0153] Although the preferred embodiments of the present disclosure have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications falling within the scope of the present disclosure.

[0154] Obviously, those skilled in the art can make various modifications and variations to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure is also intended to include these modifications and variations.

[0155] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0156] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure rather than limit the scope of its protection. Although the present disclosure has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading the present disclosure, various changes, modifications or equivalent substitutions can still be made to the specific implementation manners of the invention, but these changes, modifications or equivalent substitutions are all within the scope of protection of the claims pending for publication.

Claims

1. A method for autonomous navigation of unmanned aerial vehicles based on deep reinforcement learning, characterized in that: The method comprises: S1: The drone obtains the current state information and the state information of the previous moment based on the onboard sensors; S2: Construct the current state space of the drone based on the current state information, the previous state information, and the drone’s own information at the current moment; S3: The drone obtains the action space at the current moment based on the current state space; S4: The drone performs the desired action and obtains a reward, while updating the state information for the next moment; S5: Repeat steps S1 and S2 to construct the state space at the next moment; S6: The current state space, action space, reward, and state space at the next moment are stored as a sample in the experience replay pool for training the UAV autonomous navigation model; S7: Repeat steps S1 to S6 to train the autonomous navigation model of the UAV; S8: According to the trained UAV autonomous navigation model, repeat the actions from step S1 to step S5 to perform the UAV autonomous navigation task; The step S1 comprises: in, Indicated in t At this moment, LiDAR i The measured value of a laser beam, is the angle between the laser beam and the first-person heading of the drone, Indicates the yaw angle of the drone; The state space of the drone at the current moment in step S2 is: in, Indicates that the drone is t The state space at time, Indicates the drone's own status information, which is determined by the drone's linear speed. , angular velocity , Linear Acceleration , angular acceleration , the distance between the drone and the target And the angle between the drone's first-person perspective and the line connecting the drone and the target point composition, Indicates the difference between the current state information of the drone and the state information of the previous moment: in, Indicated in t -1 moment laser radar i The measured value of a laser beam, Indicated in t- The yaw angle of the drone at moment 1, Indicated in t The yaw angle of the drone at the moment, Indicates the difference in the yaw angle of the drone at adjacent moments; The state space of the drone contains static and dynamic information of the perceived environment, where the static information represents the observation data at the current moment, and the dynamic information represents the movement trend of the drone relative to obstacles; The reward function in step S4 is expressed as: in, Indicates distance reward, Rewards for achieving goals, Represents the reward for colliding with an obstacle, represents the free space reward, represents the step reward, Represents line speed bonus and represents the yaw rate reward; The reward function is specifically: in, Indicates the number of steps executed by the drone, arrive is a positive coefficient, and Represents a constant set according to the task, Indicates the current training round number, Indicates the maximum number of rounds set.

2. The method for autonomous navigation of a drone based on deep reinforcement learning according to claim 1, characterized in that: The airborne sensor is a laser radar, and the field of view of the laser radar is , the angular resolution is π / 18.

3. The method for autonomous navigation of a drone based on deep reinforcement learning according to claim 1, characterized in that: The method also includes limiting the actual flight speed of the drone: 。 4. A UAV autonomous navigation system based on deep reinforcement learning, characterized in that: The system comprises: A state information acquisition unit, used by the UAV to obtain current state information and state information of the previous moment based on the onboard sensor; The state space construction unit is used to construct the state space of the drone at the current moment according to the state information at the current moment, the state information at the previous moment, and the drone's own information at the current moment; the state space of the drone at the current moment is: in, Indicates that the drone is t The state space at time, Indicates the drone's own status information, which is determined by the drone's linear speed. , angular velocity , Linear Acceleration , angular acceleration , the distance between the drone and the target And the angle between the drone's first-person perspective and the line connecting the drone and the target point composition, Indicates the difference between the current state information of the drone and the state information of the previous moment: in, Indicated in t -1 moment laser radar i The measured value of a laser beam, Indicated in t- The yaw angle of the drone at moment 1, Indicated in t The yaw angle of the drone at the moment, Indicates the difference in the yaw angle of the drone at adjacent moments; The state space of the drone contains static and dynamic information of the perceived environment, where the static information represents the observation data at the current moment, and the dynamic information represents the movement trend of the drone relative to obstacles; An action space acquisition unit is used for the drone to obtain the action space at the current moment based on the current state space; The reward unit is used to construct a reward function to encourage the drone to perform the desired action and update the state information at the next moment; the reward function is expressed as: in, Indicates distance reward, Rewards for achieving goals, Represents the reward for colliding with an obstacle, represents the free space reward, represents the step reward, Represents line speed bonus and represents the yaw rate reward; The reward function is specifically: in, Indicates the number of steps executed by the drone, arrive is a positive coefficient, and Represents a constant set according to the task, Indicates the current training round number, Indicates the maximum number of rounds set; A loop unit is used to repeat the state information acquisition unit and the state space construction unit to construct the state space at the next moment; A storage unit is used to store the current state space, action space, reward, and the next state space as a sample in the experience playback pool for training the UAV autonomous navigation model; A training unit, used to repeat the actions of the state information acquisition unit to the storage unit, and train the autonomous navigation model of the UAV; The navigation unit is used to repeat the actions of the state information acquisition unit to the state information acquisition unit according to the trained UAV autonomous navigation model to perform the UAV autonomous navigation task.

5. A computer device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes a method for autonomous navigation of a drone based on deep reinforcement learning as described in any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program executes the drone autonomous navigation method based on deep reinforcement learning as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Unmanned aerial vehicle relay navigation method based on deep reinforcement learning

    CN116817909A