Reinforcement learning method for unmanned aerial vehicle path planning based on delay experience priority playback mechanism
By employing a delayed experience-first replay mechanism and reinforcement learning methods, the problem of insufficient real-time trajectory planning for UAVs in complex battlefield environments was solved, achieving millisecond-level trajectory planning and improving the success rate and real-time performance of UAV planning in dense obstacle environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2023-08-10
- Publication Date
- 2026-07-21
AI Technical Summary
Unmanned aerial vehicles (UAVs) struggle to perceive the overall situation in complex and highly dynamic battlefield environments, resulting in insufficient real-time performance in trajectory planning calculations. This increases the risk of trajectory planning failures and fails to meet the requirements of complex battlefield combat missions.
A reinforcement learning-based trajectory planning method based on delayed experience-first replay mechanism is adopted. By constructing a Markov decision process, combining the priority experience replay mechanism and the maximum entropy strategy deep reinforcement learning network model, and introducing a phased training and delayed update mechanism, the training speed and stability are improved, and millisecond-level trajectory planning is achieved.
It improves the success rate and real-time performance of UAV trajectory planning in dense obstacle environments, achieving millisecond-level trajectory planning and meeting the real-time computing needs in dynamic scenarios.
Smart Images

Figure CN116974299B_ABST