Reinforcement learning method for unmanned aerial vehicle path planning based on delay experience priority playback mechanism

By employing a delayed experience-first replay mechanism and reinforcement learning methods, the problem of insufficient real-time trajectory planning for UAVs in complex battlefield environments was solved, achieving millisecond-level trajectory planning and improving the success rate and real-time performance of UAV planning in dense obstacle environments.

CN116974299BActive Publication Date: 2026-07-21BEIJING INST OF TECH
2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2023-08-10
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Unmanned aerial vehicles (UAVs) struggle to perceive the overall situation in complex and highly dynamic battlefield environments, resulting in insufficient real-time performance in trajectory planning calculations. This increases the risk of trajectory planning failures and fails to meet the requirements of complex battlefield combat missions.

Method used

A reinforcement learning-based trajectory planning method based on delayed experience-first replay mechanism is adopted. By constructing a Markov decision process, combining the priority experience replay mechanism and the maximum entropy strategy deep reinforcement learning network model, and introducing a phased training and delayed update mechanism, the training speed and stability are improved, and millisecond-level trajectory planning is achieved.

Benefits of technology

It improves the success rate and real-time performance of UAV trajectory planning in dense obstacle environments, achieving millisecond-level trajectory planning and meeting the real-time computing needs in dynamic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116974299B_ABST
    Figure CN116974299B_ABST
Patent Text Reader

Abstract

The application discloses a kind of reinforcement learning path planning method based on delay experience priority playback mechanism, belong to path planning technical field.The application implementation method is: considering unmanned aerial vehicle dynamics, flight performance, terrain and threat constraint constructs unmanned aerial vehicle obstacle avoidance path planning problem model, and to design the state-action-reward three elements of reinforcement learning path planning problem with this;Build local path planning training and application framework based on maximum entropy strategy, reduce the calculation time of path planning under local information driving through the hierarchical mechanism of '' offline training-online planning '';Combined with non-sparse design guide reward function, gradually approach the target using local information to guide unmanned aerial vehicle.Introduce strategy delay update mechanism and priority experience playback mechanism, in the training process of network parameter, training in stages to speed up the convergence speed of reinforcement learning training.The application can improve the training speed and stability in the training process of reinforcement learning, realize millisecond level online path planning.
Need to check novelty before this filing date? Find Prior Art