The invention provides a multi-AGV path planning
algorithm based on dynamic adaptive exploration and course learning, and relates to the technical field of
automation and intelligent logistics. According to the method, although a traditional deep
reinforcement learning method is preliminarily applied to multi-AGV
system path planning, the limitations of low efficiency, poor dynamic adaptability, insufficient cooperative competition relation
processing and the like still exist, and the specific expressions are low exploration efficiency,
insufficient sample utilization,
slow convergence speed and even non-convergence. For this purpose, a multi-agent depth deterministic strategy gradient
algorithm (AECL-MADDPG) based on adaptive exploration and course learning is designed, and centralized training is adopted. A distributed execution framework is adopted, the
obstacle avoidance capability and the implicit cooperation efficiency of the AGV in a high-density environment are enhanced by sensing the environment congestion degree and decision uncertainty in real time to dynamically adjust the exploration strength, meanwhile, a course learning-based priority experience playback mechanism is constructed, a training normal form from easy to difficult is combined with key experience priority sampling, and the accuracy and the robustness of the AGV are improved. Model convergence is remarkably accelerated; and the robustness of a final strategy is improved. The
algorithm established and designed under the actual operation condition in the
automation and intelligent logistics field shows significant advantages in key indexes such as convergence speed, task success rate,
average path length and the like, and an efficient and reliable solution is provided for the multi-agent path planning problem in a complex dynamic environment.