This invention discloses a training method and
system for quadruped robots based on deep
reinforcement learning. The AMP training framework significantly enhances the
robot's adaptability to complex environments. By distinguishing the
robot's current motion sample distribution from that of a
reference sample, the training aims to maximize the probability of misjudgment and the reward for
reinforcement learning, enabling the
robot to learn motion patterns that meet actual needs. When the environment changes, the robot can adjust its strategy promptly based on new information without remodeling, solving the problem of traditional methods being sensitive to environmental changes. Furthermore, this framework introduces
imitation learning into
reinforcement learning, providing effective guidance for
robot learning, reducing the
blindness of random exploration, accelerating the learning speed of effective strategies, and improving training efficiency. Simultaneously, the integrated hierarchical LSTM Actor network can extract and aggregate motion features at multiple time scales, allowing the robot to generate actions that consider both short-term
gait and long-term trends, resulting in more natural and fluid movements. It also fully utilizes existing sample information, improving sample utilization and reducing training costs.