基于DAMP的人形机器人复杂地形运动控制方法及系统

By constructing the DAMP reinforcement learning control framework, the problems of unstable movement and unnatural actions of humanoid robots in complex terrain are solved, and more efficient autonomous adaptive control is achieved, which is suitable for a variety of application scenarios.

CN121848385BActive Publication Date: 2026-07-17SONGYAN POWER (BEIJING) TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SONGYAN POWER (BEIJING) TECHNOLOGY CO LTD
Filing Date
2026-01-13
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing humanoid robots suffer from poor motion stability, insufficient naturalness of movement, inaccurate state estimation, and inconsistent control frameworks in complex terrain motion control, resulting in poor performance in variable environments.

Method used

We construct a DAMP reinforcement learning control framework that integrates a denoised world model, adversarial motion priors, dynamic reward interpolation mechanism, and Lipschitz continuity penalty. By extracting robust states through the denoised world model, generating natural human-like actions, optimizing reward allocation and policy smoothness, we can improve the robot's autonomous adaptation ability in complex terrain.

Benefits of technology

It significantly improves the motion stability and adaptability of humanoid robots in complex terrain, generates actions that conform to human behavior patterns, enhances learning efficiency and control performance, and strengthens its application potential in service, medical and entertainment fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121848385B_ABST
    Figure CN121848385B_ABST
Patent Text Reader

Abstract

本发明提供了一种基于DAMP的人形机器人复杂地形运动控制方法及系统,属于机器人运动控制领域,包括:构建融合去噪世界模型模块、对抗运动先验模块、动态奖励插值机制及Lipschitz连续性惩罚模块的DAMP强化学习控制框架,首先利用去噪世界模型提取鲁棒的潜在状态表示并重构真实状态;其次基于人体运动捕捉数据生成多样化的类人动作;然后通过动态奖励插值机制,自适应调整模仿奖励与任务奖励的权重,优化学习策略,施加Lipschitz连续性惩罚控制策略的平滑性,减少控制输出的抖动;最后基于统一训练目标函数进行整体训练,以提升人形机器人在复杂地形中的自主运动能力,进而提高机器人在复杂地形中的运动稳定性与自然性。
Need to check novelty before this filing date? Find Prior Art