一种基于可扩展世界模型的机器人多任务持续学习方法

By constructing a modular world model that separates cognition and decision-making, and combining it with model predictive control, the problems of catastrophic forgetting and insufficient generalization in robot learning in multi-task and dynamic environments are solved, achieving efficient and stable learning results.

CN122401455APending Publication Date: 2026-07-17UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610884007.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing robots suffer from catastrophic forgetting, insufficient generalization ability, and low sample efficiency in multi-task and dynamic environments. Traditional world model methods are difficult to adapt to multi-environment and multi-task scenarios, resulting in degraded model performance and low learning efficiency.

Method used

It adopts a modular architecture based on a scalable world model, including a cognitive layer and a decision layer. The cognitive layer performs unified modeling of the environment and tasks, while the decision layer performs action sequence prediction and reward evaluation through multiple reusable skill modules. Combined with model predictive control, it achieves efficient learning.

Benefits of technology

It effectively alleviates the problem of catastrophic forgetting, improves the adaptability and generalization ability to new tasks, increases sample utilization efficiency, reduces training costs, and maintains the stability and compactness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122401455A_ABST
    Figure CN122401455A_ABST
Patent Text Reader

Abstract

本发明提出一种基于可扩展世界模型的机器人多任务持续学习方法,属于具身智能持续学习领域。方法包括:构建世界模型认知层,生成统一认知状态并匹配技能模块,筛选技能模块或初始化新技能模块;构建世界模型决策层,构建动作分布,并据此生成多组候选动作序列,对候选动作序列进行演化预测;采用模型预测控制,在每个控制时刻选取最优动作序列中的当前动作执行,并在下一时刻结合新的环境反馈重新进行规划,从而实现闭环控制;在执行动作并获得环境反馈后,对被激活技能模块的参数与结构进行更新,实现持续自进化。本发明能够提升机器人对新任务的适配能力与新环境的泛化能力,能够提高机器人持续学习系统的稳定性、可扩展性与工程部署效率。
Need to check novelty before this filing date? Find Prior Art