A multi-layered embodied capability asset humanoid robot continuous evolution method and system

CN122770002APending Publication Date: 2026-09-18CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611203504.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-10
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0005]第一,主要关注单次任务执行,缺少跨任务持续进化机制

Benefits of technology

1.本发明将人形机器人在家庭、商超、工厂等场景中执行任务的运行经验转化为结构化、标准化的多层具身能力资产,避免了每台机器人在新场景中重复学习、重复试错。随着部署本发明的机器人系统的实际运行时间的增加,机器人在处理同类任务时首次成功率、平均完成时间和用户满意度均有显著提升。在系统鲁棒性和安全性方面,本发明通过失败归因机制将每次任务失败转化为可复用的改进资产,使得机器人的故障处理能力和异常恢复能力持续增强。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122770002A_ABST
    Figure CN122770002A_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for the continuous evolution of a humanoid robot with multi-layered embodied capability assets. The method includes: S1. Receiving a user task; the system receives the task instruction input by the user through a task receiving module and parses the task into the robot's execution objective; S2. Perceiving environmental data and the robot's own state; S3. Retrieving relevant capability assets based on a capability graph; S4. Generating a task execution plan; S5. Executing the task and collecting multi-modal embodied trajectories; S6. Determining the task execution status; S7. Evaluating the task results; S8. Attributing failures or extracting successful experiences; S9. Generating candidate multi-layered embodied capability assets; S10. Updating the capability asset library and capability graph; S11. Dynamically calling capability assets in subsequent tasks. This invention not only avoids repeated learning and trial and error for each robot in new scenarios but also improves the reusability of robot capabilities and experience, reducing the cost of repeated trial and error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot intelligence technology, and in particular to a method and system for the continuous evolution of a humanoid robot with multi-layered embodied capabilities. Background Technology

[0002] With the rapid development of large language models, multimodal perception models, basic robot models, and intelligent agent technologies, humanoid robots are gradually acquiring core capabilities such as natural language understanding, environmental perception, task planning, tool invocation, motion control, and human-computer interaction. Existing embodied intelligent robot systems typically understand user tasks through large models, perceive the environment by combining information from multiple sensors such as vision, hearing, and touch, and then generate action sequences or control commands through task planning and motion control modules to drive the robot to complete specified complex tasks. Current technologies have made some progress in several directions.

[0003] In the area of ​​motion primitives or skill libraries, existing solutions combine pre-set motion primitives, micro-manipulation libraries, or skill templates to create complex actions, enabling robots to perform complex tasks such as cooking and caregiving. In the area of ​​robot skill learning, existing solutions use imitation learning, reinforcement learning, or demonstration learning to teach robots specific motion skills or control strategies. In the area of ​​embodied intelligent robot general systems, existing solutions improve the robot's understanding of the environment and user tasks by fusing large language models, memory models, multimodal perception models, and motion control models. In the area of ​​simulation teaching and program translation, existing solutions generate robot trajectories or program files in a simulation environment and translate them into executable programs for the target robot.

[0004] However, the aforementioned existing technical solutions still have the following shortcomings.

[0005] First, existing solutions primarily focus on single-task execution, lacking mechanisms for continuous evolution across tasks. While current solutions focus on how robots perceive, plan, and execute based on the current task, they lack a systematic approach to how to transform execution experience into reusable capabilities after a task is completed and continuously improve the robot's abilities in subsequent tasks.

[0006] Second, the granularity of capability representation is limited, lacking a multi-layered embodied capability asset system. Existing solutions typically represent capabilities as one of memory, motor skills, control strategies, or action primitives, which is insufficient to cover the diverse capabilities involved in real-world tasks for humanoid robots, such as environmental semantics, task strategies, motor skills, force control and tactile feedback, safety constraints, anomaly recovery, user interaction, and model adaptation.

[0007] Third, there is a lack of an abstraction mechanism from multimodal embodied trajectories to capability assets. When humanoid robots perform tasks, they generate multimodal trajectories such as vision, depth, voice, joints, torque, touch, foot pressure, collision, emergency stop, and human intervention. Existing solutions often only use this data for perception or control, without systematically extracting reusable capability assets from it.

[0008] Fourth, there is a lack of failure attribution mechanisms for continuous evolution. Robot task failures can be caused by a variety of reasons, such as errors in task understanding, errors in environmental recognition, incorrect selection of grasping points, improper force control parameters, unstable gait, insufficient safety constraints, and ambiguity in user intent. Existing solutions lack mechanisms to map failure causes to different types of capability assets.

[0009] Fifth, there is a lack of dynamic recall and adaptive combination mechanisms for capability assets. When a robot performs a new task, it should not only recall fixed skill libraries or fixed action primitives, but should dynamically combine multi-layered embodied capability assets based on task type, scenario state, user preferences, risk level, ontology state, and historical performance.

[0010] Sixth, there is a lack of capability asset mapping and traceability mechanisms. Existing solutions typically struggle to track which historical tasks a particular robot capability originated from, which failure modes it has addressed, which scenarios it is applicable to, whether it has brought performance improvements, and whether there are any safety risks.

[0011] Therefore, it is necessary to propose a method for the continuous evolution of humanoid robots based on multi-layered embodied capability assets, so that humanoid robots can continuously extract, accumulate, call and update various types of embodied capability assets from real task execution trajectories, thereby achieving continuous capability improvement across tasks, scenarios and time. Summary of the Invention

[0012] In view of the above problems, the present invention provides a method and system for the continuous evolution of humanoid robots with multi-layered embodied capability assets, which not only avoids repeated learning and trial and error for each robot in new scenarios, but also improves the reusability of robot capabilities and experience and reduces the cost of repeated trial and error.

[0013] To achieve the above and other related objectives, the technical solution provided by this invention is as follows: A method for the continuous evolution of a humanoid robot with multi-layered embodied capability assets, the method comprising: S1. Receive user tasks. The system receives the task instructions input by the user through the task receiving module and parses the task into the robot's execution target. S2. Perceive environmental data and the entity's state; S3. Retrieve relevant capability assets based on the capability map; S4. Generate a task execution plan; S5. Perform the task and collect multimodal embodied trajectories; S6. Determine the task execution status; S7. Evaluate the task results; S8. Perform failure attribution or success experience extraction; S9. Generate multi-tiered embodied ability asset candidates; S10. Update the capability asset library and capability map; S11. Dynamically call upon capability assets in subsequent tasks.

[0014] Furthermore, in step S2, the environmental data includes images, depth maps, point clouds, speech, ambient sounds, target objects, obstacles, human body positions, passable areas, and operable areas, while the body state data includes robot pose, joint angles, joint velocities, joint torques, end effector status, tactile feedback, plantar pressure, inertial measurement unit data, battery status, and actuator health status.

[0015] Furthermore, in step S3, the retrieval of relevant capability assets based on the capability graph involves the system retrieving relevant capability assets from the capability graph and capability asset library according to the task type, scene conditions, target object, user preferences, and robot state. The retrieval results include environmental semantic assets, task strategy assets, motion skill assets, force control haptic assets, safety constraint assets, anomaly recovery assets, and user interaction assets.

[0016] Further, in step S4, the task generation plan is the system combining the current task and the retrieved capability assets to generate a task execution plan. The task execution plan includes a sub-task sequence, action skill invocation order, target object and operation object, key poses, path planning constraints, force control constraints, human-machine confirmation nodes, safety constraints, and anomaly recovery strategies. In step S5, the task execution and multimodal embodied trajectory acquisition is the robot executing the task according to the task execution plan and continuously acquiring multimodal embodied trajectory during the execution process. The multimodal embodied trajectory includes task input, perception results, planning results, action skill invocation records, motion control data, force control tactile data, abnormal events, user feedback, and task results.

[0017] Furthermore, in step S6, the determination of the task execution status is the system's determination of whether the task has been successfully completed. The task execution status includes complete success, partial success, failure, cancellation by the user, manual takeover, interruption by the security mechanism, and active termination by the robot.

[0018] Furthermore, in step S7, the task result evaluation is performed by the system based on the task objectives, user feedback, automatic evaluation rules, security events, and manual review information. The evaluation indicators include task success rate, completion time, operational accuracy, object damage, human-machine distance, collision events, emergency stop events, number of manual interventions, user satisfaction, energy consumption, execution stability, and anomaly recovery effect.

[0019] Further, in step S8, the failure attribution or success experience extraction involves the system performing failure attribution when the task fails or partially succeeds, and extracting success experience when the task succeeds and the effect is better than the historical benchmark. The failure attribution and success experience extraction results are converted into evolutionary signals. The failure attribution types include task understanding failure, scene perception failure, target localization failure, object attribute judgment failure, task decomposition failure, execution order error, action skill selection error, key pose selection error, improper grasping force control, unreasonable movement path planning, insufficient human-computer interaction confirmation, insufficient safety constraint triggering, lack of abnormal recovery strategy, abnormal robot body state, and scene changes causing the original plan to fail.

[0020] Further, in step S9, the generation of multi-layered embodied capability asset candidates involves the system generating multi-layered embodied capability asset candidates based on evolutionary signals. These candidates include new environmental semantic assets, new task strategy assets, new motion skill assets, new force control haptic assets, new safety constraint assets, new anomaly recovery assets, new user interaction assets, new tool usage assets, new model adaptation assets, and new capability graph relationships. The system maps failure attribution results to different types of capability asset candidates: if the task failure is due to incorrect target object recognition, environmental semantic asset candidates are generated; if the failure is due to an unreasonable task execution order, task strategy asset candidates are generated; if the failure is due to unstable grasping points, motion skill assets or force control haptic assets candidates are generated; if the failure is due to the robot getting too close to the human body, safety constraint assets candidates are generated; if the failure is due to ambiguous user instructions, user interaction assets candidates are generated; and if manual intervention successfully completes the task after failure, anomaly recovery asset candidates are generated from the manual intervention trajectory.

[0021] Furthermore, the updating of the capability asset library and capability graph involves the system writing candidate capability assets that have undergone basic verification into the candidate asset area, and updating the relationships between tasks, scenarios, failure modes, assets, and effects in the capability graph. High-risk action skills, force control strategies, and safety constraint assets can enter a pending verification state for further processing by the safety verification module. The dynamic invocation of capability assets in subsequent tasks means that when the robot performs a new similar task, the system dynamically invokes existing capability assets based on the capability graph and asset scores, enabling the robot to continuously improve its task execution capabilities over long-term operation.

[0022] To achieve the above and other related objectives, the present invention also provides a humanoid robot continuous evolution system with multi-layered embodied capability assets, used to implement the aforementioned method for continuous evolution of humanoid robots with multi-layered embodied capability assets. The system includes: a task receiving module, a multimodal perception module, a body state acquisition module, a capability asset library, a capability graph module, a task planning module, an action execution module, a trajectory acquisition module, a task evaluation module, a failure attribution module, an asset generation module, and a dynamic invocation module. Users submit task instructions to the system through the task receiving module. Task instructions can include voice, text, gestures, images, or multimodal forms. The multimodal perception module collects and analyzes information about the robot's surrounding environment, including images, depth, point clouds, voice, tactile information, target objects, human body position, and obstacle information. The body state acquisition module collects the robot's own state, including joint angles, joint velocities, joint torques, end effector state, plantar pressure, inertial measurement unit data, battery status, and actuator health status; The capability asset library is used to store multi-layered embodied capability assets, including environmental semantic assets, task strategy assets, motion skill assets, force control and haptic assets, safety constraint assets, anomaly recovery assets, user interaction assets, tool usage assets, and model adaptation assets. The capability graph module is used to record the relationships between task types, scenario conditions, failure modes, success strategies, capability assets, robot status, and operational results. The task planning module generates a task execution plan based on the user task, environment information, ontology status, and relevant capability assets retrieved from the capability graph. The motion execution module calls the motion controller, motion skills, and tool interfaces to drive the robot to complete the task; The trajectory acquisition module records the robot's multimodal embodied trajectory during task execution. The task evaluation module evaluates the effectiveness of task execution based on task objectives, execution results, user feedback, security incidents, and manual takeover information. The failure attribution module analyzes the reasons for task failure or performance degradation and outputs the failure attribution results. The asset generation module generates multi-layered embodied ability asset candidates based on evolutionary signals; The dynamic invocation module dynamically invokes existing capability assets based on the capability graph and asset score in subsequent tasks; The modules interact with each other via an internal bus or application programming interface. The task receiving module receives multimodal user commands at its input end and outputs a structured task target vector at its output end; the multimodal perception module outputs an environmental feature matrix containing spatial coordinates and semantic labels; and the ontology state acquisition module outputs an ontology state vector containing joint state and torque information.

[0023] The present invention has the following positive effects: 1. This invention transforms the operational experience of humanoid robots performing tasks in scenarios such as homes, supermarkets, and factories into structured and standardized multi-layered embodied capability assets, avoiding repeated learning and trial-and-error for each robot in new scenarios. As the actual operating time of robot systems deployed with this invention increases, the first-time success rate, average completion time, and user satisfaction all significantly improve when handling similar tasks. Regarding system robustness and security, this invention uses a failure attribution mechanism to transform each task failure into a reusable improvement asset, continuously enhancing the robot's fault handling and anomaly recovery capabilities.

[0024] 2. This invention improves the long-term task adaptability and self-improvement capabilities of humanoid robots in real-world scenarios; reduces the impact of task failures on human living and working environments; increases the reusability of robot capabilities and experience, reducing the cost of repeated trial and error; improves the interpretability and traceability of the robot's continuous learning process; and provides a systematic capability asset foundation for multi-robot sharing and cross-ontology migration. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is an abstract mapping diagram from the multimodal embodied trajectory to capability assets in this invention; Figure 3 This is a schematic diagram of the capability map structure of the present invention; Figure 4 This is a flowchart illustrating the dynamic retrieval and adaptive combination of capability assets in this invention. Figure 5 This is a schematic diagram of the system framework of the present invention; Figure 6 This is a schematic diagram of the fusion algorithm of the hierarchical task network for capability asset graph enhancement and the two-layer Monte Carlo Tree Search (Graph-RAG-HTN-MCTS) of the present invention; Figure 7 This is a schematic diagram (I) of the multimodal embodied trajectory temporal residual test and hierarchical causal Bayesian network (CBN) algorithm of the present invention; Figure 8 This is a schematic diagram (II) of the multimodal embodied trajectory temporal residual test and hierarchical causal Bayesian network (CBN) algorithm of the present invention. Detailed Implementation

[0026] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0027] Example: Figure 1 As shown, a method for the continuous evolution of a humanoid robot with multi-layered embodied capability assets is provided, the method comprising: S1. Receive user tasks. The system receives the task instructions input by the user through the task receiving module and parses the task into the robot's execution target. S2. Perceive environmental data and the entity's state; S3. Retrieve relevant capability assets based on the capability map; S4. Generate a task execution plan; S5. Perform the task and collect multimodal embodied trajectories; S6. Determine the task execution status; S7. Evaluate the task results; S8. Perform failure attribution or success experience extraction; S9. Generate multi-tiered embodied ability asset candidates; S10. Update the capability asset library and capability map; S11. Dynamically call upon capability assets in subsequent tasks.

[0028] In this embodiment, in step S2, the environmental data includes images, depth maps, point clouds, speech, ambient sounds, target objects, obstacles, human body positions, passable areas, and operable areas. The body state data includes robot pose, joint angles, joint velocities, joint torques, end effector status, tactile feedback, plantar pressure, inertial measurement unit data, battery status, and actuator health status.

[0029] In this embodiment, in step S3, the retrieval of relevant capability assets based on the capability graph is a process in which the system retrieves relevant capability assets from the capability graph and capability asset library according to the task type, scene conditions, target object, user preferences and robot state. The retrieval results include environmental semantic assets, task strategy assets, motion skill assets, force control haptic assets, safety constraint assets, anomaly recovery assets and user interaction assets.

[0030] In this embodiment, as Figure 2As shown, this invention defines a multi-layered embodied capability asset system, which unifies and abstracts the reusable capabilities of humanoid robots into multiple types of embodied capability assets, rather than representing capabilities as a single form such as memory, motor skills, control strategies, or action primitives. Multi-layered embodied capability assets include, but are not limited to, the following types.

[0031] The first category is environmental semantic assets, which describe the spatial layout, object categories, object attributes, operable areas, obstacle areas, hazardous areas, human activity areas, and task-related semantic relationships in the robot's working environment. When the robot recognizes new environmental features or discovers previously unrecorded object categories in the environment, the system automatically generates or updates the corresponding environmental semantic assets.

[0032] The second category is task strategy assets, which describe the execution strategy, task breakdown method, task priority, operation sequence, anomaly judgment conditions, and interaction confirmation conditions for a specific task type. When the robot successfully completes a complex task or discovers a better execution path, the system extracts the strategy logic to generate task strategy assets.

[0033] The third category is motion skill assets, which describe the applicable conditions, key poses, motion constraints, and execution parameters of motion skills such as grasping, placing, delivering, carrying, opening doors, pressing buttons, pushing and pulling drawers, wiping, tidying, and inspecting.

[0034] The fourth category is force-controlled haptic assets, which are used to describe the clamping force, contact force, slip detection threshold, collision detection threshold, and force control adjustment strategy under different objects, different contact states, and different task objectives.

[0035] The fifth category is safety constraint assets, which are used to describe human-machine distance constraints, joint speed limits, joint torque limits, rules prohibiting entry into dangerous areas, collision protection rules, emergency stop triggering rules, and high-risk task constraints.

[0036] The sixth category is anomaly recovery assets, which describe recovery strategies for events such as capture failure, object drop, target loss, path obstruction, instruction ambiguity, sensor anomalies, balance anomalies, and manual takeover.

[0037] The seventh category is user interaction assets, which are used to describe different users' expression habits, confirmation preferences, task preferences, security preferences, and feedback habits.

[0038] Category 8 is tool usage assets, which describe how robots use tools or environmental facilities such as trays, grippers, brooms, rags, screwdrivers, buttons, and door handles.

[0039] The ninth category is model adaptation assets, which are used to describe retrieval indexes, prompt word templates, strategy parameters, visual model parameters, or control model parameters for specific tasks, scenarios, or users.

[0040] The tenth category is capability graph assets, which are used to describe the relationships between task types, scenario conditions, failure modes, success strategies, capability assets, and operational effects.

[0041] In this embodiment, as Figure 3 As shown, this invention proposes a failure attribution and asset mapping mechanism for continuous evolution. This mechanism does not merely determine task success, but rather performs deep failure attribution based on task trajectory, task results, safety events, and human intervention information. Failure attribution types include task comprehension failure, scene perception failure, target localization failure, object attribute judgment failure, task decomposition failure, incorrect execution order, incorrect action skill selection, incorrect key pose selection, improper gripping force control, unreasonable movement path planning, insufficient human-machine interaction confirmation, insufficient safety constraint triggering, lack of anomaly recovery strategies, robot body state abnormalities, and scene changes leading to the failure of the original plan. The system generates different types of capability asset candidates based on the failure attribution results, enabling each failure to be transformed into reusable improved assets for subsequent tasks. This mechanism differs from simple imitation learning or reinforcement learning; its focus is not on directly learning a control model, but on transforming the causes of task failure into manageable, callable, and traceable capability assets.

[0042] The system identifies the following evolutionary signals based on multimodal embodied trajectories: successful operation segments, failed operation segments, differences before and after manual intervention, differences before and after user correction, high-risk contact events, grasping failure events, target loss events, path blocking events, force control anomalies, and user negative feedback events. The system then maps these evolutionary signals to different types of embodied capability asset candidates. For example, if the task failure is due to incorrect target object recognition, environmental semantic asset candidates are generated; if the task failure is due to an unreasonable task execution order, task strategy asset candidates are generated; if the task failure is due to unstable grasping points, motion skill assets or force control haptic assets are generated; if the task failure is due to the robot getting too close to the human, safety constraint assets are generated; if the task failure is due to ambiguous user instructions, user interaction assets or confirmation strategy assets are generated; and if manual intervention successfully completes the task after failure, anomaly recovery asset candidates are generated from the manual intervention trajectory.

[0043] This invention employs a capability graph-driven continuous evolution mechanism to construct a humanoid robot capability graph, used to record the relationships between task types, scenario conditions, failure modes, success strategies, capability assets, and operational effects. The capability graph includes the following nodes: task type node, scenario type node, object type node, failure mode node, success strategy node, capability asset node, robot body node, and operational effect node. The edge relationships in the capability graph include: a certain task type invokes a certain capability asset under certain scenario conditions; a certain capability asset arises from a certain failure mode or success experience; and a certain capability asset achieves a certain operational effect on a certain robot body.

[0044] Capability graphs play a central role in continuous evolution. When a new task arrives, the system retrieves the most relevant historical capability assets from the capability graph based on the task type and scenario conditions. The system can answer questions such as: Which capability assets are most effective for this task type? What failure modes have occurred under similar scenario conditions? On which robot body has a particular capability asset been validated? Has the introduction of a particular capability asset led to performance improvements? Capability graphs support intelligent reasoning queries based on graph traversal, recommending the most relevant capability assets and strategy paths to the robot, while also supporting visual analysis and optimization decisions regarding evolutionary paths.

[0045] In this embodiment, in step S4, generating a task plan involves the system combining the current task and retrieved capability assets to generate a task execution plan. The task execution plan includes a sub-task sequence, action skill invocation order, target object and operation object, key poses, path planning constraints, force control constraints, human-machine confirmation nodes, safety constraints, and anomaly recovery strategies. In step S5, executing the task and collecting multimodal embodied trajectories involves the robot executing the task according to the task execution plan and continuously collecting multimodal embodied trajectories during the execution process. The multimodal embodied trajectories include task input, perception results, planning results, action skill invocation records, motion control data, force control tactile data, abnormal events, user feedback, and task results.

[0046] In this embodiment, in step S6, determining the task execution status means that the system determines whether the task has been successfully completed. The task execution status includes complete success, partial success, failure, cancellation by the user, manual takeover, interruption by the security mechanism, and active termination by the robot.

[0047] In this embodiment, in step S7, the task result evaluation is performed by the system based on the task objective, user feedback, automatic evaluation rules, security events, and manual review information. The evaluation indicators include task success rate, completion time, operation accuracy, object damage, human-machine distance, collision events, emergency stop events, number of manual interventions, user satisfaction, energy consumption, execution stability, and anomaly recovery effect.

[0048] In this embodiment, as Figure 4 As shown, when the robot receives a new task, the system first performs a multi-dimensional search based on the capability map. The search dimensions include: task type matching degree (the similarity between the current task and historical task types); scene condition matching degree (the similarity between the current environment and historical scene conditions); target object matching degree (the attribute matching degree between the target object and historically manipulated objects); user preference matching degree (the consistency of behavioral preferences between the current user and historical users); and robot body matching degree (the structural similarity between the current robot body and historically executing robots). Based on the comprehensive matching degree, the system retrieves the most relevant capability assets from the capability asset library.

[0049] After retrieving relevant capability assets, the system adaptively combines multiple types of capability assets according to the specific requirements of the target task. For example, when executing a delivery task, the system may simultaneously invoke environmental semantic assets to understand the current environment layout, invoke task strategy assets to determine the delivery path and operation sequence, invoke motion skill assets to perform grasping and placing operations, invoke safety constraint assets to ensure a safe distance between humans and machines, and invoke anomaly recovery assets to handle possible situations such as dropped items or blocked paths. The system continuously updates the ratings and applicable conditions of capability assets based on feedback from task execution performance, thereby achieving continuous optimization of capability assets.

[0050] In this embodiment, in step S8, the failure attribution or success experience extraction involves the system performing failure attribution when the task fails or partially succeeds, and extracting success experience when the task succeeds and the effect is better than the historical benchmark. The failure attribution and success experience extraction results are converted into evolutionary signals. The failure attribution types include task understanding failure, scene perception failure, target localization failure, object attribute judgment failure, task decomposition failure, execution order error, action skill selection error, key pose selection error, improper grasping force control, unreasonable movement path planning, insufficient human-computer interaction confirmation, insufficient safety constraint triggering, lack of abnormal recovery strategy, abnormal robot body state, and scene changes causing the original plan to fail.

[0051] In this embodiment, in step S9, the generation of multi-layered embodied capability asset candidates is a process where the system generates multi-layered embodied capability asset candidates based on evolutionary signals. These candidates include new environmental semantic assets, new task strategy assets, new motion skill assets, new force control haptic assets, new safety constraint assets, new anomaly recovery assets, new user interaction assets, new tool usage assets, new model adaptation assets, and new capability graph relationships. The system maps failure attribution results to different types of capability asset candidates: if the task failure is due to incorrect target object recognition, environmental semantic asset candidates are generated; if the failure is due to an unreasonable task execution order, task strategy asset candidates are generated; if the failure is due to unstable grasping points, motion skill assets or force control haptic assets candidates are generated; if the failure is due to the robot getting too close to the human body, safety constraint assets candidates are generated; if the failure is due to ambiguous user instructions, user interaction assets candidates are generated; and if manual intervention successfully completes the task after failure, anomaly recovery asset candidates are generated from the manual intervention trajectory.

[0052] In this embodiment, updating the capability asset library and capability graph involves the system writing candidate capability assets that have undergone basic verification into the candidate asset area and updating the relationships between tasks, scenarios, failure modes, assets, and effects in the capability graph. High-risk action skills, force control strategies, and safety constraint assets can enter a pending verification state for further processing by the safety verification module. Dynamically calling capability assets in subsequent tasks means that when the robot performs a new similar task, the system dynamically calls existing capability assets based on the capability graph and asset scores, enabling the robot to continuously improve its task execution capabilities over long-term operation.

[0053] In this embodiment, as Figure 5 As shown, the present invention provides a humanoid robot continuous evolution system with multi-layered embodied capability assets, used to realize the method for continuous evolution of humanoid robots with multi-layered embodied capability assets. The system includes: a task receiving module, a multimodal perception module, an ontology state acquisition module, a capability asset library, a capability map module, a task planning module, an action execution module, a trajectory acquisition module, a task evaluation module, a failure attribution module, an asset generation module, and a dynamic invocation module. Users submit task instructions to the system through the task receiving module. Task instructions can include voice, text, gestures, images, or multimodal forms. The multimodal perception module collects and analyzes information about the robot's surrounding environment, including images, depth, point clouds, voice, tactile information, target objects, human body position, and obstacle information. The body state acquisition module collects the robot's own state, including joint angles, joint velocities, joint torques, end effector state, plantar pressure, inertial measurement unit data, battery status, and actuator health status; The capability asset library is used to store multi-layered embodied capability assets, including environmental semantic assets, task strategy assets, motion skill assets, force control and haptic assets, safety constraint assets, anomaly recovery assets, user interaction assets, tool usage assets, and model adaptation assets. The capability graph module is used to record the relationships between task types, scenario conditions, failure modes, success strategies, capability assets, robot status, and operational results. The task planning module generates a task execution plan based on the user task, environment information, ontology status, and relevant capability assets retrieved from the capability graph. The motion execution module calls the motion controller, motion skills, and tool interfaces to drive the robot to complete the task; The trajectory acquisition module records the robot's multimodal embodied trajectory during task execution. The task evaluation module evaluates the effectiveness of task execution based on task objectives, execution results, user feedback, security incidents, and manual takeover information. The failure attribution module analyzes the reasons for task failure or performance degradation and outputs the failure attribution results. The asset generation module generates multi-layered embodied ability asset candidates based on evolutionary signals; The dynamic invocation module dynamically invokes existing capability assets based on the capability graph and asset score in subsequent tasks; The modules interact with each other via an internal bus or application programming interface. The task receiving module receives multimodal user instructions at its input end and outputs a structured task target vector at its output end. The multimodal perception module outputs an environmental feature matrix containing spatial coordinates and semantic labels. The ontology state acquisition module outputs an ontology state vector containing joint state and torque information.

[0054] In this embodiment, the data format table of the task planning module is shown in Table 1. Table 1

[0055] In this embodiment, the implementation environment and configuration of the present invention rely on a humanoid robot platform and a large language model intelligent agent system. The humanoid robot body is equipped with a multimodal sensor system, including a vision sensor, a depth sensor, a voice acquisition array, a tactile sensor, a joint torque sensor, an inertial measurement unit, and a plantar pressure sensor. The robot motion control system adopts a hierarchical control architecture, with the bottom layer for joint servo control, the middle layer for whole-body motion control, and the top layer for task planning and skill scheduling. The large language model, as the core component for task understanding and strategy generation, runs on an edge or cloud server equipped with a computing accelerator. The capability asset library adopts a hybrid storage architecture of graph database and relational database. The graph database stores the nodes and edge relationships of the capability graph, while the relational database stores the detailed content and metadata of the capability assets.

[0056] Continuous evolution in home service scenarios: In one specific embodiment, this invention is applied to a home service scenario. A humanoid robot provides services such as item delivery, desktop tidying, floor cleaning, and simple repair assistance in a home environment. In the initial deployment phase, the robot only possesses basic motion control and preset action skills, such as walking, obstacle avoidance, simple grasping, and item placement. As the robot operates in the home environment over a long period, the system begins to collect multimodal embodied trajectories during each task execution, gradually accumulating capability assets.

[0057] Taking the everyday task of fetching a bottle of mineral water from the kitchen and delivering it to the coffee table in the living room as an example, the complete execution process of this invention is explained in detail. The robot first receives user instructions through the task receiving module, the multimodal perception module identifies the positions of the refrigerator and water bottle in the kitchen environment, the body state acquisition module acquires the current joint state and battery information, and the task planning module retrieves the environmental semantic assets and grasping strategy assets from previous attempts to grasp round objects from the refrigerator based on the accumulated capability map. During the robot's task execution, the trajectory acquisition module records the complete trajectory from path planning to grasping action to delivery and placement. If the robot discovers that the bottled water is placed in a previously unrecorded location deep within the refrigerator shelf, the system records this environmental change as an evolutionary signal and generates new environmental semantic asset candidates, updating the internal spatial layout information of the refrigerator. If slippage occurs due to water condensation during grasping, the failure attribution module attributes the slippage event to insufficient force-controlled tactile parameters and generates force-controlled tactile asset candidates, updating the clamping force parameters of objects on slippery surfaces.

[0058] After a period of operation, the robot has accumulated rich experience in home services in its capability asset library. For example, it has learned the delivery locations preferred by different users: the elderly prefer to place it on the table, while children prefer to have it handed directly to them, which can be summarized as interaction assets for different users. It has mastered the grasping strategies for different types of containers: gently handling a wine glass, holding a water bottle at a moderate distance, and carrying a casserole with both hands, which can be summarized as force control tactile assets and motor skill assets. It has accumulated safety constraints for multi-room path planning: slowing down when passing through stairwells, being aware of tripping risks in carpeted areas, and strengthening obstacle avoidance in pet activity areas, which can be summarized as safety constraint assets and environmental semantic assets. It has summarized recovery methods for abnormal situations: adaptively adjusting the grasping posture and retrying when an item is dropped, replanning a detour route when the path is temporarily obstructed, and confirming and adjusting before adjusting when the user changes the task goal midway, which can be summarized as abnormal recovery assets. As shown in Table 2, Table 2

[0059] This invention proposes a multi-layered embodied capability asset continuous evolution mechanism for humanoid robots. This mechanism takes the multimodal embodied trajectories generated during the robot's real-world task execution as input, and generates environmental semantic assets, task strategy assets, motion skill assets, force control and haptic assets, safety constraint assets, anomaly recovery assets, user interaction assets, tool usage assets, model adaptation assets, and capability graph assets through task result evaluation, failure attribution, and success experience extraction. In subsequent tasks, these capability assets are dynamically invoked and updated based on task type, scene state, robot body state, and historical operational results.

[0060] In this embodiment, the task planning module of the present invention employs a fusion algorithm of "Hierarchical Task Network Enhanced with Capability Asset Graph and Two-Layer Monte Carlo Tree Search (Graph-RAG-HTN-MCTS)". This algorithm combines high-level semantic task deconstruction with low-level physical feasibility verification, and uses multi-layered embodied capability assets (task strategy assets, force control haptic assets, safety constraint assets, anomaly recovery assets, etc.) retrieved from the capability graph to dynamically prune and inject constraints into the search space, generating a high-success-rate embodied task execution plan.

[0061] In this embodiment, as Figure 6 As shown, the specific steps of the algorithm include: Step P1. High-level semantic decomposition under capability asset constraints (HTN Decomposition); 1. Input: User task instructions, current multimodal sensing environment state S env and the body state S body .

[0062] 2. According to T u With S env The target task strategy asset A with the highest matching degree is retrieved through the capability graph. strat User Interaction Asset A user .

[0063] 3. Based on the Abstract High-Level Task Network (HTN) operator, complex tasks are recursively decomposed into sub-objective nodes $g1, g2, g3, ..., g$. n The directed acyclic graph (DAG) is constructed and the preference constraints (such as delivery location preference and human-computer interaction confirmation nodes) are injected.

[0064] Step P2. Dual-Layer Monte Carlo Tree Search for Multi-Layer Asset Injection Macro-Tree: Tree node definition: Node N represents the current environment and robot state, and edge E represents the invoked action skill assets.

[0065] Selection & Expansion: Using the asset confidence-weighted UpperConfidence Bound (UCB) formula to select actions. , in, Accumulated rewards for nodes, S graph (A skill ,i) represents the historical evaluation score (evaluation result) of the skill in the current scenario in the capability graph, and λ is the asset prior weight coefficient.

[0066] Micro-Layer Feasibility Check: Expand candidate action A at the macro level skill At the same time, force-controlled haptic asset A is invoked in parallel. force (e.g., clamping force threshold, slippage judgment parameters) and safety-constrained asset A safe (such as joint angle / velocity limits, no-collision zones).

[0067] If a candidate action violates or exceeds the feasible region of the dynamics, the branch is pruned directly and a maximum penalty value is imposed.

[0068] Step P3. Failure mode prediction and contingency recovery strategy injection. 1. Traverse the historical "failure mode nodes" associated with the current subtask node in the capability graph.

[0069] 2. If a subtask has historically experienced high-frequency failures in the current scenario (failure probability P(fail) > T_{risk}), the planner automatically recovers asset A from the abnormal situation. rec Extract the corresponding preventive actions or verification nodes (such as "perform depth verification before grabbing" or "confirm hand position before delivery") and insert them as explicit child nodes into the task execution plan.

[0070] Step P4. Generate and instantiate the embodied plan; The final output is a structured task execution plan P. exec = <G DAG ,{SK k},{P ctrl,k},{C safe,k} ,{CP rec,k}>$, which contains the subtask sequence G DAG , Execution action skill set {SK k}, Force control / trajectory control parameters {Pctrl,k}、Safety constraint boundary {C safe,k} and risk recovery checkpoints {CP rec,k}

[0071] 1. Traverse the historical "failure mode nodes" associated with the current subtask node in the capability graph.

[0072] The specific algorithm flow and key parameters of the failure attribution module, and an overview of the algorithm flow and logic of the failure attribution module. The failure attribution module of this invention employs the "Multimodal Embodied Trajectory Temporal Residual Detection and Hierarchical Causal Bayesian Network (CBN)" algorithm. This module does not rely on a single modality judgment, but instead compares the high-dimensional temporal trajectories (visual pose, joint torque, tactile torque, plantar pressure, human intervention signals, etc.) synchronously collected during the robot's task execution with a standard reference trajectory to pinpoint the time window in which anomalies occur. Subsequently, through causal tree inference, physical anomalies are precisely mapped to specific capability asset defects.

[0073] like Figure 7 or Figure 8 As shown, the specific steps of the algorithm include: Step F1. Trajectory time synchronization and sliding window residual positioning; 1. Unified clock reference: The vision sensor (30Hz), joint encoder and torque sensor (500Hz), tactile / force sensor (1000Hz), IMU (200Hz) and event signals (emergency stop / manual takeover / user denial voice, 100Hz) are aligned according to a unified timestamp.

[0074] 2. Abnormal Trigger Location: Based on the task termination time or the time of manual takeover / emergency stop trigger t. end Based on the baseline, the cut length is W. win Reverse analysis time window [t] end -W win ,t end ]$.

[0075] 3. Residual Vector Calculation: Calculate the discrete residual vector between the actual execution trajectory and the standard / expected trajectory. $.

[0076] Step F2. Multimodal residual contribution and dominant mode extraction; The residual vector within the time window is normalized and weighted to calculate the anomaly contribution coefficient E for each mode. m , , Take E m The largest dominant mode combination narrows the attribution search space. For example: if rforce (Haptic residual) and If the (joint torque residual) mutation is the most significant, then the attribution tree will preferentially enter the force control / skill constraint branch.

[0077] Step F3. Hierarchical Causal Bayesian Network (CBN) Inference, constructing a three-layer causal diagnostic network: Layer 1 (Phenomenon Layer): Detect direct failure indicators (e.g., object falling, emergency stop triggered, user takeover, excessive trajectory deviation).

[0078] Layer 2 (Physical Mechanism Layer): Analyze the physical causes based on the dominant modal residuals (e.g., "slippage" caused by sudden changes in shear force and insufficient normal force; "target loss" caused by sudden changes in visual target pose; "instability" caused by ZMP exceeding the supporting polygon).

[0079] Layer 3 (Asset Defect Mapping Layer): This layer precisely maps physical mechanism causes to specific embodied capability asset classes (as shown in Table 3 below). Table 3

[0080] Step F4. Output a structured attribution report; Output a structured vector C containing the attribution results. root , $, Among them, Type asset For the target asset type to be generated / updated, Conf score For the confidence level of causal inference, P delta This refers to the suggested parameter increments / structures for correction.

[0081] List of key parameters for the failure attribution module.

[0082] In practical implementation and programming, the failure attribution module depends on the following key parameter configurations: 1. Anomaly analysis time window length (W) win ): Definition: The length of time during which multimodal trajectories are analyzed in reverse before a location failure event.

[0083] Typical value range: 1.5s-5.0s (default can be set to 3.0s).

[0084] 2. Threshold for the rate of change of shear force (γ) in slip determination slip ): Definition: The threshold value of the derivative of the shear force with respect to time detected by the end-sensing sensor, used to determine grasping slippage.

[0085] Typical numerical range: 15.0 N / s - 50.0 N / s.

[0086] 3. Vision-tactile pose consistency deviation threshold (δ) pose ): Definition: The upper limit of the Euclidean distance between the target object pose predicted by the visual method and the actual pose obtained from tactile contact feedback.

[0087] Typical numerical range: 8.0mm-20.0mm.

[0088] 4. Joint moment saturation residual threshold (σ) τ ): Definition: The threshold difference between the actual joint torque and the torque calculated by the theoretical dynamic model is used to determine whether a collision or overload has occurred.

[0089] Typical value range: 15-30% of the rated maximum torque.

[0090] 5. Zero Moment Point (ZMP) Safety Boundary Radius (R) ZMP ): Definition: The absolute safe radius within the polygon supporting the robot's foot, allowing deviation from the ZMP trajectory. Exceeding this value is considered a gait / balance anomaly.

[0091] Typical numerical range: the area where the edge of the plantar support zone is recessed inward by 2.0-5.0 cm.

[0092] 6. Pre-trigger association time for manual takeover / emergency stop (∆t) pre ): Definition: When a manual takeover or emergency stop signal is received, the earliest time delay at which the underlying machine malfunction that triggered the human intervention is determined.

[0093] Typical value range: 200ms-800ms.

[0094] 7. Modal difference weighted coefficient vector (W) residual ): Definition: The weight vector W in the normalized residual calculation formula. , Used to balance the magnitude of sensor data with different dimensions.

[0095] Typical numerical values: [0.15, 0.20, 0.25, 0.25, 0.15] (can be adaptively calibrated according to sensor accuracy).

[0096] 8. Lower limit of attribution confidence (η) conf ): Definition: The posterior probability threshold inferred by a causal Bayesian network, only Conf scoreGreater than or equal to η conf Only then will the attribution results be accepted and trigger the generation of subsequent asset candidates.

[0097] Typical value range: 0.75-0.85 (default setting is 0.80).

[0098] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for the continuous evolution of a humanoid robot with multi-layered embodied capability assets, characterized in that, The method includes: S1. Receive user tasks. The system receives the task instructions input by the user through the task receiving module and parses the task into the robot's execution target. S2. Perceive environmental data and the entity's state; S3. Retrieve relevant capability assets based on the capability map; S4. Generate a task execution plan; S5. Perform the task and collect multimodal embodied trajectories; S6. Determine the task execution status; S7. Evaluate the task results; S8. Perform failure attribution or success experience extraction; S9. Generate multi-tiered embodied ability asset candidates; S10. Update the capability asset library and capability map; S11. Dynamically call upon capability assets in subsequent tasks.

2. The method for the continuous evolution of humanoid robots with multi-layered embodied capabilities according to claim 1, characterized in that, In step S2, the environmental data includes images, depth maps, point clouds, speech, ambient sounds, target objects, obstacles, human body positions, passable areas, and operable areas. The body state data includes robot pose, joint angles, joint velocities, joint torques, end effector status, tactile feedback, plantar pressure, inertial measurement unit data, battery status, and actuator health status.

3. The method for the continuous evolution of humanoid robots with multi-layered embodied capability assets according to claim 1, characterized in that, In step S3, the process of retrieving relevant capability assets based on the capability graph involves the system retrieving relevant capability assets from the capability graph and capability asset library according to the task type, scene conditions, target object, user preferences, and robot state. The retrieval results include environmental semantic assets, task strategy assets, motion skill assets, force control haptic assets, safety constraint assets, anomaly recovery assets, and user interaction assets.

4. The method for the continuous evolution of humanoid robots with multi-layered embodied capability assets according to claim 1, characterized in that, In step S4, the task planning process involves the system generating a task execution plan by combining the current task and retrieved capability assets. The task execution plan includes a sub-task sequence, action skill invocation order, target object and operation object, key poses, path planning constraints, force control constraints, human-machine confirmation nodes, safety constraints, and anomaly recovery strategies. In step S5, the task execution and multimodal embodied trajectory acquisition process involves the robot executing the task according to the task execution plan and continuously acquiring multimodal embodied trajectories during the execution process. The multimodal embodied trajectory includes task input, perception results, planning results, action skill invocation records, motion control data, force control tactile data, abnormal events, user feedback, and task results.

5. The method for the continuous evolution of humanoid robots with multi-layered embodied capability assets according to claim 1, characterized in that, In step S6, determining the task execution status means the system determines whether the task has been successfully completed. The task execution status includes complete success, partial success, failure, cancellation by the user, manual takeover, interruption by the security mechanism, and active termination by the robot.

6. The method for the continuous evolution of a humanoid robot with multi-layered embodied capability assets according to claim 1, characterized in that, In step S7, the task result evaluation is performed by the system based on the task objectives, user feedback, automatic evaluation rules, security events, and manual review information. The evaluation indicators include task success rate, completion time, operation accuracy, object damage, human-machine distance, collision events, emergency stop events, number of manual interventions, user satisfaction, energy consumption, execution stability, and anomaly recovery effect.

7. The method for the continuous evolution of a humanoid robot with multi-layered embodied capability assets according to claim 1, characterized in that, In step S8, the failure attribution or success experience extraction involves the system performing failure attribution when the task fails or partially succeeds, and extracting success experience when the task succeeds and the effect is better than the historical benchmark. The failure attribution and success experience extraction results are converted into evolutionary signals. The failure attribution types include task understanding failure, scene perception failure, target localization failure, object attribute judgment failure, task decomposition failure, execution order error, action skill selection error, key pose selection error, improper grasping force control, unreasonable movement path planning, insufficient human-computer interaction confirmation, insufficient safety constraint triggering, lack of anomaly recovery strategy, abnormal robot body state, and scene changes causing the original plan to fail.

8. The method for the continuous evolution of a humanoid robot with multi-layered embodied capability assets according to claim 1, characterized in that, In step S9, the generation of multi-layered embodied capability asset candidates involves the system generating multi-layered embodied capability asset candidates based on evolutionary signals. These candidates include new environmental semantic assets, new task strategy assets, new motion skill assets, new force control haptic assets, new safety constraint assets, new anomaly recovery assets, new user interaction assets, new tool usage assets, new model adaptation assets, and new capability graph relationships. The system maps failure attribution results to different types of capability asset candidates: if the task failure is due to incorrect target object recognition, environmental semantic asset candidates are generated; if the failure is due to an unreasonable task execution order, task strategy asset candidates are generated; if the failure is due to unstable grasping points, motion skill assets or force control haptic assets candidates are generated; if the failure is due to the robot getting too close to the human body, safety constraint assets candidates are generated; if the failure is due to ambiguous user instructions, user interaction assets candidates are generated; and if manual takeover successfully completes the task after failure, anomaly recovery asset candidates are generated from the manual takeover trajectory.

9. The method for the continuous evolution of a humanoid robot with multi-layered embodied capability assets according to claim 1, characterized in that: The process of updating the capability asset library and capability graph involves the system writing candidate capability assets that have passed basic verification into the candidate asset area and updating the relationships between tasks, scenarios, failure modes, assets, and effects in the capability graph. High-risk action skills, force control strategies, and safety constraint assets can enter a pending verification state for further processing by the safety verification module. The process of dynamically calling capability assets in subsequent tasks involves the system dynamically calling existing capability assets based on the capability graph and asset scores when the robot performs new tasks of the same or similar nature, enabling the robot to continuously improve its task execution capabilities over long-term operation.

10. A continuously evolving system for humanoid robots with multi-layered embodied capability assets, characterized in that, The system for implementing the method for continuous evolution of a humanoid robot with multi-layered embodied capability assets as described in any one of claims 1-9 includes: a task receiving module, a multimodal perception module, an ontology state acquisition module, a capability asset library, a capability map module, a task planning module, an action execution module, a trajectory acquisition module, a task evaluation module, a failure attribution module, an asset generation module, and a dynamic invocation module. Users submit task instructions to the system through the task receiving module. Task instructions can include voice, text, gestures, images, or multimodal forms. The multimodal perception module collects and analyzes information about the robot's surrounding environment, including images, depth, point clouds, voice, tactile information, target objects, human body position, and obstacle information. The body state acquisition module collects the robot's own state, including joint angles, joint velocities, joint torques, end effector state, plantar pressure, inertial measurement unit data, battery status, and actuator health status; The capability asset library is used to store multi-layered embodied capability assets, including environmental semantic assets, task strategy assets, motion skill assets, force control and haptic assets, safety constraint assets, anomaly recovery assets, user interaction assets, tool usage assets, and model adaptation assets. The capability graph module is used to record the relationships between task types, scenario conditions, failure modes, success strategies, capability assets, robot status, and operational results. The task planning module generates a task execution plan based on the user task, environment information, ontology status, and relevant capability assets retrieved from the capability graph. The motion execution module calls the motion controller, motion skills, and tool interfaces to drive the robot to complete the task; The trajectory acquisition module records the robot's multimodal embodied trajectory during task execution. The task evaluation module evaluates the effectiveness of task execution based on task objectives, execution results, user feedback, security incidents, and manual takeover information. The failure attribution module analyzes the reasons for task failure or performance degradation and outputs the failure attribution results. The asset generation module generates multi-layered embodied ability asset candidates based on evolutionary signals; The dynamic invocation module dynamically invokes existing capability assets based on the capability graph and asset score in subsequent tasks; The modules interact with each other via an internal bus or application programming interface. The task receiving module receives multimodal user instructions at its input end and outputs a structured task target vector at its output end. The multimodal perception module outputs an environmental feature matrix containing spatial coordinates and semantic labels. The ontology state acquisition module outputs an ontology state vector containing joint state and torque information.