A mind map-based long-time task planning method and system for a home robot

CN122840186APending Publication Date: 2026-09-29SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611009379.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0002]家庭服务机器人在清洁、收纳、取放、备餐等任务中面临如下挑战:其一,家庭任务通常由多步骤组成且具有显式依赖关系,例如“打开柜门—取出物品—移动—放置—关闭柜门”;其二,家庭环境高度非结构化且动态变化,目标物体可能被遮挡或移动,导致仅依赖单轮推理的策略易出现目标漂移;其三,现有任务管理方式多以列表或简单子任务拆解为主,缺少可复用的结构化规划模板、可执行节点调度机制以及以验收条件驱动的闭环纠错机制

Benefits of technology

[0042]有益效果:与现有技术相比,本发明的显著技术效果为:通过在任务规划层引入思维图(GoT)作为显式中间结构,将家庭长时任务分解为带依赖约束与检查点的可执行节点,克服了仅依赖单轮提示驱动VLA时缺乏任务结构、难以进度跟踪的问题;通过“节点调度+提示编译”的机制,将任务语义、场景状态与机器人状态以结构化方式注入VLA输入,使模型在多阶段长程执行中保持目标一致性并提升推理稳定性;通过执行检测与失败语义驱动的动态调整策略,实现对家庭环境动态变化、常见操作失败的可恢复闭环控制,从而显著提高家庭机器人长时任务的成功率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840186A_ABST
    Figure CN122840186A_ABST
Patent Text Reader

Abstract

This invention discloses a long-term task planning method and system for home robots based on mind maps, comprising: acquiring long-term task instructions for home robots and parsing the task instructions into structured semantic descriptions; using multimodal sensors to collect scene state information and robot state information of the surrounding environment, and constructing scene state representations and robot state representations; generating a mind map based on the structured semantic descriptions, mind map templates, and scene and robot state representations, and updating the mind map when dynamic adjustments are triggered; selecting the current executable node, and compiling the structured semantic description, scene state representation, and robot state representation of the current executable node into a multimodal cue injection package; receiving the injection package, performing inference to generate an action plan, and executing it; and performing acceptance testing on the current executable node based on execution records and real-time scene states. This method can improve the consistency, traceability, and recoverability of long-term task execution for home robots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of service robots, embodied intelligence, multimodal large models and task planning, and particularly to a long-term task planning method and system for home robots based on mind mapping. Background Technology

[0002] Home service robots face the following challenges in tasks such as cleaning, tidying, retrieving and placing items, and preparing meals: First, home tasks typically consist of multiple steps with explicit dependencies, such as "opening a cabinet door—taking out an item—moving—placing—closing the cabinet door"; Second, the home environment is highly unstructured and dynamically changing, and target objects may be obscured or moved, making strategies relying solely on single-round reasoning prone to target drift; Third, existing task management methods mainly rely on lists or simple sub-task decomposition, lacking reusable structured planning templates, executable node scheduling mechanisms, and closed-loop error correction mechanisms driven by acceptance conditions.

[0003] Therefore, there is an urgent need for a task planning method and system that can explicitly express the structure of long-term tasks at the planning level and form a closed-loop collaboration with VLA inference execution. Summary of the Invention

[0004] Purpose of the invention: The present invention aims to provide a method and system for long-term task planning of home robots based on mind maps, which enables long-term home tasks to have closed-loop execution capabilities that are decomposable, schedulable, promptable, detectable, and dynamically adjustable, thereby improving the consistency and robustness of VLA in long-term tasks.

[0005] Technical solution: The method described in this invention is applied to a home robot system comprising a multimodal sensor, a robotic arm, a control terminal, and a mobile chassis. The method includes the following steps:

[0006] S1. Family Task Acquisition: Use the control terminal to acquire long-term family task instructions, and use VLA's own Large Language Model (LLM) to parse the task instructions into structured semantic descriptions for subsequent generation of mind map root nodes;

[0007] S2. Scene perception and state construction: Use multimodal sensors to collect scene state information and robot state information of the surrounding environment, and construct scene state representation and robot state representation in this way;

[0008] S3. Mind Map Generation and Update: When generating a mind map, the root node of the mind map is first obtained based on the structured semantic description. Then, the mind map template is used to generate a mind map containing a set of task nodes and a set of dependency edges by combining the scene state representation and the robot state representation. When the dynamic adjustment strategy is triggered and the mind map needs to be updated, the mind map is modified and updated accordingly based on the failure semantics in step S6.

[0009] S4. Node Scheduling and Hint Compilation: Filter the candidate set of executable nodes based on the set of dependency edges in the mind graph, and select the current executable node as the execution object for this round from the candidate set of executable nodes; compile the structured semantic description, scene state representation and robot state representation of the current executable node into a multimodal hint injection package for hint injection in step S5;

[0010] S5. VLA Reasoning and Execution: Inject the multimodal cue injection package into the Visual-Language-Motion Model (VLA), perform multimodal fusion reasoning and generate action plans, drive the robotic arm or mobile chassis to perform actions, obtain the scene state representation after execution, and form an execution record;

[0011] S6. Execution Detection and Dynamic Adjustment: Based on the execution record, perform acceptance testing on the current executable node to determine whether the current executable node is completed; when the current executable node is not completed, identify the failure semantics and trigger the dynamic adjustment strategy, return to the mind map update in step S3, and perform dynamic adjustment of the mind map; when the current executable node is completed, further check whether the root node is completed. If the root node is completed, output a task completion prompt and end the process; otherwise, return to the mind map update in step S3, update the status of the current executable node to completed, and continue to execute the task in a closed loop.

[0012] Furthermore, the structured semantic description includes one or more of the following: target semantics, target key objects, and target regions; the scene state representation includes at least one or more of the following: the set of target objects in the scene, the spatial location of the target objects, and the target state; the robot state representation includes at least one or more of the following: the pose of the robotic arm, the state of the robotic arm end effector, and the position of the mobile chassis.

[0013] Furthermore, mind maps are represented as directed graph structures. , ,in For a set of task nodes, For each task node, there is a set of dependent edges. Configure node attributes, which should include at least: node identifier, node semantics, node hierarchy depth, estimated execution cost, node status, and acceptance criteria; and a set of dependency edges. Used to describe the sequential dependencies between nodes.

[0014] Furthermore, the selection rules for the current executable nodes are as follows:

[0015] If and only if node When all predecessor nodes except the root node are in a completed state, the node It belongs to the executable node and uses the priority function. Select the currently executable node , The expression is:

[0016] ;

[0017] in, The current executable node The overall evaluation priority value; The estimated execution cost for executable nodes, The depth of the executable node level. , The weights are preset and the constraints are met. ;choose The smallest executable node is selected as the current executable node. .

[0018] Furthermore, the compilation prompts include:

[0019] Current executable node Structured semantic description ;

[0020] Current scene state representation ;

[0021] Current robot state representation ;

[0022] Create a multimodal cue injection package for use by VLA. ;

[0023] It must contain at least one or more of the following: target semantics, target key objects, and target regions; It must include at least one or more of the following: a set of target objects, the spatial location of the target objects, and the target state; It includes at least one or more of the following: robotic arm pose, robotic arm end effector state, and mobile chassis position.

[0024] Furthermore, the expected scene state representation is constructed using VLA's own Visual Language Model (VLM). And compare the expected scenario state representation with the scenario state representation after execution. Compare and define the scene state difference metric for the current executable node. Acceptance testing will be conducted.

[0025] ;

[0026] in, This represents the difference between the expected scenario state representation and the actual scenario state representation. The range of values ​​is When the expected scenario state representation perfectly matches the scenario state representation, ; Let the difference function represent the scene state. This represents the expected scenario state. This represents the scene state. for The total number of target objects in the target object set. This represents one of the target objects. Represents the target object The expected spatial location, Represents the target object Spatial location, This represents the Euclidean distance between the two. This indicates the total number of target objects whose expected state differs from the actual state; acceptance criteria. Provide a difference threshold It is dynamically generated by the Visual Language Model (VLM) based on the semantics of the currently executable node; when Determine the currently executable node at time The task is completed if successful; otherwise, it is considered a failure, and the Visual Language Model (VLM) analyzes the current scene state representation to generate failure semantics. And triggers a dynamic adjustment strategy; failure semantics This includes one or more of the following: target not found, grab failed, or placement failed.

[0027] Furthermore, dynamic adjustment strategies include:

[0028] When failure semantics When a capture or placement fails, the mind map update in step S3 is triggered, a recovery child node is added to the current executable node and the recovery child node is executed first. The recovery child node is: re-execute the current executable node once.

[0029] When failure semantics If the target is not found, the mind map update in step S3 is triggered, a recovery sub-node is added to the current executable node and the recovery sub-node is executed first. The recovery sub-node is: first locate the target in the scene, and then re-execute the current executable node.

[0030] Based on the same inventive concept, the present invention provides a mind map-based long-term task planning system for home robots, applicable to a home robot system comprising multimodal sensors, a robotic arm, a control terminal, and a mobile chassis. The task planning system includes:

[0031] The multimodal input unit is used to receive user commands and real-time scene and robot state information from multimodal sensors;

[0032] The VLA large model perception unit uses the VLA large model to perceive the input user commands, scene state information, and robot state information, and obtain structured semantic descriptions, scene state representations, and robot state representations.

[0033] The mind map unit is used to construct the mind map. First, the structured semantic description obtained by the VLA large model perception unit is used to generate the root node of the mind map. Then, the mind map containing the set of task nodes and the set of dependency edges is generated according to the mind map template, scene state representation and robot state representation.

[0034] The mind map state management unit is used to maintain the structure and content of the mind map.

[0035] The mind map node scheduling unit is used to select the currently executable node. If and only if node When all predecessor nodes except the root node are in a completed state, the node It belongs to the executable node, according to the priority function Choose the priority function The smallest executable node is selected as the current executable node. ; The estimated execution cost for executable nodes, The depth of the executable node level. , Preset weights;

[0036] Node hint injection unit, used to inject the currently executable node The structured semantic description, scene state representation, and robot state representation in the VLA are fused into a multimodal cue injection package that can be received by VLA.

[0037] The VLA large model inference unit receives multimodal prompt injection packets, performs multimodal inference, and outputs action plans.

[0038] The motion generation unit uses motion schemes to drive the robot to perform actions.

[0039] The action detection and graph correction unit is used to detect the execution result, and correct the mind map if the detection fails.

[0040] Based on the same inventive concept, the present invention provides a non-volatile storage medium for storing a computer program, wherein the computer program implements the method described herein when executed by a processor.

[0041] Based on the same inventive concept, the present invention provides a computer program product comprising a computer program / instruction that, when executed by a processor, implements the method described herein.

[0042] Beneficial Effects: Compared with existing technologies, the significant technical effects of this invention are as follows: By introducing a mind map (GoT) as an explicit intermediate structure at the task planning layer, long-term household tasks are decomposed into executable nodes with dependency constraints and checkpoints, overcoming the problems of lacking task structure and difficulty in progress tracking when relying solely on single-round prompts to drive VLA; Through the mechanism of "node scheduling + prompt compilation", task semantics, scene state, and robot state are injected into VLA input in a structured manner, enabling the model to maintain goal consistency and improve inference stability during multi-stage long-term execution; Through the dynamic adjustment strategy driven by execution detection and failure semantics, recoverable closed-loop control for dynamic changes in the home environment and common operational failures is achieved, thereby significantly improving the success rate of long-term household robot tasks. Attached Figure Description

[0043] Figure 1 This is a flowchart of the long-term task planning method for home robots according to the present invention;

[0044] Figure 2 This is a block diagram of the long-term task planning system for the home robot of the present invention. Detailed Implementation

[0045] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0046] This invention discloses a long-term task planning method and system for home robots based on mind maps, which is applicable to long-term, multi-step robot operation tasks such as storage and organization, retrieval and placement of items, and cleaning preparation in home scenarios. To address the issue of target drift and interruption in traditional end-to-end control of Vision-Language-Motion (VLA) in home scenarios due to the numerous steps, complex dependencies, and dynamic environmental changes, this invention pre-constructs a mind map template for home tasks. This template is used to generate mind maps for specific long-term home tasks, serving as an explicit intermediate structure. In the multimodal perception stage, user instructions, real-time scene information, and robot state information are perceived and acquired, resulting in a structured semantic description of the user instructions, a scene state representation, and a robot state representation. In the mind map task planning stage, the root node of the mind map is first obtained based on the structured semantic description. Then, the complete mind map is generated by combining the mind map template, scene state representation, and robot state representation. Subsequently, node scheduling and prompt compilation are performed on the generated mind map. The structured semantic description, scene state representation, and robot state representation of the currently executable node are compiled into a multimodal prompt injection package and injected into the VLA. In the inference, execution, and detection adjustment stage, the VLA generates and executes corresponding action plans based on the multimodal prompt injection package. Finally, a closed loop is formed through node completion detection and root node completion detection. When a detection is incomplete or fails, the mind map is dynamically adjusted by expansion or updating until the root task is completed. This method can improve the consistency, traceability, and recoverability of home robots in long-term task execution.

[0047] like Figure 1 As shown, a long-term task planning method for home robots based on mind mapping is applied to a home robot system including multimodal sensors, a robotic arm, a control terminal, and a mobile chassis. This method enhances the long-term reasoning and execution consistency of the visual-language-action (VLA) model in a home setting through mind mapping (GoT). The multimodal sensor can be an RGB-D camera. The method includes the following steps:

[0048] S1. Family Task Acquisition: Acquire user commands and parse them to obtain structured semantic descriptions. Specifically:

[0049] The system uses a control terminal to obtain user-inputted long-term household task instructions, and leverages VLA's own Large Language Model (LLM) to parse the task instructions into structured semantic descriptions. The structured semantic description includes one or more of the following: target semantics, target key objects, and target regions. This structured semantic description is used to subsequently generate the root node of the mind map.

[0050] For example:

[0051] ;

[0052] in, For target semantics, such as "tidy up the living room and take the trash to the door"; For key targets, such as "garbage"; For the target area, such as "living room, entrance".

[0053] S2. Scene Perception and State Construction: The robot uses multimodal sensors (an RGB-D camera in this embodiment) to acquire real-time scene state information and robot state information, and uses this information to construct a scene state representation. and robot state representation Specifically:

[0054] The robot's multimodal sensors collect scene state information and robot state information, which are then imported into the VLA's own Visual Language Model (VLM) to generate scene state representations and robot state representations. The scene state representation includes at least one or more of the following: a set of target objects in the scene, the spatial location of the target objects, and the target state. The robot state representation includes at least one or more of the following: the robot arm pose, the state of the robot arm's end effector, and the position of the mobile chassis. Robot state representation and the structured semantic description obtained in step S1 Used for generating mind maps in subsequent S3.

[0055] S3. Mind Map Generation and Update: This step, when generating the mind map, first utilizes structured semantic description. Generate the root node of the mind map, and then use the mind map template combined with the scene state representation. Robot state representation Generate a set containing task nodes With the set of dependent edges The mind map. When the dynamic adjustment strategy requires updating the mind map, it is based on the failure semantics of step S6. Make the necessary modifications and updates to the mind map.

[0056] The mind map template is a pre-written, general prompt for generating mind maps for long-term family tasks. Inputting this template into VLA's own visual language model (VLM) can constrain VLM to generate mind maps with strictly defined structures. Specifically, the mind map template is as follows:

[0057] "You are an embodied intelligent reasoning model specifically designed for long-term task planning in home robots."

[0058] Your task is to: based on the input root node and the scene state representation and robot state representation It automatically generates a mind map.

[0059] Mind map is defined as ,in For a set of task nodes, This is a set of dependent edges. It is also a set of task nodes. Each node represents a subtask unit that has clear semantics, is executable, and verifiable. The task node format is {node identifier}. Node semantics Node hierarchy depth Node expected execution cost Node status Acceptance conditions Dependency edge set This indicates the execution dependencies between tasks. The dependency edge format is: {"from": " ","to": " ",}, Dependency semantics: (A→B) means that node A must be completed before node B can be executed.

[0060] Generate a mind map strictly following these steps: Step 1: Understand the root node task, identify key target objects, and identify the target area. Step 2: Infer object reachability based on scene state. Step 3: Infer execution feasibility based on robot state. Step 4: Decompose subtasks. Step 5: Generate search nodes. Step 6: Generate navigation nodes. Step 7: Generate interactive operation nodes. Step 8: Generate transport nodes. Step 9: Generate dependencies between nodes. Step 10: Check the graph structure.

[0061] Mind map generation must be strictly constrained by the scene state, the target object state, and the robot state.

[0062] The following commands are prohibited: unreachable tasks, ungraspable tasks, spatial conflict tasks, state conflict tasks, and tasks that the robot cannot execute.

[0063] The output format must be valid JSON; outputting interpreted text is prohibited. An example of the output format is as follows:

[0064] { "nodes": [{" ": " ",

[0065] " ": "{"Locate toys and storage boxes in the living room", "Toys, storage boxes", "Living room"}",

[0066] " 0.3,

[0067] " ": 1,

[0068] " 0.7,

[0069] " ": {1},

[0070] " ": "Incomplete",

[0071] " ": [" "]}],

[0072] "edges": [{"from": "root",

[0073] to: " ",}]

[0074] }

[0075] This mind map is used to guide VLA in completing long-duration home robot tasks.

[0076] The mind map includes a set of task nodes. With the set of dependent edges This is represented as a directed graph structure. All nodes in the directed graph form a task node set. Each node is configured with its attributes during generation, and the node attributes must include at least the node identifier. Node semantics Node hierarchy depth Node Expected Execution Cost Node status and acceptance conditions This enables mind maps to form a schedulable, executable, and verifiable task structure.

[0077] The set of dependent edges Used to depict the sequential dependencies between nodes, the mind map satisfies the schedulable characteristic of "execution is only possible after dependencies are unlocked", thus forming an interpretable and traceable long-term task execution skeleton, and providing a structured basis for subsequent prompts, compilation, execution and acceptance.

[0078] For example, the root node is first generated based on the structured semantic description, and then combined with the mind map template and the scene state representation. Robot state representation To generate a complete mind map:

[0079] Root: {"tidy up the living room and take the trash to the door", "trash", "living room, doorway"};

[0080] {"Locating toys and storage boxes in the living room", "Toys, storage boxes", "Living room"};

[0081] : {"Locate empty bottles, scraps of paper, and garbage bags in the living room", "empty bottles, scraps of paper, and garbage bags", "living room"};

[0082] {"Grab toys and put them into storage boxes", "Toys, storage boxes", "Living room"};

[0083] {"Grab an empty bottle and put it in a trash bag", "empty bottle, trash bag", "living room"};

[0084] {"Grab the scraps of paper and put them in the trash bag", "Scraps of paper, trash bag", "Living room"};

[0085] : {"Grab the trash bag and move it to the doorway", "trash bag", "doorway"};

[0086] in, , , , , and It has 6 child nodes.

[0087] Simultaneously establish a set of dependent edges :(root→ (root→) ), ( → ), ( → ), ( → ), ( → ), ( → ).

[0088] S4. Node Scheduling and Compilation Hints: Based on the set of dependency edges in the mind graph, filter the candidate set of executable nodes and select the current executable node from it. As the target of this round of execution; the currently executable node. Structured semantic description Scene state representation Robot state representation The input suggestion compiler generates a multimodal suggestion injection package to drive the vision-language-action model (VLA);

[0089] Furthermore, the node scheduling and prompt compilation adopts a scheduling mechanism of "restricted candidate set + priority selection": First, an executable node candidate set is constructed based on the dependency edge set; then, the current executable node is selected based on a priority function that includes the node's expected execution cost and node level depth, and the structured semantic description, scene state representation, and robot state representation of the current executable node are passed to the prompt compiler to form a multimodal prompt injection package for the VLA model; wherein, the prompt compiler directly integrates the structured semantic description, current scene state representation, and robot state representation of the current executable node, so that VLA maintains consistent constraints on the root target in long-term tasks and reduces target drift caused by context redundancy or missing context.

[0090] The current selection rules for executable nodes are as follows:

[0091] If and only if node When all predecessor nodes except the root node are in a completed state, the node It belongs to the executable node; and according to the priority function Select the currently executable node ,in, The current executable node The overall evaluation priority value is the value of the executable node. The smaller the value, the higher the overall priority of the executable node and the more likely it is to be scheduled for execution. The estimated execution cost for an executable node is obtained by the Visual Language Model (VLM) predicting the time, distance traveled, and energy consumption required to perform the action, with values ​​ranging from [value range missing]. ; The depth of the executable node hierarchy is 0, with the depth of the root node increasing sequentially for child nodes. This depth is obtained during mind map generation. , To preset weights, , And satisfy the constraints. , and The specific value of is given by the visual language model during mind map generation. To correctly complete long-term tasks, the model prioritizes scheduling shallow nodes, and then considers time and energy consumption. , ;choose The smallest executable node is selected as the current executable node. Then, the compilation and injection are performed with prompts, resulting in:

[0092] ;

[0093] in, Inject packages for multimodal hints. For structured semantic description, This represents the scene state. For robot state representation, It must contain at least one or more of the following: target semantics, target key objects, and target regions; It must include at least one or more of the following: a set of target objects, the spatial location of the target objects, and the target state; It includes at least one or more of the following: robotic arm pose, robotic arm end effector state, and mobile chassis position.

[0094] S5, VLA Inference and Execution: Inject the multimodal hints generated in step S4 into the package. The data is fed into a Vision-Language-Motion (VLA) model for multimodal fusion inference to generate action plans, which serve as robot control commands to drive the robot arm or mobile chassis to perform actions. The robot executes the control commands to complete the current executable node. The corresponding action is executed, and the scene state representation after execution is obtained to form an execution record;

[0095] As a further technical limitation, the VLA model is a multimodal model that combines visual perception, language understanding, and action generation, aiming to improve the robot's interaction capabilities in complex environments. By deeply fusing vision, language, and action, the VLA model can directly predict continuous control commands from multimodal inputs, achieving closed-loop control from environmental understanding to physical execution. Here, a pre-trained open-source VLA model is directly used. This VLA model is used for perception, reasoning, and action generation. It has strong generalization ability in complex real-world scenarios, and can understand task semantics, break down complex task processes, and accurately execute actions.

[0096] As a further technical limitation, the VLA inference and execution steps support two action implementation methods: first, VLA directly outputs the underlying motion control sequence and drives the robot to execute it; second, VLA outputs high-level skill call instructions to drive the robot to execute actions, with the VLA's built-in skill library completing the parameterized execution of actions such as navigation, grasping, placement, and opening and closing doors and cabinets. In either of the above methods, the structured semantic description, real-time scene state representation, and robot state representation are used as joint inputs to achieve end-to-end mapping or end-to-end decision-skill collaboration from high-level task semantics to low-level motion control.

[0097] Will Input the VLA model to obtain the action plan And execute:

[0098] ;

[0099] in, It adopts a vision-language-action strategy and can output low-level control sequences or high-level skill call commands (such as navigation, grabbing, placing, opening and closing cabinet doors).

[0100] S6. Execution Detection and Dynamic Adjustment: Based on execution records, the current executable nodes are monitored and adjusted. Perform acceptance testing to determine if the current executable node has been completed; if the current executable node has not been completed, identify the failure semantics. It triggers a dynamic adjustment strategy, returns to the mind map update in step S3, and performs dynamic adjustment of the mind map. When the current executable node is completed, it further checks whether the root node is completed. If the root node is completed, it outputs a task completion prompt and ends the process. Otherwise, it returns to the mind map update in step S3, updates the status of the current executable node to complete, and continues to execute the task in a closed loop.

[0101] As a further technical limitation, the execution detection and dynamic adjustment steps map acceptance conditions to observable scene state representation criteria, utilize VLA's own Visual Language Model (VLM) to construct the expected scene state representation, and perform consistency judgment on the "expected scene state representation - scene state representation". When acceptance failure is detected, the failure is semantically classified, including one or more of the following: target not found, grasping failure, and placement failure. Based on the failure semantics, the corresponding dynamic adjustment strategy is triggered. Specifically, for grasping failure or placement failure, a recovery sub-node is added to the current executable node and executed with priority. The recovery sub-node is: re-execute the current executable node once. For the failure of target not found, a recovery sub-node is added to the current executable node and executed with priority. The recovery sub-node is: first locate the target in the scene, and then re-execute the current executable node once. This forms a closed-loop control of "execution - perception - verification - write-back - replanning".

[0102] The consistency judgment involves comparing the expected scenario state representation with the executed scenario state representation and defining a scenario state representation difference metric. For the current executable node Acceptance testing will be conducted.

[0103] ;

[0104] in, This represents the difference between the expected scenario state representation and the actual scenario state representation. The range of values ​​is When the expected scenario state representation perfectly matches the scenario state representation, ; Let the difference function represent the scene state. This represents the expected scenario state. This represents the scene state. for The total number of target objects in the target object set. This represents one of the target objects. Represents the target object The expected spatial location (three-dimensional coordinates). Represents the target object Spatial location (three-dimensional coordinates) This represents the Euclidean distance between the two. This represents the total number of target objects whose expected target state does not match the actual target state.

[0105] Acceptance conditions Provide a difference threshold The range of values ​​is It is dynamically generated by the Visual Language Model (VLM) based on the semantics of the currently executable node. For navigation or movement-type nodes, it has high spatial tolerance. Configure to a larger value, such as It allows for operational errors of tens of centimeters, but for interactive or crawling nodes, the requirements for position and state consistency are extremely high. Configure to a smaller value, such as The permissible positional error is within three to five centimeters;

[0106] like If the difference is too large, then the current executable node is determined. The task was not completed, and the Visual Language Model (VLM) analyzed the scene state representation to obtain the semantics of task failure. The task failure semantics include one or more of the following: target not found, grabbing failure, and placement failure. The following dynamic adjustment strategy is then executed:

[0107] When failure semantics When a capture or placement fails, the mind map update in step S3 is triggered, a recovery child node is added to the current executable node and the recovery child node is executed first. The recovery child node is: re-execute the current executable node once.

[0108] When failure semantics If the target is not found, the mind map update in step S3 is triggered, a recovery sub-node is added to the current executable node and the recovery sub-node is executed first. The recovery sub-node is: first locate the target in the scene, and then re-execute the current executable node.

[0109] like If the difference is acceptable, then the current executable node is considered acceptable. The task is completed, and the root node task is further checked. If the root node is completed, the process ends; otherwise, the process returns to the mind map update in step S3, updates the current executable node status to completed, and continues to execute the task in a closed loop.

[0110] like Figure 2 As shown, a mind-map-based long-term task planning system for home robots includes:

[0111] Input and perception module: used to acquire user home task instructions, collect scene state information and robot state information, and generate structured task semantic descriptions, scene state representations and robot state representations;

[0112] This module serves as the system's information entry point, receiving user commands and performing multimodal perception of the home environment. It outputs structured semantic descriptions, scene state representations, and robot state representations. This module integrates a multimodal input unit, a VLA large-model perception unit, and a mind map construction unit.

[0113] Multimodal input unit: used to receive user commands and scene state information and robot state information from multimodal sensors;

[0114] VLA Large Model Perception Unit: Uses the VLA large model to perceive the input user commands, scene state information, and robot state information respectively, and obtain a structured semantic description. Scene state representation and robot state representation ;

[0115] Constructing Mind Map Units: Structured Semantic Descriptions Obtained Using VLA Large Model Perceptual Units To generate the root node of the mind map, and then based on the mind map template and scenario state representation. Robot state representation Generate a set containing task nodes With the set of dependent edges Mind Map .

[0116] Mind Map Management and Hint Injection Module: Used to manage mind maps and select currently executable nodes during node scheduling. and the structured semantic description of the current executable node. Scene state representation and robot state representation Fusion into a VLA-acceptable multimodal cue injection package When the trigger diagram is corrected, the mind map is modified and updated accordingly.

[0117] This module serves as the core of the system's high-level planning and prompting engineering, responsible for maintaining the mind map structure and node states. During node scheduling, it selects the currently executable node and provides its structured semantic description. Scene state representation and robot state representation Fusion into a VLA-acceptable multimodal cue injection package When the mind map needs to be updated due to the trigger graph correction, the failure semantics of the action detection and graph correction unit in the action generation and detection module are used. The module allows for corresponding modifications and updates to the mind map. It includes a mind map state management unit, a mind map node scheduling unit, and a node prompt injection unit.

[0118] Mind Map Status Management Unit: Used to maintain mind maps The structure and content of the diagram can be used to create a mind map. Represented as a directed graph: ;

[0119] in, For mind maps, For a set of task nodes, For each task node, there is a set of dependent edges. It must include at least: node identifier Node semantics Node hierarchy depth Node Expected Execution Cost Node status and acceptance conditions ;

[0120] Mind graph node scheduling unit: A node is scheduled if and only if all its predecessor nodes except the root node are in a completed state. It belongs to the executable node; for all executable nodes According to the priority function Select the currently executable node ;in, The current executable node The overall evaluation priority value is the value of the executable node. The smaller the value, the higher the overall priority of the executable node and the more likely it is to be scheduled for execution. The estimated execution cost for an executable node is obtained by the Visual Language Model (VLM) predicting the time, distance traveled, and energy consumption required to perform the action, with values ​​ranging from [value range missing]. ; The depth of the executable node hierarchy is 0, with the depth of the root node increasing sequentially for child nodes. This depth is obtained during mind map generation. , To preset weights, , And satisfy the constraints. , and The specific value of is given by the visual language model during mind map generation. To correctly complete long-term tasks, the model prioritizes scheduling shallow nodes, and then considers time and energy consumption. , ;choose The smallest executable node is selected as the current executable node. ;

[0121] Node hint injection unit: used to inject the currently executing node Structured semantic description in Scene state representation and robot state representation Fusion into a VLA-acceptable multimodal cue injection package :

[0122] ;

[0123] in, Inject packages for multimodal hints. Therefore, the structured semantic description of the currently executable node, This represents the scene state. For robot state representation, It must contain at least one or more of the following: target semantics, target key objects, and target regions; It must include at least one or more of the following: a set of target objects, the spatial location of the target objects, and the target state; It includes at least one or more of the following: robotic arm pose, robotic arm end effector state, and mobile chassis position.

[0124] Action generation and detection module: Uses VLA to receive multimodal cue injection packets. It performs reasoning and outputs control commands to drive the robot to perform actions, and performs action detection on the execution results, identifies failure semantics and triggers graph correction to achieve closed-loop execution of long-term tasks;

[0125] This module serves as the system's execution and closed-loop assurance unit. It performs VLA inference, action generation, and execution result detection under prompt injection conditions, and writes the detection results back for mind graph correction. This module includes a VLA large-scale model inference unit, an action generation unit, and an action detection and graph correction unit.

[0126] VLA Large Model Inference Unit: Receives multimodal hint injection packets Performing multimodal reasoning and outputting action plan A can be represented as:

[0127] ;

[0128] in, This is a vision-language-action model strategy that can output low-level control sequences or high-level skill invocation commands (such as navigation, grasping, placing, opening and closing cabinet doors). It also acquires a representation of the scene state after the action is executed. For subsequent testing and adjustment;

[0129] Motion generation unit: Uses motion scheme A to drive the robot to perform actions.

[0130] Action Detection and Graph Correction Unit: Used to detect execution results and construct the expected scene state representation using VLA's own Visual Language Model (VLM). And represent it with the scene state after execution. Compare and define the scene state representation difference measure. For the current executable node Perform motion detection:

[0131] ;

[0132] in, This represents the difference between the expected scenario state representation and the actual scenario state representation. The range of values ​​is When the expected scenario state representation perfectly matches the scenario state representation, ; Let the difference function represent the scene state. This represents the expected scenario state. This represents the scene state. for The total number of target objects in the target object set. This represents one of the target objects. Represents the target object The expected spatial location (three-dimensional coordinates). Represents the target object Spatial location (three-dimensional coordinates) This represents the Euclidean distance between the two. This represents the total number of target objects whose expected target state does not match the actual target state.

[0133] Acceptance conditions Provide a difference threshold The range of values ​​is It is dynamically generated by the Visual Language Model (VLM) based on the semantics of the currently executable node. For navigation or movement-type nodes, it has high spatial tolerance. Configure to a larger value, such as It allows for operational errors of tens of centimeters, but for interactive or crawling nodes, the requirements for position and state consistency are extremely high. Configure to a smaller value, such as The permissible positional error is within three to five centimeters;

[0134] like If the difference in the scene state representation is too large, it is considered that the task at the currently executable node has not been completed. The Visual Language Model (VLM) analyzes the scene state representation at this time to obtain the task failure semantics. (Including target not found, grab failure, and placement failure), and trigger the following dynamic adjustment strategy to correct the graph:

[0135] When failure semantics When a capture or placement fails, return to the mind map state management unit, add a recovery sub-node to the current executable node, and execute the recovery sub-node first. The recovery sub-node is: re-execute the current executable node once.

[0136] When failure semantics If the target is not found, return to the mind map state management unit, add a recovery sub-node to the current executable node and execute the recovery sub-node first. The recovery sub-node is: first locate the target in the scene, and then re-execute the current executable node.

[0137] like If the difference in the scenario state is acceptable, the current executable node task is completed, and the root node task is further checked. If the root node task is completed, the process ends; otherwise, the process returns to the mind map state management unit, updates the current executable node state to complete, and continues to execute the task in a closed loop.

Claims

1. A long-term task planning method for home robots based on mind maps, characterized in that, The method, applied to a home robot system comprising multimodal sensors, a robotic arm, a control terminal, and a mobile chassis, includes the following steps: S1. Family Task Acquisition: Use the control terminal to acquire long-term family task instructions, and use VLA's own Large Language Model (LLM) to parse the task instructions into structured semantic descriptions for subsequent generation of mind map root nodes; S2. Scene perception and state construction: Use multimodal sensors to collect scene state information and robot state information of the surrounding environment, and construct scene state representation and robot state representation in this way; S3. Mind Map Generation and Update: When generating a mind map, the root node of the mind map is first obtained based on the structured semantic description. Then, the mind map template is used to generate a mind map containing a set of task nodes and a set of dependency edges by combining the scene state representation and the robot state representation. When the dynamic adjustment strategy is triggered and the mind map needs to be updated, the mind map is modified and updated accordingly based on the failure semantics in step S6. S4. Node Scheduling and Hint Compilation: Filter the candidate set of executable nodes based on the set of dependency edges in the mind graph, and select the current executable node as the execution object for this round from the candidate set of executable nodes; compile the structured semantic description, scene state representation and robot state representation of the current executable node into a multimodal hint injection package for hint injection in step S5; S5. VLA Reasoning and Execution: Inject the multimodal cue injection package into the Visual-Language-Motion Model (VLA), perform multimodal fusion reasoning and generate action plans, drive the robotic arm or mobile chassis to perform actions, obtain the scene state representation after execution, and form an execution record; S6. Execution Detection and Dynamic Adjustment: Based on the execution record, perform acceptance testing on the current executable node to determine whether the current executable node is completed; when the current executable node is not completed, identify the failure semantics and trigger the dynamic adjustment strategy, return to the mind map update in step S3, and perform dynamic adjustment of the mind map; when the current executable node is completed, further check whether the root node is completed. If the root node is completed, output a task completion prompt and end the process; otherwise, return to the mind map update in step S3, update the status of the current executable node to completed, and continue to execute the task in a closed loop.

2. The method according to claim 1, characterized in that, The structured semantic description includes one or more of the following: target semantics, target key objects, and target regions; the scene state representation includes at least one or more of the following: the set of target objects in the scene, the spatial location of the target objects, and the target state; the robot state representation includes at least one or more of the following: the robot arm pose, the robot arm end effector state, and the mobile chassis position.

3. The method according to claim 1, characterized in that, Mind maps are represented as directed graph structures. , ,in For a set of task nodes, For each task node, there is a set of dependent edges. Configure node attributes, which should include at least: node identifier, node semantics, node hierarchy depth, estimated execution cost, node status, and acceptance criteria; and a set of dependency edges. Used to describe the sequential dependencies between nodes.

4. The method according to claim 1, characterized in that, The current selection rules for executable nodes are as follows: If and only if node When all predecessor nodes except the root node are in a completed state, the node It belongs to the executable node and uses the priority function. Select the currently executable node , The expression is: ; in, The current executable node The overall evaluation priority value; The estimated execution cost for executable nodes, The depth of the executable node level. , The weights are preset and the constraints are met. ;choose The smallest executable node is selected as the current executable node. .

5. The method according to claim 1, characterized in that, The compilation prompts include: Current executable node Structured semantic description ; Current scene state representation ; Current robot state representation ; Create a multimodal cue injection package for use by VLA. ; It must contain at least one or more of the following: target semantics, target key objects, and target regions; It must include at least one or more of the following: a set of target objects, the spatial location of the target objects, and the target state; It includes at least one or more of the following: robotic arm pose, robotic arm end effector state, and mobile chassis position.

6. The method according to claim 1, characterized in that, Constructing the expected scene state representation using VLA's own Visual Language Model (VLM) And compare the expected scenario state representation with the scenario state representation after execution. Compare and define the scene state difference metric for the current executable node. Acceptance testing will be conducted. ; in, This represents the difference between the expected scene state representation and the actual scene state representation. The range of values ​​is When the expected scenario state representation perfectly matches the scenario state representation, ; Let the difference function represent the scene state. This represents the expected scenario state. This represents the scene state. for The total number of target objects in the target object set. This represents one of the target objects. Represents the target object The expected spatial location, Represents the target object Spatial location, This represents the Euclidean distance between the two. This indicates the total number of target objects whose expected state differs from the actual state; acceptance criteria. Provide a difference threshold It is dynamically generated by the Visual Language Model (VLM) based on the semantics of the currently executable node; when Determine the currently executable node at time The task is completed if successful; otherwise, it is considered a failure, and the Visual Language Model (VLM) analyzes the current scene state representation to generate failure semantics. And triggers a dynamic adjustment strategy; failure semantics This includes one or more of the following: target not found, grab failed, or placement failed.

7. The method according to claim 6, characterized in that, Dynamic adjustment strategies include: When failure semantics When a capture or placement fails, the mind map update in step S3 is triggered, a recovery child node is added to the current executable node and the recovery child node is executed first. The recovery child node is: re-execute the current executable node once. When failure semantics If the target is not found, the mind map update in step S3 is triggered, a recovery sub-node is added to the current executable node and the recovery sub-node is executed first. The recovery sub-node is: first locate the target in the scene, and then re-execute the current executable node.

8. A long-term task planning system for a home robot based on mind mapping, characterized in that, The task planning system is applied in home robot systems that include multimodal sensors, robotic arms, control terminals, and mobile chassis. The multimodal input unit is used to receive user commands and real-time scene and robot state information from multimodal sensors; The VLA large model perception unit uses the VLA large model to perceive the input user commands, scene state information, and robot state information, and obtain structured semantic descriptions, scene state representations, and robot state representations. The mind map unit is used to construct the mind map. First, the structured semantic description obtained by the VLA large model perception unit is used to generate the root node of the mind map. Then, the mind map containing the set of task nodes and the set of dependency edges is generated according to the mind map template, scene state representation and robot state representation. The mind map state management unit is used to maintain the structure and content of the mind map. The mind map node scheduling unit is used to select the currently executable node. If and only if node When all predecessor nodes except the root node are in a completed state, the node It belongs to the executable node, according to the priority function Choose the priority function The smallest executable node is selected as the current executable node. ; The estimated execution cost for executable nodes, The depth of the executable node level. , Preset weights; Node hint injection unit, used to inject the currently executable node The structured semantic description, scene state representation, and robot state representation in the VLA are fused into a multimodal cue injection package that can be received by VLA. The VLA large model inference unit receives multimodal prompt injection packets, performs multimodal inference, and outputs action plans. The motion generation unit uses motion schemes to drive the robot to perform actions. The action detection and graph correction unit is used to detect the execution result, and correct the mind map if the detection fails.

9. A non-volatile storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the method as described in any one of claims 1-8.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method described in any one of claims 1-8.