Double-arm robot long time sequence planning method based on tree-shaped hierarchical decomposition
The tree-structured hierarchical decomposition method for robot planning effectively addresses logic span and context redundancy issues, enhancing task success and adaptability in complex environments by breaking down tasks into sub-targets and using real-time feedback.
Patent Information
- Application Number
- CN202510579358.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art has problems such as large logic span, context redundancy and insufficient dynamic adaptability in robot long-term task planning, making it difficult to effectively deal with complex tasks.
A method based on tree-like hierarchical decomposition is adopted to decompose the task into a sub-target tree structure. Through hierarchical decomposition and closed-loop feedback mechanisms, fine-grained sub-targets are gradually generated, and planning strategies are adjusted in real time to adapt to environmental changes.
It significantly improves the success rate and accuracy of task planning, enhances the dynamic adaptability of robots in complex environments, and is able to handle tasks involving hidden objects and dynamic changes.
Smart Images

Figure CN120307292A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot task planning and execution, and particularly relates to a long-time sequence planning method for a dual-arm robot based on tree-like hierarchical decomposition. Background Art
[0002] In the prior art, researchers have tried various methods to solve the problem of long-time sequence task planning for robots, but there are certain limitations. The following are several typical existing solutions:
[0003] (1) Task Decomposition and Memory Management Method
[0004] To address the reasoning limitations of the LLM in long-time sequence tasks, some research has introduced task decomposition and memory management mechanisms. For example, FLTRNN[2] and LLM-State[3] improve the performance of the LLM in complex tasks by gradually decomposing tasks and managing memory during the reasoning process. However, these methods still rely on the direct reasoning of the LLM, and the performance will significantly decline as the task complexity increases.
[0005] (2) Tree Structure-based Planning Method
[0006] Inspired by the "Tree of Thought" (ToT) framework, some research has attempted to optimize the planning process by constructing a tree-like structure. For example, Tree-Planner[4] and RAP[5] improve the task success rate by sampling multiple plans and selecting the optimal path at each decision node. However, these methods still require the LLM to generate a complete plan in each inference, and do not fundamentally solve the problems of logical span and context redundancy.
[0007] (3) Feedback-based Reinforcement Learning Method
[0008] Some research has evaluated the plans generated by the LLM by introducing reinforcement learning. For example, Ahn et al. proposed "Do as I Can, Not as I Say"[6], which evaluates the accuracy of the plan through reinforcement learning. However, this method requires a large amount of data to train the network and has poor real-time performance in complex environments.
[0009] (4) Code Generation-based Planning Method
[0010] Some research has attempted to convert natural language tasks into code generation problems to utilize the code generation ability of the LLM. For example, Chaffin et al. proposed PPL-MCTS[7], which guides the LLM to generate a feasible plan by generating constraint text. However, this method limits the flexibility of planning and is difficult to adapt to complex tasks in open environments.
[0011] Through the analysis of the prior art, the following main drawbacks can be summarized:
[0012] 1. Logical span problem: There are significant difficulties in the prior art in transforming abstract natural language instructions into specific executable actions. For example, direct reasoning methods based on large language models (LLMs) (such as EmbodiedGPT) perform poorly in complex tasks because LLMs have difficulty handling the reasoning process from high-level instructions to primitive actions.
[0013] 2. Context redundancy problem: In long-term tasks, the task sequence contains a large amount of redundant information, making it difficult for LLMs to effectively focus on key information during the reasoning process. For example, in direct reasoning methods based on LLMs, redundant context information interferes with the attention mechanism of LLMs, reducing the task success rate. In addition, although existing task decomposition methods (such as FLTRNN and LLM-State) introduce memory management, they still cannot effectively reduce the impact of redundant context on reasoning.
[0014] 3. Lack of dynamic adaptability: The prior art lacks a real-time feedback mechanism for the environmental state and has difficulty dynamically adjusting the planning strategy to adapt to complex environments. For example, although tree-structure-based planning methods (such as Tree-Planner) optimize the path through sampling and selection and can perform dynamic switching between branches, their strategy planning still depends on a single inference of the LLM, and the dynamic adjustment range is difficult to cover areas outside the sampling, resulting in difficulty in adapting to dynamic scenarios in the real environment. Summary of the Invention
[0015] In view of this, the present invention provides a long-term planning method for a dual-arm robot based on tree-shaped hierarchical decomposition, which can ensure the successful execution of tasks, and at the same time, the entire planning process is coherent and dynamic.
[0016] The technical solution of the present invention is implemented as follows:
[0017] In a first aspect, a long-term planning method for a dual-arm robot based on tree-shaped hierarchical decomposition according to the present invention, the specific process is as follows:
[0018] Step 1: Initialize the input task instruction as the root node of the sub-goal tree, and decompose the root node using a sub-goal decomposition model to construct the first-level sub-goals;
[0019] Step 2: Further decompose the current sub-goal based on the current sub-goal, parent node information, and the current environment to generate new sub-goals;
[0020] Step 3: Take the finest-grained sub-goals as leaf nodes, evaluate each leaf node to determine whether it meets the termination condition, and further decompose it when it does not;
[0021] Step 4: According to the primitive action sequence generated by the leaf nodes, the robot executes specific actions, obtains environmental feedback during the execution process, and performs dynamic adjustment planning of the sub-goal tree;
[0022] Step 5: When the primitive actions corresponding to all leaf nodes are executed and the task objective is achieved, the task is completed.
[0023] Optionally, the leaf nodes of the present invention evaluate and determine whether they meet the termination conditions, and the termination conditions include:
[0024] Mappability: Whether the sub-goal can be directly mapped to a primitive action;
[0025] Consistency: Whether the sub-goal conforms to the current environment and the capabilities of the robot;
[0026] When both mappability and consistency are met, it is considered that the termination condition is satisfied.
[0027] Optionally, the present invention makes a termination decision based on the evaluation results:
[0028] If the leaf node meets the termination condition, it is mapped to a primitive action and executed.
[0029] If the mappability is not met but the consistency is met, the sub-goal continues to be decomposed.
[0030] If the consistency is not met, re-plan its parent node.
[0031] Optionally, if the robot fails to execute a certain action or the environment changes, the present invention returns to Step 2 to decompose the sub-goal again.
[0032] Optionally, the sub-goal decomposition model of the present invention is a large language model.
[0033] In a second aspect, a long-time sequence planning device for a dual-arm robot based on tree-like hierarchical decomposition according to the present invention includes:
[0034] A task initialization module, configured to initialize the input task instruction as the root node of the sub-goal tree, and decompose the root node by using the sub-goal decomposition model to construct the first-layer sub-goals;
[0035] A sub-goal decomposition module, based on the current sub-goal, parent node information, and the current environment, further decomposes the current sub-goal to generate new sub-goals;
[0036] A leaf node evaluation module, which records the finest-grained sub-goals as leaf nodes, evaluates each leaf node to determine whether it meets the termination conditions, and further decomposes it when it does not meet the conditions;
[0037] The action execution and feedback module is used to execute specific actions by the robot according to the primitive action sequence generated by the leaf nodes, obtain environmental feedback during the execution process, and perform dynamic adjustment and planning of the sub-goal tree;
[0038] The task completion judgment module is used to output the task completion result when the primitive actions corresponding to all leaf nodes are executed and the task goal is achieved.
[0039] Beneficial effects:
[0040] The core improvement of the present invention lies in gradually decomposing complex long-time tasks into finer-grained sub-goals through a hierarchical sub-goal tree structure, thereby reducing the logical span and context redundancy. Specifically:
[0041] First, reduce the logical span: By constructing a hierarchical sub-goal tree structure, complex high-level tasks are gradually decomposed into finer-grained sub-goals, thereby reducing the reasoning difficulty from abstract instructions to specific actions and improving the success rate of task planning.
[0042] Second, reduce context redundancy: Utilize the structural characteristics of the hierarchical sub-goal tree, taking the parent node as the input context instead of the entire task sequence, thereby reducing the interference of redundant information on reasoning and enhancing the reasoning ability of the LLM in long-time tasks.
[0043] Third, enhance dynamic adaptability: Through a closed-loop feedback mechanism, the executability of sub-goals is evaluated in real time, and the planning strategy is dynamically adjusted according to the environmental state to ensure the flexibility and reliability of task planning. Description of the drawings
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 It is the flowchart of the method of the present invention;
[0046] Figure 2 It is the schematic diagram of the sub-goal tree structure. Detailed implementation manners
[0047] The embodiments of the present invention will be described in detail below with reference to the drawings.
[0048] It should be noted that, without conflict, the following embodiments and the features in the embodiments may be combined with each other; moreover, based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.
[0049] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement an apparatus and / or practice a method. Additionally, this apparatus may be implemented and this method may be practiced using other structures and / or functionality in addition to one or more of the aspects described herein.
[0050] An embodiment of the present application provides a long-time sequence planning method for a dual-arm robot based on tree-shaped hierarchical decomposition. By constructing a sub-goal tree structure, complex tasks are gradually decomposed into executable sub-goals, and a closed-loop feedback mechanism is used to dynamically adjust the planning strategy. The following is a detailed implementation plan of this method, including a method flow chart and a detailed description of each step.
[0051] As Figure 1 shown, an embodiment of the present application provides a long-time sequence planning method for a dual-arm robot based on tree-shaped hierarchical decomposition, including key steps such as the construction of a sub-goal tree, sub-goal decomposition, leaf node evaluation, and task execution. There are clear time and logical relationships between each step, ensuring the coherence and dynamics of the entire planning process.
[0052] Step 1: Task initialization and sub-goal tree construction
[0053] Initialize the input task instruction as the root node of the sub-goal tree, and use the sub-goal decomposition model to decompose the task instruction to construct the sub-goal tree; the specific process of this step is as follows:
[0054] (1) Task initialization
[0055] Input a high-level task instruction (such as "clean the desktop"), and use it as the root node of the sub-goal tree (the structure is shown in the appendix Figure 2 ). Initialize the sub-goal tree structure and record the initial state and environmental information of the task.
[0056] (2) Sub-goal tree construction
[0057] Using a sub-goal decomposition model (based on LLM), decompose the high-level task into multiple sub-goals. For example, decompose "clean the desktop" into sub-goals such as "collect the clutter on the desktop" and "put it in the trash can". These sub-goals serve as the first-level nodes of the tree.
[0058] Step 2: Sub-goal Decomposition
[0059] Based on the current sub-goal, parent node information, and the current environment, further decompose the current sub-goal and evaluate the mapability and consistency of the decomposed sub-goals;
[0060] (1) Decomposition Model Input
[0061] The input of the sub-goal decomposition model includes the current sub-goal, parent node information, and the current environmental observation. For example, the current sub-goal is "collect the clutter on the desktop", the parent node is "clean the desktop", and the environmental observation includes the current state of the desktop.
[0062] (2) Sub-goal Generation
[0063] Based on the reasoning ability of the LLM, generate new sub-goals. For example, "collect the clutter on the desktop" can be further decomposed into "pick up the pen" and "pick up the eraser". These sub-goals serve as the next-level nodes of the tree.
[0064] Step 3: Leaf Node Termination Evaluation
[0065] (1) Leaf Node Evaluation
[0066] Evaluate each leaf node (i.e., the most fine-grained sub-goal) to determine whether it meets the termination conditions. The termination conditions include:
[0067] Mapability: Whether the sub-goal can be directly mapped to a primitive action (such as "pick up the pen").
[0068] Consistency: Whether the sub-goal conforms to the current environment and the robot's capabilities (such as whether the robot can pick up the pen).
[0069] (2) Termination Decision
[0070] If the leaf node meets the termination conditions, map it to a primitive action and execute it.
[0071] If it does not meet the mapability but meets the consistency, continue to decompose the sub-goal.
[0072] If it does not meet the consistency, re-plan the parent node.
[0073] Step 4: Task Execution and Feedback
[0074] (1) Action Execution
[0075] Based on the primitive action sequence generated from leaf nodes, the robot performs specific actions. For example, "pick up the pen" corresponds to the grasping action of the robot.
[0076] (2) Environmental feedback
[0077] During the execution process, environmental feedback is obtained in real time (such as whether the pen is successfully picked up), and subsequent plans are adjusted according to the feedback.
[0078] (3) Dynamic adjustment
[0079] If a certain action fails or the environment changes, re-evaluate the sub-goal tree and dynamically adjust the planning strategy.
[0080] Step 5: Task completion
[0081] (1) Task completion judgment
[0082] When the primitive actions corresponding to all leaf nodes are executed and the task goal is achieved, the task is completed.
[0083] (2) Result output
[0084] Output the task execution result, and record information such as the success rate and execution time of the task.
[0085] Compared with the prior art, the present invention has the following effects:
[0086] (1) Effectively reduce the logical span
[0087] In the prior art, when converting abstract high-level task instructions into specific executable actions, the logical reasoning process is complex and the success rate is low. For example, direct reasoning methods based on LLM (such as EmbodiedGPT) are difficult to handle the complex reasoning process from high-level instructions to primitive actions. In addition, although the tree structure-based planning method (such as Tree-Planner) optimizes the path through sampling and selection, it still relies on the single reasoning of LLM and fails to fundamentally solve the complexity of logical reasoning.
[0088] Improvements of the present application:
[0089] The present application constructs a hierarchical sub-goal tree structure, gradually decomposing complex high-level tasks into finer-grained sub-goals. Each sub-goal is evaluated and verified to ensure that it can be directly mapped to a primitive action. This hierarchical decomposition method significantly reduces the complexity of logical reasoning, enabling LLM to more efficiently handle the reasoning process from abstract instructions to specific actions. For example, the "clean the desktop" task is gradually decomposed into "collect the sundries on the desktop" and "put them into the trash can", and further decomposed into fine-grained sub-goals such as "pick up the pen" and "pick up the eraser", thus effectively reducing the logical span.
[0090] Technical effects:
[0091] By reducing the logical span, this application significantly improves the success rate of task planning, especially in complex tasks and long-time series tasks.
[0092] (2) Significantly reduce context redundancy
[0093] In the prior art, when dealing with long-time series tasks, a large amount of redundant information is included in the task sequence, making it difficult for the LLM to effectively focus on key information during the reasoning process. For example, in the direct reasoning method based on the LLM, redundant context information will interfere with the attention mechanism of the LLM, reducing the task success rate. In addition, although the existing task decomposition methods (such as FLTRNN and LLM-State) introduce memory management, they still cannot effectively reduce the impact of redundant context on reasoning.
[0094] Improvements of this application:
[0095] This application utilizes the structural characteristics of the hierarchical sub-goal tree, taking the parent node as the input context instead of the entire task sequence. This structural design reduces the interference of redundant information on reasoning, enabling the LLM to focus more on the decomposition and reasoning process of the current sub-goal. For example, when decomposing the task of "cleaning the desktop", only "collecting the sundries on the desktop" is used as the current context input instead of the entire task sequence, thus significantly reducing context redundancy.
[0096] Technical effects:
[0097] By reducing context redundancy, this application significantly improves the reasoning ability of the LLM in long-time series tasks, and improves the accuracy and success rate of task planning.
[0098] (3) Enhance dynamic adaptability
[0099] The prior art lacks a real-time feedback mechanism for the environmental state and is difficult to dynamically adjust the planning strategy to adapt to complex environments. For example, although the reinforcement learning-based evaluation methods (such as Do as I Can, Not as I Say) introduce a feedback mechanism, they require a large amount of data for training and have poor real-time performance in complex environments.
[0100] Improvements of this application:
[0101] This application introduces a closed-loop feedback mechanism. The leaf node termination model evaluates the executability of sub-goals in real time and dynamically adjusts the planning strategy according to the environmental state. For example, if a certain sub-goal is infeasible in the current environment (such as a robot being unable to pick up an object), the leaf node termination model will trigger replanning and adjust the sub-goal tree structure to adapt to environmental changes. In addition, this application also adjusts the task execution strategy in real time through environmental feedback to ensure the smooth progress of the task.
[0102] Technical effect:
[0103] By enhancing dynamic adaptability, this application significantly improves the success rate of a robot in performing tasks in complex environments, especially excelling in tasks involving hidden objects and dynamic changes.
[0104] (4) Improve adaptability to complex tasks
[0105] The prior art performs poorly in handling long-term tasks involving hidden objects or complex scenarios. For example, code-generation-based planning methods (such as PPLMCTS) improve the logic of planning but limit the flexibility of planning and are difficult to adapt to complex tasks in open environments.
[0106] Improvements of this application:
[0107] This application enables a robot to better handle long-term tasks involving hidden objects and complex scenarios through the dynamic construction and adjustment of a hierarchical sub-goal tree. For example, in the "clean the table" task, if an object is hidden in a drawer, the robot can complete the task by gradually decomposing sub-goals (such as "open the drawer" and "pick up the hidden object"). This dynamic adjustment ability significantly improves the adaptability of the robot to complex tasks.
[0108] Technical effect:
[0109] By improving the adaptability to complex tasks, this application significantly enhances the task execution ability of a robot in a real environment and can handle a wider range of task types.
[0110] In a second aspect, an embodiment of this application provides a long-term planning device for a dual-arm robot based on tree-like hierarchical decomposition, including:
[0111] A task initialization module, configured to initialize an input task instruction as the root node of a sub-goal tree, decompose the root node by using a sub-goal decomposition model, and construct the first-layer sub-goals;
[0112] A sub-goal decomposition module, configured to further decompose a current sub-goal based on the current sub-goal, parent node information, and the current environment to generate new sub-goals;
[0113] A leaf node evaluation module, configured to record the finest-grained sub-goals as leaf nodes, evaluate each leaf node to determine whether it meets the termination condition, and further decompose it when it does not;
[0114] An action execution and feedback module, configured to, according to a primitive action sequence generated by a leaf node, enable a robot to execute specific actions, obtain environmental feedback during the execution, and perform dynamic adjustment planning of the sub-goal tree;
[0115] The task completion judgment module is used to output the task completion result when the primitive actions corresponding to all leaf nodes are executed and the task objective is achieved.
[0116] With the settings of the sub-goal decomposition module and the action execution and feedback module in this device, it can effectively handle long-term tasks involving hidden objects and complex scenarios, and improve the task execution ability of the robot in the real environment.
[0117] The sub-goal decomposition module of this device effectively reduces the logical span and context redundancy by gradually decomposing high-level tasks into fine-grained sub-goals, and improves the success rate of long-term task planning.
[0118] This device uses a large language model (LLM) to dynamically decompose complex tasks into sub-goals, and generates new sub-goals based on the parent node information and environmental observations to ensure the executability and consistency of the sub-goals.
[0119] The leaf node evaluation module of this device dynamically determines whether to terminate the decomposition of the sub-goal tree or re-plan by evaluating the mapability and consistency of the sub-goals, ensuring that each leaf node can be directly mapped to a primitive action.
[0120] In summary, the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A long-term timing planning method for a dual-arm robot based on tree-like hierarchical decomposition, characterized in that The specific process is as follows: Step 1: Initialize the input task instruction as the root node of the sub-goal tree, and use the sub-goal decomposition model to decompose the root node to construct the first-level sub-goals; Step 2: Based on the current sub-goal, parent node information, and the current environment, further decompose the current sub-goal to generate new sub-goals; Step 3: Take the finest-grained sub-goals as leaf nodes, evaluate each leaf node to determine whether it meets the termination conditions, and further decompose it when it does not; Step 4: According to the primitive action sequence generated by the leaf nodes, the robot executes specific actions, obtains environmental feedback during the execution process, and performs dynamic adjustment planning of the sub-goal tree; Step 5: When the primitive actions corresponding to all leaf nodes are executed and the task goal is achieved, the task is completed.
2. The long-time sequence planning method for a dual-arm robot based on tree-like hierarchical decomposition according to claim 1, wherein The leaf nodes determine whether they meet the termination conditions, and the termination conditions include: Mapping ability: Whether the sub-goal can be directly mapped to a primitive action; Consistency: Whether the sub-goal conforms to the current environment and the capabilities of the robot; When both mapping ability and consistency are met, it is considered that the termination conditions are met.
3. The long-term timing planning method for a dual-arm robot based on tree-like hierarchical decomposition according to claim 2, wherein Make a termination decision based on the evaluation results: If the leaf node meets the termination conditions, map it to a primitive action and execute it; If it does not meet the mapping ability but meets the consistency, continue to decompose the sub-goal; If it does not meet the consistency, re-plan its parent node.
4. The method for long-time sequence planning of a two-armed robot based on tree-like hierarchical decomposition according to claim 3, wherein If the robot fails to execute a certain action or the environment changes, return to Step 2 to re-decompose the sub-goal.
5. The long-term timing planning method for a dual-arm robot based on tree-like hierarchical decomposition according to claim 1, wherein The sub-goal decomposition model is a large language model.
6. An apparatus for long-time sequence planning of a dual-arm robot based on tree-like hierarchical decomposition, characterized in that, It includes: A task initialization module, which is used to initialize the input task instruction as the root node of the sub-goal tree, and use the sub-goal decomposition model to decompose the root node to construct the first-level sub-goals; A sub-goal decomposition module, which further decomposes the current sub-goal based on the current sub-goal, parent node information, and the current environment to generate new sub-goals; A leaf node evaluation module, which records the finest-grained sub-goals as leaf nodes, evaluates each leaf node to determine whether it meets the termination conditions, and further decomposes it when it does not; An action execution and feedback module, which is used to make the robot execute specific actions according to the primitive action sequence generated by the leaf nodes, obtain environmental feedback during the execution process, and perform dynamic adjustment planning of the sub-goal tree; A task completion judgment module, which is used to output the task completion result when the primitive actions corresponding to all leaf nodes are executed and the task goal is achieved.
7. The long-time sequence planning device for a dual-arm robot based on tree-like hierarchical decomposition according to claim 6, wherein The leaf nodes determine whether they meet the termination conditions, and the termination conditions include: Mapping ability: Whether the sub-goal can be directly mapped to a primitive action; Consistency: Whether the sub-goal conforms to the current environment and the capabilities of the robot; When both mapping ability and consistency are met, it is considered that the termination conditions are met.
8. The long-term timing planning method for a dual-arm robot based on tree-like hierarchical decomposition according to claim 7, wherein Make a termination decision based on the evaluation results: If the leaf node meets the termination conditions, map it to a primitive action and execute it; If it does not meet the mapping ability but meets the consistency, continue to decompose the sub-goal; If it does not meet the consistency, re-plan its parent node.
9. The method for long-time sequence planning of a dual-arm robot based on tree-like hierarchical decomposition according to claim 8, wherein If the robot fails to execute a certain action or the environment changes, trigger the sub-goal decomposition module to re-decompose the sub-goal.
10. The long-time sequence planning method for a dual-arm robot based on tree-like hierarchical decomposition according to claim 6, wherein The sub-goal decomposition model is a large language model.
Citation Information
Patent Citations
Predictive controller, vehicle and method for controlling system
CN111670415A
Humanoid robot flexible assembly method and system based on thinking chain and medium
CN119238612A
Industrial robot path planning and execution method based on layered Monte Carlo tree search
CN119369408A
Portable fabric cutting device with handle length control
KR1020250173247A
Semi-optimal path finding in a wholly unknown environment
WO2001078951A1