Training Method and System for Embodied Intelligence Task Planner Based on Multimodal Large Model
By adopting a combination method of multimodal large model and behavior tree in embodied intelligent task planning, the problem of lack of environmental feedback and insufficient utilization of multimodal information in the prior art is solved, and the executability and success rate of task planning are improved.
Patent Information
- Application Number
- CN202410250472.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-03-05
AI Technical Summary
The existing methods of embodied intelligent task planning lack feedback mechanisms with the environment, and the insufficient utilization of multimodal information leads to the inability of the planning to be reliable enough.
The embodied intelligent task planner training method based on multimodal large model is adopted, and the output of embodied planning data is concisely used to express the task planning problem in the behavior tree, and the feasibility verification is used to enhance the executability of the planning.
The planning success rate of the embodied intelligent task planner is improved, and the executability of the robot generated by the data fine-tuning and model understanding is reduced, and the planning complexity is improved, and the rationality and reliability of the planning is improved.
Smart Images

Figure CN118036750B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of embodied intelligence, and in particular to a training method and system for an embodied intelligence task planner based on a multimodal large model. Background Art
[0002] Embodied intelligence is an intelligent system based on a physical robot body, which realizes perception and action through interaction with the environment. This system obtains information, understands problems, makes decisions and executes corresponding actions through the interaction between the robot agent and the surrounding environment, thus demonstrating intelligent behavior and adaptability. Embodied intelligence task planning refers to converting the perceived information, including task instructions, target objects and the current scene, into appropriate decisions and formulating step-by-step plans. This process needs to consider the execution ability of the robot and the changes in the environment to ensure that the planned actions can be successfully executed during implementation.
[0003] Currently, common planning methods include using search algorithms and heuristic hierarchical task decoupling algorithms. However, both of these methods belong to chain planning, which is prone to falling into local optima and ignoring the global optimal path. Moreover, the above two planning methods lack a feedback mechanism with the environment and insufficient utilization of multimodal information. Therefore, enabling the robot to make better plans based on the current environmental situation and making the executability of embodied planning more reliable is still a huge challenge. Summary of the Invention
[0004] Embodiments of this application solve the technical problems in the prior art that embodied intelligence task planning lacks a feedback mechanism with the environment and insufficient utilization of multimodal information by providing a training method and system for an embodied intelligence task planner based on a multimodal large model, realizing fine-tuning on data in the field of embodied planning, enabling the model to generate the executability of a robot that can plan based on domain knowledge understanding, and improving the planning success rate of the embodied intelligence task planner.
[0005] Embodiments of this application provide a training method for an embodied intelligence task planner based on a multimodal large model, including the following steps: reformatting multiple original datasets in the embodied field, where the dataset format is an image-text-robot action trajectory pair; outputting concise embodied planning data to unify the granularity of the embodied planning data; formalizing the embodied task planning problem; and representing the process of continuously reasoning using reusable information as a search on a behavior tree.
[0006] Furthermore, the dataset format of the image-text-robot action trajectory pair respectively corresponds to the instruction description of an embodied intelligence task, all first-person view image frames during the robot's execution process, and a detailed description of each step of the robot's behavior.
[0007] Further, each node of the behavior tree is a state, which represents a partial solution with the input and the sequence of intermediate steps so far. This state is represented as s = [x, z 1...i , where the root node s = [x] represents the problem input, and each intermediate step z i is a single-step executable action expressed in text. If the z contained in the state s n is the last step required to complete the input x, then all intermediate steps are combined to obtain y = z 1 ,..., z n , as the result of the final task planning step.
[0008] Further, when generating the node state on the behavior tree, the step generator G(p θ , s, k) is used, where k is a hyperparameter representing the number of intermediate steps to be generated. In the current state s, k possible intermediate steps z (j) ~p θ = (z i+1 |s, z 1…i-1 ) = p θ (z i+1 |x, z 1…i )(j = 1...k) are obtained by sampling the probability distribution of the intermediate steps.
[0009] Further, when evaluating the node state on the behavior tree, the probability generated by the large model for the state s = [x, z 1...i is used as the score, that is and a threshold T is set. If the obtained state score score(s) >= T, it is considered that the plan given by this state can be realized, and the corresponding plan y = z 1 ,..., z n is output.
[0010] Further, the feasibility verification of the embodied task planning tree is enhanced by using the multi-modal large model itself. The steps for enhancing the feasibility verification include: setting the prompt for state evaluation as prompt v ; the set of states to be evaluated currently is S, which contains the state s = [x, z 1...i ∈ S; after inputting prompt v (S) and the visual observation O to the multi-modal large model, it will serve as a state evaluator V(p θ , S) to evaluate the score V(p θ , S)(s) of each state in S; the evaluator V(p θ , S) will score each state independently. Let Regarding its scoring method, the specific value v represents the score of state s; the states stored in set S will be continuously optimized according to the score. When the state score reaches the threshold, the requirements of input x are fulfilled.
[0011] Furthermore, the breadth-first search algorithm is used for continuous iterative search to obtain the state with the highest score.
[0012] This application also provides a training system for an embodied intelligent task planner based on a multimodal large model. The training system for the embodied intelligent task planner includes: a reformatting unit configured to reformat multiple original datasets in the embodied domain, and the dataset format is an image-text-robot action trajectory pair; a granularity unification unit configured to output concise embodied planning data to make the granularity of the embodied planning data unified; a formalization unit configured to formalize the embodied task planning problem; a behavior tree unit, and the behavior tree search unit is configured to represent the process of continuously reasoning by repeatedly using available information and perform a search on the behavior tree.
[0013] Furthermore, it also includes an enhanced unit for feasibility verification. The enhanced unit for feasibility verification includes: a preset sub-unit configured to set the prompt for state evaluation as prompt v ; a state set sub-unit configured to set the set of states to be evaluated currently as S, including state s = [x, z 1...i ∈ S; an evaluation sub-unit configured to input prompt v (S) and the visual observation O into the multimodal large model, and it will be used as a state evaluator V(p θ , S) to evaluate the scores V(p θ , S)(s) of each state in S; a scoring sub-unit configured to the evaluator V(p θ , S) will score each state independently. Let be its scoring method, then the specific value v represents the score of state s; an optimization sub-unit configured to the states stored in set S will be continuously optimized according to the score. When the state score reaches the threshold, the requirements of input x are fulfilled.
[0014] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0015] 1. Due to the adoption of a behavior-tree-based multi-modal large model embodied task planner, all plans are no longer generated at once. Instead, the planning process is modeled in a tree structure, and the task planning is decomposed into multiple levels. Each level is responsible for solving phased subtasks, thereby reducing the complexity of the planning and improving the rationality of the planning.
[0016] 2. And a multi-modal large model embodied task planning tree with enhanced feasibility verification is adopted, implementing a verification method before action execution. Combining the idea of the behavior tree, it attempts to explore the robot's assessment and analysis of the surrounding environment before executing an action, predict possible interaction results, and evaluate the risks of the action. Brief Description of the Drawings
[0017] Figure 1 It is a flowchart of the training method for the embodied intelligent task planner based on the multi-modal large model in the embodiments of the present application;
[0018] Figure 2 It is a schematic diagram of the behavior tree in the embodiments of the present application;
[0019] Figure 3 It is a flowchart of the steps for enhancing the feasibility verification in the embodiments of the present application;
[0020] Figure 4 It is a structural diagram of the training system for the embodied intelligent task planner based on the multi-modal large model in the embodiments of the present application;
[0021] Figure 5 It is a schematic structural diagram of the enhanced unit for feasibility verification in the embodiments of the present application. Detailed Embodiments
[0022] To make the basic method of the embodiments of the present application more obvious and understandable, the following will give a detailed description of the specific embodiments of the present application with reference to the drawings.
[0023] Embodiment 1
[0024] Figure 1 It is a training method for an embodied intelligent task planner based on a multi-modal large model in the embodiments of the present application, which will be described in detail through the following specific steps.
[0025] S11. Reformatting multiple original datasets in the embodied domain, and the dataset format is image-text-robot action trajectory pairs.
[0026] The ideal data generally required for embodied planning is: the input is an RGB image and a text task instruction, where the image represents the observation of the visual module of the embodied intelligent robot at the current moment, and the output is a text step plan, representing the text description corresponding to the current step.
[0027] In the specific implementation of this application, the reformatting of multiple original datasets in the embodied domain is taken as an example of the ALFRED dataset. The format of the ALFRED dataset can be set as an image-text-robot action trajectory pair. The dataset format of the image-text-robot action trajectory pair corresponds to the instruction description of an embodied intelligence task, all the first-person view image frames during the robot's execution process, and the detailed description of each step of the robot's behavior respectively.
[0028] Therefore, for the input part of the data, in the embodiments of this application, the instruction description is input as a text instruction, and the previous frame at the beginning of each step is input as an image, so that the model can understand that the planned text content to be output at this frame should be the content of this step.
[0029] S12. Output of concise embodied planning data to make the granularity of the embodied planning data unified.
[0030] For the output part of the data, although the ALFRED dataset has step descriptions in text form, its step descriptions are too detailed and are datasets used to assist robot navigation. For example, for an ideal step like "find an apple", the description in its dataset is "turn left, walk out of the bedroom, turn right, pick up the apple on the table". For the embodied planning model, its description is by no means so detailed. The task of the model is to give an executable and easy-to-understand step, hand it over to the lower-level execution module, and play the role of disassembling the task using semantic prior knowledge to generate single-step sub-goals.
[0031] Therefore, in the specific implementation of this application, for a text description with too fine a granularity, the text description and the corresponding task prompt can be handed over to GPT for processing to summarize the core purpose of this text description, and then the summarized concise steps are used as the output of the embodied planning data. For example, the detailed description of "turn left, walk out of the bedroom, turn right, pick up the apple on the table" is summarized as "go find an apple". In this way, the granularity of the obtained data output part is unified, meeting the specific requirements of the data meaning in the embodied planning field, thus completing the reformatting process.
[0032] S13. Formalize the embodied task planning problem.
[0033] In the specific implementation, formalize the embodied task planning problem: use p θ to represent a large model with pre-trained parameters θ. The lowercase letters x, y, z, s represent a natural language sequence, and the uppercase letter S represents a set of natural language sequences, that is, s ∈ S. The model p θThe obtained multimodal input x consists of a prompt message prompt(I) containing the input text instruction I (instruction) and the visual observation image O (observation) of the robot, and the output is a sequence of text planning steps y.
[0034] S14. Represent the process of continuously reasoning by reusing available information as a search on the behavior tree.
[0035] In a specific implementation, each node of the behavior tree is a state, which represents a partial solution with the input and the intermediate step sequence so far, and this state is represented as s = [x, z 1...i , where the root node is s = [x], representing the problem input, and each intermediate step z i is a single-step executable action expressed in text. If the z n contained in the state s is the last step required to complete the requirements of the input x, then all intermediate steps are combined to obtain y = z 1 ,..., z n , as the result of the final task planning steps. A schematic diagram of the behavior tree is as Figure 2 shown.
[0036] Among them, when generating the node state on the behavior tree, use the step generator G(p θ , s, k), where k is a hyperparameter representing the number of intermediate steps to be generated. Under the current state s, by sampling the probability distribution of intermediate steps, k possible intermediate steps z (j) ~ p θ = (z i+1 | s, z 1…i-1 ) = p θ (z i+1 | x, z 1…i )(j = 1... k).
[0037] And, when evaluating the node state on the behavior tree, use the probability generated by the large model to obtain the state s = [x, z 1...i as the score, that is And set a threshold T. If the obtained state score score(s) >= T, it is considered that the plan given by this state can be realized, and the corresponding plan y = z 1 ,..., z n is output.
[0038] Referring to Figure 3 , use the multimodal large model itself to enhance the feasibility verification of the embodied task planning tree. The steps of enhancing the feasibility verification include:
[0039] S21. Set the prompt for state evaluation as prompt v .
[0040] S22. The set of states to be evaluated currently is S, which contains the state s = [x, z 1...i ∈ S.
[0041] S23. After inputting prompt v (S) and the visual observation O into the multimodal large model, it will serve as a state evaluator V(p θ , S), and evaluate the score V(p θ , S)(s) of each state in S.
[0042] S24. The evaluator V(p θ , S) will score each state independently. Let be its scoring method, then the specific value v represents the score of state s.
[0043] S25. The states saved in the set S will be continuously optimized according to the scores. When the state score reaches the threshold, the requirement of the input x will be achieved.
[0044] Among them, based on the above evaluator, the breadth - first search algorithm can be used to continuously iterate and search to obtain the state with the highest score. Let the iteration limit step be T, and b of the most likely successful states be maintained at each step, then the size of the state set |S| = b.
[0045] In summary, due to the adoption of the multimodal large - model embodied task planner based on the behavior tree, all plans are not generated at once, but the planning process is modeled in a tree - like structure, and the task planning is decomposed into multiple levels. Each level is responsible for solving stage - specific subtasks, thereby reducing the complexity of the planning and improving the rationality of the planning.
[0046] Moreover, the multimodal large - model embodied task planning tree with enhanced feasibility verification is adopted, and a verification method based on before - action execution is implemented. Combining the idea of the behavior tree, it attempts to explore the robot's evaluation and analysis of the environment before executing an action, predict possible interaction results, and evaluate the risk of the action.
[0047] To enable those skilled in the art to better understand and implement the embodiments of the present application, the following is a corresponding introduction to a training system for an embodied intelligent task planner based on a multimodal large model. Figure 4 For an embodied intelligent task planner based on a multimodal large model.
[0048] Embodiment 2
[0049] Refer to Figure 4As shown, the embodiment of the present application provides a training system for an embodied intelligent task planner based on a multimodal large model. The embodied intelligent task planner training system includes:
[0050] A reformatter unit configured to reformat multiple original datasets in the embodied domain, and the dataset format is an image-text-robot action trajectory pair.
[0051] A granularity unification unit configured to output concise embodied planning data to make the granularity of the embodied planning data unified.
[0052] A formalization unit configured to formalize the embodied task planning problem.
[0053] A behavior tree unit, and the behavior tree search unit is configured to represent the process of continuously reasoning by reusing available information and search on the behavior tree.
[0054] In a specific implementation, as Figure 5 shown, it further includes an enhanced unit for feasibility verification, and the enhanced unit for feasibility verification includes:
[0055] A preset subunit configured to set the prompt for state evaluation as prompt v .
[0056] A state set subunit configured to set the state set to be evaluated currently as S, including the state s = [x, z 1...i ∈ S.
[0057] An evaluation subunit configured to input prompt v (S) and the visual observation O into the multimodal large model, and it will be used as a state evaluator V(p θ , S) to evaluate the score V(p θ , S)(s) of each state in S.
[0058] A scoring subunit configured to the evaluator V(p θ , S) will score each state independently. Let be its scoring method, and the specific value v represents the score of the state s.
[0059] An optimization subunit configured to the states saved in the set S will be continuously optimized according to the score. When the state score reaches the threshold, the requirement of the input x is achieved.
[0060] All the variations and specific implementations of the training method of the embodied intelligent task planner based on the multimodal large model in the foregoing Embodiment 1 are equally applicable to the training system of the embodied intelligent task planner based on the multimodal large model in this embodiment. Through the foregoing detailed description of the training method of the embodied intelligent task planner based on the multimodal large model, those skilled in the art can clearly know the training system of the embodied intelligent task planner based on the multimodal large model in this embodiment. Therefore, for the sake of brevity of the specification, it will not be elaborated herein.
[0061] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0062] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a system for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0063] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these changes and modifications.
Claims
1. A training method for an embodied intelligent task planner based on a multimodal large model, characterized in that: The following steps are involved: Reformatting of multiple original datasets in the embodied domain into image-text-robot motion trajectory pairs; The output of embodied planning data is concise, making the granularity of embodied planning data uniform; Formalize the embodied task planning problem; The process of repeatedly using available information to continuously reason is represented by searching on the behavior tree; Each node of the behavior tree is a state that represents a partial solution with the input so far and the sequence of intermediate steps. The state is represented as s = [x, z 1…i ], where the root node is s = [x], which represents the problem input, and each intermediate step z i is a textually expressed single-step executable action. If the state s contains z n To complete the last step required for the input x, all the intermediate steps are combined to obtain y = z1,…,z n , as the final mission planning step result; When generating the node state on the behavior tree, use the step generator G(p θ ,s,k), where p θ represents the parameters of the multimodal large model, where k is a hyperparameter representing the number of intermediate steps that need to be generated. Under the current state s, k possible intermediate steps are obtained by sampling the possibility distribution of the intermediate steps. Where j = 1.....k; When evaluating the node state on the behavior tree, the large model is used to generate the state s = [x, z 1…i ] as the score, that is Among them, p θ To represent a large model with parameters θ after pre-training, and set a threshold T. If the state score score(s)>=T, it is considered that the plan given by the state can be realized, and the corresponding plan y=z1,…,z is output. n ; The feasibility check of the embodied task behavior tree is enhanced by using the multimodal large model itself. The steps of the feasibility check enhancement include: Let prompt used for status evaluation be prompt v ; The current state set to be evaluated is S, which contains the state s = [x, z 1…i ]∈S; The prompt v (S) and visual observation O are input to the multimodal large model, which will serve as a state estimator V(p θ ,S), evaluate the score V(p) of each state in S θ ,S)(s); Evaluator V(p θ ,S) will score each state independently. is its scoring method, then the specific value v represents the score of state s; The states saved in set S will be continuously optimized according to the score. When the state score reaches the threshold, the input x requirement is met.
2. The method for training an embodied intelligent task planner based on a multimodal large model as claimed in claim 1, characterized in that: The dataset format of image-text-robot motion trajectory pairs corresponds to all first-person image frames during a robot's execution process, the instruction description of the embodied intelligence task, and a detailed description of each step of the robot's behavior.
3. The method for training an embodied intelligent task planner based on a multimodal large model as claimed in claim 1, characterized in that: The breadth-first search algorithm is used to iterate the search to obtain the state with the highest score.
4. An embodied intelligent task planner training system based on a multimodal large model, characterized in that: The embodied intelligent task planner training system comprises: A reformatting unit, wherein the reformatting unit is configured to reformat a plurality of original datasets of embodied domains, wherein the dataset formats are image-text-robot action trajectory pairs; a granularity unification unit, the granularity unification unit being configured to simplify the output of the embodied planning data so that the granularity of the embodied planning data is unified; a formalization unit configured to formalize an embodied task planning problem; A behavior tree unit, wherein the behavior tree search unit is configured to repeatedly use available information to continuously reason about a process representation and search on the behavior tree; Each node of the behavior tree in the behavior tree unit is a state, which represents a partial solution with the input and intermediate step sequence so far, and the state is represented by s = [x, z 1…i ], the root node is s = [x], which represents the problem input, and each intermediate step z i is a textually expressed single-step executable action. If the state s contains z n To complete the last step required for the input x, all the intermediate steps are combined to obtain y = z1,…,z n , as the final mission planning step result; When generating the node state on the behavior tree, use the step generator G(p θ ,s,k), where p θ represents the parameters of the multimodal large model, where k is a hyperparameter representing the number of intermediate steps that need to be generated. Under the current state s, k possible intermediate steps are obtained by sampling the possibility distribution of the intermediate steps. Where j = 1.....k; When evaluating the node state on the behavior tree, the large model is used to generate the state s = [x, z 1…i ] as the score, that is Among them, p θ To represent a large model with parameters θ after pre-training, and set a threshold T. If the state score score(s)>=T, it is considered that the plan given by the state can be realized, and the corresponding plan y=z1,…,z is output. n ; The embodied intelligent task planner training system further includes a feasibility verification enhancement unit, and the feasibility verification enhancement unit includes: A preset subunit is configured to set a prompt for status evaluation as prompt v ; The state set subunit is configured such that the state set currently to be evaluated is S, including the state s = [x, z 1…i ]∈S; The evaluation subunit is configured to prompt v (S) and visual observation O are input to the multimodal large model, which will serve as a state estimator V(p θ ,S), evaluate the score V(p) of each state in S θ ,S)(s); The scoring subunit is configured as an evaluator V(p θ ,S) will score each state independently. is its scoring method, then the specific value v represents the score of state s; The optimization subunit is configured such that the state stored in the set S will be continuously optimized according to the score. When the state score reaches a threshold, the requirement of inputting x is realized.
Citation Information
Patent Citations
Robot task planning method, system and device, storage medium and program product
CN116911552A
Visual language navigation technical scheme based on multi-modal perception model and large language model
CN117073701A
Cited By
Aeronautical manufacturing long-time-history action execution method based on large model task planning
CN122264979A