Multi-robot collaborative task planning method based on large language model

Through the multi-robot collaborative task planning method based on the large language model (LLM), the flexibility and efficiency problems of task planning of multi-robot systems in complex dynamic environments are solved, efficient task decomposition, skill matching and resource optimization are achieved, and the flexibility and adaptability of task execution are improved.

CN120816487APending Publication Date: 2025-10-21DONGHUA UNIV

Patent Information

Application Number
CN202511076493.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing multi-robot task planning methods lack flexibility and efficiency in complex and dynamically changing environments, making it difficult to optimize task allocation and coordination between robots in real time, affecting the system's execution efficiency and task completion quality.

Method used

A multi-robot collaborative task planning method based on a large language model (LLM) is adopted. Through four stages of task decomposition, skill matching, task allocation and task execution, LLM is used to parse high-level task instructions and generate accurate task plans, and task execution is monitored and adjusted in real time.

Benefits of technology

It significantly improves the flexibility and efficiency of task execution of multi-robot systems in dynamic environments, supports multi-robot collaboration to complete complex tasks, and is particularly suitable for task decomposition and resource optimization of heterogeneous robot teams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120816487A_ABST
    Figure CN120816487A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-robot cooperative task planning method based on a large language model, and the method comprises the steps: converting a high-level task instruction into a multi-robot executable task plan through a large language model LLM, and sequentially carrying out the task decomposition: firstly receiving the high-level task instruction through a natural language understanding module, the task is then decomposed into a plurality of sub-tasks. And skill matching: after task decomposition, the system selects appropriate robots to form a robot team according to task requirements and robot skills. And task allocation: then, the system reasonably allocates robots to each sub-task according to a task decomposition result through a task allocation processing flow. And the system executes the result through the task generation and instruction execution module. The problem of how to effectively plan and coordinate multiple robots to complete diversified tasks in a complex and dynamically changing environment is solved, accurate decomposition and efficient execution of complex tasks are achieved, and the task execution flexibility and efficiency can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of task planning for multi-robot systems, and in particular to a method for parsing and converting high-level task instructions into specific multi-robot execution plans using large language models (LLMs). Background Art

[0002] Currently, the world is in the midst of a technological revolution characterized by high levels of automation and intelligence. Multi-robot systems have been widely used in a variety of fields, including industrial production, the service industry, and operations in complex environments. In particular, how to efficiently plan and coordinate multiple robots to complete complex tasks in highly dynamic and unpredictable environments has become a critical issue. Existing task planning methods, such as traditional methods based on strict rules and procedures, generally perform well in static or predefined environments, but often fall short when faced with scenarios that require rapid response and real-time decision-making. The limitations of these methods are particularly evident when performing complex tasks such as navigation, handling, assembly, and monitoring, as they lack sufficient flexibility and efficiency to cope with dynamically changing environments.

[0003] With the development of large language models (LLMs), these models have shown great potential in understanding and processing natural language, providing new solutions for task planning in multi-robot systems. LLMs can convert complex natural language task instructions into specific operational steps that can be executed by robots, thereby significantly improving the flexibility of task planning and the adaptability of the system. However, how to effectively integrate LLMs into multi-robot systems and design a task planning framework that can fully utilize the capabilities of LLMs remains a technical challenge. These challenges include how to efficiently handle the parsing of high-level task instructions, optimize task decomposition, achieve effective coordination between robots, and make real-time adjustments during task execution.

[0004] This invention aims to address a key technical problem encountered in task planning for multi-robot systems: how to effectively plan and coordinate multiple robots to complete diverse tasks in complex and dynamically changing environments. Traditional task planning methods rely on pre-set static plans and programs. While these methods perform well in static or predictable environments, they often lack sufficient flexibility and efficiency when faced with constantly changing task requirements and environmental conditions. Existing methods struggle to optimize task allocation and inter-robot coordination in real time, particularly when performing complex tasks, such as those in industrial automation, disaster response, or everyday service robotics applications. This impacts system efficiency and the quality of task completion. Summary of the Invention

[0005] To address the challenge of effectively planning and coordinating multiple robots to complete diverse tasks in complex and dynamically changing environments, a multi-robot collaborative task planning method based on a large language model is proposed. This method introduces an innovative multi-robot task planning framework that utilizes LLMs to process and translate high-level task instructions, enabling precise decomposition and efficient execution of complex tasks. This framework is particularly well-suited for dynamic and complex operating environments, significantly improving the flexibility and efficiency of task execution and providing a novel solution for the application of multi-robot systems.

[0006] The technical solution of the present invention is:

[0007] A multi-robot collaborative task planning method based on a large language model (LLM) uses the large language model (LLM) to convert high-level task instructions into a multi-robot executable task plan, completing the four stages of task decomposition, skill matching, and task allocation and execution.

[0008] Step 1. Task decomposition: In this stage, the system receives a high-level natural language instruction I from the user. It is assumed that the task can be performed in a given environment E. The environment E contains multiple entities and objects. The states and interactions of the elements of these entities and objects will directly affect the execution of the task. The system first uses the large language model LLM to parse the task and decompose it into a series of subtasks T = {T1, T2, ..., T K}, K represents the number of subtasks, each of which contains clear execution steps and required skill sets The system uses a preset environment model E and the robot’s skill set Δ to determine the skills required for each subtask, taking into account the available objects in the environment and the robot’s skill limitations;

[0009] Step 2. Skill matching: After the task is decomposed, the system enters the skill matching stage. In this stage, the system's goal is to match the skill set required for each subtask. Select the most suitable robot or robot team to perform the corresponding subtask; if the skill set required for a subtask cannot be completed by a single robot, the system will combine multiple robots to complete the subtask according to the task requirements;

[0010] Step 3. Task allocation: In this stage, the system will allocate tasks according to the decomposed subtasks T={T1,T2,...,T KThe system then determines the skill matching results between the robot team and the robot team to make reasonable task allocations. Task allocation not only considers the execution order and parallelism of tasks, but also the robot's own capabilities and the complexity of the tasks. Based on the candidate robot collaboration combinations formed during the skill matching phase, the system determines whether a robot can complete a subtask independently or whether the task requires the joint completion of multiple robots. The system then assigns each robot its own subtask and plans the task sequence to ensure that tasks are executed in parallel as much as possible while meeting timing constraints.

[0011] Step 4. Task Execution: In this stage, the system processes the task execution plan generated in the task allocation stage and transmits the corresponding task instructions to each robot or robot group. The robot team then begins to execute their respective subtasks according to the plan. Each robot will execute its assigned subtask by invoking its low-level skills. These low-level skills enable the robot to move around in the environment, interact, or perform specific actions. The robot will provide real-time feedback on its status through the communication interface, and the system will dynamically monitor and adjust its strategy based on this feedback. If a subtask is detected to have failed or deviated from the expected path, the system will re-analyze the task execution status based on the current status and, with the help of the large language model (LLM), generate a new task adjustment plan to complete subtask reallocation or local action replanning to ensure the stable and efficient completion of the overall task.

[0012] Finally, the system summarizes and analyzes the task execution results, identifies the critical paths, bottlenecks, and resource conflicts that affect the success of the task, and provides feedback for subsequent task planning and system optimization.

[0013] Furthermore, during the task decomposition process, a natural language understanding module is built into the system. At this stage, the system receives and parses high-level task instructions through the natural language understanding module; this module uses LLM to convert instructions into task details that the machine can understand; through this module, LLM can identify key elements in task instructions and generate corresponding subtasks based on these elements.

[0014] Furthermore, during the task allocation process, the system integrates a task allocation processing flow based on semantic reasoning, which includes a task semantic parsing unit, a robot capability matching unit and a scheduling timing generation unit; this process calls the large language model LLM as the core of semantic understanding and reasoning, and models the scheduling relationship and allocates logical planning for tasks based on the subtask description, robot skill structure and matching results of the previous stage; among them, the task semantic parsing unit converts the subtask into a structured instruction representation, the robot capability matching unit retrieves the executable robot combination, and the scheduling timing generation unit generates a feasible task allocation table based on the inter-task dependency graph and robot availability; through this processing flow, the system can output a complete scheduling plan including "subtask→robot allocation mapping" and "execution timing table", which supports the smooth execution of tasks under resource constraints.

[0015] Furthermore, during the task execution process, a task generation and instruction execution module is built into the system, which is responsible for converting the results of the task decomposition and task allocation stages into executable action sequences; the system converts the semantic plan generated by calling the LLM into structured task instructions in the program generation module; this module supports compiling the task logic into a parsable program code form based on preset templates or rules, and sends it to the virtual or physical robot through the interface to drive it to execute the task process and complete the automated operation.

[0016] Furthermore, the task decomposition is implemented as follows:

[0017] Step 1.1: The system constructs structured prompts to guide the large language model. These prompts include several typical task examples and their subtask breakdown information. Through this structured input, the language model is guided to perform reasoning and generation according to a specific task decomposition method. The prompts include a natural language task description, the involved environmental objects, the target operations and their sequence, and the skill call elements.

[0018] Step 1.2: The system provides the large language model with object information and state descriptions of the current environment E, as well as the definition and constraints of the robot's skill set Δ. This information can be embedded in the prompt content in a parameterized manner to assist the model in considering the actual execution environment and robot capability limitations when decomposing tasks, ensuring that the generated subtasks are executable and compatible.

[0019] Step 1.3: The system organizes multiple known tasks and their subtask decomposition results to establish a mapping relationship between tasks, skills, and actions, and incorporates this into the prompt content in a structured form. This structured sample set also includes annotated descriptions of the task logic, which are used to characterize the operation objects, skill call sequence, and task intent. With the above sample input, the large language model can perform analogical reasoning based on existing mapping examples, automatically complete subtask decomposition when receiving new task instructions, and generate a subtask sequence and its corresponding skill set that conforms to the preset format.

[0020] Furthermore, the specific implementation of skill matching is as follows:

[0021] For the task set T={T1,T2,...,T K Each subtask T in k , where 1≤k≤K, analyze the skills required And each robot R n Skill Set S n ;

[0022] By combining robot skills with task requirements, the system determines the appropriate robot team. If a robot cannot complete a task independently, the system will select multiple robots to form a team based on the principle of complementary skills to ensure that the task can be completed efficiently.

[0023] The system calls the Large Language Model (LLM) to conduct a multi-dimensional evaluation of each candidate collaborative combination based on the semantic description of subtasks, the composition of collaborative combination members, and their skill distribution. The evaluation indicators include skill redundancy, collaborative efficiency, and task adaptability. The LLM infers and outputs a comprehensive score based on natural language prompt templates, thereby assisting the system in screening high-quality collaborative combination structures.

[0024] Furthermore, the task allocation is specifically implemented as follows:

[0025] The system decomposes the subtasks T={T1,T2,...,T K Combined with the established robot task matching strategy to determine the specific task allocation plan;

[0026] The task allocation process includes determining which subtasks need to be executed in parallel and which subtasks have dependencies and need to be executed sequentially;

[0027] During this process, the system uses LLM to perform semantic reasoning on the relationships between tasks and robot capabilities, assisting in generating reasonable scheduling logic;

[0028] Ultimately, the system outputs a structured task allocation result, indicating which robot or robot combination will complete each subtask, ensuring that the task is executed efficiently and accurately within time and resource constraints.

[0029] Furthermore, the task execution is specifically implemented as follows:

[0030] The system uses the task execution plan generated by LLM reasoning to send structured task instructions to the target robot or robot combination through the interface API;

[0031] During the task execution process, the system will monitor the progress of the task in real time and make dynamic adjustments based on the feedback from the task execution to ensure that the task is completed on time and efficiently;

[0032] If it is detected that the tasks are not executed in the predetermined order or manner, the system will make immediate adjustments, including dynamically reallocating unfinished subtasks, or modifying the order of tasks and calling alternative execution robot operation strategies.

[0033] Furthermore, it also includes system evaluation of execution effects, and the key indicators used include: task success rate SR, task completion rate TCR, target condition recall rate GCR, robot utilization rate RU and task executableness Exe; among them, the calculation of robot utilization rate RU is based on the "number of state transitions" during task execution, that is, the total number of operations in which the robot switches from one functional state to another in the behavioral process; the system compares the actual number of transitions with the theoretical transition path in the task plan to evaluate the robot's execution efficiency and the rationality of resource scheduling.

[0034] Preferably, a benchmark dataset for natural language task planning in multi-robot scenarios is constructed, and a series of simulation and real-robot experiments are conducted to establish experimental setup, evaluation criteria, and comparative analysis. This dataset, derived from AI2-THOR, includes 36 high-level instructions describing different tasks and their corresponding AI2-THOR floor plans, providing the spatial context required for task execution. The system records the execution of each task instance in real time, collecting key metrics including task completion status, robot utilization, number of execution interruptions, and task execution duration, thereby constructing a quantifiable performance evaluation system. By analyzing the evaluation results, the system further optimizes its internal natural language task planning model.

[0035] The beneficial effects of the present invention are:

[0036] This paper proposes a multi-robot task planning framework based on a large-scale language model, which effectively addresses the lack of flexibility and adaptability of existing task planning methods in complex and dynamic environments. Traditional planning methods based on rules and procedures often suffer from low execution efficiency when faced with uncertain environments and multi-robot coordination due to a lack of real-time responsiveness and task decomposition accuracy. The multi-stage task planning framework designed by this invention, comprising four phases: task decomposition, skill matching, task allocation, and task execution, can comprehensively improve the task execution efficiency of multi-robot systems.

[0037] This invention utilizes a large language model (LLM) as its core technology, parsing high-level task instructions through natural language and generating precise task plans, significantly improving the flexibility and applicability of task planning. By introducing a skill matching mechanism, it supports multi-robot collaboration to complete complex tasks, making it particularly suitable for task decomposition and resource optimization within heterogeneous robot teams. Real-time monitoring and adjustments during task execution further ensure accuracy and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a flow chart of the multi-robot collaborative task planning method based on a large language model of the present invention;

[0039] Figure 2 This is a schematic diagram of the task decomposition process of the present invention;

[0040] Figure 3 This is a schematic diagram of the robot skill matching of the present invention;

[0041] Figure 4 A schematic diagram of the task allocation process of the present invention;

[0042] Figure 5 It is the purpose of the task execution phase of the present invention. DETAILED DESCRIPTION

[0043] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0044] The present invention relates to an innovative framework for multi-robot collaborative task planning, namely intelligent multi-robot task planning based on large language models. The method leverages the powerful capabilities of large language models (LLMs) to convert high-level task instructions into task plans executable by multiple robots. By executing a series of key stages, including task decomposition, skill matching, and task allocation and execution, all processes are guided by programmed LLM prompts under a small number of example prompting paradigms. In order to validate the multi-robot task planning problem, the present invention provides a benchmark dataset covering four categories of high-level task instructions of different complexity. Evaluation experiments cover both simulation and actual scenarios, and the experimental results show that the proposed model can effectively generate multi-robot task plans and achieve remarkable results.

[0045] This paper proposes a multi-robot task planning method based on large language models (LLMs), aiming to optimize the efficiency of multi-robot systems in complex dynamic environments. Through rational task decomposition, robot skill matching, task allocation, and execution, this method effectively addresses key issues in multi-robot collaborative tasks. The method comprises four main phases: task decomposition, robot skill matching, task allocation, and task execution. The following describes the implementation steps and specific operations for each phase in detail.

[0046] For specific steps, see Figure 1 :Flowchart of the multi-robot collaborative task planning method based on a large language model. The multi-robot collaborative task planning method based on a large language model of the embodiment of the present application comprises the following steps:

[0047] Step 1. Task decomposition (such as Figure 2 As shown in Figure 2): Task decomposition is the first step in multi-robot task planning. At this stage, the system receives a high-level natural language instruction I from the user, which usually describes a high-level goal or task, assuming that the task can be performed in a given environment E. The environment E contains multiple entities and objects, and the states and interactions of these elements (entities and objects) directly affect the execution of the task. The system first uses a large language model (LLM) to parse the task and decompose it into a series of subtasks T = {T1, T2, ..., T K}, K represents the number of subtasks, each of which contains clear execution steps and required skill sets

[0048] For example, if the given instruction is "clean the desktop," the system will decompose the task into several subtasks, such as "clean the items on the desktop," "wipe the desktop surface," etc. During the task decomposition process, the system uses the preset environment model E and the robot's skill set Δ to determine the skills required for each subtask, while taking into account the available objects in the environment and the robot's skill limitations.

[0049] The system has a natural language understanding module built into it. During this phase, the system receives and interprets high-level task instructions through the module. This module uses the LLM to convert these instructions into machine-understandable task details. Through this module, the LLM can identify key elements in the task instructions, such as the operation object and the task sequence, and generate corresponding subtasks based on these elements. The specific implementation is as follows:

[0050] Step 1.1: The system constructs structured prompts to guide the large language model. These prompts include several typical task examples and their subtask breakdowns. This structured input guides the language model through reasoning and generation according to a specific task decomposition method. The prompts include a natural language task description, the environmental objects involved, the target operations and their sequence, and the skill calls. This provides a task representation capable of pattern induction, enabling effective subtask decomposition of new tasks.

[0051] In step 1.2, the system provides the large language model with object information and state descriptions of the current environment E, as well as the definition and constraints of the robot's skill set Δ. This information can be embedded in the prompt content in a parameterized manner to assist the model in factoring the actual execution environment and the robot's capability limitations when decomposing tasks, ensuring that the generated subtasks are executable and compatible.

[0052] In step 1.3, the system organizes multiple known tasks and their subtask decomposition results to establish a mapping relationship between tasks, skills, and actions, and incorporates this mapping relationship into the prompt content in a structured form. This structured sample set also includes annotated descriptions of the task logic, which characterize the operation objects, skill invocation sequence, and task intent. Using this sample input, the large language model can perform analogical reasoning based on existing mapping examples. When receiving new task instructions, it automatically completes the subtask decomposition and generates a subtask sequence and corresponding skill set that conforms to the pre-set format.

[0053] Step 2. Skill matching (such as Figure 3 After the task is decomposed, the system enters the skill matching stage. In this stage, the system aims to match the required skill sets for each subtask. The system selects the most suitable robot or robot team to perform the corresponding subtask. If a subtask requires a skill set that cannot be completed by a single robot, the system will combine multiple robots to complete the subtask based on the task requirements.

[0054] For example, if a robot is unable to independently complete a heavy lifting task due to insufficient skills (e.g., weight restrictions), the system will select multiple robots to jointly perform the task, with each robot taking on a specific subtask. This skill matching not only maximizes the robots' efficiency but also ensures successful completion of the task.

[0055] Specific implementation:

[0056] For the task set T={T1,T2,...,T K Each subtask T in k , (where 1≤k≤K) analysis required skills And each robot R n Skill Set Sn .

[0057] The system determines the appropriate robot team by combining the robot's skills with the task requirements. If a robot cannot complete the task independently, the system will select multiple robots to form a team based on the principle of complementary skills to ensure the task can be completed efficiently.

[0058] In a preferred implementation, the system utilizes a large language model (LLM) to perform a multi-dimensional evaluation of candidate collaboration combinations based on subtask semantic descriptions, the composition of collaboration combinations, and their skill distribution. Evaluation metrics include skill redundancy, collaborative efficiency, and task adaptability. The LLM infers and outputs a comprehensive score based on natural language prompt templates, assisting the system in selecting high-quality collaboration combinations.

[0059] Step 3. Task allocation (such as Figure 4 As shown in Figure 2): The task allocation phase is carried out after the skill matching is completed. In this phase, the system will assign tasks to the subtasks T = {T1, T2, ..., T K The robot team's skills match results and reasonable task allocation is carried out. Task allocation should not only consider the execution order and parallelism of tasks, but also the robot's own capabilities and the complexity of the task.

[0060] Based on the candidate robot collaboration combinations generated during the skill matching phase, the system determines whether a particular robot can independently complete a subtask or whether the task requires the collaboration of multiple robots. The system then assigns each robot its own subtask and plans the order of tasks, ensuring that tasks are executed as concurrently as possible while meeting timing constraints. This maximizes resource utilization and minimizes robot idleness and conflicts.

[0061] The system integrates a task allocation process based on semantic reasoning, comprising a task semantic parsing unit, a robot capability matching unit, and a scheduling sequence generation unit. This process utilizes a large language model (LLM) as the core of semantic understanding and reasoning. Based on subtask descriptions, the robot's skill structure, and the matching results from the previous phase, it models the scheduling relationships and performs logical planning for task allocation.

[0062] The task semantic parsing unit converts subtasks into structured instruction representations, the robot capability matching unit retrieves executable robot combinations, and the scheduling sequence generation unit generates a feasible task allocation table based on the inter-task dependency graph and robot availability. Through this process, the system can output a complete scheduling solution, including a "subtask → robot allocation mapping" and an "execution sequence table," to ensure smooth task execution within resource constraints.

[0063] Specific implementation:

[0064] The system decomposes the subtasks T={T1,T2,...,T K} Combined with the established robot task matching strategy, determine the specific task allocation plan.

[0065] The task allocation process includes determining which subtasks need to be executed in parallel and which subtasks have dependencies and need to be executed sequentially.

[0066] During this process, the system uses LLM to perform semantic reasoning on the relationships between tasks and robot capabilities, and assists in generating reasonable scheduling logic.

[0067] Ultimately, the system outputs a structured task allocation result, indicating which robot or robot combination will complete each subtask, ensuring that the task is executed efficiently and accurately within time and resource constraints.

[0068] Step 4. Task execution (such as Figure 5 Task execution is the final step in the system. In this phase, the system processes the task execution plan generated in the previous task assignment phase and transmits the corresponding task instructions to each robot or robot group. The robot team then begins executing its subtasks according to the plan. Each robot executes its assigned subtask by invoking its low-level skills (such as GoToLocation and ClickPicture). These low-level skills enable the robot to move around the environment, interact with it, or perform specific actions.

[0069] During task execution, the robot provides real-time feedback on its status (such as current location, task completion status, and abnormal conditions) through a communication interface. The system uses this feedback to dynamically monitor and adjust its strategies. If a subtask fails or deviates from the expected path, the system reanalyzes the task execution based on the current status and, using a large language model (LLM), generates a new task adjustment plan. This involves reassigning subtasks or replanning local actions to ensure stable and efficient completion of the overall task.

[0070] The system includes a task generation and instruction execution module, responsible for converting the results of the task decomposition and assignment phases into executable action sequences. The system invokes the semantic plan generated by the LLM and converts it into structured task instructions in the program generation module. This module compiles task logic into parseable program code, such as Python, based on pre-set templates or rules. This code is then distributed to virtual or physical robots through an interface, driving them to execute task flows and complete automated operations.

[0071] Specific implementation:

[0072] The system uses the task execution plan generated by LLM reasoning to send structured task instructions to the target robot or robot combination through the interface API.

[0073] During the task execution process, the system will monitor the progress of the task in real time and make dynamic adjustments based on the feedback from task execution to ensure that the task is completed on time and efficiently.

[0074] If it is detected that the tasks are not executed in the predetermined order or manner, the system will make immediate adjustments, including dynamically reallocating unfinished subtasks, or modifying the order of tasks, calling alternative execution robots and other operational strategies.

[0075] The key indicators used by the system to evaluate execution effects include: Success Rate (SR), Task Completion Rate (TCR), Goal Condition Recall (GCR), Robot Utilization (RU) and Task Executability (Exe).

[0076] Robot Utilization (RU) is calculated based on the number of "state transitions" during task execution. This refers to the total number of times the robot transitions from one functional state (such as navigation, recognition, or manipulation) to another during its behavioral flow. The system compares the actual number of transitions with the theoretical transition paths in the task plan to assess robot execution efficiency and resource scheduling rationality.

[0077] Finally, the system summarizes and analyzes the task execution results, identifies the critical paths, bottlenecks, and resource conflicts that affect the success of the task, and provides feedback for subsequent task planning and system optimization.

[0078] This paper presents a method for multi-robot task planning, particularly suitable for multi-robot collaborative execution of tasks in scenarios such as domestic environments. To validate the effectiveness and applicability of this method, we construct a benchmark dataset and conduct a series of simulation and real-robot experiments, providing a detailed experimental setup, evaluation criteria, and comparative analysis.

[0079] To evaluate the performance of our task planning methods and quantitatively compare them with other baseline methods, we created a benchmark dataset for natural language task planning in multi-robot scenarios. This dataset is derived from AI2-THOR, a deterministic simulation platform for simulating typical household activities. The dataset consists of 36 high-level instructions describing different tasks and their corresponding AI2-THOR floor plans, providing the spatial context required for task execution.

[0080] The system records the execution of each task instance in real time, collecting key metrics such as task completion status, robot utilization, execution interruption counts, and task execution duration, thereby building a quantifiable performance evaluation system. By analyzing the evaluation results, the system further optimizes its internal natural language task planning model.

[0081] The task planning model refers to the multi-stage task scheduling mechanism proposed in this paper, driven by a large language model (LLM). It consists of four phases: task decomposition, skill matching, task allocation, and task execution. During the optimization process, the system dynamically adjusts module parameters such as the granularity of task decomposition, skill matching strategy, robot combination method, and execution sequence based on actual task performance to adapt to different scenarios and improve overall task completion efficiency.

[0082] Through horizontal comparison on a unified evaluation dataset, the system can conduct a fair performance comparison with other benchmark methods, verifying the comprehensive performance advantages of the present invention in multi-robot task planning from multiple dimensions such as natural language comprehension ability, task completion quality, and robot resource utilization efficiency.

[0083] The tasks in the dataset are divided into four categories:

[0084] Elemental Tasks: Applicable to a single robot, assuming that the robot has the skills and capabilities required to perform all subtasks, so there is no need to coordinate with other robots.

[0085] Simple Tasks: These involve multiple objects and can be broken down into sequential or parallel subtasks, but not simultaneous execution. All robots possess the necessary skills to complete the task.

[0086] Compound Tasks: Similar to simple tasks, they offer greater flexibility, allowing tasks to be executed sequentially, in parallel, or in a hybrid fashion. Robots are heterogeneous, possessing different skills and attributes, allowing them to execute subtasks based on skill matching.

[0087] Complex Tasks: Designed for heterogeneous robot teams, these tasks typically need to be broken down into multiple subtasks and completed collaboratively by multiple robots. Unlike complex tasks, a single robot cannot independently perform certain subtasks and must rely on the collaborative efforts of the robot team to complete the task.

[0088] Specific implementation:

[0089] Task decomposition: In the experiment, the system first receives a high-level task instruction and then decomposes the task into multiple subtasks. Task decomposition can be performed sequentially or in parallel, depending on the nature of the task. Different task prompts were used in the experiment to simulate various task decomposition scenarios.

[0090] Skill Matching: After the task is broken down, the system selects appropriate robots to form a team based on the task requirements and robot skills. During the skill matching process, the robot team must be properly configured based on the task characteristics to ensure smooth task execution.

[0091] Task Allocation: The system then rationally assigns robots to each subtask based on the task decomposition results. The key to task allocation is how to efficiently allocate resources to complete tasks and minimize time wasted during task execution.

[0092] In all experiments, task complexity and the collaborative approach of the robot team significantly impacted task completion efficiency. Within a single task, the system efficiently decomposed the task and assigned appropriate subtasks to the robots. However, in some tasks, the accuracy of the task decomposition was limited, hindering the robots' performance.

[0093] The system demonstrated strong parallel execution capabilities for both simple and complex tasks, efficiently coordinating multiple robots to perform tasks simultaneously. However, certain complex tasks require coordinating the skills of heterogeneous robots, making coordination of robot teams more difficult.

[0094] For complex tasks, the system demonstrated a high success rate in task decomposition and team collaboration. In particular, in the selection and coordination of robot teams, the system was able to flexibly adjust execution strategies to ensure successful task completion.

[0095] Through a series of simulations and actual experiments, the method proposed in this invention has demonstrated its advantages in multi-robot task planning. This method can effectively perform task decomposition, skill matching, and team collaboration, and is particularly suitable for complex tasks involving heterogeneous robots. Experimental results show that this method has obvious advantages in improving task execution efficiency, optimizing robot utilization, and enhancing task adaptability. This method is not only applicable to simulation environments, but can also be deployed in actual robot systems, with strong flexibility and scalability. Because this method does not rely on modifications to specific tasks or environments, it is highly scalable and adaptable and can cope with more complex task requirements in the future.

[0096] The above-described embodiment merely represents one embodiment of the present invention. While the description is relatively specific and detailed, it should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. A multi-robot collaborative task planning method based on a large language model, characterized in that: The following steps are involved: Step 1. Task decomposition: In this stage, the system receives a high-level natural language instruction I from the user. It is assumed that the task can be performed in a given environment E. The environment E contains multiple entities and objects. The states and interactions of the elements of these entities and objects will directly affect the execution of the task. The system first uses the large language model LLM to parse the task and decompose it into a series of subtasks T = {T1, T2, ..., T K }, K represents the number of subtasks, each of which contains clear execution steps and required skill sets The system uses a preset environment model E and the robot’s skill set Δ to determine the skills required for each subtask, taking into account the available objects in the environment and the robot’s skill limitations; Step 2. Skill matching: After the task is decomposed, the system enters the skill matching stage; At this stage, the system aims to Select the most suitable robot or robot team to perform the corresponding subtask; if the skill set required for a subtask cannot be completed by a single robot, the system will combine multiple robots to complete the subtask according to the task requirements; Step 3. Task allocation: In this stage, the system will allocate tasks according to the decomposed subtasks T={T1,T2,...,T K The system then determines the skill matching results between the robot team and the robot team to make reasonable task allocations. Task allocation not only considers the execution order and parallelism of tasks, but also the robot's own capabilities and the complexity of the tasks. Based on the candidate robot collaboration combinations formed during the skill matching phase, the system determines whether a robot can complete a subtask independently or whether the task requires the joint completion of multiple robots. Subsequently, the system assigns each robot the subtask it is required to perform and plans the task sequence. Step 4. Task Execution: In this stage, the system processes the task execution plan generated in the task allocation stage and transmits the corresponding task instructions to each robot or robot group. The robot team then begins executing its own subtasks according to the plan. Each robot executes its assigned subtask by invoking its low-level skills. These low-level skills enable the robot to move around in the environment, interact, or perform specific actions. The robot provides real-time feedback on its status through the communication interface, and the system dynamically monitors and adjusts its strategy based on this feedback. If a subtask is detected to have failed or deviated from the expected path, the system will re-analyze the task execution status based on the current status and, with the help of the large language model (LLM), generate a new task adjustment plan to complete subtask reallocation or local action replanning. Finally, the system summarizes and analyzes the task execution results, identifies situations that affect the success of the task, and provides feedback for subsequent task planning and system optimization.

2. The multi-robot collaborative task planning method based on a large language model according to claim 1 is characterized in that: During the task decomposition process, a natural language understanding module is built into the system. At this stage, the system receives and parses high-level task instructions through the natural language understanding module. This module uses LLM to convert the instructions into task details that the machine can understand. Through this module, LLM can identify the key elements in task instructions and generate corresponding subtasks based on these elements.

3. The multi-robot collaborative task planning method based on a large language model according to claim 1 is characterized in that: During the task allocation process, the system integrates a task allocation processing flow based on semantic reasoning, which includes a task semantic parsing unit, a robot capability matching unit, and a scheduling sequence generation unit. This process uses the Large Language Model (LLM) as the core of semantic understanding and reasoning. Based on the subtask description, robot skill structure, and matching results from the previous stage, it models the scheduling relationship and performs logical allocation planning for tasks. The task semantic parsing unit converts subtasks into structured instruction representations, the robot capability matching unit retrieves executable robot combinations, and the scheduling sequence generation unit generates a feasible task allocation table based on the inter-task dependency graph and robot availability. Through this processing flow, the system can output a complete scheduling plan including "subtask→robot allocation mapping" and "execution sequence table", supporting the smooth execution of tasks under resource constraints.

4. The multi-robot collaborative task planning method based on a large language model according to claim 1 is characterized in that: During the task execution process, a task generation and instruction execution module is built into the system, which is responsible for converting the results of the task decomposition and task allocation stages into executable action sequences; the system calls the semantic plan generated by LLM and converts it into structured task instructions in the program generation module; this module supports compiling task logic into a parsable program code form based on preset templates or rules, and sends it to the virtual or physical robot through the interface to drive it to execute the task process and complete the automated operation.

5. The multi-robot collaborative task planning method based on a large language model according to claim 1 is characterized in that: The specific implementation of task decomposition is as follows: Step 1.1: The system constructs structured prompts to guide the large language model. These prompts include several typical task examples and their subtask breakdown information. Through this structured input, the language model is guided to perform reasoning and generation according to a specific task decomposition method. The prompts include a natural language task description, the involved environmental objects, the target operations and their sequence, and the skill call elements. Step 1.2: The system provides the large language model with object information and state descriptions of the current environment E, as well as the definition and constraints of the robot's skill set Δ. This information can be embedded in the prompt content in a parameterized manner to assist the model in considering the actual execution environment and robot capability limitations when decomposing tasks, ensuring that the generated subtasks are executable and compatible. Step 1.3: The system organizes multiple known tasks and their subtask decomposition results to establish a mapping relationship between tasks, skills, and actions, and incorporates this into the prompt content in a structured form. The structured sample set also includes annotated descriptions of the task logic, which are used to characterize the operation objects, skill calling sequence, and task intent. With the above sample input, the large language model can perform analogical reasoning based on existing mapping examples, automatically complete subtask decomposition when receiving new task instructions, and generate subtask sequences and their corresponding skill sets that conform to the preset format.

6. The multi-robot collaborative task planning method based on a large language model according to claim 1 is characterized in that: The specific implementation of skill matching is as follows: For the task set T={T1,T2,...,T K Each subtask T in k , where 1≤k≤K, analyze the skills required And each robot R n Skill Set S n ; By combining robot skills with task requirements, the system determines the appropriate robot team. If a robot cannot complete a task independently, the system will select multiple robots to form a team based on the principle of complementary skills to ensure that the task can be completed efficiently. The system calls the Large Language Model (LLM) to conduct a multi-dimensional evaluation of each candidate collaborative combination based on the semantic description of subtasks, the composition of collaborative combination members, and their skill distribution. The evaluation indicators include skill redundancy, collaborative efficiency, and task adaptability. The LLM infers and outputs a comprehensive score based on natural language prompt templates, thereby assisting the system in screening high-quality collaborative combination structures.

7. The multi-robot collaborative task planning method based on a large language model according to claim 1 is characterized in that: The specific implementation of task allocation is as follows: The system decomposes the subtasks T={T1,T2,...,T K Combined with the established robot task matching strategy to determine the specific task allocation plan; The task allocation process includes determining which subtasks need to be executed in parallel and which subtasks have dependencies and need to be executed sequentially; During this process, the system uses LLM to perform semantic reasoning on the relationships between tasks and robot capabilities, assisting in generating reasonable scheduling logic; Ultimately, the system outputs a structured task allocation result, indicating which robot or robot combination will complete each subtask, ensuring that the task is executed efficiently and accurately within time and resource constraints.

8. The multi-robot collaborative task planning method based on a large language model according to claim 1 is characterized in that: The specific implementation of the task is as follows: The system uses the task execution plan generated by LLM reasoning to send structured task instructions to the target robot or robot combination through the interface API; During the task execution process, the system will monitor the progress of the task in real time and make dynamic adjustments based on the feedback from the task execution to ensure that the task is completed on time and efficiently; If it is detected that the tasks are not executed in the predetermined order or manner, the system will make immediate adjustments, including dynamically reallocating unfinished subtasks, or modifying the order of tasks and calling alternative execution robot operation strategies.

9. The multi-robot collaborative task planning method based on a large language model according to claim 1, characterized in that: It also includes system evaluation of execution effects, and the key indicators used include: task success rate SR, task completion rate TCR, target condition recall rate GCR, robot utilization rate RU and task executable Exe; The calculation of robot utilization (RU) is based on the "number of state transitions" during task execution, that is, the total number of operations in which the robot switches from one functional state to another in the behavioral process. The system compares the actual number of transitions with the theoretical transition path in the task plan to evaluate the robot's execution efficiency and the rationality of resource scheduling.

10. The multi-robot collaborative task planning method based on a large language model according to claim 1, characterized in that: A benchmark dataset for natural language task planning in multi-robot scenarios was constructed, and a series of simulation and real-robot experiments were conducted to establish experimental setup, evaluation criteria, and comparative analysis. This dataset, derived from AI2-THOR, includes 36 high-level instructions describing different tasks and their corresponding AI2-THOR floor plans, providing the spatial context required for task execution. The system records the execution of each task instance in real time, collecting key metrics including task completion status, robot utilization, execution interruptions, and task execution duration, thereby constructing a quantifiable performance evaluation system. By analyzing the evaluation results, the system further optimizes its internal natural language task planning model.

Citation Information

Patent Citations

  • Multi-robot online task allocation and execution method and device and storage medium

    CN115284288A

  • Multi-agent cooperation method and system

    CN118917632A

  • Large model and knowledge graph linkage multi-robot collaborative assembly system and method

    CN119644933A

  • Complex task decomposition and dynamic optimization method and device based on large language model

    CN119883549A

  • Robot control system, robot control device, control method, and program

    JP2025107080A

Cited By

  • Task processing method based on intelligent agent and electronic equipment

    CN121501438A

  • An agent-based task processing method and electronic device

    CN121501438B

  • Robot task generation method and system and computer equipment

    CN121973215A

  • Task execution methods and systems

    CN122411533A