A hybrid multi-robot collaboration method and system driven by a large language model
Through a hybrid architecture driven by a large language model, efficient collaboration of multiple robot systems under natural language task instructions is achieved. By centralizing task decomposition and distributive negotiation, the problem of low efficiency in multi-robot collaboration in existing technologies is solved, and the flexibility and efficiency of task understanding and execution are improved.
Patent Information
- Application Number
- CN202511037729.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-28
AI Technical Summary
Existing technologies struggle to achieve efficient collaboration among multiple robot systems under natural language task instructions, particularly lacking centralized semantic decomposition and decentralized negotiation capabilities, resulting in low collaboration efficiency in complex task scenarios.
A hybrid architecture driven by a large language model is adopted. Through centralized task decomposition and distributed task negotiation, the large language model for task decomposition is used for semantic parsing to generate a set of structured subtasks. Then, the robot ontology negotiation large language model is used for decentralized negotiation to achieve task allocation and collaborative execution.
It significantly improves the ability of multi-robot systems to understand natural language tasks and their execution flexibility, and enhances collaboration efficiency and task completion quality in complex task scenarios.
Smart Images

Figure CN120542462B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-robot task technology, specifically to a hybrid multi-robot collaboration method and system driven by a large language model. Background Technology
[0002] With the development of artificial intelligence and robotics, multi-robot systems are widely used in complex scenarios such as home services, warehousing and logistics, and disaster relief. Users expect to issue task instructions in natural language, which the multi-robot system will automatically understand and efficiently collaborate to complete the corresponding operations. This requires multi-robot systems to possess multiple intelligent capabilities, including natural language understanding, task decomposition, heterogeneous capability matching, and task collaboration.
[0003] Existing natural language-driven robot control methods primarily focus on executing single robot tasks. Typical approaches rely on Large Language Models (LLMs) to semantically parse user commands, generating operation sequences to drive execution modules such as perception, navigation, and grasping. These methods have achieved some success in semantic understanding and basic task execution. However, when faced with multi-robot application scenarios with complex task structures and collaborative relationships between subtasks, existing methods generally lack hierarchical task modeling and decomposition capabilities. This makes it difficult to support a deep understanding of complex tasks and the matching of heterogeneous capabilities, thus limiting the realization of efficient multi-robot task collaboration.
[0004] Meanwhile, traditional multi-robot collaborative systems typically allocate tasks based on fixed rule templates or centralized scheduling strategies. Task inputs are often in structured formats, and the collaborative process relies on manual modeling and rule systems, making it difficult to support direct access to natural language task instructions. Furthermore, they lack flexibility when facing task diversity or robot heterogeneity. These systems generally lack semantic-based decentralized task negotiation capabilities.
[0005] Therefore, existing technologies are still insufficient to achieve the ability to automatically complete the parsing and collaborative division of labor among multiple robots based on natural language task instructions. In particular, there is a lack of a general method that can perform both centralized semantic decomposition and support decentralized negotiation mechanisms to meet the high-efficiency execution requirements of multi-robot systems for complex natural language tasks. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a hybrid multi-robot collaboration method and system driven by a large language model. This system enables multi-robot systems to semantically decompose and collaboratively execute complex task instructions under natural language command guidance. It employs a hybrid architecture of centralized task decomposition and distributed task negotiation: the upper layer deploys a large language model for task decomposition, semantically parsing natural language task instructions to generate a structured set of subtasks containing skill requirements; the lower layer consists of multiple robots with heterogeneous capabilities, each integrating an ontology-negotiated large language model, achieving task allocation and collaborative execution through a decentralized language negotiation mechanism. This invention significantly enhances the understanding and execution flexibility of multi-robot systems for natural language tasks, improving their collaboration efficiency and task completion quality in complex task scenarios.
[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a hybrid multi-robot collaboration method driven by a large language model, employing a hybrid architecture of centralized task decomposition and distributed task negotiation, specifically including:
[0009] It receives natural language task instructions from users and performs semantic parsing through a task decomposition large language model to generate a task set containing multiple structured subtasks. Each subtask includes a natural language description and skill requirements.
[0010] Multiple robots, through their respective integrated ontology negotiation big language models, respond to and compete for each sub-task based on the perceived environmental state, their own skill space, and information about sub-tasks, and determine the task allocation result when a single robot acts as the execution subject.
[0011] When the skill requirements of a subtask exceed the skill space of any single robot, a collaborative task team building process based on a decentralized language negotiation mechanism is triggered. The ontology negotiation of the large language model of each robot participates in the process, and a multi-robot collaborative team is automatically built as the execution subject.
[0012] Each implementing entity performs the assigned sub-tasks based on the negotiation results.
[0013] In one embodiment, the multiple robots form a heterogeneous robot system, with different robots possessing different skill spaces; the executing entity is a single robot or a collaborative team of multiple robots, and the skill space of the executing entity satisfies the skill set required for the sub-task.
[0014] In one embodiment, each operational skill in each robot skill space is explicitly modeled through a preset skill description, which includes the type of executable operational skill and operational skill parameters; the skill space includes, but is not limited to, the following operational skills: recognition, navigation, grasping, and monitoring; the robot matches the skill description with the skill requirements of the sub-task.
[0015] In one embodiment, the semantic parsing through a task-decomposition large language model specifically includes:
[0016] The task decomposition large language model uses structured prompt word templates to perform semantic parsing; the structured prompt words include:
[0017] 1) Skill space information for all robots, used to describe the operational skills possessed by each robot;
[0018] 2) Current environment semantic information, used to provide information about the objects existing in the environment and their corresponding states;
[0019] 3) Several examples of task decomposition based on the few-shot learning paradigm and their corresponding standard response formats;
[0020] 4) The natural language task instructions to be executed.
[0021] In one embodiment, the use of a decentralized negotiation mechanism to respond to and compete for each subtask specifically includes:
[0022] Each robot, based on its integrated ontology negotiation language model, generates a structured response through comprehensive analysis of its own skill space, state perception information, and sub-task skill requirement information. The response includes:
[0023] 1) Assessment of the operable skills and their compatibility with sub-tasks;
[0024] 2) Estimate the cost required to complete the candidate subtasks, including but not limited to time and resources required;
[0025] For each subtask, collect all responses submitted by the robots, compare the execution costs of candidate robots that meet the skill requirements, and prioritize assigning the subtask to the robot with the lowest estimated cost.
[0026] In one embodiment, triggering the collaborative task team building process based on a decentralized language negotiation mechanism specifically includes:
[0027] Initiate a collaborative bidding process based on the skill space of a multi-robot system, and broadcast the set of operational skills required for the sub-task;
[0028] Each robot assesses the operational skills it can provide and the corresponding costs through an integrated ontology negotiation big language model, and submits a structured response.
[0029] Based on the principles of minimizing the completeness of skill coverage and the cost of collaborative execution, robot collaboration teams that meet the conditions are automatically assembled.
[0030] Secondly, this invention provides a hybrid multi-robot collaborative system driven by a large language model, employing a hybrid architecture of centralized task decomposition and distributed task negotiation, including:
[0031] The task decomposition module receives natural language task instructions input by the user and performs semantic parsing through the task decomposition large language model to generate a task set containing multiple structured subtasks. Each subtask includes a natural language description and skill requirements.
[0032] The multi-robot negotiation module involves multiple robots using their integrated ontology negotiation language models. Based on the perceived environmental state, their own skill space, and information about sub-tasks, they respond to and compete for each sub-task using a decentralized negotiation mechanism to determine the task allocation result when a single robot acts as the executor. When the skill requirements of a sub-task exceed the skill space of any single robot, a collaborative task team building process based on the decentralized language negotiation mechanism is triggered. The ontology negotiation language models of each robot participate collaboratively to automatically build a multi-robot collaborative team as the executor. Each executor executes the assigned sub-task according to the negotiation result.
[0033] In one embodiment, the semantic parsing through a task-decomposition large language model specifically includes:
[0034] The task decomposition large language model uses structured prompt word templates to perform semantic parsing; the structured prompt words include:
[0035] 1) Skill space information for all robots, used to describe the operational skills possessed by each robot;
[0036] 2) Current environment semantic information, used to provide information about the objects existing in the environment and their corresponding states;
[0037] 3) Several examples of task decomposition based on the few-shot learning paradigm and their corresponding standard response formats;
[0038] 4) The natural language task instructions to be executed.
[0039] In one embodiment, the use of a decentralized negotiation mechanism to respond to and compete for each subtask specifically includes:
[0040] Each robot, based on its integrated ontology negotiation language model, generates a structured response through comprehensive analysis of its own skill space, state perception information, and sub-task skill requirement information. The response includes:
[0041] 1) Assessment of the operable skills and their compatibility with sub-tasks;
[0042] 2) Estimate the cost required to complete the candidate subtasks;
[0043] For each subtask, collect all responses submitted by the robots, compare the execution costs of candidate robots that meet the skill requirements, and prioritize assigning the subtask to the robot with the lowest estimated cost.
[0044] In one embodiment, triggering the collaborative task team building process based on a decentralized language negotiation mechanism specifically includes:
[0045] Initiate a collaborative bidding process based on the skill space of a multi-robot system, and broadcast the set of operational skills required for the sub-task;
[0046] Each robot assesses the operational skills it can provide and the corresponding costs through an integrated ontology negotiation big language model, and submits a structured response.
[0047] Based on the principles of minimizing the completeness of skill coverage and the cost of collaborative execution, robot collaboration teams that meet the conditions are automatically assembled.
[0048] The system and method in this invention correspond to each other; the specific technical solutions applicable to the method are also applicable to the system.
[0049] Compared with the prior art, the beneficial technical effects of the present invention are:
[0050] 1. This invention proposes a hybrid multi-robot task collaboration method based on a large language model, enabling users to directly drive multi-robot collaborative execution through natural language commands, significantly reducing the threshold for human-machine interaction.
[0051] 2. This invention proposes a structured prompt word template for task decomposition, which realizes the reasonable decomposition of complex natural language instructions into structured subtasks, and significantly improves the accuracy and detail of task understanding.
[0052] 3. This invention proposes a distributed multi-robot task negotiation mechanism, which utilizes the robot's own integrated ontology negotiation large language model to autonomously complete task allocation, thereby enhancing the system's flexibility and robustness. Attached Figure Description
[0053] Figure 1This is a flowchart illustrating the hybrid multi-robot collaboration method driven by a large language model in an embodiment of the present invention.
[0054] Figure 2 This is a schematic diagram of the robot negotiation response and task allocation process in an embodiment of the present invention.
[0055] Figure 3 This is a schematic diagram illustrating the process of forming a robot collaborative team in an embodiment of the present invention.
[0056] Figure 4 This is a schematic diagram of a system in an embodiment of the present invention. Detailed Implementation
[0057] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.
[0058] like Figure 1 As shown, this invention provides a hybrid multi-robot collaboration method driven by a large language model, employing a hybrid architecture of centralized task decomposition and distributed task negotiation, specifically including the following steps:
[0059] S1 receives natural language task instructions input by the user and performs semantic parsing through a task decomposition large language model to generate a task set containing multiple structured subtasks. Each subtask includes a natural language description and skill requirements.
[0060] S2, multiple robots, through their respective integrated ontology negotiation big language models, respond to and compete for each sub-task based on the perceived environmental state, their own skill space and sub-task information, and determine the task allocation result when a single robot acts as the execution subject.
[0061] S3. When the skill requirements of a subtask exceed the skill space of any single robot, a collaborative task team building process based on a decentralized language negotiation mechanism is triggered. The ontology negotiation of the large language model of each robot participates in the process, and a multi-robot collaborative team is automatically built as the execution subject.
[0062] S4, each implementing entity performs the assigned sub-tasks according to the negotiation results.
[0063] In one embodiment, step S1, which involves receiving natural language task instructions input by the user and performing semantic parsing using a task decomposition large language model to generate a task set containing multiple structured subtasks, specifically includes:
[0064] Users input natural language task instructions (such as "Please turn on the TV and put the drinks on the table into the refrigerator") via voice or text. The Task Parser LLM model performs structured semantic parsing on the input natural language task instructions.
[0065] To improve the accuracy of semantic understanding and structure generation, the task decomposition large language model introduces structured cue word templates as auxiliary information during the reasoning process, which mainly includes the following four types of elements:
[0066] 1) Current environmental semantic information: Provide natural language descriptions through perception information, such as "The TV is located in the living room and is turned off, the beverage is on the table in the living room, and the refrigerator is located on the right side of the kitchen and is turned off".
[0067] 2) All robot skill space information: obtained through the registration of skill capabilities, described in a structured form. Robot-A, Robot-B, and Robot-C are the identifiers (IDs) of the three robots. walk, find, switchon, switchoff, grab, putin, open, and close represent basic operation skills such as "move," "find," "turn on appliances," "turn off appliances," "grab," "place," "open," and "close," respectively. <obj>This refers to the specific object being operated on (such as "television", "beverage", "refrigerator" etc.).
[0068] The robot's skill space information in JSON format is as follows: [
[0070] {
[0071] "ID": "Robot-A",
[0072] "Skill Space": ["walk"] <obj>", "find <obj>", "switchon <obj>", "switchoff <obj>"]
[0073] },
[0074] {
[0075] "ID": "Robot-B",
[0076] "Skill Space": ["walk"] <obj>", "find <obj>", "grab <obj>", "putin <obj> <obj>"]
[0077] },
[0078] {
[0079] "ID": "Robot-C",
[0080] "Skill Space": ["walk"] <obj>", "find <obj>","switchon <obj>", "switchoff <obj>", "open <obj>", "close <obj>"]
[0081] }
[0082] ].
[0083] 3) Example of task decomposition based on the few-shot learning paradigm: that is, the known task instructions and their standard subtask decomposition structure pairs are used to prompt the model to learn the decomposition format.
[0084] 4) The natural language task instruction to be executed: such as "Please put the drinks on the table into the refrigerator and turn on the TV".
[0085] With the support of the above task context, the upper-level Task Parser LLM can output a structured collection of subtasks in JSON format: [
[0087] {
[0088] "ID": "Subtask 1",
[0089] Task Description: "Turn on the TV in the living room".
[0090] "Skill Requirements": ["walk"] <obj>", "find <obj>", "switchon <obj>"]
[0091] },
[0092] {
[0093] "ID": "Subtask 2",
[0094] Task Description: "Put the drinks on the living room table into the kitchen refrigerator."
[0095] "Skill Requirements": ["walk"] <obj>", "find <obj>", "grab <obj>", "putin <obj> <obj>", "open <obj>", "close <obj>"]
[0096] }
[0097] ] .
[0098] This set of structured subtasks will serve as the basis for subsequent robot negotiation and assignment.
[0099] In one embodiment, such as Figure 2 As shown, step S2 specifically includes:
[0100] After multiple robots receive a set of subtasks from the upper layer, each robot generates response information for each subtask using its locally deployed Ontology Negotiation Large Language Model (Agent LLM).
[0101] For each received subtask, the negotiation language model generates the following structured response:
[0102] 1) Assess the suitability of the subtask (based on its skill space and current state). If the suitability is low, return "unsuitable" with an explanation of the reason.
[0103] 2) Estimate the overall cost of performing the task to support subsequent task allocation decisions.
[0104] The response format of the ontology negotiation large language model is shown below.
[0105] When the robot possesses the corresponding execution capabilities, the response of the ontology negotiation language model in JSON format is as follows:
[0106] {
[0107] "Robot ID": "Robot-A",
[0108] Subtask ID: Subtask 1
[0109] Task Description: "Turn on the TV in the living room".
[0110] "Compatibility": "Suitable"
[0111] Reason for suitability: Possesses relevant skills, and is currently available for scheduling.
[0112] Cost estimate: Medium
[0113] }
[0114] When the robot lacks the corresponding execution capabilities, the response of the ontology negotiation language model in JSON format is as follows:
[0115] {
[0116] "Robot ID": "Robot-A",
[0117] Subtask ID: Subtask 2
[0118] Task Description: "Put the drinks on the living room table into the kitchen refrigerator."
[0119] "Compatibility": "Not suitable"
[0120] Reason for compatibility: "Lack of 'grab', 'putin', 'open', and 'close' skills".
[0121] "Cost estimate": null
[0122] }
[0123] Collect the response content of the ontology negotiation large language model of all robots, and complete the task assignment based on the following strategy:
[0124] 1) If only one robot is available for matching, assign the robot directly;
[0125] 2) If multiple robots meet the skill requirements, select the one with the lowest cost to perform the task;
[0126] 3) If no robot can meet the requirements independently, a collaborative team should be formed.
[0127] Based on the above strategy, if a subtask can be completed by a single robot, the executor of the subtask can be determined. For example, for "subtask 1", its executor can be identified as Robot-A.
[0128] In one embodiment, such as Figure 3 As shown, step S3 specifically includes:
[0129] If no single robot can complete a sub-task independently, a collaboration mechanism will be automatically triggered to organize multiple robots to work together.
[0130] For example, for "Subtask 2: Put the drinks on the living room table into the kitchen refrigerator", the skill requirements include multiple operational skills such as grab, putin, open, and close. If no single robot possesses all the required skills, such as Robot-B only being able to grab and put, and Robot-C being able to open and close the refrigerator, then this subtask will be marked as "requires collaborative execution".
[0131] In collaborative mode, the system broadcasts the skill requirements of the subtask to all robots that have not been assigned subtasks, inviting them to submit collaborative responses based on their own skills and current status. Each robot, using an ontology negotiation large language model (Agent LLM), generates a structured collaborative intention in JSON format.
[0132] {
[0133] "Robot ID": "Robot-B",
[0134] Subtask ID: Subtask 2
[0135] "Skills to be executed": ["grab"] <obj>", "putin <obj> <obj>"],
[0136] "Compatibility": "Suitable"
[0137] "Willingness to cooperate": "Participation"
[0138] Cost estimate: "Low"
[0139] }
[0140] After summarizing all robot response information, the collaborative execution entities for sub-tasks are formed based on factors such as skill matching completeness and collaborative execution cost, and the specific robot skill division is clarified. For example, for "sub-task 2", it is determined that it will be completed jointly by [Robot-B and Robot-C].
[0141] In one embodiment, step S4 specifically includes:
[0142] Each executor (including independent robots or collaborative teams of multiple robots) performs its assigned sub-tasks based on the negotiation and task allocation results in steps S2 and S3, in order to achieve a complete response to natural language instructions.
[0143] For example, for the task instruction "Please turn on the TV and put the beverage on the table into the refrigerator," the system breaks it down into two sub-tasks: Sub-task 1, "Turn on the TV in the living room," is executed by Robot-A, and Sub-task 2, "Put the beverage on the living room table into the refrigerator in the kitchen," is completed collaboratively by the team [Robot-B, Robot-C]. By having each executor autonomously complete its respective sub-task based on its own skills, the complete execution of the original task instruction can be achieved.
[0144] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or stages, which are not necessarily completed at the same time, but may be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0145] Based on the description of the above method embodiments, the present invention also provides a system. The system may be a system that uses software (applications), modules, components, servers, clients, etc., using the methods described in the embodiments of this specification, combined with necessary implementation hardware. Based on the same innovative concept, the systems in one or more embodiments provided in this disclosure are as described in the following embodiments. Since the implementation schemes and methods for solving the system problem are similar, the specific system implementations in the embodiments of this specification can refer to the implementations of the foregoing methods, and repeated details will not be repeated. As used below, the term "module" or "module group" refers to a combination of software and / or hardware capable of implementing a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.
[0146] like Figure 4 As shown, the present invention proposes a hybrid multi-robot task collaboration system driven by a large language model, comprising:
[0147] The task decomposition module receives natural language task instructions input by the user and performs semantic parsing through the task decomposition large language model to generate a task set containing multiple structured subtasks. Each subtask includes a natural language description and skill requirements.
[0148] The multi-robot negotiation module involves multiple robots using their integrated ontology negotiation language models. Based on the perceived environmental state, their own skill space, and information about sub-tasks, they respond to and compete for each sub-task using a decentralized negotiation mechanism to determine the task allocation result when a single robot acts as the executor. When the skill requirements of a sub-task exceed the skill space of any single robot, a collaborative task team building process based on the decentralized language negotiation mechanism is triggered. The ontology negotiation language models of each robot participate collaboratively to automatically build a multi-robot collaborative team as the executor. Each executor executes the assigned sub-task according to the negotiation result.
[0149] Among them, the Task Parser LLM has powerful natural language understanding and task structure generation capabilities. It can receive natural language task instructions input by users and transform them into a set of structured subtasks by combining environmental semantics and robot capability context information.
[0150] The locally integrated ontology negotiation big language model (Agent LLM) can obtain the ontology's skill space and real-time status information, has the ability to judge and negotiate task adaptability, can generate response opinions for tasks, and can negotiate and divide tasks with other robots.
[0151] Preferably, the system of the present invention may further include:
[0152] The subtask management module is used to aggregate all subtask response information and make subtask allocation decisions based on set strategies (such as priority, cost, and load).
[0153] The perception information sharing module is responsible for collecting environmental information from various robot perception devices (such as cameras, LiDAR, and semantic mapping modules), extracting key semantic features (such as "drink location" and "target area occupancy status"), and transmitting them to the task decomposition module to assist the large language model in completing semantically accurate task parsing and reasoning.
[0154] The skill and capability registration module is used to manage the skill space of all robots in the system in a unified manner, and to provide standardized and structured skill and capability descriptions to the upper-level task decomposition module.
[0155] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0156] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0157] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.< / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj> < / obj>
Claims
1. A hybrid multi-robot collaborative method driven by a large language model, characterized in that, A hybrid architecture combining centralized task decomposition and distributed task negotiation is adopted, specifically including: It receives natural language task instructions from users and performs semantic parsing through a task decomposition large language model to generate a task set containing multiple structured subtasks. Each subtask includes a natural language description and skill requirements. Multiple robots, through their respective integrated ontology negotiation language models, respond to and compete for each subtask based on perceived environmental states, their own skill spaces, and subtask information, employing a decentralized negotiation mechanism. Specifically, each robot, based on its integrated ontology negotiation language model, generates a structured response through comprehensive analysis of its own skill space, state perception information, and subtask skill requirements. This response includes: 1) an assessment of executable operational skills and their compatibility with the subtask; and 2) an estimated cost to complete the candidate subtask. For each subtask, all robot responses are collected, and the execution costs of candidate robots meeting the skill requirements are compared. The subtask is preferentially assigned to the robot with the lowest estimated cost. The task allocation result is then determined when a single robot acts as the execution entity. When the skill requirements of a subtask exceed the skill space of any single robot, a collaborative task team building process based on a decentralized language negotiation mechanism is triggered. This process includes: initiating a collaborative bidding process based on the skill space of a multi-robot system, broadcasting the set of operational skills required for the subtask; each robot evaluating the operational skills it can provide and their corresponding costs through an integrated ontology negotiation language model, and submitting a structured response; automatically combining robot collaboration teams that meet the conditions based on the principle of minimizing skill coverage completeness and collaborative execution costs; and automatically building a multi-robot collaboration team as the execution entity through the collaborative participation of the ontology negotiation language models of each robot. Each implementing entity performs the assigned sub-tasks based on the negotiation results.
2. The hybrid multi-robot collaborative method driven by a large language model according to claim 1, characterized in that, The multiple robots form a heterogeneous robot system, with different robots possessing different skill sets; the executing entity is a single robot or a collaborative team of multiple robots, and the skill set of the executing entity satisfies the skill set required for the sub-task.
3. The hybrid multi-robot collaborative method driven by a large language model according to claim 1, characterized in that, Each operational skill in the robot's skill space is explicitly modeled through a preset skill description, which includes the type of executable operational skill and the operational skill parameters. The skill space includes the following operational skills: recognition, navigation, grasping, and monitoring; The robot matches skill descriptions with the skill requirements of sub-tasks.
4. The hybrid multi-robot collaborative method driven by a large language model according to claim 1, characterized in that, The semantic parsing through task decomposition and large language model specifically includes: The task decomposition large language model uses structured prompt word templates to perform semantic parsing; the structured prompt words include: 1) Skill space information for all robots, used to describe the operational skills possessed by each robot; 2) Current environment semantic information, used to provide information about the objects existing in the environment and their corresponding states; 3) Several examples of task decomposition based on the few-shot learning paradigm and their corresponding standard response formats; 4) The natural language task instructions to be executed.
5. A hybrid multi-robot collaborative system driven by a large language model, characterized in that, A hybrid architecture combining centralized task decomposition and distributed task negotiation is adopted, including: The task decomposition module receives natural language task instructions input by the user and performs semantic parsing through the task decomposition large language model to generate a task set containing multiple structured subtasks. Each subtask includes a natural language description and skill requirements. The multi-robot negotiation module involves multiple robots, each with its own integrated ontology negotiation language model, responding to and competing for each sub-task based on perceived environmental states, their own skill spaces, and sub-task information, using a decentralized negotiation mechanism. Specifically, each robot, based on its integrated ontology negotiation language model, generates a structured response through a comprehensive analysis of its own skill space, state perception information, and sub-task skill requirements. This response includes: 1) an assessment of executable operational skills and their suitability for the sub-task; and 2) an estimated cost to complete the candidate sub-task. For each sub-task, all robot responses are collected, and the execution costs of candidate robots meeting the skill requirements are compared. The sub-task is preferentially assigned to the robot with the lowest estimated cost. The module then determines which robot will perform the task. The task allocation result when a single robot acts as the executor; when the skill requirements of a subtask exceed the skill space of any single robot, a collaborative task team building process based on a decentralized language negotiation mechanism is triggered, specifically including: initiating a collaborative bidding based on the skill space of a multi-robot system, broadcasting the set of operational skills required for the subtask; each robot assessing the operational skills it can provide and their corresponding costs through an integrated ontology negotiation large language model, and submitting a structured response; automatically combining robot collaboration teams that meet the conditions according to the principle of minimizing skill coverage completeness and collaborative execution costs; automatically constructing a multi-robot collaboration team as the executor through the collaborative participation of the ontology negotiation large language models of each robot; and each executor executing the assigned subtask according to the negotiation results.
6. The hybrid multi-robot collaborative system driven by a large language model according to claim 5, characterized in that, The semantic parsing through task decomposition and large language model specifically includes: The task decomposition large language model uses structured prompt word templates to perform semantic parsing; the structured prompt words include: 1) Skill space information for all robots, used to describe the operational skills possessed by each robot; 2) Current environment semantic information, used to provide information about the objects existing in the environment and their corresponding states; 3) Several examples of task decomposition based on the few-shot learning paradigm and their corresponding standard response formats; 4) The natural language task instructions to be executed.
Citation Information
Patent Citations
Task planning method and equipment for heterogeneous multi-robot system driven by large language model
CN118567222A