Multi-agent cooperative task reasoning and robot scheduling system and method
Through the multi-agent collaboration architecture, combined with large language model and robot feedback information, the problem of inefficient task execution in complex environments is solved, and efficient and intelligent human-machine collaboration and task scheduling are achieved.
Patent Information
- Application Number
- CN202510616928.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-15
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-29
AI Technical Summary
Existing service robots lack deep natural language understanding and dynamic reasoning capabilities in complex environments and are unable to actively seek human assistance, resulting in inefficient or failure in task execution.
The multi-agent collaboration architecture is adopted, including perception, planning, decision-making and reflection agents, and task perception, planning, and decision-making are carried out through large language models, and dynamic scheduling and reflection are carried out in combination with robot feedback information to achieve natural language understanding and human-machine collaboration.
It improves the task execution efficiency and adaptability of service robots in complex environments, can actively request information or collaborate with users to solve problems, and achieve efficient and intelligent human-machine collaboration.
Smart Images

Figure CN120552040A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of large language models and intelligent robots, and specifically relates to a multi-agent collaborative embodied task reasoning and robot scheduling system and method, which is particularly suitable for application scenarios where robots and humans collaborate to perform tasks in complex office environments. Background Art
[0002] Existing service robots are typically limited to executing pre-set instructions when handling real-world tasks. They lack the deep understanding and dynamic reasoning capabilities of complex natural language commands, making them difficult to handle tasks in uncertain and volatile environments. Furthermore, when faced with task failure or information loss, traditional service robots often fail to proactively seek human assistance, resulting in inefficient or even complete failure.
[0003] While large language models have made significant progress in natural language processing in recent years, combining them with embodied robots and applying them to real-world office scenarios still faces numerous challenges. For example, existing service robots typically only follow pre-set instructions when performing real-world tasks. They lack the deep understanding and dynamic reasoning capabilities for complex natural language, making them difficult to cope with the uncertainty and variability of their environments. Furthermore, these robots are unable to proactively seek human assistance when task failures or information loss occur, resulting in inefficient execution and even mission failure. Despite the significant progress achieved in natural language processing with large language models, combining them with embodied robots and applying them to real-world office scenarios still faces multiple challenges, including embodied task perception and planning, efficient request mechanisms for human-robot collaboration, system integration and resource optimization, and robot deployment. Addressing these technical bottlenecks is crucial for realizing intelligent, adaptable, safe, and reliable service robots. Therefore, a method that combines embodied task reasoning with embodied task execution is urgently needed to enhance the task reasoning, decision-making, and scheduling capabilities of service robots in complex environments and meet practical application requirements. Summary of the Invention
[0004] The present invention aims to solve one of the technical problems existing in the existing related technologies to at least a certain extent.
[0005] To this end, the present invention provides a multi-agent collaborative embodied task reasoning and robot scheduling system and method, introduces a multi-agent collaborative architecture based on a large language model, and by simulating the four thinking steps of humans in dealing with problems - perception, planning, execution and reflection, endows service robots with natural language understanding, dynamic task reasoning, autonomous decision-making and human-computer collaboration capabilities in complex environments, and can efficiently handle complex task instructions and dynamic situations.
[0006] To achieve the above-mentioned purpose, the present invention adopts the following technical solutions:
[0007] A first aspect of the present invention provides a multi-agent collaborative embodied task reasoning and robot scheduling system, comprising:
[0008] an interaction unit, configured to obtain natural language instructions issued by the user and send them to the intelligent processing unit, interact with the intelligent processing unit, receive task execution results and robot embodied state information, and issue a request to supplement missing task information when the intelligent processing unit determines that task information is missing;
[0009] The intelligent processing unit is configured to dynamically perceive the environment and task status based on user instructions sent by the interaction unit, embodied state information and environmental information fed back by the robot, and the memory stored in the intelligent processing unit itself, and to dynamically plan and make decisions based on the perception results and reflection results on the task execution results, thereby realizing task reasoning and robot scheduling;
[0010] The robot cluster performs tasks according to the action sequence issued by the intelligent processing unit and feeds back its own embodied state information through the interface to help the intelligent processing unit dynamically perceive the environmental state and update the memory of the intelligent processing unit.
[0011] In some embodiments, the intelligent processing unit adopts a multi-agent collaborative architecture based on a large language model, including a memory module, a perception agent, a planning agent, a decision agent, and a reflection agent. Each agent interacts through natural language and shares the memory stored in the memory module.
[0012] The perception agent is used to dynamically perceive the environment and task status based on the user instructions obtained by the interaction unit, combined with the robot's embodied state information and historical task execution logs, and generate a task perception package containing task goals, location information and real-time task execution feedback;
[0013] The planning agent is used to dynamically reason about task decomposition and execution scheme based on the task perception package and the reflection results provided by the reflection agent, and generate a task plan;
[0014] The decision-making agent is used to analyze the task plan, generate action sequences, and perform robot scheduling to guide the robots to complete the task;
[0015] The reflective agent is used to evaluate the task execution results. If the task execution results do not meet expectations, it provides feedback for the next round of task reasoning to replan or adjust the execution plan.
[0016] In some embodiments, the memory module stores long-term memory and short-term memory;
[0017] The long-term memory includes an environmental semantic map, a user's default location, static object location information, public device location information, user preferences, and historical task execution logs;
[0018] The short-term memory includes real-time status information and dynamic context information during the current task execution process. The real-time status information includes user instructions, current task execution progress, the robot's embodied status information, dynamic environment information, the reflection results of the reflective intelligent agent, and user feedback records during task execution.
[0019] In some embodiments, the task reasoning modes provided by the intelligent processing unit include reactive, proactive, and proactive;
[0020] Reactive: When the perception agent determines that the task information is complete based on the user instructions sent by the interaction unit and the embodied status feedback from the robot, it extracts the task objectives and operation requirements from the task information and generates a task perception package containing the task objectives, location information and real-time feedback; the planning agent generates task execution steps based on the task perception package and the reflection results provided by the reflection agent to form a task plan. The task plan should meet the most basic executable conditions; the decision-making agent generates an action sequence based on the task plan issued by the planning agent and the embodied capabilities of the robot cluster;
[0021] Active: When the perception agent determines that the task information is incomplete based on the user instructions sent by the interaction unit and the embodied state feedback from the robot, it sends a request for supplementary task information to the planning agent to address the information gap. The planning agent forms a corresponding task plan based on the request and the reflection results provided by the reflection agent. The decision-making agent generates an action sequence based on the task plan and obtains the required supplementary information by searching the relevant information in the memory module or issuing a user-initiated query request to the instruction through the interaction unit;
[0022] Proactive: When the current task cannot be completed after multiple rounds of active reasoning, the perception agent sends a multi-party assistance request to the planning agent. The planning agent forms a corresponding task plan based on the request and the reflection results provided by the reflection agent. The decision-making agent generates an action sequence based on the task plan and initiates query requests to other users in the environment through the interaction unit to expand the scope of information acquisition and ensure the smooth execution of the task.
[0023] In some embodiments, the method of initiating a query request to other users in the environment through the interaction unit includes:
[0024] Individual contact: identifying potential assisting users and sending requests to obtain necessary information. The potential assisting users include persons who are not directly involved in the task or persons whose records do not exist in the memory module;
[0025] Group broadcast: Publish assistance requests in the work scenario group and coordinate multiple resources to obtain the information required for the task.
[0026] In some embodiments, the robot comprises:
[0027] The robot knowledge container is used to store the environment semantic map, target user location and object coordinate information, and supports real-time updates;
[0028] A high-level robot function container, which provides function tools for obtaining and controlling the robot's embodied state, including navigation, object interaction, and real-time status feedback.
[0029] A robot execution module, configured to execute corresponding tasks according to the action sequence generated by the decision-making agent, including navigation, item interaction, and sending online messages;
[0030] The embodied state feedback interface is used to feed back the robot's current physical state and task progress to the perceptual agent to update the task state information.
[0031] In some embodiments, the action sequence generated by the decision-making agent includes online operations and physical actions. The decision-making agent sends action instructions to the robot execution module through an interface, and dynamically adjusts the execution plan during the task execution process.
[0032] In some embodiments, the process of the decision-making agent scheduling the robot includes:
[0033] The decision-making agent breaks down the task plan output by the planning agent into multiple subtasks based on the function tools available to the robot, and generates a complete action sequence; the decision-making agent selects one or more subtasks to be executed in the next step from the subtasks based on the robot's load, position and physical capabilities, determines the most suitable robot through a load balancing strategy, and sends specific action instructions to the corresponding robot execution module using a unified communication protocol; during the process of the robot executing the action instructions, the decision-making agent detects whether the action is proceeding smoothly by continuously receiving the operating status, power, execution progress or error reports provided by each robot through the interface, thereby realizing continuous tracking of the execution status; at the same time, the decision-making agent identifies abnormal situations based on rules or learning models, and once an abnormality is detected, it triggers error correction or interrupts task execution.
[0034] In some embodiments, the reflective agent determines whether the task is successfully completed by analyzing the real-time task execution feedback recorded in the task perception package, and records the execution status to achieve real-time evaluation of the task execution results; if the task is not completed or fails to execute, the reflective agent diagnoses the cause of the failure, provides optimization suggestions to the decision-making agent, and stores the reflection results in the memory module during subsequent information transmission; the reflective agent also optimizes the reasoning and execution process of future tasks through historical reflection results.
[0035] In some embodiments, the reflection results include a judgment on the success or failure of task execution, an analysis of the causes of abnormalities, and improvement suggestions.
[0036] A second aspect of the present invention provides an embodied task reasoning and robot scheduling method, comprising the following steps:
[0037] Obtain natural language instructions from the user, as well as the execution results of the task and the robot's embodied state information, and issue a request to supplement the missing task information if it is determined that the task information is missing;
[0038] The system dynamically perceives the environment and task status based on the natural language instructions, the embodied state information fed back by the robot, the perceived environmental information, and the stored memory, and performs dynamic planning and decision-making based on the perception results and the reflection results on the task execution results to achieve task reasoning and robot scheduling. The embodied state information is fed back by the robot cluster after executing the task according to the received total work sequence, and is used to dynamically perceive the environmental status and update the stored memory.
[0039] In some embodiments, the method supports three task interaction modes:
[0040] Reactive interaction: Directly execute the natural language instructions with complete information;
[0041] Active interaction: When task information is missing, the user requests instructions or searches its own stored memory to supplement the missing task information;
[0042] Proactive interaction: When active interaction fails to obtain missing information for a task, multiple parties collaborate to supplement the missing information by sending requests to other users in the environment, either one-on-one or by broadcasting in a group, seeking help.
[0043] The characteristics and beneficial effects of the present invention are:
[0044] 1) Enable service robots to dynamically reason and efficiently execute complex tasks through reactive, proactive, and active interaction mechanisms;
[0045] 2) Combined with the embodied state information of the robot feedback module, the task plan is optimized in real time to improve task adaptability and execution efficiency;
[0046] 3) Based on a combination of "large-scale training data + high-dimensional parameters + iterative optimization," the large model continuously learns linguistic patterns and inter-conceptual connections within a layered network structure, thereby achieving deep natural language understanding and reasoning capabilities. In terms of a multi-agent architecture, by dividing multiple large-model agents with different functions or strategies into roles and collaborating under a shared communication mechanism (such as natural language communication or a unified task protocol), each agent can leverage its assigned strengths and achieve a more complete task solution through collaborative reasoning across the "perception-planning-execution-reflection" phase. This division of labor and collaboration helps to "simplify complexity," applying the large model's powerful language understanding and reasoning capabilities to multi-step, multi-role task scenarios. Lower-level functional modules such as perception and execution do not need to master the full language and reasoning logic, but are instead coordinated by higher-level large language model agents. Ultimately, through efficient communication and complementary roles across the collaborative architecture, dynamic collaboration and deep reasoning on complex tasks are possible, giving the system greater adaptability and problem-solving capabilities.
[0047] Based on the above analysis, the multi-agent collaboration framework proposed in this paper has deep natural language understanding and dynamic human-machine collaboration capabilities, and can actively initiate information requests or collaborate with users to solve problems during execution;
[0048] 4) Introducing a reflection mechanism. Introducing a "reflection agent" into the multi-agent collaborative framework can be seen as a meta-reasoning mechanism. It is linked to the specific perception, planning, and execution processes and is specifically responsible for monitoring, evaluating, and dynamically optimizing the system's overall task execution process, thereby ensuring the accuracy of task results.
[0049] 5) Suitable for complex and dynamic environments, and can provide efficient, intelligent and humanized services in real office scenarios.
[0050] In summary, the present invention proposes a multi-agent collaborative embodied task reasoning and robot scheduling system and method, which utilizes the powerful natural language understanding and reasoning capabilities of the large language model to efficiently process complex task instructions and dynamic situations. Multi-agent collaboration enhances the flexibility and task decomposition capabilities of the system, allowing each agent to communicate in human language and collaborate to complete complex tasks, which not only increases the interpretability of the algorithm, but also improves the overall execution efficiency. By coordinating large language model agents, manual and robots, the framework achieves seamless information sharing and real-time decision-making, ensuring the intelligence and adaptability of task execution in environments such as offices. This integration method not only improves the adaptability and robustness of the system, but also significantly enhances the effect of human-computer collaboration. It is suitable for variable and complex actual application environments, and achieves efficient execution of embodied tasks and robot scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the overall framework of the multi-agent collaborative embodied task reasoning and robot scheduling system provided by the embodiment of the first aspect of the present invention.
[0052] Figure 2 It is a flowchart of the multi-agent internal collaborative reasoning based on a large model in an embodiment of the present invention.
[0053] Figure 3 It is a task execution flow chart of the intelligent agent processing unit in the reactive, proactive and proactive modes in a specific embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solutions and advantages of this application more clearly understood, this application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0055] On the contrary, this application covers any alternatives, modifications, equivalents, and solutions made within the spirit and scope of this application as defined by the claims. Furthermore, to facilitate a better understanding of this application, certain specific details are described in detail below in the detailed description of this application. Those skilled in the art will be able to fully understand this application without these details.
[0056] This invention provides a system and method for embodied task reasoning and robot scheduling based on large-scale multi-agent collaboration. Through multi-agent collaboration and dynamic reasoning, it enables efficient execution of complex tasks and real-time scheduling of robot embodied states. The following describes the execution process and methods of each module of the system in detail, using implementation examples.
[0057] See also Figure 1 、 Figure 2 The first embodiment of the present invention proposes a multi-agent collaborative embodied task reasoning and robot scheduling system, comprising an interconnected interaction unit 100, an intelligent processing unit 200, and an embodied robot cluster 300. The embodied robot cluster 300 includes multiple task-performing robots 310. The various units in the system work together to ensure that the robots can adapt to dynamic and changing task requirements through three interaction modes: reactive, proactive, and proactive. Specifically:
[0058] The interaction module 100 is configured to receive natural language instructions from the user and send them to the intelligent processing unit 200. The interaction module 100 is also configured to interact with the intelligent processing unit 200, receive task execution results and information about the robot 310's physical state, and provide a means of requesting the user to supplement key task information when task information is missing.
[0059] The intelligent processing unit 200 is used to dynamically perceive the environment and task status based on user instructions sent by the interaction unit 100, embodied state information and environmental information fed back by the robot, and the memory stored in the intelligent processing unit 100. Based on the perception results and reflection on the task execution results, dynamic planning and decision-making are carried out to achieve task reasoning and robot scheduling;
[0060] The embodied robot cluster 300 executes tasks according to the action sequence issued by the intelligent processing unit 200, and feeds back its own embodied status through the interface, helping the intelligent processing unit 200 to dynamically perceive the environmental status and update the memory of the intelligent processing unit 200.
[0061] In some embodiments, the interaction unit 100 interacts with the user through a commonly used social platform (such as WeChat), receives natural language instructions from the user, provides feedback on task execution results, and proactively requests additional information from the user when task information is missing. The natural language instructions issued by the user can be text instructions or voice instructions, which the interaction unit 100 recognizes and transmits to the intelligent processing unit 200.
[0062] In some embodiments, the intelligent processing unit 200 adopts a multi-agent collaborative architecture based on a large language model, including a memory module 210, a perception agent 220, a planning agent 230, a decision agent 240, and a reflection agent 250. Task reasoning and action sequence generation are performed through multi-agent collaboration, and each agent interacts through natural language.
[0063] The memory module 210 includes a long-term memory 211 for storing semantic maps, user locations, and item information, and a short-term memory 212 for storing real-time status information of task execution. The memory module 210 is shared among all agents.
[0064] The perception agent 220 is used to dynamically perceive the environment and task status based on the natural language instructions issued by the user obtained by the interaction unit 100, combined with the robot's embodied state information and historical task execution logs, and generate a task perception package containing task goals, location information and real-time feedback;
[0065] The planning agent 230 is used to dynamically reason about task decomposition and execution plans based on the task perception package and the reflection results provided by the reflection agent 250, and generate a task plan;
[0066] The decision agent 240 is used to analyze the task plan, generate an action sequence including specific online operations and physical actions, and perform robot scheduling to guide the robot 310 to complete the task;
[0067] The reflective agent 250 is used to evaluate the task execution results (including the reasoning content of each agent and the results of the robot's task execution). If the task execution results do not meet expectations, it provides feedback for the next round of task reasoning to replan or adjust the execution plan.
[0068] In some embodiments, each robot 310 includes:
[0069] The robot knowledge container 311 is used to store the environment semantic map, target user location and object coordinate information, and supports real-time updates;
[0070] The robot high-level function container 312 is used to provide function tools for obtaining and controlling the embodied state of the robot 310, including navigation, object interaction, and real-time status feedback;
[0071] The robot execution module 313 is used to execute corresponding tasks according to the action sequence generated by the decision-making agent 240, such as navigation, item interaction, and sending online messages;
[0072] The embodied state feedback interface 314 is used to feed back the current physical state (such as current position, power status, etc.) and task progress of the robot 310 to the perception agent 220, helping the perception agent 220 to dynamically perceive the embodied state of the robot, providing real-time data support for task reasoning and planning, and thus updating the task status information.
[0073] In some embodiments, the intelligent processing unit 200 is the core of this application. This unit relies on the collaborative interaction between the perception agent 220, the planning agent 230, the decision-making agent 240, and the reflection agent 250 to generate action sequences and perform dynamic scheduling. Through multimodal information fusion, multi-agent collaboration, and real-time human-computer interaction, it comprehensively analyzes and infers user natural language commands, robot embodied status, and environmental perception data, thereby providing the system with efficient and flexible task execution capabilities. The following is a detailed description of the various components of the intelligent processing unit 200.
[0074] 1. Memory module
[0075] The memory module 210 is the data support core of the entire system and consists of long-term memory 211 and short-term memory 212. The memory module 210 supports real-time updates. If the environment changes during task execution (such as the location of an object being moved), the perception agent 220 can perceive the latest environmental information, ensuring the system's environmental adaptability and memory retention capabilities during task execution, providing information support for task reasoning and decision-making.
[0076] Long-term Memory 211: This is used to store the system's relatively stable environmental semantic map, user default locations, static item locations, public device locations, user preferences, and task execution logs. Long-term Memory 211 uses key-value databases (such as Redis) and graph databases (such as Neo4j) for persistent storage of spatial semantic information and incorporates spatial indexing algorithms to accelerate retrieval.
[0077] Short-term memory 212: This is used to dynamically store real-time status and dynamic context information during the current task execution process, including user instructions, the current task execution progress, the robot's embodied state, dynamic environmental information, the reflection results of the reflective agent 250, and human feedback records during task execution. Short-term memory 212 uses an in-memory database and memory cache mechanism (such as the Key-Value memory model) to achieve fast reading and writing. Unlike long-term memory 211, the information stored in short-term memory 212 is cleared or transferred to the history record after the task is completed to maintain available system memory space.
[0078] 2. Perceptual Agent
[0079] The perception agent 220 receives the user's natural language instructions and the robot's current state information, and parses and integrates them to form a task perception package, which serves as the input of the planning agent 230. The perception process of the perception agent 220 is divided into three modes: reactive perception, proactive perception, and proactive perception, to respond to different user instructions:
[0080] Reactive perception: The perception agent 220 receives the user's natural language instructions sent by the interaction unit 100 and the robot's current embodied state fed back by the embodied state feedback interface 314. When the perception agent 220 determines that the task information is complete (clear user instructions are received and the environmental information is complete), it directly extracts the task objectives and operation requirements from the task information, and directly generates a task perception package in combination with the robot's current embodied state and environmental information.
[0081] Active perception: When the perception agent 220 determines that the task information is incomplete, the perception agent 220 sends a request for supplementary task information to the planning agent 230 based on the missing information. The planning agent 230 generates a corresponding task plan based on the request and the reflection results provided by the reflection agent 250, and the decision-making agent 240 generates a specific action sequence. The perception agent 220 actively checks the memory module 210 and / or actively initiates inquiries to the user who issued the instruction based on the action sequence to obtain the required supplementary information, and generates a task perception package based on the obtained task information and combined with the robot's current embodied state and environmental information.
[0082] Proactive perception: In complex tasks, the perception agent 220 combines the physical state information of the memory module 210 and the robot's embodied state feedback interface to dynamically perceive the current embodied state. After active perception and planning fail, the perception agent 230 sends a multi-party assistance request to the planning agent 230. The planning agent 230 generates a corresponding task plan based on the request and the reflection results provided by the reflection agent 250, and the decision-making agent 240 generates a specific action sequence. The perception agent 220 seeks assistance from other people or groups in the environment based on the action sequence to obtain auxiliary information to supplement the task information, providing a broader planning space for subsequent decision-making. The perception agent 220 generates a task perception package based on the acquired task information and combined with the robot's current embodied state and environmental information.
[0083] In this embodiment, the specific steps for the perception agent 220 to generate a task perception package include:
[0084] Step 1) Input data collection:
[0085] The input data collected by the perception agent 220 includes user natural language commands, the robot's embodied state, and environmental information. User natural language commands are text commands transmitted by the interaction unit 100. These text commands can be original text commands issued by the user or the recognition results of voice commands obtained by converting the user's voice commands through a speech-to-text model. The robot's embodied state includes the robot's real-time location information, power status, and fault reports transmitted by the robot's embodied state feedback interface. Environmental information includes sensor data and historical records. Sensor data includes images collected by the robot's onboard camera and radar data collected in real time by the robot's onboard ranging radar (such as a lidar). Historical records include long-term and short-term memories stored in the memory module 210, such as user preferences, target object coordinates, task execution logs, task completion status, abnormal status reports, and other information. The perception agent 220 stores the collected real-time data in the short-term memory 212 of the memory module 210.
[0086] Step 2) Pattern determination and information gap detection:
[0087] First, based on the pre-trained large language model GPT4o, the user's input command is analyzed for intent recognition, key element extraction (including location and item name), and context understanding to generate preliminary command parsing results.
[0088] Subsequently, the perception agent 220 determines whether the task information composed of the instruction parsing result and the environmental information is complete. The integrity of the instruction parsing result is determined by determining whether the instruction parsing result has a clear target object and operation requirement. For example, the instruction parsing result "Please give me the file on the table in office A" has a clear target object, execution action and target location, which is a complete instruction parsing result. When the perception agent 220 determines that the task information is incomplete (such as the target location is unknown, the target object attributes are unclear, or there is a conflict in key elements), the active perception mode is first triggered, and a request for supplementary task information is sent to the planning agent 230 for the information gap. The planning agent 230 generates a corresponding task plan based on the request and the reflection result provided by the reflection agent 250. The decision agent 240 generates a specific action sequence based on the task plan. In response to the action sequence, the perception agent 220 actively searches for relevant information in the memory module 210 or actively sends a request to the instruction agent 230 through the interaction unit 100. The issuing user initiates an inquiry request (the perception agent 220 gives priority to searching the memory module 210. If the missing information cannot be obtained after the search, the interaction unit 100 initiates an inquiry request to the instruction issuing user) to obtain the required supplementary information; if the task decomposition cannot be completed after multiple active perceptions or planning corrections, the perception agent 220 triggers the forward-looking perception mode and sends a multi-party assistance request to the planning agent 230. The planning agent 230 generates a corresponding task plan based on the request and the reflection result provided by the reflection agent 250. The decision-making agent 240 generates a specific action sequence based on the task plan. The perception agent 220 responds to the action sequence to further expand the information source, including calling more online resources, contacts or groups, obtaining richer contextual support, and generating auxiliary information; once the perception agent 220 determines that the task information is complete, the reactive perception mode is triggered to directly extract the task objectives and operation requirements from the complete task information to generate a preliminary perception package.
[0089] Step 3) Information integration:
[0090] The perception agent 220 integrates the generated preliminary perception package with the robot's embodied state and environmental information into a task perception package, which is used to provide a directly operable structured description for the planning agent 230.
[0091] Step 4) Output and callback:
[0092] The perception agent 220 transmits the task-based perception package to the planning agent 230 and keeps monitoring in the subsequent process. Once new information or user instruction updates are detected during the robot 310's task execution phase, the task perception package is supplemented or corrected in real time.
[0093] 3. Planning Agent
[0094] Planning Agent 230 leverages the multi-step reasoning capabilities of a large language model, combined with the long-term and short-term memory of Memory Module 210, to analyze the user's intent and environmental information in the task perception package and generate a specific task plan. Planning Agent 230's reasoning process includes the following three modes, each addressing different task complexities and execution states.
[0095] Reactive planning: When the planning agent 230 determines that the task perception package meets the most basic executable conditions, the planning agent 230 quickly generates task execution steps based on the task perception package to form a task plan, ensuring that the task objectives are clear and the path is clear, and passes the task plan to the decision-making agent 240 to directly guide the robot 310 to perform the task.
[0096] Active Planning: When the Planning Agent 230 receives a request for additional task information from the Perception Agent 220, it generates a task plan that automatically searches for relevant information in the Memory Module 210 or issues a request for additional information. During task execution, the Planning Agent 230 dynamically adjusts the task plan by combining the reflection results of the Reflection Agent 250 with the task perception package provided by the Perception Agent 220 to ensure smooth progress.
[0097] Forward-looking planning: When proactive information supplementation fails, the planning agent 230 can foresee that continuing to generate a task plan will lead to task execution failure. When the existing information is insufficient to support the generation of a feasible plan, the planning agent 230 will proactively initiate the forward-looking planning mode. Based on the received task perception package and the reflection results provided by the reflection agent 250, it will generate a task plan that requests collaboration or assistance from other human team members in the environment, obtain the necessary information, and thus ensure the smooth progress of the task and improve the success rate of the execution of user instructions. Specifically, it includes:
[0098] 1) Individual contact: accurately locate potential assisting personnel and send a request to obtain necessary information; wherein, potential assisting personnel refers to personnel who are not directly involved in the task and personnel who do not have any records in the memory module 210.
[0099] 2) Group Broadcast: Publish assistance requests in the work scenario group and coordinate multiple resources to obtain the information required for the task.
[0100] Through real-time human-machine collaboration and physical state feedback, the planning agent 230 has the ability to adaptively correct task plans and dynamically adjust the task execution path, thereby improving the reliability and efficiency of the system in complex task environments.
[0101] In this embodiment, the specific steps for the planning agent 230 to generate a task plan include:
[0102] Step 1) Input collection and initialization:
[0103] The planning agent 230 obtains the task perception package from the perception agent 220, calls the long-term memory and short-term memory stored in the memory module 210 for verification, and confirms the currently available resources (such as available resources for robots when they are currently idle) and prior knowledge (such as the location information of people and objects involved in the current task).
[0104] Step 2) Mode determination and mission plan generation:
[0105] When the planning agent 230 receives the task perception package sent by the perception agent 220 in the reactive mode, the reactive planning mode is triggered. The planning agent 230 uses the multi-step reasoning capability of the pre-trained large language model GPT4o to realize the semantic decomposition and sub-task generation of the instruction parsing results in the task perception package. The large language model uses the prompt engineering (PE) and chain thinking (Chain-of-Thought) mechanism to enable the model to explain why a certain sub-task strategy is adopted, thereby improving interpretability. The planning agent 230 passes the task plan output by the large language model (including the task execution process, the personnel and resource requirements and paths involved, etc.) to the decision-making agent 240 for the decision-making agent 240 to generate the next executable action. Among them, the task plan generated by the planning agent 230 should meet the most basic executable conditions, that is, it should meet the requirements of completeness, consistency, executableness and unambiguity at the same time. Completeness refers to the fact that the generated task plan includes the necessary elements for completing the task, including the target object, target location, and operation, and each necessary element should be clear and unambiguous. Specifically, for the target object or target location, the task plan clearly states the target object or target location to be operated on or traveled to (e.g., "the document in Office A"); for the operation, the task plan specifies the core operations that the robot needs to perform (e.g., "pick up," "move to," "hand it to someone"). Consistency refers to the consistency between the task objectives and environmental information in the task plan and the current resource availability, without major conflicts (e.g., the absence of obstacles or robot failures). Executability refers to the theoretical ability to complete the current task given the environment (i.e., the environmental information perceived by the perceptual agent), the robot's capabilities, and the time limit. Unambiguity refers to the absence of multiple meanings or ambiguities in the instruction parsing results and context, which would prevent the robot from determining the specific operation object or process. For example, when the instruction parsing results involve prior steps or prerequisites (e.g., permissions, item location), and they are guaranteed to have been met or do not require reconfirmation, unambiguity is considered met.
[0106] When the planning agent 230 receives a request for supplementary task information from the perception agent 220, the active planning mode is triggered. The planning agent 230 generates a task plan by spontaneously searching for relevant information in the memory module 210 or issuing a user-initiated inquiry request to the instruction through the interactive unit 100 in combination with the reflection results provided by the reflection agent 250. After the missing information is supplemented, the task perception package newly generated by the perception agent and the reflection results provided by the reflection agent 250 are used to form a task plan using multi-step reasoning based on a large language model, and then the task plan is passed to the decision-making agent 240.
[0107] When the planning agent 230 receives a multi-party assistance request sent by the perception agent 220 or the task perception package still cannot meet the most basic executable conditions after multiple rounds of active planning, and it is predicted that continuing to generate the task plan according to the current situation will lead to task execution failure, the forward-looking planning mode is triggered. The planning agent 230 generates a task plan based on the received task perception package and the reflection results provided by the reflection agent 250 to request cooperation or help from other human team members in the environment (including online groups and dedicated collaboration) to obtain a wider range of information sources. After the task perception package meets the most basic executable conditions, the multi-step reasoning based on the large language model is used to form a task plan and pass the task plan to the decision-making agent 240. Among them, when any one or more of the following situations occur during the task execution or planning process and cannot be resolved through further collaboration or external help, it is determined that continuing to generate the task plan according to the current situation will lead to task execution failure:
[0108] ① Information is missing or conflicting and difficult to fill: Actively requesting information from the user or memory module 210 times still fails to fill the key gap, or there is conflicting environmental information that cannot be resolved;
[0109] ② Multiple plans fail to materialize: After trying multiple alternative task decomposition solutions, no executable action sequence can be generated (e.g., the path is completely blocked by obstacles);
[0110] ③ Resource or physical limitations: The robot is low on battery and has no available charging solution; the robotic arm is damaged and has no way to repair it; or the external equipment required for the task is unavailable;
[0111] ④ Time or priority conflict: The task must be completed within the specified time, but during planning, it is found that the time limit cannot be met, or other higher-priority tasks occupy resources, causing the current task to be shelved and unable to be completed.
[0112] In this embodiment, the interaction and collaboration between the planning agent 230 and other agents include:
[0113] With the perception agent 220: obtain updated task perception packages in real time; if planning encounters a bottleneck, request the perception agent 220 to query more environmental, sensor, or social platform data. The perception agent 220 stores the queried data in the memory module 210 for the planning agent 230 to call, ensuring that the planning agent 230 obtains the latest information;
[0114] With the decision-making agent 240: the generated task plan is handed over to the decision-making agent 240 to determine the specific robot action instructions for the next step and monitor the execution feedback;
[0115] With the Reflective Agent 250: When errors or uncertainties occur while the robot is performing a task, the Reflective Agent 250 will make correction suggestions or re-planning requests, and the Planning Agent 230 will adjust the task plan or re-plan accordingly.
[0116] 4. Decision-making Agent
[0117] Based on the task plan generated by the planning agent 230, the decision agent 240 converts it into a specific action sequence that the robot can execute, including both online operations and physical actions. The decision agent 240 has three decision modes, which are selected based on the planning mode used by the planning agent 230:
[0118] Reactive decision-making: When the planning agent 230 adopts the reactive planning mode, the decision-making agent 240 selects the reactive decision-making mode, directly decomposes the task plan into sub-tasks, generates an action sequence, and assigns it to the most suitable one or more robots 310 to guide the robots 310 to complete the task.
[0119] Active decision-making: When the planning agent 230 adopts the active planning mode, the decision-making agent 240 selects the active decision-making mode, generates a search memory module 210 according to the task plan generated by the planning agent 230, or sends an action sequence of the user-initiated request for supplementary task information to the instruction through the interactive unit 100 to obtain the missing task information. The perception agent 220, the planning agent 230 and the decision-making agent 240 collaborate to perform perception, planning and decision-making in turn, thereby ensuring the smooth execution of the user's instructions.
[0120] Proactive Decision-Making: When Planning Agent 230 uses the proactive planning mode, Decision Agent 240 selects the proactive decision-making mode and generates tasks based on the task plan generated by Planning Agent 230. Combining the task perception package with proactive planning, Decision Agent 240 not only generates tasks such as robot actions and online notifications based on the established plan, but also proactively collaborates with Planning Agent 230 to conduct online inquiries via individual or group chats, proactively obtaining additional information to optimize subsequent decisions. By dynamically adjusting the order and strategy of task execution, Decision Agent 240 proactively avoids erroneous actions, improves the success rate of user-command completion, and ensures efficient and accurate task execution.
[0121] In this embodiment, the specific decision-making process of the decision agent 240 includes:
[0122] Step 1) Input collection and initialization:
[0123] The decision agent 240 obtains the task plan including the task execution process, personnel involved, resource requirements, etc. from the planning agent 230.
[0124] Step 2) Mode determination and action sequence generation and collaborative allocation:
[0125] The decision agent 240 triggers its own decision mode according to the attributes of the task plan generated by the planning agent 240.
[0126] Reactive decision-making: If the task plan is a complete execution plan for the original instruction, the reactive decision-making mode is triggered to generate an action sequence and perform specific sub-task allocation. Specifically: according to the function tools available to the robot 310, the task plan (corresponding to the high-level task steps) output by the planning agent 230 is decomposed into multiple sub-tasks using the reasoning ability of the large language model (specifically, the sub-task decomposition can be achieved through thinking chains and prompt words plus a few sample decomposition examples). The sub-tasks use action instructions or function interfaces that can be directly called by the robot 310, such as "move to coordinates (x, y)", "grab item A", "talk to a user", API calls, etc.; a template is established for basic action instructions (such as "grab", "move", etc.), and a complete action sequence is generated by filling parameters (target coordinates, target objects, etc.) into the template. Subsequently, for multi-robot scenarios, the decision-making agent 240 extracts one or more subtasks that should be executed in the next step from the refined subtasks based on indicators such as the load, position and embodied capabilities of the robot 310, determines the most suitable robot through a load balancing strategy based on the Hungarian Algorithm, and sends specific action instructions to the corresponding robot execution module 313 using a unified ROS Topic communication protocol; during the process of the robot executing the action instructions, the decision-making agent 240 continuously receives the operating status, power, execution progress or error report provided by each robot 310 through the embodied state feedback interface 314 to detect whether the action is proceeding smoothly, thereby achieving continuous tracking of the execution status; at the same time, the decision-making agent 240 quickly identifies abnormal situations (such as mechanical failures or accidental blockages of the path, etc.) based on rules or learning models, and once an abnormality is detected, it triggers error correction or interrupts the process.
[0127] Active decision-making: If the task plan is generated based on active planning, that is, the planning agent 230 finds that the task information is missing when generating the task, the decision-making agent 250 enters the active decision-making mode. In this mode, the robot will not be deployed directly. Instead, the user who issues the instruction will initiate an active inquiry for the active task plan. Specifically, online operations can be generated by calling the information sending method of the robot (specifically a smart phone), or by searching the memory module 210 to obtain the missing information.
[0128] Proactive Decision-Making: If the planning agent 230 enters proactive planning mode, meaning proactive information supplementation fails, the planning agent 230 will issue a task plan that requires initiating a single or group chat to obtain a wider range of assistance information to facilitate the smooth execution of the instructions. At this point, the decision-making agent 250 enters proactive decision-making mode. In this mode, the robot (specifically, a smartphone) is invoked to generate an online operation based on the proactive task plan, using the robot's (specifically, a smartphone's) message sending method. Based on the user or group chat name specified in the task plan, the corresponding online request for assistance message is initiated.
[0129] In this embodiment, the interaction and collaboration between the decision agent 240 and other agents include:
[0130] Interaction with the Perception Agent 220: Acquire the latest task perception package in real time to ensure that decisions are based on current environmental information for task decomposition and allocation. When the decision-making process encounters a bottleneck, the Perception Agent 220 can be requested to further query environmental, sensor, or social platform data, and the query results can be stored in the memory module 210 for the decision-making agent 240 to access and optimize the task decomposition and allocation process.
[0131] Collaboration with the planning agent 230: Based on the task plan provided by the planning agent 230, extract the specific tasks that need to be performed at the current moment;
[0132] Interaction with the reflective agent 250: After the task is completed, the decision-making agent 240 submits the decomposition and allocation results to the reflective agent 250, assists the reflective agent 250 in analyzing the execution effect and potential problems, generates reflection results, and integrates optimization suggestions into the decomposition and allocation process of future tasks.
[0133] 5. Reflective Agents
[0134] The reflective agent 250 evaluates the task execution results based on the execution log, task completion status, abnormal status report and other information obtained from the task perception package of the perception agent 220, and provides the reflection results for system optimization to ensure the accuracy and completeness of the task. The specific steps include:
[0135] Task evaluation: The reflective agent 250 analyzes the results of the robot execution module feedback recorded in the task perception package, determines whether the task is successfully completed, and records the execution status.
[0136] Problem diagnosis: If the task is not completed or fails to execute, the reflective agent diagnoses the cause of the failure, provides optimization suggestions to the decision-making agent, and stores the reflection results in the memory unit during the subsequent information transmission process.
[0137] Process optimization: By leveraging historical reflection data, we can optimize the reasoning and execution processes of future tasks and improve the success rate of tasks.
[0138] Step 1) Input collection and initialization:
[0139] The reflective agent 250 obtains the current task execution results from the task perception package of the perception agent 220, including execution logs, task completion status, abnormal status reports and other information, and integrates them. The reflective agent 250 also calls long-term and short-term memories by accessing the memory module 210 to obtain the preset task goals and possibly related historical execution records (such as the completion status of similar tasks in the past). Ultimately, these two parts form a complete reflective closed-loop information chain for reflecting on the current step. Through unified timestamp management during the reflection process, the reflective agent 250 can accurately find the sequence and causal relationship of each step in the task execution, thereby improving diagnostic efficiency.
[0140] Step 2) Task Evaluation:
[0141] The reflective agent 250 compares the current task execution results with the preset task objectives. If the current task execution results meet or exceed the preset task objectives, the task is deemed successful and execution details (such as time and resource consumption) are recorded in the task execution log. If the current task execution results do not meet the preset task objectives, such as missing steps, incorrect item coordinates, or unsatisfactory user feedback, the task is deemed to have deviated or failed, and the next step of the problem diagnosis process is entered.
[0142] Step 3) Problem diagnosis:
[0143] First, the reflection agent 250 collects the task perception package generated by the perception agent 220, the task plan generated by the planning agent 230, and the action sequence generated by the decision agent 240 as evidence, and comprehensively compares the environmental information before and after the robot executes the action sequence;
[0144] Subsequently, the reflective agent 250 utilizes the chain reasoning of the large language model, such as providing a textual description of the cause of the error and the error correction strategy, so that the reflective agent can output the judgment process and improvement plan in natural language to improve interpretability, or use rule-based reasoning to locate the specific links where task deviation or failure occurs and infer the cause. Among them, the inferred causes of task deviation include: any one or more of incomplete perception data, planning logic conflicts, decision-making instruction errors, robot hardware failures, and temporary changes in user needs.
[0145] Finally, the reflective agent 250 provides a textual description of the problem diagnosis results, outputs the problem diagnosis results in natural language to improve interpretability, and proceeds to the next step of the process.
[0146] Step 4) Feedback generation and process optimization:
[0147] If the diagnosed problem can be fixed locally (such as simply re-grasping or replacing the robot), the reflective agent 250 sends a local error correction instruction to the decision-making agent 240; if the diagnosed problem involves the overall reconstruction of the task, the reflective agent 250 passes the cause of failure and the need for additional information to the planning agent 240; if the diagnosed problem is caused by the lack of key information or requires user decision-making, the reflective agent 250 passes the information that it needs to call the interactive unit 100 to request user intervention (clarification of requirements, provision of additional resources, etc.) to the planning agent 230, and the planning agent 230 generates a corresponding task plan, and the decision-making agent 240 generates a specific online inquiry action sequence, and re-performs perception, planning and decision-making after the robot responds.
[0148] In some embodiments, the present invention achieves deep collaboration with humans through three task execution modes: reactive, proactive, and active:
[0149] 1) Responsive interaction: When user instructions are clear and information is complete, the system directly executes the task and provides feedback, improving task execution efficiency.
[0150] 2) Active interaction: When task information is incomplete or execution is blocked, the system initiates inquiries to the user or team members through the interaction unit 100 to obtain missing information or request assistance, and dynamically update the task execution plan.
[0151] 3) Proactive Interaction: In complex tasks, the system combines embodied state feedback information and environmental perception to proactively identify potential problems or predictively determine future resource needs, collaborating with users or team members to optimize execution plans and ensure smooth task completion.
[0152] See also Figure 1 、 2 The second embodiment of the present invention provides an embodied task reasoning and robot scheduling method based on the above system, including:
[0153] Obtain natural language instructions from the user, as well as the execution results of the task and the robot's embodied state information, and issue a request to supplement the missing task information if it is determined that the task information is missing;
[0154] Based on the acquired natural language instructions, the embodied state information fed back by the robot, the perceived environmental information, and the stored memory, the environment and task status are dynamically perceived. Dynamic planning and decision-making are performed based on the perception results and the reflection results on the task execution results to achieve task reasoning and robot scheduling. Among them, the embodied state information is fed back by the robot cluster after executing the task according to the received total work sequence, which is used to dynamically perceive the environmental status and update the stored memory.
[0155] Furthermore, the method of the embodiment of the present invention specifically includes the following steps:
[0156] a) Receiving user instructions: receiving the user's natural language instructions through the interaction unit 100 and passing the instructions to the intelligent processing unit 200;
[0157] b) Generate task perception package: The perception agent 220 in the intelligent processing unit 200 parses the user's natural language instructions and generates a task perception package by combining long-term memory, environmental information, and the robot's embodied state information;
[0158] c) Task Reasoning and Planning: When receiving a user instruction for the first time, the planning agent 230 will directly reason on the task perception package, decompose the task, and generate a task plan without combining it with the reflection results. When receiving a user instruction for a different period of time, the planning agent 230 will generate a task plan based on the task perception package and the reflection results provided by the reflection agent 250. The planning agent 230 will also dynamically adjust the task plan based on the environmental information fed back by the perception agent 220 and the robot's embodied state information during task execution.
[0159] d) Task decision and interaction: The decision agent 240 generates an executable action sequence based on the task plan generated by the planning agent 230, including specific online tasks and physical tasks, and sends it to the corresponding robot execution module. When information is missing, it actively interacts with the instruction issuing user to supplement the required information;
[0160] e) Task execution and status feedback: The robot execution module receives the action sequence issued by the decision agent 240, executes the task through high-level functions, and provides real-time feedback of the current embodied state to the perception agent 220. The perception agent 220 then re-perceives the state change information, generates a new task perception package, and updates the memory module 210.
[0161] f) Reflection and optimization: The reflection agent 250 evaluates the task execution results based on the dynamic environmental information in the task perception packages at the previous and next moments. If the results do not meet expectations, problem diagnosis and process optimization are performed to form a reflection result. The planning agent and decision-making agent re-plan and make decisions based on the reflection result to ensure that the task is completed. After the task is completed, the system updates the memory module information to provide a reference for future task execution.
[0162] The following combination Figure 3 The embodiments of the present invention are described with specific task execution examples.
[0163] Example task: "Zhang San: Take this document to Li Si for signature, and then give it to Wang Wu."
[0164] Deployable embodied robots: a mobile chassis robot with a smart lock (hereinafter named "Robot No. 1"), a robot dog (hereinafter named "Robot No. 2"), a robotic arm (hereinafter named "Robot No. 3"), and a smartphone (with WeChat and other chat software installed and accounts configured).
[0165] For the above example tasks, the task reasoning and robot scheduling process of each agent in the embodiment of the present invention is as follows:
[0166] Time step one:
[0167] 1. Perception Agent
[0168] The perceptual agent 220 perceives a new user instruction in WeChat, receives the user instruction, and parses the preliminary information: "The file needs to be obtained from Zhang San and sent to Li Si first, and finally to Wang Wu."
[0169] The perception agent 220 calls the memory module 210 to query the location information of objects in the environment, Zhang San, Li Si, Wang Wu and others, as well as the user's intention.
[0170] The perception agent 220 uses the queried information, the original user instructions and the perception summary information to generate a task perception package and pass it to the reflection agent 250 (if it is the first execution, it skips the reflection agent 250 and passes it directly to the planning agent 230).
[0171] 2. Planning Agent:
[0172] The planning agent 230 decomposes the original user instructions in the task perception package into three steps:
[0173] ① Get the document from Zhang San first;
[0174] ② Go to Li Si's location and send him a signature;
[0175] ③Finally, go to Wang Wu’s position and hand it over to him.
[0176] The planning agent 230 generates a complete task plan based on the above task decomposition results. If important information constituting the above task plan is missing at this time, it starts to generate a backup task plan.
[0177] 3. Decision-making Agent:
[0178] The decision agent 240 compares the task plan generated by the planning agent 230 with the capabilities of different robots in the embodied robot cluster and infers and generates a specific action sequence:
[0179] ①Physical action: <Robot No. 1: Move to Zhang San>.
[0180] ②Online operation: <Mobile phone: send a notification and a QR code for unlocking to Zhang San>.
[0181] The decision-making agent 240 calls the robot high-level function container 312 according to the generated action sequence, controls robot No. 1 to navigate to the first target position (where Zhang San is), and waits for Zhang San to complete the operation of placing the item.
[0182] Robot No. 1 transmits its embodied feedback information to the perception agent 220, which perceives new environmental information and embodied status, such as the placement of items.
[0183] Time step two:
[0184] 1. Perception Agent
[0185] The perception agent 220 perceives that Robot No. 1 has completed the first step of the task and successfully obtained the file.
[0186] The memory module 210 queries the location of Li Si and updates the execution status of the previous task into the task perception package. The task perception package is updated as follows:
[0187] Current status: The file has been retrieved, and Robot 1 is at Zhang San's location.
[0188] The perception agent 220 passes the updated task perception package to the planning agent 230.
[0189] 2. Reflective Agent:
[0190] The reflective agent 250 evaluates the task execution result based on the generated plan, the specific task generated by the decision-making agent 240 in the previous time step, and the task perception package of the two state changes before and after the task execution, and confirms that Robot No. 1 has obtained the file; if it fails midway (such as Zhang San did not put in the item), the reflective agent 250 will reflect on the reasons for the failure of task execution.
[0191] After finishing its thinking, the reflective agent 250 will pass its reflection results to the downstream planning agent 230 and decision-making agent 240 to guide them in the next step of planning and decision-making, that is, whether to continue to execute the specific tasks corresponding to the remaining part of the user instructions, or to re-execute the specific tasks that failed in the previous step.
[0192] 3. Planning Agent:
[0193] If the execution result of the previous step is correct after reflection by the reflection agent 250, the planning agent 230 decomposes the unfinished part into steps according to the user instructions and the content of the task perception package:
[0194] ① Take the document to Li Si's location and ask him to sign;
[0195] ② Go to Wang Wu’s position and hand it over to him.
[0196] If the result of the previous step is reflected by the reflection agent 250 and is found to be an error, the planning agent 230 will generate a supplementary plan based on the reflection result. Assuming that the reason for the error in task execution is that Zhang San did not scan the code to put the file into Robot No. 1, the planning agent 230 will update the task plan as follows:
[0197] Send a confirmation message to Zhang San, asking him whether he has put the file into the robot locker;
[0198] 4. Decision-making Agent:
[0199] The decision agent 240 generates a specific action sequence based on the task plan issued by the planning agent 230 and matching it with the robot cluster capabilities:
[0200] Physical action: <Robot No. 1: Move to Li Si's position>.
[0201] Online operation: <Smartphone: Send a notification to Li Si, reminding him to prepare to sign>.
[0202] The decision-making agent 240 calls the high-level function container 312 of robot No. 1 according to the generated action sequence, controls robot No. 1 to navigate to where Li Si is, and sends a notification to him using a smartphone.
[0203] After arriving at Li Sisi according to the navigation, Robot No. 1 waits for the signing process to be completed, monitors the document status in real time, and transmits the embodied status to the perceptual intelligent agent 220 through its own embodied status feedback interface 314.
[0204] Time step three:
[0205] 1. Perception Agent
[0206] The perceptual agent 220 receives feedback from robot No. 1 and confirms that the document has been signed by Li Sisi.
[0207] The perception agent 220 calls the memory module 210 to query Wang Wu's location and environmental information.
[0208] The execution status of the previous task is updated to the task perception package. The content of the updated task perception package is: Current status: The document has been signed, and Robot No. 1 is at Li Si's position.
[0209] The perceptual agent 220 passes the updated task perception package to the reflective agent 250.
[0210] 2. Reflective Agent:
[0211] The reflective agent 250 evaluates the task execution result based on the task plan generated in the previous step, the specific task, and the task perception package of the two state changes before and after the task execution: if the document has been signed, the current task is confirmed to be successful and the next stage of the task will continue; if the document has not been signed (such as Li Si did not sign), the reflective agent 250 will feedback the reason.
[0212] After completing the reflection, the reflection agent 250 passes the reflection results to the downstream planning agent 230 and decision-making agent 240, guiding them to continue to execute subsequent instructions or generate remedial measures.
[0213] 3. Planning Agent:
[0214] Based on the task perception package provided by the perception agent 220, the planning agent 230 determines the unfinished task part and generates the following task plan: take the signed document to Wang Wu's location and deliver the document to Wang Wu.
[0215] If there is a failure in the previous task, the planning agent 230 generates a remediation plan based on the feedback from the reflection agent 250. For example, if the document is not signed, the planning agent 230 will insert a remediation step (such as having the robot resubmit the document to Li Si).
[0216] 4. Decision-making Agent:
[0217] The decision agent 240 generates a specific action sequence based on the task plan issued by the planning agent 230 and the capabilities of the robot cluster:
[0218] ①Physical action: <Robot No. 1: Move to Wang Wu’s position>.
[0219] ②Online operation: <Smartphone: Send a notification to Wang Wu, reminding him to prepare to receive the file>.
[0220] The decision-making agent 240 calls the high-level function container 312 of robot No. 1, controls robot No. 1 to navigate to Wang Wu's location, and sends a notification to Wang Wu via the smartphone.
[0221] After Robot No. 1 arrives at the destination, it checks the delivery status to ensure that the file has been successfully delivered to Wang Wu; Robot No. 1 transmits its embodied feedback information to the perception agent 220 to report the file delivery status.
[0222] 5. Mission Completed:
[0223] If the file is successfully delivered to Wang Wu, the perception agent 220 will update the final task completion status to the task perception package, and then generate an action sequence to notify the user that the task has been completed after passing through the reflection agent 250, the planning agent 230 and the decision-making agent 240. The smartphone responds to the action sequence and notifies the user that the task has been completed.
[0224] If the file is not delivered successfully, the reflection agent 250 will provide feedback on the reasons for the failure and guide the planning agent 230 and the decision-making agent 240 to generate remedial measures to ensure the successful completion of the task.
[0225] The embodiment of the present invention embodies three task processing modes during task execution:
[0226] 1. Reactive task processing: After receiving user instructions, it automatically begins to perceive the environment, plan problems, generate specific decision-making tasks, reflect on task behaviors and results, and finally complete user instructions like a human assistant.
[0227] 2. Active task completion: When information is incomplete, proactively interact with the user to complete the information to ensure smooth task execution.
[0228] 3. Proactive help-seeking: The robot monitors the execution process in real time through embodied feedback, predicts the results of instruction execution, and proactively seeks human help to ensure the correct implementation of the task.
[0229] In summary, the embodiments of the present invention provide a multi-agent collaborative embodied task reasoning and robot scheduling system and method with the following advantages:
[0230] 1. Multi-agent collaboration: Through the collaborative work of perception, planning, decision-making and reflection agents, intelligent reasoning and efficient execution of complex tasks are achieved.
[0231] 2. Dynamic reasoning and interaction: It has reactive, proactive, and proactive task interaction modes, effectively solving the problems of incomplete task information and dynamic environmental changes.
[0232] 3. Closed-loop feedback and optimization: Dynamically evaluate execution results through reflective agents to continuously optimize system performance and task execution results.
[0233] Therefore, the present invention effectively improves the task execution capability of service robots in real scenarios and provides an intelligent and dynamic solution for the processing of complex tasks.
[0234] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0235] Although the embodiments of the present disclosure have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and alterations may be made to the embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the claims and their equivalents.
Claims
1. A multi-agent collaborative embodied task reasoning and robot scheduling system, characterized by: include: An interaction unit, configured to obtain natural language instructions issued by the user and send them to the intelligent processing unit, interact with the intelligent processing unit, receive task execution results and robot embodied state information, and issue a request for supplementary key task information when the intelligent processing unit determines that task information is missing; The intelligent processing unit is configured to dynamically perceive the environment and task status based on user instructions sent by the interaction unit, embodied state information and environmental information fed back by the robot, and the memory stored in the intelligent processing unit itself, and to dynamically plan and make decisions based on the perception results and reflection results on the task execution results, thereby realizing task reasoning and robot scheduling; The robot cluster performs tasks according to the action sequence issued by the intelligent processing unit and feeds back its own embodied state information through the interface to help the intelligent processing unit dynamically perceive the environmental state and update the memory of the intelligent processing unit.
2. The system according to claim 1, wherein: The intelligent processing unit adopts a multi-agent collaborative architecture based on a large language model, including a memory module, a perception agent, a planning agent, a decision-making agent, and a reflection agent. Each agent interacts through natural language and shares the memory stored in the memory module. The perception agent is used to dynamically perceive the environment and task status based on the user instructions obtained by the interaction unit, combined with the robot's embodied state information and historical task execution logs, and generate a task perception package containing task goals, location information and real-time task execution feedback; The planning agent is used to dynamically reason about task decomposition and execution scheme based on the task perception package and the reflection results provided by the reflection agent, and generate a task plan; The decision-making agent is used to analyze the task plan, generate action sequences, and perform robot scheduling to guide the robots to complete the task; The reflective agent is used to evaluate the task execution results. If the task execution results do not meet expectations, it provides feedback for the next round of task reasoning to replan or adjust the execution plan.
3. The system according to claim 2, characterized in that The memory module stores long-term memory and short-term memory; The long-term memory includes an environmental semantic map, a user's default location, static object location information, public device location information, user preferences, and historical task execution logs; The short-term memory includes real-time status information and dynamic context information during the current task execution process. The real-time status information includes user instructions, current task execution progress, the robot's embodied status information, dynamic environment information, the reflection results of the reflective intelligent agent, and user feedback records during task execution.
4. The system according to claim 2, wherein: The task reasoning modes of the intelligent processing unit include reactive, proactive and proactive; Reactive: When the perception agent determines that the task information is complete based on the user instructions sent by the interaction unit and the embodied state feedback from the robot, it extracts the task objectives and operation requirements from the task information and generates a task perception package containing the task objectives, location information, and real-time feedback. The planning agent generates task execution steps based on the task perception package and the reflection results provided by the reflection agent to form a task plan. The task plan should meet the most basic executable conditions. The decision agent generates an action sequence based on the task plan issued by the planning agent and matching it with the embodied capabilities of the robot cluster; Active: When the perception agent determines that the task information is incomplete based on the user instructions sent by the interaction unit and the embodied state feedback from the robot, it sends a request for supplementary task information to the planning agent to address the information gap. The planning agent forms a corresponding task plan based on the request and the reflection results provided by the reflection agent. The decision-making agent generates an action sequence based on the task plan and obtains the required supplementary information by searching the relevant information in the memory module or issuing a user-initiated query request to the instruction through the interaction unit; Proactive: When the current task cannot be completed after multiple rounds of active reasoning, the perception agent sends a multi-party assistance request to the planning agent. The planning agent forms a corresponding task plan based on the request and the reflection results provided by the reflection agent. The decision-making agent generates an action sequence based on the task plan and initiates query requests to other users in the environment through the interaction unit to expand the scope of information acquisition and ensure the smooth execution of the task.
5. The system according to claim 4, characterized in that The method of initiating a query request to other users in the environment through the interaction unit includes: Individual contact: identifying potential assisting users and sending requests to obtain necessary information. The potential assisting users include persons who are not directly involved in the task or persons whose records do not exist in the memory module; Group broadcast: Publish assistance requests in the work scenario group and coordinate multiple resources to obtain the information required for the task.
6. The system according to claim 2, wherein: The robot comprises: The robot knowledge container is used to store the environment semantic map, target user location and object coordinate information, and supports real-time updates; A high-level robot function container, which provides function tools for obtaining and controlling the robot's embodied state, including navigation, object interaction, and real-time status feedback. A robot execution module, configured to execute corresponding tasks according to the action sequence generated by the decision-making agent, including navigation, item interaction, and sending online messages; The embodied state feedback interface is used to feed back the robot's current physical state and task progress to the perceptual agent to update the task state information.
7. The system according to claim 6, characterized in that The action sequence generated by the decision-making agent includes online operations and physical actions. The decision-making agent sends action instructions to the robot execution module through an interface and dynamically adjusts the execution plan during the task execution process.
8. The system according to claim 6, wherein: The process of the decision-making agent scheduling the robot includes: The decision-making agent breaks down the task plan output by the planning agent into multiple subtasks based on the function tools available to the robot, and generates a complete action sequence; the decision-making agent selects one or more subtasks to be executed in the next step from the subtasks based on the robot's load, position and physical capabilities, determines the most suitable robot through a load balancing strategy, and sends specific action instructions to the corresponding robot execution module using a unified communication protocol; during the process of the robot executing the action instructions, the decision-making agent detects whether the action is proceeding smoothly by continuously receiving the operating status, power, execution progress or error reports provided by each robot through the interface, thereby realizing continuous tracking of the execution status; at the same time, the decision-making agent identifies abnormal situations based on rules or learning models, and once an abnormality is detected, it triggers error correction or interrupts task execution.
9. The system according to claim 2, wherein: The reflective agent determines whether the task is successfully completed by analyzing the real-time task execution feedback recorded in the task perception package, and records the execution status to achieve real-time evaluation of the task execution results; if the task is not completed or fails to execute, the reflective agent diagnoses the cause of the failure, provides optimization suggestions to the decision-making agent, and stores the reflection results in the memory module during subsequent information transmission; the reflective agent also optimizes the reasoning and execution process of future tasks through historical reflection results.
10. The system according to claim 2, wherein: The reflection results include the judgment of the success or failure of the task execution, the analysis of the causes of abnormalities and improvement suggestions.
11. A method for embodied task reasoning and robot scheduling, characterized in that: The following steps are involved: Obtain natural language instructions from the user, as well as the execution results of the task and the robot's embodied state information, and issue a request to supplement the missing task information if it is determined that the task information is missing; The system dynamically perceives the environment and task status based on the natural language instructions, the embodied state information fed back by the robot, the perceived environmental information, and the stored memory, and performs dynamic planning and decision-making based on the perception results and the reflection results on the task execution results to achieve task reasoning and robot scheduling. The embodied state information is fed back by the robot cluster after executing the task according to the received total work sequence, and is used to dynamically perceive the environmental status and update the stored memory.
12. The method according to claim 11, characterized in that The method supports three task interaction modes: Reactive interaction: Directly execute the natural language instructions with complete information; Active interaction: When task information is missing, the user requests instructions or searches its own stored memory to supplement the missing task information; Proactive interaction: When active interaction fails to obtain missing information for a task, multiple parties collaborate to supplement the missing information by sending requests to other users in the environment, either one-on-one or by broadcasting in a group, seeking help.
Citation Information
Cited By
Multi-agent collaborative system based on large model and bionic human brain structure
CN120745685A
Multi-agent collaborative system based on large model and bionic brain structure
CN120745685B
Multi-modal large model body planning method and system based on iterative feedback
CN121480559A
Multi-agent collaborative embodied task reasoning and robot scheduling system and method
WO2026152575A1