Multi-agent cooperative task reasoning and robot scheduling system and method

Through the multi-agent collaboration architecture, the human thinking steps are simulated, and the service robots are given natural language understanding and dynamic reasoning capabilities, which solves the problem that service robots in the existing technology are difficult to cope with complex natural language instructions and dynamic environments, and achieves efficient and highly adaptable task execution.

CN120023807AInactive Publication Date: 2025-05-23TSINGHUA UNIVERSITY
View PDF 0 Cites 29 Cited by

Patent Information

Application Number
CN202510065041.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing service robots lack deep understanding and dynamic reasoning capabilities when dealing with complex natural language instructions, find it difficult to cope with uncertainty and variability in the actual environment, and cannot actively seek human assistance when the task fails or information is missing, resulting in inefficient task execution.

Method used

Using a multi-agent collaboration architecture, the service robot is empowered with natural language understanding, dynamic task reasoning, autonomous decision-making and human-machine collaboration capabilities by simulating four thinking steps in humans’ problem-solving—perception, planning, execution and reflection. The architecture includes memory module, perceptual agent, planning agent, decision-making agent and reflection agent. Each agent interacts through natural language and shares the information of the memory module.

Benefits of technology

It realizes the natural language understanding, dynamic task reasoning and independent decision-making capabilities of service robots in complex environments, improves the efficiency and adaptability of task execution, and can efficiently handle complex task instructions and dynamic situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120023807A_ABST
    Figure CN120023807A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent cooperative task inference and robot scheduling system, which comprises an interaction unit used for acquiring a natural language instruction issued by a user, sending the natural language instruction to an intelligent processing unit and interacting with the intelligent processing unit, receiving an execution result of a task and body state information of a robot, and sending the received execution result to the intelligent processing unit; sending a task missing information supplementing request when the task information is missing; the intelligent processing unit is used for dynamically sensing an environment and a task state according to a user instruction, a body state and environment information fed back by the robot and stored memory, and performing dynamic planning and decision making based on a sensing result and an inverse result of a task execution result to realize task reasoning and robot scheduling; and the robot cluster executes tasks according to the action sequence and feeds back own body state information and environment information. According to the invention, through a multi-agent cooperation framework based on a large model, the robot is endowed with dynamic reasoning and active interaction capabilities, and the requirements of efficient task reasoning and cooperative execution in a dynamic environment are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of large language models and intelligent robots, and specifically relates to a multi-agent collaborative embodied task reasoning and robot scheduling system and method, which is particularly suitable for application scenarios in which robots and humans collaborate to perform tasks in complex office environments. Background Art

[0002] Existing service robots can usually only execute preset instructions when handling real-world tasks. They lack the ability to deeply understand and dynamically reason about complex natural language instructions, and are unable to cope with tasks that are full of uncertainty and variability in real environments. In addition, when faced with task failure or information loss, traditional service robots are often unable to actively seek human assistance, resulting in inefficient task execution or even complete failure.

[0003] Although large language models have made significant progress in natural language processing in recent years, there are still many challenges in combining them with embodied robots and applying them to real office scenarios. For example, when performing real tasks, existing service robots can usually only follow preset instructions, lack deep understanding of complex natural languages ​​and dynamic reasoning capabilities, and have difficulty coping with uncertainty and variability in the environment. In addition, these robots cannot actively seek human assistance when tasks fail or information is missing, resulting in low execution efficiency or even task failure. Although large language models have made significant progress in natural language processing, combining them with embodied robots and applying them to real office scenarios still faces multiple challenges, including perception and planning of embodied tasks, efficient request mechanisms in human-machine collaboration, and system integration and resource optimization, robot deployment, and other issues. Solving these technical bottlenecks is of great significance for realizing intelligent, adaptable, safe and reliable service robots. Therefore, there is an urgent need for a method that can combine embodied task reasoning with embodied task execution to improve the task reasoning, decision-making and scheduling capabilities of service robots in complex environments and meet practical application needs. Summary of the invention

[0004] The present invention aims to solve one of the technical problems existing in the prior art to at least some extent.

[0005] To this end, the present invention provides a multi-agent collaborative embodied task reasoning and robot scheduling system and method, introduces a multi-agent collaborative architecture based on a large language model, and endows service robots with natural language understanding, dynamic task reasoning, autonomous decision-making and human-computer collaboration capabilities in complex environments by simulating the four thinking steps of humans in dealing with problems - perception, planning, execution and reflection, so that they can efficiently handle complex task instructions and dynamic situations.

[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme:

[0007] The first aspect of the present invention provides a multi-agent collaborative embodied task reasoning and robot scheduling system, comprising:

[0008] An interaction unit, used to obtain a natural language instruction issued by a user and send it to the intelligent processing unit, interact with the intelligent processing unit, receive the execution result of the task and the embodied state information of the robot, and issue a request to supplement the missing task information when the intelligent processing unit determines that the task information is missing;

[0009] The intelligent processing unit is used to dynamically perceive the environment and task status according to the user instructions sent by the interactive unit, the embodied state information and environmental information fed back by the robot, and the memory stored in the intelligent processing unit itself, and to dynamically plan and make decisions based on the perception results and the reflection results of the task execution results to realize task reasoning and robot scheduling;

[0010] The robot cluster executes tasks according to the action sequence issued by the intelligent processing unit, and feeds back its own embodied state information through the interface, so as to help the intelligent processing unit dynamically perceive the environmental state and update the memory of the intelligent processing unit.

[0011] In some embodiments, the intelligent processing unit adopts a multi-agent collaborative architecture based on a large language model, including a memory module, a perception agent, a planning agent, a decision agent, and a reflection agent, and each agent interacts through natural language and shares the memory stored in the memory module;

[0012] The perception agent is used to dynamically perceive the environment and task status according to the user instructions obtained by the interaction unit, combined with the robot's embodied state information and historical task execution logs, and generate a task perception package containing task goals, location information and real-time task execution feedback;

[0013] The planning agent is used to dynamically reason about task decomposition and execution scheme and generate a task plan according to the task perception package and the reflection results provided by the reflection agent;

[0014] The decision-making agent is used to analyze the task plan, generate action sequences, and perform robot scheduling to guide the robot to complete the task;

[0015] The reflective agent is used to evaluate the task execution results. If the task execution results do not meet expectations, it provides feedback for the next round of task reasoning to re-plan or adjust the execution plan.

[0016] In some embodiments, the memory module stores long-term memory and short-term memory;

[0017] The long-term memory includes an environment semantic map, a user default location, static location information of items, location information of public devices, user preferences, and historical task execution logs;

[0018] The short-term memory includes real-time status information and dynamic context information during the current task execution process. The real-time status information includes user instructions, current task execution progress, the robot's embodied status information, dynamic environment information, the reflection results of the reflective agent and user feedback records during the task execution process.

[0019] In some embodiments, the task reasoning modes of the intelligent processing unit include reactive, proactive and proactive;

[0020] Reactive: When the perception agent determines that the task information is complete according to the user instructions sent by the interaction unit and the embodied state fed back by the robot, it extracts the task objectives and operation requirements from the task information and generates a task perception package containing the task objectives, location information and real-time feedback; the planning agent generates the task execution steps according to the task perception package and the reflection results provided by the reflection agent to form a task plan, which should meet the most basic executable conditions; the decision-making agent generates an action sequence according to the task plan issued by the planning agent and matches it with the embodied capabilities of the robot cluster;

[0021] Active: When the perception agent determines that the task information is incomplete according to the user instruction sent by the interaction unit and the embodied state fed back by the robot, it sends a request for supplementary task information to the planning agent for the information gap. The planning agent forms a corresponding task plan based on the request and the reflection result provided by the reflection agent. The decision agent generates an action sequence according to the task plan and obtains the required supplementary information by searching the relevant information in the memory module or sending a user-initiated query request to the instruction through the interaction unit;

[0022] Proactive: When the current task cannot be completed after multiple rounds of active reasoning, the perception agent sends a multi-party assistance request to the planning agent. The planning agent forms a corresponding task plan based on the request and the reflection results provided by the reflective agent. The decision-making agent generates an action sequence based on the task plan and initiates inquiry requests to other users in the environment through the interaction unit to expand the scope of information acquisition and ensure the smooth execution of the task.

[0023] In some embodiments, the method of initiating a query request to other users in the environment through the interaction unit includes:

[0024] Individual contact: Identify potential assisting users and send requests to obtain necessary information. The potential assisting users include persons who are not directly involved in the task or persons whose records do not exist in the memory module;

[0025] Group broadcast: Post assistance requests in the work scene group and coordinate multiple resources to obtain the information required for the task.

[0026] In some embodiments, the robot comprises:

[0027] The robot knowledge container is used to store the environment semantic map, target user location and object coordinate information, and supports real-time updates;

[0028] Robot high-level function container, used to provide function tools for obtaining and controlling the robot's embodied state, including navigation, item interaction, and real-time state feedback;

[0029] A robot execution module, used to execute corresponding tasks according to the action sequence generated by the decision-making agent, including navigation, item interaction and sending online messages;

[0030] The embodied state feedback interface is used to feed back the robot's current physical state and task progress to the perceptual agent to update the task state information.

[0031] In some embodiments, the action sequence generated by the decision-making agent includes online operations and physical actions. The decision-making agent sends action instructions to the robot execution module through an interface and dynamically adjusts the execution plan during task execution.

[0032] In some embodiments, the process of the decision agent scheduling the robot includes:

[0033] The decision-making agent breaks down the task plan output by the planning agent into multiple subtasks based on the function tools available to the robot, and generates a complete action sequence; the decision-making agent selects one or more subtasks that should be executed in the next step from the subtasks based on the robot's load, position and embodied capabilities, determines the most suitable robot through a load balancing strategy, and sends specific action instructions to the corresponding robot execution module using a unified communication protocol; during the process of the robot executing the action instructions, the decision-making agent detects whether the action is proceeding smoothly by continuously receiving the operating status, power, execution progress or error reports provided by each robot through an interface, thereby achieving continuous tracking of the execution status; at the same time, the decision-making agent identifies abnormal situations based on rules or learning models, and once an abnormality is detected, it triggers error correction or interrupts task execution.

[0034] In some embodiments, the reflective agent determines whether the task is successfully completed by analyzing the real-time task execution feedback recorded in the task perception package, and records the execution status to achieve real-time evaluation of the task execution results; if the task is not completed or the execution fails, the reflective agent diagnoses the cause of the failure, provides optimization suggestions to the decision-making agent, and stores the reflection results to the memory module during subsequent information transmission; the reflective agent also optimizes the reasoning and execution process of future tasks through historical reflection results.

[0035] In some embodiments, the reflection results include a judgment of the success or failure of task execution, an analysis of abnormal causes, and improvement suggestions.

[0036] A second aspect of the present invention provides an embodied task reasoning and robot scheduling method, comprising the following steps:

[0037] Obtain natural language instructions issued by the user, as well as the execution results of the task and the robot's embodied state information, and issue a request to supplement the missing task information when it is determined that the task information is missing;

[0038] The environment and task status are dynamically perceived according to the natural language instructions, the embodied state information fed back by the robot, the perceived environmental information, and the stored memory, and dynamic planning and decision-making are performed based on the perception results and the reflection results on the task execution results to achieve task reasoning and robot scheduling, wherein the embodied state information is fed back by the robot cluster after executing the task according to the received total work sequence, and is used to dynamically perceive the environmental status and update the stored memory.

[0039] In some embodiments, the method supports three task interaction modes:

[0040] Reactive interaction: directly execute the natural language instructions with complete information;

[0041] Active interaction: When task information is missing, the user requests the instructions or searches its own stored memory to supplement the missing task information;

[0042] Proactive interaction: When active interaction fails to obtain the missing information of the task, multiple parties collaborate to supplement the missing information of the task by sending requests to other users in the environment, either in a one-to-one contact or by broadcasting in a group, seeking help.

[0043] The characteristics and beneficial effects of the present invention are:

[0044] 1) Through reactive, proactive and active interaction mechanisms, service robots can dynamically reason and efficiently execute complex tasks;

[0045] 2) Combine the embodied state information of the robot feedback module to optimize the task plan in real time and improve task adaptability and execution efficiency;

[0046] 3) Based on the combination of "large-scale training data + high-dimensional parameters + iterative optimization", the large model continuously learns the relationship between language rules and concepts in the layered network structure, thereby obtaining deep natural language understanding and reasoning capabilities. In terms of multi-agent architecture, by dividing the roles of multiple large model agents with different functions or strategies and working together under a shared communication mechanism (such as natural language communication or a unified task protocol), each agent can play the set advantages and obtain a more complete task solution through collaborative reasoning in stages such as "perception-planning-execution-reflection". This division of labor and collaboration helps to "simplify the complex" and apply the powerful language understanding and reasoning capabilities of the large model to multi-step, multi-role task scenarios: the lower-level functional modules such as perception and execution do not need to master all the language and reasoning logic, but are coordinated by the higher-level large language model agents. Ultimately, under the efficient communication and role complementarity of the entire collaborative architecture, dynamic collaboration and deep reasoning of complex tasks can be achieved, so that the system has better adaptability and problem-solving capabilities.

[0047] Based on the above analysis, the multi-agent collaboration framework proposed in this invention has deep natural language understanding and dynamic human-machine collaboration capabilities, and can actively initiate information requests or collaborate with users to solve problems during execution;

[0048] 4) Introducing the reflection mechanism. Introducing the "Reflection Agent" in the multi-agent collaborative framework can be regarded as a meta-reasoning mechanism, which is connected in series with the specific perception, planning and execution processes, and is specifically responsible for monitoring, evaluating and dynamically optimizing the overall task execution process of the system, thereby ensuring the accuracy of the task results;

[0049] 5) Suitable for complex dynamic environments, and able to provide efficient, intelligent and humanized services in real office scenarios.

[0050] In summary, the present invention proposes a multi-agent collaborative embodied task reasoning and robot scheduling system and method, which utilizes the powerful natural language understanding and reasoning capabilities of a large language model to efficiently process complex task instructions and dynamic situations. Multi-agent collaboration enhances the flexibility and task decomposition capabilities of the system, allowing each agent to communicate in human language and collaborate to complete complex tasks, which not only increases the interpretability of the algorithm, but also improves the overall execution efficiency. By coordinating large language model agents, manual and robots, the framework achieves seamless information sharing and real-time decision-making, ensuring the intelligence and adaptability of task execution in environments such as offices. This integration method not only improves the adaptability and robustness of the system, but also significantly enhances the effect of human-computer collaboration. It is suitable for variable and complex actual application environments, and achieves efficient execution of embodied tasks and robot scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a schematic diagram of the overall framework of the multi-agent collaborative embodied task reasoning and robot scheduling system provided in the embodiment of the first aspect of the present invention.

[0052] Figure 2 It is a flowchart of the internal collaborative reasoning of multiple agents based on a large model in an embodiment of the present invention.

[0053] Figure 3 It is a task execution flow chart of the intelligent agent processing unit in the reactive, proactive and proactive modes in a specific embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] On the contrary, the present application covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present application as defined by the claims. Further, in order to make the public have a better understanding of the present application, some specific details are described in detail in the detailed description of the present application below. Those skilled in the art can fully understand the present application without the description of these details.

[0056] The present invention provides a system and method for embodied task reasoning and robot scheduling based on large-scale multi-agent collaboration, which realizes efficient execution of complex tasks and real-time scheduling of robot embodied states through multi-agent collaboration and dynamic reasoning. The following is a detailed description of the execution process and method of each module of the system in combination with implementation examples.

[0057] See also Figure 1 , Figure 2 The first embodiment of the present invention proposes a multi-agent collaborative embodied task reasoning and robot scheduling system, including an interconnected interactive unit 100, an intelligent processing unit 200 and an embodied robot cluster 300, wherein the embodied robot cluster 300 contains multiple robots 310 that perform tasks. Each unit in the system works together to ensure that the robot can adapt to dynamic and changing task requirements through three interactive modes: reactive, proactive and proactive. Among them:

[0058] The interaction unit 100 is used to obtain the natural language instructions issued by the user and send them to the intelligent processing unit 200; the interaction module 100 is also used to interact with the intelligent processing unit 200, receive the execution results of the task and the embodied state information of the robot 310, and provide measures to send a request to the user for supplementary key task information when the task information is missing;

[0059] The intelligent processing unit 200 is used to dynamically perceive the environment and task status according to the user instructions sent by the interaction unit 100, the embodied state information and environmental information fed back by the robot, and the memory stored in the intelligent processing unit 100 itself, and to dynamically plan and make decisions based on the perception results and the reflection results of the task execution results, so as to realize task reasoning and robot scheduling;

[0060] The embodied robot cluster 300 executes tasks according to the action sequence issued by the intelligent processing unit 200, and feeds back its own embodied state through the interface, helping the intelligent processing unit 200 to dynamically perceive the environmental state and update the memory of the intelligent processing unit 200.

[0061] In some embodiments, the interaction unit 100 interacts with the user through a commonly used social platform (such as WeChat), receives natural language instructions issued by the user, feedbacks the task execution results, and actively sends a request for supplementary information to the user when the task information is missing. The natural language instructions issued by the user can be text instructions or voice instructions, and the interaction unit 100 recognizes the instructions and transmits them to the intelligent processing unit 200.

[0062] In some embodiments, the intelligent processing unit 200 adopts a multi-agent collaboration architecture based on a large language model, including a memory module 210, a perception agent 220, a planning agent 230, a decision agent 240, and a reflection agent 250, and performs task reasoning and action sequence generation through multi-agent collaboration, and each agent interacts through natural language;

[0063] The memory module 210 includes a long-term memory storage 211 for storing semantic maps, user locations and item information, and a short-term memory storage 212 for storing real-time status information of task execution. The memory module 210 is shared among the agents;

[0064] The perception agent 220 is used to dynamically perceive the environment and task status according to the natural language instructions issued by the user obtained by the interaction unit 100, combined with the robot's embodied state information and historical task execution logs, and generate a task perception package containing task goals, location information and real-time feedback;

[0065] The planning agent 230 is used to dynamically reason about task decomposition and execution schemes and generate a task plan based on the task perception package and the reflection results provided by the reflection agent 250;

[0066] The decision agent 240 is used to analyze the task plan, generate an action sequence including specific online operations and physical actions, and perform robot scheduling to guide the robot 310 to complete the task;

[0067] The reflective agent 250 is used to evaluate the task execution results (including the reasoning content of each agent and the results of the robot's task execution). If the task execution results do not meet expectations, it provides feedback for the next round of task reasoning to re-plan or adjust the execution plan.

[0068] In some embodiments, each robot 310 includes:

[0069] The robot knowledge container 311 is used to store the environment semantic map, the target user location and the object coordinate information, and supports real-time update;

[0070] A robot high-level function container 312 is used to provide function tools for obtaining and controlling the embodied state of the robot 310, including navigation, item interaction, and real-time state feedback;

[0071] The robot execution module 313 is used to execute corresponding tasks according to the action sequence generated by the decision-making agent 240, such as navigation, object interaction, and sending online messages;

[0072] The embodied state feedback interface 314 is used to feed back the current physical state (such as current position, power status, etc.) and task progress of the robot 310 to the perception agent 220, helping the perception agent 220 to dynamically perceive the embodied state of the robot, and provide real-time data support for task reasoning and planning, thereby updating the task status information.

[0073] In some embodiments, the intelligent processing unit 200 is the core of the present application. The unit generates action sequences and performs dynamic scheduling based on the collaborative interaction between the perception agent 220, the planning agent 230, the decision agent 240, and the reflection agent 250. Through multimodal information fusion, multi-agent collaboration, and real-time human-computer interaction, the unit performs comprehensive analysis and reasoning on the user's natural language instructions, the robot's embodied state, and the environmental perception data, thereby providing the system with efficient and flexible task execution capabilities. The following is a detailed description of each component in the intelligent processing unit 200.

[0074] 1. Memory module

[0075] The memory module 210 is the data support core of the entire system, and is composed of a long-term memory storage 211 and a short-term memory storage 212. The memory module 210 supports real-time updates. If there are environmental changes during task execution (such as the location of an object being moved), the perception agent 220 can perceive the latest environmental information to ensure that the system has environmental adaptability and memory retention capabilities during task execution, providing information support for task reasoning and decision-making.

[0076] Long-term memory storage 211: used to store the relatively stable environment semantic map in the system, the user's default location and the static location information of the items, the location information of the public equipment, the user's preferences and the task execution log, etc. The long-term memory storage 211 performs persistent storage of spatial semantic information based on the key-value database (such as Redis) and the graph database (such as Neo4j), and combines the spatial index algorithm to accelerate retrieval.

[0077] Short-term memory storage 212: used to dynamically store the real-time status and dynamic context information during the execution of the current task, including user instructions, the progress of the current task execution, the robot's embodied state, dynamic environmental information, the reflection results of the reflective agent 250, and human feedback records during the task execution. The short-term memory storage 212 uses a memory database and a memory cache mechanism (such as a Key-Value memory model) to achieve fast reading and writing. Unlike the long-term memory storage 211, the information stored in the short-term memory storage 212 will be cleared or transferred to the historical record after the task is completed to maintain the system's available memory space.

[0078] 2. Perception Agent

[0079] The perception agent 220 is used to receive the user's natural language instructions and the robot's current state information, and to form a task perception package through parsing and integration as the input of the planning agent 230. The perception process of the perception agent 220 is divided into three modes: reactive perception, proactive perception, and proactive perception, in order to respond to different user instructions:

[0080] Reactive perception: The perception agent 220 receives the natural language instructions of the user sent by the interaction unit 100 and the current embodied state of the robot fed back by the embodied state feedback interface 314. When the perception agent 220 determines that the task information is complete (the received explicit user instructions are clear and the environmental information is complete), it directly extracts the task objective and operation requirements from the task information, and directly generates a task perception package in combination with the current embodied state of the robot and the environmental information.

[0081] Proactive perception: When the perception agent 220 determines that the task information is incomplete, the perception agent 220 sends a supplementary task information request to the planning agent 230 according to the missing information. The planning agent 230 generates a corresponding task plan according to the request and the reflection result provided by the reflection agent 250, and the decision-making agent 240 generates a specific action sequence. The perception agent 220 actively checks the memory module 210 and / or actively asks the user who issued the instruction to obtain the required supplementary information based on the action sequence, and generates a task perception package based on the obtained task information in combination with the current embodied state of the robot and the environmental information.

[0082] Prospective perception: In complex tasks, the perception agent 220 combines the physical state information of the memory module 210 and the embodied state feedback interface of the robot to dynamically perceive the current embodied state. After the proactive perception and planning fail, the perception agent 230 sends a multi-party assistance request to the planning agent 230. The planning agent 230 generates a corresponding task plan according to the request and the reflection result provided by the reflection agent 250, and the decision-making agent 240 generates a specific action sequence. The perception agent 220 seeks assistance from other people or groups in the environment according to the action sequence to obtain auxiliary information to supplement the task information, providing a broader planning space for subsequent decision-making. The perception agent 220 generates a task perception package based on the obtained task information in combination with the current embodied state of the robot and the environmental information.

[0083] In this embodiment, the specific steps for the perception agent 220 to generate a task perception package include:

[0084] Step 1) Input data collection:

[0085] The input data collected by the perception agent 220 include user natural language instructions, the robot's embodied state and environmental information; wherein the user natural language instructions are text instructions transmitted by the interaction unit 100, which can be instructions in the form of text originally issued by the user, or can be the recognition result of voice instructions obtained by converting the voice instructions issued by the user through a voice-to-text model; the robot's embodied state includes the robot's real-time position information, power status and fault report transmitted by the robot's embodied state feedback interface; environmental information includes sensor data and historical records, sensor data includes images collected by the robot's camera and radar data collected in real time by the robot's range-finding radar (such as a laser radar), and historical records include long-term memory and short-term memory stored in the memory module 210, such as user preferences, target object coordinates, task execution logs, task completion status, abnormal status reports and other information. The perception agent 220 stores the above-collected real-time data in the short-term memory storage 212 of the memory module 210.

[0086] Step 2) Mode determination and information gap detection:

[0087] First, based on the pre-trained large language model GPT4o, the user's input command is recognized for intent, key elements are extracted (including location and item name, etc.), and context is understood to generate preliminary command parsing results.

[0088] Subsequently, the perception agent 220 determines whether the task information composed of the instruction parsing result and the environmental information is complete, wherein the integrity judgment of the instruction parsing result is to determine whether the instruction parsing result has a clear target object and operation requirements. For example, the instruction parsing result "Please give me the file on the table A in the office" has a clear target object, execution action and target position, which is a complete instruction parsing result; when the perception agent 220 determines that the task information is incomplete (such as the target position is unknown, the target object attributes are unclear, or there is a conflict in key elements), the active perception mode is first triggered, and a request for supplementary task information is sent to the planning agent 230 for the information gap. The planning agent 230 generates a corresponding task plan based on the request and the reflection result provided by the reflection agent 250, and the decision-making agent 240 generates a specific action sequence based on the task plan. In response to the action sequence, the perception agent 220 actively searches for relevant information in the memory module 210 or actively sends a request to the instruction agent 230 through the interaction unit 100. The user who issues the command initiates an inquiry request (the perception agent 220 gives priority to searching the memory module 210. If the missing information cannot be obtained after the search, the interaction unit 100 initiates an inquiry request to the user who issues the command) to obtain the required supplementary information; if the task decomposition cannot be completed after multiple active perceptions or planning corrections, the perception agent 220 triggers the forward-looking perception mode and sends a multi-party assistance request to the planning agent 230. The planning agent 230 generates a corresponding task plan based on the request and the reflection result provided by the reflection agent 250. The decision-making agent 240 generates a specific action sequence based on the task plan. The perception agent 220 responds to the action sequence to further expand the information source, including calling more online resources, contacts or groups, obtaining richer contextual support, and generating auxiliary information. Once the perception agent 220 determines that the task information is complete, the reactive perception mode is triggered to directly extract the task objectives and operation requirements from the complete task information to generate a preliminary perception package.

[0089] Step 3) Information integration:

[0090] The perception agent 220 integrates the generated preliminary perception package with the robot's embodied state and environmental information into a task perception package, which is used to provide a directly operable structured description for the planning agent 230.

[0091] Step 4) Output and callback:

[0092] The perception agent 220 transmits the task-based perception package to the planning agent 230 and keeps listening in the subsequent process. Once new information or user instruction updates are detected during the robot 310's task execution phase, the task perception package is supplemented or corrected in real time.

[0093] 3. Planning Agent

[0094] The planning agent 230 uses the multi-step reasoning capability of the large language model and combines the long-term and short-term memory of the memory module 210 to analyze the user intention and environmental information in the task perception package and generate a specific task plan. The reasoning process of the planning agent 230 includes the following three modes, which respectively deal with different task complexities and execution states.

[0095] Reactive planning: When the planning agent 230 determines that the task perception package meets the most basic executable conditions, the planning agent 230 quickly generates task execution steps based on the task perception package to form a task plan, ensuring that the task objectives and paths are clear, and passes the task plan to the decision-making agent 240 to directly guide the robot 310 to perform the task.

[0096] Active planning: When the planning agent 230 receives a request for supplementary task information sent by the perception agent 220, the planning agent 230 generates a task plan that spontaneously searches for relevant information in the memory module 210 or sends a user request for information supplement to the instruction. During the task execution process, the planning agent 230 can dynamically adjust the task plan by combining the reflection results of the reflection agent 250 and the task perception package fed back by the perception agent 220 to ensure the smooth progress of the task.

[0097] Forward-looking planning: When proactive information supplementation fails, the planning agent 230 can foresee that continuing to generate a task plan will lead to task execution failure. When the existing information is insufficient to support the generation of a feasible plan, the planning agent 230 will proactively start the forward-looking planning mode, and generate a task plan to request collaboration or help from other human team members in the environment based on the received task perception package and the reflection results provided by the reflection agent 250, and obtain the necessary information, thereby ensuring that the task can proceed smoothly and improving the success rate of the execution of user instructions. Specifically including:

[0098] 1) Individual contact: accurately locate potential assisting personnel and send a request to obtain necessary information; wherein, potential assisting personnel refers to personnel who are not directly involved in the task and personnel for whom there is no record in the memory module 210.

[0099] 2) Group broadcast: Post assistance requests in the work scene group and coordinate multiple resources to obtain the information required for the task.

[0100] Through real-time human-machine collaboration and physical state feedback, the planning agent 230 has the ability to adaptively correct the task plan and dynamically adjust the task execution path, thereby improving the reliability and efficiency of the system in complex task environments.

[0101] In this embodiment, the specific steps of the planning agent 230 generating a task plan include:

[0102] Step 1) Input collection and initialization:

[0103] The planning agent 230 obtains the task perception package from the perception agent 220, calls the long-term memory and short-term memory stored in the memory module 210 for verification, and confirms the current available resources (such as available resources for robots when they are currently idle) and prior knowledge (such as the location information of people and objects involved in the current task).

[0104] Step 2) Mode determination and task plan generation:

[0105] When the planning agent 230 receives the task perception package sent by the perception agent 220 in the reactive mode, the reactive planning mode is triggered. The planning agent 230 uses the multi-step reasoning capability of the pre-trained large language model GPT4o to realize the semantic decomposition and sub-task generation of the instruction parsing results in the task perception package. The large language model uses the prompt engineering (PE) and chain-of-thought mechanism to enable the model to explain why a certain sub-task strategy is adopted to improve interpretability. The planning agent 230 passes the task plan (including the task execution process, the personnel and resource requirements and paths involved, etc.) output by the large language model to the decision agent 240 for the decision agent 240 to generate the next executable action. Among them, the task plan generated by the planning agent 230 should meet the most basic executable conditions, that is, it should meet the requirements of completeness, consistency, executableness and unambiguity at the same time. Among them, completeness means that the generated task plan contains the necessary elements to complete the task, including the target object, target location and operation, and each necessary element should be clear and unambiguous; specifically, for the target object or target location, the task plan clearly states the target object or target location to be operated or traveled to (such as "files in office A"); for the operation, the task plan specifies the core operations that the robot needs to perform (such as "pick up", "move to", "hand it to someone"). Consistency means that the task objectives and environmental information in the task plan are consistent with the current resource availability, and there are no major conflicts (such as no obstacles or robot failures). Executability means that the current task can be completed in theory under the given environment (i.e., the environmental information perceived by the perceptual agent), robot capabilities and time requirements. Unambiguity means that there are no multiple meanings or ambiguities in the instruction parsing results and context, which makes the robot unable to determine the specific operation object or process; for example, when the instruction parsing results involve previous steps or prerequisites (such as permissions, item location), it is ensured that they have been met or do not need to be confirmed again, then it is judged that unambiguity is met.

[0106] When the planning agent 230 receives a request for supplementary task information from the perception agent 220, the active planning mode is triggered. The planning agent 230 generates a task plan by spontaneously searching for relevant information in the memory module 210 or issuing a user-initiated inquiry request to the instruction through the interactive unit 100 in combination with the reflection results provided by the reflective agent 250. After the missing information is supplemented, the task perception package newly generated by the perception agent and the reflection results provided by the reflective agent 250 are used to form a task plan using multi-step reasoning based on a large language model, and then the task plan is passed to the decision-making agent 240.

[0107] When the planning agent 230 receives a multi-party assistance request sent by the perception agent 220 or the task perception package still cannot meet the most basic executable conditions after multiple rounds of active planning, and it is predicted that continuing to generate the task plan according to the current situation will lead to the failure of task execution, the forward-looking planning mode is actively triggered. The planning agent 230 generates a task plan to request cooperation or help from other human team members in the environment (including online groups and dedicated collaboration) based on the received task perception package and the reflection results provided by the reflection agent 250, so as to obtain a wider range of information sources. After the task perception package meets the most basic executable conditions, the multi-step reasoning based on the large language model is used to form a task plan, and the task plan is passed to the decision-making agent 240. Among them, when any one or more of the following situations occur again during the task execution or planning process and cannot be resolved through further collaboration or external help, it is determined that continuing to generate the task plan according to the current situation will lead to the failure of task execution:

[0108] ① Information is missing or conflicting and difficult to make up for: Actively sending instructions to the user or memory module 210 times to request information still cannot fill the key gap, or there is conflicting environmental information and it cannot be resolved;

[0109] ② Multiple plans cannot be implemented: After trying multiple alternative task decomposition solutions, it is still impossible to generate an executable action sequence (for example, the path is completely blocked by obstacles);

[0110] ③ Resource or physical limitations: The robot is low on power and has no available charging solution, the robot arm is damaged and has no way to repair it; or the external equipment required for the task is unavailable;

[0111] ④ Time or priority conflict: The task must be completed within the specified time, but during planning it is found that the time limit cannot be met, or other higher priority tasks occupy resources, causing the current task to be shelved and unable to be completed.

[0112] In this embodiment, the interaction and collaboration between the planning agent 230 and other agents include:

[0113] With the perception agent 220: obtain updated task perception packages in real time; if the planning encounters a bottleneck, request the perception agent 220 to query more environment, sensor or social platform data, and the perception agent 220 stores the queried data in the memory module 210 for the planning agent 230 to call, so as to ensure that the planning agent 230 obtains the latest situation;

[0114] With the decision-making agent 240: the generated task plan is handed over to the decision-making agent 240 to decide the specific robot action instructions for the next step, and monitor the execution feedback;

[0115] With the reflective agent 250: When errors or uncertainties occur while the robot is performing a task, the reflective agent 250 will make correction suggestions or re-planning requests, and the planning agent 230 will adjust the task plan or re-plan accordingly.

[0116] 4. Decision-making Agent

[0117] The decision agent 240 converts the task plan generated by the planning agent 230 into a specific action sequence that can be executed by the robot, including two types of online operations and physical actions. The decision agent 240 has three decision modes, and the corresponding decision mode is selected according to the planning mode adopted by the planning agent 230:

[0118] Reactive decision-making: When the planning agent 230 adopts the reactive planning mode, the decision-making agent 240 selects the reactive decision-making mode, directly decomposes the task plan into sub-tasks, generates action sequences, and allocates them to the most suitable one or more robots 310 to guide the robots 310 to complete the task.

[0119] Active decision-making: When the planning agent 230 adopts the active planning mode, the decision-making agent 240 selects the active decision-making mode, generates a search memory module 210 according to the task plan generated by the planning agent 230, or sends an action sequence of the user's request for supplementary task information to the instruction through the interactive unit 100 to obtain the missing task information. The perception agent 220, the planning agent 230 and the decision-making agent 240 collaborate to perform perception, planning and decision-making in turn, thereby ensuring the smooth execution of the user's instructions.

[0120] Proactive decision-making: When the planning agent 230 adopts the proactive planning mode, the decision agent 240 selects the proactive decision-making mode and generates the decision agent 240 based on the task plan generated by the planning agent 230. The decision agent 240 combines the task perception package with proactive planning, not only generates tasks such as robot actions or online notifications based on the established plan, but also actively cooperates with the planning agent 230 to conduct online inquiries through single chat or group chat, and proactively obtains additional information to optimize subsequent decisions. By dynamically adjusting the task execution sequence and strategy, the decision agent 240 proactively avoids the generation of erroneous actions, improves the success rate of user instruction completion, and ensures the efficiency and accuracy of task execution.

[0121] In this embodiment, the specific decision-making process of the decision agent 240 includes:

[0122] Step 1) Input collection and initialization:

[0123] The decision agent 240 obtains the task plan including the task execution process, the personnel involved, and the resource requirements from the planning agent 230 .

[0124] Step 2) Mode determination and action sequence generation and collaborative allocation:

[0125] The decision agent 240 triggers its own decision mode according to the attributes of the task plan generated by the planning agent 240.

[0126] Reactive decision-making: If the task plan is a complete execution plan for the original instruction, the reactive decision-making mode is triggered to generate an action sequence and perform specific sub-task allocation. Specifically: according to the function tools available to the robot 310, the task plan (corresponding to the high-level task steps) output by the planning agent 230 is decomposed into multiple sub-tasks using the reasoning ability of the large language model (specifically, the sub-task decomposition can be achieved through thinking chains and prompt words plus a few sample decomposition examples) and the sub-tasks use action instructions or function interfaces that can be directly called by the robot 310, such as "move to coordinates (x, y)", "grab item A", "talk to a user", API calls, etc.; templates are established for basic action instructions (such as "grab", "move", etc.), and a complete action sequence is generated by filling parameters (target coordinates, target objects, etc.) into the template. Subsequently, for multi-robot scenarios, the decision-making agent 240 extracts one or more subtasks that should be executed in the nearest step from the refined subtasks based on indicators such as the load, position and embodied capabilities of the robot 310, determines the most suitable robot through a load balancing strategy based on the Hungarian Algorithm, and sends specific action instructions to the corresponding robot execution module 313 using a unified ROS Topic communication protocol; during the robot's execution of the action instruction, the decision-making agent 240 detects whether the action is proceeding smoothly by continuously receiving the operating status, power, execution progress or error report provided by each robot 310 through the embodied state feedback interface 314, thereby achieving continuous tracking of the execution status; at the same time, the decision-making agent 240 quickly identifies abnormal situations (such as mechanical failures or unexpected path blockages) based on rules or learning models, and once an abnormality is detected, it triggers error correction or interrupts the process.

[0127] Active decision-making: If the task plan is generated based on active planning, that is, the planning agent 230 finds that the task information is missing when generating the task, the decision-making agent 250 enters the active decision-making mode. In this mode, the robot will not be deployed directly. Instead, the user who issues the instruction will initiate an active inquiry for the active task plan. Specifically, online operations can be generated by calling the robot's (specifically a smart phone) information sending method, or by searching the memory module 210 to obtain the missing information.

[0128] Proactive decision-making: If the planning agent 230 enters the proactive planning mode, that is, when the proactive information supplement fails, the planning agent 230 will issue a task plan that requires initiating a single chat or a group chat, so as to obtain a wider range of assistance information to promote the smooth execution of the instructions. At this time, the decision-making agent 250 enters the proactive decision-making mode. In this mode, the robot will not be directly deployed. Instead, the robot (specifically a smart phone) will be called to generate an online operation for the proactive task plan. According to the user or group chat name given in the task plan, the corresponding online request for assistance message is initiated.

[0129] In this embodiment, the interaction and collaboration between the decision agent 240 and other agents include:

[0130] Interaction with the perception agent 220: Real-time acquisition of the latest task perception package to ensure that decisions are made based on current environmental information for task decomposition and allocation; when the decision-making process encounters a bottleneck, the perception agent 220 can be requested to further query the environment, sensor or social platform data, and the query results can be stored in the memory module 210 for the decision-making agent 240 to call to optimize the task decomposition and allocation process;

[0131] Collaboration with the planning agent 230: Based on the task plan provided by the planning agent 230, extract the specific tasks that need to be performed at the current moment;

[0132] Interaction with the reflective agent 250: After the task is completed, the decision-making agent 240 submits the decomposition and allocation results to the reflective agent 250, assists the reflective agent 250 in analyzing the execution effect and potential problems, generates reflection results, and integrates optimization suggestions into the decomposition and allocation process of future tasks.

[0133] 5. Reflective Agents

[0134] The reflective agent 250 evaluates the task execution results based on the execution log, task completion status, abnormal status report and other information obtained from the task perception package of the perception agent 220, and provides the reflection results for system optimization to ensure the accuracy and completeness of the task. The specific steps include:

[0135] Task evaluation: The reflective agent 250 analyzes the results of the robot execution module feedback recorded in the task perception package, determines whether the task is successfully completed, and records the execution status.

[0136] Problem diagnosis: If the task is not completed or fails to execute, the reflective agent diagnoses the cause of the failure, provides optimization suggestions to the decision-making agent, and stores the reflection results in the memory unit during the subsequent information transmission process.

[0137] Process optimization: Through historical reflection data, optimize the reasoning and execution process of future tasks and improve the success rate of tasks.

[0138] Step 1) Input collection and initialization:

[0139] The reflective agent 250 obtains the current task execution results from the task perception package of the perception agent 220, including execution logs, task completion status, abnormal status reports and other information, and integrates them; the reflective agent 250 also calls long-term and short-term memory by accessing the memory module 210 to obtain the preset task goals and possible related historical execution records (such as the completion status of similar tasks in the past). Finally, these two parts form a complete reflective closed-loop information chain for reflecting on the current step. Through unified timestamp management during the reflection process, the reflective agent 250 can accurately find the sequence and causal relationship of each step in the task execution, thereby improving the diagnostic efficiency.

[0140] Step 2) Task Evaluation:

[0141] The reflective agent 250 compares the current task execution result with the preset task goal. If the current task execution result meets or exceeds the preset task goal, the task is judged to be successful and the execution details (such as time, resource consumption, etc.) are recorded in the task execution log. If part of the execution result of the current task does not meet the preset task goal, such as missing steps, incorrect object coordinates, unsatisfactory user feedback, etc., the task is judged to be deviation or failure, and the next step of the problem diagnosis process is entered.

[0142] Step 3) Problem diagnosis:

[0143] First, the reflection agent 250 retrieves the task perception package generated by the perception agent 220, the task plan generated by the planning agent 230, and the action sequence generated by the decision agent 240 as collected evidence, and comprehensively compares the environmental information before and after the robot executes the action sequence;

[0144] Subsequently, the reflective agent 250 utilizes the chain reasoning of the large language model, such as textual description of the cause of the error and the error correction strategy, so that the reflective agent can output the judgment process and improvement plan in natural language to improve interpretability, or rule-based reasoning to locate the specific links where the task deviation or failure occurs and infer the cause. Among them, the inferred causes of task deviation include: any one or more of incomplete perception data, planning logic conflicts, decision-making instruction errors, robot hardware failures and temporary changes in user needs.

[0145] Finally, the reflective agent 250 describes the problem diagnosis results in text form, outputs the problem diagnosis results in natural language, improves interpretability, and enters the next step of the process.

[0146] Step 4) Feedback generation and process optimization:

[0147] If the diagnosed problem can be fixed locally (such as simply re-grasping or replacing the robot), the reflective agent 250 sends a local error correction instruction to the decision-making agent 240; if the diagnosed problem involves the overall reconstruction of the task, the reflective agent 250 passes the failure cause and the need for additional information to the planning agent 240; if the diagnosed problem is caused by the lack of key information or requires user decision-making, the reflective agent 250 passes the information that it needs to call the interactive unit 100 to request user intervention (clarify requirements, provide additional resources, etc.) to the planning agent 230, and the planning agent 230 generates a corresponding task plan, and the decision-making agent 240 generates a specific online inquiry action sequence, and re-perceives, plans and makes decisions after the robot responds.

[0148] In some embodiments, the present invention achieves deep collaboration with humans through three task execution modes: reactive, proactive, and active:

[0149] 1) Responsive interaction: When the user's instructions are clear and the information is complete, the system directly executes the task and provides feedback, thus improving the efficiency of task execution.

[0150] 2) Active interaction: When task information is incomplete or execution is blocked, the system initiates inquiries to the user or team members through the interaction unit 100 to obtain missing information or request assistance, and dynamically updates the task execution plan.

[0151] 3) Proactive interaction: In complex tasks, the system combines embodied state feedback information and environmental perception to proactively discover potential problems or predictively determine future resource needs, and collaborates with users or team members to optimize execution plans to ensure smooth completion of tasks.

[0152] See also Figure 1 , 2 The second aspect of the present invention provides an embodied task reasoning and robot scheduling method based on the above system, including:

[0153] Obtain natural language instructions issued by the user, as well as the execution results of the task and the robot's embodied state information, and issue a request to supplement the missing task information when it is determined that the task information is missing;

[0154] The system dynamically perceives the environment and task status based on the acquired natural language instructions, the embodied state information fed back by the robot, the perceived environmental information, and the stored memory. It makes dynamic planning and decisions based on the perceived results and the reflection results on the task execution results to achieve task reasoning and robot scheduling. The embodied state information is fed back by the robot cluster after executing the task according to the received total work sequence. It is used to dynamically perceive the environmental status and update the stored memory.

[0155] Furthermore, the method of the embodiment of the present invention specifically includes the following steps:

[0156] a) Receiving user instructions: receiving the user's natural language instructions through the interaction unit 100, and passing the instructions to the intelligent processing unit 200;

[0157] b) Generate task perception package: The perception agent 220 in the intelligent processing unit 200 parses the user's natural language instructions and generates a task perception package by combining long-term memory, environmental information and the robot's embodied state information;

[0158] c) Task reasoning and planning: For the first time receiving instructions from the user, the planning agent 230 will directly reason on the task perception package, decompose the task and generate a task plan without combining the reflection results; for the non-first time receiving instructions from the user, the planning agent 230 will generate a task plan based on the task perception package and the reflection results provided by the reflection agent 250, and dynamically adjust the task plan according to the environmental information fed back by the perception agent 220 and the robot's embodied state information during the task execution;

[0159] d) Task decision and interaction: The decision agent 240 generates an executable action sequence based on the task plan generated by the planning agent 230, including specific online tasks and physical tasks, and sends it to the corresponding robot execution module. When information is missing, it actively interacts with the instruction issuing user to supplement the required information;

[0160] e) Task execution and status feedback: The robot execution module receives the action sequence issued by the decision agent 240, executes the task through high-level functions, and feeds back the current embodied state to the perception agent 220 in real time. The perception agent 220 re-perceives the state change information to generate a new task perception package and updates the memory module 210;

[0161] f) Reflection and optimization: The reflective agent 250 evaluates the task execution results based on the dynamic environment information in the task perception packages at two moments. If the results do not meet expectations, it will perform problem diagnosis and process optimization to form a reflective result. The planning agent and decision-making agent will re-plan and make decisions based on the reflective result to ensure that the task is completed. After the task is completed, the system updates the memory module information to provide a reference for future task execution.

[0162] Combine the following Figure 3 The embodiments of the present invention are described with reference to the accompanying drawings and specific task execution examples.

[0163] Sample task: "Zhang San: Take this document to Li Si for signature, and then give it to Wang Wu."

[0164] Deployable embodied robots: a mobile chassis robot with a smart lock (hereinafter named "Robot No. 1"), a robot dog (hereinafter named "Robot No. 2"), a robotic arm (hereinafter named "Robot No. 3"), and a smart phone (with WeChat and other chat software installed and accounts configured).

[0165] For the above example tasks, the task reasoning and robot scheduling process of each agent in the embodiment of the present invention is as follows:

[0166] Time step one:

[0167] 1. Perception Agent:

[0168] The perceptual agent 220 senses the appearance of a new user instruction in WeChat, receives the user instruction, and parses out preliminary information: "The file needs to be obtained from Zhang San and sent to Li Si first, and finally to Wang Wu."

[0169] The perception agent 220 calls the memory module 210 to query the location information of objects, Zhang San, Li Si, Wang Wu and others in the environment as well as the user's intention.

[0170] The perception agent 220 generates a task perception package using the queried information, the original user instructions, and the perception summary information, and passes it to the reflection agent 250 (if it is the first execution, the reflection agent 250 is skipped and directly passed to the planning agent 230).

[0171] 2. Planning Agent:

[0172] The planning agent 230 decomposes the original user instructions in the task perception package into three steps:

[0173] ① Get the file from Zhang San first;

[0174] ② Go to Li Si's location again and send him a signature;

[0175] ③Finally, go to Wang Wu’s position and hand it over to him.

[0176] The planning agent 230 generates a complete task plan based on the above task decomposition results. If important information constituting the above task plan is missing at this time, it starts to generate a backup task plan.

[0177] 3. Decision-making Agent:

[0178] The decision agent 240 compares the task plan generated by the planning agent 230 with the capabilities of different robots in the embodied robot cluster and generates a specific action sequence by reasoning:

[0179] ①Physical action: <Robot No. 1: Move to Zhang San>.

[0180] ②Online operation: <Mobile phone: send a notification and a QR code for unlocking to Zhang San>.

[0181] The decision agent 240 calls the robot high-level function container 312 according to the generated action sequence, controls robot No. 1 to navigate to the first target position (i.e. where Zhang San is), and waits for Zhang San to complete the operation of putting the item in.

[0182] Robot No. 1 transmits its embodied feedback information to the perception agent 220, which perceives new environmental information and embodied status, such as the placement of items.

[0183] Time step two:

[0184] 1. Perception Agent:

[0185] The perception agent 220 perceives that Robot No. 1 has completed the first step of the task and successfully obtained the file.

[0186] The memory module 210 queries the location of Li Si, and updates the execution status of the previous task into the task perception package, which is updated as follows:

[0187] Current status: The file has been obtained, and Robot No. 1 is at Zhang San's position.

[0188] The perception agent 220 passes the updated task perception package to the planning agent 230.

[0189] 2. Reflective Agents:

[0190] The reflective agent 250 evaluates the task execution result based on the generated plan, the specific task generated by the decision-making agent 240 in the previous time step, and the task perception package of the two state changes before and after the task execution, and confirms that Robot No. 1 has obtained the file; if it fails midway (such as Zhang San did not put in the item), the reflective agent 250 will reflect on the reasons for the failure of task execution.

[0191] After finishing thinking, the reflective agent 250 will pass its reflection results to the downstream planning agent 230 and decision-making agent 240 to guide them to carry out the next step of planning and decision-making, that is, to continue to execute the specific tasks corresponding to the remaining part of the user's instructions, or to re-execute the specific tasks that failed in the previous step.

[0192] 3. Planning Agent:

[0193] If the execution result of the previous step is correct after reflection by the reflection agent 250, the planning agent 230 decomposes the unfinished part into steps according to the user instructions and the content of the task perception package:

[0194] ① Take the document to Li Si's location and ask him to sign;

[0195] ②Go to Wang Wu’s position and give it to him.

[0196] If the result of the previous step is reflected by the reflection agent 250 and it is found that the execution is wrong, the planning agent 230 will generate a supplementary plan based on the reflection result. Assuming that the reason for the error in task execution is that Zhang San did not scan the code to put the file into robot No. 1, the planning agent 230 updates the task plan as follows:

[0197] Send a confirmation message to Zhang San, asking him whether he has put the file into the robot locker;

[0198] 4. Decision-making Agent:

[0199] The decision agent 240 generates a specific action sequence based on the task plan issued by the planning agent 230 and matching it with the robot cluster capabilities:

[0200] Physical action: <Robot No. 1: Move to Li Si's position>.

[0201] Online operation: <Smartphone: send a notification to Li Si, reminding him to prepare to sign>.

[0202] The decision agent 240 calls the high-level function container 312 of Robot No. 1 according to the generated action sequence, controls Robot No. 1 to navigate to the location of Li Si, and sends a notification to him using a smart phone.

[0203] After arriving at Li Sisi according to the navigation, Robot No. 1 waits for the signing process to be completed, monitors the document status in real time, and transmits the embodied status to the perceptual intelligent agent 220 through its own embodied status feedback interface 314.

[0204] Time step three:

[0205] 1. Perception Agent:

[0206] Perception agent 220 receives feedback from robot No. 1 and confirms that the document has been signed by Li Sisi.

[0207] The perception agent 220 calls the memory module 210 to query Wang Wu's location and environmental information.

[0208] The execution status of the previous task is updated to the task perception package. The content of the updated task perception package is: Current status: The document has been signed, and Robot No. 1 is at Li Si's position.

[0209] The perceptual agent 220 passes the updated task perception package to the reflective agent 250.

[0210] 2. Reflective Agents:

[0211] The reflection agent 250 evaluates the task execution result based on the task plan, specific task, and task perception package of the state changes before and after task execution generated in the previous step: If the document has been signed, it confirms that the current task is successful and continues to execute the tasks in the next stage; If the document is not signed (e.g., Li Si did not perform the signing operation), the reflection agent 250 will provide feedback on the reasons.

[0212] After completing the reflection, the reflection agent 250 passes the reflection result to the downstream planning agent 230 and decision-making agent 240 to guide them to continue executing subsequent instructions or generate remedial measures.

[0213] 3. Planning agent:

[0214] Based on the task perception package provided by the perception agent 220, the planning agent 230 determines the unfinished part of the task and generates the following task plan: Take the signed document to Wang Wu's location and deliver the document to Wang Wu.

[0215] If the previous task fails, the planning agent 230 generates a remedial plan based on the feedback from the reflection agent 250. For example, if the document is not signed, the planning agent 230 will insert a remedial step (such as asking the robot to hand the document to Li Si again).

[0216] 4. Decision-making agent:

[0217] The decision-making agent 240 generates a specific action sequence based on the task plan issued by the planning agent 230 and the capabilities of the robot cluster:

[0218] ① Physical action: <Robot No. 1: Move to Wang Wu's location>.

[0219] ② Online operation: <Smartphone: Send a notification to Wang Wu to remind him to prepare to receive the document>.

[0220] The decision-making agent 240 calls the advanced function container 312 of Robot No. 1 to control Robot No. 1 to navigate to the location where Wang Wu is and send a notification to Wang Wu through the smartphone.

[0221] After Robot No. 1 reaches the target, it checks the delivery status to ensure that the document has been successfully delivered to Wang Wu; Robot No. 1 transmits its embodied feedback information to the perception agent 220 to report the document delivery status.

[0222] 5. Task completion:

[0223] If the document is successfully delivered to Wang Wu, the perception agent 220 updates the state of the final task completion to the task perception package. After passing through the reflection agent 250, planning agent 230, and decision-making agent 240 in sequence, an action sequence for notifying the user that the task has been completed is generated. The smartphone responds to this action sequence to notify the user that the task has been completed.

[0224] If the file is not successfully delivered, the reflection agent 250 will provide feedback on the reasons for the failure and guide the planning agent 230 and the decision-making agent 240 to generate remedial measures to ensure the successful completion of the task.

[0225] The embodiment of the present invention embodies three task processing modes during the task execution process:

[0226] 1. Reactive task processing: After receiving user instructions, it automatically begins to perceive the environment, plan problems, generate specific decision-making tasks, reflect on task behaviors and results, and finally complete user instructions like a human assistant.

[0227] 2. Active task completion: When information is incomplete, actively interact with users to complete the information to ensure smooth execution of the task.

[0228] 3. Proactive help seeking: The robot monitors the execution process in real time through embodied feedback, predicts the results of instruction execution, and proactively seeks human help to ensure the correct implementation of the task.

[0229] In summary, the embodiment of the present invention provides a multi-agent collaborative embodied task reasoning and robot scheduling system and method, which has the following advantages:

[0230] 1. Multi-agent collaboration: Through the collaborative work of perception, planning, decision-making and reflection agents, intelligent reasoning and efficient execution of complex tasks can be achieved.

[0231] 2. Dynamic reasoning and interaction: It has reactive, proactive and proactive task interaction modes, which can effectively solve the problems of incomplete task information and dynamic changes in the environment.

[0232] 3. Closed-loop feedback and optimization: Dynamically evaluate the execution results through reflective agents to continuously optimize system performance and task execution results.

[0233] Therefore, the present invention effectively improves the task execution capability of the service robot in real scenarios and provides an intelligent and dynamic solution for the processing of complex tasks.

[0234] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0235] Although embodiments of the present disclosure have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and alterations may be made to the embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the claims and their equivalents.

Claims

1. A multi-agent collaborative embodied task reasoning and robot scheduling system, characterized in that: include: An interaction unit, used to obtain a natural language instruction issued by a user and send it to the intelligent processing unit, interact with the intelligent processing unit, receive the execution result of the task and the embodied state information of the robot, and provide a request for supplementary key task information when the intelligent processing unit determines that the task information is missing; The intelligent processing unit is used to dynamically perceive the environment and task status according to the user instructions sent by the interactive unit, the embodied state information and environmental information fed back by the robot, and the memory stored in the intelligent processing unit itself, and to dynamically plan and make decisions based on the perception results and the reflection results of the task execution results to realize task reasoning and robot scheduling; The robot cluster executes tasks according to the action sequence issued by the intelligent processing unit, and feeds back its own embodied state information through the interface, so as to help the intelligent processing unit dynamically perceive the environmental state and update the memory of the intelligent processing unit.

2. The system according to claim 1, characterized in that The intelligent processing unit adopts a multi-agent collaborative architecture based on a large language model, including a memory module, a perception agent, a planning agent, a decision agent, and a reflection agent. Each agent interacts through natural language and shares the memory stored in the memory module; The perception agent is used to dynamically perceive the environment and task status according to the user instructions obtained by the interaction unit, combined with the robot's embodied state information and historical task execution logs, and generate a task perception package containing task goals, location information and real-time task execution feedback; The planning agent is used to dynamically reason about task decomposition and execution scheme and generate a task plan according to the task perception package and the reflection results provided by the reflection agent; The decision-making agent is used to analyze the task plan, generate action sequences, and perform robot scheduling to guide the robot to complete the task; The reflective agent is used to evaluate the task execution results. If the task execution results do not meet expectations, it provides feedback for the next round of task reasoning to re-plan or adjust the execution plan.

3. The system according to claim 2, characterized in that The memory module stores long-term memory and short-term memory; The long-term memory includes an environment semantic map, a user default location, static location information of items, location information of public devices, user preferences, and historical task execution logs; The short-term memory includes real-time status information and dynamic context information during the current task execution process. The real-time status information includes user instructions, current task execution progress, the robot's embodied status information, dynamic environment information, the reflection results of the reflective agent and user feedback records during the task execution process.

4. The system according to claim 2, characterized in that The task reasoning modes of the intelligent processing unit include reactive, proactive and proactive; Reactive: When the perception agent determines that the task information is complete according to the user instructions sent by the interaction unit and the embodied state fed back by the robot, it extracts the task objectives and operation requirements from the task information and generates a task perception package containing the task objectives, location information and real-time feedback; the planning agent generates the task execution steps according to the task perception package and the reflection results provided by the reflection agent to form a task plan, which should meet the most basic executable conditions; The decision agent generates an action sequence after matching the task plan issued by the planning agent with the embodied capabilities of the robot cluster; Active: When the perception agent determines that the task information is incomplete according to the user instruction sent by the interaction unit and the embodied state fed back by the robot, it sends a request for supplementary task information to the planning agent for the information gap. The planning agent forms a corresponding task plan based on the request and the reflection result provided by the reflection agent. The decision agent generates an action sequence according to the task plan and obtains the required supplementary information by searching the relevant information in the memory module or sending a user-initiated query request to the instruction through the interaction unit; Proactive: When the current task cannot be completed after multiple rounds of active reasoning, the perception agent sends a multi-party assistance request to the planning agent. The planning agent forms a corresponding task plan based on the request and the reflection results provided by the reflective agent. The decision-making agent generates an action sequence based on the task plan and initiates inquiry requests to other users in the environment through the interaction unit to expand the scope of information acquisition and ensure the smooth execution of the task.

5. The system according to claim 4, characterized in that The method of initiating a query request to other users in the environment through the interaction unit includes: Individual contact: Identify potential assisting users and send requests to obtain necessary information. The potential assisting users include persons who are not directly involved in the task or persons whose records do not exist in the memory module; Group broadcast: Post assistance requests in the work scene group and coordinate multiple resources to obtain the information required for the task.

6. The system according to claim 2, characterized in that The robot comprises: The robot knowledge container is used to store the environment semantic map, target user location and object coordinate information, and supports real-time updates; Robot high-level function container, used to provide function tools for obtaining and controlling the robot's embodied state, including navigation, item interaction, and real-time state feedback; A robot execution module, used to execute corresponding tasks according to the action sequence generated by the decision-making agent, including navigation, item interaction and sending online messages; The embodied state feedback interface is used to feed back the robot's current physical state and task progress to the perceptual agent to update the task state information.

7. The system according to claim 6, characterized in that The action sequence generated by the decision-making agent includes online operations and physical actions. The decision-making agent sends action instructions to the robot execution module through an interface and dynamically adjusts the execution plan during the task execution process.

8. The system according to claim 6, characterized in that The process of the decision-making agent scheduling the robot includes: The decision-making agent breaks down the task plan output by the planning agent into multiple subtasks based on the function tools available to the robot, and generates a complete action sequence; the decision-making agent selects one or more subtasks that should be executed in the next step from the subtasks based on the robot's load, position and embodied capabilities, determines the most suitable robot through a load balancing strategy, and sends specific action instructions to the corresponding robot execution module using a unified communication protocol; during the process of the robot executing the action instructions, the decision-making agent detects whether the action is proceeding smoothly by continuously receiving the operating status, power, execution progress or error reports provided by each robot through an interface, thereby achieving continuous tracking of the execution status; at the same time, the decision-making agent identifies abnormal situations based on rules or learning models, and once an abnormality is detected, it triggers error correction or interrupts task execution.

9. The system according to claim 2, characterized in that The reflective agent determines whether the task is successfully completed by analyzing the real-time task execution feedback recorded in the task perception package, and records the execution status to achieve real-time evaluation of the task execution results; if the task is not completed or the execution fails, the reflective agent diagnoses the cause of the failure, provides optimization suggestions to the decision-making agent, and stores the reflection results in the memory module during subsequent information transmission; the reflective agent also optimizes the reasoning and execution process of future tasks through historical reflection results.

10. The system according to claim 2, characterized in that The reflection results include the judgment of the success or failure of task execution, analysis of abnormal causes and suggestions for improvement.

11. An embodied task reasoning and robot scheduling method, characterized in that: The following steps are involved: Obtain natural language instructions issued by the user, as well as the execution results of the task and the robot's embodied state information, and issue a request to supplement the missing task information when it is determined that the task information is missing; The environment and task status are dynamically perceived according to the natural language instructions, the embodied state information fed back by the robot, the perceived environmental information, and the stored memory, and dynamic planning and decision-making are performed based on the perception results and the reflection results on the task execution results to achieve task reasoning and robot scheduling, wherein the embodied state information is fed back by the robot cluster after executing the task according to the received total work sequence, and is used to dynamically perceive the environmental status and update the stored memory.

12. The method according to claim 11, characterized in that The method supports three task interaction modes: Reactive interaction: directly execute the natural language instructions with complete information; Active interaction: When task information is missing, the user requests the instructions or searches its own stored memory to supplement the missing task information; Proactive interaction: When active interaction fails to obtain the missing information of the task, multiple parties collaborate to supplement the missing information of the task by sending requests to other users in the environment, either in a one-to-one contact or by broadcasting in a group, seeking help.

Citation Information

Cited By

  • Intelligent agent interpretability information processing apparatus, method, device, and medium

    CN120234400A

  • Multi-agent task planning method and device for body environment

    CN120235420A

  • Intelligent yield management method and system based on multi-agent network

    CN120373972A

  • Space intelligent reasoning method and system

    CN120471182A

  • Robot control method and device, equipment and computer readable storage medium

    CN120503204A