Multi-agent architecture-based method and system for realizing intelligent control of tool body

By decomposing embodied intelligent control tasks into multiple sub-tasks through a multi-agent architecture, which are then executed collaboratively by multiple intelligent agent modules, the coupling and robustness issues of monolithic models are resolved, resulting in more efficient and flexible embodied intelligent control.

CN121069822APending Publication Date: 2025-12-05WULINGXIN (HAINAN) INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511292752.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing single-unit embodied intelligent control methods suffer from problems such as high model coupling, poor scalability, insufficient robustness, and limited task processing capabilities.

Method used

A multi-agent architecture is adopted to decompose complex control tasks into multiple sub-tasks, which are executed collaboratively by specialized intelligent agent modules, including a planning intelligent agent module, a specialized execution intelligent agent module, and a task scheduling and communication module, to achieve task decomposition, assignment, collaboration, and status monitoring.

Benefits of technology

It improves the modularity and scalability of the system, significantly enhances task processing capabilities, strengthens the system's robustness and fault tolerance, and enables the system to complete complex tasks more efficiently and accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121069822A_ABST
    Figure CN121069822A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for realizing intelligent control of a tool body based on a multi-agent architecture, and belongs to the technical field of artificial intelligence and robots. The method aims to solve the problems of single task processing capability, poor robustness and difficulty in expansion of the existing single-body intelligent model. The core thought of the method is to introduce a multi-agent collaboration architecture, decouple'high-level task planning 'and'bottom-level skill execution', and decompose complex control tasks to specialized agents with different functions. The method comprises the following steps: receiving a high-order task instruction; decomposing the high-order task instruction into a plurality of subtasks by a planning agent; the planning agent assigns at least one corresponding execution agent to the subtask from a set comprising a plurality of specialized execution agents according to the attribute of each subtask; and the assigned execution agent cooperatively executes the allocated sub-task so as to complete the high-order task. According to the invention, the modularization, the expandability and the task execution robustness of the intelligent system are obviously improved, and the method is suitable for various intelligent scenes such as industrial manufacturing and home service.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to embodied AI and robot control technology, and more particularly to a method and system for implementing embodied AI control using a multi-agent collaboration architecture. BACKGROUND

[0002] Currently, one of the main research directions in the field of embodied AI is to develop monolithic models, that is, to build a large, usually end-to-end neural network model that is responsible for handling the entire process from environmental perception, task understanding, decision planning to action execution.

[0003] However, this monolithic paradigm faces many insurmountable bottlenecks in engineering practice and commercialization: High coupling and low scalability: monolithic models highly couple all functions (such as navigation, grasping, and interaction) together. When a new skill needs to be added to the system or it needs to adapt to a new environment, costly retraining or complex fine-tuning of the entire large model is often required, lacking modular design and poor scalability; Poor robustness and fault tolerance: since all functional modules are tightly coupled, any error in a single module (such as a small error in a perception module) can cause the entire task chain to fail through error propagation. This "single point of failure" problem makes the system vulnerable in complex and variable environments, lacking effective fault tolerance and recovery mechanisms; Limited task processing capability and high training cost: a single model cannot master multiple professional skills with large differences at the same time. Trying to make a model "omniscient" not only leads to suboptimal performance in any single skill, but also dramatically increases model complexity and the amount of data required for training, resulting in high development threshold and long cycle.

[0004] Therefore, it is a key technical problem in the field to build an embodied AI control architecture that can coordinate multiple professional skills to complete complex tasks, while avoiding the inherent defects of monolithic models, and is more modular, scalable, and robust. SUMMARY

[0005] The present application aims to solve the problems of high model coupling, poor scalability, insufficient robustness, and limited task processing capability in existing monolithic embodied AI control methods.

[0006] To solve the above problems, the first object of the present application is to provide a method for realizing embodied intelligent control based on multi-agent architecture, the core idea of which is "divide and conquer", that is, a huge and complex control problem is transformed into multiple smaller and more focused sub-problems, and each field "expert" agent is responsible for solving the problem. S1, instruction receiving and decomposition: a planning agent receives a high-level task instruction and decomposes the high-level task instruction into multiple ordered sub-tasks; S2, task assignment: the planning agent matches and assigns at least one corresponding execution agent for each sub-task according to the attributes of the sub-task from a set containing multiple specialized execution agents; S3, cooperative execution: the assigned execution agent is instructed to call its internal encapsulated professional skills to cooperatively execute the assigned sub-task; S4, state monitoring and feedback: the planning agent monitors the task execution state of each execution agent, and confirms the completion of the high-level task after all sub-tasks are completed.

[0007] The second object of the present application is to provide an embodied intelligent control system based on multi-agent architecture, characterized in that it comprises: a planning agent module for receiving a high-level task instruction, understanding and planning the task, decomposing the instruction into multiple sub-tasks, and generating an execution plan containing task assignment information; a set of specialized execution agent modules containing multiple independent execution agent modules encapsulating different underlying algorithms or professional skills, which are used to receive and execute specific sub-task instructions according to the execution plan; a task scheduling and communication module for analyzing the execution plan, sequentially calling the assigned execution agent modules in the set, and being responsible for transmitting state information and coordinating instructions between the planning agent module and each execution agent module.

[0008] Compared with the prior art, the multi-agent architecture of the present application has the following advantages: Excellent modularity and scalability: by decomposing complex control tasks into different functional specialized agents, each agent can be developed, optimized and upgraded as an independent "skill plug-in". When a new skill is needed (such as adding a "door opening" skill), only a new specialized agent needs to be developed and integrated, without changing the core architecture, greatly improving the scalability of the system.

[0009] Significant improvement in task processing capability: Each execution agent can focus on its area of expertise (such as visual positioning, pushing manipulation), achieving higher skill proficiency and precision. Through the collaboration of these experts in various fields, the system can more efficiently and accurately complete complex tasks that require a combination of multiple capabilities.

[0010] Very high robustness and fault tolerance: Distributed execution of tasks avoids single-point failures. When an execution agent fails, the planning agent can detect the failure and initiate fault tolerance mechanisms, such as reassigning tasks to other agents or adjusting the original plan, significantly improving the reliability of the entire system in the face of uncertainty. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed for the description of the embodiments will be briefly introduced below.

[0012] Figure 1 The system architecture diagram for implementing embodied intelligent control based on multi-agent architecture in an embodiment of the present application; Figure 2 The method flowchart for implementing embodied intelligent control based on multi-agent architecture in an embodiment of the present application.

[0013] ATTACHMENT Figure 1 Label: 10, embodied intelligent control system; 20, physical environment; 30, planning agent module; 40, set of specialized execution agent modules; 41, visual positioning agent module; 42, pushing manipulation agent module; 50, task scheduling and communication module. DETAILED DESCRIPTION

[0014] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all embodiments.

[0015] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another component, it can be directly on the other component or indirectly disposed on the other component; when a component is referred to as being "connected to" another component, it can be directly connected to the other component or indirectly connected to the other component.

[0016] The core idea of the present application is to build a multi-agent system that works collaboratively in a divide-and-conquer manner. The following is an example of a system deployed on a robot with a mechanical arm and a table lamp (referred to as "table lamp robot" for short), which illustrates the complete process of executing the high-level task of "please push the water cup on the table to the right coaster".

[0017] 1. System architecture details Referring to Figure 1 The embodiment provides a somatic intelligent control system 10 based on a multi-agent architecture. The core of the system 10 is a collaborative network composed of multiple different agent modules. In the embodiment, the network mainly includes the following core modules: 1.1 Planning agent module (reference numeral 30) Role: "brain" and "commander" of the system.

[0018] Function: used to receive external high-level task instructions from the user (for example, "push the water cup on the table to the right coaster"). After receiving the instructions, the module performs task understanding and planning, decomposes the complex high-level task into a series of simpler and more specific ordered subtasks, and generates an execution plan containing task assignment information. For example, the following plan is generated: Step 1: assign [visual positioning agent 41] to locate the coordinates of "water cup" and "coaster"; Step 2: assign [push manipulation agent 42] to plan a push path according to the coordinates; Step 3: assign [push manipulation agent 42] to execute the push action; Step 4: assign [push manipulation agent 42] to reset the mechanical arm.

[0019] 1.2 Set of specialized execution agent modules (40) Role: "expert toolbox" or "skill library" of the system.

[0020] Composition: composed of multiple independent execution agent modules encapsulating different underlying algorithms or professional skills. Each module is an expert in its field and provides services to the outside through a standardized communication interface. In the embodiment, the set 40 at least includes: Visual positioning agent module (41): internally encapsulating object detection and three-dimensional positioning algorithms, responsible for executing all subtasks related to "finding and positioning objects"; Push manipulation agent module (42): internally encapsulating motion planning, visual servoing, and force control algorithms, responsible for executing all non-grasping object manipulation tasks such as pushing, squeezing, etc.

[0021] 1.3 Task scheduling and communication module (50) Role: The "nerve center" and "dispatcher" of the system.

[0022] Function: To parse the execution plan from the planning agent module (30), and sequentially invoke the corresponding module assigned in the specialized execution agent module set (40). It is responsible for managing the life cycle of each execution agent module, and as a communication bus, reliably transferring state information (such as "task success", "task failure", "current coordinates") and coordination instructions between the planning agent module (30) and each execution agent module (41, 42).

[0023] 2. Workflow details: Pushing the water cup task Referring to The complete method flow of the table lamp robot executing the "please push the water cup on the table to the right coaster" task is as follows: Step S10: Receive high-level task instruction. The system (10) receives a high-level task instruction through the user interface; Step S20: Task decomposition and planning. The instruction is passed to the planning agent module (30). This module performs semantic understanding on the instruction and decomposes it into the aforementioned four ordered subtasks, generating an execution plan; Step S30: Task assignment and scheduling. The planning agent module (30) sends the execution plan to the task scheduling and communication module (50). The task scheduling and communication module (50) parses the first step of the plan, and then invokes the visual positioning agent module (41) in the set (40), and issues the "position the water cup and coaster" instruction to it; Step S40: Subtask execution. The visual positioning agent module (41) is activated, which processes data from the camera, finds the water cup and coaster in the physical environment (20), calculates their coordinates, and then feeds back the results to the planning agent module (30) through the task scheduling and communication module (50); Step S50: Monitoring and state transition. After the planning agent module (30) receives the successful execution result, it judges that subtask 1 is completed and updates the state of the execution plan. The task scheduling and communication module (50) then continues to invoke the pushing control agent module (42) according to the plan to execute the subsequent path planning, pushing and resetting tasks;

[0024] Step S60: Cycle and fault tolerance. In the process of pushing, if the pushing control intelligent agent module (42) feedbacks "pushing blocked", the failure state will be reported to the planning intelligent agent module (30) through the task scheduling and communication module (50). The planning module (30) can start the fault tolerance mechanism, for example, generate a new plan containing "adjust the pushing angle", and issue it to the task scheduling and communication module (50) again for execution. Until all subtasks are confirmed as "yes" (all completed) in step S60, the whole high-level task is declared to be completed.

[0024] As can be seen from the above embodiments, the spirit of the present application is to transform a huge and complex control problem into multiple smaller and more focused sub-problems, and to hand them over to "expert" intelligent agent modules in their respective fields to solve them collaboratively through a clear scheduling mechanism, thereby realizing a more efficient, flexible and reliable embodied intelligent control paradigm as a whole.

[0025] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for realizing embodied intelligent control based on a multi-agent architecture, characterized in that, Comprising the following steps: S1, Instruction receiving and decomposition: by a planning agent, receiving a high-level task instruction and decomposing the high-level task instruction into a plurality of ordered sub-tasks; S2, Task assignment: the planning agent matches and assigns at least one corresponding execution agent for each sub-task from a set of specialized execution agents according to the attributes of each sub-task; S3, Collaborative execution: instructing the assigned execution agent to invoke its internal encapsulated professional skills to collaboratively execute the assigned sub-task; S4, State monitoring and feedback: the planning agent monitors the task execution state of each execution agent and confirms the completion of the high-level task after all sub-tasks are completed.

2. The method of claim 1, wherein, Further comprising: a perception agent responsible for processing sensor data from the physical environment and providing environment state information to the planning agent and the execution agent.

3. The method according to claim 1 or 2, characterized in that, Further comprising: a communication agent responsible for transmitting state information and coordinating instructions between the planning agent and the execution agent, as well as between different execution agents.

4. The method of claim 1, wherein, The specialized execution agent includes at least one of a visual positioning agent and a pushing control agent.

5. The method of claim 1, wherein, The planning agent is further configured to: re-plan or re-assign the failed sub-task when monitoring the failure of sub-task execution.

6. The method of claim 1, wherein, The S2 step further comprises: According to the complexity of the high-level task instruction, dynamically instantiate new execution agents and add them to the set.

7. A body-aware intelligent control system based on a multi-agent architecture, characterized in that, Comprising: a planning agent module for receiving a high-level task instruction, task understanding and planning, decomposing the instruction into a plurality of sub-tasks, and generating an execution plan containing task assignment information; a set of specialized execution agent modules containing a plurality of independent execution agent modules encapsulating different underlying algorithms or professional skills, which are used to receive and execute specific sub-task instructions according to the execution plan; a task scheduling and communication module for parsing the execution plan, sequentially invoking the assigned execution agent modules in the set, and responsible for transmitting state information and coordinating instructions between the planning agent module and each execution agent module.

8. The system of claim 7, wherein, Further comprising: a perception agent module configured to process sensor data and provide environment state information to the planning agent module and the set of specialized execution agent modules.

9. The system of claim 7, wherein, The task scheduling and communication module interacts with the planning agent module and the set of specialized execution agent modules through an internal communication bus.

10. A computer-readable storage medium having stored thereon a computer program, the program being executed by a processor to implement the method of any one of claims 1 to 6.