Multi-agent service collaboration method based on collaboration diagram and large language model
By building a collaboration map in smart homes and using large language models and reinforcement learning technology for task planning and optimization, the problem of lack of global coordination in the collaboration of multiple smart home devices is solved, and efficient intelligent collaboration and housework tasks are achieved.
Patent Information
- Application Number
- CN202510143475.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology lacks effective global coordination and task planning methods in the collaboration of multiple smart home devices, resulting in the equipment being prone to conflicts or inconsistencies and being unable to achieve efficient human-machine collaboration.
A multi-agent service collaboration method based on collaboration graphs and large language models is adopted, high-level task planning is carried out through large language models, and low-level skill movements are optimized in combination with reinforcement learning technology to achieve efficient collaboration between different agents.
It realizes unified cooperation among multiple agents, improves the efficiency and flexibility of housework execution, enhances the robustness and safety of the system, and provides convenient, comfortable and intelligent services for family life.
Smart Images

Figure CN120068922A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent devices, and particularly relates to a multi-agent service collaboration method based on a collaboration graph and a large language model. Background Art
[0002] With the development of artificial intelligence and Internet of Things technologies, household appliances and service robots have been continuously improved in intelligence and gradually possess certain autonomous operation capabilities. For example, a sweeping robot can perform local path planning based on sensors; a dishwasher can automatically adjust the washing mode according to the dirt level of tableware; a washing machine can receive remote commands through the network. However, when multiple smart home devices and robots need to collaborate simultaneously to complete more complex household scenarios (such as cleaning the living room while the dishwasher is washing dishes, the washing machine is doing laundry, and other robots are carrying items), there is currently a lack of effective global coordination and task planning means.
[0003] The "intelligence" of individual devices is prone to conflicts or inconsistencies in multi-device collaboration, such as overlapping cleaning areas, device idleness, or resource preemption, and efficient human-machine collaboration cannot be achieved. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a multi-agent service collaboration method based on a collaboration graph and a large language model. High-level task planning is performed through the large language model, and reinforcement learning technology is used to optimize low-level skill actions, so that different types of agents can collaborate efficiently and provide convenience and comfort for family life.
[0005] To achieve the above object, the technical solution adopted by the present invention is: a multi-agent service collaboration method based on a collaboration graph and a large language model, including the steps of:
[0006] S10, collecting environmental perception information, human-machine collaboration relationship information, and user requirements including multiple agents, and performing preprocessing;
[0007] S20, constructing a collaboration graph from the preprocessed data to obtain a collaboration graph;
[0008] S30, performing large language model planning on the collaboration graph to obtain a high-level task plan;
[0009] S40, optimizing the high-level task plan through reinforcement learning skills to obtain specific actions or execution instructions;
[0010] S50, sending the specific actions or execution instructions to the corresponding agents for execution.
[0011] Furthermore, in step S10, data collection includes:
[0012] The environmental perception information includes: home layout data, voice or text instructions, and environmental sensor data;
[0013] The human - machine collaboration relationship information: family member needs, mutual exclusion or dependency relationships between intelligent agent devices, and historical interaction logs;
[0014] User needs: task request instructions sent by the user to the intelligent agent device.
[0015] Furthermore, in step S10, the data pre - processing includes:
[0016] Data cleaning: removing outliers and noise from the sensor data;
[0017] Data alignment: aligning the data from different sensors and user inputs in terms of time;
[0018] Data standardization: normalizing the data from different sensors to the same scale;
[0019] Feature extraction: extracting useful features from the images captured by the camera, extracting speech features from the voice signals received by the microphone, and extracting key features from the device status sensors;
[0020] Data fusion: Data fusion is to comprehensively process the data from different sensors and user inputs to generate a comprehensive data view;
[0021] Data annotation: For some common tasks and scenarios, data annotation is carried out.
[0022] Furthermore, in step S20, during the construction of the collaboration graph, the steps include:
[0023] S201, obtaining the sensor and home appliance status data, and obtaining the prior information of household devices and users;
[0024] S202, storing the processed data in the form of an adjacency matrix or a graph database, and constructing a collaboration graph;
[0025] In the collaboration graph, people, various intelligent agents, and key locations or resources in the home environment are defined as a node, and are attached with attribute information; at the same time, the collaboration graph also records the collaboration edges between nodes, indicating the reachability, interaction relationships, or task dependency relationships between devices, and is attached with attribute information.
[0026] Furthermore, according to the changes in device status, the updates of user needs, and possible fault or energy consumption replenishment events, the collaboration graph is dynamically updated.
[0027] Furthermore, in step S30, the large - language model planning is carried out, including the steps:
[0028] S301. Obtain the established collaboration graph and obtain the user intention or task instruction;
[0029] S302. Adopt the GraphRAG principle, retrieve the collaboration graph G to obtain the context information Context related to the user query UserQuery, and input this information together with the user query into the large language model LLM to generate a detailed high-level task plan Plan;
[0030] Context = Retrieve(G, UserQuery)
[0031] Plan = LLM(UserQuery + Context);
[0032] Among them, Retrieve represents the search function:
[0033] The high-level task plan includes: a multi-step task sequence, which clearly indicates the operations that each agent should perform at a specific time; the scheduling result and subtask allocation.
[0034] Furthermore, when generating the plan, the LLM formulates an optimal or sub-optimal task execution plan according to the constraints of multi-device collaboration.
[0035] Furthermore, according to the user's new instruction or the change of the device state, retrieve the collaboration graph again and dynamically adjust the original plan.
[0036] Furthermore, in step S40, the reinforcement learning skill optimization process includes:
[0037] S401. Obtain the high-level task plan and obtain the graph embedding information;
[0038] S402. Multi-device reinforcement learning strategy: Each device or robot has its exclusive or shared policy network, and the graph embedding is fused into the state representation during training;
[0039] S403. The output content is specific action or device control instructions, which are generated according to the subtasks decomposed from the high level and the graph embedding information.
[0040] Furthermore, evaluate and adjust the execution results of the agent and give feedback, and optimize and adjust in the value preprocessing stage, collaboration graph construction stage, large language model planning stage, and reinforcement learning skill optimization stage respectively.
[0041] Beneficial effects of adopting this technical solution:
[0042] Through constructing a collaboration graph and using a large language model for task planning, and combining reinforcement learning techniques to optimize low-level skill actions, this system can dynamically schedule and control multiple intelligent devices and robots according to user needs and the status of the home environment, and complete complex household tasks such as cleaning, tidying up, and cooking assistance. It aims to solve problems such as conflicts, inconsistencies, and lack of effective global coordination in multi-device collaboration in the prior art, improve the efficiency, flexibility, and robustness of household task execution, and provide more convenient, comfortable, and intelligent services for family life.
[0043] This invention can achieve unified collaboration of multiple intelligent agents: breaking the limitation of only focusing on mobile robots in the traditional sense, incorporating household appliances such as floor sweepers, dishwashers, and washing machines into the collaboration framework, realizing efficient collaboration between multiple devices and between humans and machines, and completing various household tasks such as cleaning, tidying up, and cooking assistance.
[0044] This invention can improve the efficiency and flexibility of household task execution: by constructing a collaboration graph, comprehensively understanding the capabilities and status of each device, avoiding duplicate work or resource conflicts, and significantly reducing the overall time and energy consumption of household chores. At the same time, using a large language model for high-level task planning can generate the most reasonable task allocation plan, adapting to complex and dynamic household scenarios. For example, when a user temporarily adds a new task or a device fails, the system can dynamically update the collaboration graph and regenerate the plan.
[0045] This invention can enhance the robustness and security of the system: adopting a hierarchical structure, combining high-level planning and low-level skill optimization, enabling the system to still reconstruct the plan in a timely manner when there are local failures or individual device anomalies, without affecting the overall household chore process. Reinforcement learning autonomously adapts to environmental changes in control details, reduces misoperations and damage to household items, and improves the overall robustness and security of the system.
[0046] This invention can provide personalized and scalable services: the system can reflect personalized arrangements in the collaboration graph according to the preferences and schedules of different family members, meeting the diverse needs of users. In addition, this invention has good scalability. When adding or replacing household appliances, only need to register new device nodes and their skill attributes in the collaboration graph, and the system can automatically incorporate them into the collaboration network, applicable to various scenarios such as families, nursing homes, medical places, and public places.
[0047] This invention can optimize multi-device control and action execution: after each smart home device receives a high-level task, it uses a reinforcement learning strategy to select optimal or sub-optimal actions, improving the overall efficiency and ensuring safety. Through graph embedding information and reinforcement learning strategies, this invention can guide the device to make autonomous decisions and real-time adjustments in a complex environment, such as precise control of a floor sweeper when avoiding obstacles and a robotic arm when carrying items.
[0048] The present invention can achieve dynamic adjustment and real-time feedback: The system can perceive environmental changes and device status in real time, dynamically update the collaboration graph, and ensure the real-time and accuracy of task planning and execution. Users can adjust task requirements at any time through voice or text instructions. The system will regenerate the task plan and issue execution commands according to the latest collaboration graph and user instructions, realizing true dynamic adjustment and real-time feedback. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic flow chart of a multi-agent service collaboration method based on a collaboration graph and a large language model of the present invention;
[0050] Figure 2 It is a schematic diagram of collaboration graph construction in an embodiment of the present invention;
[0051] Figure 3 It is a schematic diagram of large language model planning in an embodiment of the present invention;
[0052] Figure 4 It is a schematic diagram of reinforcement learning skill optimization in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.
[0054] The present invention relates to a multi-agent collaboration method using a collaboration graph and a large language model in a smart home scenario. The agents here include not only mobile robots (such as wheeled robots, robotic arm robots, etc.), but also household appliances with certain intelligent control functions (such as floor sweepers, dishwashers, washing machines, etc.), so as to achieve efficient cooperation between multiple devices and between humans and machines. It can complete various household tasks (such as cleaning, tidying up, cooking assistance, etc.) according to user needs in a home environment, and make dynamic adjustments according to changes in the environment and needs.
[0055] In this embodiment, referring to Figure 1 as shown, the present invention proposes a multi-agent service collaboration method based on a collaboration graph and a large language model, including the steps of:
[0056] S10, Collect environmental perception information, human-machine collaboration relationship information and user needs including multiple agents, and perform preprocessing;
[0057] S20, Construct a collaboration graph from the preprocessed data to obtain a collaboration graph;
[0058] S30, Perform large language model planning on the collaboration graph to obtain a high-level task plan;
[0059] S40. Optimize the high-level task planning with reinforcement learning skills to obtain specific actions or execution instructions;
[0060] S50. Send the specific actions or execution instructions to the corresponding agents for execution.
[0061] As an optimized solution for the above embodiments, in step S10, data collection includes:
[0062] Environmental perception information includes: home layout data (room area, furniture location), voice or text instructions, and environmental sensor data (cameras, microphones, device status sensors, etc.); specifically, the home layout information is generated by cameras and lidar (LIDAR) sensors installed in various corners of the home. The system can accurately record the area of each room and the placement location of furniture. This information helps agents perform effective path planning and space utilization when executing tasks. In addition, multiple sensors are integrated, including cameras, microphones, and device status sensors, for real-time monitoring of various dynamic changes in the home environment. Cameras are used for visual monitoring, capturing images and videos in the home environment to identify objects, people, and scene changes; microphones are used for speech recognition, receiving user voice instructions and performing real-time speech parsing; device status sensors are used to monitor the operating status of home appliances and robots, such as battery power, fault information, working mode, etc., to ensure the normal operation of the devices.
[0063] Human-machine collaboration relationship information: family member needs, mutual exclusion or dependency relationships between agent devices (e.g., a dishwasher needs water, which may conflict with the water intake function of a floor sweeper), and historical interaction logs. The system collects the specific needs and preferences of family members in various ways. For example, user A may prefer to work in a quiet environment and does not want the washing machine to run during working hours. At the same time, the system identifies and manages the mutual exclusion or dependency relationships between different devices. For example, when the dishwasher is running, it needs to use water resources, which may conflict with the water intake function of the floor sweeper. The system needs to coordinate such resource competitions to ensure that devices can use resources as needed. In addition, the system records all interaction histories between users and agents, including task requests, execution results, and user feedback. These logs help it better understand the user's behavior patterns and preferences, thereby providing more personalized and efficient services.
[0064] User Requirements: Task request instructions sent by users to intelligent agent devices. For example, "Please clean the living room and wash the dirty dishes by the way" or "I need to wash my clothes before 4 pm". The system receives voice signals through the microphone and uses advanced speech recognition technology to convert them into text instructions. Subsequently, the system performs semantic parsing on the text instructions to extract key information such as task content, time requirements, priority, etc. For example, from "Please clean the living room and wash the dirty dishes by the way", two tasks are parsed: cleaning the living room and washing the dirty dishes, and it is recognized that these two tasks can be executed in parallel. In addition, the system not only supports voice instructions but also combines visual inputs, such as users entering text instructions through a mobile application or the touch screen of an intelligent device. For example, users can select specific rooms to clean in the application or set the washing mode and time of the washing machine. During the task execution process, the system prompts the user about the progress of the task through voice or text and receives the user's feedback. For example, users can confirm whether the task is completed or evaluate the task result, and the system makes adjustments and optimizations based on the feedback.
[0065] In step S10, data preprocessing includes:
[0066] Data cleaning: Removing outliers and noise from sensor data; for example, processing noise in camera images through a filtering algorithm or removing background noise in microphone voice signals through signal processing techniques. For missing data, interpolation or other statistical methods are used to fill it to ensure data integrity. For example, if the data of a device status sensor is missing, the data at adjacent time points can be used for interpolation to ensure data continuity and reliability.
[0067] Data alignment: Aligning data from different sensors and user inputs in terms of time; data alignment ensures the consistency of data from different sources in terms of time and space. For example, aligning the images captured by the camera with the voice instructions received by the microphone to accurately parse the user's task request. In addition, for data related to spatial positions, such as home layout information and device locations, spatial alignment is performed to ensure the accuracy of the data in space. For example, aligning the 3D map generated by lidar with the camera image to more accurately identify objects and paths.
[0068] Data standardization: Normalizing data from different sensors to the same scale; for example, normalizing the data of device status sensors (such as battery level, temperature, etc.) to the range of 0 to 1 for subsequent analysis and processing. In addition, standardizing the data eliminates the dimensional differences between different data sources. For example, normalizing the pixel values of camera images to a distribution with a mean of 0 and a standard deviation of 1 for more effective feature extraction and improving the efficiency and accuracy of data processing.
[0069] Feature extraction: Extract useful features from the images captured by the camera, extract speech features from the speech signals received by the microphone, and extract key features from the device status sensors; such as the shape, color, texture of objects, etc. Deep learning techniques such as Convolutional Neural Network (CNN) can be used for feature extraction for subsequent object recognition and scene understanding. Extract speech features from the speech signals received by the microphone, such as Mel Frequency Cepstral Coefficients (MFCC), etc., for speech recognition and semantic parsing. At the same time, extract key features from the device status sensors, such as the operating status, fault information, power level of the device, etc., for device management and task scheduling.
[0070] Data fusion: Data fusion is to comprehensively process the data from different sensors and user inputs to generate a comprehensive data view; the system fuses the camera images, microphone speech signals and device status data to generate a comprehensive home environment model. In addition, context information such as the user's historical interaction logs, the needs of family members, and the mutual exclusion or dependency relationships between devices is fused into the data so that the system can better understand the background and requirements of the task, thereby providing more personalized and efficient services.
[0071] Data annotation: For some common tasks and scenarios, data annotation is carried out. For example, the objects captured by the camera are automatically annotated as "chair", "table", etc. For complex or uncommon tasks and scenarios, manual annotation can be carried out to ensure the accuracy and reliability of the data. For example, manually annotate the specific task content and time requirements in the user's voice commands so that the system can understand and execute the user's commands more accurately.
[0072] Through the above data preprocessing process, the system can ensure the quality and usability of the input data, providing a solid data foundation for the efficient cooperation and task execution of multiple agents. This not only improves the accuracy and efficiency of task execution, but also enhances the robustness and adaptability of the system, ensuring stable and reliable services in a complex and changing home environment.
[0073] As an optimized solution of the above embodiment, in step S20, as Figure 2 shown, in the process of constructing the collaboration graph, it includes the steps:
[0074] S201, Obtain sensor and home appliance status data, and obtain prior information of household devices and users;
[0075] 1. Sensor and home appliance status data:
[0076] The position information of the sweeping robot, the current task progress (whether it is cleaning, whether the power is sufficient, etc.);
[0077] The current washing mode of the dishwasher and the estimated remaining time;
[0078] The working status (washing / spinning / done) of the washing machine and its availability;
[0079] The real-time position, skill set and fault warning of home robots (such as robotic arms, wheeled service robots);
[0080] Environmental sensors (room temperature, humidity, illuminance, etc.);
[0081] The content of the user's voice input and the semantic analysis result.
[0082] 2. Home appliances and user prior information:
[0083] The list of family members and their preferences: For example, user A likes quietness and does not want the washing machine to run late at night; user B likes to complete cleaning quickly;
[0084] The skill lists of home appliances and robots: The floor sweeper only has the ability to clean the floor, the dishwasher can wash tableware, the washing machine can wash clothes, the robotic arm can perform handling operations, etc.;
[0085] The home layout and the positions of devices: The spatial distribution of the living room, kitchen, laundry room, etc.
[0086] S202, the processed data is stored in the form of an adjacency matrix or a graph database to construct a collaboration graph;
[0087] In the collaboration graph, people, various intelligent agents, and key positions or resources in the home environment are defined as nodes, and are attached with attribute information, such as "device type = washing machine", "working status = standby", "position = laundry room", etc. These attribute vectors detail key information such as the type, current status, and position of the device. At the same time, the collaboration graph also records the collaboration edges between nodes, indicating the reachability, interaction relationship, or task dependency between devices, such as "washing machine - robotic arm" indicating that the robotic arm can assist in picking up and placing clothes and is attached with information;
[0088] The collaboration graph data structure, which is one of the core components, is used to comprehensively describe and manage the relationships between people and various intelligent agents (including floor sweepers, dishwashers, washing machines, robotic arms, service robots, etc.) and key environmental positions or resources in the home environment. The collaboration graph consists of three main parts:
[0089] 1. Nodes: Nodes represent people, various intelligent agents, and key positions or resources in the home environment. These nodes are the basic units of the collaboration graph, covering all entities participating in household tasks, such as family members, various household appliances and robots, and important environmental positions, such as the living room, kitchen, laundry room, etc.
[0090] 2. Edge: An edge represents the reachability, interaction relationship, or task dependency between nodes. For example, the edge "washing machine - robotic arm" indicates that the robotic arm can assist in picking up and placing clothes in the washing machine, reflecting the collaborative ability between devices. Through these edges, the system can clearly identify and manage the interactions and dependencies between different devices, thereby optimizing the task execution process.
[0091] 3. Attribute information: Each node and edge is accompanied by detailed attribute information, which provides specific descriptions of the node or edge. For nodes, attributes may include skill types (such as cleaning, washing, handling, etc.), current working status (such as running, standby, faulty, etc.), location information, etc. For edges, attributes may describe the type (such as assistance, dependency, conflict, etc.) and intensity of the interaction. These attribute information provide rich context for the system, enabling it to perform task planning and resource allocation more intelligently.
[0092] Through this structured collaboration graph, the system can comprehensively perceive the dynamic changes in the home environment, adjust task allocation and execution strategies in real time, ensure the efficient collaborative work among multiple agents, and meet the diverse needs of family members.
[0093] Preferably, in order to maintain the real-time nature and accuracy of the collaboration graph, the collaboration graph is dynamically updated according to changes in device status, updates in user requirements, and possible fault or energy consumption replenishment events. This real-time update mechanism ensures that the system can respond to environmental changes in a timely manner and provide the latest information support for the efficient collaboration of multiple agents.
[0094] Examples of application scenarios for collaboration graph construction:
[0095] The user says, "I need to wash clothes first and then use the floor sweeper to clean the kitchen." The system integrates information such as "the washing machine is in an idle state" and "the floor sweeper is already on the charging dock" into the graph, obtaining "the washing machine is available" and "the floor sweeper can be started immediately", and establishing nodes and edges such as "User A - the washing machine is available - the floor sweeper can be started".
[0096] As an optimized solution for the above embodiment, in step S30, as Figure 3 shown, large language model planning is performed, including the steps of:
[0097] S301, obtaining the established collaboration graph and obtaining user intentions or task instructions;
[0098] S302, using the GraphRAG principle, retrieving the collaboration graph G to obtain context information Context related to the user query UserQuery, and inputting this information together with the user query into the large language model LLM to generate a detailed high-level task plan Plan;
[0099] Context = Retrieve(G, UserQuery)
[0100] Plan = LLM(UserQuery + Context);
[0101] Among them, Retrieve represents: the search function.
[0102] High-level task planning includes: multi-step task sequences, clearly indicating the operations that each agent should perform at a specific time; scheduling results and subtask assignments. For example, "The sweeper: starts cleaning the living room at 12:00; the robotic arm: puts the dirty clothes into the washing machine at 11:50; the dishwasher: automatically starts the quick wash mode at 12:30", etc. These detailed instructions ensure that each agent can work together to efficiently complete various tasks in the home environment.
[0103] In this process, according to the user's specific request, such as "arrange housework", the most relevant node information is retrieved from the collaboration graph, including idle devices, potential conflicts, and user preferences, etc., and this context information is provided to the LLM to generate a comprehensive and accurate plan.
[0104] Preferably, when generating a plan, the LLM formulates an optimal or sub-optimal task execution plan according to the constraints of multi-device collaboration, such as the skill matching of different household appliances and robots, specific time period restrictions, and energy consumption constraints, etc.
[0105] Preferably, according to the user's new instructions or changes in the device status, the collaboration graph is retrieved again and the original plan is dynamically adjusted to make it have the ability to update the plan. For example, if it is necessary to postpone the start time of the washing machine or change the execution order of the devices, the system can flexibly make corresponding adjustments to ensure that the task plan always meets the latest home environment and user needs.
[0106] Examples of application scenarios for large language model planning:
[0107] The user says "I'm going out at 3 pm and hope you can finish cleaning the living room, washing clothes, and tidying up the table before 2 pm". After retrieving the collaboration graph, the LLM Planner finds that the dishwasher will be idle before 1 pm, the washing machine can be started immediately, and the sweeper has sufficient power, and then generates the following schedule:
[0108] 1. Washing machine: Immediately start the standard wash, expected to end at 1:40;
[0109] 2. Sweeper: Clean the living room from 1:00 to 2:00;
[0110] 3. Robotic arm: Assist in tidying up the table or moving obstacles at 1:30;
[0111] 4. Dishwasher: Start it according to the situation after the above is completed.
[0112] This parallel scheduling can meet the user's time requirements.
[0113] As an optimized solution for the above embodiment, in step S40, as Figure 4 shown, the process of optimizing the reinforcement learning skill includes:
[0114] S401, Obtain the high-level task plan and obtain the graph embedding information;
[0115] The main inputs of this module include two key parts: one is the subtasks decomposed at the high level, which are specific operation tasks generated by the large language model planner according to the user's instructions, such as "the sweeper goes to the living room for cleaning", "the robotic arm moves the dirty dishes from the table to the dishwasher", and "the washing machine selects the appropriate washing mode", etc., which clarify the specific actions that each agent needs to perform; the other is the graph embedding information, which is the embedding vector extracted from the collaboration graph through the graph neural network (GNN), and it can reflect the state of the current home environment, the relationship between devices, and other relevant context information. These input information together provide the necessary data support for the low-level skill optimization module of the system, enabling the system to further optimize the specific execution actions of the agents according to the real-time home environment and device status, and ensuring that the tasks can be completed efficiently and accurately.
[0116] S402, Multi-device reinforcement learning strategy: Each device or robot has its own exclusive or shared policy network, and the graph embedding (such as device location, task priority, surrounding environment information) is fused into the state representation during training.
[0117] Reward design: For the sweeper: The higher the cleaning coverage rate, the greater the reward, and the lower the reward for colliding with furniture or having a high number of repeated coverage; for the robotic arm: A positive reward can be obtained for stable posture, high grasping success rate, and short time during the handling process; for the dishwasher: There is an additional bonus for high washing cleanliness and minimized water and electricity consumption; for the washing machine: The reward is higher if it is completed within the specified time and resources are saved.
[0118] Training and deployment: The reinforcement learning training can be first completed in a simulation environment (such as a housework simulator), and after the policy tends to converge, it is deployed to the real environment, and continuous fine-tuning is carried out in combination with the actual execution feedback.
[0119] S403, the output content is specific action or device control instructions, which are generated based on the subtasks decomposed from the high level and the graph embedding information, so as to ensure that each agent can accurately execute its assigned tasks. For a floor sweeper, the system output includes driving trajectory planning and obstacle avoidance instructions, such as adjusting the speed and direction of the wheel drive to efficiently complete the cleaning task. For a robotic arm, the system provides low-level action instructions such as joint angles and grasping forces, enabling it to accurately perform handling and operation tasks. The dishwasher automatically selects an appropriate washing process according to the system instructions, such as "pre-rinse - main wash - rinse", and adds detergent at the appropriate time. The control instructions for the washing machine include matching an appropriate washing mode (such as quick wash, standard wash or low-temperature wash), and automatically determining the required water volume and rotation speed according to the weight and dirt level of the clothes. These specific instructions enable each agent to work together and efficiently complete various tasks in the home environment.
[0120] Example of application scenario:
[0121] When the floor sweeper is performing "go to the living room to clean", the reinforcement learning strategy will guide it to avoid the robotic arm that is still working in the home environment (to prevent collision), and dynamically plan the cleaning order (first suck up the visible debris, and then process the area under the sofa). If the ground is partially wet, the strategy may suggest reducing the driving speed or taking a detour.
[0122] Finally, the overall output is a specific instruction sequence for the agents (robots, home appliances), such as the floor sweeper starts and cleans the living room, the robotic arm puts the dirty dishes into the dishwasher, the washing machine automatically dries after washing, or controls the air conditioner to adjust the indoor temperature when necessary.
[0123] As an optimization scheme of the above embodiment, as Figure 1 shown, the execution results of the agents are evaluated and adjusted, and feedback is provided. Optimization and adjustment are carried out in the value preprocessing stage, the collaborative graph construction stage, the large language model planning stage, and the reinforcement learning skill optimization stage respectively.
[0124] Specific embodiment: Home cleaning and parallel operation of home appliances.
[0125] Suppose the user inputs via voice through the smart speaker "Please quickly complete the cleaning of the living room floor, and at the same time help me wash the dishes and clothes clean", the following is a possible implementation process of the system:
[0126] Collaborative graph construction and update:
[0127] The system first collects the status information of each intelligent device. For example, the sweeping robot has sufficient current power, the robotic arm is idle and available, there are a small number of dishes to be washed in the dishwasher, and the washing machine is in standby mode. At the same time, it obtains the home layout information and learns that the living room is adjacent to the kitchen, while the laundry room is at the other end. Subsequently, the system integrates the status, location, and their skill information of these devices into a collaboration graph. By analyzing the collaboration graph, the system detects that the dishwasher and the washing machine can operate simultaneously without resource conflicts, providing an accurate basis for subsequent task planning and scheduling.
[0128] LLM Planning and Scheduling:
[0129] By retrieving the collaboration graph, the system confirms that the sweeping robot can immediately start cleaning the living room, and at the same time, the washing machine and the dishwasher can also be started simultaneously without resource conflicts. Based on this information, the system outputs a high-level task plan as follows: The sweeping robot will immediately start cleaning the living room; the dishwasher will start the quick wash mode and is expected to complete the dishwashing task within 15 minutes; the washing machine will start the standard wash mode and is expected to take 30 minutes to complete the laundry task; in addition, the robotic arm is on standby and will provide assistance at any time if obstacles need to be moved during the task execution.
[0130] Reinforcement Learning Low-Level Execution:
[0131] The sweeping robot uses reinforcement learning strategies to determine the most efficient cleaning path and the most suitable suction settings to ensure the thoroughness and efficiency of the cleaning work. The dishwasher can automatically select the best washing process according to the type and degree of dirt of the tableware to achieve precise cleaning. When starting, the washing machine uses built-in sensors to detect the weight of the clothes and dynamically adjusts the required water volume and rotation speed during washing to achieve the best washing effect while saving resources.
[0132] The intelligent agent home service collaboration method proposed by the present invention is applicable to a variety of scenarios, aiming to improve efficiency and convenience through the collaborative work of intelligent household appliances and service robots. In the home scenario, the system can coordinate multiple intelligent household appliances and robots to jointly complete household tasks. In nursing homes and medical places, the system can integrate multiple auxiliary devices, such as medicine delivery robots, laundry equipment, and care robots, to improve the efficiency of nursing work. In public places, such as shopping malls or airports, the system can dispatch a queue of sweeping robots for zoned cleaning, use large automatic dishwashers to centrally process tableware, and cooperate with other transportation robots, thus significantly improving the overall service efficiency.
[0133] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A multi-agent service collaboration method based on collaboration graph and large language model, characterized in that: Includes steps: S10, collecting environmental perception information including multiple intelligent agents, human-machine collaboration relationship information and user needs, and performing preprocessing; S20, constructing a collaboration graph for the preprocessed data to obtain a collaboration graph; S30, planning the collaboration graph with a large language model to obtain a high-level task plan; S40, optimizing the high-level task planning through reinforcement learning skills to obtain specific actions or execution instructions; S50, sending specific actions or execution instructions to the corresponding agent for execution.
2. According to claim 1, a multi-agent service collaboration method based on collaboration graph and large language model is characterized in that: In step S10, data collection includes: Environmental perception information includes: home layout data, voice or text commands, and environmental sensor data; Human-machine collaboration relationship information: family members’ needs, mutual exclusion or dependency between intelligent devices, and historical interaction logs; User demand: The task request instruction issued by the user to the intelligent device.
3. According to claim 1, a multi-agent service collaboration method based on collaboration graph and large language model is characterized in that: In step S10, data preprocessing includes: Data cleaning: removing outliers and noise from sensor data; Data alignment: Time-align data from different sensors and user inputs; Data normalization: normalize the data from different sensors to the same scale; Feature extraction: extract useful features from images captured by the camera, extract voice features from voice signals received by the microphone, and extract key features from device status sensors; Data fusion: Data fusion is the process of combining data from different sensors and user inputs to generate a comprehensive data view; Data labeling: perform data labeling for some common tasks and scenarios.
4. According to claim 1, a multi-agent service collaboration method based on collaboration graph and large language model is characterized in that: In step S20, during the collaboration diagram construction process, the steps include: S201, obtaining sensor and home appliance status data, and obtaining home appliance and user prior information; S202, storing the processed data in the form of an adjacency matrix or a graph database to construct a collaboration graph; In the collaboration graph, people, various intelligent agents, and key locations or resources in the home environment are defined as a node with attached attribute information; at the same time, the collaboration graph also records the collaboration edges between nodes, indicating the reachability, interaction relationship or task dependency between devices, with attached attribute information.
5. A multi-agent service collaboration method based on collaboration graph and large language model according to claim 4, characterized in that: The collaboration diagram is dynamically updated based on changes in device status, updates to user requirements, and possible failures or energy replenishment events.
6. The multi-agent service collaboration method based on collaboration graph and large language model according to claim 1, characterized in that: In step S30, a large language model planning is performed, including the steps of: S301, obtaining the established collaboration diagram, obtaining the user intention or task instruction; S302, using the GraphRAG principle, retrieve the collaborative graph G to obtain the context information Context related to the user query UserQuery, and input this information together with the user query into the large language model LLM to generate a detailed high-level task plan Plan; Context=Retrieve(G,UserQuery) Plan=LLM(UserQuery+Context); Among them, Retrieve represents the search function: High-level task planning includes: a multi-step task sequence that specifies what actions each agent should perform at a specific time; scheduling results and subtask allocation.
7. A multi-agent service collaboration method based on collaboration graph and large language model according to claim 6, characterized in that: When generating a plan, LLM develops an optimal or suboptimal task execution plan based on the constraints of multi-device collaboration.
8. The multi-agent service collaboration method based on collaboration graph and large language model according to claim 6, characterized in that: Based on new instructions from the user or changes in equipment status, the collaboration diagram is re-called and the original plan is dynamically adjusted.
9. The multi-agent service collaboration method based on collaboration graph and large language model according to claim 1, characterized in that: In step S40, the reinforcement learning skill optimization process includes: S401, obtaining high-level task planning and graph embedding information; S402, multi-device reinforcement learning strategy: Each device or robot has its own exclusive or shared strategy network, and the graph embedding is integrated into the state representation during training; S403, the output content is a specific action or device control instruction, which is generated based on the high-level decomposed subtasks and graph embedding information.
10. The multi-agent service collaboration method based on collaboration graph and large language model according to claim 1, characterized in that: The execution results of the intelligent agent are evaluated and adjusted, and feedback is provided to optimize and adjust the feedback value preprocessing stage, collaboration graph construction stage, large language model planning stage, and reinforcement learning skill optimization stage.
Citation Information
Cited By
Multi-agent interaction intention understanding and cooperative control method based on large model
CN120952060A
Long-range planning method and system of open body environment fusion model and symbol solver
CN120996196A