Cooperative control method, device, equipment, medium and program product of intelligent agent
By adopting decentralized architecture and iterative task management methods in multiple embodied agent systems, the behavior adjustment problem of embodied agents in unforeseen circumstances is solved, and the task success rate and flexibility are improved.
Patent Information
- Application Number
- CN202510105513.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
The existing multi-embolic agent collaboration frameworks rely too much on predefined behavior patterns and central control units, making it difficult for embolic agents to adjust their behavior in unforeseen circumstances, affecting the successful execution of tasks.
Adopting a decentralized intelligent body architecture, each embodied intelligent body is divided into an interactive end and a task end. Through efficient internal communication interface collaboration, the task end of the embodied intelligent body adopts a closed-loop system design of task planning, single-task execution and task re-planning to achieve iterative task management.
It improves the success rate of task completion, enhances the ability of the embodied agent to deal with emergencies, and improves the efficiency and flexibility of task execution.
Smart Images

Figure CN120066712A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, equipment, medium and program product for collaborative control of embodied agents. Background Art
[0002] In today's rapidly developing technological environment, the technology of embodied agents has become an important part of the field of artificial intelligence. An embodied agent refers to a software or hardware entity that can operate autonomously in a specific environment and make appropriate responses according to environmental changes. With the improvement of computing power and the progress of algorithms, multi-agent systems (MAS) have gradually become a research hotspot, especially showing great potential in complex task processing, distributed computing, and human-computer interaction.
[0003] Currently, a variety of multi-agent collaborative frameworks have been proposed and applied in different scenarios. These frameworks usually rely on predefined rules and strategies to guide the collaborative behavior between embodied agents. For example, some frameworks utilize a centralized control mechanism, where one or more central nodes are responsible for coordinating the actions of each embodied agent.
[0004] However, the inventors found that the prior art has at least the following problems: due to excessive reliance on pre-set behavior patterns or central control units, once encountering unforeseen situations or environmental changes, the embodied agents may not be able to effectively adjust their own behaviors to adapt to new challenges, thus affecting the successful execution of the overall task. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a method, device, equipment, medium and program product for collaborative control of embodied agents, which adopts a decentralized agent architecture and optimizes the task execution process of each embodied agent, improving the success rate of task completion and enhancing the ability of embodied agents to handle emergencies.
[0006] To achieve the above purpose, the embodiments of the present invention provide a method for collaborative control of embodied agents, which is applied to a single embodied agent. The method includes:
[0007] Receiving user input information, and processing the user input information to generate a task request corresponding to the user input information;
[0008] Decomposing and preliminarily planning the task request according to preset individual cognition and preset individual capabilities to generate a sub-task queue; wherein, the sub-task queue includes a number of sub-tasks;
[0009] Sequentially executing each sub-task in the sub-task queue;
[0010] After each subtask is completed, update the subtask queue according to the completion status of the subtask, the preliminary plan, and the environmental changes, until all subtasks in the subtask queue are executed.
[0011] As an improvement to the above solution, after all subtasks in the subtask queue are executed, the method further includes:
[0012] Summarize the completion status of the subtask queue, generate task result feedback information, and output the task result feedback information.
[0013] As an improvement to the above solution, the individual cognition includes individual role cognition and individual environment cognition; the individual abilities include but are not limited to query ability, observation ability, movement ability, and communication ability.
[0014] As an improvement to the above solution, after receiving the user input information, the method further includes:
[0015] Process the user input information, generate a reply message corresponding to the user input information, and output the reply message.
[0016] As an improvement to the above solution, the method further includes:
[0017] After each task request is completed, iteratively optimize the behavior model of the embodied agent itself according to historical processing data.
[0018] As an improvement to the above solution, the method further includes:
[0019] When the knowledge content in the individual knowledge base of itself meets the preset knowledge sharing condition, summarize and generalize the knowledge content in the individual knowledge base through a large language model, and send the summarized and generalized knowledge content to the group knowledge base in the form of a message pool; where the group knowledge base refers to the knowledge base of a group composed of multiple embodied agents.
[0020] An embodiment of the present invention further provides an embodied agent collaborative control device, which is applied to a single embodied agent. The device includes an interaction end and a task end. The interaction end includes a text generation module; the task end includes a task planning module, a task execution module, and a task replanning module;
[0021] The text generation module is used to receive user input information and process the user input information to generate a task request corresponding to the user input information;
[0022] The task planning module is used to decompose and preliminarily plan the task request according to the preset individual cognition and the preset individual ability to generate a sub-task queue; wherein, the sub-task queue includes a number of sub-tasks;
[0023] The task execution module is used to sequentially execute each sub-task in the sub-task queue;
[0024] The task replanning module is used to update the sub-task queue according to the completion status of the sub-task, the preliminary plan and the environmental change situation after each sub-task is completed until all the sub-tasks in the sub-task queue are executed.
[0025] An embodiment of the present invention further provides an embodied agent collaborative control device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the embodied agent collaborative control method described in any one of the above is implemented.
[0026] An embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the embodied agent collaborative control method described in any one of the above.
[0027] An embodiment of the present invention further provides a computer program product, which includes a computer program or computer instructions. When the computer program or the computer instructions are executed by a processor, the embodied agent collaborative control method described in any one of the above is implemented.
[0028] Compared with the prior art, the embodied agent collaborative control method, device, equipment, medium and program product disclosed in the present invention adopt a decentralized architecture. Each embodied agent is divided into two main parts: an interaction end and a task end. The two cooperate closely through an efficient internal communication interface, which can provide a seamless human-computer interaction experience and an efficient task completion ability. Moreover, the task end of the embodied agent adopts a closed-loop system design of task planning, single-task execution and task replanning, so that a new round of evaluation and adjustment will be triggered after each sub-task is completed. This iterative task management method not only improves the success rate of task completion, but also enhances the ability of the embodied agent to cope with emergencies. Each embodied agent has its own individual cognition and domain knowledge, and independently conducts task planning and decision-making according to its own expertise and environmental conditions. Based on individual cognition and domain knowledge, each embodied agent can reasonably allocate tasks according to its own expertise, avoiding resource waste and significantly improving the efficiency and flexibility of the embodied agent to complete tasks. Description of the Drawings
[0029] Figure 1 It is a schematic flowchart of a method for collaborative control of embodied agents provided by an embodiment of the present invention;
[0030] Figure 2 It is a schematic flowchart of the working process of an embodied agent in an embodiment of the present invention;
[0031] Figure 3 It is a schematic structural diagram of a device for collaborative control of embodied agents provided by an embodiment of the present invention;
[0032] Figure 4 It is a schematic structural diagram of a device for collaborative control of embodied agents provided by an embodiment of the present invention. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0034] In the description of the present application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.
[0035] The terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "plurality" is two or more.
[0036] In the description of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0037] See Figure 1 , which is a schematic flowchart of a method for collaborative control of embodied agents provided by an embodiment of the present invention. An embodiment of the present invention provides a method for collaborative control of embodied agents, which is applied to a single embodied agent. The method specifically includes steps S11 to S14:
[0038] S11. Receive user input information, process the user input information, and generate a task request corresponding to the user input information;
[0039] S12. Decompose and preliminarily plan the task request according to preset individual cognition and preset individual capabilities to generate a sub-task queue; wherein, the sub-task queue includes several sub-tasks;
[0040] S13. Execute each sub-task in the sub-task queue in sequence;
[0041] S14. After each sub-task is completed, update the sub-task queue according to the completion status of the sub-task, the preliminary plan, and the environmental change situation until all sub-tasks in the sub-task queue are executed.
[0042] In an embodiment of the present invention, in an embodied environment where humans and machines coexist, an innovative multi-agent collaborative framework is proposed in an embodiment of the present invention, so as to equip each user with a specially designed embodied agent. The embodied agent is divided into two main parts: an interaction end and a task end.
[0043] The interaction end focuses on providing personal services to users, including understanding user instructions, providing consultations, answering questions, and assisting in decision-making, etc. It can accurately capture user needs and respond in the most intuitive way. In addition, the interaction end is also responsible for converting user intentions into specific task requests and transmitting them to the task end.
[0044] The task end is the part that acts and executes tasks in the actual environment. According to the task request received from the interaction end, the task end can autonomously plan paths, select tools or devices, and cooperate with other embodied agents to complete the specified work. This end also has the ability to perceive environmental changes and can quickly adjust the action plan when encountering obstacles or emergencies.
[0045] In addition, each embodied agent has individual cognition. Preferably, the individual cognition includes individual role cognition and individual environment cognition, that is, clearly understanding one's own role, position, and scope of responsibilities in the entire system, and also having information about the service object (user). This kind of cognition helps the embodied agent better understand the user's needs, so as to provide more personalized services and support. Moreover, each embodied agent has corresponding individual capabilities, which include but are not limited to query capabilities, observation capabilities, movement capabilities, and communication capabilities. By dispatching these individual capabilities, corresponding tasks can be completed.
[0046] See Figure 2 , which is a schematic diagram of the workflow of the embodied agent in the embodiment of the present invention. Specifically, the interaction end includes a text generation module. As the main communication channel between the user and the embodied agent, this module is responsible for processing user input information, such as questions, commands, etc. When it is determined that there is a task execution requirement after processing the user input information, a task request corresponding to the user input information is generated and sent to the task end.
[0047] The task end includes a task planning module, a task execution module, and a task replanning module. After receiving the task request, the task planning module will decompose and preliminarily plan the task request based on its preset individual cognition and preset individual capabilities, combined with specific task requirements, and create a sub-task queue including several sub-tasks.
[0048] Taking the auxiliary teaching scenario as an example, the individual role cognition of the embodied agent includes roles predefined in the individual knowledge base, such as the secretary of the teacher, the student counselor, etc. The individual environment cognition is the basic cognition of the environment, such as being in a certain school now, being responsible for assisting a certain teacher, and having knowledge about certain classrooms under its jurisdiction, including which students are in them, etc. The individual capabilities are the tools that the embodied agent can call, such as querying the Internet, moving, communicating through a microphone and a speaker, and visual observation through a camera. Then, knowing these inputs, the task planning module makes a preliminary decomposition of the task. For example, for the user input information of the teacher "Check if there is anyone in the classroom?", it is decomposed into 4 steps, including: 1. Move to the classroom; 2. Observe if there is anyone in the classroom; 3. Move back to the teacher's office; 4. Tell the teacher what was observed.
[0049] Next, the task planning module hands over the subtask queue to the task execution module, and the task execution module executes each subtask in the subtask queue in sequence. To ensure flexibility, the task execution module only executes one subtask at a time and reports the result to the task replanning module after completion. After each subtask is completed, the subtask queue is updated according to the completion status of the subtask, the preliminary plan, and the environmental changes until all subtasks in the subtask queue are executed.
[0050] The task replanning module receives the subtask completion status and the preliminary plan from the task execution module, analyzes the current environmental changes, resource availability, and other factors that may affect the task progress, determines whether it is necessary to replan the remaining task sequence, and when it is determined that it is necessary, performs subtask replanning, which can adjust the original plan or add new subtasks, and updates the subtask queue to ensure the successful completion of the entire task.
[0051] As an example, for the task request "Go to the bedroom to see if A is there. If so, tell him to come to the meeting. If not, go to the office and ask B to find him." obtained after processing, the result of the subtask queue formed in the preliminary plan is:
[0052] 1. Go to the bedroom;
[0053] 2. Observe whether A is in the bedroom;
[0054] 3. If A is in the bedroom, tell him to come to the meeting;
[0055] 4. If A is not in the bedroom, go to the office;
[0056] 5. Find B in the office;
[0057] 6. Tell B to find A for the meeting.
[0058] When the embodied agent completes step 1, the subtask queue is updated to:
[0059] 1. Observe whether A is in the bedroom;
[0060] 2. If A is in the bedroom, tell him to come to the meeting;
[0061] 3. If A is not in the bedroom, go to the office;
[0062] 4. Find B in the office;
[0063] 5. Tell B to find A for the meeting.
[0064] When the embodied agent completes the new step 1, since the subsequent task planning is affected by the execution result of the current step, such as "A is not there", the task replanning module replans the subtask queue as:
[0065] 1. Go to the office;
[0066] 2. Find B in the office;
[0067] 3. Instruct B to go to a meeting with A.
[0068] The final task will be completed through such iterations.
[0069] Understandably, the input of the task replanning module is similar to that of the task planning module, with the addition of the historical planning of the task and the planned completion status, and a new plan is driven by the large language model LLM for output.
[0070] Adopting the technical means of the embodiments of the present invention has the following beneficial effects:
[0071] First, a decentralized dual - end embodied intelligent agent architecture:
[0072] The present invention proposes a decentralized task - oriented multi - embodied intelligent agent system, in which each embodied intelligent agent is divided into two main parts: an interaction end and a task end. The interaction end understands the user's intention through a large model (LLM) and provides a high - quality human - machine interaction experience; the task end is responsible for the execution of actual tasks, including physical operation capabilities such as path planning and tool use. The two cooperate closely through an efficient internal communication interface to ensure seamless connection from communication to action. The specific implementation method of this dual - end architecture, especially the design of the efficient communication protocol between the interaction end and the task end and their co - action mechanism, can provide a seamless human - machine interaction experience and an efficient task - completion ability. Moreover, due to the adoption of a decentralized architecture, the present invention can easily adapt to application scenarios of different scales and complexities. When adding new embodied intelligent agents or modifying the existing structure, there is no need to make large - scale adjustments to the entire system, and only relevant modules need to be updated.
[0073] Second, a flexible task execution mechanism:
[0074] The task end of the embodied intelligent agent adopts a closed - loop system design of task planning, single - task execution, and task replanning, so that a new round of evaluation and adjustment will be triggered after each subtask is completed. This iterative task management method not only improves the success rate of task completion but also enhances the ability of the embodied intelligent agent to cope with emergencies.
[0075] Third, improving the task completion rate:
[0076] Each embodied intelligent agent has its own individual cognition and domain knowledge, and independently conducts task planning and decision - making according to its own expertise and environmental conditions. Based on individual cognition and domain knowledge, each embodied intelligent agent can reasonably allocate tasks according to its own expertise, avoiding resource waste and significantly improving the efficiency and flexibility of the embodied intelligent agent in completing tasks.
[0077] It should be noted that existing multi-embodied agent systems often lack the ability to understand and respond to user intentions, resulting in a poor user experience and poor interactivity. In practical applications, when manual intervention or adjustment is required, the flexibility and adaptability of the system are significantly insufficient.
[0078] To solve the problem of poor interactivity of existing embodied agents, as a preferred implementation, the embodiments of the present invention are further implemented on the basis of any of the above embodiments. In step S14, that is, after all subtasks in the subtask queue are completed, the method further includes step S15:
[0079] S15. Summarize the completion status of the subtask queue, generate task result feedback information, and output the task result feedback information.
[0080] As a preferred implementation, the embodiments of the present invention are further implemented on the basis of any of the above embodiments. After receiving the user input information, the method further includes step S16:
[0081] S16. Process the user input information, generate a reply message corresponding to the user input information, and output the reply message.
[0082] The embodiments of the present invention are designed with a dedicated interaction port to solve the problem of poor interaction between embodied agents and human users in traditional frameworks. The interaction port allows each embodied agent to understand and respond to user instructions and feedback, thereby greatly enhancing the flexibility of the system and the user experience. Through this method, the embodied agent can more accurately capture user intentions and adjust its behavior strategy accordingly.
[0083] Specifically, the text generation module serves as the main communication channel between the user and the embodied agent. This module is responsible for processing user input information. When it is determined to be a simple interaction requirement after processing the user input information, the text generation module can generate an appropriate response as needed. It can understand natural language and output answers or instructions that conform to the context logic.
[0084] In the embodiments of the present invention, a feedback loop is also provided. Once all subtasks are completed, the final result will be sent back to the text generation module at the interaction end through the feedback loop. At this time, the text generation module can summarize according to the task completion status and provide detailed feedback information to the user, informing the result of the task and any noteworthy details.
[0085] By adopting the technical means of the embodiments of the present invention, the human-computer interaction experience is enhanced. By introducing a dedicated interaction terminal, including a text generation module, the embodied agent can provide more natural, fluent and personalized services, promptly reply to the user's questions and information such as task completion status, and improve the user's experience.
[0086] As a preferred embodiment, the embodiments of the present invention are further implemented on the basis of any of the above embodiments, and the method further includes step S17:
[0087] S17. After each task request is completed, according to the historical processing data, iteratively optimize the behavior model of the embodied agent itself.
[0088] It should be noted that the existing embodied agent systems lack the ability of self-evolution and learning. Most traditional frameworks do not integrate advanced self-learning mechanisms, which makes it difficult for embodied agents to learn lessons from past experiences and achieve continuous improvement in performance. In the long run, the system will be difficult to cope with the increasingly complex real-world problems.
[0089] In the embodiments of the present invention, the interaction terminal further includes a self-reflection module, which is mainly used to improve the behavior models of the text generation module and the task execution module by using the reflection results, so as to ensure that the embodied agent continuously evolves in a dynamic environment. Whether it is the interaction terminal or the task terminal, it can benefit from the reflection and continuously improve the interaction quality with the user and the task execution effect.
[0090] On the one hand, the self-reflection module can monitor all activities in the interaction process, continuously collect data and evaluate its own behavior. It will form thoughts on the current dialogue state and feedback these thoughts to the text generation module to assist it in more accurately understanding and responding to the user's needs.
[0091] On the other hand, the self-reflection module can also review the entire process after each task is completed, analyze the successes and deficiencies, and optimize the behavior pattern of the task terminal accordingly. These lessons learned become part of the long-term memory to guide subsequent operations.
[0092] Specifically, different from the previous fixed rule sets or limited learning algorithms, the present invention endows each embodied agent with long-term memory storage ability and autonomous learning function. With the execution of each task, the embodied agent will accumulate experience and use this data to continuously improve its own behavior model. This continuous learning process not only helps the embodied agent adapt to newly emerging task types, but also enables them to make more efficient and accurate responses when facing similar problems. In particular, when encountering complex or unforeseen situations, the embodied agent can find the best solution by reviewing past cases or develop new coping strategies.
[0093] To achieve self - evolution, both the interaction end and the task end are given the ability to conduct phased summaries and reflections. This means that after each task is completed, the embodied intelligent agent will review the entire process, analyze the successes and deficiencies, and optimize future behavior patterns accordingly. These lessons learned will become part of the long - term memory and be saved for guiding subsequent operations. The summaries and reflections for both ends include: The summary is manifested as a summary of the history of the phased context.
[0094] As an example, during the process of interacting with the user and finally completing the purchase of medicine, for the task execution end, its historical context starts from receiving the task sent by the interaction end, including the content of the subtask queue output by the task planning module each time, the execution content of the tools used by the embodied intelligent agent, and the subtasks re - planned by the task re - planning module, until a complete process is finally completed. The history of this entire process is handed over to a large - model LLM responsible for summarization for summary, and the summarized result is saved as a reference for input when performing tasks in the next round. The history obtained for reflection is the same as that for summarization, but the thinking angle through the LLM is different. Compared with the event - based summary of the summary, reflection pays more attention to the knowledge and experience formed during the execution process that can optimize future behavior. For example, for the matter of buying medicine, the result reflected by the task end may be that "in the future, when the user has a fever, there is no need to consult at the front desk again, and you can directly register in the internal medicine clinic." Another example is that when a user at the interaction end hopes to forward a certain notice, the interaction end may tend to ask for details clearly, which will lead to a cumbersome process and reduce the user experience. Therefore, the content of the reflection may be that "for notice - type tasks, try to reduce the number of interactions and do not ask unnecessary details."
[0095] By adopting the technical means of the embodiments of the present invention, by introducing a self - reflection module at the interaction end, continuously monitoring and optimizing the interaction process, ensuring that each conversation can more accurately capture the user's intention, the user experience is greatly improved. The present invention particularly emphasizes the self - evolution characteristic of the embodied intelligent agent, that is, improving its own behavior pattern through long - term memory and experience accumulation. Whether at the individual cognitive level or the group knowledge level, the embodied intelligent agent can learn from past experiences and apply this knowledge to future task processing, thereby achieving continuous improvement in performance. The embodied intelligent agent can not only operate in a preset environment but also quickly respond to environmental changes. Through continuous self - reflection and learning, it can maintain an efficient working state under dynamic conditions.
[0096] As a preferred implementation manner, the embodiments of the present invention are further implemented on the basis of any of the above - mentioned embodiments, and the method further includes step S18:
[0097] S18. When the knowledge content in its own individual knowledge base meets the preset knowledge sharing conditions, summarize and generalize the knowledge content of the individual knowledge base through a large language model, and send the summarized knowledge content to the group knowledge base in the form of a message pool; wherein, the group knowledge base refers to the knowledge base of a group composed of multiple embodied agents.
[0098] Preferably, the knowledge sharing conditions are that the knowledge content of the individual knowledge base reaches a preset word count threshold or reaches a preset sharing time period.
[0099] Specifically, the embodiments of the present invention also strengthen the combination of individual and collective wisdom. In addition to the learning and memory of a single embodied agent, the importance of knowledge sharing and collective wisdom of the entire group is also emphasized. When an embodied agent obtains new insights or skills, it can share them with other members to promote the common progress of the whole system. This feature is crucial for building a highly flexible and fast-responsive multi-embodied agent network, especially in a dynamically changing environment.
[0100] Over time, each embodied agent will accumulate a large amount of task experience and environmental information, forming rich individual cognitions. When these individual cognitions converge, they constitute group knowledge. Group knowledge covers all relevant information of the entire embodied environment, as well as the knowledge points mastered by all individual embodied agents. This not only promotes knowledge sharing among individuals. For both individual embodied agents and the group, knowledge bases are set up (stored in the form of a database). When the new knowledge content in the individual's knowledge base reaches a certain threshold (such as 10 items or 500 words), or after a certain period of time has passed (such as one day), the new individual knowledge will be further summarized and generalized through the LLM, and then sent to the database of group knowledge in the form of a message pool. After receiving the new knowledge, the database of group knowledge will update itself. Each individual has access to the group knowledge. They use the user's commands and conversation history as queries to access the group knowledge base, and retrieve a certain amount of relevant knowledge through technologies such as RAG (such as semantic similarity matching) as part of the input for the next interaction or task execution, and also enables the entire system to continuously learn and develop as a whole, adapting to a wider range of application scenarios.
[0101] It should be noted that whether at the individual or group level, embodied agents can gradually improve their own performance through continuous practice and communication. More importantly, this learning is not static, but dynamically responds to changes in the external environment, ensuring that the embodied agents are always in the best state and ready to meet new challenges at any time.
[0102] By adopting the technical means of the invention embodiments, a dynamically updated group knowledge base is formed among all embodied agents, which contains information about the entire embodied environment and service experience. This enables the embodied agents to share information through an efficient communication mechanism, ensuring that all members can promptly obtain the latest task progress and changes. The mutual promotion between individual cognition and group knowledge enables the embodied agents to quickly adapt to new situations and enhance the wisdom of the overall system by sharing information. The conversion mechanism of individual cognition to group knowledge and the dynamic update mechanism of group knowledge can ensure that it always reflects the latest environmental changes and service requirements, thereby optimizing the overall collaboration effect, improving the success rate of tasks, and simultaneously enhancing the adaptability of the system and the efficiency of problem-solving.
[0103] See Figure 3 , which is a schematic structural diagram of an embodied agent collaborative control device provided by the embodiments of the present invention. The embodiments of the present invention also provide an embodied agent collaborative control device 10, which is applied to a single embodied agent. The device 10 includes an interaction end 11 and a task end 12. The interaction end includes a text generation module 111; the task end includes a task planning module 121, a task execution module 122, and a task replanning module 123.
[0104] The text generation module 111 is configured to receive user input information and process the user input information to generate a task request corresponding to the user input information.
[0105] The task planning module 121 is configured to decompose and preliminarily plan the task request according to preset individual cognition and preset individual capabilities to generate a sub-task queue; wherein, the sub-task queue includes a plurality of sub-tasks.
[0106] The task execution module 122 is configured to sequentially execute each sub-task in the sub-task queue.
[0107] The task replanning module 123 is configured to update the sub-task queue according to the completion situation of the sub-task, the preliminary plan, and environmental changes after each sub-task is completed until all sub-tasks in the sub-task queue are executed.
[0108] As a preferred embodiment, the text generation module 111 is further configured to summarize the completion situation of the sub-task queue, generate task result feedback information, and output the task result feedback information.
[0109] As a preferred embodiment, the text generation module 111 is further configured to process the user input information to generate a reply information corresponding to the user input information, and output the reply information.
[0110] As a preferred embodiment, the interaction end further includes a self-reflection module 112;
[0111] The self-reflection module 112 is used to iteratively optimize the behavior model of the embodied agent itself according to historical processing data after each task request is completed.
[0112] It should be noted that an embodied agent collaborative control device provided in an embodiment of the present invention is used to execute all the process steps of an embodied agent collaborative control method in the above embodiment. The working principles and beneficial effects of the two correspond one by one, so they will not be elaborated here.
[0113] See Figure 4 FIG. is a schematic structural diagram of an embodied agent collaborative control device provided in an embodiment of the present invention. An embodiment of the present invention further provides an embodied agent collaborative control device 20, including a processor 21, a memory 22, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the embodied agent collaborative control method described in any one of the above embodiments.
[0114] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the embodied agent collaborative control method described in any one of the above embodiments.
[0115] An embodiment of the present invention further provides a computer program product. The computer program product includes a computer program or computer instructions. When the computer program or the computer instructions are executed by a processor, they implement the embodied agent collaborative control method described in any one of the above embodiments.
[0116] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0117] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A collaborative control method for an embodied intelligent body, characterized in that: Applied to a single embodied agent, the method comprises: Receive user input information, process the user input information, and generate a task request corresponding to the user input information; According to the preset individual cognition and the preset individual ability, the task request is decomposed and preliminarily planned to generate a subtask queue; wherein the subtask queue includes a plurality of subtasks; Executing each of the subtasks in the subtask queue in sequence; After each subtask is completed, the subtask queue is updated according to the completion status of the subtask, the preliminary plan and the environmental changes, until all subtasks in the subtask queue are executed.
2. The embodied intelligent body collaborative control method according to claim 1, characterized in that: After all subtasks in the subtask queue are executed, the method further includes: The completion status of the subtask queue is summarized, task result feedback information is generated, and the task result feedback information is output.
3. The embodied intelligent body collaborative control method according to claim 1, characterized in that: The individual cognition includes individual role cognition and individual environment cognition; the individual ability includes but is not limited to inquiry ability, observation ability, movement ability and communication ability.
4. The embodied intelligent body collaborative control method according to claim 1, characterized in that: After receiving the user input information, the method further includes: The user input information is processed, reply information corresponding to the user input information is generated, and the reply information is output.
5. The embodied intelligent body collaborative control method according to any one of claims 1 to 4, characterized in that: The method further comprises: After each task request is completed, the behavior model of the embodied intelligent agent itself is iteratively optimized based on historical processing data.
6. The embodied intelligent body collaborative control method according to claim 1, characterized in that: The method further comprises: When the knowledge content of one's own individual knowledge base meets the preset knowledge sharing conditions, the knowledge content of the individual knowledge base is summarized and generalized through a large language model, and the summarized and generalized knowledge content is sent to the group knowledge base in the form of a message pool; wherein the group knowledge base refers to the knowledge base of a group composed of multiple embodied intelligent agents.
7. An embodied intelligent collaborative control device, characterized in that: Applied to a single embodied intelligent agent, the device comprises an interaction end and a task end, the interaction end comprises a text generation module; the task end comprises a task planning module, a task execution module and a task re-planning module; The text generation module is used to receive user input information, process the user input information, and generate a task request corresponding to the user input information; The task planning module is used to decompose and preliminarily plan the task request according to the preset individual cognition and the preset individual ability to generate a subtask queue; wherein the subtask queue includes a plurality of subtasks; The task execution module is used to execute each of the subtasks in the subtask queue in sequence; The task re-planning module is used to update the sub-task queue according to the completion status of the sub-task, the preliminary plan and the environmental changes after each sub-task is completed, until all sub-tasks in the sub-task queue are executed.
8. An embodied intelligent collaborative control device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the embodied intelligent body collaborative control method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the embodied intelligent body collaborative control method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product includes a computer program or computer instructions, and when the computer program or the computer instructions are executed by a processor, the embodied intelligent body collaborative control method as described in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Processing methods for embodied agents based on skill graph-enhanced industrial large-scale models
CN122547504A