A humanoid robot dynamic agent control architecture and method

CN122807897APending Publication Date: 2026-09-25LINGNAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611096576.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-23
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提供一种人型机器人动态智能体控制架构及方法,旨在解决现有机器人控制架构在处理需要长期规划、实时交互且易被中断的复杂任务时,存在的规划能力与实时响应能力割裂、多设备协同缺乏物理感知和动态编排能力的问题

Benefits of technology

[0030]1.本申请通过设置快慢模型协同处理单元和决策机制,能够在执行需要深度思考的复杂任务的同时,对用户的简单提问或交互给予快速、流畅的回应,解决了传统智能体技术响应慢的问题,极大地提升了人机交互的自然性和连续性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807897A_ABST
    Figure CN122807897A_ABST
Patent Text Reader

Abstract

The application discloses a human type robot dynamic intelligent agent control architecture and method, and belongs to the fields of robot technology and artificial intelligence. The architecture aims to solve the problem of the separation of the robot planning ability and the real-time response ability. The architecture comprises a processing unit, a message bus for receiving general interaction instructions, and a device bus for transmitting device control instructions. The processing unit processes the instructions on the device bus with a higher priority than the message bus, and the processing unit comprises a fast model, an Agentic slow model, and a decision mechanism. Through the fast-slow model cooperation and the double-bus priority design, the application unifies the robot planning ability and the real-time performance, realizes the intelligent and autonomous arrangement of multiple devices, and enhances the task robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of robotics and artificial intelligence, specifically to a control architecture and method for a humanoid robot dynamic intelligent agent. Background Technology

[0002] The current design of the control core of humanoid robots, their "interactive brain," exhibits limitations and varying interaction methods when interacting with the physical world, humans, and the robot's own multiple devices. Existing technologies primarily rely on visual-language-action models (VLAMA) and intelligent agent technology. VLAMA can perform continuous physical actions based on simple, clearly defined instructions, such as picking up an object. However, it cannot understand complex instructions with ambiguous intentions, struggles with long-term planning, and often blocks other interaction channels during action execution, preventing the robot from handling other tasks simultaneously. Intelligent agent technology excels at long-term task planning and is widely used in software and information processing. However, when directly applied to robots, it suffers from poor real-time interaction capabilities and insufficient understanding of the physical world. Its inherent loop processing mechanism often leads to high response latency and fails to adequately consider the physical state dependencies and temporal causal relationships between the robot's various hardware devices.

[0003] Furthermore, while some existing robot operating systems allow for modular function deployment and can manage tasks of varying priorities, they are essentially task execution frameworks. They lack a high-level cognitive core to unify the understanding of complex human instructions and to enable autonomous planning and real-time interaction in dynamic, uncertain real-world environments. These shortcomings collectively result in existing robots' inability to effectively translate advanced, complex human instructions into coherent, intelligent behavior when handling complex tasks requiring long-term planning, real-time interaction, and prone to interruption. This leads to a disconnect between planning capabilities and real-time response capabilities, as well as a lack of physical perception and dynamic orchestration capabilities in multi-device collaboration. Summary of the Invention

[0004] The purpose of this invention is to provide a humanoid robot dynamic intelligent agent control architecture and method, which aims to solve the problems of existing robot control architectures in handling complex tasks that require long-term planning, real-time interaction and are easily interrupted, such as the separation of planning ability and real-time response ability, and the lack of physical perception and dynamic orchestration ability for multi-device collaboration.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A control architecture for a humanoid robot dynamic intelligent agent includes:

[0007] The processing unit is configured to execute the agent loop;

[0008] The message bus is used to receive general interactive commands;

[0009] The device bus is used to transmit commands related to device control.

[0010] The processing unit is configured to process device control-related instructions on the device bus with a higher priority than processing general interactive instructions on the message bus.

[0011] Furthermore, the processing unit includes:

[0012] Fast model;

[0013] The Agentic slow model and the decision-making mechanism serve as a voice playback acceleration module, used to coordinate the output of the fast model and the Agentic slow model based on a streaming block processing mechanism. The fast model prioritizes the output of transitional dialogue to achieve a fast response, while the Agentic slow model performs task processing in the background and generates detailed data to fill the transitional dialogue in blocks, thereby masking the processing delay.

[0014] As a further aspect of the present invention, the Agentic slow model is further configured to parse the device control-related instructions into at least two device tasks, and call a dynamic orchestration model to dynamically sort or rearrange the execution order of the at least two device tasks.

[0015] As a further aspect of the present invention: the Agentic slow model is also configured to invoke a toolset, the toolset including a Visual Language Action (VLA) model service, and the Agentic slow model performs physical actions by sending instructions to the device bus to invoke the VLA model service.

[0016] As a further aspect of the present invention, it also includes a multidimensional memory system, which is used to support the Agentic slow model in executing the device task and / or to resume the device task after the device task is interrupted.

[0017] As a further aspect of the present invention, the multidimensional memory system includes at least one of the following: a device operating status record for task recovery, a user session history record for personalized interaction, or a visual memory system for long-term environmental perception.

[0018] As a further aspect of the present invention, the dynamic orchestration model is configured to: during task execution, in response to detected environmental changes, trigger a rearrangement of the task execution order of the device.

[0019] As a further aspect of the present invention: the dynamic orchestration model is configured to: when planning a task sequence, identify the preceding device state required to execute subsequent tasks, determine the current device state, and when the current device state conflicts with the preceding device state, automatically insert a state rollback task to make the device state satisfy the preceding device state.

[0020] As a further aspect of the present invention: the decision-making mechanism is configured to, after determining that the fast-slow collaborative processing mode has been entered, schedule the fast model to output a response template based on the local context, and schedule the Agentic slow model to execute tool calls based on the global historical context, so as to obtain and stream-fill the detailed data.

[0021] This invention also provides a method for controlling a humanoid robot's dynamic intelligent agent, comprising the following steps:

[0022] It receives general interactive commands through a message bus and transmits device control-related commands through a device bus.

[0023] Device control-related instructions on the device bus are processed with a higher priority than general interaction instructions on the message bus.

[0024] After determining that the fast-slow collaborative processing mode has been entered, the streaming block collaborative processing logic is activated; the fast model generates and outputs a transitional voice response template based on the local context; the Agentic slow model performs task processing in the background and streams the generated detailed data into the voice response template to generate the final response and mask the processing delay.

[0025] As a further aspect of the present invention, the method further includes:

[0026] During the execution of the first task processed by the Agentic slow model, a second instruction is received as an interrupt instruction.

[0027] The second instruction is distributed to the fast model for processing and a response is output.

[0028] After the fast model completes its response, context information representing the state of the first task before it was interrupted is read from a memory system. The context information includes the task execution steps and related parameters. Based on the context information, the Agentic slow model continues to execute the first task.

[0029] Compared with the prior art, the present invention has at least the following beneficial effects:

[0030] 1. By setting up a fast and slow model collaborative processing unit and decision-making mechanism, this application can provide a fast and smooth response to simple questions or interactions from users while performing complex tasks that require deep thinking. This solves the problem of slow response in traditional intelligent agent technology and greatly improves the naturalness and continuity of human-computer interaction.

[0031] 2. This application addresses the lack of physical perception and dynamic planning capabilities in multi-device robot collaboration through a dynamic orchestration model and a high-priority device bus. The robot can autonomously and dynamically plan and adjust its operation sequence based on device status and task dependencies to adapt to dynamic and easily interrupted real-world environments, without requiring manual pre-setting of all processes.

[0032] 3. This application encapsulates physical execution capabilities such as visual language action models into tools that can be called by the Agentic slow model, thus opening up a complete link from high-level semantic understanding and long-term task planning to the execution of specific physical actions, forming a unified embodied intelligent brain that can effectively transform advanced and complex human instructions into coherent and intelligent robot behaviors.

[0033] 4. Enhanced task robustness. This application combines a multidimensional memory system and a dynamic orchestration model, enabling the robot to not only resume the task after it is interrupted by external factors, but also intelligently adjust subsequent steps according to environmental changes, significantly improving the robot's success rate in completing tasks in complex scenarios. Attached Figure Description

[0034] The invention will now be further described with reference to the accompanying drawings.

[0035] Figure 1 This is a schematic diagram of the system architecture of a humanoid robot dynamic intelligent agent control architecture provided in an embodiment of this application.

[0036] Figure 2 This is a flowchart illustrating a dynamic intelligent agent control method for a humanoid robot provided in an embodiment of this application.

[0037] Figure 3 This is a timing diagram of signaling interaction for task execution and interruption recovery provided in an embodiment of this application. Detailed Implementation

[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] In the description of this invention, it should be understood that the terms "upper", "lower", "left", "right", "front", "rear", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or a specific orientational structure and operation. Therefore, they should not be construed as limitations on this invention.

[0040] Furthermore, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking," etc., should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0041] Example 1

[0042] Please see Figure 1 As shown, this embodiment provides a dynamic intelligent agent control architecture for a humanoid robot. This architecture can be physically deployed on a combination of the humanoid robot's local computing unit (e.g., an onboard computer) and a cloud server, or deployed independently of either. Figure 1 As shown, the architecture includes an input gateway, a message bus (input), an Agent Loop processing unit, a device bus, a message bus (output), a Feynman device orchestration module, a visual language action model module, a memory module, and an output layer.

[0043] The input gateway, serving as the entry point for the system to interact with the external world, is configured to receive input signals from various channels. For example, the input gateway could be a microphone array with integrated speech recognition capabilities to receive user voice commands; it could also be a software interface connected to an instant messaging application to receive text message commands; or it could be an interface connecting to sensors inside the robot, such as cameras, timed task triggers, or heartbeat signals used to detect specific visual intentions.

[0044] The message bus-in and message bus-out together constitute a general message passing channel. As an optional implementation, this message passing channel can be based on a first-in, first-out queue. The message bus-in is responsible for collecting and buffering all general interactive commands and data from the input gateway, ready for processing by the Agent Loop processing unit. Correspondingly, the message bus-out is responsible for receiving the processing results from the Agent Loop processing unit and distributing them to the corresponding output modules.

[0045] The device bus is a dedicated channel for transmitting commands related to robot hardware control. Unlike general-purpose message buses, device buses have a higher processing priority. The Agent Loop processing unit is configured to prioritize processing commands on the device bus, thereby ensuring timely and reliable execution of control over critical hardware such as the robot arm, camera, and mobile chassis. It should be noted that this dual-bus priority design effectively isolates the device control flow, which has extremely high real-time requirements, from the interactive information flow, which has higher versatility requirements, and is one of the core features of this application.

[0046] The Agent Loop processing unit, as the core processing unit of this architecture, is responsible for executing the agent loop. It continuously retrieves instructions from the message bus and device bus, and performs decisions and processing. To balance real-time interaction with planning depth, this Agent Loop processing unit integrates an innovative fast-slow model collaborative processing mechanism. Specifically, this mechanism includes a fast model, an agentic slow model, and a PK decision-making mechanism.

[0047] The Fast Model is a lightweight, optimized language model designed to achieve rapid response and excels at handling simple interactions that do not require complex reasoning or tool calls, such as casual conversation and simple question-and-answer sessions.

[0048] The Agentic slow model is a more powerful agent model with planning and reasoning capabilities. It handles complex tasks requiring multi-step planning, invoking external tools, or deep thinking. To accomplish such tasks, the Agentic slow model has access to a toolset. This toolset encapsulates various capabilities, including but not limited to: web search tools for information retrieval, interpreter tools for executing code, and device control tools for controlling the physical movements of robots.

[0049] The PK decision mechanism acts as a streaming output collaboration module (or voice playback acceleration module) between the fast and slow models. When the Agent Loop processing unit receives new instructions from the bus and completes intent and attribute analysis, if it determines to enter the fast-slow collaborative processing mode, the PK decision mechanism is responsible for scheduling the output. Specifically, the PK decision mechanism controls the fast model to prioritize the output of transitional dialogue for a rapid response. After the Agentic slow model calls tools (e.g., performs physical actions, queries external information) to obtain in-depth data, it streams detailed data into the fast model's response, thus effectively masking the background processing latency of complex tasks.

[0050] The Feynman device orchestration module, visual language action model module, and memory module are key functional modules that provide advanced capability support for the Agent Loop processing unit.

[0051] The visual language motion model module is a service specifically designed to perform concrete physical actions. In one embodiment of this application, the visual language motion model module can be deployed as an independent service, receiving instructions from the device bus and converting them into control signals for the robot's joint motors, thereby driving the robot to complete physical actions such as grasping, moving, and manipulating. The Agentic slow model interacts with the visual language motion model module through device control tools in its toolset, in the form of tool calls, thereby decoupling and connecting high-level task planning with low-level physical execution.

[0052] The memory module provides the robot with both long-term and short-term memory capabilities. It is a multi-dimensional memory system that can include at least: device operating status records to support recovery after task interruption; user session history records to provide personalized interactive experiences; and a visual memory system for long-term environmental perception, such as recording facial information and object locations.

[0053] The Feynman device orchestration module, another core innovation of this application, is responsible for the intelligent and dynamic orchestration of complex tasks involving multiple devices. After the Agentic slow model decomposes a complex task into a series of subtasks (i.e., device tasks), it invokes the Feynman device orchestration module to determine the optimal execution order of these subtasks. This module understands the physical dependencies and state constraints between devices, thereby generating a reasonable and efficient execution plan.

[0054] The output layer is responsible for presenting the system's processing results to the user or environment, such as through a text-to-speech module for voice broadcasting or through an instant messaging interface to reply to messages.

[0055] The following will combine Figure 2 and Figure 3 The workflow of the above architecture will be described through a specific scenario. Figure 2 This is a flowchart illustrating the method described in this application. Figure 3 This is a timing diagram of the signaling interaction between task execution and interruption recovery in this scenario.

[0056] Scenario: A user stands in front of a humanoid robot and says to it, "Please explain this coffee machine to me and then make me a latte."

[0057] Receive and parse instructions (corresponding) Figure 2 (S100): The user's voice is captured by the microphone of the input gateway and converted into text by the speech recognition module: "Please explain this coffee machine to me and then make me a latte." This text instruction is encapsulated into a message and placed in the message bus -in.

[0058] Intent recognition and pattern determination (corresponding) Figure 2 (S200): The Agent Loop processing unit obtains the instruction from the message bus input and identifies that the instruction contains complex intentions involving multiple steps such as "explain" and "make latte", and determines to enter the fast and slow collaborative processing mode.

[0059] Fast and slow collaborative processing paths (corresponding) Figure 2 (S210 and S400): After determining that the process has entered cooperative mode, the process is executed in parallel in two paths:

[0060] 1. Fast model streaming output transition response (corresponding to) Figure 2 S210): The PK decision-making mechanism scheduling fast model is based on local context and instantaneously streams transitional voice response templates (such as "Okay, no problem") to achieve real-time interaction.

[0061] 2. Agentic slow model processing (corresponding) Figure 2 S400 (China): Meanwhile, the Agentic slow model synchronously performs task processing in the background, including calling the Feynman model for minimum-energy global orchestration (corresponding to...). Figure 2 (S420) to determine the optimal sequence for physical execution.

[0062] Detailed data streaming and output (corresponding) Figure 2 S450 and S500 (in China): After the Agentic slow model obtains the detailed data, the process merges into the detailed data streaming fill step (corresponding to...). Figure 2 (S450). The decision-making mechanism controls the streaming filling of detailed data generated by the slow model into the aforementioned transitional response template, generating the final response and outputting it (corresponding to...). Figure 2 (S500), thereby fully masking background processing delays by utilizing a parallel collaboration mechanism.

[0063] Complex task planning and orchestration (corresponding) Figure 2 (S410 and S420): After receiving a task, the Agentic slow model begins to think and plan. First, it breaks down the high-level instruction into a sequence of device tasks, such as: Subtask 1: Move to the coffee machine; Subtask 2: Activate the vision module to identify coffee machine parts; Subtask 3: Generate and organize explanatory text; Subtask 4: Perform the explanatory actions (combining voice and body language); Subtask 5: Pick up the cup; Subtask 6: Operate the coffee machine to make a latte; Subtask 7: Deliver the made latte to the user. After decomposition, the Agentic slow model calls the Feynman device orchestration module to perform minimum-energy global orchestration (corresponding to...). Figure 2(S420). This module dynamically sorts or rearranges the task sequence by comprehensively calculating a global energy function that incorporates multi-dimensional parameters such as power consumption prediction, spatiotemporal location, equipment status, and subjective costs, to ensure the optimality and continuity of the physical execution path.

[0064] Execute the first subtask (corresponding to) Figure 2 (S430): After planning is complete, the Agentic slow model begins executing its first subtask, "Move to the coffee machine." It generates a command to the visual-language-motion model module using the device control tools in its toolset. This command is sent to the high-priority device bus. After receiving the command from the device bus, the visual-language-motion model module begins controlling the robot's chassis to move towards the coffee machine. This process... Figure 3 In this context, the Agent Loop processing unit sends a message "Invocation: Execute subtask A1" to the visual language action model service through its internal Agentic slow model.

[0065] Handling User Interruptions: During the robot's movement, the user issues a new voice command: "What's the weather like today?". This new command enters the message bus again through the input gateway. The Agent Loop processing unit receives this new command and performs intent analysis. It determines that this is a simple information query, requiring no complex tools and having high real-time requirements. Therefore, this command is directly handed over to the fast model for rapid processing. This decision corresponds to the processing path in Figure 2 that switches to S300 after S200. This interruption process is represented in Figure 3 by the user sending the message "req: Ask a simple question B (interruption)" to the Agent Loop processing unit.

[0066] Fast interruption response: The fast model quickly handles the issue, potentially calling an internal weather query function—a lightweight function call not falling under the "tool call" category of the Agentic slow model—and generating the response "Today's weather is sunny, 25 degrees Celsius." This response is then broadcast aloud via the text-to-speech module in the output layer. The entire interruption handling process is extremely short, with virtually no impact on the user experience. Figure 3 In this context, the Agent Loop processing unit distributes question B to the fast model, the fast model returns answer B, and the Agent Loop then outputs answer B to the user. During the fast model's processing, the main task (making coffee) handled by the Agentic slow model is temporarily suspended.

[0067] Task Resumption and Continuation: After answering the weather question, the system needs to resume the main task that was interrupted. At this point, the Agentic slow model accesses the memory module. Understandably, the memory module records complete context information about the main task before it was interrupted, such as the currently executing subtask ("move to the coffee machine"), the overall planned sequence of the task, and relevant parameters (e.g., the coordinates of the target location). Based on this context information read from the memory module, the Agentic slow model seamlessly continues execution of the main task. Figure 3 In this scenario, after the interruption is handled, the visual language action model service completes the movement task and returns a "callback: Subtask A1 completed" message to the Agentic slow model. The Agentic slow model then restores the context from its memory system and continues to call the visual language action model service to execute subsequent steps of the task. The robot continues to move to the coffee machine and sequentially performs the subsequent explanation and coffee-making tasks until all user instructions are completed.

[0068] As can be seen from this embodiment, the architecture proposed in this application successfully transforms a high-level, complex instruction into a series of coherent and intelligent behaviors of the robot, while smoothly handling real-time interruptions from the user and seamlessly resuming the main task after the interruption ends, demonstrating high interactivity and strong task execution robustness.

[0069] Example 2

[0070] This embodiment focuses on how the Feynman device orchestration module dynamically orchestrates device tasks based on the lowest energy prediction in the physical world when facing a highly disrupted physical environment for a humanoid robot. It also demonstrates its advanced embodied intelligence characteristics, achieving continuous task execution and short-term "defense" of unreasonable instructions through energy function evaluation while ensuring physical rationality. The system architecture of this embodiment is consistent with the architecture shown in Figure 1 of Embodiment 1.

[0071] Scenario: Physical interaction and interference-resistant orchestration based on minimum energy prediction

[0072] The user gave the robot the instruction: "Go pick up those aluminum cans over there and throw them away."

[0073] 1. Command reception and intent parsing (corresponding to S100 and S200 in Figure 2)

[0074] User commands enter the system through the input gateway. The Agent Loop processing unit analyzes the commands and identifies the intent, determining that the command contains complex intents involving multi-physical device collaboration, such as "move," "identify," "grab," and "deploy." During this process, the PK decision mechanism intervenes as a voice playback acceleration module. It is not responsible for command analysis and intent identification, but instead uses a streaming block processing mechanism to schedule the fast model to output transitional voice response templates (such as "Okay, I'll handle it right away") to respond to the user immediately, masking the delay of the subsequent deep planning by the Agentic slow model.

[0075] 2. Task decomposition (executed by the Agentic slow model, corresponding to S410 in Figure 2)

[0076] Upon receiving the task objective, the Agentic slow model breaks it down into a series of executable tasks. In this embodiment, these tasks include activating the depth camera, invoking chassis navigation, and driving the robotic arm to perform grasping. It should be noted that the Agentic slow model is only responsible for constraining high-level instructions into task slices of device actions, and not for the specific execution timing order.

[0077] 3. Feynman energy function arrangement (corresponding to S420 in Figure 2)

[0078] The disassembled set of device tasks is input into the Feynman device orchestration module. This module simulates the execution process as a game on a 3D chessboard, which includes environmental objects and the robot's own devices (equipment). Each device is configured to perform only a limited set of actions within its action set.

[0079] The Feynman device orchestration module constructs and optimizes a global energy function by comprehensively calculating multi-dimensional parameters such as target spatiotemporal location, device status (represented by vectors), action set, time sequence, power consumption prediction, and subjective economic cost, in order to predict the lowest energy execution order in the physical world.

[0080] Multi-order energy modeling: Taking "grabbing an aluminum can" as an example, if the depth camera detects a distant object, the 2-pair energy function of "visual localization" and "hand grasping" is evaluated. This will result in an extremely high energy cost.

[0081] Path dynamic generation: To minimize energy, the Feynman model automatically inserts navigation nodes into the sequence by calculating 3-pair energy combinations. It was found that shortening the physical distance by moving the chassis can significantly reduce the total energy consumption. As a result, the Feynman model generates a logically sound and physically economical initial sequence: [visual recognition] -> [chassis navigation to target point] -> [robotic arm grasping].

[0082] 4. Mechanism for handling environmental interference and task "defiant" (corresponding to "detecting interruption or environmental change" in Figure 2): During the robot's "chassis navigation to the target point", the user suddenly issues an interfering command: "Forget about that, come and sweep the floor for me right now!"

[0083] Real-time energy reassessment: The interrupt command is quickly broken down into a set of "sweeping" tasks by the Agentic slow model and sent to the Feynman module. The Feynman model immediately compares the energy consumption of the two paths: "interrupting the current task to switch to sweeping" and "completing the current closed loop before sweeping".

[0084] Physical rationality decision: Calculations show that if the interruption command is executed immediately, the economic and subjective costs of forcibly braking, switching toolsets, and abandoning the located target will be extremely high because the robot is in motion inertia, which will lead to a significant reduction in power utilization efficiency.

[0085] Short-term "disobedience": Based on the principle of minimum global energy, the Feynman model determines that the current interruption is unreasonable. Unlike the traditional logic of blindly following human instructions, this architecture exhibits short-term "disobedience" behavior here, that is, it does not immediately execute the "sweeping" action, but maintains the continuity of the current "discarding the can" task.

[0086] Collaborative Feedback: At this point, the PK decision-making mechanism once again plays a role in accelerating the response, with a streaming voice output: "I will first finish cleaning up the cans, and then sweep the floor for you immediately," thus maintaining smooth human-computer interaction while ensuring physical rationality.

[0087] This embodiment addresses the issues of fragmented actions and energy waste caused by frequent interruptions in the dynamic physical world by utilizing the energy assessment mechanism of the Feynman device orchestration module. The robot is no longer passively receiving verbal commands, but can proactively predict the energy cost of physical interactions, ensuring successful and efficient task processing even in highly disruptive environments through scientific step-by-step planning.

[0088] Example 3

[0089] This embodiment will analyze in detail the PK decision-making mechanism inside the Agent Loop processing unit, explaining how it intelligently and efficiently coordinates between the fast model and the Agentic slow model based on different types of user commands, thereby achieving the optimal balance between response speed and processing depth in the system at a macro level. The description in this embodiment will focus on the internal logic and interaction of the message bus input, PK decision-making mechanism, fast model, Agentic slow model, and their toolset.

[0090] The core function of the PK decision-making mechanism is as a collaborative scheduling core based on streaming and segmentation. When processing user instructions, this mechanism does not wait for either model's output to finish before making a decision and providing feedback; instead, it works collaboratively by having fast and slow models work alternately.

[0091] Specifically, to achieve a balance between efficiency and context: the fast model only acquires local context messages from the current and recent steps, quickly generating a response framework using preset transition phrases and template strategies; while the agentic slow model acquires the complete context including all historical messages, progressively executing deep planning and tool calls in the background. As the agentic slow model gradually acquires detailed data, the decision-making mechanism populates these data blocks into the response template generated by the fast model. Through this alternating fast and slow streaming output method, the system can effectively mask the various delays caused by background tool calls, network requests, and deep inference, significantly improving the user experience.

[0092] The following is a specific scenario to illustrate the working process of this streaming PK decision-making mechanism:

[0093] Scenario: Streaming speed coordination and latency masking. User initiates a command: "What's the weather like tomorrow?"

[0094] 1. Fast Model Instantaneous Response: After the instruction enters the processing unit, the fast model instantly generates a transitional response based on a very short context: "Okay, fine." This response is immediately output to the user, stabilizing the interaction expectation.

[0095] 2. Agentic slow model starts in the background: At the same time, the Agentic slow model begins to parse the intent and executes the Web Tool call to query the weather data of the target city.

[0096] 3. Fast model continues to mask delays: During the gap when the slow model is waiting for the tool to return results, the fast model continues to output templated phrases: "I am retrieving local weather information for you...".

[0097] 4. Slow model details: The background tool call is completed and returns specific data: "Cloudy, 15-22 degrees Celsius".

[0098] 5. Detail Filling and Final Output: The fast model receives the detailed data transmitted by the slow model, fills it into the preset response template, and completes the final output: "Tomorrow's weather will be cloudy, 15-22 degrees Celsius, please keep warm."

[0099] The following three consecutive scenarios will further illustrate the working process of this streaming collaboration mechanism under tasks of different complexities:

[0100] Scene 1: Simple Greeting

[0101] The user says to the robot, "Hello."

[0102] 1. The instruction enters the processing unit, and the fast model captures the general greeting based on its extremely short context window.

[0103] 2. Since this type of instruction does not involve detailed data or deep reasoning, the Agentic slow model does not need to execute tool calls after background evaluation.

[0104] 3. The fast model directly outputs a templated response: "Hello, it's a pleasure to help you. How can I assist you?"

[0105] In such simple scenarios, the system exhibits an instantaneous response with extremely low latency, and the fast-slow coordination mechanism naturally transitions into the dominant output of the fast model.

[0106] Scenario 2: Complex queries that require calling tools

[0107] The user then said, "Could you check the flights from Beijing to Shanghai for me tomorrow?"

[0108] 1. Fast Model Instant Transition: The fast model immediately outputs a transitional template: "Okay, checking flights from Beijing to Shanghai for you tomorrow, please wait...", to cover up the subsequent network request delay.

[0109] 2. Slow Model Background Call: Meanwhile, the Agentic slow model parses the instructions in the background, and after clarifying the intent, calls an external tool named flight_search_tool.

[0110] 3. Detailed segmentation and filling: After the Agentic slow model obtains the specific flight data (such as CA1837, price, etc.) returned by the tool, it transmits this detailed data in segments.

[0111] 4. Fast Model Fusion Output: The fast model receives detailed data from the slow model, seamlessly populates it into the response template, and continuously outputs: "Found it, recommends airline CA1837, departing at 8 am, economy class price 850 yuan. Do you need more detailed information?"

[0112] Scenario 3: Open-ended problems requiring lengthy contextual reasoning

[0113] After hearing the flight information, the user continued, "Which flight do you think offers the best value for money?"

[0114] 1. Fast Model Instantaneous Response and Short Context: The fast model generates a transition template based on a short context (containing only the current question and a list of immediately adjacent flights): "This is a good question, I'm comparing the cost-effectiveness of these flights for you..."

[0115] 2. Deep Reasoning and Long Context in the Slow Model: The Agentic slow model can acquire all long context records, including historical preferences. It not only analyzes the currently acquired flight prices and times, but also combines the implicit needs in the long context to perform deep reasoning and scoring.

[0116] 3. Alternating fast and slow output: The slow model transmits the core conclusions derived from reasoning in chunks, while the decision-making mechanism controls the fast model to complete the final wording assembly and output: "Considering both price and departure time, I believe that flight CA1837 of a certain airline has the highest cost performance."

[0117] Through the above progressive scenarios, this embodiment demonstrates how the streaming PK decision mechanism, as the core of intelligent scheduling, can precisely control the isolation of long and short contexts, and through the alternating output and detail filling of fast and slow models, perfectly mask the latency of background processing when faced with diverse user inputs, thereby optimizing the overall interactive performance and user experience.

[0118] Example 4

[0119] This embodiment aims to demonstrate in detail the advanced capabilities of the Feynman device orchestration module in handling situations involving multiple devices with physical preconditions and state conflicts, particularly how it resolves conflicts by autonomously inserting "state rollback tasks" to ensure the smooth execution of complex task flows. This embodiment will focus on the device bus, the Feynman device orchestration module, the memory module (especially the device state recording section), and the representation and management of the robot's own state.

[0120] Scenario: A user in a meeting room equipped with a projector and an electronic whiteboard gives the robot the instruction: "Play my presentation on the projector and then write down the key points on the whiteboard."

[0121] 1. Task Decomposition and Preliminary Planning: This instruction is processed by the Agentic slow model. The Agentic slow model decomposes it into two main device tasks:

[0122] Task A: Operate the projector to play the file and Task B: Write on the whiteboard.

[0123] At the same time, it will call the Feynman device orchestration module to perform preliminary sequence planning for the two tasks, that is, to execute task A first and then task B.

[0124] 2. Execute task A and cause a state change:

[0125] The Agentic slow model begins executing task A, sending instructions to the visual language action model module via the device bus.

[0126] The robot first moves to the table, and after visual recognition, picks up the projector's remote control.

[0127] The moment the robot successfully picks up the remote control, an important state change event occurs: one of the robot's internal states—the "hand state"—changes from "idle" to "occupied by the remote control".

[0128] It should be noted that this status information is updated in real time and recorded in the device operation status record of the memory module, providing a key basis for subsequent intelligent decision-making.

[0129] Next, the robot uses its occupied hand to operate the remote control, turning on the projector and selecting to play the user-specified presentation. This completes the main part of Task A.

[0130] 3. Prepare to execute task B and detect a state conflict:

[0131] After completing task A, the Agentic slow model prepares to start task B – “writing on the whiteboard”.

[0132] Before issuing specific action instructions, it consults the Feynman device orchestration module again to confirm whether the conditions for executing task B are met.

[0133] A core capability of Feynman's device orchestration module is its ability to identify and manage the preconditions for task execution. For the task of "writing on a whiteboard," its model stores a necessary "precondition device state": the robot's "hand state" must be "idle" in order to pick up the whiteboard marker.

[0134] The Feynman device orchestration module then queries the memory module for the "current device status" of the robot's "hand" and reads that the status is "occupied by remote control".

[0135] At this point, a conflict occurs: the current device state (occupied) does not match the "preceding device state" (idle) required to perform subsequent tasks.

[0136] 4. Automatically insert rollback tasks in the current state:

[0137] Faced with such state conflicts, a system lacking advanced intelligence might crash or report an error. However, the Feynman device orchestration module of this application is configured to proactively resolve such problems.

[0138] After identifying a conflict, based on the principle of global minimum energy and cost prediction, it assesses that the subjective economic cost of forcibly reporting an error or terminating the task is too high. Therefore, it will not immediately attempt to execute task B, but will automatically and dynamically insert a new, necessary intermediate task (i.e., the calculated minimum cost remedy path) between task A and task B. The sole purpose of this task is to resolve the state conflict.

[0139] This inserted task can be called a "state rollback task," and in this scenario, it specifically manifests as: "Put the remote control back on the table." The goal of this task is to roll back the robot's "hand state" from "occupied" to "idle."

[0140] The Feynman device orchestration module places this newly generated "state rollback task" instruction on the device bus and sets it to the highest priority.

[0141] 5. Execute the state rollback task and meet the prerequisites:

[0142] The Agentic slow model and visual-language-action model modules execute this newly inserted task. The robot moves back to the table and securely places the remote control.

[0143] After the task is completed, the robot's "hand status" is updated to "idle" again and recorded in the memory module.

[0144] At this point, the Feynman device orchestration module checks the prerequisites for task B again and finds that the "current device state" (idle) matches the "predecessor device state" (idle), thus eliminating the conflict.

[0145] 6. Continue with the original task:

[0146] Once the prerequisites are met, the Feynman device orchestration module sends a "can continue" signal to the Agentic slow model.

[0147] Accordingly, the Agentic slow model begins to execute Task B normally. The robot walks to the whiteboard, picks up the whiteboard marker (since its hand is now free), and begins writing down the key points based on the presentation.

[0148] This embodiment profoundly demonstrates the deep understanding of the complexity of the physical world by the architecture proposed in this application. It does not merely execute a series of isolated actions, but rather models and reasons about tasks, device states, and physical constraints as a whole. Through the Feynman device orchestration module's ability to identify preconditions, detect state conflicts, and autonomously insert state rollback tasks, the robot exhibits intelligent behavior similar to humans—"preparing" before action and "cleaning up" after action—enabling it to autonomously and smoothly complete complex, interconnected tasks in the real world.

[0149] The preferred embodiments of the present invention have been described in detail above and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A control architecture for a humanoid robot's dynamic intelligent agent, characterized in that, include: The processing unit is configured to execute the agent loop; The message bus is used to receive general interactive commands; The device bus is used to transmit commands related to device control. The processing unit is configured to process device control-related instructions on the device bus with a higher priority than processing general interactive instructions on the message bus. Furthermore, the processing unit includes: Fast model; Agentic slow model; and A decision-making mechanism is configured to coordinate the fast model and the Agentic slow model based on a streaming block processing mechanism. The fast model acquires local context information to generate a response template and transitional response, while the Agentic slow model acquires all historical context information to execute tool calls and obtain detailed data. The decision-making mechanism controls the fast model and the Agentic slow model to output alternately, progressively filling the response template with the detailed data in blocks to mask task processing delays.

2. The humanoid robot dynamic intelligent agent control architecture according to claim 1, characterized in that, The Agentic slow model is further configured to: parse the device control-related instructions into at least two device tasks, and invoke a dynamic orchestration model to dynamically sort or rearrange the execution order of the at least two device tasks.

3. The humanoid robot dynamic intelligent agent control architecture according to claim 1, characterized in that, The Agentic slow model is also configured to invoke a toolset that includes a Visual Language Action (VLA) model service, and the Agentic slow model performs physical actions by sending instructions to the device bus to invoke the VLA model service.

4. The humanoid robot dynamic intelligent agent control architecture according to claim 2, characterized in that, It also includes a multidimensional memory system for supporting the Agentic slow model in executing the device task and / or resuming the device task after it is interrupted.

5. The humanoid robot dynamic intelligent agent control architecture according to claim 4, characterized in that, The multidimensional memory system includes at least one of the following: a device operating status record for task recovery, a user session history record for personalized interaction, or a visual memory system for long-term environmental perception.

6. The humanoid robot dynamic intelligent agent control architecture according to claim 2, characterized in that, The dynamic orchestration model is configured to trigger a rearrangement of the device task execution order in response to detected environmental changes during task execution.

7. A humanoid robot dynamic intelligent agent control architecture according to claim 2 or 6, characterized in that, The dynamic orchestration model is configured to: when planning a task sequence, identify the preceding device states required to execute subsequent tasks, determine the current device state, and when the current device state conflicts with the preceding device state, automatically insert a state rollback task to make the device state satisfy the preceding device state.

8. The humanoid robot dynamic intelligent agent control architecture according to claim 3, characterized in that, The decision-making mechanism is configured to, upon determining that the fast-slow collaborative processing mode has been entered, schedule the fast model to output a response template based on the local context, and schedule the Agentic slow model to execute tool calls based on the global historical context, in order to obtain and stream-fill the detailed data.

9. A method for controlling a humanoid robot's dynamic intelligent agent, characterized in that, Includes the following steps: It receives general interactive commands through a message bus and transmits device control-related commands through a device bus. Device control-related instructions on the device bus are processed with a higher priority than general interaction instructions on the message bus. After determining that the fast and slow collaborative processing mode has been entered, the streaming block collaborative processing logic is activated. The fast model generates and outputs a transitional voice response template based on the local context; The Agentic slow model performs task processing in the background and streams the generated detailed data into the voice response template to generate the final response and mask the processing delay.

10. A method for controlling a humanoid robot's dynamic intelligent agent according to claim 9, characterized in that, The method further includes: During the execution of the first task processed by the Agentic slow model, a second instruction is received as an interrupt instruction. The second instruction is distributed to the fast model for processing and a response is output. After the fast model completes its response, context information representing the state of the first task before it was interrupted is read from a memory system. The context information includes the task execution steps and related parameters. Based on the context information, the Agentic slow model continues to execute the first task.