Robot body-equipped brain system based on cooperation of large and small models

By using a robot embodied brain system that combines large and small models, the problem of insufficient autonomy and intelligence in traditional robots has been solved. This system enables precise environmental perception, complex task planning, and multimodal human-machine interaction, thereby enhancing the robot's autonomous and intelligent capabilities.

CN121018546APending Publication Date: 2025-11-28杭州智元研究院有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511215119.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Traditional robots have low levels of autonomy and intelligence, lack the ability to accurately integrate and perceive complex environments and the ability to make autonomous decisions and plans for complex tasks, and the practical application of large model technology is constrained by factors such as data, computing power, and interpretability.

Method used

The robot adopts a robot embodied brain system based on the collaboration of large and small models. Through the collaboration of large and small models, it realizes environmental perception, decision-making and planning, control execution and human-computer interaction. It integrates heterogeneous data from multiple sensors, accurately perceives physical environment information, and supports multiple interaction methods to improve the level of autonomous intelligence.

Benefits of technology

It significantly improves the robot's autonomous intelligence level, enabling it to achieve optimal perception results and efficient task planning in complex environments, and providing a smooth and comfortable human-computer interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121018546A_ABST
    Figure CN121018546A_ABST
Patent Text Reader

Abstract

The invention discloses a robot body brain system based on cooperation of large and small models. The robot body brain system has the functions of complex task planning / re-planning, action control execution, visual perception reasoning and the like. The system comprises four functional modules of environment perception, decision planning, control execution and man-machine interaction, and can be realized based on cooperation of large and small models: different sensors are called as required according to specific tasks, and the sensors return original data containing external environment information to the system for subsequent processing; an advanced task instruction issued by the user is received, and a task execution result is fed back to the user in a text or voice mode; an upper-layer control instruction conforming to the specification is sent to the robot body or the robot load, and meanwhile, a task execution result returned by the robot body or the robot load can be received.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of embodied intelligence and robotics, particularly a robot embodied brain system based on the collaboration of large and small models. Background Technology

[0002] With the rapid development of the global software and hardware industries, key technologies for robots, such as environmental perception, motion control, energy and power, and structural design, have gradually matured. Currently, robots can assist or replace humans in completing heavy and repetitive tasks and are widely used in various scenarios. However, towards the ultimate goal of general artificial intelligence, traditional robots still suffer from problems such as low levels of autonomy and intelligence, a lack of accurate fusion perception capabilities in complex environments, and an inability to autonomously make decisions and plan for complex tasks.

[0003] Large models are machine learning models in the field of artificial intelligence with massive parameters and complex structures, possessing functional characteristics such as natural interactivity, learning and growth capabilities, and multi-tasking ability. Currently, significant progress has been made in core technologies such as large-scale cluster pre-training, efficient parameter fine-tuning, and cross-modal data alignment in the field of large models, and future development will trend towards increasing parameter scale, improved inference performance, and an increase in modality types. Compared with traditional artificial intelligence algorithms, large models have achieved disruptive breakthroughs in knowledge reserves, interactivity, and task versatility, and have been widely applied in civilian fields such as medicine, law, and finance. However, current large model-related technologies and products are mostly limited to the software algorithm level, and the practical application of large model technology is still constrained by factors such as data, computing power, and interpretability, requiring further research and application verification.

[0004] Therefore, the concept of "embodied intelligence," which advocates a complementary combination of robots and large-scale models, has rapidly emerged. Embodied intelligence refers to an intelligent system that interacts with its environment through physical entities such as robots, enabling it to perceive the environment, recognize information, make autonomous decisions, and take actions, and to achieve intelligent growth and adaptive behavior based on experiential feedback. Embodied intelligent robots mainly consist of three parts: the brain, the cerebellum, and the body. While the cerebellum and body technologies are relatively mature, the embodied intelligent brain technology is far from mature. Currently, embodied intelligence has become a focus of academic and industrial circles both domestically and internationally, and is an effective path to achieving general artificial intelligence. However, embodied brain technology is in its early stages of rapid development, and many challenges remain in terms of algorithms, data, and hardware / software. Summary of the Invention

[0005] The purpose of this invention is to address the problems existing in the prior art by providing a robot embodied brain system based on the collaboration of large and small models.

[0006] The technical solution to achieve the purpose of this invention is: a robot embodied brain system based on the collaboration of large and small models. The system is based on the collaboration of large and small models to realize: different sensors are called as needed according to specific tasks, and the sensors return raw data containing external environmental information to the system for subsequent processing.

[0007] Furthermore, based on the collaboration of large and small models, the system can also receive advanced task instructions issued by the user and provide feedback on the task execution results to the user in the form of text or voice.

[0008] Furthermore, based on the collaboration of large and small models, the system can also send compliant upper-level control commands to the robot body or robot payload, while simultaneously receiving the returned task execution results.

[0009] Furthermore, the system includes four major functional modules: an environment perception module, a decision planning module, a control execution module, and a human-computer interaction module, all of which are implemented through the collaboration of large and small models.

[0010] The environmental perception module is used to fuse heterogeneous data from multiple sensors and accurately perceive physical environment information based on a method of large and small model collaboration.

[0011] The decision planning module completes task planning and reasoning analysis based on a large model, and generates step-by-step decision instructions to control actions.

[0012] The control execution module is used to receive decision instructions, send control instructions to the underlying controller, and control the robot to perform specific actions;

[0013] The human-computer interaction module is used to accurately understand human intentions, provide feedback on human commands, and supports multiple interaction methods;

[0014] All decision-making and summary analysis operations in each module are performed by the large model.

[0015] Furthermore, the environmental perception module, based on a method of large and small model collaboration, accurately perceives physical environment information, specifically including:

[0016] Large models are used to align heterogeneous data from different sensors in order to capture the correlations and complementarities between the data.

[0017] For perception tasks lacking data, a method of scheduling a large model with a dedicated small model is used to obtain better perception results.

[0018] Furthermore, the environmental perception module uses sensors mounted on the robot to perceive and model the surrounding environment, and fully understands the objective laws and inherent relationships of the surrounding environment based on the perception model. The environmental perception module constructs a spatial mapping with the actual environment, providing the robot with reliable environmental cognition and understanding capabilities. In the environmental perception process, different modal perception data are aligned to a unified feature space through group coding, realizing the cognitive association and fusion of different modal information. Furthermore, the model output is converted into the required modal format through multimodal decoding.

[0019] Furthermore, the decision planning module, after receiving environmental perception information and user instruction input, completes advanced task planning and reasoning analysis, and generates decision instructions for controlling the robot's actions; the large and small models of this decision planning module are implemented collaboratively, specifically as follows:

[0020] The decision planning model is assigned based on the user input instructions. For simple task instructions, a small decision planning model is used to directly obtain executable subtasks; for complex task instructions, a large decision planning model is used to plan and decompose the instructions to obtain executable subtasks.

[0021] Furthermore, the decision planning module has the ability to replan on the fly, and can adjust the decision planning in real time according to changes in the environment and task requirements.

[0022] Furthermore, the large and small models of the control execution module are implemented collaboratively, specifically as follows:

[0023] The large model sends high-level control commands to the low-level control algorithm of the low-level controller, i.e., the small model. The robot schedules the corresponding small model control algorithm to execute the task.

[0024] Furthermore, the large and small models of the human-computer interaction module are implemented collaboratively, specifically as follows:

[0025] Use either a large or small model to handle single-modal interactions;

[0026] The results of the single-modal interactions are then aggregated into a large model for fusion processing to obtain the final interaction results.

[0027] Compared with the prior art, the significant advantages of this invention are:

[0028] (1) Compared with traditional intelligent robot systems based solely on small or large models, this invention combines the advantages of both large and small models and designs a robot embodied brain system based on the concept of large and small model collaboration. The method of large and small model collaboration runs through the four major modules of the system: perception, decision-making, execution, and interaction, which can significantly improve the autonomous intelligence level of the robot compared with existing technologies.

[0029] (2) There are two mature methods for calling small models from large models: function calling and MCP. This invention realizes a more systematic and refined arrangement and collaboration between large and small models. Based on the existing methods for calling small models from large models, the large model also plays a central role in decision-making and summary analysis. Before the task is executed, it can determine whether to use the large model itself or call certain small models. After the task is executed, it can comprehensively analyze the multi-path results of each small model.

[0030] (3) Thanks to the innovative method of large and small model collaboration, compared with the existing technology, the present invention has improved in environmental perception, decision planning and human-computer interaction capabilities: In terms of environmental perception, the present invention can take the optimal path to obtain the optimal perception result for traditional and novel targets; In terms of decision planning, the present invention can adaptively and efficiently complete task planning for simple and complex tasks; In terms of human-computer interaction, the present invention can summarize and analyze multiple interaction results to obtain a smoother and more comfortable interaction experience.

[0031] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description

[0032] Figure 1 This is a diagram showing the relationship between the four functional modules in one embodiment: the environment perception module, the decision planning module, the control execution module, and the human-computer interaction module.

[0033] Figure 2 This is an example of an external data flow diagram of the system.

[0034] Figure 3 This is a data flow diagram of the system's internal processes in one embodiment.

[0035] Figure 4 This is a task flow diagram in one embodiment.

[0036] Figure 5 This is a system state transition diagram in one embodiment. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0038] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0039] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0040] In one embodiment, a robot embodied brain system based on large and small model collaboration is provided. The system is based on large and small model collaboration to call different sensors (cameras, lidar, etc.) as needed according to specific tasks. The sensors return raw data containing external environment information to the system for subsequent processing.

[0041] The system can also receive advanced task instructions from users via text, voice, gestures, etc., and provide feedback on the task execution results to users via text or voice.

[0042] The system can also send compliant upper-level control commands to the robot body or robot payload, and receive the task execution results returned by the robot.

[0043] Furthermore, in one embodiment, the system includes four functional modules: an environment perception module, a decision planning module, a control execution module, and a human-computer interaction module, all of which are implemented through the collaboration of large and small models.

[0044] The environmental perception module is used to fuse heterogeneous data from multiple sensors and accurately perceive physical environment information based on a method of large and small model collaboration.

[0045] The decision planning module completes task planning and reasoning analysis based on a large model, and generates step-by-step decision instructions to control actions.

[0046] The control execution module is used to receive decision instructions, send control instructions to the underlying controller, and control the robot to perform specific actions;

[0047] The human-computer interaction module is used to accurately understand human intentions, provide feedback on human commands, and supports multiple interaction methods such as text and voice.

[0048] All decision-making and summary analysis operations in each module are performed by the large model.

[0049] Preferably, in some embodiments, the environment perception module accurately perceives physical environment information based on a method of large and small model collaboration, specifically including:

[0050] Large models are used to align heterogeneous data from different sensors in order to capture the correlations and complementarities between the data.

[0051] For perception tasks lacking data, such as identifying novel targets, a method of scheduling a large model with a dedicated small model can yield better perception results.

[0052] Preferably, in some embodiments, the environment perception module uses sensors mounted on the robot to perceive and model the surrounding environment, and fully understands the objective laws and inherent relationships of the surrounding environment based on the perception model; the environment perception module constructs a spatial mapping with the actual environment through target detection and tracking, target recognition and other technologies, providing the robot with reliable environmental cognition and understanding capabilities; in the environment perception process, different modal perception data such as images and text are aligned to a unified feature space through group coding, realizing the cognitive association and fusion of different modal information, and further converting the model output into the required modal format through multimodal decoding.

[0053] Preferably, in some embodiments, the decision planning module, after receiving environmental perception information and user instruction input, completes advanced task planning and reasoning analysis, and generates decision instructions for controlling robot actions; the large and small models of this decision planning module are implemented collaboratively, specifically as follows:

[0054] The decision planning model is assigned based on the user input instructions. For simple task instructions, a small decision planning model is used to directly obtain executable subtasks; for complex task instructions, a large decision planning model is used to plan and decompose the instructions to obtain executable subtasks.

[0055] Here, the decision planning module receives environmental perception information and user command input, completes advanced task planning and reasoning analysis, and generates decision commands to control the robot's actions;

[0056] The decision-making and planning module can enhance the accuracy and controllability of robot actions, while providing necessary input information for the control and execution module;

[0057] The specific implementation of the decision planning module is machine intelligent decision-making with a large model as its core. It is responsible for receiving various information from the environmental perception module, and after understanding the task objectives, it formulates specific action strategies.

[0058] The decision planning module can adjust its decision planning in real time according to changes in the environment and task requirements. It can continuously learn and optimize decisions by acquiring perceived information and action experience. It can also effectively coordinate and control other modules to ensure decision efficiency.

[0059] Preferably, in some embodiments, the large and small models of the control execution module are implemented collaboratively, specifically as follows:

[0060] The large model sends high-level control commands to the low-level control algorithm of the low-level controller, i.e., the small model. The robot schedules the corresponding small model control algorithm to execute the task.

[0061] Here, the main tasks of the control execution module include navigation, object manipulation, and object interaction, and its main components include motion control and actuator driving.

[0062] Meanwhile, the control execution module integrates commonly used embodied control methods in the industry, including: embodied control execution based on large model-assisted reinforcement learning, embodied control execution based on visual language action large model, and embodied control execution based on imitation learning.

[0063] Preferably, in some embodiments, the large and small models of the human-computer interaction module are implemented collaboratively, specifically as follows:

[0064] Use either a large or small model to handle single-modal interactions;

[0065] The results of the single-modal interactions are then aggregated into a large model for fusion processing to obtain the final interaction results.

[0066] Here, the human-computer interaction module parses the user's intent information into an information format that can be directly processed by the large model, and feeds back the output information of the large model to the user through multiple channels, serving as a communication bridge between the user and the robot.

[0067] Here, the human-computer interaction module breaks through the limitations of single-modal human-computer interaction, realizing multi-modal human-computer interaction that integrates multiple interaction methods such as text, voice, and gestures, improving the naturalness and smoothness of human-computer interaction, and ensuring seamless communication and efficient collaboration between humans and machines.

[0068] Combination Figures 1 to 5 The following will provide a detailed explanation of the implementation of the large model central module, the environmental perception module, the decision planning module, the control execution module, and the human-computer interaction module.

[0069] (1) The large model is the core hub of this system. To use the environmental perception, decision planning, control execution, and human-computer interaction modules, the large model hub needs to be built and started first. The specific implementation is as follows:

[0070] Read the configuration items in the YAML configuration file, including the operating mode (OPERATING_MODE), agent parameters (AGENT_PARAMS), planner parameters (PLANNER_PARAMS), ROS version (ROS_TYPE), voice broadcast switch (ENABLE_SPEAKER), visual model parameters (VISUAL_PARAMS), image input format (IMAGE_FORMAT), and topic information (TOPIC).

[0071] Define a large model skill library, which contains various tools that the large model can call (such as GoSomewhereTool for stationary movement, InstructionControlTool for instruction control, TakePhotosTool for taking photos, DanceTool for dancing, etc.);

[0072] Define the control and execution agent, including model parameters (api_key, base_url, model, temperature), prompt words, and bound skill library, etc.

[0073] Define the system's flow state variables, which maintain key variables in the system's execution process. These variables will flow and be updated in different functional modules during system operation, including: user input, real-time planning results, past steps, and final response.

[0074] Define the decision planning / replanning agent, including model parameters (api_key, base_url, model, temperature) and prompt words;

[0075] Define and compile the system's langgraph, including the graph's state, nodes, and edges;

[0076] Initialize the ROS node of the system hub, define a ROS topic Subscriber to subscribe to user input information, and call the spin function to start the large model hub.

[0077] (2) The specific implementation of the environmental perception module is as follows:

[0078] The large model central control continuously monitors ROS topic messages called by the environment perception module, with message type std_msgs / String;

[0079] Once the corresponding ROS topic message is received, the environment awareness module extracts the prompt words and task ID from the message;

[0080] The environment awareness module calls the wait_for_message function to obtain the ROS topic message of the camera's real-time frame, and further calls the imgmsg_to_cv2 function to convert the ROS Image format message into an OpenCV format message;

[0081] The environment awareness module has an optional image display function, which can display the image of the current frame to the user using the imshow function;

[0082] The environment perception module calls the imencode function to encode the image into binary data in JPEG format; it then calls the b64encode function to convert the binary data into a Base64 string.

[0083] Use the OpenAI library's client.chat.completions.create function to send environment-aware messages to a large model;

[0084] After obtaining the text results of the environment perception, the environment perception module will feed back the results to the central node or user in the form of ROS topics.

[0085] (3) The specific implementation of the decision-making and planning module is as follows;

[0086] The central system of the large model continuously monitors ROS topic messages called by the decision planning module, with message type std_msgs / String.

[0087] Once the corresponding ROS topic message is received, the decision planning module extracts the prompt words and task ID from the message;

[0088] The decision-making and planning agent defined in the central model breaks down user instructions into sub-task units that can be executed by the robot, and sends the sub-task sequence to the control execution module.

[0089] Whenever the control execution module completes a subtask, the replanning agent will replan the task based on the real-time progress of the task and send the replanning result to the control execution module. This process is repeated until the task objective is achieved or the task fails.

[0090] After the task objectives are achieved, the decision planning module will feed back the results to the central node or user via ROS topics.

[0091] The decision planning module features text presentation and voice broadcast functions (optional), which can provide real-time feedback on the results of task planning / replanning to the user, so that the user can keep track of the task progress or intervene in the task process in a timely manner.

[0092] (3) The specific implementation method of the control execution module is as follows;

[0093] After receiving the single-step task instruction from the decision planning module, the control execution module relies on the understanding and reasoning ability and tool calling ability of the large model to call the most suitable control execution tool in the skill library to execute the corresponding task. The specific task execution process may involve the robot's underlying controller and underlying hardware, but these factors are unrelated to the decoupling of this system.

[0094] After a single-step task is completed, the control execution module feeds back the task execution result to the large model hub and the replanning agent so that subsequent actions can be completed.

[0095] (4) The specific implementation of the human-computer interaction module is as follows;

[0096] The human-computer interaction module supports multiple input methods. For text input, it can be processed directly by the large model. For voice input, a small speech-to-text (STT) model is first used to convert the voice information into text information before processing. For gesture input, a small gesture recognition model is first used to convert the gesture image into the corresponding text meaning information before processing.

[0097] The embodied brain system executes tasks based on input instructions, and the process and results of task execution can be fed back in multiple ways;

[0098] Text feedback: The task process and results are presented on the display in text form. Voice feedback: The task process and results are first converted into speech information using a text-to-speech (TTS) model, and then the information is broadcast through a speaker.

[0099] As a specific example, the invention is illustrated in one embodiment.

[0100] Take two typical tasks as examples: complex task planning and visual perception reasoning.

[0101] The user input for the complex task planning was: "Move forward 30 meters, then go to the playground to take a picture. Never mind, let's change the picture location to the east gate." The Embodied Brain system received and understood the user input through the human-computer interaction module, and then completed the task planning decomposition based on the decision-making and planning module. The result was:

[0102] {'plan':['Control the robot to move forward 30 meters','Control the robot to go to the East Gate','Take a picture of the East Gate']}

[0103] After the control execution module completes the first subtask (controlling the robot to move forward 30 meters), it feeds back the result of the subtask to the embodied brain system: "Task steps: Move forward 30 meters, completed, execution result: success@ Reached 30 meters ahead." The system stores the historical steps in the memory module.

[0104] {'past_steps':[('Control the robot to move forward 30 meters','The robot has successfully moved forward 30 meters. The next step is to continue the plan. Please tell me the next instructions.')]}

[0105] The embodied brain system's decision-making and planning module performs ad-hoc replanning based on the real-time progress of the task, resulting in an updated plan:

[0106] {'plan':['Control the robot to go to the East Gate','Take and save a photo of the East Gate']}

[0107] After the control execution module completes the second subtask (controlling the robot to go to the East Gate), it feeds back the execution result of the subtask to the embodied brain system: "Task steps: Go to the East Gate, completed, execution result: success@Arrived at the East Gate." The system stores the historical steps in the memory module.

[0108] {'past_steps':[('Control the robot to move forward 30 meters','The robot has successfully moved forward 30 meters. The next step is to continue the plan. Please tell me the next instruction.'),('Control the robot to go to the East Gate','The robot has successfully gone to the East Gate. The next step is to choose to have the robot take a picture or perform other operations. If you want to continue the original plan now, please tell me.')]}

[0109] The embodied brain system's decision-making and planning module performs ad-hoc replanning based on the real-time progress of the task, resulting in an updated plan:

[0110] {'plan':['Take and save photos of the East Gate']}

[0111] After the control execution module completes the last subtask (taking and saving a photo of the East Gate), it feeds back the result of the subtask to the embodied brain system: "Task steps: Take a photo, completed, execution result: success@Photo taken." The system stores the historical steps in the memory module.

[0112] {'past_steps':[('Control the robot to move forward 30 meters','The robot has successfully moved forward 30 meters. The next step is to continue the plan. Please tell me the next instruction.'),('Control the robot to go to the East Gate','The robot has successfully gone to the East Gate. The next step is to choose to have the robot take a picture or perform other operations. If you want to continue the original plan now, please tell me.'),('Take a picture of the East Gate and save it','Taking a picture of the East Gate is complete.')]}

[0113] The embodied brain system's decision-making and planning module determines that the overall task objective has been achieved and returns the final response to the user through the human-computer interaction module:

[0114] {'response':'The task has been completed, including going to the East Gate and taking photos of it.'}

[0115] The user input for visual perception reasoning is: "How many workers are there ahead, and describe their clothing characteristics?" The embodied brain system receives and understands the user input through the human-computer interaction module, extracts the image data transmitted back in real time by the robot's camera, and forms a text-image pair input of "question input + image input", which is then handed over to the environmental perception module for visual perception reasoning.

[0116] After completing the reasoning, the environmental perception module returns the results to the embodied brain system and the user:

[0117] {'visual_result:'The image shows four workers. They are wearing work clothes, protective helmets, carrying backpacks, and holding tools.'}

[0118] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention without departing from its spirit and scope should be included within the protection scope of the present invention.

Claims

1. A robot embodied brain system based on the collaboration of large and small models, characterized in that, The system is based on the collaborative implementation of large and small models: different sensors are called as needed according to specific tasks, and the sensors return raw data containing external environmental information to the system for subsequent processing.

2. The robot embodied brain system based on large and small model collaboration according to claim 1, characterized in that, Based on the collaboration of large and small models, the system can also receive advanced task instructions issued by users and provide feedback on the task execution results to users in the form of text or voice.

3. The robot embodied brain system based on large and small model collaboration according to claim 1, characterized in that, Based on the collaboration of large and small models, the system can also send compliant upper-level control commands to the robot body or robot payload, and receive the task execution results returned by the robot.

4. The robot embodied brain system based on large and small model collaboration according to any one of claims 1 to 3, characterized in that, The system comprises four major functional modules: an environment perception module, a decision planning module, a control execution module, and a human-computer interaction module, all implemented through the collaboration of large and small models. The environmental perception module is used to fuse heterogeneous data from multiple sensors and accurately perceive physical environment information based on a method of large and small model collaboration. The decision planning module completes task planning and reasoning analysis based on a large model, and generates step-by-step decision instructions to control actions. The control execution module is used to receive decision instructions, send control instructions to the underlying controller, and control the robot to perform specific actions; The human-computer interaction module is used to accurately understand human intentions, provide feedback on human commands, and supports multiple interaction methods; All decision-making and summary analysis operations in each module are performed by the large model.

5. The robot embodied brain system based on large and small model collaboration according to claim 4, characterized in that, The environmental perception module, based on a method of large and small model collaboration, accurately perceives physical environment information, specifically including: Large models are used to align heterogeneous data from different sensors in order to capture the correlations and complementarities between the data. For perception tasks lacking data, a method of scheduling a large model with a dedicated small model is used to obtain better perception results.

6. The robot embodied brain system based on large and small model collaboration according to claim 5, characterized in that, The environmental perception module uses sensors mounted on the robot to perceive and model the surrounding environment, and fully understands the objective laws and inherent relationships of the surrounding environment based on the perception model. The environmental perception module constructs a spatial mapping with the actual environment, providing the robot with reliable environmental cognition and understanding capabilities. In the environmental perception process, different modal perception data are aligned to a unified feature space through grouping and encoding, realizing the cognitive association and fusion of different modal information. Furthermore, the model output is converted into the required modal format through multimodal decoding.

7. The robot embodied brain system based on large and small model collaboration according to claim 4, characterized in that, The decision-making and planning module receives environmental perception information and user command input, completes advanced task planning and reasoning analysis, and generates decision commands to control the robot's actions. This module utilizes a large-scale and a small-scale model in a collaborative manner, specifically as follows: The decision planning model is assigned based on the user input instructions. For simple task instructions, a small decision planning model is used to directly obtain executable subtasks; for complex task instructions, a large decision planning model is used to plan and decompose the instructions to obtain executable subtasks.

8. The robot embodied brain system based on large and small model collaboration according to claim 7, characterized in that, The decision planning module has the ability to replan on the fly, and can adjust the decision planning in real time according to changes in the environment and task requirements.

9. The robot embodied brain system based on large and small model collaboration according to claim 4, characterized in that, The control execution module is implemented through the coordinated use of large and small models, specifically as follows: The large model sends high-level control commands to the low-level control algorithm of the low-level controller, i.e., the small model. The robot schedules the corresponding small model control algorithm to execute the task.

10. The robot embodied brain system based on large and small model collaboration according to claim 4, characterized in that, The large and small models of the human-computer interaction module are implemented collaboratively, specifically as follows: Use either a large or small model to handle single-modal interactions; The results of the single-modal interactions are then aggregated into a large model for fusion processing to obtain the final interaction results.

Citation Information

Cited By

  • Robot for detecting illumination light environment

    CN121589838A

  • Robot control method and system based on dual system, training method and system

    CN122401441A