A somatic intelligent robot control method and device, a terminal and a storage medium

By generating chain-like reasoning text using large-scale pre-trained language models and multimodal information, and combining it with a whole-body visual motion strategy learning algorithm, an atomic skill library is constructed. This solves the problem of collaborative planning between the robot's whole-body movements and local operations, enabling the robot to perform tasks efficiently and safely in complex environments.

CN120244990BActive Publication Date: 2026-03-31SHENZHEN YOUMI TECHNOLOGY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as inconsistent planning, unsmooth action switching, and insufficient information fusion in the collaborative planning of robot whole-body movements and local operations, which cannot meet the high requirements of task execution, especially in complex dynamic environments.

Method used

By employing a large-scale pre-trained language model and multimodal prompts, chain-like reasoning text is generated. An atomic skill library is constructed through a full-body visual motion strategy learning algorithm to achieve full-body coordinated movements of the robot. Combined with dynamic feedback and replanning mechanisms, this ensures the robot's efficient task execution in complex environments.

Benefits of technology

It improves the success rate and safety of robots in complex environments, achieves seamless integration of whole-body movement and local operation, and ensures the continuity and robustness of tasks in dynamic changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244990B_ABST
    Figure CN120244990B_ABST
Patent Text Reader

Abstract

The application provides a body intelligent robot control method and device, a terminal and a storage medium, and relates to the technical field of robot autonomous task planning and motion control. The method comprises the following steps: collecting multi-modal prompt information of a robot to be controlled; based on a large-scale pre-training language model, multi-modal prompt information and an atomic skill library suitable for long sequence whole body linkage motion of the robot, generating a corresponding chain reasoning text when the robot to be controlled executes a task description, and decomposing the task description into an atomic skill calling sequence according to the chain reasoning text, wherein the atomic skill calling sequence comprises a plurality of atomic skills in calling order; after each atomic skill is called and activated one by one in the calling order, the robot to be controlled is controlled to execute each atomic skill in the calling order. The application can realize seamless connection of whole body motion and local operation of the robot, so that the robot can efficiently complete a long sequence task in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot autonomous task planning and motion control technology, and in particular to a control method, device, terminal and storage medium for an embodied intelligent robot. Background Technology

[0002] In recent years, with the rapid development of deep learning technology and large-scale pre-trained models in the fields of vision, natural language processing, and robot control, robot task planning and motion control have gradually shifted from traditional rule-based systems to semantic understanding and data-driven approaches. Existing research utilizes visual language models to translate natural language instructions into specific operational steps. Through pre-trained target detection, grasping posture generation, and motion planning algorithms, robots have achieved automatic decomposition and execution of local tasks. Furthermore, the emergence of embodied intelligence provides an important technological and theoretical foundation for enhancing current robot cognitive abilities and moving towards general intelligence. Through interaction with the environment and humans, robots can obtain feedback from humans and the environment in real physical or virtual digital spaces, thereby better completing tasks. However, existing technologies mainly focus on the implementation of local actions, often planning only for robotic arms or local execution units, lacking an organic connection between the robot's whole-body movements and detailed local operations.

[0003] Previous research proposed a method combining multimodal information with pre-trained models for robot task planning and motion control through a pre-defined atomic skill library. While this method has achieved some success in target detection, grasping, and placement, it still suffers from problems such as inconsistent planning, unsmooth motion transitions, and insufficient information fusion when facing real-world scenarios involving full-body movements (e.g., global navigation, center of gravity adjustment, cross-platform collaboration). Furthermore, the chain-like reasoning and failure detection and replanning mechanisms in existing technologies also have room for improvement, failing to fully meet the high demands of task execution in complex dynamic environments. Summary of the Invention

[0004] This application provides a control method, device, terminal, and storage medium for an embodied intelligent robot to solve the problem of the lack of coordinated planning of robot whole-body movements and local operations in the prior art.

[0005] In a first aspect, this application provides a method for controlling an embodied intelligent robot, comprising:

[0006] Collect multimodal prompting information of the robot to be controlled, including the current full-body state of the robot to be controlled, environmental visual information, and task description input by the user;

[0007] Based on a large-scale pre-trained language model, the multimodal cue information, and an atomic skill library suitable for long-sequence full-body coordinated movements of the robot, a chain-like inference text corresponding to the task description is generated when the robot to be controlled executes the task description. The task description is then decomposed into an atomic skill invocation sequence based on the chain-like inference text. The atomic skill invocation sequence includes multiple atomic skills in the order of invocation. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm and includes target detection skills, navigation skills, object grasping and placement skills, remote object grasping and placement skills, and target following skills.

[0008] After calling and activating each atomic skill in the order of invocation, the robot to be controlled is controlled to execute each atomic skill in the order of invocation.

[0009] Secondly, this application provides a control device for an embodied intelligent robot, comprising:

[0010] The information acquisition module is used to acquire multimodal prompting information of the robot to be controlled. The multimodal prompting information includes the current full-body state of the robot to be controlled, environmental visual information, and task description input by the user.

[0011] The generation module is used to generate chained reasoning text corresponding to the task description when the robot to be controlled executes the task description, based on a large-scale pre-trained language model, the multimodal prompting information, and an atomic skill library suitable for long-sequence full-body linkage movements of the robot. The module also decomposes the task description into an atomic skill call sequence based on the chained reasoning text. The atomic skill call sequence includes multiple atomic skills in the order of call. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm and includes target detection skills, navigation skills, object grasping and placement skills, remote object grasping and placement skills, and target following skills.

[0012] The control module is used to control the robot to be controlled to execute each atomic skill in the order of invocation after calling and activating each atomic skill in the order of invocation.

[0013] Thirdly, this application provides a terminal including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method as described in the first aspect or any possible implementation of the first aspect above.

[0014] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method as described in the first aspect or any possible implementation of the first aspect.

[0015] This application provides a method, device, terminal, and storage medium for controlling an embodied intelligent robot. The method involves collecting multimodal cue information from the robot to be controlled, including the robot's current full-body state, environmental visual information, and a task description input by the user. Based on a large-scale pre-trained language model, the multimodal cue information, and an atomic skill library suitable for long-sequence full-body coordinated movements of the robot, a chain-like inference text corresponding to the robot's execution of the task description is generated. The task description is then decomposed into an atomic skill invocation sequence based on this inference text. This sequence includes multiple atomic skills in the order of invocation. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm and includes target detection skills, navigation skills, object grasping and placement skills, remote object grasping and placement skills, and target following skills. After invoking and activating each atomic skill in the order of invocation, the robot is controlled to execute each atomic skill in that order. This application deeply integrates whole-body status with environmental visual information, enabling real-time adjustment of action strategies in the face of dynamic changes and uncertainties, significantly improving task success rate and safety. Furthermore, by combining large-scale pre-trained language models and multimodal prompts, this application generates chained reasoning text and decomposes tasks into atomic skill invocation sequences. This breaks down complex tasks into a series of executable atomic skills, allowing the robot to systematically plan task execution steps, improving task processing efficiency and accuracy. Simultaneously, the atomic skill library is specifically designed for long-sequence whole-body coordinated movements of the robot, achieving seamless integration of whole-body motion and local operations. This allows the robot to efficiently complete long-sequence tasks in complex environments, with global navigation skills ensuring the robot quickly reaches the target area from any initial position, while local skills guarantee precise manipulation of the target object. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the structure of an embodied intelligent robot platform with full-body movement capabilities provided in an embodiment of this application;

[0018] Figure 2 This is a flowchart illustrating the implementation of the embodied intelligent robot control method provided in the embodiments of this application;

[0019] Figure 3 This is a schematic diagram of the task execution flow provided in the embodiments of this application;

[0020] Figure 4 This is a schematic diagram of the structure of the embodied intelligent robot control device provided in the embodiments of this application;

[0021] Figure 5 This is a schematic diagram of the terminal provided in the embodiments of this application. Detailed Implementation

[0022] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0024] To address the limitations of previous methods where atomic skills are confined to local actions, lacking comprehensive robot-wide coordinated actions, multimodal data perception supplementation, and insufficient detail in task planning / replanning, this application proposes an embodied intelligent robot control method based on a multimodal large model and a comprehensive coordinated atomic skill library. By integrating global and local information, employing a large-scale pre-trained language model for chain reasoning, and designing dynamic feedback and replanning mechanisms, a comprehensive coordinated embodied intelligent robot atomic skill library is constructed, thereby enhancing the robot's autonomous collaborative operation capability and robustness in complex environments.

[0025] The control method in this application is based on Figure 1 The illustrated embodied intelligent robot platform, capable of full-body movement, is developed based on an open robot operating system and integrates vision processing, path planning, and large-scale pre-trained model inference modules. All modules use a unified interface for data interaction. The platform mainly includes a multimodal input module, a task planning module (i.e., a large-scale pre-trained language model), an atomic skill library, a task execution and feedback closed-loop module, and a task memory storage module. The multimodal input module is connected to the task planning module, which is connected to both the task execution and feedback closed-loop module and the task memory storage module. The task execution and feedback closed-loop module and the task memory storage module are connected to the atomic skill library.

[0026] In addition, the main hardware of the platform includes:

[0027] (1) Multifunctional mobile chassis, supporting global navigation.

[0028] (2) The robotic arm and end effector are responsible for local operations.

[0029] (3) Camera, depth sensor, lidar and inertial measurement unit, responsible for collecting environmental and robot status data.

[0030] (4) Force and torque sensors provide motion feedback.

[0031] Figure 2 The implementation flowchart of the embodied intelligent robot control method based on an embodied intelligent robot platform with full-body motion capabilities, provided in the embodiments of this application, is described in detail below:

[0032] In step 101, multimodal prompting information of the robot to be controlled is collected. The multimodal prompting information includes the current full-body state of the robot to be controlled, environmental visual information, and task description input by the user.

[0033] In this embodiment, a multimodal input module is used to collect the current full-body state of the robot under control, environmental visual information, and a task description in natural language input by the user. The collected information is then integrated to obtain multimodal prompting information for the robot. Specifically, multiple sensors, such as cameras, depth sensors, LiDAR, and inertial measurement units, are used to collect the current full-body state and environmental visual information of the robot under control. The environmental visual information is represented by images and depth maps, and the full-body state includes pose information. The collected images, depth maps, and the pose information of the robot under control are organically integrated with the task description input by the user to form a unified form of multimodal prompting information.

[0034] In addition, before the task begins, this embodiment first requires the use of various sensors in the multimodal input module to collect panoramic images, depth maps, and full-body state data of the robot to be controlled. Environmental mapping is performed using LiDAR and camera data, and the sensor data is time-synchronized and calibrated to form a unified data stream. The full-body state data includes pose, center of gravity distribution, motion parameters, etc., and is transmitted as a memory message to the task planning module.

[0035] This application embodiment deeply integrates vision, language, whole-body status and environmental map information, enabling the robot to adjust its action strategy in real time when facing dynamic changes and uncertainties, significantly improving the success rate and safety of tasks.

[0036] In step 102, based on a large-scale pre-trained language model, multimodal cue information, and an atomic skill library suitable for long-sequence full-body linkage movements of the robot, a chain-like inference text corresponding to the task description executed by the robot to be controlled is generated. The task description is then decomposed into an atomic skill call sequence based on the chain-like inference text. The atomic skill call sequence includes multiple atomic skills in the order of call. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm and includes target detection skills, navigation skills, object grasping and placement skills, remote object grasping and placement skills, and target following skills.

[0037] The atomic skill library in this application extends existing research. Besides traditional local operations such as target detection, grasping posture generation, arm motion planning, and object placement and release, it implements long-sequence atomic skills based on a whole-body visual motion strategy learning algorithm. These include long-sequence whole-body coordinated actions such as object grasping and placement, and remote object grasping and placement. The whole-body visual motion strategy learning algorithm (WB-VIMA) is an innovative algorithm specifically designed for learning the whole-body visual motion strategy of robots, aiming to improve the robot's whole-body manipulation capabilities in the home environment, especially performing exceptionally well in complex daily household tasks.

[0038] The specific generation process of the atomic skill library is as follows: The WB-VIMA algorithm first predicts the future trajectory of the mobile chassis through its autoregressive denoising process. This prediction serves as a condition for subsequent torso and arm movement predictions. The WB-VIMA algorithm primarily relies on the robot's own multimodal sensor data, combined with the robot's body state (such as joint positions and chassis speed). This information is fused in real time to generate coordinated navigation and full-body movement commands. Specifically, the atomic skill library includes, but is not limited to, the following:

[0039] Object detection skills: Extract target regions and generate target segmentation masks based on task descriptions and image information using a large-scale pre-trained language model.

[0040] Navigation skills: Utilizing global maps and sensor data, global path planning algorithms enable global navigation from the current location to the target area, with dynamic adjustments made in conjunction with visual and inertial data.

[0041] Object grasping and placement skills: The robot achieves the ability to grasp and place objects through the coordinated movement of its entire body.

[0042] Remote object grasping and placement skills: Through the robot's full-body linkage and long-sequence navigation capabilities, the robot can smoothly complete the task of remotely placing objects after grasping them.

[0043] Target following skills: Through target tracking, full-body robot coordination, and long-term navigation capabilities, the robot can ensure long-term target tracking and provide better service.

[0044] In this embodiment, each atomic skill adopts a unified interface format, and the input parameters include task description, visual data, depth information, sensor feedback, and motion constraints. The modular design of the atomic skill library allows each skill to be executed independently or flexibly combined according to task requirements, providing a standardized calling interface for the task planning module.

[0045] In this embodiment, the task planning module includes both skill invocation for local operations and full-body motion planning such as global navigation and center of gravity adjustment. The overall planning process is as follows:

[0046] A unified prompt is constructed based on the multimodal input, and the multimodal prompt information is concatenated into a standardized input.

[0047] Detailed chained reasoning text is generated by large-scale pre-trained language models, decomposing task descriptions into atomic skill call sequences.

[0048] The output is structured, including the skill call order and the specific parameters of each skill.

[0049] Specifically, the reasoning method based on the large-scale pre-trained language model in the task planning module utilizes a pre-defined chain-thinking template and memory context information to input multimodal prompts and atomic skills suitable for long-sequence full-body coordinated movements of the robot into the large-scale pre-trained language model. The large-scale pre-trained language model performs chain-thinking according to the preset template, generating a structured planning output, which is the chain-thinking text corresponding to the robot's execution of the current task description. This chain-thinking text details the chain-thinking process of task decomposition. Based on this chain-thinking text, the current task description is decomposed into a sequence of atomic skill calls. This sequence includes multiple atomic skills arranged in the order of call, listing a continuous sequence of operations from global navigation, target detection, grasping posture generation, local motion planning, full-body center of gravity adjustment to specific grasping and placement. Furthermore, the output of the task planning module in this embodiment is organized through a preset unified interface format, facilitating the invocation of the task execution and feedback closed-loop module.

[0050] The embodiments of this application adopt an atomic skill library design with a unified interface, which enables the robot platform to flexibly integrate new skills and support the changing needs of different task scenarios.

[0051] In one possible implementation, based on a large-scale pre-trained language model, multimodal cue information, and an atomic skill library suitable for long-sequence full-body coordinated movements of the robot, chain-like inference text corresponding to the task description performed by the robot to be controlled is generated, which may include:

[0052] Multimodal prompts and preset chained thinking templates are used as constant information, and the robot's historical full-body state, historical task records, and execution feedback are used as memory information.

[0053] Constant information, memory information, task description, and an atomic skill library suitable for long-sequence full-body coordinated movements of the robot are input into a large-scale pre-trained language model, which outputs the chain-like reasoning text corresponding to the task description when the robot to be controlled performs it.

[0054] In this embodiment, based on the acquired images, depth maps, and overall body status, the task description, sensor data, and the contents of a pre-set atomic skill library are integrated into standardized multimodal prompt information. This multimodal prompt information can be divided into four parts:

[0055] (1) Constant information, embedded with detailed chain-thinking templates, guides large-scale pre-trained language models to logically decompose task descriptions. Specifically, it includes visual perception information such as RGB images, depth images and point clouds, and preset chain-thinking templates, requiring large-scale pre-trained language models to decompose task descriptions step by step, while clarifying the skills to be called and their parameter formats for each step.

[0056] (2) Memory information: Record the robot's recent state and historical task information. Specifically, this includes the robot's most recent overall state, historical task records, and execution feedback, which are used to maintain continuity in multi-stage tasks.

[0057] (3) Task description, that is, the specific task objectives described by the user in natural language.

[0058] (4) Atomic Skills Library: Lists all currently available atomic skills and their function descriptions.

[0059] Optionally, constant information, memory information, task description, and atomic skill library are input into a large-scale pre-trained language model to generate a structured output, namely the chained reasoning text corresponding to the task description performed by the robot to be controlled.

[0060] This application embodiment relies on the reasoning ability of a large-scale pre-trained language model and a preset thinking template to clearly explain the task decomposition logic and replan based on real-time feedback during task execution, ensuring the continuity and robustness of the entire task process.

[0061] In one possible implementation, after decomposing the task description into a sequence of atomic skill calls based on the chained reasoning text, the method may further include:

[0062] Based on the task description, assign corresponding skill parameters to each atomic skill in the atomic skill invocation sequence.

[0063] Optionally, after obtaining the atomic skill invocation sequence, specific input parameters are provided for each atomic skill in the atomic skill invocation sequence, such as target position, visual data, motion range, image data, depth data, global map information, center of gravity state, and motion control parameters.

[0064] In the process of generating chained reasoning text and atomic skill call sequences, the embodiments of this application ensure that the planned output has both a transparent reasoning process and meets the interface requirements for subsequent specific action execution.

[0065] In step 103, after calling and activating each atomic skill in the order of invocation, the robot to be controlled executes each atomic skill in the order of invocation.

[0066] In this embodiment, the skill invocation sequence output by the task planning module is passed to the task execution and feedback closed-loop module. The task execution and feedback closed-loop module invokes and activates each atomic skill one by one according to the invocation order, so as to control the robot to be controlled to execute according to each atomic skill.

[0067] The embodiments of this application enable seamless integration of the robot's whole-body motion and local operations, allowing the robot to efficiently complete long sequences of tasks in complex environments. Global navigation skills ensure that the robot can quickly reach the target area from any initial position, while local skills guarantee precise manipulation of the target object.

[0068] In one possible implementation, after calling and activating each atomic skill in the order of invocation, and after controlling the robot to be controlled to execute each atomic skill in the order of invocation, the method may further include:

[0069] The robot's motion status and environmental feedback data are collected in real time. The robot's motion status includes the robot's overall state information after each atomic skill is executed, and the environmental feedback data includes the environmental visual information after each atomic skill is executed.

[0070] Optionally, the task is executed according to the atomic skill invocation sequence planned by the large-scale pre-trained language model, and the robot's action status and environmental feedback data are collected in real time during the execution process.

[0071] Throughout the entire execution process, the closed-loop feedback mechanism in this application embodiment ensures that the robot can continuously adjust its motion planning according to real-time changes in complex and dynamic environments, achieving seamless integration and robust control of global navigation and local operations, and improving the robustness and stability of task execution.

[0072] After real-time acquisition of robot motion status and environmental feedback data, a feedback adjustment mechanism is implemented. The feedback mechanism in this embodiment includes two methods: sensor feedback detection and state monitoring and fault replanning.

[0073] Sensor feedback detection determines whether the target has undergone the expected change by comparing images and depth data before and after each atomic skill is executed.

[0074] Condition monitoring and fault replanning utilize force and pose sensors to detect whether the robot's end effector or whole-body motion has reached a predetermined state.

[0075] Then, when an action fails or does not meet expectations, the current status, the cause of the failure, and historical operation records are used to form a reflection prompt, which is fed back to the task planning module. The chain reasoning is then re-performed to generate a new skill call sequence and re-plan subsequent actions.

[0076] The embodiments of this application utilize the aforementioned closed-loop feedback mechanism, which has dynamic adjustment and adaptive recovery capabilities, significantly improving the overall task success rate.

[0077] In one possible implementation, after real-time acquisition of robot motion state and environmental feedback data, the method may further include:

[0078] For each atomic skill, perform the following steps:

[0079] The difference between the robot's action state before and after the execution of the atomic skill is taken as the first difference;

[0080] Determine whether the first difference is not greater than the first threshold;

[0081] If the first difference is greater than the first threshold, it is determined that the atomic skill has not achieved the expected effect. The robot's action state after the execution of the atomic skill is combined with the corresponding fault information to form a reflection prompt, and the process returns to the step of collecting multimodal prompt information of the robot to be controlled to continue execution.

[0082] If the first difference is not greater than the first threshold, it is determined that the atomic skill has achieved the expected effect, and the robot's action state after the execution of the atomic skill is stored.

[0083] Optionally, this process is a state monitoring and fault planning feedback mechanism, namely: for each atomic skill, the difference between the robot's action state before and after the execution of the atomic skill is taken as the first difference, and it is determined whether the first difference is not greater than a first threshold. If it is greater, it is determined that the robot under control has failed to achieve the expected effect under the execution of the atomic skill, such as the target not reaching the predetermined position, insufficient grasping force, or the robot losing balance. Then, the fault information and the current state data are used to construct a reflection prompt, and after the reflection prompt is passed to the task planning module, the process returns to the multimodal prompt information collection step to continue execution. A new skill call sequence is generated by a large-scale pre-trained language model to remedy the failure. If it is not greater, it is determined that the robot under control has achieved the expected effect under the execution of the atomic skill, and the robot's action state after the execution of the atomic skill that achieved the expected effect is stored in the task memory storage module.

[0084] In one possible implementation, after real-time acquisition of robot motion state and environmental feedback data, the method may further include:

[0085] For each atomic skill, perform the following steps:

[0086] The difference between the environmental visual information before and after the execution of the atomic skill is used as the second difference.

[0087] Determine whether the second difference is not greater than the second threshold;

[0088] If the second difference is greater than the second threshold, it is determined that the action of the atomic skill has not achieved the expected effect, and the process returns to the step of collecting multimodal prompt information of the robot to be controlled to continue execution;

[0089] If the second difference is not greater than the second threshold, it is determined that the action of the atomic skill has achieved the expected effect, and the environmental visual information after the execution of the atomic skill is stored.

[0090] Optionally, this process employs a sensor feedback detection mechanism: for each atomic skill, the difference in environmental visual information before and after the execution of that atomic skill is used as a second difference. This includes comparing changes in the depth map, force sensor feedback, and positional deviation before and after execution, and determining whether the second difference is not greater than a second threshold. If it is greater, it is determined that the robot under control has failed to achieve the expected effect under the execution of that atomic skill, and the process directly returns to the multimodal prompt information acquisition step for re-execution. If it is not greater, it is determined that the robot under control has achieved the expected effect under the execution of that atomic skill, and the robot's action state after the execution of that atomic skill, indicating that the expected effect has been achieved, is stored in the task memory storage module.

[0091] The introduction of the task memory storage module in this embodiment helps to continuously optimize the planning strategy during long-term operation, enabling self-learning and evolution.

[0092] For example, refer to Figure 3 The user inputs a task description of "get me a bottle of water." The multimodal input module then inputs the task description, the robot's current overall state, and environmental visual information into the task planning module. This task description is broken down into three actions: "get water," "navigate to the user," and "hand over the water." These three actions are then passed to the task execution and feedback closed-loop module, where the corresponding atomic skills are executed in the order they are called. The module determines whether each atomic skill was successfully completed. If successful, the result is stored directly in the task memory module; if it fails, the fault information and current data form a reflection prompt, which is then sent back to the task planning module for replanning.

[0093] This application provides a method for controlling an embodied intelligent robot. The method involves collecting multimodal cue information from the robot to be controlled, including the robot's current full-body state, environmental visual information, and a task description input by the user. Based on a large-scale pre-trained language model, the multimodal cue information, and an atomic skill library suitable for long-sequence full-body coordinated movements of the robot, a chain-like inference text corresponding to the robot's execution of the task description is generated. The task description is then decomposed into an atomic skill invocation sequence based on this inference text. This sequence includes multiple atomic skills in the order of invocation. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm and includes target detection skills, navigation skills, object grasping and placement skills, remote object grasping and placement skills, and target following skills. After invoking and activating each atomic skill in the order of invocation, the robot is controlled to execute each atomic skill in that order. This application deeply integrates whole-body status with environmental visual information, enabling real-time adjustment of action strategies in the face of dynamic changes and uncertainties, significantly improving task success rate and safety. Furthermore, by combining large-scale pre-trained language models and multimodal prompts, this application generates chained reasoning text and decomposes tasks into atomic skill invocation sequences. This breaks down complex tasks into a series of executable atomic skills, allowing the robot to systematically plan task execution steps, improving task processing efficiency and accuracy. Simultaneously, the atomic skill library is specifically designed for long-sequence whole-body coordinated movements of the robot, achieving seamless integration of whole-body motion and local operations. This allows the robot to efficiently complete long-sequence tasks in complex environments, with global navigation skills ensuring the robot quickly reaches the target area from any initial position, while local skills guarantee precise manipulation of the target object.

[0094] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0095] The following are device embodiments of this application. For details not described in detail, please refer to the corresponding method embodiments described above.

[0096] Figure 4 A schematic diagram of the structure of the embodied intelligent robot control device provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown, and are described in detail below:

[0097] like Figure 4 As shown, the embodied intelligent robot control device 4 includes:

[0098] The information acquisition module 41 is used to acquire multimodal prompting information of the robot to be controlled. The multimodal prompting information includes the current full-body status of the robot to be controlled, environmental visual information, and task description input by the user.

[0099] The generation module 42 is used to generate chained reasoning text corresponding to the task description performed by the robot to be controlled, based on a large-scale pre-trained language model, multimodal prompting information, and an atomic skill library suitable for long-sequence full-body linkage movements of the robot. The task description is decomposed into an atomic skill call sequence based on the chained reasoning text. The atomic skill call sequence includes multiple atomic skills in the order of call. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm and includes target detection skills, navigation skills, object grasping and placement skills, remote object grasping and placement skills, and target following skills.

[0100] The control module 43 is used to control the robot to be controlled to execute each atomic skill in the order of invocation after calling and activating each atomic skill one by one.

[0101] This application provides an embodied intelligent robot control device. It collects multimodal cue information from the robot to be controlled, including the robot's current full-body state, environmental visual information, and a task description input by the user. Based on a large-scale pre-trained language model, the multimodal cue information, and an atomic skill library suitable for long-sequence full-body coordinated movements of the robot, it generates chained reasoning text corresponding to the robot's execution of the task description. The task description is then decomposed into an atomic skill invocation sequence based on this chained reasoning text. This sequence includes multiple atomic skills in the order of invocation. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm and includes target detection skills, navigation skills, object grasping and placement skills, remote object grasping and placement skills, and target following skills. After invoking and activating each atomic skill in the order of invocation, the device controls the robot to execute each atomic skill in that order. This application deeply integrates whole-body status with environmental visual information, enabling real-time adjustment of action strategies in the face of dynamic changes and uncertainties, significantly improving task success rate and safety. Furthermore, by combining large-scale pre-trained language models and multimodal prompts, this application generates chained reasoning text and decomposes tasks into atomic skill invocation sequences. This breaks down complex tasks into a series of executable atomic skills, allowing the robot to systematically plan task execution steps, improving task processing efficiency and accuracy. Simultaneously, the atomic skill library is specifically designed for long-sequence whole-body coordinated movements of the robot, achieving seamless integration of whole-body motion and local operations. This allows the robot to efficiently complete long-sequence tasks in complex environments, with global navigation skills ensuring the robot quickly reaches the target area from any initial position, while local skills guarantee precise manipulation of the target object.

[0102] In one possible implementation, the generation module can be used for:

[0103] Multimodal prompts and preset chain thinking templates are used as constant information, and the robot's historical full-body state, historical task records, and execution feedback are used as memory information.

[0104] Constant information, memory information, task description, and atomic skill library are input into a large-scale pre-trained language model, which outputs the chained reasoning text corresponding to the task description when the robot to be controlled performs it.

[0105] In one possible implementation, the device may further include a parameter enabling module, which can be used to:

[0106] Based on the task description, assign corresponding skill parameters to each atomic skill in the atomic skill invocation sequence.

[0107] In one possible implementation, the device may further include a data acquisition module, which can be used for:

[0108] The robot's motion status and environmental feedback data are collected in real time. The robot's motion status includes the robot's overall state information after each atomic skill is executed, and the environmental feedback data includes the environmental visual information after each atomic skill is executed.

[0109] In one possible implementation, the device may further include a first determination module, which may be used to:

[0110] For each atomic skill, perform the following steps:

[0111] The difference between the robot's action state before and after the execution of the atomic skill is taken as the first difference;

[0112] Determine whether the first difference is not greater than the first threshold;

[0113] If the first difference is greater than the first threshold, it is determined that the atomic skill has not achieved the expected effect. The robot's action state after the execution of the atomic skill is combined with the corresponding fault information to form a reflection prompt, and the process returns to the step of collecting multimodal prompt information of the robot to be controlled to continue execution.

[0114] If the first difference is not greater than the first threshold, it is determined that the atomic skill has achieved the expected effect, and the robot's action state after the execution of the atomic skill is stored.

[0115] In one possible implementation, the device may further include a second determining module, which may be used to:

[0116] For each atomic skill, perform the following steps:

[0117] The difference between the environmental visual information before and after the execution of the atomic skill is used as the second difference.

[0118] Determine whether the second difference is not greater than the second threshold;

[0119] If the second difference is greater than the second threshold, it is determined that the action of the atomic skill has not achieved the expected effect, and the process returns to the step of collecting multimodal prompt information of the robot to be controlled to continue execution;

[0120] If the second difference is not greater than the second threshold, it is determined that the action of the atomic skill has achieved the expected effect, and the environmental visual information after the execution of the atomic skill is stored.

[0121] Figure 5 This is a schematic diagram of the terminal provided in an embodiment of this application. For example... Figure 5As shown, the terminal 5 in this embodiment includes: a processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50. When the processor 50 executes the computer program 52, it implements the steps described in the various embodiments of the embodied intelligent robot control method, for example... Figure 1 Steps 101 to 103 are shown. Alternatively, when the processor 50 executes the computer program 52, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 4 The functions of each module are shown.

[0122] For example, the computer program 52 can be divided into one or more modules / units, which are stored in the memory 51 and executed by the processor 50 to complete this application. The one or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program 52 in the terminal 5. For example, the computer program 52 can be divided into... Figure 4 The modules shown.

[0123] The terminal 5 can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The terminal 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will understand that... Figure 5 This is merely an example of terminal 5 and does not constitute a limitation on terminal 5. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal may also include input / output devices, network access devices, buses, etc.

[0124] The processor 50 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0125] The memory 51 can be an internal storage unit of the terminal 5, such as a hard disk or memory of the terminal 5. The memory 51 can also be an external storage device of the terminal 5, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal 5. Furthermore, the memory 51 can include both internal storage units and external storage devices of the terminal 5. The memory 51 is used to store the computer program and other programs and data required by the terminal. The memory 51 can also be used to temporarily store data that has been output or will be output.

[0126] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0127] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0128] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0129] In the embodiments provided in this application, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0130] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0131] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0132] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various embodiments of the intelligent robot control method described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0133] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A somatic intelligent robot control method, characterized by, The method comprises the following steps: collecting multi-modal prompt information of a robot to be controlled, wherein the multi-modal prompt information comprises current full-body state and environmental visual information of the robot to be controlled, and a task description input by a user, and the full-body state comprises a pose, a center of gravity distribution, and motion parameters; generating a chain reasoning text corresponding to execution of the task description by the robot to be controlled based on a large-scale pre-trained language model, the multi-modal prompt information, and an atomic skill library applicable to long-sequence full-body linkage actions of the robot, and decomposing the task description into an atomic skill calling sequence according to the chain reasoning text, wherein the atomic skill calling sequence comprises a plurality of atomic skills in calling order, and the atomic skill library is constructed by using a full-body visual motion strategy learning algorithm, and the atomic skill library comprises a target detection skill, a navigation skill, an object grasping and placing skill, a remote grasping and placing object skill, and a target following skill; controlling the robot to be controlled to execute each atomic skill in the calling order after each atomic skill is called and activated in the calling order; wherein the generating of the chain reasoning text corresponding to the execution of the task description by the robot to be controlled based on the large-scale pre-trained language model, the multi-modal prompt information, and the atomic skill library applicable to long-sequence full-body linkage actions of the robot comprises: taking the multi-modal prompt information and a preset chain thinking template as constant information, and taking historical full-body state, historical task records, and execution feedback of the robot to be controlled as memory information; inputting the constant information, the memory information, the task description, and the atomic skill library applicable to long-sequence full-body linkage actions of the robot into the large-scale pre-trained language model, and outputting the chain reasoning text corresponding to the execution of the task description by the robot to be controlled.

2. The somatically intelligent robot control method of claim 1, wherein, After the task description is decomposed into the atomic skill calling sequence according to the chain reasoning text, the method further comprises: assigning a corresponding skill parameter to each atomic skill in the atomic skill calling sequence according to the task description.

3. The somatically intelligent robot control method of claim 1, wherein, After the robot to be controlled is controlled to execute each atomic skill in the calling order after each atomic skill is called and activated in the calling order, the method further comprises: collecting robot action state and environmental feedback data in real time, wherein the robot action state comprises full-body state information of the robot after execution of each atomic skill, and the environmental feedback data comprises environmental visual information after execution of each atomic skill.

4. The somatically intelligent robot control method of claim 3, wherein, After the robot action state and the environmental feedback data are collected in real time, the method further comprises: for each atomic skill, the following steps are performed: taking a difference between the robot action state before and after execution of the atomic skill as a first difference; determining whether the first difference is not greater than a first threshold value; if the first difference is greater than the first threshold value, it is determined that the atomic skill does not achieve an expected effect, a robot action state after execution of the atomic skill and corresponding failure information are combined to form a reflection prompt, and the step of collecting the multi-modal prompt information of the robot to be controlled is returned for continuous execution. If the first difference is not greater than the first threshold, it is determined that the atomic skill achieves the expected effect, and the robot action state after the execution of the atomic skill is stored.

5. The somatically intelligent robot control method of claim 3, wherein, After the real-time collection of the robot action state and the environmental feedback data, the method further includes: For each atomic skill, the following steps are performed: The difference between the environmental visual information before and after the execution of the atomic skill is taken as a second difference; It is determined whether the second difference is not greater than a second threshold; If the second difference is greater than the second threshold, it is determined that the action of the atomic skill does not achieve the expected effect, and the step of collecting the multi-modal prompt information of the robot to be controlled is returned to continue execution; If the second difference is not greater than the second threshold, it is determined that the action of the atomic skill achieves the expected effect, and the environmental visual information after the execution of the atomic skill is stored.

6. A somatic intelligent robot control device, characterized by, It includes: An information collection module for collecting multi-modal prompt information of a robot to be controlled, the multi-modal prompt information including a current full-body state and environmental visual information of the robot to be controlled and a task description input by a user; A generation module for generating a chain reasoning text corresponding to the execution of the task description by the robot to be controlled based on a large-scale pre-trained language model, the multi-modal prompt information and an atomic skill library applicable to long sequence full-body linkage actions of a robot, and decomposing the task description into an atomic skill calling sequence according to the chain reasoning text, the atomic skill calling sequence including a plurality of atomic skills in calling order, the atomic skill library being constructed using a full-body visual motion strategy learning algorithm, and the atomic skill library including a target detection skill, a navigation skill, an object grasping and placing skill, a remote grasping and placing object skill and a target following skill; A control module for controlling the robot to be controlled to execute each atomic skill in the calling order after each atomic skill is called and activated in the calling order; The generation module is configured to: Take the multi-modal prompt information and a preset chain thinking template as constant information, and take historical full-body states, historical task records and execution feedback of the robot to be controlled as memory information; Input the constant information, the memory information, the task description and the atomic skill library into the large-scale pre-trained language model to output the chain reasoning text corresponding to the execution of the task description by the robot to be controlled.

7. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the somatic intelligent robot control method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, the computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 7. The computer program is executed by the processor to implement the steps of the somatic intelligent robot control method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for realizing man-machine interaction inspection of mobile robot by using large language model

    CN116483977A

  • Body-fitting intelligent execution and training method and device based on edge-end collaborative large model

    CN119416817A