Control method and device for intelligent robot with body, terminal and storage medium

By building an atomic skill library based on large-scale pre-trained language model and full-body visual motion strategy learning algorithm, the seamless connection between robot full-body motion and local operation is achieved, and the problem of incoherence in task execution of robots in complex environments is solved, and the task success rate and safety are improved.

CN120244990AActive Publication Date: 2025-07-04SHENZHEN YOUMI TECHNOLOGY TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510630129.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-04
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

In the prior art, the robot's whole-body movement and local operations lack coordinated planning, resulting in incoherent task execution, unsmooth motion switching and insufficient information fusion in complex environments, which cannot meet the high-demand task execution.

Method used

By collecting multimodal prompt information of the robot, atomic skills library is constructed using large-scale pre-trained language models and whole-body visual motion strategy learning algorithms, chain reasoning text is generated and tasks are decomposed into atomic skill call sequences, and seamless connection between robot full-body movement and local operations is achieved.

Benefits of technology

It improves the success rate and security of the robot's task in complex environments, ensures that global navigation quickly reaches the target area and performs precise operations, and improves task processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244990A_ABST
    Figure CN120244990A_ABST
Patent Text Reader

Abstract

The invention provides a control method and device for an intelligent robot with a body, a terminal and a storage medium, and relates to the technical field of robot autonomous task planning and motion control. The method comprises the steps of collecting multi-mode prompt information of a to-be-controlled robot; based on a large-scale pre-training language model, multi-modal prompt information and an atomic skill library suitable for robot long-sequence whole-body linkage actions, generating a corresponding chain reasoning text when a to-be-controlled robot executes task description, and decomposing the task description into atomic skill calling sequences according to the chain reasoning text, the atomic skill calling sequence comprises a plurality of atomic skills according to a calling sequence; and after calling and activating the atomic skills one by one according to the calling sequence, controlling the to-be-controlled robot to execute the atomic skills according to the calling sequence. According to the invention, seamless connection of whole-body motion and local operation of the robot can be realized, so that the robot can efficiently complete a long-sequence task in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of robot autonomous task planning and motion control, and particularly to a control method, device, terminal and storage medium for an embodied intelligent robot. Background Art

[0002] In recent years, with the rapid development of deep learning technology and large-scale pre-trained models in the fields of vision, natural language processing, and robot control, robot task planning and motion control have gradually shifted from traditional rule-based systems to semantic understanding and data-driven directions. Existing research uses vision-language models to convert natural language instructions into specific operation steps, and through pre-trained object detection, grasping pose generation, and motion planning algorithms, the robot realizes the automatic decomposition and execution of local tasks. Moreover, the emergence of embodied intelligence provides an important technical and theoretical basis for enhancing the current cognitive ability of robots and moving towards general intelligence. Through interaction with the environment and humans, the robot can obtain human and environmental feedback from the real physical or virtual digital space, thereby better completing tasks. However, existing technologies mainly focus on the implementation of local actions, often only planning for the robotic arm or local execution units, lacking the organic connection between the robot's whole-body actions and local detailed operations.

[0003] In previous research, researchers proposed a method that combines multi-modal information with pre-trained models to perform task planning for robots and achieve motion control through a preset atomic skill library. Although this method has achieved certain results in object detection, grasping, and placement, etc., when facing scenarios involving whole-body movements (such as global navigation, center-of-gravity adjustment, cross-platform collaboration) in practical applications, there are still problems such as discontinuous planning, uneven action switching, and insufficient information fusion. At the same time, there is also room for improvement in the chain reasoning of the task planning process and the failure detection and replanning mechanism in existing technologies, which cannot fully meet the high requirements of task execution in complex dynamic environments. Summary of the Invention

[0004] This application provides a control method, device, terminal and storage medium for an embodied intelligent robot to solve the problem that existing technologies lack the collaborative planning of the robot's whole-body actions and local operations.

[0005] In a first aspect, this application provides a control method for an embodied intelligent robot, including:

[0006] Collect multi-modal prompt information of the robot to be controlled, where the multi-modal prompt information includes the current whole-body state and environmental visual information of the robot to be controlled and the task description input by the user;

[0007] Based on a large-scale pre-trained language model, the multi-modal prompt information, and an atomic skill library applicable to long-sequence full-body linkage actions of the robot, generate a chain reasoning text corresponding to the task description for the to-be-controlled robot to execute, and decompose the task description into an atomic skill call sequence according to the chain reasoning text. The atomic skill call sequence includes multiple atomic skills in the call order. The atomic skill library is constructed using a full-body vision motion strategy learning algorithm, and the atomic skill library includes object detection skills, navigation skills, object grasping and placing skills, remote object grasping and placing skills, and target following skills;

[0008] After sequentially calling and activating each atomic skill according to the call order, control the to-be-controlled robot to execute each atomic skill according to the call order.

[0009] In a second aspect, the present application provides an embodied intelligent robot control device, including:

[0010] An information collection module, configured to collect multi-modal prompt information of the to-be-controlled robot, where the multi-modal prompt information includes the current full-body state and environmental vision information of the to-be-controlled robot and a task description input by a user;

[0011] A generation module, configured to generate a chain reasoning text corresponding to the task description for the to-be-controlled robot to execute based on a large-scale pre-trained language model, the multi-modal prompt information, and an atomic skill library applicable to long-sequence full-body linkage actions of the robot, and decompose the task description into an atomic skill call sequence according to the chain reasoning text. The atomic skill call sequence includes multiple atomic skills in the call order. The atomic skill library is constructed using a full-body vision motion strategy learning algorithm, and the atomic skill library includes object detection skills, navigation skills, object grasping and placing skills, remote object grasping and placing skills, and target following skills;

[0012] A control module, configured to control the to-be-controlled robot to execute each atomic skill according to the call order after sequentially calling and activating each atomic skill according to the call order.

[0013] In a third aspect, the present application provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described in the first aspect or any possible implementation manner of the first aspect above are implemented.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect or any possible implementation manner of the first aspect above are implemented.

[0015] The present application provides a method, apparatus, terminal, and storage medium for controlling an embodied intelligent robot. By collecting multimodal prompt information of the robot to be controlled, the multimodal prompt information includes the current full-body state and environmental visual information of the robot to be controlled, as well as the task description input by the user. Based on a large-scale pre-trained language model, multimodal prompt information, and an atomic skill library applicable to the long-order full-body linkage actions of the robot, a chain reasoning text corresponding to the task description when the robot to be controlled executes the task description is generated, and the task description is decomposed into an atomic skill call sequence according to the chain reasoning text. The atomic skill call sequence includes multiple atomic skills in the call order. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm, and the atomic skill library includes object detection skills, navigation skills, object grasping and placing skills, remote object grasping and placing skills, and target following skills. After sequentially calling and activating each atomic skill in the call order, the robot to be controlled is controlled to execute each atomic skill in the call order. The present application deeply integrates the full-body state and environmental visual information, and can adjust the action strategy in real time when facing dynamic changes and uncertainties, significantly improving the task success rate and safety. In addition, the present application combines a large-scale pre-trained language model and multimodal prompt information to generate a chain reasoning text and decompose the task into an atomic skill call sequence, which can break down complex tasks into a series of executable atomic skills, enabling the robot to methodically plan the task execution steps, improving the efficiency and accuracy of task processing. At the same time, the atomic skill library is specifically constructed for the long-order full-body linkage actions of the robot, which can achieve seamless connection between the full-body movement and local operations of the robot, enabling the robot to efficiently complete long-sequence tasks in a complex environment, and the global navigation skill ensures that the robot can quickly reach the target area from any initial position, while the local skill ensures precise operation on the target object. Brief Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0017] Figure 1 It is a schematic structural diagram of an embodied intelligent robot platform with full-body movement ability provided by an embodiment of the present application;

[0018] Figure 2 It is a flowchart of the implementation of the method for controlling an embodied intelligent robot provided by an embodiment of the present application;

[0019] Figure 3 It is a schematic diagram of the task execution process provided by an embodiment of the present application;

[0020] Figure 4 It is a schematic structural diagram of the embodied intelligent robot control device provided by an embodiment of the present application;

[0021] Figure 5 It is a schematic diagram of the terminal provided by an embodiment of the present application. Detailed implementation manners

[0022] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0023] To make the objectives, technical solutions, and advantages of the present application clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.

[0024] To solve the problems that the atomic skills of past methods are limited to local actions, lack the whole-body linkage actions of the robot, the multi-modal data perception supplement, and the task planning / re-planning is not detailed enough, the present application proposes an embodied intelligent robot control method based on a multi-modal large model and a whole-body linkage atomic skill library. By integrating global and local information, using a large-scale pre-trained language model for chain reasoning, and designing a dynamic feedback and re-planning mechanism, a whole-body linkage embodied intelligent robot atomic skill library is constructed, thereby improving the autonomous collaborative operation ability and robustness of the robot in complex environments.

[0025] The control method in the present application is executed according to Figure 1 the embodied intelligent robot platform with whole-body movement ability shown. This platform is developed based on the open robot operating system and integrates visual processing, path planning, and large-scale pre-trained model inference modules. All use a unified interface for data interaction, mainly including a multi-modal input module, a task planning module (i.e., a large-scale pre-trained language model), an atomic skill library, a task execution and feedback closed-loop module, and a task memory storage module. The multi-modal input module is connected to the task planning module, the task planning module is respectively connected to the task execution and feedback closed-loop module and the task memory storage module, and the task execution and feedback closed-loop module and the task memory storage module are respectively connected to the atomic skill library.

[0026] In addition, the main hardware of this platform includes:

[0027] (1) A multi-functional mobile chassis that supports global navigation.

[0028] (2) A robotic arm and an end effector responsible for local operations.

[0029] (3) Camera, depth sensor, lidar and inertial measurement unit, responsible for collecting environmental and robot state information.

[0030] (4) Force and torque sensor, providing motion feedback.

[0031] Figure 2 The following is the implementation flowchart of the embodied intelligent robot control method executed based on the embodied intelligent robot platform with full-body motion ability provided by the embodiments of the present application:

[0032] In step 101, multi-modal prompt information of the robot to be controlled is collected. The multi-modal prompt information includes the current full-body state and environmental visual information of the robot to be controlled and the task description input by the user.

[0033] In the embodiments of the present application, the multi-modal input module is used to collect the current full-body state, environmental visual information and the task description of the natural language input by the user of the robot to be controlled, and the collected information is integrated to obtain the multi-modal prompt information of the robot to be controlled. Specifically, a variety of sensors such as cameras, depth sensors, lidars, and inertial measurement units are used to collect the current full-body state and environmental visual information of the robot to be controlled. Among them, the environmental visual information is represented by images and depth maps. The full-body state includes pose information, and the collected images, depth maps and the pose information of the robot to be controlled are organically integrated with the task description input by the user to form multi-modal prompt information in a unified form.

[0034] In addition, before the task starts, the embodiments of the present application first need to collect panoramic environmental images, depth maps and the full-body state data of the robot to be controlled through the sensors in the multi-modal input module. The lidar and camera data are used for environmental mapping, and the sensor data are synchronized and calibrated in time to form a unified data stream. The full-body state includes pose, center of gravity distribution, motion parameters, etc., and is passed to the task planning module as a memory message.

[0035] The embodiments of the present application deeply integrate visual, language, full-body state and environmental map information, enabling the robot to adjust its action strategy in real time when facing dynamic changes and uncertainties, and significantly improving the task success rate and safety.

[0036] In step 102, based on a large-scale pre-trained language model, multi-modal prompt information, and an atomic skill library applicable to long-sequence whole-body linkage actions of the robot, a chain reasoning text corresponding to the task description to be controlled for the robot to execute is generated, and the task description is decomposed into an atomic skill call sequence according to the chain reasoning text. The atomic skill call sequence includes multiple atomic skills in the call order. The atomic skill library is constructed using a whole-body visuo-motor attention learning algorithm and includes object detection skills, navigation skills, object grasping and placement skills, remote object grasping and placement skills, and target following skills.

[0037] Among them, the atomic skill library in the embodiments of this application is extended based on existing research. In addition to traditional local operations such as object detection, grasping pose generation, arm motion planning, object placement and release, etc., based on the whole-body visuo-motor attention learning algorithm, long-sequence atomic skills of whole-body linkage are realized, including long-sequence whole-body linkage actions such as object grasping and placement, remote object grasping and placement. Among them, the whole-body visuo-motor attention learning algorithm (WB-VIMA) is an innovative algorithm specifically used to learn the whole-body visuo-motor strategy of the robot, aiming to improve the whole-body operation ability of the robot in the home environment, especially performing well when executing complex daily home tasks.

[0038] The specific generation process of the atomic skill library is as follows: Using the WB-VIMA algorithm, first predict the future trajectory of the mobile chassis through its autoregressive denoising process, and this prediction serves as the condition for subsequent torso and arm motion predictions. The WB-VIMA algorithm mainly relies on the robot's own multi-modal sensing data and combines the robot's body state (such as joint positions and chassis speed). After real-time fusion of this information, it is used to generate coordinated navigation and whole-body action instructions. Specifically, the atomic skill library includes but is not limited to the following:

[0039] Object detection skill: Through a large-scale pre-trained language model, extract the target area according to the task description and image information, and generate a target segmentation mask.

[0040] Navigation skill: Utilize the global map and sensor data to achieve global navigation from the current position to the target area through a global path planning algorithm, and perform dynamic adjustment by combining vision and inertial data.

[0041] Object grasping and placement skill: Through the linkage of the whole body of the robot, achieve the ability of object grasping and placement.

[0042] Remote object grasping and placement skill: Through the whole-body linkage of the robot + long-sequence navigation ability, ensure that the robot can smoothly complete the task during the process of remotely placing the item after grasping.

[0043] Target following skill: Through target tracking + whole-body linkage of the robot + long-sequence navigation ability, ensure that the robot can track the target in a long sequence and provide better service.

[0044] In the embodiments of the present application, each atomic skill adopts a unified interface format, and the input parameters include task description, visual data, depth information, sensor feedback, and motion constraints, etc. The modular design of the atomic skill library enables each skill to be executed independently or flexibly combined according to task requirements, providing a standardized call interface for the task planning module.

[0045] In the embodiments of the present application, the task planning module includes both skill calls for local operations and whole-body motion planning such as global navigation and center-of-gravity adjustment. The overall process of the planning steps is as follows:

[0046] Construct a unified prompt based on multi-modal input and splice the multi-modal prompt information into a standardized input.

[0047] Generate a detailed chain reasoning text through a large-scale pre-trained language model and decompose the task description into a sequence of atomic skill calls.

[0048] Output a structured result, including the skill call order and the specific parameters of each skill.

[0049] Specifically, based on the reasoning method of the large-scale pre-trained language model in the task planning module, using a preset chain-of-thought template and memory context information, input the multi-modal prompt information and atomic skills applicable to the long-sequence whole-body linkage actions of the robot into the large-scale pre-trained language model. The large-scale pre-trained language model performs chain reasoning according to the preset template to generate a structured planning output, that is, the chain reasoning text corresponding to the robot to execute the current task description. Among them, the chain reasoning text details the chain reasoning process of task decomposition. And decompose the current task description into a sequence of atomic skill calls according to the chain reasoning text. Among them, the atomic skill call sequence includes multiple atomic skills arranged in the call order, and the call order lists the continuous operations from global navigation, target detection, grasping pose generation, local motion planning, whole-body center-of-gravity adjustment to specific grasping and placement. And the output result of the task planning module in the embodiments of the present application is organized through a preset unified interface format, which is convenient for the call of the task execution and feedback closed-loop module.

[0050] The design of the atomic skill library with a unified interface in the embodiments of the present application enables the robot platform to flexibly integrate new skills and support the demand changes in different task scenarios.

[0051] In a possible implementation, based on a large-scale pre-trained language model, multi-modal prompt information, and an atomic skill library applicable to long-sequence full-body linkage actions of a robot, a chain reasoning text corresponding to the task description to be controlled for the robot to execute can be generated, which may include:

[0052] Regarding the multi-modal prompt information and the preset chain of thought template as constant information, and regarding the historical full-body state, historical task record, and execution feedback of the robot to be controlled as memory information;

[0053] Input the constant information, memory information, task description, and the atomic skill library applicable to long-sequence full-body linkage actions of the robot into the large-scale pre-trained language model, and output the chain reasoning text corresponding to the task description to be controlled for the robot to execute.

[0054] Among them, in the embodiments of the present application, according to the collected images, depth maps, and full-body states, the task description is integrated with the content of each sensor data and the preset atomic skill library into standardized multi-modal prompt information. Among them, the multi-modal prompt information can be divided into four parts, namely:

[0055] (1) Constant information, embedding a detailed chain of thought template to guide the large-scale pre-trained language model to logically decompose the task description. Specifically, it includes visual perception information such as RGB images, depth images, and point clouds, and the preset chain of thought template, requiring the large-scale pre-trained language model to gradually decompose the task description and clarify the skills to be called and their parameter formats for each step.

[0056] (2) Memory information, recording the recent state and historical task information of the robot. Specifically, it includes the robot's last full-body state, historical task record, and execution feedback, which are used to maintain continuity in multi-stage tasks.

[0057] (3) Task description, that is, the specific task objective described by the user in natural language.

[0058] (4) Atomic skill library, listing all the current atomic skills and their function descriptions.

[0059] Optionally, input the constant information, memory information, task description, and atomic skill library into the large-scale pre-trained language model to generate a structured output, that is, the chain reasoning text corresponding to the task description to be controlled for the robot to execute.

[0060] The embodiments of the present application rely on the reasoning ability of the large-scale pre-trained language model and the preset thinking template, can clearly explain the task decomposition logic, and perform re-planning according to real-time feedback during the task execution process to ensure the continuity and robustness of the entire task process.

[0061] In a possible implementation, after decomposing the task description into a sequence of atomic skill calls according to the chain reasoning text, the method may further include:

[0062] According to the task description, corresponding skill parameters are assigned to each atomic skill in the atomic skill call sequence.

[0063] Optionally, after obtaining the atomic skill call sequence, specific input parameters are provided for each atomic skill in the atomic skill call sequence, such as target position, visual data, action amplitude, image data, depth data, global map information, center of gravity state, and motion control parameters.

[0064] In the embodiments of the present application, during the process of generating the chain reasoning text and the atomic skill call sequence, it is ensured that the planning output not only has a transparent reasoning process but also meets the interface requirements for subsequent specific action execution.

[0065] In step 103, after calling and activating each atomic skill one by one in the calling order, the robot to be controlled is controlled to execute each atomic skill in the calling order.

[0066] In the embodiments of the present application, the skill call sequence output by the task planning module is passed to the task execution and feedback closed-loop module. The task execution and feedback closed-loop module calls and activates each atomic skill one by one according to the calling order to control the robot to be controlled to execute according to each atomic skill.

[0067] The embodiments of the present application can achieve seamless connection between the whole-body movement and local operation of the robot, enabling the robot to efficiently complete long-sequence tasks in a complex environment. The global navigation skill ensures that the robot quickly reaches the target area from any initial position, while the local skill guarantees precise operation on the target object.

[0068] In a possible implementation, after calling and activating each atomic skill one by one in the calling order and controlling the robot to be controlled to execute each atomic skill in the calling order, the method may further include:

[0069] Robot action state and environmental feedback data are collected in real time. The robot action state includes the whole-body state information of the robot after the execution of each atomic skill, and the environmental feedback data includes the environmental visual information after the execution of each atomic skill.

[0070] Optionally, the task is executed according to the atomic skill call order planned by the large-scale pre-trained language model, and the robot action state and environmental feedback data are collected in real time during the execution process.

[0071] During the entire execution process of the embodiments of the present application, the closed-loop feedback mechanism ensures that the robot can continuously adjust its motion planning according to real-time changes in a complex and dynamic environment, achieving seamless connection and robust control between global navigation and local operations, and enhancing the robustness and stability of task execution.

[0072] After real-time acquisition of the robot's motion state and environmental feedback data, a feedback adjustment mechanism is carried out. The feedback mechanism in the embodiments of the present application includes two means, namely sensor feedback detection and status monitoring and fault replanning.

[0073] Sensor feedback detection is to judge whether the target has an expected change by comparing the front and rear images and depth data after each atomic skill is executed.

[0074] Status monitoring and fault replanning is to use force and pose sensors to detect whether the end effector or the whole body motion of the robot reaches the predetermined state.

[0075] Then when it is detected that the action fails or does not meet the expectation, the current state, the cause of the fault, and the historical operation records are formed into a reflection prompt, which is fed back to the task planning module, and chain reasoning is carried out again to generate a new skill call sequence and replan the subsequent actions.

[0076] The embodiments of the present application utilize the above-mentioned closed-loop feedback mechanism, have the ability of dynamic adjustment and adaptive recovery, and greatly improve the overall task success rate.

[0077] In a possible implementation manner, after real-time acquisition of the robot's motion state and environmental feedback data, the method may further include:

[0078] For each atomic skill, the following steps are executed:

[0079] Taking the difference between the robot's motion states before and after the execution of the atomic skill as the first difference;

[0080] Judging whether the first difference is not greater than the first threshold;

[0081] If the first difference is greater than the first threshold, it is determined that the atomic skill does not achieve the expected effect, and the robot's motion state after the execution of the atomic skill and the corresponding fault information are formed into a reflection prompt, and the step of collecting multi-modal prompt information of the robot to be controlled is returned to continue execution;

[0082] If the first difference is not greater than the first threshold, it is determined that the atomic skill achieves the expected effect, and the robot's motion state after the execution of the atomic skill is stored.

[0083] Optionally, the process is a state monitoring and fault planning feedback mechanism, that is: for each atomic skill, the difference in the robot's action state before and after the execution of the atomic skill by the robot to be controlled is used as the first difference, and it is judged whether the first difference is not greater than the first threshold. If it is greater, it is determined that the robot to be controlled fails to achieve the expected effect under the execution of the atomic skill, such as the target not reaching the predetermined position, insufficient grasping force, the robot losing balance, etc. Then, the fault information and the current state data are constructed into a reflection prompt, and after the reflection prompt is passed to the task planning module, the multi-modal prompt information collection step is returned to continue execution, and a new skill call sequence is generated by the large-scale pre-trained language model to remedy the failed part; if it is not greater, it is determined that the robot to be controlled achieves the expected effect under the execution of the atomic skill, and the robot's action state after the execution of the atomic skill that achieves the expected effect is stored in the task memory storage module.

[0084] In a possible implementation manner, after the robot's action state and environmental feedback data are collected in real time, the method may further include:

[0085] For each atomic skill, the following steps are performed:

[0086] The difference in the environmental visual information before and after the execution of the atomic skill is used as the second difference;

[0087] Judge whether the second difference is not greater than the second threshold;

[0088] If the second difference is greater than the second threshold, it is determined that the action of the atomic skill fails to achieve the expected effect, and the step of collecting the multi-modal prompt information of the robot to be controlled is returned to continue execution;

[0089] If the second difference is not greater than the second threshold, it is determined that the action of the atomic skill achieves the expected effect, and the environmental visual information after the execution of the atomic skill is stored.

[0090] Optionally, the process is a sensor feedback detection mechanism, that is: for each atomic skill, the difference in the environmental visual information before and after the execution of the atomic skill by the robot to be controlled is used as the second difference, such as comparing the change in the depth map before and after execution, the force sensor feedback, the position deviation, etc., and it is judged whether the second difference is not greater than the second threshold. If it is greater, it is determined that the robot to be controlled fails to achieve the expected effect under the execution of the atomic skill, and at this time, the multi-modal prompt information collection step is directly returned to execute again; if it is not greater, it is determined that the robot to be controlled achieves the expected effect under the execution of the atomic skill, and the robot's action state after the execution of the atomic skill that achieves the expected effect is stored in the task memory storage module.

[0091] The introduction of the task memory storage module in the embodiments of the present application helps to continuously optimize the planning strategy during long-term operation and achieve self-learning and evolution.

[0092] Exemplarily, referring to Figure 3 , the task description input by the user is "Get me a bottle of water". Then, the task description, the current full-body state of the robot, and the environmental visual information are input into the task planning module through the multimodal input module, and the task description is decomposed into three actions: "fetch water", "navigate to the user", and "hand over the water". Then, the above three actions are passed to the task execution and feedback closed-loop module, and the corresponding atomic skills are executed in the call order, and it is judged whether each atomic skill is successfully completed. If successful, the result of successful completion is directly stored in the task memory storage module; if failed, the fault information and the current data are combined to form a reflection prompt and passed back to the task planning module for re-planning.

[0093] The present application provides a method for controlling an embodied intelligent robot. By collecting multimodal prompt information of the robot to be controlled, the multimodal prompt information includes the current full-body state of the robot to be controlled, environmental visual information, and a task description input by the user; based on a large-scale pre-trained language model, multimodal prompt information, and an atomic skill library applicable to the long-order full-body linkage actions of the robot, a chain reasoning text corresponding to the robot to be controlled when executing the task description is generated, and the task description is decomposed into an atomic skill call sequence according to the chain reasoning text. The atomic skill call sequence includes multiple atomic skills in the call order. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm. The atomic skill library includes object detection skills, navigation skills, object grasping and placing skills, remote object grasping and placing skills, and target following skills; after calling and activating each atomic skill one by one in the call order, the robot to be controlled is controlled to execute each atomic skill in the call order. The present application deeply integrates the full-body state and environmental visual information, and can adjust the action strategy in real time in the face of dynamic changes and uncertainties, significantly improving the task success rate and safety; in addition, the present application combines a large-scale pre-trained language model and multimodal prompt information to generate a chain reasoning text and decompose the task into an atomic skill call sequence, which can disassemble complex tasks into a series of executable atomic skills, enabling the robot to methodically plan the task execution steps, improving the efficiency and accuracy of task processing; at the same time, the atomic skill library is specifically constructed for the long-order full-body linkage actions of the robot, which can achieve seamless connection between the full-body movement and local operations of the robot, enabling the robot to efficiently complete long-sequence tasks in a complex environment, and the global navigation skill ensures that the robot quickly reaches the target area from any initial position, while the local skill ensures precise operation on the target object.

[0094] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0095] The following is an apparatus embodiment of the present application. For details not described in detail, reference may be made to the corresponding method embodiment above.

[0096] Figure 4 The structural schematic diagram of the embodied intelligent robot control device provided by the embodiment of the present application is shown. For the convenience of description, only the parts related to the embodiment of the present application are shown and are described in detail as follows:

[0097] As Figure 4 shown, the embodied intelligent robot control device 4 includes:

[0098] An information acquisition module 41, configured to acquire multi-modal prompt information of the robot to be controlled, where the multi-modal prompt information includes the current whole-body state and environmental visual information of the robot to be controlled and the task description input by the user;

[0099] A generation module 42, configured to generate a chain reasoning text corresponding to the task description when the robot to be controlled executes the task description based on a large-scale pre-trained language model, multi-modal prompt information, and an atomic skill library applicable to the long-order whole-body linkage actions of the robot, and decompose the task description into an atomic skill call sequence according to the chain reasoning text. The atomic skill call sequence includes multiple atomic skills in the call order. The atomic skill library is constructed by using a whole-body visual motion strategy learning algorithm, and the atomic skill library includes object detection skills, navigation skills, object grasping and placing skills, remote object grasping and placing skills, and target following skills;

[0100] A control module 43, configured to control the robot to be controlled to execute each atomic skill in the call order after calling and activating each atomic skill in the call order one by one.

[0101] The present application provides an embodied intelligent robot control device, which collects multi-modal prompt information of the robot to be controlled. The multi-modal prompt information includes the current full-body state and environmental visual information of the robot to be controlled and the task description input by the user. Based on a large-scale pre-trained language model, the multi-modal prompt information, and an atomic skill library applicable to the long-order full-body linkage actions of the robot, a chain reasoning text corresponding to the robot to be controlled when executing the task description is generated, and the task description is decomposed into an atomic skill call sequence according to the chain reasoning text. The atomic skill call sequence includes multiple atomic skills in the call order. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm and includes object detection skills, navigation skills, object grasping and placing skills, remote object grasping and placing skills, and target following skills. After sequentially calling and activating each atomic skill in the call order, the robot to be controlled is controlled to execute each atomic skill in the call order. The present application deeply integrates the full-body state and environmental visual information, and can adjust the action strategy in real time when facing dynamic changes and uncertainties, significantly improving the task success rate and safety. In addition, the present application combines a large-scale pre-trained language model and multi-modal prompt information to generate a chain reasoning text and decompose the task into an atomic skill call sequence, which can break down complex tasks into a series of executable atomic skills, enabling the robot to methodically plan the task execution steps and improve the efficiency and accuracy of task processing. At the same time, the atomic skill library is specifically constructed for the long-order full-body linkage actions of the robot, enabling seamless connection between the full-body movement and local operations of the robot, enabling the robot to efficiently complete long-sequence tasks in a complex environment, and the global navigation skill ensures that the robot quickly reaches the target area from any initial position, while the local skills ensure precise operation of the target object.

[0102] In a possible implementation manner, the generation module can be used for:

[0103] Regarding the multi-modal prompt information and the preset chain of thought template as constant information, and regarding the historical full-body state, historical task records, and execution feedback of the robot to be controlled as memory information;

[0104] Inputting the constant information, memory information, task description, and atomic skill library into the large-scale pre-trained language model, and outputting the chain reasoning text corresponding to the robot to be controlled when executing the task description.

[0105] In a possible implementation manner, the device may further include a parameter empowerment module, and the parameter empowerment module can be used for:

[0106] According to the task description, corresponding skill parameters are assigned to each atomic skill in the atomic skill call sequence.

[0107] In a possible implementation, the device may further include a data acquisition module, and the data acquisition module may be used for:

[0108] Real-time collect the robot action state and environmental feedback data. The robot action state includes the whole-body state information of the robot after each atomic skill is executed, and the environmental feedback data includes the environmental visual information after each atomic skill is executed.

[0109] In a possible implementation, the device may further include a first judgment module, and the first judgment module may be used for:

[0110] For each atomic skill, perform the following steps:

[0111] Take the difference between the robot action states before and after the execution of the atomic skill as the first difference;

[0112] Judge whether the first difference is not greater than the first threshold;

[0113] If the first difference is greater than the first threshold, determine that the atomic skill has not achieved the expected effect, and form a reflection prompt by combining the robot action state after the execution of the atomic skill with the corresponding fault information, and return to the step of collecting multi-modal prompt information of the robot to be controlled to continue execution;

[0114] If the first difference is not greater than the first threshold, determine that the atomic skill has achieved the expected effect, and store the robot action state after the execution of the atomic skill.

[0115] In a possible implementation, the device may further include a second judgment module, and the second judgment module may be used for:

[0116] For each atomic skill, perform the following steps:

[0117] Take the difference between the environmental visual information before and after the execution of the atomic skill as the second difference;

[0118] Judge whether the second difference is not greater than the second threshold;

[0119] If the second difference is greater than the second threshold, determine that the action of the atomic skill has not achieved the expected effect, and return to the step of collecting multi-modal prompt information of the robot to be controlled to continue execution;

[0120] If the second difference is not greater than the second threshold, determine that the action of the atomic skill has achieved the expected effect, and store the environmental visual information after the execution of the atomic skill.

[0121] Figure 5 It is a schematic diagram of the terminal provided by the embodiments of the present application. As Figure 5As shown, the terminal 5 of this embodiment includes: a processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50. When the processor 50 executes the computer program 52, it implements the steps in the above embodiments of various embodied intelligent robot control methods, such as Figure 1 the steps 101 to 103 shown. Alternatively, when the processor 50 executes the computer program 52, it implements the functions of each module / unit in the above device embodiments, such as Figure 4 the functions of each module shown.

[0122] Exemplarily, the computer program 52 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 51 and executed by the processor 50 to complete this application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 52 in the terminal 5. For example, the computer program 52 can be divided into Figure 4 the various modules shown.

[0123] The terminal 5 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art can understand that Figure 5 this is only an example of the terminal 5 and does not constitute a limitation on the terminal 5. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the terminal may further include input / output devices, network access devices, a bus, etc.

[0124] The so-called processor 50 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0125] The memory 51 may be an internal storage unit of the terminal 5, such as the hard disk or memory of the terminal 5. The memory 51 may also be an external storage device of the terminal 5, such as a plug-in hard disk equipped on the terminal 5, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 51 may also include both the internal storage unit of the terminal 5 and external storage devices. The memory 51 is used to store the computer program and other programs and data required by the terminal. The memory 51 may also be used to temporarily store data that has been output or is to be output.

[0126] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0127] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0128] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0129] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0130] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0131] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0132] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned embodiment methods of the embodied intelligent robot control method can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0133] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A control method for an embodied intelligent robot, characterized in that, Including: Collecting multimodal prompt information of the robot to be controlled, where the multimodal prompt information includes the current full-body state and environmental visual information of the robot to be controlled and the task description input by the user; Based on a large-scale pre-trained language model, the multimodal prompt information, and an atomic skill library applicable to the long-order full-body linkage actions of the robot, generating a chain reasoning text corresponding to the robot to be controlled when executing the task description, and decomposing the task description into an atomic skill call sequence according to the chain reasoning text. The atomic skill call sequence includes multiple atomic skills in the call order. The atomic skill library is constructed using a full-body visual motion strategy learning algorithm, and the atomic skill library includes object detection skills, navigation skills, object grasping and placing skills, remote object grasping and placing skills, and target following skills; After sequentially calling and activating each atomic skill according to the call order, controlling the robot to be controlled to execute each atomic skill according to the call order.

2. The embodied intelligent robot control method according to claim 1, wherein The generating, based on a large-scale pre-trained language model, the multimodal prompt information, and an atomic skill library applicable to the long-order full-body linkage actions of the robot, a chain reasoning text corresponding to the robot to be controlled when executing the task description includes: Regarding the multimodal prompt information and a preset chain thinking template as constant information, and regarding the historical full-body state, historical task record, and execution feedback of the robot to be controlled as memory information; Inputting the constant information, the memory information, the task description, and an atomic skill library applicable to the long-order full-body linkage actions of the robot into the large-scale pre-trained language model, and outputting a chain reasoning text corresponding to the robot to be controlled when executing the task description.

3. The embodied intelligent robot control method according to claim 1, characterized in that After decomposing the task description into an atomic skill call sequence according to the chain reasoning text, the method further includes: Assigning corresponding skill parameters to each atomic skill in the atomic skill call sequence according to the task description.

4. The embodied intelligent robot control method according to claim 1, wherein, After controlling the robot to be controlled to execute each atomic skill according to the call order after sequentially calling and activating each atomic skill according to the call order, the method further includes: Real-time collecting the robot action state and environmental feedback data. The robot action state includes the full-body state information of the robot after each atomic skill is executed, and the environmental feedback data includes the environmental visual information after each atomic skill is executed.

5. The embodied intelligent robot control method according to claim 4, wherein, After the real-time collecting the robot action state and environmental feedback data, the method further includes: For each atomic skill, performing the following steps: Taking the difference between the robot action states before and after the execution of this atomic skill as the first difference; Judging whether the first difference is not greater than a first threshold; If the first difference is greater than the first threshold, determining that this atomic skill fails to achieve the expected effect, and forming a reflection prompt with the robot action state after the execution of this atomic skill and the corresponding fault information, and returning to the step of collecting the multimodal prompt information of the robot to be controlled to continue execution; If the first difference is not greater than the first threshold, it is determined that the atomic skill achieves the expected effect, and the action state of the robot after the execution of the atomic skill is stored.

6. The embodied intelligent robot control method according to claim 4, wherein After the real-time acquisition of the robot action state and environmental feedback data, the method further includes: For each atomic skill, the following steps are performed: Taking the difference between the environmental visual information before and after the execution of the atomic skill as the second difference; Judging whether the second difference is not greater than the second threshold; If the second difference is greater than the second threshold, it is determined that the action of the atomic skill does not achieve the expected effect, and the step of collecting multimodal prompt information of the robot to be controlled is returned to continue execution; If the second difference is not greater than the second threshold, it is determined that the action of the atomic skill achieves the expected effect, and the environmental visual information after the execution of the atomic skill is stored.

7. An embodied intelligent robot control device, characterized in that, Including: An information acquisition module for acquiring multimodal prompt information of the robot to be controlled, where the multimodal prompt information includes the current whole-body state and environmental visual information of the robot to be controlled and the task description input by the user; A generation module for generating a chain reasoning text corresponding to the execution of the task description by the robot to be controlled based on a large-scale pre-trained language model, the multimodal prompt information, and an atomic skill library applicable to the long-order whole-body linkage actions of the robot, and decomposing the task description into an atomic skill call sequence according to the chain reasoning text, where the atomic skill call sequence includes multiple atomic skills in the call order, and the atomic skill library is constructed using a whole-body visual motion strategy learning algorithm, and the atomic skill library includes object detection skills, navigation skills, object grasping and placing skills, remote object grasping and placing skills, and target following skills; A control module for controlling the robot to be controlled to execute each atomic skill in the call order after calling and activating each atomic skill in the call order one by one.

8. The embodied intelligent robot control device according to claim 7, wherein The generation module is used for: Taking the multimodal prompt information and a preset chain thinking template as constant information, and taking the historical whole-body state, historical task record, and execution feedback of the robot to be controlled as memory information; Inputting the constant information, the memory information, the task description, and the atomic skill library into the large-scale pre-trained language model, and outputting the chain reasoning text corresponding to the execution of the task description by the robot to be controlled.

9. A terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the embodied intelligent robot control method according to any one of claims 1 to 6 above are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the embodied intelligent robot control method according to any one of claims 1 to 6 above are implemented.

Citation Information

Patent Citations

  • Robot based on multi-modal fusion in complex limited environment and operation method

    CN115319764A

  • Method for realizing man-machine interaction inspection of mobile robot by using large language model

    CN116483977A

  • Body-fitting intelligent execution and training method and device based on edge-end collaborative large model

    CN119416817A

  • Robot control method based on multi-modal large model

    CN119897864A

  • Action learning method, medium, and electronic device

    US20220203523A1

Cited By

  • Robot control method and device and storage medium

    CN120921403A

  • Multi-modal large model body planning method and system based on iterative feedback

    CN121480559A