Control method and equipment of intelligent robot with body and storage medium

By integrating environmental perception data in real time to generate task semantic information and using an implicit planner to generate action tags, the embodied intelligent robot can flexibly respond to environmental changes, solve the coupling problem between action instructions and predefined scenarios, and improve the stability and success rate of task execution.

CN120663323APending Publication Date: 2025-09-19YOUDI ROBOT (WUXI) CO LTD
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510993644.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

When embodied intelligent robots face changes in environmental parameters or perform new tasks, the deep coupling of action instructions with predefined scenarios leads to a high probability of task failure.

Method used

By real-time integration of visual, tactile and other environmental perception data, task semantic information is generated, and the implicit planner is used to parse the information, generate implicit action tags, decompose the task into multiple task sub-goals, and generate and execute action instruction sequences.

Benefits of technology

The coupling degree between action instructions and predefined scenarios is reduced, the robot's adaptability to environmental changes and the robustness of task execution are improved, and the probability of task execution failure is significantly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120663323A_ABST
    Figure CN120663323A_ABST
Patent Text Reader

Abstract

The invention discloses a control method and equipment for an intelligent robot with a body and a storage medium, and belongs to the technical field of robots. The method comprises the steps of receiving a task instruction and environment perception data, performing fusion processing on the task instruction and the environment perception data, generating task semantic information associated with the task instruction, analyzing the task semantic information through an implicit planner, generating an implicit action mark, and based on the implicit action mark, generating an implicit action. And decomposing a to-be-executed task corresponding to the task instruction into a plurality of task sub-targets layer by layer, generating an action instruction sequence based on the task sub-targets, and executing a control action corresponding to the action instruction sequence. According to the method, based on the constraint of the implicit action mark and the action instruction sequence, the robot with the body can flexibly respond to environment parameter changes such as object position deviation or new task requirements in the task execution process, the coupling degree of the action instruction of the robot with the body and the predefined scene is reduced, and high stability of task execution is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robotics technology, and in particular to a control method, device, and storage medium for an embodied intelligent robot. Background Art

[0002] An embodied intelligent robot is one that can grow in intelligence through interaction with its environment. In related technologies, embodied intelligent robots have pre-defined mapping relationships stored in them. The robot matches collected environmental data and task instructions with predefined scenarios and obtains the action instruction sequence associated with the predefined scenario. The robot then sequentially executes the control actions corresponding to the action instructions in the action instruction sequence.

[0003] However, in this scenario, the motion commands generated by the embodied intelligent robot are deeply coupled to the predefined scenario. When environmental parameters change (such as object position shifts) or when performing a new task, the mismatch between the motion command sequence and the real-time environment can lead to anomalies such as collisions during robot movement, resulting in a high probability of task failure.

[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a control method, device and storage medium for an embodied intelligent robot, aiming to solve the technical problem that the probability of failure in task execution of the embodied intelligent robot is high.

[0006] To achieve the above objectives, the present application provides a control method for an embodied intelligent robot, the method comprising the following steps:

[0007] receiving a task instruction and environmental perception data, wherein the task instruction is input via language and / or visual modality, and the environmental perception data includes at least one modality input of vision and touch;

[0008] fusing the task instruction and the environmental perception data to generate task semantic information associated with the task instruction;

[0009] Parsing the task semantic information through an implicit planner to generate implicit action tags;

[0010] Based on the implicit action tag, decomposing the to-be-executed task corresponding to the task instruction into multiple task sub-goals layer by layer;

[0011] An action instruction sequence is generated based on the task sub-goal, and a control action corresponding to the action instruction sequence is executed.

[0012] In one embodiment, the step of fusing the task instruction and the environmental perception data to generate task semantic information associated with the task instruction includes:

[0013] extracting multi-scale visual features from the environmental perception data through a convolutional neural network in a visual encoder;

[0014] Identifying action objects and constraints in the task instructions through an attention mechanism in a language model to obtain the natural language semantics of the task instructions;

[0015] The multi-scale visual features are aligned with the natural language semantics through cross-modal attention to generate the task semantic information.

[0016] In one embodiment, the step of parsing the task semantic information by an implicit planner to generate implicit action tags includes:

[0017] Inputting the task semantic information into the task decomposition model in the implicit planner, modeling the dependency relationship of the task semantic information through a graph neural network, and generating a task decomposition path;

[0018] According to the task decomposition path, matching adapted action units in a predefined high-level action template library, wherein the high-level action template library includes action patterns and corresponding action constraint information;

[0019] The action unit is encoded as an implicit action tag, wherein the tag includes at least one of the action mode, execution priority, and environment constraint parameters.

[0020] In one embodiment, the step of decomposing the to-be-executed task corresponding to the task instruction into a plurality of task sub-goals layer by layer based on the implicit action tag further includes:

[0021] According to the action mode in the implicit action tag, mapping the task to be executed into at least two action levels through a hierarchical state machine;

[0022] In each of the action levels, combining the environmental perception data and the action constraint information, adjusting the action execution parameters of the action level to form a corresponding atomic action;

[0023] The task sub-goal is determined according to the atomic action.

[0024] In one embodiment, the step of generating an action instruction sequence based on the task sub-goal and executing a control action corresponding to the action instruction sequence includes:

[0025] The task sub-goals are received by the action expert agent, and the inverse kinematics solver in the robot dynamics model is called to calculate the target pose of each joint;

[0026] generating an end-effector trajectory through a path planning algorithm based on the spatial distribution of obstacles in the environmental perception data;

[0027] Encoding the target posture and the end effector trajectory into the action instruction sequence including joint angles, velocities and / or accelerations in time sequence;

[0028] The control actions corresponding to the action instructions in the action instruction sequence are executed in sequence.

[0029] In one embodiment, after the steps of generating an action instruction sequence based on the task sub-goal and executing the control action corresponding to the action instruction sequence, the method further includes:

[0030] Capturing object posture change data through visual sensors, and / or collecting contact force feedback data through tactile sensors;

[0031] Inputting the object pose change data and / or the contact force feedback data into an implicit planner, and updating the environmental constraint parameters corresponding to the implicit action markers through an optimization algorithm;

[0032] According to the updated environmental constraint parameters, the replanning action of the task sub-goal is executed, and the action instruction sequence is updated.

[0033] In one embodiment, before the step of parsing the task semantic information by an implicit planner to generate implicit action tags, the step further includes:

[0034] Receiving a training dataset, wherein the training dataset includes task semantic information samples and labeled true values ​​of corresponding implicit actions;

[0035] Inputting the task semantic information sample into a graph neural network to obtain a predicted label of the implicit action, and calculating the semantic similarity loss between the predicted label and the true value of the label;

[0036] The graph neural network is iterated based on the semantic similarity loss, and the implicit planner is obtained based on the iterated graph neural network.

[0037] In one embodiment, the step of parsing the task semantic information by an implicit planner to generate implicit action tags further includes:

[0038] Obtaining internet video data, extracting spatiotemporal features of a target action segment from the internet video data, and encoding the target action segment into an implicit action pattern vector based on the spatiotemporal features;

[0039] Acquiring cross-ontology operation data of the target robot and extracting action semantics of the cross-ontology operation data;

[0040] The action semantics and / or the implicit action pattern vector are mapped into action units, and the action units are stored in a predefined high-level action template library.

[0041] In addition, to achieve the above-mentioned purpose, the present application also provides a control device for an embodied intelligent robot, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the control method for the embodied intelligent robot as described above.

[0042] In addition, to achieve the above-mentioned purpose, the present application also provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the control method of the embodied intelligent robot as described above are implemented.

[0043] One or more technical solutions proposed in this application have at least the following technical effects:

[0044] The present application fuses environmental perception data including at least one modal input such as visual and tactile data in real time, and performs fusion processing to generate task semantic information associated with task instructions. In the planning stage, implicit action tags corresponding to the task semantic information are formed, and the task to be executed is decomposed layer by layer into multiple task sub-goals. Action instruction sequences are generated and executed, so that the action instruction sequences can flexibly respond to changes in environmental parameters such as object position offset or new task requirements, reduce the coupling degree between the action instructions of the embodied intelligent robot and predefined scenes, and avoid abnormal situations such as collisions caused by mismatch between the instruction sequence and the real-time environmental status, thereby significantly improving the environmental adaptability and task robustness of the embodied robot, ensuring high stability of task execution in dynamic scenes, and reducing the probability of failure. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1This is a flow chart of the first embodiment of the control method of the embodied intelligent robot of the present application;

[0048] Figure 2 This is a flow chart of the second embodiment of the control method of the embodied intelligent robot of the present application;

[0049] Figure 3 This is a flow chart of the third embodiment of the control method of the embodied intelligent robot of the present application;

[0050] Figure 4 This is a flow chart of a fourth embodiment of the control method of the embodied intelligent robot of the present application;

[0051] Figure 5 It is a structural diagram of the control device of the embodied intelligent robot in the hardware operating environment involved in the embodiment of the present application.

[0052] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0053] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0054] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0055] The main solution of the embodiment of the present application is: receiving task instructions and environmental perception data, wherein the task instructions are input through language and / or visual modalities, and the environmental perception data includes at least one modal input of vision and touch; fusing the task instructions and the environmental perception data to generate task semantic information associated with the task instructions; parsing the task semantic information through an implicit planner to generate implicit action tags; based on the implicit action tags, decomposing the task to be executed corresponding to the task instruction into multiple task sub-goals layer by layer; generating an action instruction sequence based on the task sub-goals, and executing the control action corresponding to the action instruction sequence.

[0056] Because in the prior art, the embodied intelligent robot has a pre-defined mapping relationship. The robot matches the collected environmental data and task instructions with the predefined scene, and obtains the action instruction sequence associated with the predefined scene, so that the robot can execute the control actions corresponding to the action instructions in sequence according to the action instruction sequence. However, in this case, the action instructions generated by the embodied intelligent robot are deeply coupled with the predefined scene. When the environmental parameters change (such as the position offset of the object) or when performing a new task, due to the mismatch between the action instruction sequence and the real-time environmental state, the robot is prone to abnormal situations such as collisions during movement, which leads to a higher probability of task execution failure.

[0057] The present application fuses environmental perception data including at least one modal input such as visual and tactile data in real time, and performs fusion processing to generate task semantic information associated with task instructions. In the planning stage, implicit action tags corresponding to the task semantic information are formed, and the task to be executed is decomposed layer by layer into multiple task sub-goals. Action instruction sequences are generated and executed, so that the action instruction sequences can flexibly respond to changes in environmental parameters such as object position offset or new task requirements, reduce the coupling degree between the action instructions of the embodied intelligent robot and predefined scenes, and avoid abnormal situations such as collisions caused by mismatch between the instruction sequence and the real-time environmental status, thereby significantly improving the environmental adaptability and task robustness of the embodied robot, ensuring high stability of task execution in dynamic scenes, and reducing the probability of failure.

[0058] To better understand the above technical solutions, exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0059] It should be noted that the execution subject of this embodiment can be the control system of an embodied robot, or a computing service device with data processing, network communication, and program execution capabilities, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the aforementioned functions, or a control device for an embodied intelligent robot, etc. This embodiment does not specifically limit this. The following uses the control system of an embodied robot as an example to illustrate this embodiment and the following embodiments.

[0060] Based on this, the embodiment of the present application provides a control method for an embodied intelligent robot, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the control method of the embodied intelligent robot of the present application.

[0061] In this embodiment, the control method of the embodied intelligent robot includes steps S10 to S40:

[0062] Step S10: receiving a task instruction and environmental perception data, wherein the task instruction is input via language and / or visual modality, and the environmental perception data includes at least one modality input of vision and touch;

[0063] In this embodiment, the user can input task instructions (i.e., instructions for specific actions or tasks that the robot needs to complete) into the embodied robot through either a language modality or a visual modality, such as voice commands or gesture commands. Simultaneously, the embodied robot can obtain information about its surroundings through visual, laser, and / or infrared sensors as environmental perception data. Visual environmental perception data can include images or videos captured by a camera, containing information such as the shape, position, and color of an object. Tactile environmental perception data can include information such as the hardness, texture, and contact force of an object's surface, obtained through a tactile sensor.

[0064] It should be noted that an embodied robot is a robot that has applied embodied intelligence technology. It can interact with the environment in real time through physical entities to achieve a closed loop of perception, cognition, decision-making and action. Its technical system covers multiple fields such as machine vision, natural language understanding, and robotics. Among them, embodied intelligence is a cutting-edge field at the intersection of artificial intelligence and robotics. It emphasizes that intelligent agents achieve autonomous learning and evolution through dynamic interaction between the body and the environment. Its core lies in the deep integration of perception, action and cognition. Optionally, the control system can be installed in the embodied robot to control the robot in real time, or it can be deployed on a cloud server to achieve robot control through remote connection.

[0065] Specifically, the robot first receives the user's voice commands through the voice recognition module. This module uses deep speech recognition technology to convert the user's voice signals into text. Simultaneously, the robot captures images of the environment through a camera and uses computer vision algorithms to preprocess the images and extract key features from the images. For tactile information, the robot perceives the surface characteristics of objects through tactile sensors mounted on its arms or grippers. For example, when receiving a task instruction, the voice recognition module activates, recording the user's voice instructions in real time and converting them into text. It then makes a preliminary judgment of the instruction's intent based on grammatical rules and semantic understanding. Regarding environmental perception, the camera captures images of the environment at a preset frequency. Image processing algorithms perform grayscale conversion, noise removal, and edge detection on each frame to extract basic information such as the object's outline and position. The tactile sensors transmit the electrical signals generated when the robot touches the object in real time to the control unit, where tactile perception data is generated after signal conversion and processing.

[0066] Step S20: fusing the task instruction and the environmental perception data to generate task semantic information associated with the task instruction;

[0067] In this embodiment, fusion processing refers to integrating data from different modalities (such as language, vision, and touch) to obtain more comprehensive and accurate environmental and task information. Task semantic information is the task-related semantic content extracted from the fused data, including key information such as task objectives, operation objects, and operation methods.

[0068] Specifically, the control system inputs the acquired language modality task instructions and the environmental perception data from the visual and tactile modalities into the data fusion module. This data fusion module synchronizes and spatially aligns the data from different modalities to ensure their temporal and spatial consistency. Multimodal fusion algorithms in deep learning, such as fusion networks based on attention mechanisms, are used to extract and fuse features from the data of each modality. For example, when fusing language and visual data, the attention mechanism calculates the correlation weights between language features and visual features, automatically focusing on the visual areas relevant to the task.

[0069] As an optional implementation, the control system uses a convolutional neural network in the visual encoder to extract multi-scale visual features from the environmental perception data. Simultaneously, the attention mechanism in the language model identifies the action objects and constraints in the task instructions, thereby deriving the natural language semantics of the task instructions. The control system then performs cross-modal attention alignment on the multi-scale visual features and the natural language semantics to generate task semantic information that the embodied robot can understand.

[0070] As another optional implementation, a multimodal large model base is provided in the embodied robot, which can process multiple modal inputs such as vision, language, and touch, and generate unified action labels.

[0071] Optionally, the embodied robot aligns information from different modalities based on a unified action tagging scheme. After recognizing the task instructions, the robot combines the action tagging scheme with the task semantics. Alternatively, the embodied robot can directly provide the action tagging scheme based on the unified standard to the implicit planner, which then predicts the implicit task tagging scheme.

[0072] Step S30: parsing the task semantic information through an implicit planner to generate implicit action tags;

[0073] In this embodiment, the implicit planner is a planning module based on machine learning and artificial intelligence algorithms. It generates implicit action tags based on task semantics. These implicit action tags implicitly represent the action sequence and strategy that the robot needs to execute. The implicit action tags contain information such as the action type, sequence, priority, and environmental constraints, and are used to encode them in a compact format that can be understood and executed by the robot.

[0074] It should be noted that the implicit planner is the core intelligent hub between task understanding and action generation. It receives fused task semantic information and models the logical dependencies between task steps through its internal pre-trained graph neural network (GNN). For example, "grasping an object" must follow "moving to the object's location," and "placing an object" must follow "grasping." It then analyzes the execution of tasks using a predefined library of high-level action templates, rather than directly outputting specific, low-level joint motion instructions. Based on this analysis and matching, the implicit planner outputs a structured intermediate representation, the implicit action label. Implicit action labels are abstract, semantic descriptions of action units, and are generalizable, high-level action units. They are used to constrain embodied robot actions without specifically fixing them. During the execution of an action, the robot can automatically adjust action parameters based on real-time environmental changes and within the constraints of the implicit action label, based on its "generalization" nature.

[0075] In one example, implicit action tagging includes three key elements: action mode, execution priority, and environmental constraint parameters. The action mode specifies the core action type to be performed, directly corresponding to the action unit used as an action template in the high-level action template library. The execution priority indicates the relative importance and execution order of the action in the overall task sequence. The environmental constraint parameters define the key conditions or restrictions that must be met to execute the action, such as "target coordinates = (x, y, z)", "maximum grasping force = 5N", and "obstacle avoidance distance = 0.1m". These are derived from task semantic parsing and will be specifically instantiated in subsequent hierarchical decompositions combined with real-time perception data. Based on these environmental constraint parameters, the action constraint information of the embodied robot can be obtained.

[0076] As an optional implementation, the embodied robot's control system inputs task semantic information into a task decomposition model within an implicit planner, then uses a graph neural network to model dependencies within the task semantic information and generate a task decomposition path. The implicit planner converts the task semantic information into a graph structure, where key elements such as action goals, operation objects, and operation methods are used as nodes, and relationships between these elements, such as operation sequence and goal dependencies, are used as edges. The task decomposition model uses a graph neural network to decompose complex tasks into a series of subtasks in a reasonable sequence and structure, forming a task decomposition path. Based on this task decomposition path, the control system matches appropriate action units from a predefined high-level action template library. The high-level action template library contains action patterns and corresponding action constraint information. The control system matches the task decomposition path against the high-level template library step by step, obtaining the action patterns and action constraint information therein and generating corresponding action units. This is achieved by encoding the action units as implicit action tags, which contain at least one of the action pattern, execution priority, and environmental constraint parameters.

[0077] Alternatively, the embodied robot employs a hybrid expert system consisting of an implicit planner and an action expert agent to perform task planning and motion control. The implicit planner is responsible for transferring motion patterns to the robot task to be performed, enabling the prediction of implicit action labels and decomposing complex tasks into generalizable high-level action units. The action expert agent, on the other hand, generates specific action sequences based on the implicit action labels.

[0078] Step S40: Based on the implicit action tag, decompose the to-be-executed task corresponding to the task instruction into multiple task sub-goals layer by layer;

[0079] In this embodiment, after obtaining the implicit action tag based on the parsing result of the task semantic information, the control system can determine the actions to be performed by multiple robots based on information such as the action mode contained in the implicit action tag, thereby forming task sub-goals.

[0080] Specifically, based on the action patterns in the implicit action tags, a hierarchical state machine maps the task to be executed into at least two action hierarchies. Within each action hierarchical level, the embodied robot combines environmental perception data with the action constraints corresponding to the implicit task tags to adjust the action execution parameters of the action hierarchical level, forming the corresponding atomic actions. Based on these atomic actions, multiple task sub-goals can be determined.

[0081] Alternatively, the embodied robot can employ a three-part architecture consisting of a visual encoder, a language model, and an action decoder. The visual encoder processes visual input and extracts environmental information. The language model understands natural language instructions and generates intermediate representations, such as implicit action labels, in conjunction with an implicit planner. The action decoder generates specific action instructions based on these intermediate representations.

[0082] Step S50: generating an action instruction sequence based on the task sub-goal, and executing a control action corresponding to the action instruction sequence.

[0083] In this embodiment, based on the obtained task sub-goals, the control system of the embodied robot generates a specific action instruction sequence based on the execution order of the task sub-goals to control the embodied robot to execute the control action corresponding to the action instruction and complete the task to be executed corresponding to the task instruction.

[0084] As an optional implementation, the embodied robot also includes an action expert agent, which can independently complete decision-making tasks such as action instruction generation based on the agent model in the embodied robot.

[0085] Specifically, the control system receives task subgoals through an expert agent and invokes the inverse kinematics solver in the robot's dynamics model to calculate the target pose for each joint. Based on the spatial distribution of obstacles in the environmental perception data, the expert agent generates an end-effector trajectory using a path planning algorithm. This encodes the target pose and end-effector trajectory in a time-sequential manner into a sequence of action instructions containing joint angles, velocities, and / or accelerations. The control system sequentially executes the embodied robot control actions corresponding to the action instructions in the action instruction sequence.

[0086] As another optional implementation, the control system can also generate specific action instructions based on the embodied robot's action decoder based on intermediate representations such as implicit action tags. Alternatively, the control system can call the action decoder through the action expert agent to complete the compilation of action instructions.

[0087] It should be noted that based on the action constraint parameters of the implicit task label, the control system can dynamically adjust the action parameters of the task instructions in the task instruction sequence within the range of the action constraint parameters based on the environmental parameters at the current moment during the task execution of the embodied robot, thereby improving the flexibility of the embodied robot in task execution.

[0088] For example, the motion model architecture of an embodied intelligent robot is based on a multimodal large model base, a hybrid expert system, and an end-to-end architecture using a three-part "visual encoder + language model + action decoder" architecture. The multimodal large model base can process multimodal inputs such as vision, language, and touch, generating unified action labels. The implicit planner in the hybrid expert system is responsible for predicting implicit action labels, breaking down complex tasks into generalizable high-level action units. The action expert agent then generates specific action sequences based on the implicit action labels.

[0089] The embodiment of the present application fuses environmental perception data including at least one modal input of visual, tactile and other data in real time, and performs fusion processing to generate task semantic information associated with task instructions, forms implicit action tags corresponding to the task semantic information in the planning stage, decomposes the task to be executed layer by layer into multiple task sub-goals, generates and executes action instruction sequences, and enables the embodied robot to flexibly respond to changes in environmental parameters such as object position offset or new task requirements based on the constraints of the implicit action tags and the action instruction sequence, thereby reducing the coupling degree between the action instructions of the embodied intelligent robot and the predefined scenes, avoiding abnormal situations such as collisions caused by mismatch between the instruction sequence and the real-time environmental status, significantly improving the environmental adaptability and task robustness of the embodied robot, ensuring high stability of task execution in dynamic scenes, and reducing the probability of failure.

[0090] Based on the same inventive concept, this application also provides a second embodiment, referring to Figure 2 , Figure 2 This is a flow chart of the second embodiment of the control method of the embodied intelligent robot of the present application.

[0091] In this embodiment, after generating an action instruction sequence based on the task sub-goal as described in step S50 and executing the control action corresponding to the action instruction sequence, steps S61 to S63 are further included:

[0092] Step S61: capturing object posture change data through a visual sensor, and / or collecting contact force feedback data through a tactile sensor;

[0093] Step S62: inputting the object posture change data and / or the contact force feedback data into an implicit planner, and updating the environmental constraint parameters corresponding to the implicit action mark through an optimization algorithm;

[0094] Step S63: executing the replanning action of the task sub-goal according to the updated environmental constraint parameters, and updating the action instruction sequence.

[0095] In this embodiment, while the embodied robot is executing a task, it can use sensors to collect real-time environmental information. These include, but are not limited to, visual sensors capturing object pose change data and / or tactile sensors collecting contact force feedback data. Based on this environmental information, the control system detects command changes, environmental changes, and other information, thereby adjusting the embodied robot's motion instructions and corresponding execution actions in real time.

[0096] In one embodiment, after receiving the object posture change data and / or contact force feedback data, the implicit planner updates the environmental constraint parameters corresponding to the implicit action mark through an optimization algorithm, and executes the re-planning action of the task sub-goal through the action expert executor to update the action instruction sequence.

[0097] In another optional embodiment, the control system can also compile the action instruction sequence based on implicit action tags through the action expert agent and / or action encoder without changing the environmental constraint parameters to update the action parameters corresponding to the action instruction sequence to adjust the action amplitude of the robot's action.

[0098] For example, the embodied robot collects multiple modal inputs, such as vision, speech, and touch, through a multimodal large model base to generate real-time action signatures. The control system uses an implicit planner to determine whether a task change or an update to the implicit action signature is triggered. If not, the control system uses an action expert agent and / or action encoder to compile action instructions to adjust action execution parameters. If so, the control system uses the implicit planner to update the environmental constraint parameters corresponding to the implicit action signature and / or uses the action expert agent to update the action instruction sequence.

[0099] The embodiment of the present application is based on implicit task labeling and further reduces the probability of task execution failure by dynamically adjusting the action instruction sequence, thereby improving the stability and flexibility of the robot operation.

[0100] Since the system described in Example 2 of this application is the system used to implement the method of Example 1 of this application, those skilled in the art will be able to understand the specific structure and variations of the system based on the method described in Example 1 of this application, and therefore, no further description is given here. All systems used in the method of Example 1 of this application fall within the scope of protection to be provided by this application.

[0101] Based on the same inventive concept, this application also provides a third embodiment, referring to Figure 3 , Figure 3 This is a flow chart of the third embodiment of the control method of the embodied intelligent robot of the present application.

[0102] In this embodiment, the control method of the embodied intelligent robot further includes steps S01 to S03:

[0103] Step S01: receiving a training dataset, wherein the training dataset includes task semantic information samples and labeled true values ​​of corresponding implicit actions;

[0104] Step S02: Inputting the task semantic information sample into the graph neural network to obtain the predicted label of the implicit action, and calculating the semantic similarity loss between the predicted label and the true value of the label;

[0105] Step S03: Iterating the graph neural network based on the semantic similarity loss, and obtaining the implicit planner based on the iterated graph neural network.

[0106] In this embodiment, the training dataset is a collection of data used to train a machine learning model. Task semantic information samples and the corresponding implicit action label true values ​​provide the input-output pairs necessary for model learning. Receiving such a training dataset provides the foundational data for training the graph neural network, enabling the model to learn the mapping between task semantic information and implicit action labels.

[0107] Furthermore, the embodied robot inputs task semantic information samples into the graph neural network to obtain predicted labels for implicit actions. By calculating the semantic similarity loss between the predicted labels and the true labels, the model's prediction error can be quantified, providing direction for model optimization. Semantic similarity loss compares the semantic differences between the predicted labels and the true values, ensuring that the model accurately learns the relationship between task semantic information and implicit action labels. The embodied robot trains the graph neural network based on the semantic similarity loss using iterative optimization algorithms such as stochastic gradient descent, gradually adjusting the model's parameters to minimize the difference between the predicted labels and the true labels. This allows the graph neural network to learn the mapping relationship between task semantic information and implicit action labels after multiple iterations, thereby forming an implicit planner.

[0108] Since the system described in Example 3 of this application is the system used to implement the method of Example 1 of this application, those skilled in the art will be able to understand the specific structure and variations of the system based on the method described in Example 1 of this application, and therefore will not be described in detail here. All systems used in the method of Example 1 of this application fall within the scope of protection to be provided by this application.

[0109] Based on the same inventive concept, this application also provides a fourth embodiment, referring to Figure 4 , Figure 4 This is a flow chart of the fourth embodiment of the control method of the embodied intelligent robot of the present application.

[0110] In this embodiment, the control method of the embodied intelligent robot includes steps S04 to S06:

[0111] Step S04: obtaining Internet video data, extracting spatiotemporal features of a target action segment from the Internet video data, and encoding the target action segment into an implicit action pattern vector based on the spatiotemporal features;

[0112] Step S05: Acquire cross-entity operation data of the target robot, and extract action semantics of the cross-entity operation data;

[0113] Step S06: Mapping the action semantics and / or the implicit action pattern vector into action units, and storing the action units into a predefined high-level action template library.

[0114] In this embodiment, the action patterns in Internet videos and cross-ontology operation data are transferred to the robot task through implicit coding.

[0115] It's important to note that internet video data contains a rich collection of human action examples. By analyzing these videos, we can extract the spatiotemporal features of target action segments—that is, the patterns of temporal and spatial variations of the action. Encoding these spatiotemporal features as implicit action pattern vectors converts complex actions into a manageable numerical form, making them easier to store and apply. This provides robots with an additional source of action patterns, expanding their action repertoire and enabling them to perform a wider variety of tasks.

[0116] In one embodiment, video data related to robotic tasks is collected from the internet. The video data is preprocessed, such as by unifying the frame rate and adjusting the resolution. Target action segments, such as "grab an object" or "place an object," are identified. Computer vision techniques, such as optical flow or 3D convolutional neural networks, are used to extract spatiotemporal features of the target action segments. The extracted spatiotemporal features are converted into implicit action pattern vectors using dimensionality reduction and encoding algorithms. These vectors are stored for subsequent mapping into action units.

[0117] In another embodiment, the control system can collect operational data of the target robot in various mission scenarios. The data is preprocessed, such as by removing noise and standardizing the data format. Action semantics are extracted using data mining and machine learning techniques, such as by identifying common action patterns through cluster analysis. The extracted action semantics are encoded and represented for subsequent mapping into action units. The extracted action semantic information is stored and combined with other action data.

[0118] Furthermore, the robot control system can convert the action semantics and implicit action pattern vectors into a unified action unit representation based on preset mapping rules. According to the mapping rules, the extracted action semantics and implicit action pattern vectors are converted into action units, and the properties of the action units, such as action mode, execution priority and environmental constraint parameters, are determined, so that the generated action units are stored in a predefined high-level action template library.

[0119] The embodiments of the present application are based on generalized transferability and can significantly improve the action generation and execution efficiency of embodied intelligent robots. Among them, the training effect and training efficiency of embodied robots are improved through Internet learning or cross-entity operation data.

[0120] Since the system described in Example 4 of this application is the system used to implement the method of Example 1 of this application, those skilled in the art will be able to understand the specific structure and variations of the system based on the method described in Example 1 of this application, and therefore will not be described in detail here. All systems used in the method of Example 1 of this application fall within the scope of protection to be provided by this application.

[0121] The present application provides a control device for an embodied intelligent robot, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the control method for the embodied intelligent robot in the above-mentioned embodiment one.

[0122] Reference below Figure 5 , which shows a schematic diagram of the structure of a control device suitable for implementing an embodied intelligent robot in an embodiment of the present application. The control device of the embodied intelligent robot in an embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The control device of the embodied intelligent robot shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0123] like Figure 5As shown, the control device of the embodied intelligent robot may include a processing device 1001 (e.g., a core processor, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the control device of the embodied intelligent robot are also stored in the random access memory 1004. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the control device of the embodied intelligent robot to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows the control device of the embodied intelligent robot with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems can be implemented or have instead.

[0124] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0125] The control device for an embodied intelligent robot provided in this application utilizes the control method for an embodied intelligent robot described in the aforementioned embodiments, thereby resolving the technical issue of the high probability of task execution failure for an embodied intelligent robot. Compared to the prior art, the beneficial effects of the control device for an embodied intelligent robot provided in this application are the same as those of the control method for an embodied intelligent robot described in the aforementioned embodiments. Other technical features of the control device for an embodied intelligent robot are the same as those disclosed in the aforementioned embodiments and are not further elaborated upon here.

[0126] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0127] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0128] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer program) stored thereon, and the computer-readable program instructions are used to execute the control method of the embodied intelligent robot in the above-mentioned embodiment.

[0129] The computer-readable storage medium provided in this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF, Radio Frequency), etc., or any suitable combination thereof.

[0130] The above-mentioned computer-readable storage medium may be included in the control device of the embodied intelligent robot; or it may exist independently without being assembled into the control device of the embodied intelligent robot.

[0131] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the control device of the embodied intelligent robot, the control device of the embodied intelligent robot: receives task instructions and environmental perception data, wherein the task instructions are input through language and / or visual modalities, and the environmental perception data includes at least one modal input of vision and touch; fuses the task instructions and the environmental perception data to generate task semantic information associated with the task instructions; parses the task semantic information through an implicit planner to generate implicit action tags; based on the implicit action tags, decomposes the task to be executed corresponding to the task instruction into multiple task sub-goals layer by layer; generates an action instruction sequence based on the task sub-goals, and executes the control action corresponding to the action instruction sequence.

[0132] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0134] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0135] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the control method of the embodied intelligent robot described above. This computer-readable storage medium can solve the technical problem of the high probability of task execution failure of the embodied intelligent robot due to the mismatch between the action instruction sequence and the real-time environmental state. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the control method of the embodied intelligent robot provided in the above embodiment, and will not be elaborated here.

[0136] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A control method for an embodied intelligent robot, characterized in that: The method comprises the following steps: receiving a task instruction and environmental perception data, wherein the task instruction is input via language and / or visual modality, and the environmental perception data includes at least one modality input of vision and touch; fusing the task instruction and the environmental perception data to generate task semantic information associated with the task instruction; Parsing the task semantic information through an implicit planner to generate implicit action tags; Based on the implicit action tag, decomposing the to-be-executed task corresponding to the task instruction into multiple task sub-goals layer by layer; An action instruction sequence is generated based on the task sub-goal, and a control action corresponding to the action instruction sequence is executed.

2. The method according to claim 1, wherein The step of fusing the task instruction and the environmental perception data to generate task semantic information associated with the task instruction includes: extracting multi-scale visual features from the environmental perception data through a convolutional neural network in a visual encoder; Identifying action objects and constraints in the task instructions through an attention mechanism in a language model to obtain the natural language semantics of the task instructions; The multi-scale visual features are aligned with the natural language semantics through cross-modal attention to generate the task semantic information.

3. The method according to claim 1 or 2, wherein: The step of parsing the task semantic information by an implicit planner to generate implicit action tags includes: Inputting the task semantic information into the task decomposition model in the implicit planner, modeling the dependency relationship of the task semantic information through a graph neural network, and generating a task decomposition path; According to the task decomposition path, matching adapted action units in a predefined high-level action template library, wherein the high-level action template library includes action patterns and corresponding action constraint information; The action unit is encoded as an implicit action tag, wherein the tag includes at least one of the action mode, execution priority, and environment constraint parameters.

4. The method according to claim 3, wherein The step of decomposing the to-be-executed task corresponding to the task instruction into a plurality of task sub-goals layer by layer based on the implicit action tag further includes: According to the action mode in the implicit action tag, mapping the task to be executed into at least two action levels through a hierarchical state machine; In each of the action levels, combining the environmental perception data and the action constraint information, adjusting the action execution parameters of the action level to form a corresponding atomic action; The task sub-goal is determined according to the atomic action.

5. The method according to claim 1, wherein The step of generating an action instruction sequence based on the task sub-goal and executing a control action corresponding to the action instruction sequence includes: The task sub-goals are received by the action expert agent, and the inverse kinematics solver in the robot dynamics model is called to calculate the target pose of each joint; generating an end-effector trajectory through a path planning algorithm based on the spatial distribution of obstacles in the environmental perception data; Encoding the target posture and the end effector trajectory into the action instruction sequence including joint angles, velocities and / or accelerations in time sequence; The control actions corresponding to the action instructions in the action instruction sequence are executed in sequence.

6. The method according to claim 1, wherein After the steps of generating an action instruction sequence based on the task sub-goal and executing the control action corresponding to the action instruction sequence, the method further includes: Capturing object posture change data through visual sensors, and / or collecting contact force feedback data through tactile sensors; Inputting the object pose change data and / or the contact force feedback data into an implicit planner, and updating the environmental constraint parameters corresponding to the implicit action markers through an optimization algorithm; According to the updated environmental constraint parameters, the replanning action of the task sub-goal is executed, and the action instruction sequence is updated.

7. The method according to claim 1, wherein Before the step of parsing the task semantic information by the implicit planner to generate implicit action tags, the method further includes: Receiving a training dataset, wherein the training dataset includes task semantic information samples and labeled true values ​​of corresponding implicit actions; Inputting the task semantic information sample into a graph neural network to obtain a predicted label of the implicit action, and calculating the semantic similarity loss between the predicted label and the true value of the label; The graph neural network is iterated based on the semantic similarity loss, and the implicit planner is obtained based on the iterated graph neural network.

8. The method according to claim 1, wherein The step of parsing the task semantic information by the implicit planner to generate implicit action tags further includes: Obtaining internet video data, extracting spatiotemporal features of a target action segment from the internet video data, and encoding the target action segment into an implicit action pattern vector based on the spatiotemporal features; Acquiring cross-ontology operation data of the target robot and extracting action semantics of the cross-ontology operation data; The action semantics and / or the implicit action pattern vector are mapped into action units, and the action units are stored in a predefined high-level action template library.

9. A control device for an embodied intelligent robot, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the control method of the embodied intelligent robot according to any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the control method of the embodied intelligent robot according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Natural language driven mechanical arm control method, device, equipment, medium and product

    CN120886271A

  • Robot control method and device and storage medium

    CN120921403A

  • Multi-mode-based body robot control method and device, electronic equipment, readable storage medium and program product

    CN121043159A

  • Multimodal-based embodied robot control method and device, electronic equipment, readable storage medium and program product

    CN121043159B

  • Robot, operation method thereof, operation device, storage medium, and program product

    CN121424349A