Robot control method and system and storage medium

By using target strategies generated through imitation learning and reinforcement learning, combined with multi-dimensional control commands, the operational capabilities and stability of humanoid robots in three-dimensional space are improved. This solves the problems of insufficient flexibility and control precision in existing technologies, and achieves control precision and flexibility in narrow or complex environments.

CN121018567APending Publication Date: 2025-11-28GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511319190.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

In existing technologies, the teleoperation control of humanoid robots in three-dimensional space, especially in narrow or structurally complex environments, lacks flexibility and control precision, making it difficult to meet the needs of high-degree-of-freedom operation.

Method used

By acquiring training instruction sets and reference action datasets, and combining imitation learning and reinforcement learning, target strategies for robots with preset shapes are generated, including multi-dimensional control instructions for upper limbs, lower limbs, base height, and upper body three-dimensional posture, enabling robots to operate with high degrees of freedom in complex environments.

Benefits of technology

It improves the robot's operational capabilities and stability in three-dimensional space, especially in narrow or complex environments, enhances control precision and flexibility, and solves the problems of low control precision and poor flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121018567A_ABST
    Figure CN121018567A_ABST
Patent Text Reader

Abstract

The invention discloses a robot control method and system and a storage medium. The method comprises the steps that a training instruction set and a reference action data set are acquired, and the reference action data set is used for controlling a robot with a preset form to conduct imitation learning on actions of a teaching object; controlling the robot with the preset form to perform imitation learning on the reference action data set, and learning action knowledge from the reference action data set; and controlling the preset form robot to perform reinforcement learning on the training instruction set by using the action knowledge, and generating a target strategy to be used by the preset form robot. The technical problems of low control precision and poor flexibility of a robot control mode in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a robot control method and system and a storage medium. BACKGROUND

[0002] Humanoid robots have unique application potential in various fields due to their close-to-human body structure, especially in teleoperation control, which has become a research focus. With the increasing complexity of use scenarios, robots not only need to have the ability to walk and operate with both arms, but also need to be able to move freely in three-dimensional environments to access more inaccessible locations and complete tasks such as kitchen organization and precision assembly. However, it is not easy to achieve the above high-degree-of-freedom operations, and traditional control methods are often limited by the difficulties of complex modeling and parameter adjustment. In related technologies, the teleoperation control of humanoid robots mainly relies on dual-arm operation on a fixed base or a simple mobile platform. Although methods combining mobile carriers and dual-arm control have been developed subsequently, improving the operating range of robots, the flexibility and operating precision of robots in three-dimensional space, especially in narrow or complex structural environments, still need to be improved. In addition, due to the limitations of the robot body design, such as interference problems of biped walking, picking up objects from the ground is also a challenge. Therefore, although the robot control method in related technologies has expanded the operating range to some extent, it still cannot meet the precise control requirements in three-dimensional environments with rich longitudinal structures and narrow spaces, resulting in poor flexibility and low control precision of robots.

[0003] At present, there is no effective solution to the above problems. SUMMARY

[0004] Embodiments of the present application provide a robot control method, system and storage medium to at least solve the technical problems of low control precision and poor flexibility of the robot control method in related technologies.

[0005] According to an aspect of an embodiment of the present application, a robot control method is provided, comprising: obtaining a training instruction set and a reference action data set, wherein the reference action data set is used to control a preset form robot to mimic learning of a teaching object action; controlling the preset form robot to mimic learning of the reference action data set, learning action knowledge from the reference action data set; and controlling the preset form robot to use the action knowledge to perform reinforcement learning on the training instruction set, generating a target strategy to be used by the preset form robot.

[0006] Optionally, the training instruction set comprises a plurality of dimensional control instructions, wherein the plurality of dimensional control instructions comprise: an upper limb action control instruction of the preset morphology robot; a lower limb movement control instruction of the preset morphology robot; a base height control instruction of the preset morphology robot; and an upper body three-dimensional posture control instruction of the preset morphology robot.

[0007] Optionally, obtaining the reference action data set comprises: obtaining demonstration object action data; and mapping the demonstration object action data to the body configuration of the preset morphology robot to obtain the reference action data set.

[0008] Optionally, mapping the demonstration object action data to the body configuration of the preset morphology robot to obtain the reference action data set comprises: reorienting the demonstration object action data based on the body configuration of the preset morphology robot, converting the demonstration object action data from a key point coordinate system to a robot coordinate system to obtain a conversion result; and performing inverse kinematics processing on the conversion result to obtain the reference action data set.

[0009] Optionally, controlling the preset morphology robot to perform imitation learning on the reference action data set to learn action knowledge from the reference action data set comprises: controlling the preset morphology robot to perform imitation learning on the reference action data set to learn explicit knowledge and implicit knowledge from the reference action data set, wherein the explicit knowledge is used to determine a plurality of joint angles and rotational speeds, a floating baseline and angular velocity, and a related point position of the preset morphology robot, and the implicit knowledge is used to determine an action balance ability, a foot contact, and switching and connection between different actions of the preset morphology robot.

[0010] Optionally, controlling the preset morphology robot to perform reinforcement learning on the training instruction set using the action knowledge to generate a target strategy to be used by the preset morphology robot comprises: controlling the preset morphology robot to perform instruction analysis on the training instruction set to output an inferred action sequence; determining a body state observation result of the preset morphology robot according to the inferred action sequence; and controlling the preset morphology robot to perform reinforcement learning on the training instruction set using the action knowledge based on the body state observation result to generate the target strategy to be used by the preset morphology robot.

[0011] Optionally, controlling the preset morphology robot to perform reinforcement learning on the training instruction set using the action knowledge based on the body state observation result to generate a target strategy to be used by the preset morphology robot comprises: performing comparative learning based on the body state observation result and a reference action sequence in the reference action data set to obtain a learning result; controlling the preset morphology robot to perform instruction following on the training instruction set using the action knowledge based on the body state observation result to obtain a following result; performing strategy evaluation on an original strategy of the preset morphology robot according to the learning result and the following result to obtain an evaluation result; and performing strategy improvement on the original strategy according to the evaluation result to generate the target strategy to be used by the preset morphology robot.

[0012] According to another aspect of the embodiments of this application, a robot control method is also provided, including: acquiring a target control instruction; controlling a preset-form robot to execute the target control instruction according to a target strategy, and obtaining an instruction execution result; wherein the target strategy is generated according to any one of the robot control methods in the embodiments of this application.

[0013] According to another aspect of the embodiments of this application, a robot control system is also provided, including: a memory storing an executable program; and a controller for running the program, wherein the program executes the methods in various embodiments of this application when it runs.

[0014] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0015] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0016] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.

[0017] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.

[0018] In this embodiment, by acquiring a training instruction set and a reference action dataset, the reference action dataset is used to control a preset-form robot to imitate and learn the actions of the taught object. This, in turn, controls the preset-form robot to imitate and learn from the reference action dataset, learning action knowledge. Finally, the preset-form robot uses this action knowledge to perform reinforcement learning on the training instruction set, generating a target strategy for the preset-form robot to use. This not only skips the complex controller modeling stage, reducing development difficulty and time costs, but also ensures high fidelity and reliability of the robot's actions. By integrating imitation learning and reinforcement learning, control optimization under high degrees of freedom and complex instructions is achieved for the preset-form robot, thereby enhancing the robot's operational capabilities and stability in three-dimensional space, especially in narrow or complex environments. This further improves the robot's control accuracy and flexibility, thus solving the technical problems of low control accuracy and poor flexibility in related robot control methods. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0020] Figure 1 This is a flowchart of a robot control method according to an embodiment of this application;

[0021] Figure 2 This is a flowchart of another robot control method according to an embodiment of this application;

[0022] Figure 3 This is a schematic diagram of a robot control method according to an embodiment of this application;

[0023] Figure 4 This is a schematic diagram of another robot control method according to an embodiment of this application;

[0024] Figure 5 This is a structural block diagram of a robot control device according to an embodiment of this application;

[0025] Figure 6 This is a structural block diagram of another robot control device according to an embodiment of this application. Detailed Implementation

[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] According to an embodiment of this application, a method embodiment for robot control is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0029] Figure 1 This is a flowchart of a robot control method according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:

[0030] Step S11: Obtain the training instruction set and the reference action dataset, wherein the reference action dataset is used to control the preset shape robot to imitate and learn the actions of the teaching object.

[0031] Step S12: Control the robot in the preset shape to imitate and learn from the reference action dataset, and learn action knowledge from the reference action dataset;

[0032] Step S13: Control the preset shape robot to perform reinforcement learning on the training instruction set using motion knowledge, and generate the target strategy to be used by the preset shape robot.

[0033] The aforementioned training instruction set can be a series of instructions designed by the operator based on specific task scenarios, aiming to guide a robot in a preset form to perform specific actions or tasks, such as walking or picking up objects. The training instruction set not only includes movement instructions for the two arms, but also involves the control of base height adjustment, upper body posture changes, and robot walking speed, comprehensively covering the multi-degree-of-freedom control requirements of the robot in the preset form.

[0034] The aforementioned reference motion dataset can be a dataset composed of human-taught actions, used for imitation learning of robots with pre-defined shapes. After retargeting and inverse kinematics processing, the reference motion dataset is transformed into information such as joint angles, velocities, and linear and angular velocities of the floating base of the robot with the pre-defined shape, providing a motion template with human dynamic characteristics for imitation learning. The introduction of the reference motion dataset significantly improves the realism and naturalness of the robot's movements with the pre-defined shape, avoiding the reliance on complex controller modeling in traditional control methods.

[0035] During the imitation learning process, by allowing the pre-formed robot to learn human movements from a reference action dataset and imitate human action patterns, the pre-formed robot can master explicit and implicit knowledge of movements, including the specific angles and speeds of joints, as well as advanced skills such as how to adjust foot contact to maintain balance, thereby achieving the smoothness and stability of robot movements.

[0036] During the imitation learning phase, a pre-defined robot can learn and internalize a series of movement rules and execution principles, thus acquiring motion knowledge. Motion knowledge includes, but is not limited to, how to coordinate joints to maintain balance, how to effectively adjust body posture to adapt to different tasks, and how to smoothly transition between walking and maneuvering. The accumulation of motion knowledge is the foundation for a pre-defined robot to achieve high degrees of freedom in its movements.

[0037] Reinforcement learning is a process in which a robot with a pre-defined form, after acquiring basic motion knowledge, optimizes its motion execution strategy based on a training instruction set through interaction with the environment. The goal of reinforcement learning is to enable the robot with a pre-defined form to accurately execute operator commands based on its motion knowledge, while also considering environmental factors and the continuity of the motion. Through reinforcement learning, the robot with a pre-defined form makes more intelligent and adaptive decisions when performing tasks, effectively improving the accuracy and efficiency of its movements.

[0038] The aforementioned target strategy can be a detailed motion plan for a pre-defined robot to execute a training instruction set, optimized through reinforcement learning. This target strategy transforms the motion knowledge acquired by the pre-defined robot during the imitation learning phase, combined with the optimization results of reinforcement learning, into a motion execution plan for actual operation. This enables the pre-defined robot to flexibly and accurately complete tasks in complex and ever-changing environments, based on the operator's intentions.

[0039] Based on steps S11 to S13 above, by acquiring a training instruction set and a reference action dataset, the reference action dataset is used to control the preset-shape robot to imitate and learn the actions of the taught object. This, in turn, controls the preset-shape robot to imitate and learn from the reference action dataset, learning action knowledge. Finally, the preset-shape robot uses this action knowledge to perform reinforcement learning on the training instruction set, generating the target strategy to be used by the preset-shape robot. This not only skips the complex controller modeling stage, reducing development difficulty and time costs, but also ensures high fidelity and reliability of the robot's actions. By integrating imitation learning and reinforcement learning, control optimization under high degrees of freedom and complex instructions is achieved for the preset-shape robot, thereby enhancing the robot's operational capabilities and stability in three-dimensional space, especially in narrow or complex environments. This further improves the robot's control accuracy and flexibility, thus solving the technical problems of low control accuracy and poor flexibility in related robot control methods.

[0040] The robot control method in the embodiments of this application will be further described below.

[0041] In one optional embodiment, the training instruction set includes: multi-dimensional control instructions, wherein the multi-dimensional control instructions include: upper limb movement control instructions for the preset shape robot; lower limb movement control instructions for the preset shape robot; base height control instructions for the preset shape robot; and upper body three-dimensional posture control instructions for the preset shape robot.

[0042] In this embodiment, the training instruction set is a comprehensive set of instructions used to guide a robot with a preset form to perform dynamic learning and operation with multiple degrees of freedom, aiming to achieve the requirement of freely performing tasks in complex environments. The training instruction set is subdivided into several key parts, each targeting different moving parts or functions of the robot, reflecting the meticulous and comprehensive control of the robot as a whole.

[0043] The upper limb motion control commands for the pre-form robot focus on controlling the robot's arms, including the independent and coordinated movements of joints such as fingers, wrists, elbows, and shoulders. These commands cover delicate movements such as arm extension, grasping, and rotation, providing the robot with human-like arm manipulation capabilities, enabling it to perform fine tasks such as grasping and placing objects in teleoperation mode.

[0044] The lower limb movement control commands for the pre-form robot are used to control its walking and gait adjustment, including setting parameters such as foot lifting, landing timing, and stride length. These commands ensure stable movement of the pre-form robot on different terrains, enabling it to perform basic and advanced movement skills such as walking straight, turning, and going up and down slopes, providing it with a wide range of motion and adaptability.

[0045] The base height control command for the preset-form robot allows the operator to adjust the height of the robot's base, which is crucial for the preset-form robot to operate on workbenches at different heights or to pick up items from the ground. Dynamic adjustment of the base height not only increases the operational flexibility of the preset-form robot but also avoids limitations on certain tasks caused by a fixed height, such as picking up items from the floor.

[0046] The three-dimensional posture control commands for the upper body of the preset-shape robot are used to control the posture of the upper body in three-dimensional space, including the forward and backward tilting, left and right swaying, and rotation of the head and torso. These upper-body three-dimensional posture control commands enable the preset-shape robot to adapt to changing environmental requirements, such as avoiding obstacles by moving sideways in narrow spaces or bending its body to pick up objects at low altitudes, greatly expanding its operational range and efficiency.

[0047] Based on the above optional embodiments, the training instruction set realizes all-round remote operation control of the preset shape robot by including control instructions in multiple dimensions such as upper limbs, lower limbs, base height and upper body three-dimensional posture. It fully covers the motion requirements that the preset shape robot may encounter in the three-dimensional environment. The combination of control instructions in multiple dimensions greatly improves the flexibility and adaptability of the preset shape robot in the remote operation mode, enabling it to perform highly free and precise tasks in complex environments.

[0048] In an optional embodiment, step S11, obtaining the reference action dataset includes:

[0049] Step S111: Obtain the action data of the teaching object;

[0050] Step S112: Map the motion data of the teaching object to the body configuration of the preset shape robot to obtain a reference motion dataset.

[0051] When acquiring motion data of the teachable object, sensors can capture or record the raw motion data of the human teacher, such as the position, angle, and speed of skeletal joints. This teachable object motion data serves as a learning template for the pre-defined robot form, containing rich kinematic information and human dynamics characteristics. In specific application scenarios, such as factory training or teaching home service robots, operators demonstrate the tasks to be performed to the robot through natural body movements, such as picking up tools, opening and closing doors, and cleaning tabletops. The acquisition of teachable object motion data typically relies on motion capture equipment, such as wearable sensor suits or depth cameras, to track the movement trajectory of human joints with high precision, ensuring the accuracy of subsequent learning.

[0052] Furthermore, the motion data of the taught object is mapped to the body configuration of the preset-form robot to obtain a reference motion dataset. Specifically, this involves converting human-taught actions into a format suitable for execution by the preset-form robot. Due to the differences in body structure between humans and the preset-form robot, the motion data of the taught object cannot be directly applied to the preset-form robot. In the mapping process, key points on the human skeletal model are first redirected to the positions of corresponding parts on the preset-form robot. Then, using inverse kinematics algorithms, the positions of the key points are converted into angle parameters of the joints of the preset-form robot, forming a sequence of motion instructions that the preset-form robot can understand and execute. This provides the preset-form robot with the basic data for learning, namely the reference motion dataset. The reference motion dataset contains the converted robot motion information, ensuring that the preset-form robot can correctly execute every detail of human teaching during the imitation learning process, while also taking into account the limitations and dynamic characteristics of the preset-form robot's body structure.

[0053] Based on the above optional embodiments, by acquiring the action data of the teaching object and then mapping the action data of the teaching object to the body configuration of the preset form robot, a reference action dataset is obtained. The preset form robot can quickly learn a series of complex and high-precision human actions, thereby exhibiting more natural and smooth behavior when performing human-like operation tasks, greatly improving the flexibility of the preset form robot in remote operation control mode.

[0054] In an optional embodiment, in step S112, the teaching object's motion data is mapped to the body configuration of a preset-shape robot to obtain a reference motion dataset, which includes:

[0055] Step S1121: Based on the body configuration of the preset shape robot, redirect the motion data of the teaching object, transform the motion data of the teaching object from the key point coordinate system to the robot coordinate system, and obtain the transformation result;

[0056] Step S1121: Perform inverse kinematics processing on the conversion result to obtain the reference motion dataset.

[0057] The process of redirecting the motion data of a teachable object based on the body configuration of a pre-defined robot aims to adjust the joint positions and movements of the teachable object to match the physical structure and range of motion of the pre-defined robot. The teachable object's motion data is initially collected in the teachable object's coordinate system, but the pre-defined robot has its own corresponding coordinate system based on its specific joint structure and motion patterns. Therefore, redirection transforms the teachable object's joint position information from its keypoint coordinate system to the pre-defined robot's coordinate system. This transformation ensures that the pre-defined robot can understand and simulate human movements, even when the pre-defined robot's body configuration is not entirely identical to that of a human.

[0058] In practical applications, such as teaching a pre-defined robot how to pick up a cup from a table, a human instructor would naturally bend their arm, palm, and fingers to grasp the cup. Through redirection, the angles of the pre-defined robot's arm and finger joints are adjusted to closely approximate the human instructor's movements, while taking into account the physical limitations of the robot's arm and the range of motion of its joints, to generate the transformation result.

[0059] Furthermore, the transformation result is processed using inverse kinematics to obtain a reference motion dataset. Inverse kinematics (IK) is a mathematical method used to calculate the angles that the joints of a robot with a preset shape need to reach so that its end effector or other key points can reach the specified target position. Inverse kinematics processing is applied to the transformation result, i.e., the motion data of the taught object mapped to the coordinate system of the robot with the preset shape. Using the inverse kinematics algorithm, the angle and velocity of each joint of the robot with the preset shape can be calculated to accurately execute the taught motion.

[0060] For example, a robot in a preset form needs to learn how to bend over and pick up a heavy object from the ground. After retargeting, the positional information of the robot's key points has been transformed into the robot's coordinate system. Inverse kinematics processing will calculate the precise angles of the waist, leg, and arm joints of the robot in the preset form, ensuring that the robot can bend over smoothly like a human instructor while maintaining balance and avoiding falls or damage.

[0061] Based on the above optional embodiments, the motion data of the teaching object is redirected by the body configuration of the robot with a preset shape, and the motion data of the teaching object is transformed from the key point coordinate system to the robot coordinate system to obtain the transformation result. Then, the transformation result is processed by inverse kinematics to obtain a reference motion dataset, which improves the anthropomorphism and stability of the motion in the remote operation control mode and lays a solid foundation for subsequent reinforcement learning and task execution.

[0062] In an optional embodiment, in step S12, the robot with a preset shape is controlled to perform imitation learning on a reference action dataset, and the action knowledge learned from the reference action dataset includes:

[0063] The robot with a preset shape learns by imitating a reference action dataset. It learns explicit and implicit knowledge from the reference action dataset. The explicit knowledge is used to determine the joint angles and rotation speeds, floating baselines and angular velocities, and associated point positions of the robot with the preset shape. The implicit knowledge is used to determine the robot's motion balance, foot contact, and the switching and connection between different actions.

[0064] Imitation learning is a machine learning technique that enhances the ability of pre-defined robots to perform tasks by observing and imitating reference action datasets. It is suitable for learning complex, unstructured movements, such as natural human walking or fine hand manipulation, without explicit programming or the creation of complex mathematical models. For pre-defined robots, imitation learning allows them to directly learn from human actions, enabling them to master multi-degree-of-freedom movements.

[0065] Explicit knowledge refers to the specific information that a pre-defined robot can directly obtain from a reference action dataset during the imitation learning process. This includes, but is not limited to, the angles and rotational speeds of each joint of the pre-defined robot, the linear and angular velocities of the floating baseline, and the positions of associated points directly related to the action execution. For example, when learning to walk, explicit knowledge would include the flexion and extension angles of the leg joints, the swing amplitude of the arms, and the trajectory of the body's center of gravity. The acquisition of explicit knowledge provides the pre-defined robot with direct guidance to perform human-like actions, contributing to the high fidelity of the movements.

[0066] Compared to explicit knowledge, implicit knowledge is more difficult to observe and quantify directly, but it plays a crucial role in the feasibility and coherence of actions. Implicit knowledge involves how a pre-defined robot maintains balance, adjusts foot contact to adapt to changes in the ground, and smoothly switches between different actions, such as from standing to walking to squatting. Mastering implicit knowledge enables a pre-defined robot to autonomously adjust its actions according to the environment and task requirements, ensuring stability and coherence in complex and changing situations, and achieving natural and smooth movements.

[0067] Based on the above optional embodiments, by controlling a robot with a preset shape to imitate and learn from a reference action dataset, explicit and implicit knowledge can be learned from the reference action dataset. This enables the robot with the preset shape to maintain the continuity and stability of its movements in complex environments. Even when faced with challenges such as uneven ground or narrow spaces, it can respond flexibly and further improve the naturalness and efficiency of its operation.

[0068] In an optional embodiment, in step S13, controlling the preset shape robot to perform reinforcement learning on the training instruction set using action knowledge to generate a target strategy to be used by the preset shape robot includes:

[0069] Step S131: Control the robot in the preset form to perform instruction analysis on the training instruction set and output the reasoning action sequence;

[0070] Step S132: Determine the body state observation results of the robot with the preset shape based on the reasoning action sequence;

[0071] Step S133: Based on the ontological state observation results, control the preset shape robot to perform reinforcement learning on the training instruction set using action knowledge, and generate the target strategy to be used by the preset shape robot.

[0072] During the process of controlling the preset-form robot to analyze the training instruction set, the preset-form robot receives the training instruction set, which contains a series of actions that the operator wants the robot to perform, such as walking, bending over, and raising its hand. Through instruction analysis, the preset-form robot decomposes high-level instructions into executable sequences, i.e., inferred action sequences. These inferred action sequences contain detailed joint angles, speeds, and other necessary control information to guide the preset-form robot to accurately execute the instructions. For example, upon receiving the instruction "walk forward and pick up an object on the ground," the preset-form robot will analyze the specific walking gait, bending angle, and hand grasping motion, ensuring that each action is smoothly connected.

[0073] Based on the inferred action sequence, the pre-defined robot will execute corresponding actions while monitoring its own state through built-in sensors, including information such as joint angles, speed, torque, and overall body posture and displacement. These monitoring results constitute the body state observation results, which are key data for evaluating whether the robot's actions meet expectations and are stable. For example, when the robot's hand grasps a heavy object in the pre-defined form, the sensors will monitor the force on the hand joints and the actual position of the object, thereby evaluating the accuracy and safety of the action execution.

[0074] The pre-defined robot continuously adjusts and optimizes its action execution strategy based on its own state observations. Reinforcement learning, a machine learning method, uses reward and punishment mechanisms to teach the robot to make the best choice in a specific situation to maximize rewards. In this process, the robot applies its previously learned action knowledge, trying different action combinations until it finds a target strategy that accurately executes instructions while maintaining its own stability. The final generated target strategy represents the optimal action plan for the robot to perform similar tasks, accurately following the operator's instructions while fully considering the robot's own dynamics and environmental factors, ensuring both safety and efficiency.

[0075] Based on the above optional embodiments, by controlling a preset-form robot to analyze the training instruction set and output a reasoning action sequence, the body state observation results of the preset-form robot are determined based on the reasoning action sequence. Finally, based on the body state observation results, the preset-form robot is controlled to use action knowledge to perform reinforcement learning on the training instruction set to generate a target strategy to be used by the preset-form robot. This not only significantly improves the preset-form robot's ability to perform high-degree-of-freedom teleoperation tasks, but also greatly reduces the need for human intervention through automated adjustment and optimization of strategies, enhancing the adaptability and intelligence level of the preset-form robot in complex environments.

[0076] In an optional embodiment, in step S133, based on the ontological state observation results, the preset shape robot is controlled to perform reinforcement learning on the training instruction set using action knowledge to generate the target strategy to be used by the preset shape robot, including:

[0077] Step S1331: Based on the ontology state observation results and the reference action sequences in the reference action dataset, perform comparative learning to obtain the learning results;

[0078] Step S1332: Based on the ontological state observation results, control the robot in the preset shape to follow the training instruction set using motion knowledge to obtain the following results;

[0079] Step S1333: Based on the learning results and following results, evaluate the original strategy of the robot with the preset shape to obtain the evaluation results.

[0080] Step S1334: Improve the original strategy according to the evaluation results and generate a target strategy for the robot with a preset shape to use.

[0081] By comparing the ontological state observation results with the reference action sequences in the reference action dataset, the learning results are obtained. The real-time state of the preset-shape robot when performing an action, i.e., the ontological state observation results, is compared with the ideal action sequences stored in the reference action dataset. Contrastive learning aims to improve the accuracy and stability of the preset-shape robot's actions by analyzing the differences between the current action of the preset-shape robot and the standard action, identifying and correcting deviations. For example, if the preset-shape robot exhibits slight wobbling when attempting to replicate a human-taught bending action, contrastive learning will identify this difference and adjust the joint rotation speeds and torque distribution during the execution of the preset-shape robot's action, making the next attempt closer to the standard action, thus optimizing the learning results.

[0082] Based on ontological state observations, the robot in a preset form uses motion knowledge to follow the training instruction set. Utilizing motion knowledge acquired through imitation learning and combined with current ontological state observations, the robot precisely tracks and executes each instruction in the training instruction set. This instruction following process emphasizes the robot's immediate responsiveness to operator commands and the accuracy of its motion execution, ensuring that the robot can smoothly complete the predetermined task without unnecessary movements or deviations from the expected path. The final output is the following result for each instruction.

[0083] After comparing learning and instruction following, the effectiveness of the robot's current original strategy is evaluated based on the learning and following results. The evaluation process not only focuses on the accuracy of the robot's movements but also examines its stability, safety, and energy consumption during execution, comprehensively determining whether the original strategy needs adjustment. The evaluation results will clearly indicate which aspects performed well and which need improvement, providing direction for further strategy optimization.

[0084] Based on the evaluation results, the original strategy is fine-tuned and optimized. This may involve adjusting the order of joint motion sequences, adding or removing certain actions, and adjusting the speed and force of the actions. Through continuous iteration, the control strategy of the preset-shape robot is gradually adjusted, and the generated target strategy can ensure that the preset-shape robot is both accurate and stable when performing high-degree-of-freedom teleoperation tasks, further improving the success rate and efficiency of the tasks.

[0085] Based on the above optional embodiments, the learning results are obtained by comparing the ontology state observation results with the reference action sequences in the reference action dataset. Then, based on the ontology state observation results, the preset shape robot is controlled to follow the training instruction set using action knowledge to obtain the following results. Subsequently, based on the learning results and the following results, the original strategy of the preset shape robot is evaluated to obtain the evaluation results. Finally, the original strategy is improved according to the evaluation results to generate the target strategy to be used by the preset shape robot. It can autonomously adjust the action strategy according to the real-time observation results and task requirements, achieving more natural, efficient and stable action execution.

[0086] Figure 2 This is a flowchart of another robot control method according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:

[0087] Step S21: Obtain the target control command;

[0088] Step S22: Control the preset form machine to execute the target control command according to the target strategy and obtain the command execution result; wherein, the target strategy is generated according to any one of the robot control methods in the embodiments of this application.

[0089] Based on the above optional embodiments, by acquiring target control instructions, the preset-form robot is controlled to execute the target control instructions according to the target strategy, and the execution result is obtained. In the process of generating the target strategy, a training instruction set and a reference action dataset are acquired. The reference action dataset is used to control the preset-form robot to imitate and learn the actions of the taught object, and then to control the preset-form robot to imitate and learn the reference action dataset, learning action knowledge from the reference action dataset. Finally, the preset-form robot uses the action knowledge to perform reinforcement learning on the training instruction set to generate the target strategy to be used by the preset-form robot. This not only skips the complex controller modeling stage, reducing development difficulty and time cost, but also ensures the high fidelity and reliability of robot actions. By integrating imitation learning and reinforcement learning, control optimization under high degrees of freedom and complex instructions is achieved for the preset-form robot, thereby enhancing the robot's operational ability and stability in three-dimensional space, especially in narrow or complex environments, further improving the robot's control accuracy and flexibility, and thus solving the technical problems of low control accuracy and poor flexibility in robot control methods in related technologies.

[0090] Figure 3 This is a schematic diagram of a robot control method according to an embodiment of this application, such as... Figure 3 As shown, the motion data of the teaching object is acquired. Based on the body configuration of the robot with a preset shape, the motion data of the teaching object is redirected, transforming the motion data from the key point coordinate system to the robot coordinate system, and obtaining the transformation result. Then, the transformation result is processed by inverse kinematics to obtain a reference motion dataset. The robot with the preset shape is controlled to imitate and learn from the reference motion dataset, learning explicit knowledge and implicit knowledge. Explicit knowledge is used to determine the joint angles and rotational speeds, floating baselines and angular velocities, and associated point positions of the robot with the preset shape. Implicit knowledge is used to determine the robot's motion balance ability, foot contact, and the switching and connection between different actions.

[0091] Figure 4 This is a schematic diagram of another robot control method according to an embodiment of this application, such as... Figure 4 As shown, the robot in a preset form analyzes the training instruction set and outputs a reasoning action sequence. Based on this sequence, the robot's ontological state observation results are determined. The ontological state observation results are compared with reference action sequences in the reference action dataset to obtain learning results. Based on these results, the robot uses action knowledge to follow the training instruction set, obtaining following results. Then, based on the learning and following results, the robot's original strategy is evaluated to obtain evaluation results. Finally, the original strategy is improved according to the evaluation results to generate the target strategy for the robot to use.

[0092] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0093] According to an embodiment of this application, an apparatus embodiment for a robot control method is provided. It should be noted that the apparatus can be used to execute the above-described robot control method.

[0094] Figure 5 This is a structural block diagram of a robot control device according to an embodiment of this application, such as... Figure 5 As shown, the device includes:

[0095] The acquisition module 501 is used to acquire the training instruction set and the reference action dataset, wherein the reference action dataset is used to control the preset shape robot to imitate and learn the actions of the teaching object.

[0096] The learning module 502 is used to control the robot with the preset shape to imitate and learn from the reference action dataset, and learn action knowledge from the reference action dataset.

[0097] The generation module 503 is used to control the preset shape robot to perform reinforcement learning on the training instruction set using motion knowledge, and generate the target strategy to be used by the preset shape robot.

[0098] Optionally, the training instruction set includes: multi-dimensional control instructions, wherein the multi-dimensional control instructions include: upper limb movement control instructions for the robot with a preset shape; lower limb movement control instructions for the robot with a preset shape; base height control instructions for the robot with a preset shape; and upper body three-dimensional posture control instructions for the robot with a preset shape.

[0099] Optionally, the acquisition module 501 is further configured to: acquire the action data of the teaching object; map the action data of the teaching object to the body configuration of the preset shape robot to obtain a reference action dataset.

[0100] Optionally, the acquisition module 501 is further configured to: redirect the teaching object motion data based on the body configuration of the preset morphological robot, transform the teaching object motion data from the key point coordinate system to the robot coordinate system, and obtain the transformation result; perform inverse kinematics processing on the transformation result to obtain a reference motion dataset.

[0101] Optionally, the learning module 502 is also used to: control the preset shape robot to imitate and learn from the reference action dataset, and learn explicit knowledge and implicit knowledge from the reference action dataset. The explicit knowledge is used to determine the joint angles and rotation speeds, floating baselines and angular velocities, and associated point positions of the preset shape robot. The implicit knowledge is used to determine the motion balance ability, foot contact, and switching and connection between different actions of the preset shape robot.

[0102] Optionally, the generation module 503 is further configured to: control the preset shape robot to perform instruction analysis on the training instruction set and output a reasoning action sequence; determine the ontological state observation results of the preset shape robot based on the reasoning action sequence; and, based on the ontological state observation results, control the preset shape robot to perform reinforcement learning on the training instruction set using action knowledge to generate a target strategy to be used by the preset shape robot.

[0103] Optionally, the generation module 503 is further configured to: perform comparative learning based on the ontology state observation results and the reference action sequences in the reference action dataset to obtain learning results; control the preset shape robot to follow the training instruction set using action knowledge based on the ontology state observation results to obtain following results; evaluate the original strategy of the preset shape robot based on the learning results and following results to obtain evaluation results; improve the original strategy according to the evaluation results to generate the target strategy to be used by the preset shape robot.

[0104] Figure 6 This is a structural block diagram of another robot control device according to an embodiment of this application, such as... Figure 6 As shown, the device includes:

[0105] Acquisition module 601 is used to acquire target control commands;

[0106] The execution module 602 is used to control the preset form machine to execute the target control instruction according to the target strategy and obtain the instruction execution result; wherein, the target strategy is generated according to any one of the robot control methods in the embodiments of this application.

[0107] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0108] Embodiments of this application also provide a robot control system, including: a memory storing an executable program; and a controller for running the program, wherein the program executes the methods described in various embodiments of this application during runtime.

[0109] Optionally, in this embodiment, the controller can be configured to perform the following steps via a computer program:

[0110] S1, Obtain the training instruction set and reference action dataset, wherein the reference action dataset is used to control the robot with the preset shape to imitate and learn the actions of the teaching object;

[0111] S2, control the robot in the preset shape to imitate and learn from the reference action dataset, and learn action knowledge from the reference action dataset;

[0112] S3 controls the preset-form robot to use motion knowledge to perform reinforcement learning on the training instruction set, and generates the target strategy to be used by the preset-form robot.

[0113] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0114] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:

[0115] S1, Obtain the training instruction set and reference action dataset, wherein the reference action dataset is used to control the robot with the preset shape to imitate and learn the actions of the teaching object;

[0116] S2, control the robot in the preset shape to imitate and learn from the reference action dataset, and learn action knowledge from the reference action dataset;

[0117] S3 controls the preset-form robot to use motion knowledge to perform reinforcement learning on the training instruction set, and generates the target strategy to be used by the preset-form robot.

[0118] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0119] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.

[0120] Embodiments of this application also provide a computer program that, when executed by a processor, implements the methods described in the various embodiments of this application.

[0121] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0122] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0123] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0124] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0125] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0126] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A robot control method, characterized in that, include: Acquire a training instruction set and a reference action dataset, wherein the reference action dataset is used to control a robot with a preset shape to imitate and learn the actions of the teaching object; The robot in the preset form is controlled to imitate and learn from the reference action dataset, and learn action knowledge from the reference action dataset; The robot in the preset form is controlled to perform reinforcement learning on the training instruction set using the action knowledge, thereby generating a target strategy to be used by the robot in the preset form.

2. The robot control method according to claim 1, characterized in that, The training instruction set includes: control instructions in multiple dimensions, wherein the control instructions in multiple dimensions include: The upper limb movement control commands of the robot in the preset form; The lower limb movement control commands of the robot in the preset shape; The base height control command for the robot in the preset shape; The three-dimensional posture control commands for the upper body of the robot in the preset form.

3. The robot control method according to claim 1, characterized in that, Obtaining the reference action dataset includes: Obtain the action data of the teaching object; The motion data of the teaching object is mapped to the body configuration of the preset-form robot to obtain the reference motion dataset.

4. The robot control method according to claim 3, characterized in that, The reference motion dataset is obtained by mapping the motion data of the teaching object to the body configuration of the preset-shape robot, including: Based on the body configuration of the robot with the preset shape, the motion data of the teaching object is redirected, and the motion data of the teaching object is transformed from the key point coordinate system to the robot coordinate system to obtain the transformation result; The conversion result is subjected to inverse kinematics processing to obtain the reference motion dataset.

5. The robot control method according to claim 3, characterized in that, Controlling the preset-shape robot to imitate and learn from the reference action dataset, and learning the action knowledge from the reference action dataset includes: The robot in the preset form is controlled to learn by imitation from the reference action dataset, and learn explicit knowledge and implicit knowledge from the reference action dataset. The explicit knowledge is used to determine the joint angles and rotation speeds, floating baselines and angular velocities, and associated point positions of the robot in the preset form. The implicit knowledge is used to determine the robot's motion balance ability, foot contact, and the switching and connection between different actions.

6. The robot control method according to claim 1, characterized in that, Controlling the preset-form robot to perform reinforcement learning on the training instruction set using the action knowledge, and generating the target strategy to be used by the preset-form robot includes: The robot in the preset form is controlled to perform instruction analysis on the training instruction set and output a sequence of reasoning actions. The body state observation results of the preset form robot are determined based on the reasoning action sequence; Based on the observed state of the body, the robot in the preset form is controlled to perform reinforcement learning on the training instruction set using the action knowledge, thereby generating the target strategy to be used by the robot in the preset form.

7. The robot control method according to claim 6, characterized in that, Based on the observed ontological state, the preset-form robot is controlled to perform reinforcement learning on the training instruction set using the action knowledge, generating the target strategy to be used by the preset-form robot, including: The learning results are obtained by comparing the ontology state observation results with the reference action sequences in the reference action dataset. Based on the observed body state, the robot in the preset form is controlled to follow the training instruction set using the motion knowledge to obtain the following result. Based on the learning results and the following results, the original strategy of the robot with the preset shape is evaluated to obtain the evaluation result; Based on the evaluation results, the original strategy is improved to generate the target strategy to be used by the preset-form robot.

8. A robot control method, characterized in that, include: Obtain target control commands; The machine in the preset form is controlled to execute the target control command according to the target strategy, and the command execution result is obtained; The target strategy is generated according to the robot control method described in any one of claims 1 to 7.

9. A robot control system, characterized in that, include: Memory, which stores executable programs; A controller for running the program, wherein the program executes the robot control method according to any one of claims 1 to 8 when it runs.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the robot control method according to any one of claims 1 to 8.