A humanoid robot object grasping method and device based on reinforcement learning control

By using a policy model trained with reinforcement learning algorithms and combining rewards for stable standing and grasping target poses, the stability and posture problems of humanoid robots during the grasping process are solved, thereby improving grasping efficiency and stability.

CN117961888BActive Publication Date: 2026-08-25ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410029777.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2026-08-25
Estimated Expiration
2044-01-08

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the stability and grasping posture issues of humanoid robots during grasping processes. In particular, robots with unstable bases are prone to tipping over when performing grasping tasks, leading to grasping failures.

Method used

A reinforcement learning algorithm is used to train the policy model. By introducing a first reward for the robot to stand stably and a second reward for grasping the target pose, an end-to-end policy model is trained to control the movement of each joint of the robot to achieve stable grasping.

Benefits of technology

It simplifies the planning process for grasping tasks, improves the robot's real-time grasping efficiency and stability, and avoids the problem of humanoid robots tipping over during the grasping process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117961888B_ABST
    Figure CN117961888B_ABST
Patent Text Reader

Abstract

The specification discloses a humanoid robot object grasping method and device based on reinforcement learning control. Through the training mode of reinforcement learning, a first reward for stabilizing the robot standing and a second reward for the robot grasping the object to reach a target pose are introduced in the reward, an end-to-end strategy model is trained, and the adaptive task requirements of robot standing balance and the grasping task requirements are realized at the same time. In the actual execution process of the grasping task, the method of directly controlling the trained strategy model replaces the steps of grasping point detection and whole body motion coordination planning in the conventional grasping method, simplifies the planning process of the grasping task, shortens the planning time, and improves the implementation efficiency of the real-time grasping of the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of robotics, and in particular to a method and apparatus for grasping objects using a humanoid robot based on reinforcement learning control. Background Technology

[0002] Recently, with the rapid development of information technology and the widespread application of intelligent hardware and automated systems, significant breakthroughs have been achieved in robotics technology after years of research both domestically and internationally. Today, robots are widely used in various fields of industry and daily life services. Among these, the task of grasping objects is one of the most fundamental and important aspects of robot operation skills; therefore, how to achieve robot grasping has become one of the most pressing problems to be solved.

[0003] In most robotic grasping solutions, the robot's base is fixed, such as serial palletizing robotic arms and parallel workpiece picking robotic arms. The fixed base ensures the robot's stability, so only the grasping posture of the actuator end effector needs to be monitored. The same applies to mobile robotic arm platforms; whether wheeled or tracked, the robot's overall center of gravity is low, eliminating the risk of robot instability during arm movement.

[0004] However, for legged robots, especially humanoid robots, the grasping process requires consideration not only of the dexterous hand's grasping posture and gripping method, but also of the entire body's coordination to ensure the stability of the humanoid robot. Current base-fixed robot grasping solutions are difficult to directly transfer to humanoid robots.

[0005] Based on this, this specification provides a method for humanoid robot object grasping based on reinforcement learning control. Summary of the Invention

[0006] This specification provides a method and apparatus for grasping humanoid robots based on reinforcement learning control, in order to partially solve the aforementioned problems existing in the prior art.

[0007] The following technical solution is adopted in this specification:

[0008] This specification provides a method for humanoid robot object grasping based on reinforcement learning control, including:

[0009] The attributes of a reference object and a simulated robot are pre-acquired. Using these attributes as states and the joint angles of each joint of the simulated robot when it grasps the reference object as actions, the states are input into a policy model to be trained. This yields the actions output by the policy model, and the reward for the simulated robot performing the actions in the given states is determined. The policy model is then trained with maximizing these rewards as the training objective, resulting in a fully trained policy model. The rewards include at least a first reward for enabling the simulated robot to stand stably and a second reward for enabling the simulated robot to grasp the reference object and achieve a target pose.

[0010] Deploy the trained policy model on the robot;

[0011] When the real-time attributes of the object to be grasped are obtained through the vision device configured in the robot, the real-time attributes of the object to be grasped and the real-time attributes of the robot are input into the pre-trained strategy model to obtain the target joint angles of each joint of the robot output by the strategy model.

[0012] The robot is controlled to grasp the object to be grasped based on the target joint angles of each joint.

[0013] Optionally, before training the policy model based on the reinforcement learning algorithm, the method further includes:

[0014] Obtain the robot's geometric parameters;

[0015] On a pre-built virtual platform, a simulated robot with the same geometric parameters as the actual robot is constructed.

[0016] The simulated robot is optimized using a convex decomposition optimization algorithm.

[0017] Optionally, determining the reward for the simulated robot performing the action in the stated state specifically includes:

[0018] The first reward is determined based on the current pose of a designated part of the simulated robot and the current speed of the designated part of the simulated robot when the simulated robot performs the action in the state.

[0019] The second reward is determined based on the current coordinates of multiple key points on the three-dimensional bounding box of the reference object and the target pose of the reference object when the simulated robot performs the action in the state.

[0020] The third reward is determined based on the current pose of the palm of the dexterous hand configured for grasping of the simulated robot when the simulated robot performs the action in the state, and the current pose of the center of the reference object.

[0021] The fourth reward is determined based on the current coordinates of the fingertips of the dexterous hand configured for grasping of the simulated robot when the robot performs the action in the state, and the current pose of the center of the reference object.

[0022] The reward for the simulated robot to perform the action in the state is determined based on the first reward, the second reward, the third reward, and the fourth reward.

[0023] Optionally, determining the reward for the simulated robot performing the action in the stated state specifically includes:

[0024] Determine the current angular velocity of each joint of the simulated robot when the simulated robot performs the action in the state;

[0025] Based on the current angular velocity of each joint of the simulated robot, a joint velocity penalty term is determined;

[0026] The reward for the simulated robot to perform the action in the state is determined based on the first reward, the second reward, the third reward, the fourth reward, and the joint speed penalty.

[0027] Optionally, a first reward is determined based on the current pose of a designated part of the simulated robot and the current velocity of that designated part when the simulated robot performs the action in the stated state. This specifically includes:

[0028] When the simulated robot performs the action in the state, determine the current pose of a specified part of the simulated robot and the current velocity of the specified part of the simulated robot;

[0029] Based on the current pose of the specified part of the simulated robot, determine the current height of the centroid of the specified part of the simulated robot;

[0030] Determine the difference between the current height of the centroid of a specified part of the simulated robot and a first preset height;

[0031] Based on the velocity component of the current velocity of a specified part of the simulated robot in a specified direction, a stable standing penalty is determined;

[0032] The first reward is determined based on the first reward parameter, the difference between the current height of the center of mass of the specified part of the simulated robot and the first preset height, and the stable standing penalty item.

[0033] Optionally, based on the current coordinates of multiple key points on the three-dimensional bounding box of the reference object when the simulated robot performs the action in the stated state, and the target pose of the reference object, a second reward is determined, specifically including:

[0034] When the simulated robot performs the action in the state, determine the current coordinates of multiple key points on the three-dimensional bounding box of the reference object, and the current pose of the center of the reference object;

[0035] Determine the grasping step function based on the current pose of the center of the reference object;

[0036] Obtain the target coordinates of multiple key points on the 3D bounding box of the reference object under the target pose;

[0037] Based on the current coordinates of the multiple key points and the current difference between the current coordinates of the multiple key points and the target coordinates of the multiple key points, the current distance to be moved is determined;

[0038] The second reward is determined based on the second reward parameter, the current distance to be moved, the preset distance threshold, and the first sparse reward.

[0039] Optionally, based on the current pose of the palm of the dexterous hand configured for grasping of the simulated robot when the simulated robot performs the action in the stated state, and the current pose of the center of the reference object, a third reward is determined, specifically including:

[0040] When the simulated robot performs the action in the state, determine the current pose of the palm of the dexterous hand configured for grasping of the simulated robot, and the current pose of the center of the reference object;

[0041] Based on the difference between the current pose of the palm of the dexterous hand configured for grasping in the simulated robot and the current pose of the center of the reference object, the current distance between the palm of the dexterous hand and the center of the reference object is determined.

[0042] Obtain the minimum historical distance between the palm of the dexterous hand and the center of the reference object;

[0043] The third reward is determined based on the third reward parameter, the current distance, and the minimum historical distance.

[0044] Optionally, based on the current coordinates of the fingertips of the simulated robot's dexterous hand configured for grasping when the simulated robot performs the action in the stated state, and the current pose of the center of the reference object, a fourth reward is determined, specifically including:

[0045] When the simulated robot performs the action in the state, determine the current coordinates of each fingertip of the dexterous hand configured for grasping of the simulated robot, and the current pose of the center of the reference object;

[0046] Determine the grasping step function based on the current pose of the center of the reference object;

[0047] Based on the current coordinates of each fingertip of the dexterous hand and the current pose of the center of the reference object, determine the current distance between each fingertip and the center of the reference object;

[0048] Obtain the minimum historical distance between each fingertip of the dexterous hand and the center of the reference object, and determine the difference between the current distance between each fingertip and the center of the reference object and the minimum historical distance between each fingertip of the dexterous hand and the center of the reference object, as a reference difference;

[0049] The fourth reward is determined based on the grasping step function, the fourth reward parameter, the reference difference, the fifth reward parameter, and the second sparse reward.

[0050] Optionally, the simulated robot includes multiple simulated humanoid robots constructed on a virtual platform. The different simulated humanoid robots correspond to different reference objects with different attributes. The attributes of the reference objects include the size of the reference object, the weight of the reference object, the initial coordinates of multiple key points on the three-dimensional bounding box of the reference object, and the target coordinates of the multiple key points on the three-dimensional bounding box of the reference object for the target pose.

[0051] This specification provides a humanoid robot object grasping device based on reinforcement learning control, including:

[0052] A training module is used to pre-acquire the attributes of a reference object and the attributes of a simulated robot, and input the attributes of the reference object and the simulated robot as states, and the joint angles of each joint of the simulated robot when it grasps the reference object as actions. The state is then input into a policy model to be trained to obtain the actions output by the policy model, and the reward for the simulated robot to perform the actions in the given state is determined. The policy model is trained with maximizing the reward as the training objective to obtain a trained policy model. The reward includes at least a first reward for enabling the simulated robot to stand stably, and a second reward for enabling the simulated robot to grasp the reference object to achieve a target pose.

[0053] The deployment module is used to deploy the trained policy model on the robot;

[0054] The target joint angle determination module is used to input the real-time attributes of the object to be grasped and the real-time attributes of the robot into a pre-trained strategy model when the real-time attributes of the object to be grasped are obtained through the vision device configured in the robot, so as to obtain the target joint angles of each joint of the robot output by the strategy model.

[0055] The control module is used to control the robot to grasp the object to be grasped based on the target joint angles of each joint of the robot.

[0056] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described humanoid robot object grasping method based on reinforcement learning control.

[0057] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described humanoid robot object grasping method based on reinforcement learning control.

[0058] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:

[0059] This specification provides a humanoid robot object grasping method based on reinforcement learning control. Through reinforcement learning training, a first reward for enabling the robot to stand stably and a second reward for achieving the target pose while grasping the object are incorporated into the reward system. This trains an end-to-end policy model, simultaneously achieving both the adaptive task requirements for robot standing balance and the grasping task requirements. It is evident that using the policy model for direct control replaces the steps of grasp point detection and whole-body motion coordination planning required in conventional grasping methods, simplifying the planning process, shortening planning time, and improving the efficiency of real-time robot grasping. Attached Figure Description

[0060] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:

[0061] Figure 1 This is a flowchart illustrating a humanoid robot object grasping method based on reinforcement learning control as described in this specification.

[0062] Figure 2 This is a schematic diagram of a robot that simulates grasping a reference object based on reinforcement learning training, as described in this specification.

[0063] Figure 3This is a flowchart illustrating a humanoid robot object grasping method based on reinforcement learning control as described in this specification.

[0064] Figure 4 This is a schematic diagram of a parallel training method based on multiple simulated robots as described in this specification;

[0065] Figure 5 This specification provides a schematic diagram of a humanoid robot object grasping device based on reinforcement learning control.

[0066] Figure 6 The corresponding information provided in this specification Figure 1 A schematic diagram of an electronic device. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0068] Additionally, it should be noted that all actions involving the acquisition of signals, information, or data in this manual are performed in accordance with the relevant data protection regulations and policies of the locality and with authorization from the owner of the relevant device.

[0069] It should be noted that, unless otherwise specified, the features in the following embodiments and implementation methods can be combined with each other.

[0070] As mentioned earlier, most current robotic grasping solutions use a fixed base, eliminating the risk of robot instability. Therefore, current fixed-base-based grasping solutions cannot be directly applied to humanoid robots, potentially leading to instability and tipping over during grasping tasks. It is therefore crucial to consider the robot's overall coordination and ensure stability while performing grasping tasks in legged robots, especially humanoid robots.

[0071] Furthermore, existing robot grasping technologies generally fall into two categories: one is teaching-based grasping, which relies on pre-set fixed robot motion trajectories and pre-set end-effector movements to perform the grasping task; the other is machine vision-based grasping, which involves target localization, pose estimation, grasping detection, and robot motion planning. The former has poor adaptability; if the actual scene changes (e.g., workpieces on an assembly line have different postures or unequal distances between them), the grasping will fail. The latter has a longer planning process, resulting in longer planning time and lower real-time grasping efficiency.

[0072] Based on this, this specification provides a humanoid robot object grasping method based on reinforcement learning control. When training the policy model based on the reinforcement learning algorithm, a first reward for making the robot stand stably and a second reward for making the robot grasp the object to the target pose are introduced into the reward. This allows the trained policy model to be applicable to both the adaptive task requirements of robot standing balance and the grasping task requirements, and simplifies the planning process of the planning task, thereby improving the efficiency of real-time grasping.

[0073] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0074] Figure 1 This document presents a flowchart illustrating a humanoid robot object grasping method based on reinforcement learning control.

[0075] S100: The attributes of the reference object and the attributes of the simulated robot are obtained in advance. The attributes of the reference object and the attributes of the simulated robot are used as the state, and the joint angles of each joint of the simulated robot when it grasps the reference object are used as the action. The state is input into the policy model to be trained to obtain the action output by the policy model to be trained. The reward for the simulated robot to perform the action in the state is determined. The policy model to be trained is trained with the maximization of the reward as the training objective to obtain the trained policy model. The reward includes at least a first reward for making the simulated robot stand stably and a second reward for making the simulated robot grasp the reference object to reach the target pose.

[0076] In this specification, reinforcement learning is used to train the policy model. The policy model is used to output the joint angles of each joint of the simulated robot based on the attributes of the input reference object (the object to be grasped in the application) and the attributes of the simulated robot (a physical robot in the application) (or the reference attributes of the reference object relative to the simulated robot). Thus, based on the joint angles of each joint of the simulated robot output by the policy model, the simulated robot is controlled to perform corresponding behaviors to perform the task of grasping the reference object.

[0077] Based on the above training approach, the overall reinforcement learning grasping process of the simulated robot in this specification can be regarded as a discrete-time sequential decision-making process. In each decision sub-step, the reference attributes of the reference object relative to the simulated robot are defined as the state, and the joint angles of each joint of the simulated robot output by the policy model are defined as actions. The reward is the reward that the simulated robot receives after performing the current action output by the policy model in the current state in each sub-step.

[0078] The attributes of the reference object include: the coordinates of multiple key points on the 3D bounding box of the reference object, the pose and velocity of the center of the reference object, and the target pose of the target object and the target coordinates of the multiple key points on the 3D bounding box of the reference object.

[0079] The attributes of the simulated robot include: the joint angles and angular velocities of each joint of the simulated robot, the pose and velocity of a specified part of the simulated robot, the pose and velocity of the palm of the dexterous hand configured for grasping, and the coordinates of each fingertip of the dexterous hand.

[0080] The reinforcement learning algorithm used to train the policy model in this specification may be Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), Q-Learning, etc. The model structure of the policy model can be any existing type of neural network structure. The joint angles output by the policy model correspond to joints that can simulate multiple parts of a robot, such as a multi-fingered dexterous hand, arm, and foot. This specification does not limit the actual reinforcement learning algorithm, policy model structure, or joint type used.

[0081] In one optional embodiment of this specification, the algorithm used to train the policy model is the PPO algorithm, which belongs to the Actor-Critic algorithm framework. This is a reinforcement learning algorithm that combines the characteristics of both value-based and policy-based algorithms. The policy model has a structure consisting of a cascaded structure of a one-layer Long Short-Term Memory (LSTM) recurrent neural network and a three-layer multilayer perceptron. The joints corresponding to the joint angles of the simulated robot output by the policy model can be the joints of the entire body, including the two legs (8-DoF), waist (1-DoF), one arm (6-DoF), and multi-fingered dexterous hand (14-DoF), totaling 29 degrees of freedom. Since only single-arm grasping is considered, the joints of the other arm remain fixed.

[0082] During policy model training, the simulated robot can perform movement and object grasping tasks based on the joint angles of each joint output by the policy model. In other words, the simulated robot performs actions within a given state. Since the simulated robot changes state after performing an action, in each sub-step, the policy model determines the current action based on the current state. This current action controls the simulated robot to perform the task. During task execution, the reward for that sub-step is obtained, and the current state is updated. The updated state becomes the state for the next sub-step. Based on maximizing the reward for that sub-step, the parameters of the policy model are adjusted. The adjusted policy model then serves as the policy model for determining the action of the next sub-step. By iteratively executing this process, the policy model can be trained.

[0083] Therefore, determining the reward is crucial. In this specification, the reward for the simulated robot performing actions in a given state includes at least a first reward for enabling the simulated robot to learn to stand stably, and a second reward for enabling the simulated robot to learn to grasp a reference object from an initial pose to a target pose. The first reward is introduced to train the joint angles of each joint output by the policy model to control the simulated robot to maintain its balance and prevent tipping during task execution. The second reward is introduced to train the joint angles of each joint output by the policy model to control the simulated robot to perform the task of grasping and moving a reference object. Thus, after the trained policy model is deployed on a physical robot, in actual grasping tasks, the robot can not only perform grasping and moving tasks but also overcome the problem of humanoid robots easily tipping over, improving the efficiency and stability of grasping tasks.

[0084] It should also be noted that in this specification, the simulated robot is actually a simulated 3D model of the robot built on a virtual platform, with the same geometric parameters as the physical robot. Similarly, the reference object is also built on a virtual platform. Both the simulated robot and the reference object are placed in the simulation environment of the virtual platform, such as... Figure 2 As shown in the diagram, a simulated desktop is also placed in the simulation environment, and the reference object is placed on the simulated desktop. The actions output by the simulated robot based on the policy model ultimately aim to grasp the reference object and move it to the target pose.

[0085] In an optional embodiment of this specification, the simulated robot can be constructed as follows: The geometric parameters of the robot are obtained; a simulated robot with the same geometric parameters as the physical robot is constructed on a pre-built virtual platform; and the simulated robot is optimized using a convex decomposition optimization algorithm. Specifically, before training, a virtual robot 3D model with the same geometric parameters as the physical robot needs to be constructed as the simulated robot, converted into a file of a specified format (such as a URDF file), and it is ensured that the joint steering and maximum joint angle / joint angular velocity limits of the simulated robot are consistent with those of the physical robot. Furthermore, the 3D model used for training involves collision detection during training, but an overly detailed virtual robot 3D model will significantly increase the distance calculation time, leading to a decrease in training speed. Therefore, V-HACD convex decomposition can be used to optimize the virtual robot 3D model, balancing model accuracy and training speed.

[0086] The conditions for completing the training of the strategy model can be that the number of iterations is greater than the preset first threshold, or that the number of times the simulated robot successfully grasps the reference object (grabs the reference object, moves to the target pose and holds it for a period of time) reaches the preset second threshold. This manual does not limit this.

[0087] S102: Deploy the trained policy model on the robot.

[0088] After the strategy model is trained, it can be deployed in the robot's controller. Then, when the robot needs to perform a grasping task, it can obtain the joint angles of multiple joints based on the strategy model deployed in the controller, and send these angles to the motion controller to control the robot's joints to execute the grasping task.

[0089] S104: When the real-time attributes of the object to be grasped are obtained through the vision device configured on the robot, the real-time attributes of the object to be grasped and the real-time attributes of the robot are input into the pre-trained strategy model to obtain the target joint angles of each joint of the robot output by the strategy model.

[0090] Generally, robots are equipped with vision devices and machine vision processing systems to identify and detect objects in their vicinity. In this specification, the robot's vision devices acquire images of the object to be grasped. The robot's vision processing system then uses these images to locate and estimate the object's pose, thereby determining the object's real-time attributes relative to the robot. These real-time attributes include the coordinates of multiple key points on the object's 3D bounding box, the pose and velocity of the object's center, and the target coordinates of these key points on the object's 3D bounding box.

[0091] In this specification, the machine vision processing system configured on the robot can use a single-stage multi-object pose estimation method to complete the target localization and pose estimation steps in the same neural network. For example, using YOLOPose V2, it is possible to simultaneously estimate the 3D bounding box, classification label, and coordinates of multiple key points (3D bounding box corners) of the object to be grasped in the image. The vision device configured on the robot can be a binocular vision acquisition device, and the robot vision processing system can solve the real-time attributes of the object to be grasped based on a binocular stereo matching algorithm.

[0092] Of course, in an optional embodiment of this specification, the binocular vision method can be replaced with other vision-based methods. For example, if the object to be grasped can be marked with a fluorescent ball, a 3D motion capture system can be used to capture the motion of the fluorescent ball in real time and convert it into the current state information of the object to be grasped. Similarly, when the surface of the object to be grasped is flat, ArUco QR codes can be pasted on different surfaces, and a camera can be used to perform real-time pose estimation of the object to be grasped. This specification does not limit the specific types of vision devices and corresponding machine vision processing systems configured on the robot in actual scenarios.

[0093] In this step, the real-time state of the object to be grasped, obtained through the vision device configured on the robot, and the real-time state of the robot are used as observation inputs into the trained policy model. The model can then output the joint angles of each joint based on the current state of the robot and the current state of the object to be grasped.

[0094] S106: Control the robot to grasp the object to be grasped according to the target joint angles of each joint of the robot.

[0095] In practical applications, a robot performing a grasping task can still be viewed as a sequential decision-making process, where each decision sub-step corresponds to controlling the robot to move one step or perform one action. In each decision sub-step, the policy model outputs the joint angles of each joint of the robot at the current decision sub-step, based on the robot's real-time attributes and the real-time attributes of the object to be grasped. That is, the policy model outputs the joint angles of each joint of the robot multiple times, allowing the robot to control different joints based on the joint angles of each decision sub-step, thus achieving the complete robot grasping task.

[0096] This specification provides a humanoid robot object grasping method based on reinforcement learning control. Through reinforcement learning training, a first reward for stable robot standing and a second reward for achieving the target pose while grasping the object are incorporated into the reward system. This trains an end-to-end policy model, simultaneously achieving both the adaptive task requirements for robot standing balance and the grasping task requirements. In the actual execution of the grasping task, the method of directly controlling the robot using the trained policy model replaces the steps of grasp point detection and whole-body motion coordination planning required in conventional grasping methods. This simplifies the planning process, shortens the planning time, and improves the efficiency of real-time robot grasping.

[0097] In one or more embodiments of this specification, when determining the reward obtained by the simulated robot in performing actions under a certain state during the reinforcement learning algorithm training of the policy model in S100, in addition to the first reward for enabling the simulated robot to stand stably and the second reward for enabling the simulated robot to grasp and move the reference object, other types of rewards can also be introduced so that the actions output by the policy model can better control the simulated robot to smoothly and seamlessly perform the grasping task, thereby improving the task execution accuracy and efficiency of the physical robot. The specific scheme is as follows: Figure 3 As shown:

[0098] S200: Determine the first reward based on the current pose of a designated part of the simulated robot and the current speed of the designated part of the simulated robot when the simulated robot performs the action in the state.

[0099] Specifically, the first reward is used to encourage the control simulation robot to stand stably at a relatively stable preset height. The reward for controlling the simulation robot to stand stably can be determined based on the current velocity of a specified part of the simulation robot, and the reward for controlling the simulation robot to maintain a relatively stable preset height can be determined based on the difference between the pose of the specified part of the simulation robot and the preset height.

[0100] Therefore, the first reward can be obtained through the following scheme:

[0101] Step 1: Determine the current pose of a specified part of the simulated robot and the current velocity of the specified part of the simulated robot when the simulated robot performs the action in the state.

[0102] During the training of the policy model, the attributes of both the simulated robot and the reference object can be directly obtained through the virtual platform. Therefore, when the simulated robot performs actions in a given state, the current pose and current velocity of a specified part of the simulated robot can be obtained from its attributes.

[0103] Since the first reward is used to encourage the robot to stand stably, the designated part of the robot can generally be the torso of the robot that does not directly participate in the grasping task. Optionally, the pelvis may be used as an example of the designated part in this specification, and the specific technical solution will be described later.

[0104] Step 2: Determine the current height of the centroid of the specified part of the simulated robot based on the current pose of the specified part of the simulated robot.

[0105] Specifically, such as Figure 2 As shown, during the training of the policy model, the grasping task performed by the simulated robot involves grasping a reference object placed on the simulated table and moving it to the target pose. Therefore, during the controlled completion of the grasping task, it is generally expected that the simulated robot maintains a stable standing position. Thus, the first reward can be related to the current height of the simulated robot's pelvic center of mass. To this end, the current height of the simulated robot's pelvic center of mass can be obtained from the current pose analysis of a specified part of the simulated robot.

[0106] Step 3: Determine the difference between the current height of the centroid of the specified part of the simulated robot and the first preset height.

[0107] Step 4: Determine the stable standing penalty based on the velocity component of the current velocity of the specified part of the simulated robot in the specified direction.

[0108] In practice, if the horizontal velocity of the simulated robot's torso is high, the robot's standing posture may be unstable and prone to swaying. To ensure that the joint angles output by the policy model contribute to the robot's tension stability, a stability penalty term can be determined based on the horizontal velocity component of the robot's pelvis. This stability penalty term is directly proportional to the horizontal velocity component of the robot's pelvis; that is, the larger the horizontal velocity component of the robot's pelvis, the larger the stability penalty term. Introducing a stability penalty term into the first reward ensures that the joint angles output by the policy model control the horizontal velocity of the robot's pelvis, thereby reducing its horizontal velocity and maintaining stability.

[0109] Step 5: Determine the first reward based on the first reward parameter, the difference between the current height of the center of mass of the specified part of the simulated robot and the first preset height, and the stable standing penalty item.

[0110] Specifically, using the first reward parameter as the weight, the difference between the current height of the center of mass of a designated part of the simulated robot and the first preset height is weighted, and the first reward is obtained by subtracting the stable standing penalty from the weighted result. The first reward parameter can be a pre-set fixed value or it can change during the training process; this specification does not limit its implementation.

[0111] One possible implementation of the first reward is as follows:

[0112]

[0113] Where, r 1(stand) (s) is the first reward, i.e., the standing stability reward, α stand Characterizing the first reward parameter, h p h is the current height at which the robot is executing at a specified part. sp Represents the first preset height, The penalty for maintaining a stable standing position.

[0114] S202: Determine the second reward based on the current coordinates of multiple key points on the three-dimensional bounding box of the reference object when the simulated robot performs the action in the state, and the target pose of the reference object.

[0115] In this specification, the difference between the current coordinates of each key point on the 3D bounding box of the reference object and the target coordinates of each key point on the 3D bounding box of the reference object in the target pose is monitored to determine whether the reference object has reached the target pose, or how far it is from reaching the target pose. Generally, in each decision sub-step, when the simulated robot performs an action in the state, the smaller the difference between the current coordinates of multiple key points on the 3D bounding box of the reference object and the target coordinates of multiple key points on the 3D bounding box of the reference object in the target pose, the greater the secondary reward. In this way, the simulated robot can be isolated to grasp the reference object and move it towards the target pose.

[0116] Therefore, the second reward can be obtained through the following scheme:

[0117] Step 1: Determine the current coordinates of multiple key points on the 3D bounding box of the reference object, and the current pose of the center of the reference object, when the simulated robot performs the action in the state.

[0118] Specifically, during the training of the strategy model, since the simulated robot performs tasks within the simulated environment of the virtual platform, it can directly obtain the attributes of reference objects that are also in the simulated environment. These reference object attributes include the current coordinates of multiple key points on the reference object's 3D bounding box.

[0119] In addition, multiple keypoints can be sampled on the 3D bounding box of the reference object. Generally, the corner points (up to eight) of the 3D bounding box can be used as keypoints. To further reduce the computational load, the corner points on any two opposite edges of the 3D bounding box can also be used as multiple keypoints on the 3D bounding box of the reference object.

[0120] Step 2: Determine the grasping step function based on the current pose of the center of the reference object.

[0121] In this specification, when the height of the reference object's center, determined by its current position, is greater than a preset height threshold, the grab step function is set to 1. When the height of the reference object's center is not greater than the preset height threshold, the grab step function is set to 0.

[0122] Step 3: Obtain the target coordinates of multiple key points on the 3D bounding box of the reference object under the target pose.

[0123] This step is similar to the first step described above. The target pose is the pre-set target of the grasping task I proposed before training. The 3D bounding box information of the reference object under the target pose can be directly obtained in the simulation environment, including the target coordinates of multiple keypoints on the 3D bounding box. It is important to note that the multiple keypoints corresponding to the target coordinates in the second step are the same multiple keypoints on the 3D bounding box of the reference object as the multiple keypoints corresponding to the current coordinates in the first step. During the iterative training of the policy model using reinforcement learning, the target pose of the reference object remains unchanged in each iteration; the positions of the multiple keypoints on the 3D bounding box also remain unchanged. However, the aforementioned information may or may not change during different training iterations; this specification does not impose any restrictions on this.

[0124] Step 4: Determine the current distance to be moved based on the current coordinates of the multiple key points and the current difference between the current coordinates of the multiple key points and the target coordinates of the multiple key points.

[0125] Specifically, the current difference between the current coordinates of each key point and the target coordinates can be calculated. The current distance to be moved can be obtained by statistically analyzing the current differences of each key point, such as taking the maximum value, summing, averaging, variance, etc. Alternatively, the current distance of each key point can be used as the current distance to be moved.

[0126] Step 5: Determine the second reward based on the second reward parameter, the current distance to be moved, the preset distance threshold, and the first sparse reward.

[0127] The purpose of the second reward is to provide a larger reward when the simulated robot grasps and moves a reference object, bringing the reference object to the target pose. Therefore, in this specification, if the current distance to be moved is less than or equal to a preset distance threshold, the second reward is the larger value of the first sparse reward. If the current distance to be moved is greater than the preset distance threshold, no second reward is given, i.e., the second reward is 0. Therefore, the relationship between the current distance to be moved and the preset distance threshold serves as the weight of the first sparse reward, indicating whether the first sparse reward exists in the second reward under different circumstances.

[0128] The second reward parameter can be a pre-set fixed value or it can change during the training process; this manual does not limit this.

[0129] Optionally, the second reward may also incorporate the minimum value among multiple current distances to be moved, determined by the differences between the current coordinates of various key points of the reference object and the target coordinates, as determined in multiple decision sub-steps prior to the current decision sub-step; that is, the historical minimum distance to be moved. Specifically, the difference between the historical minimum distance to be moved and the current distance to be moved is introduced into the second reward to encourage the simulated robot to grasp the reference object and move in a direction closer to the target pose.

[0130] Therefore, one possible implementation of the second reward is as follows:

[0131]

[0132]

[0133]

[0134] Where, r 2(targ) (s) is the second reward, that is, the reward for reaching the target pose, α targ Characterizing the second reward parameter, I picked It is a step function; when the height of the reference object is higher than a preset value, I... picked A value of 1 indicates that the reference object has moved to a certain height above the simulated tabletop. ε * This is a preset distance threshold used to characterize the difference threshold between the current pose of the reference object and the target pose. This represents the current distance to be moved (the maximum difference between all keypoints), used to measure the difference between the current pose of the reference object and the target position. The formula for determining the current distance to be moved includes... Let i be the target pose of the i-th key point. Let i be the reference pose of the i-th keypoint. The historical minimum distance to be moved for each key point of the reference object.

[0135] S204: Determine the third reward based on the current pose of the palm of the dexterous hand configured for grasping of the simulated robot when the simulated robot performs the action in the state, and the current pose of the center of the reference object.

[0136] In practice, when a robot's dexterous hand grasps a reference object, in order to enhance the stability of the grasp and prevent the reference object from falling during movement, the palm of the dexterous hand can be placed close to the center of the reference object during the grasping process, thus enabling the dexterous hand to grasp the reference object more stably.

[0137] Therefore, the current coordinates of the dexterous hand's palm can be obtained by analyzing the current pose of the dexterous hand's palm used for grasping, which is included in the attributes of the simulated robot. The current coordinates of the reference object's center can also be obtained by analyzing its current pose, which is included in the attributes of the reference object. Thus, a third reward is determined based on the difference between the current coordinates of the dexterous hand's palm and the current coordinates of the reference object's center. Generally, the difference between the current coordinates of the dexterous hand's palm and the current coordinates of the reference object's center is inversely proportional to the third reward; that is, the smaller the difference, the larger the third reward.

[0138] Therefore, the third reward can be obtained through the following scheme:

[0139] Step 1: Determine the current pose of the palm of the dexterous hand configured for grasping of the simulated robot, and the current pose of the center of the reference object, when the simulated robot performs the action in the state.

[0140] Step 2: Determine the current distance between the palm of the dexterous hand and the center of the reference object based on the difference between the current pose of the palm of the dexterous hand configured for grasping in the simulated robot and the current pose of the center of the reference object.

[0141] Step 3: Obtain the minimum historical distance between the palm of the dexterous hand and the center of the reference object.

[0142] Optionally, in this specification, in addition to determining the third reward based on the current distance between the palm of the dexterous hand and the center of the reference object, a historical distance between the palm of the dexterous hand and the center of the reference object may also be introduced. Here, the historical distance between the palm of the dexterous hand and the center of the reference object refers to the distance between the palm of the dexterous hand and the reference object after the simulated robot performed an action in one or more decision sub-steps prior to the current decision sub-step.

[0143] Introducing the historical distance between the dexterous hand's palm and the center of the reference object to determine the third reward is to prevent the distance between the dexterous hand's palm and the center of the reference object from increasing. Therefore, during training, the task objective of the grasping task is to simulate the robot stably grasping the reference object to reach the target pose using its dexterous hand. If the distance between the dexterous hand's palm and the center of the reference object increases, unstable grasping or even loss may occur. Therefore, incorporating the historical distance between the dexterous hand's palm and the center of the reference object into the third reward limits the direction of distance growth between the dexterous hand's palm and the center of the reference object, thereby improving grasping stability.

[0144] In this specification, the third reward is determined by the minimum historical distance between the palm of the dexterous hand and the reference object after the robot performs an action in one or more decision sub-steps before the current decision sub-step. This can further limit the development direction of the distance between the palm of the dexterous hand and the center of the reference object.

[0145] Step 4: Determine the third reward based on the third reward parameter, the current distance, and the minimum historical distance.

[0146] Specifically, first determine the current distance d between the palm of the dexterous hand and the center of the reference object, and the minimum historical distance d between the palm of the dexterous hand and the center of the reference object. closest The difference between them. When (d closest When -d) is greater than 0, it indicates that the current distance is smaller than the minimum historical distance, meaning that the direction of the distance development between the palm of the dexterous hand and the center of the reference object is correct, and a larger reward can be given. When (d) closest -d) If it is not greater than 0, it means that the current distance is greater than or equal to the minimum historical distance. This means that the direction of the distance between the palm of the dexterous hand and the center of the reference object is wrong. The third reward needs to be set to zero or even punished (the third reward is less than 0).

[0147] One possible implementation of the third reward is as follows:

[0148] r3(reach) (s)=α reach *max(d closest -d,0)

[0149] Where, r 3(reach) (s) is the third reward, α reach Characterizing the second reward parameter, d closest The minimum historical distance between the palm of the dexterous hand and the center of the reference object is represented by d, and the current distance between the palm of the dexterous hand and the center of the reference object is represented by d.

[0150] The third reward parameter can be a pre-set fixed value or it can change during the training process; this manual does not limit this.

[0151] S206: Determine the fourth reward based on the current coordinates of the fingertips of the dexterous hand configured for grasping of the simulated robot when the simulated robot performs the action in the state, and the current pose of the center of the reference object.

[0152] In practice, when a robot's dexterous hand grasps a reference object, in order to enhance the stability of the grasp and prevent the reference object from falling during movement, the fingertips of the dexterous hand can be positioned closer to the center of the reference object during the grasping process. This allows the dexterous hand to surround and encircle the reference object, thus grasping it more stably.

[0153] To this end, the current coordinates of each fingertip of the dexterous hand configured for grasping in the simulated robot can be obtained from the robot's attributes. The current coordinates of the reference object's center can then be obtained by analyzing its current pose, which is included in the reference object's attributes. Based on the difference between the current coordinates of the dexterous hand's fingertips and the current coordinates of the reference object's center, a fourth reward can be determined.

[0154] Generally, the difference between the current coordinates of the fingertips of a dexterous hand and the current coordinates of the center of the reference object is inversely proportional to the fourth reward; that is, the smaller the difference between the current coordinates of the fingertips of a dexterous hand and the current coordinates of the center of the reference object, the greater the fourth reward.

[0155] Additionally, the fourth reward may include a grasping step function related to the height of the reference object from the simulated tabletop. The grasping step function is used by the dexterous hand to grasp the object and move it up the platform to a height that moves the reference object away from the simulated tabletop.

[0156] Therefore, the fourth reward can be obtained through the following scheme:

[0157] One possible implementation of the fourth reward is as follows:

[0158] Step 1: Determine the current coordinates of each fingertip of the dexterous hand configured for grasping of the simulated robot, and the current pose of the center of the reference object, when the simulated robot performs the action in the state.

[0159] Step 2: Determine the grasping step function based on the current pose of the center of the reference object.

[0160] Specifically, the current coordinates of the reference object's center are determined based on its current pose. The difference between these coordinates and the height of the simulated tabletop is then calculated. If this difference exceeds a preset height threshold, it indicates that the reference object has been grasped and lifted away from the simulated tabletop by the simulated robot. In this case, the grasping step function is set to 1.

[0161] If the difference between the current coordinates of the center of the reference object and the height of the simulated table is not greater than a preset height threshold, it means that the reference object has not yet been grasped and lifted away from the simulated table by the simulated robot. In this case, the grasping step function is set to 0.

[0162] Step 3: Based on the current coordinates of each fingertip of the dexterous hand and the current pose of the center of the reference object, determine the current distance between each fingertip and the center of the reference object.

[0163] Determine the current coordinates of the center of the reference object based on its current pose.

[0164] In this specification, the distance between the current coordinates of each fingertip of the dexterous hand and the current coordinates of the center of the reference object is determined. Then, the distances between the current coordinates of each fingertip of the dexterous hand and the current coordinates of the center of the reference object are summed to obtain the current distance between each fingertip and the center of the reference object.

[0165] Step 4: Obtain the minimum historical distance between each fingertip of the dexterous hand and the center of the reference object, and determine the difference between the current distance between each fingertip and the center of the reference object and the minimum historical distance between each fingertip of the dexterous hand and the center of the reference object, as a reference difference.

[0166] Furthermore, in addition to determining the fourth reward based on the current distance between each fingertip of the dexterous hand and the center of the reference object, this specification can also introduce the historical distance between each fingertip of the dexterous hand and the center of the reference object. Here, the historical distance between each fingertip of the dexterous hand and the center of the reference object refers to the distance between each fingertip of the dexterous hand and the reference object after the simulated robot performed an action in one or more decision sub-steps prior to the current decision sub-step.

[0167] Introducing the historical distances between the fingertips of the dexterous hand and the center of the reference object to determine the fourth reward is to prevent the distances between the fingertips and the center of the reference object from increasing. During training, the task objective of the grasping task is to simulate the robot stably grasping the reference object to reach the target pose using its dexterous hand. Stable grasping can involve the dexterous hand surrounding and encircling the reference object. If the distances between the fingertips of the dexterous hand and the center of the reference object increase, problems such as incomplete encirclement, unstable grasping, or even detachment may occur. Therefore, incorporating the historical distances between the fingertips of the dexterous hand and the center of the reference object into the fourth reward is used to limit the direction of the increase in the distances between the fingertips of the dexterous hand and the center of the reference object, thereby improving grasping stability.

[0168] In this specification, the fourth reward is determined by the minimum historical distance between each fingertip of the dexterous hand and the reference object after the robot performs an action in one or more decision sub-steps before the current decision sub-step. This can further limit the development direction of the distance between each fingertip of the dexterous hand and the center of the reference object.

[0169] Furthermore, the reference difference is determined based on the difference between the current distance between each fingertip and the center of the reference object and the minimum historical distance between each fingertip of the dexterous hand and the center of the reference object.

[0170] Step 5: Determine the fourth reward based on the grasping step function, the fourth reward parameter, the reference difference, the fifth reward parameter, and the second sparse reward.

[0171] The fourth reward parameter is used to weight the difference from the reference object, and the fifth reward parameter is used to weight the distance between the current coordinates of the reference object and the height of the simulated tabletop. The fourth and fifth reward parameters can be preset fixed values ​​or can be changed during training; this specification does not impose any restrictions on this.

[0172] One possible implementation of the fourth reward is as follows:

[0173] r 4(finger)(s)=(1-I picked )*(α finger *r finger +α pick *h t )+I picked *r picked

[0174]

[0175] Where, r 4(finger) (s) is the fourth reward, α finger Characterizing the fourth reward parameter, α pick The fifth reward parameter, r finger Characterizing the reference difference, h t This represents the height of the current coordinates of the reference object relative to the simulated desktop. picked To capture the step function, r picked This is the second sparse matrix. The minimum historical distance d is used to characterize the individual fingertips of a dexterous hand from the center of a reference object. tipsN Used to represent the current distance between the fingertips of a dexterous hand and the center of a reference object.

[0176] S208: Determine the reward for the simulated robot to perform the action in the state based on the first reward, the second reward, the third reward, and the fourth reward.

[0177] Based on the sum of the first, second, third, and fourth rewards determined above, the reward for the simulated robot to perform an action in the given state can be determined.

[0178] The following is one possible implementation method:

[0179] r(s,a)=r 1(stand) (s)+r 2(targ) (s)+r 3(reach) (s)+r 4(pick) (s)

[0180] Furthermore, in order to make the motion control simulation robot output by the strategy model smoother, a joint speed penalty term for the simulated robot can be introduced into the reward. This joint speed penalty term is used to ensure that the angles of each joint of the simulated robot are not too large, so as not to produce excessively large movements and promote smoother motion.

[0181] Specifically, S208 can be:

[0182] Step 1: Determine the current angular velocity of each joint of the simulated robot when the simulated robot performs the action in the state.

[0183] Step 2: Determine the joint velocity penalty term based on the current angular velocity of each joint of the simulated robot.

[0184] Since the joint velocity penalty term is used to control the joint angular velocity of each joint of the simulated robot to prevent it from becoming too large, the current angular velocity of each joint of the simulated robot is directly proportional to the joint velocity penalty term. That is, the larger the current angular velocity of each joint of the simulated robot, the larger the joint velocity penalty term.

[0185] Step 3: Determine the reward for the simulated robot to perform the action in the state based on the first reward, the second reward, the third reward, the fourth reward, and the joint speed penalty.

[0186] The reward that includes a joint speed penalty smooths out the speed changes of each joint when controlling the robot to grasp, reduces joint impact, and is more conducive to the transfer to physical robots.

[0187] With r vel (a) Characterizing the joint velocity penalty term, one possible implementation of the reward obtained by the simulated robot when performing actions in a certain state is as follows:

[0188] r(s,a)=r 1(stand) (s)+r 2(targ) (s)+r 3(reach) (s)+r 4(pick) (s)-r vel (a)

[0189] It is important to note that when the simulated robot performs the task of grasping a reference object, and controls the robot's actions step-by-step based on the joint angles of multiple joints output by the strategy model in each decision sub-step, the actual process is that the robot's dexterous hand first moves to the vicinity of the reference object and grasps it. At this point, since the height of the reference object's center from the simulated tabletop is less than a preset height threshold, the grasping step function is 0. Furthermore, the current distance to be moved is generally greater than a preset distance threshold; therefore, the second reward is 0, meaning there is no second reward. The fourth reward can be implemented as follows:

[0190] r 4(pick) (s)=α finger *r finger +α pick *h t

[0191] The reward at this point is actually:

[0192] r(s,a)=r 1(stand) (s)+r 3(reach)(s)+α finger *r finger +α pick *h t -r vel (a)

[0193] When the simulated robot's dexterous hand grasps and moves a reference object, and the height of the reference object's center from the simulated tabletop is greater than a preset height threshold, but the distance to be moved is greater than a preset distance threshold, the grasping step function is set to 1. The actual implementation of the second reward can be as follows:

[0194]

[0195] The fourth reward can be implemented in the following way:

[0196] r 4(finger) (s)=r picked

[0197] That is, when the height of the center of the reference object from the simulated tabletop is greater than the preset height threshold, the fourth reward becomes a fixed value, namely the second sparse reward.

[0198] The reward at this point is actually:

[0199]

[0200] Furthermore, when the current distance to be moved is not greater than a preset distance threshold, the actual implementation of the second reward can be as follows:

[0201]

[0202] The reward at this point is actually:

[0203]

[0204] Therefore, it can be seen that the second and fourth rewards take different forms at different stages.

[0205] In one or more embodiments of this specification, the highly parallel GPU-accelerated physics engine of Isaac Gym can be used for multi-robot reinforcement learning training, thereby accelerating the training process of the policy model and greatly speeding up the policy learning process, such as... Figure 4 As shown.

[0206] Specifically, multiple simulated humanoid robots are constructed on a virtual platform, and the geometric parameters of these simulated humanoid robots are the same as those of the physical humanoid robots. However, the reference objects grasped by different simulated humanoid robots can be diverse. That is, the attributes of the reference objects corresponding to different simulated humanoid robots are different. The attributes of the reference objects include the size of the reference object, the weight of the reference object, the initial coordinates of multiple key points on the 3D bounding box of the reference object, and the target coordinates of the multiple key points on the 3D bounding box of the reference object for the target pose.

[0207] In an optional embodiment of this specification, during the training of the policy model such as S100 using reinforcement learning, if the simulated robot grasps the reference object and reaches the target pose, and the difference between the current pose and the target pose of the reference object is not greater than the tolerance within a duration τ, the grasping attempt is considered successful. If the reference object falls during the attempt, fails to reach the target pose within a preset time, or although the reference object reaches the target pose within the preset time, but the difference between the current pose and the target pose of the reference object is greater than the tolerance within a duration τ, the grasping attempt is deemed to have failed. Once the grasping attempt fails, or N consecutive grasping attempts succeed, the grasping attempt is considered unsuccessful. max Next, it is necessary to reset the joint angles of each joint of the simulated robot, the properties of the reference object, and the simulation environment. Based on this, the following three preset conditions are set. When the reference object meets any of the following conditions, the simulated robot is reset.

[0208] First preset condition: The current difference between the current coordinates of multiple key points on the 3D bounding box of the reference object and the target coordinates of the multiple key points on the 3D bounding box of the reference object under the target pose is greater than a preset difference threshold.

[0209] The second preset condition is: the moment when the simulated robot grasps the reference object is taken as the start time, and the moment when the current difference between the current coordinates of multiple key points on the three-dimensional bounding box of the reference object and the target coordinates of the multiple key points on the three-dimensional bounding box of the reference object under the target pose is less than the tolerance is taken as the end time. The duration of the time period from the start time to the end time is greater than the preset duration.

[0210] The third preset condition is that the current difference between the current coordinates of multiple key points on the 3D bounding box of the reference object and the target coordinates of the multiple key points on the 3D bounding box of the reference object under the target pose is less than the tolerance, so that the duration is less than the duration.

[0211] Fourth preset condition: The number of successful captures is greater than the preset threshold.

[0212] The above describes one or more embodiments of a humanoid robot object grasping method based on reinforcement learning control, as provided in this specification. Based on the same approach, this specification also provides corresponding humanoid robot object grasping devices based on reinforcement learning control, such as… Figure 5 As shown.

[0213] Figure 5 This specification provides a schematic diagram of a humanoid robot object grasping device based on reinforcement learning control, which specifically includes:

[0214] The training module 300 is used to pre-acquire the attributes of a reference object and the attributes of a simulated robot, and input the attributes of the reference object and the simulated robot as states, and the joint angles of each joint of the simulated robot when it grasps the reference object as actions, into a policy model to be trained, to obtain the actions output by the policy model to be trained, and to determine the reward for the simulated robot to perform the actions in the state, and to train the policy model to be trained with maximizing the reward as the training objective, to obtain a trained policy model; wherein, the reward includes at least a first reward for making the simulated robot stand stably, and a second reward for making the simulated robot grasp the reference object to achieve a target pose.

[0215] Deployment module 302 is used to deploy the trained strategy model on the robot;

[0216] The target joint angle determination module 304 is used to input the real-time attributes of the object to be grasped and the real-time attributes of the robot obtained by the vision device configured by the robot into a pre-trained strategy model to obtain the target joint angles of each joint of the robot output by the strategy model when the real-time attributes of the object to be grasped are obtained by the vision device configured by the robot.

[0217] The control module 306 is used to control the robot to grasp the object to be grasped according to the target joint angles of each joint of the robot.

[0218] Optionally, the device further includes:

[0219] The simulated robot construction module 308 is specifically used to obtain the geometric parameters of the robot; construct a simulated robot with the same geometric parameters as the robot on a pre-built virtual platform; and optimize the simulated robot using a convex decomposition optimization algorithm.

[0220] Optionally, the training module 300 is specifically configured to: determine a first reward based on the current pose and current velocity of a designated part of the simulated robot when the simulated robot performs the action in the state; determine a second reward based on the current coordinates of multiple key points on the three-dimensional bounding box of the reference object and the target pose of the reference object when the simulated robot performs the action in the state; determine a third reward based on the current pose of the palm of the dexterous hand configured for grasping of the simulated robot and the current pose of the center of the reference object when the simulated robot performs the action in the state; determine a fourth reward based on the current coordinates of each fingertip of the dexterous hand configured for grasping of the simulated robot and the current pose of the center of the reference object when the simulated robot performs the action in the state; and determine the reward for the simulated robot performing the action in the state based on the first reward, the second reward, the third reward, and the fourth reward.

[0221] Optionally, the training module 300 is specifically configured to: determine the current angular velocity of each joint of the simulated robot when the simulated robot performs the action in the state; determine a joint velocity penalty term based on the current angular velocity of each joint of the simulated robot; and determine the reward for the simulated robot to perform the action in the state based on the first reward, the second reward, the third reward, the fourth reward and the joint velocity penalty term.

[0222] Optionally, the training module 300 is specifically configured to: determine the current pose of a designated part of the simulated robot and the current velocity of the designated part of the simulated robot when the simulated robot performs the action in the state; determine the current height of the center of mass of the designated part of the simulated robot based on the current pose of the designated part of the simulated robot; determine the difference between the current height of the center of mass of the designated part of the simulated robot and a first preset height; determine a stable standing penalty item based on the velocity component of the current velocity of the designated part of the simulated robot in a designated direction; and determine a first reward based on a first reward parameter, the difference between the current height of the center of mass of the designated part of the simulated robot and the first preset height, and the stable standing penalty item.

[0223] Optionally, the training module 300 is specifically configured to: determine the current coordinates of multiple key points on the three-dimensional bounding box of the reference object and the current pose of the center of the reference object when the simulated robot performs the action in the state; determine the grasping step function based on the current pose of the center of the reference object; obtain the target coordinates of the multiple key points on the three-dimensional bounding box of the reference object under the target pose; determine the current distance to be moved based on the current difference between the current coordinates of the multiple key points and the target coordinates of the multiple key points; and determine the second reward based on the second reward parameter, the current distance to be moved, the preset distance threshold, and the first sparse reward.

[0224] Optionally, the training module 300 is specifically configured to: determine the current pose of the palm of the dexterous hand configured for grasping of the simulated robot and the current pose of the center of the reference object when the simulated robot performs the action in the state; determine the current distance between the palm of the dexterous hand and the center of the reference object based on the difference between the current pose of the palm of the dexterous hand configured for grasping of the simulated robot and the current pose of the center of the reference object; obtain the minimum historical distance between the palm of the dexterous hand and the center of the reference object; and determine a third reward based on a third reward parameter, the current distance, and the minimum historical distance.

[0225] Optionally, the training module 300 is specifically configured to: determine the current coordinates of each fingertip of the dexterous hand configured for grasping of the simulated robot when the simulated robot performs the action in the state, and the current pose of the center of the reference object; determine a grasping step function based on the current pose of the center of the reference object; determine the current distance between each fingertip and the center of the reference object based on the current coordinates of each fingertip of the dexterous hand and the current pose of the center of the reference object; obtain the minimum historical distance between each fingertip of the dexterous hand and the center of the reference object, and determine the difference between the current distance between each fingertip and the center of the reference object and the minimum historical distance between each fingertip of the dexterous hand and the center of the reference object as a reference difference; and determine a fourth reward based on the grasping step function, a fourth reward parameter, the reference difference, a fifth reward parameter, and a second sparse reward.

[0226] Optionally, the simulated robot includes multiple simulated humanoid robots constructed on a virtual platform. The different simulated humanoid robots correspond to different reference objects with different attributes. The attributes of the reference objects include the size of the reference object, the weight of the reference object, the initial coordinates of multiple key points on the three-dimensional bounding box of the reference object, and the target coordinates of the multiple key points on the three-dimensional bounding box of the reference object for the target pose.

[0227] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 The method shown is a humanoid robot object grasping method based on reinforcement learning control.

[0228] This instruction manual also provides Figure 6 The diagram shows a schematic structural representation of the electronic device. Figure 6 At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for the business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 1 The method shown is a humanoid robot object grasping method based on reinforcement learning control. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.

[0229] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0230] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0231] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0232] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.

[0233] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0234] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0235] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0236] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0237] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0238] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0239] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0240] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0241] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0242] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0243] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0244] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.

Claims

1. A method for grasping objects using a humanoid robot based on reinforcement learning control, characterized in that, include: Get the properties of the reference object and the properties of the simulated robot; The state is defined by the properties of the reference object and the properties of the simulated robot, and the action is defined by the joint angles of each joint of the simulated robot when the simulated robot grasps the reference object. The state is input into the policy model to be trained to obtain the action output by the policy model to be trained; Based on the current pose of a designated part of the simulated robot and the current velocity of the designated part of the simulated robot when the simulated robot performs the action in the state, a first reward for making the simulated robot stand stably is determined. Based on the current coordinates of multiple key points on the three-dimensional bounding box of the reference object when the simulated robot performs the action in the state, and the target pose of the reference object, a second reward is determined to enable the simulated robot to grasp the reference object and reach the target pose. The third reward is determined based on the current pose of the palm of the dexterous hand configured for grasping of the simulated robot when the simulated robot performs the action in the state, and the current pose of the center of the reference object. The fourth reward is determined based on the current coordinates of the fingertips of the dexterous hand configured for grasping of the simulated robot when the robot performs the action in the state, and the current pose of the center of the reference object. Based on the first reward, the second reward, the third reward, and the fourth reward, determine the reward for the simulated robot to perform the action in the state; The policy model to be trained is trained with the goal of maximizing the reward, and the trained policy model is obtained. Deploy the trained policy model on the robot; When the real-time attributes of the object to be grasped are obtained through the vision device configured in the robot, the real-time attributes of the object to be grasped and the real-time attributes of the robot are input into the pre-trained strategy model to obtain the target joint angles of each joint of the robot output by the strategy model. The robot is controlled to grasp the object to be grasped based on the target joint angles of each joint.

2. The method as described in claim 1, characterized in that, The acquisition of the simulated robot's attributes includes: Obtain the robot's geometric parameters; On a pre-built virtual platform, a simulated robot with the same geometric parameters as the actual robot is constructed. The simulated robot is optimized using a convex decomposition optimization algorithm.

3. The method as described in claim 1, characterized in that, Determining the reward for the simulated robot performing the action in the stated state specifically includes: Determine the current angular velocity of each joint of the simulated robot when the simulated robot performs the action in the state; Based on the current angular velocity of each joint of the simulated robot, a joint velocity penalty term is determined; The reward for the simulated robot to perform the action in the state is determined based on the first reward, the second reward, the third reward, the fourth reward, and the joint speed penalty.

4. The method as described in claim 1, characterized in that, The first reward is determined based on the current pose and current velocity of a designated part of the simulated robot when the simulated robot performs the action in the state. Specifically, this includes determining the current pose and current velocity of a designated part of the simulated robot when the simulated robot performs the action in the state. Based on the current pose of the specified part of the simulated robot, determine the current height of the centroid of the specified part of the simulated robot; Determine the difference between the current height of the centroid of a specified part of the simulated robot and a first preset height; Based on the velocity component of the current velocity of a specified part of the simulated robot in a specified direction, a stable standing penalty is determined; The first reward is determined based on the first reward parameter, the difference between the current height of the center of mass of the specified part of the simulated robot and the first preset height, and the stable standing penalty item.

5. The method as described in claim 1, characterized in that, The second reward is determined based on the current coordinates of multiple key points on the three-dimensional bounding box of the reference object and the target pose of the reference object when the simulated robot performs the action in the state. Specifically, this includes determining the current coordinates of multiple key points on the three-dimensional bounding box of the reference object and the current pose of the center of the reference object when the simulated robot performs the action in the state. Determine the grasping step function based on the current pose of the center of the reference object; Obtain the target coordinates of multiple key points on the 3D bounding box of the reference object under the target pose; The current distance to be moved is determined based on the current coordinates of the multiple key points and the current difference between them and the target coordinates of the multiple key points; the second reward is determined based on the second reward parameter, the current distance to be moved, the preset distance threshold, and the first sparse reward.

6. The method as described in claim 1, characterized in that, Based on the current pose of the palm of the dexterous hand configured for grasping of the simulated robot when the robot performs the action in the stated state, and the current pose of the center of the reference object, a third reward is determined, specifically including: When the simulated robot performs the action in the state, determine the current pose of the palm of the dexterous hand configured for grasping of the simulated robot, and the current pose of the center of the reference object; Based on the difference between the current pose of the palm of the dexterous hand configured for grasping in the simulated robot and the current pose of the center of the reference object, the current distance between the palm of the dexterous hand and the center of the reference object is determined. Obtain the minimum historical distance between the palm of the dexterous hand and the center of the reference object; The third reward is determined based on the third reward parameter, the current distance, and the minimum historical distance.

7. The method as described in claim 1, characterized in that, Based on the current coordinates of the fingertips of the simulated robot's dexterous hand configured for grasping when the simulated robot performs the action in the stated state, and the current pose of the center of the reference object, a fourth reward is determined, specifically including: When the simulated robot performs the action in the state, determine the current coordinates of each fingertip of the dexterous hand configured for grasping of the simulated robot, and the current pose of the center of the reference object; Determine the grasping step function based on the current pose of the center of the reference object; Based on the current coordinates of each fingertip of the dexterous hand and the current pose of the center of the reference object, determine the current distance between each fingertip and the center of the reference object; Obtain the minimum historical distance between each fingertip of the dexterous hand and the center of the reference object, and determine the difference between the current distance between each fingertip and the center of the reference object and the minimum historical distance between each fingertip of the dexterous hand and the center of the reference object, as a reference difference; The fourth reward is determined based on the grasping step function, the fourth reward parameter, the reference difference, the fifth reward parameter, and the second sparse reward.

8. The method according to any one of claims 1 to 7, characterized in that, The simulated robot includes multiple simulated humanoid robots constructed on a virtual platform. The different simulated humanoid robots correspond to different reference objects with different attributes. The attributes of the reference objects include the size of the reference object, the weight of the reference object, the initial coordinates of multiple key points on the three-dimensional bounding box of the reference object, and the target coordinates of the multiple key points on the three-dimensional bounding box of the reference object for the target pose.

9. A humanoid robot object grasping device based on reinforcement learning control, characterized in that, include: The training module is used to acquire the properties of the reference object and the simulated robot, and The state is defined by the properties of the reference object and the properties of the simulated robot, and the action is defined by the joint angles of each joint of the simulated robot when the simulated robot grasps the reference object. The state is input into the policy model to be trained, and the action output by the policy model is obtained. Based on the current pose of a designated part of the simulated robot and the current velocity of the designated part of the simulated robot when the simulated robot performs the action in the state, a first reward for making the simulated robot stand stably is determined. Based on the current coordinates of multiple key points on the three-dimensional bounding box of the reference object when the simulated robot performs the action in the state, and the target pose of the reference object, a second reward is determined to enable the simulated robot to grasp the reference object and reach the target pose. The third reward is determined based on the current pose of the palm of the dexterous hand configured for grasping of the simulated robot when the simulated robot performs the action in the state, and the current pose of the center of the reference object. The fourth reward is determined based on the current coordinates of the fingertips of the dexterous hand configured for grasping of the simulated robot when the robot performs the action in the state, and the current pose of the center of the reference object. Based on the first reward, the second reward, the third reward, and the fourth reward, the reward for the simulated robot performing the action in the stated state is determined. The policy model to be trained is trained with the goal of maximizing the reward, and the trained policy model is obtained. The deployment module is used to deploy the trained policy model on the robot; The target joint angle determination module is used to input the real-time attributes of the object to be grasped and the real-time attributes of the robot into a pre-trained strategy model when the real-time attributes of the object to be grasped are obtained through the vision device configured in the robot, so as to obtain the target joint angles of each joint of the robot output by the strategy model. The control module is used to control the robot to grasp the object to be grasped based on the target joint angles of each joint of the robot.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 8.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for learning hand and object interactive motion control from RGBD videos

    CN112720504A

  • Robot turning and squatting stable standing control method and device and related equipment

    CN117301061A