Robot turning and squatting stable standing control method and device and related equipment

CN117301061BActive Publication Date: 2026-08-18IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311357479.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-19
Publication Date
2026-08-18
Estimated Expiration
2043-10-19

AI Technical Summary

Technical Problem

[0003]以往的对人形机器人的动作控制通常采用一些比较传统的算法,如PID算法等,在测试阶段需要在真机上验证并调优算法,这个过程不可避免的会由于算法的缺陷给机器人带来不利影响,严重时甚至给机器人硬件带来不可逆的损伤,极大浪费了硬件设备资源

Benefits of technology

[0045]Using the above technical solution, this application employs a reinforcement learning training method to train the policy network in a simulator. The network input includes the proprioceptive information of the target humanoid robot, such as the specified rotation angle when performing a turning task and the specified pelvic height when performing a squatting task. Further reinforcement learning reward functions can include a first function to encourage the robot to squat to the specified pelvic height, a second function to encourage the robot's pelvic motion speed in the z-axis direction to approach the set desired squatting speed, and a third function to encourage the robot to turn to the specified rotation angle. By setting these reward functions, the policy network can be trained in the simulator, enabling it to control the robot to perform turning and squatting tasks effectively and maintain stable standing during the task process. After completing the simulation training of the policy network, it can be deployed in a real environment to obtain the proprioceptive information of the target humanoid robot. This proprioceptive information is then input into the simulation-trained policy network to obtain the n-dimensional motion output by the policy network. A proportional-derivative (PD) controller converts the n-dimensional motion into joint torques, which drive the corresponding joint motors to perform the turning and squatting tasks. This embodiment trains the policy network in a simulator using reinforcement learning, with the robot's proprioceptive information as input. A predefined reward function is used to train the policy network through reinforcement learning. The simulated-trained policy network can be deployed to a real-world environment without requiring algorithm optimization testing on the robot in a real-world setting, thus avoiding wear and tear on the robot's hardware. Furthermore, the reinforcement learning training method allows for trial and error in the simulation environment, with rewards used as feedback for reinforcement learning. This leads to a higher-performance policy network, enabling stable standing control during robot turns and crouching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117301061B_ABST
    Figure CN117301061B_ABST
Patent Text Reader

Abstract

The application discloses a robot turning and squatting stable standing control method and device, equipment and storage medium, obtains the body perception information of a target humanoid robot, including a target rotation angle when a turning task is performed and a target pelvis height when a squatting task is performed; the body perception information is input into a strategy network to obtain an output action, the strategy network is trained in a simulator by using reinforcement learning, a reward function includes a first function for encouraging the robot to squat to a specified pelvis height, a second function for encouraging the movement speed of the pelvis in the z-axis direction during the squatting process to be close to an expected squatting speed, and a third function for encouraging the robot to turn to a specified angle; and the output action is converted into joint torque by a controller to control the action of the robot. The simulation training avoids the loss of the robot hardware equipment. The strategy network with better performance can be obtained by using the reinforcement learning training, and the stable standing control of the robot turning and squatting process is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control technology, and more specifically, to a method, device, equipment, and storage medium for controlling a robot's turning, squatting, and stable standing. Background Technology

[0002] With the application and development of robotics technology, robots are gradually penetrating from the industrial field into service, entertainment, education, and other areas. Currently, humanoid robots are a popular research area. Humanoid robots, through program control, can perform pre-programmed actions.

[0003] Traditional motion control for humanoid robots typically employs conventional algorithms such as PID control. During the testing phase, these algorithms need to be verified and optimized on a real machine. This process inevitably leads to adverse effects on the robot due to algorithmic flaws, sometimes even causing irreversible damage to the robot's hardware, resulting in a significant waste of hardware resources. Furthermore, traditional algorithms are not very effective, especially during robot turns and crouching, failing to maintain stable standing. Summary of the Invention

[0004] In view of the above problems, this application is made to provide a method, device, equipment, and storage medium for controlling the stability of a robot's turning and squatting movements, so as to achieve control over the stability of a humanoid robot's standing position during turning and squatting. The specific solution is as follows:

[0005] Firstly, a method for controlling a robot's turning, squatting, and stable standing is provided, including:

[0006] Acquire the body perception information of the target humanoid robot, including the target rotation angle when performing a turning task and the target pelvic height when performing a squatting task;

[0007] The proprioceptive information is input into the policy network trained in the simulator to obtain the n-dimensional action output by the policy network, where n is equal to the number of joint degrees of freedom of the target humanoid robot. The policy network is trained in the simulator using a reinforcement learning method. The input of the policy network includes the proprioceptive information, and the reward function includes: a first function that encourages the robot to squat to a specified pelvic height, a second function that encourages the robot to squat with the pelvis moving at a speed close to the set desired squatting speed in the z-axis direction, and a third function that encourages the robot to turn to a specified rotation angle.

[0008] The n-dimensional motion is converted into joint torques by a proportional-derivative (PD) controller, and the corresponding joint motors are driven according to the joint torques to perform the turning and squatting tasks.

[0009] Preferably, the training process of the policy network in the simulator includes:

[0010] Import the target humanoid robot's universal robot description format urdf file into the simulator;

[0011] Configure a policy network and an evaluation network for reinforcement learning. Set the input of the policy network to include the ontology-aware information, and set the input of the evaluation network to include the ontology-aware information and privileged information, wherein the privileged information is other observable information different from the ontology-aware information.

[0012] The reward function for reinforcement learning includes: the first function, the second function, and the third function.

[0013] Preferably, the privileged information includes terrain height information.

[0014] Preferably, the ontology-aware information includes:

[0015] The linear velocity of the robot body in the x, y, and z directions, the angular velocity of the robot body in the x, y, and z directions, the three-dimensional gravity vector, the linear velocity command of the robot body in the x, y, and z directions, the difference between the current position and the default joint position of each joint of the target humanoid robot, the velocity of each joint of the target humanoid robot, and the predicted position of each joint output at the previous time step.

[0016] Preferably, the n-dimensional actions output by the policy network include: the relative position changes of each of the n joints relative to the default joint position.

[0017] Preferably, the reward function for reinforcement learning further includes any one or more of the following reward functions, which are weighted and summed to form the total reward function:

[0018] The fourth function is used to encourage the robot to maintain a given standing posture during the turning process;

[0019] The fifth function is used to encourage the robot to keep both feet in the same posture;

[0020] The sixth function is used to encourage the robot to align its pelvic position with its feet in the x-direction.

[0021] The seventh function is used to encourage the robot to maintain a set linear velocity in the x and y directions.

[0022] The eighth function is used to encourage the robot to approach the set angular velocity command of the robot body;

[0023] The ninth function is used to encourage the robot to choose the action for the next time step with smaller action differences;

[0024] The tenth function is used to encourage the robot to maintain a stable joint speed;

[0025] The eleventh function is used to encourage the robot to remain stable and still after completing the turning and crouching tasks;

[0026] The twelfth function is used to encourage the robot to move within the set joint limit positions;

[0027] The thirteenth function is used to encourage robots to perform actions with less joint torque;

[0028] The fourteenth function is used to encourage the robot to keep its feet on the ground;

[0029] The fifteenth function is used to encourage the robot to reduce its speed along the z-axis.

[0030] The sixteenth function is used to encourage the robot to stand still until the end of the round.

[0031] Preferably, in the simulator, the training process is simulated by using a domain randomization method to randomize the dynamic parameters of a set type.

[0032] And / or,

[0033] A thrust is applied to a random position of the target humanoid robot, and a random thrust is applied to the torso position of the target humanoid robot;

[0034] And / or,

[0035] Noise is added to the ontological perception information.

[0036] Preferably, the dynamic parameters of the set type include any one or more of the following: ground friction coefficient, joint friction coefficient, target humanoid robot torso mass, torso center of mass position, and motor strength.

[0037] Secondly, a robot turning, squatting, and stable standing control device is provided, including:

[0038] The body perception information acquisition unit is used to acquire the body perception information of the target humanoid robot. The body perception information includes the target rotation angle when performing a turning task and the target pelvic height when performing a squatting task.

[0039] The policy network computing unit is used to input the ontology perception information into the policy network trained in the simulator to obtain the n-dimensional action output by the policy network, where n is equal to the number of joint degrees of freedom of the target humanoid robot. The policy network is trained in the simulator using a reinforcement learning training method. The input of the policy network includes the ontology perception information, and the reward function includes: a first function that encourages the robot to squat down to the target pelvic height, a second function that encourages the robot to squat down with the pelvic motion speed in the z-axis direction close to the set desired squat speed, and a third function that encourages the robot to turn to the target rotation angle.

[0040] The motion execution unit is used to convert the n-dimensional motion into joint torques through a proportional-derivative (PD) controller, and drive the corresponding joint motors according to the joint torques to perform the turning task and the squatting task.

[0041] Thirdly, a robot turning, squatting, and standing stably control device is provided, including: a memory and a processor;

[0042] The memory is used to store programs;

[0043] The processor is used to execute the program to implement each step of the aforementioned robot turning and squatting stable standing control method.

[0044] Fourthly, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the various steps of the aforementioned robot turning and squatting stable standing control method.

[0045] Using the above technical solution, this application employs a reinforcement learning training method to train the policy network in a simulator. The network input includes the proprioceptive information of the target humanoid robot, such as the specified rotation angle when performing a turning task and the specified pelvic height when performing a squatting task. Further reinforcement learning reward functions can include a first function to encourage the robot to squat to the specified pelvic height, a second function to encourage the robot's pelvic motion speed in the z-axis direction to approach the set desired squatting speed, and a third function to encourage the robot to turn to the specified rotation angle. By setting these reward functions, the policy network can be trained in the simulator, enabling it to control the robot to perform turning and squatting tasks effectively and maintain stable standing during the task process. After completing the simulation training of the policy network, it can be deployed in a real environment to obtain the proprioceptive information of the target humanoid robot. This proprioceptive information is then input into the simulation-trained policy network to obtain the n-dimensional motion output by the policy network. A proportional-derivative (PD) controller converts the n-dimensional motion into joint torques, which drive the corresponding joint motors to perform the turning and squatting tasks. This embodiment trains the policy network in a simulator using reinforcement learning, with the robot's proprioceptive information as input. A predefined reward function is used to train the policy network through reinforcement learning. The simulated-trained policy network can be deployed to a real-world environment without requiring algorithm optimization testing on the robot in a real-world setting, thus avoiding wear and tear on the robot's hardware. Furthermore, the reinforcement learning training method allows for trial and error in the simulation environment, with rewards used as feedback for reinforcement learning. This leads to a higher-performance policy network, enabling stable standing control during robot turns and crouching. Attached Figure Description

[0046] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0047] Figure 1 This is a schematic flowchart of a robot turning and squatting stable standing control method provided in an embodiment of this application;

[0048] Figure 2 An example is provided: a schematic diagram of a robot's body coordinate system.

[0049] Figure 3 A schematic diagram illustrating a simulation training process based on the PPO reinforcement learning algorithm is provided.

[0050] Figure 4A schematic diagram of a robot turning and squatting stable standing control device provided in an embodiment of this application;

[0051] Figure 5 This is a schematic diagram of a robot turning, squatting, and standing stably control device provided in an embodiment of this application. Detailed Implementation

[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0053] This application provides a control scheme for a humanoid robot to stand stably during the turning and squatting process. It can be applied to the process control of various humanoid robots performing turning and squatting tasks, ensuring that the robot's torso can stand stably during the turning and squatting process, and avoiding shaking, tilting or even falling.

[0054] The proposed solution can be implemented based on a humanoid robot control device. This control device can be a controller installed inside the robot, or other control terminals that communicate with the humanoid robot, such as handheld control terminals or servers.

[0055] This application employs a simulation training method to pre-train the policy network for controlling the robot. Reinforcement learning is used in the simulator to train the policy network. The trained policy network is then deployed in a real-world environment to control the motion of the target humanoid robot, specifically to achieve stable standing control during the robot's turning and crouching processes.

[0056] In reinforcement learning, the state space S is the set of all states, and the action space A is the set of all action ranges. At each time step t, the environment state is S. t The robot selects an action A t To change the state of the environment, the environment provides the next state S. t+1 And reward R t+1 and discount factor γ t+1 The discount factor, γ∈[0,1], is used to calculate the discount on the reward. The robot then selects its next action A based on the feedback. t+1 The probability distribution of actions in each state is denoted by policy π, and the robot selects actions based on policy π. t |s t ) indicates that in state s t Choose action a tThe probability, the transition function P(s) t+1 |s t a t ) indicates that in state s t Select action a t Under the condition of obtaining the next state s t+1 The probability, r(s) t a t ) indicates that in state s t The reward for choosing action 'a'.

[0057] This embodiment can employ various types of reinforcement learning methods, such as reinforcement learning methods based on PPO (Proximal Policy Optimization), etc.

[0058] The principle of reinforcement learning methods is as follows:

[0059] During the training phase, the algorithm's goal is to maximize the reward obtained by the policy, i.e.:

[0060]

[0061] Where θ is the network parameter, π θ This represents a strategy with network parameters. This represents the expected reward obtained from the strategy.

[0062] Taking the PPO algorithm as an example, the advantage function can be defined as the advantage of action a relative to the average action, that is:

[0063]

[0064] Where V is the value used to evaluate the network output. This represents the state value obtained by the strategy at time t+1.

[0065] The policy gradient update method solves the optimization objective through continuous iteration. The gradient is obtained, thus yielding the optimal parameters θ of the network:

[0066]

[0067] Furthermore, since the policy gradient update method requires online resampling for each parameter update, it has an unavoidable drawback: slow update speed. After a large batch of sampling and updates, the data is discarded. Resampling again for the next update is inefficient and results in low data utilization. Therefore, the importance sampling technique can be introduced into the PPO algorithm, using another distribution π2 to approximate the original distribution π1, thereby reusing sampling experience and improving update efficiency. The policy gradient is then expressed as:

[0068]

[0069]

[0070] The introduced distribution must approximate the original distribution, therefore an implicit constraint needs to be added:

[0071]

[0072] Based on the reinforcement learning algorithms introduced above, the following will combine... Figure 1 The robot turning and crouching stable standing control method described in this application may include the following steps:

[0073] Step S100: Obtain the body perception information of the target humanoid robot. The body perception information includes the target rotation angle when performing the turning task and the target pelvic height when performing the squatting task.

[0074] Specifically, proprioceptive information can be acquired through various types of sensors on the target humanoid robot. This proprioceptive information can be partially observable information. To control the target humanoid robot's turning and squatting tasks, the proprioceptive information can be set to include the target rotation angle specified when the robot performs a turning task, and the target pelvic height specified when the robot performs a squatting task.

[0075] Optionally, in addition to the two types of information mentioned above, the ontology-aware information may also include other observable information, including but not limited to the following:

[0076] (1) The linear velocity of the fuselage in the x, y, and z directions;

[0077] (2) Angular velocities of the fuselage in the x, y, and z directions;

[0078] (3) Commands for three-dimensional gravity vector and linear velocity of the fuselage in the x, y, and z directions;

[0079] (4) The difference between the current position of each joint of the target humanoid robot and the default joint position;

[0080] (5) The speed of each joint of the target humanoid robot;

[0081] (6) The predicted positions of each joint output in the previous time step;

[0082] (7) Two-dimensional clock information, etc.

[0083] Combination Figure 2 The example illustrates a schematic diagram of a robot's body coordinate system. The body coordinate system follows a right-handed perspective, with the vertically upward direction as the z-axis and the direction the robot's head faces as the x-axis.

[0084] Step S110: Input the ontology perception information into the policy network trained in the simulator to obtain the action output by the policy network. The policy network is trained in the simulator using reinforcement learning method, and its input includes ontology perception information and the reward function includes a set function.

[0085] Specifically, as mentioned above, this application can pre-train the policy network in a simulator using reinforcement learning methods. The input to the policy network includes the robot's proprioceptive information, as described in the previous step. Furthermore, the reward function for reinforcement learning can include multiple functions. To achieve the target humanoid robot's stable turning and crouching tasks, the target task and stable standing can be optimized by setting reward incentives. The reward function can include:

[0086] First function r h Used to encourage the robot to squat to a specified pelvic height:

[0087]

[0088] Among them, h cur h represents the robot's current pelvic height. ref This indicates the desired pelvic height that the robot will reach when squatting.

[0089] The second function r pv : Used to encourage the robot's pelvic movement speed along the z-axis to closely approximate the set desired squatting speed during the squatting process:

[0090]

[0091] Among them, pv cur pv represents the current velocity of the pelvis along the z-axis. ref This represents the desired squatting speed that the robot will achieve, when the difference between the pelvic height and the desired height exceeds a threshold h. thr Rewards are executed at specific times, where h thr The value can be set by the user, for example, 0.005m.

[0092] The third function r torso Used to encourage the robot to turn to a specified rotation angle.

[0093]

[0094] Where, j torso j represents the current rotation angle of the robot's torso. ref This indicates the desired angle the robot will turn to.

[0095] The reward functions in the above example can be weighted and summed to obtain the total reward function, and reinforcement learning training can be performed according to the total reward function.

[0096] The output action of the policy network can include an n-dimensional action, where n equals the number of degrees of freedom of the target humanoid robot. This embodiment provides an optional structure for the humanoid robot, which can include 19 joint degrees of freedom: three at the hips, one at the knees, one at the ankles, and one at the waist for each leg; three at the shoulders and one at the elbows for each arm, totaling 19 joints. Of course, the number of robot joints can be increased or decreased based on this.

[0097] Step S120: The actions output by the strategy network are converted into joint torques by the proportional-derivative PD controller, and the corresponding joint motors are driven according to the joint torques to perform the turning and squatting tasks.

[0098] In one alternative approach, the n-dimensional action output by the policy network can specifically be the relative position changes of each of the n joints relative to the default joint position. For each dimension of the output action, the sum of the output action and the default joint position is the target joint position of the target humanoid robot.

[0099] The target joint position can be converted into joint torque τ using a PD controller.

[0100] τ=KP×(pos target -pos current )-KD×vel current

[0101] Where KP is the proportional adjustment coefficient, used to reflect the proportion of the difference between the expected target and the actual target in the algorithm output. KD represents the derivative adjustment coefficient, which counteracts the motion oscillations caused by KP, and pos target Indicates the target joint position, pos curent Indicates the current joint position, vel current This indicates the current joint velocity.

[0102] The method provided in this application employs reinforcement learning to train a policy network in a simulator. The network input includes the target humanoid robot's proprioceptive information, such as a specified rotation angle when performing a turning task and a specified pelvic height when performing a squatting task. Further reinforcement learning reward functions can include a first function encouraging the robot to squat to the specified pelvic height, a second function encouraging the robot's pelvic motion speed in the z-axis direction to approach the set desired squatting speed, and a third function encouraging the robot to turn to the specified rotation angle. By setting these reward functions, the policy network can be trained in the simulator, enabling it to control the robot to perform turning and squatting tasks effectively and maintain stable standing during the task process. After completing the simulation training of the policy network, it can be deployed in a real environment to acquire the target humanoid robot's proprioceptive information. This proprioceptive information is then input into the simulation-trained policy network to obtain the n-dimensional motion output by the policy network. A proportional-derivative (PD) controller converts the n-dimensional motion into joint torques, which drive the corresponding joint motors to perform the turning and squatting tasks. This embodiment trains the policy network in a simulator using reinforcement learning, with the robot's proprioceptive information as input. A predefined reward function is used to train the policy network through reinforcement learning. The simulated-trained policy network can be deployed to a real-world environment without requiring algorithm optimization testing on the robot in a real-world setting, thus avoiding wear and tear on the robot's hardware. Furthermore, the reinforcement learning training method allows for trial and error in the simulation environment, with rewards used as feedback for reinforcement learning. This leads to a higher-performance policy network, enabling stable standing control during robot turns and crouching.

[0103] like Figure 3 The example illustrates a simulation training process based on the PPO reinforcement learning algorithm.

[0104] Combination Figure 3 The training process of the policy network in the simulator is introduced:

[0105] First, import the target humanoid robot's universal robot description format urdf file into the simulator to build the simulation environment.

[0106] Furthermore, when using the PPO reinforcement learning algorithm, it is necessary to configure the policy network and the evaluation network, and set the inputs for the policy network and the evaluation network respectively.

[0107] The policy network and the evaluation network can both be structured as a multilayer perceptron (MLP) with multiple hidden layers. For example, with three hidden layers, the number of nodes in each layer can be 512, 256, or 128, respectively. The policy network and the evaluation network can have the same structure but do not share parameters.

[0108] In this embodiment, the input to the policy network can be set as ontology-aware information. The input to the evaluation network can be supplemented with privileged information used only in simulation training. That is, the input to the policy network is only partially observable information, while the input to the evaluation network has more observable information than the input to the policy network.

[0109] Optionally, when training the robot to turn around, the input specified rotation angle can be a random value within a specified angle range, such as a random value in [-π, π]. When training the robot to squat, the input specified pelvic height can be a random value within a specified height range, such as a random value in [-0.65, 1.06].

[0110] In one alternative example, the privileged information may include terrain height information. By adding this terrain height information to the input of the evaluation network, the evaluation network can reward and incentivize the policy network based on more comprehensive observation information. This allows the policy network to be applied to terrain scenarios with different heights in real-world environments, i.e., the policy network is more robust.

[0111] After training in simulation, the policy network only needs to be deployed on a real device in a real environment. Therefore, no privileged information is required on the real device, reducing the difficulty of collecting observable information. Taking terrain height information as an example, no hardware device for collecting terrain height information needs to be set up on the real device, reducing equipment costs and machine complexity, while still being applicable to terrain scenarios with different heights in real environments.

[0112] Furthermore, the reward function for reinforcement learning includes at least the first function, the second function, and the third function described in the foregoing embodiments.

[0113] In some embodiments of this application, in order to achieve better stable standing control of the robot during the process of performing turning and squatting tasks, additional reward functions can be set, as shown in the following examples:

[0114] Fourth function r still This is used to encourage the robot to maintain a given standing posture during the turning process. During the robot's turn, only the leg joints are used to constrain the lower body to remain standing throughout the turn.

[0115]

[0116] Here, pos represents all leg joint positions, such as the hip, knee, and ankle joints. default This indicates the default standing leg joint position.

[0117] Fifth function r ankle This is used to encourage robots to keep both feet in the same posture.

[0118] r ankle =-|ankle l -ankle r |

[0119] Among them, ankle l with ankle r These indicate the joint positions of the left and right ankles, respectively.

[0120] The sixth function r stab It is used to constrain the standing stability of the robot. By calculating the difference between the pelvic position and the position of the feet in the x-direction, it avoids the robot from tilting forward or backward, that is, it encourages the robot to keep the pelvic position close to the position of the feet in the x-direction.

[0121]

[0122] Where, pos pelvis The x-axis position of the pelvis is represented by pos. lfoot and pos rfoot These represent the x-direction positions of both feet.

[0123] The seventh function r v It is used to track the linear velocity of the robot body, so that the robot can get close to the given linear velocity command, that is, to encourage the robot to get close to the set linear velocity command of the robot body in the x and y directions.

[0124]

[0125] in, This represents the linear velocity of the fuselage in the x and y directions in the fuselage coordinate system. Commands representing the linear velocity of the fuselage in the x and y directions in the fuselage coordinate system.

[0126] Eighth function r ω It is used to track the angular velocity of the robot body, so that the robot can get close to the given angular velocity command, that is, to encourage the robot to get close to the set angular velocity command of the robot body.

[0127]

[0128] Where ω base ω represents the fuselage angular velocity in the fuselage coordinate system. com This command represents the fuselage angular velocity in the fuselage coordinate system.

[0129] Ninth function r a It is used to constrain the robot's motion changes and prevent the motion changes from being too rapid, that is, to encourage the robot to choose the action of the next time step with smaller motion differences.

[0130] r a =-||a n -a l || 2

[0131] Among them, a n Indicates the action selected at the current time step, a l This indicates the action selected at the previous time step.

[0132] The tenth function r acc It is used to constrain the joint acceleration of a robot, prevent the robot from making drastic changes in movement when completing a task, and maintain a stable joint speed, that is, to encourage the robot to maintain a stable joint speed.

[0133]

[0134] Where v n Indicates the current joint velocity, v l dt represents the joint velocity at the previous time step, and dt represents the time interval.

[0135] The eleventh function r v It is used to constrain the speed of robot joints, encouraging the robot to remain stable and still after completing a specified task, that is, to encourage the robot to remain stable and still after completing the tasks of turning and squatting.

[0136]

[0137] Among them, v thr This represents the speed threshold, which can be a small value, such as 0.1 or other values.

[0138] The twelfth function r lim It is used to constrain the extreme positions of robot joints to prevent damage to the hardware caused by joint movement to the extreme position, that is, to encourage the robot to move within the set joint extreme positions.

[0139]

[0140] in, Indicates the lowest limit position of the joint. This indicates the highest limit position of the joint.

[0141] The thirteenth function r torqueIt is used to constrain the torque of a robot to prevent excessive torque from increasing energy consumption, that is, to encourage the robot to perform actions with smaller joint torque.

[0142] r torque =-||τ|| 2

[0143] The fourteenth function r no_fly It is used to keep the robot's feet from leaving the ground, and to determine whether it is stepping on the ground by the ground contact force, that is, to encourage the robot's feet to stay on the ground.

[0144]

[0145] in, and F represents the ground contact force of the left foot and the right foot, respectively. thr This represents the contact force threshold, which can be a small value, such as 0.1 or other values.

[0146] The fifteenth function r z This is used to penalize the speed of the z-axis in the body coordinate system, preventing the robot from shaking up and down, that is, encouraging the robot to reduce the speed of the z-axis.

[0147]

[0148] in This represents the velocity along the z-axis of the fuselage in the fuselage coordinate system, which follows the right-hand rule, with the vertically upward direction as the z-axis and the direction the fuselage nose faces as the x-axis.

[0149] The sixteenth function is used to encourage the robot to stand steadily until the end of the round. If the robot stands steadily until the end of the round, it is given a positive reward; otherwise, it is given a negative penalty if the robot falls down.

[0150]

[0151] Among them, reset_buf indicates whether the state needs to be reset. It is 1 only when the robot falls or the round length is greater than the set value. episode_length indicates the round duration. max_length indicates the set round duration threshold, which can be equal to or slightly less than the duration of a complete round.

[0152] The total reward can be defined as the weighted sum of all the above rewards, as shown in the following formula:

[0153] r = w1·r v +…+w 16 ·r termin

[0154] Among them, the weight coefficients w1,…,w16 These are hyperparameters that can be adjusted according to the actual task.

[0155] By weighted summation of multiple rewards and using the total reward for reinforcement learning training, the target humanoid robot can achieve stable standing when performing turning and squatting tasks.

[0156] In some embodiments of this application, considering that the modeling of the target humanoid robot's URDF may have errors, and that it is usually difficult to accurately model the physical characteristics of the real environment, in order to effectively transfer the target humanoid robot's stable turning and squatting functions from the simulation environment to the real world, this embodiment can also actively add interference during the simulation training process to reduce the modeling error from simulation to the real world.

[0157] An alternative way to add disturbance is to use a domain randomization method to randomize the dynamic parameters of a set type, so as to randomize the simulation data to cover the real data distribution.

[0158] In this embodiment, the randomized dynamic parameters include, but are not limited to: ground friction coefficient, joint friction coefficient, target humanoid robot torso mass, torso center of mass position, motor strength, etc.

[0159] Another alternative way to add perturbation is to apply a thrust to a random position of the target humanoid robot during the simulation training process, or to apply a random thrust to the torso position of the target humanoid robot, in order to simulate the random external perturbations that the target humanoid robot may be subjected to in the real world.

[0160] Another optional way to add noise is to add noise to the ontology perception information input to the policy network and the evaluation network, including but not limited to: adding noise to the fuselage linear velocity, fuselage angular velocity, gravity vector, linear velocity command, joint position, joint velocity, and the joint position predicted in the previous moment.

[0161] The above treatments can improve the robustness of the target humanoid robot.

[0162] The simulation training process of this application can be based on The system was powered by a Gold 5118 CPU at 2.30GHz and an NVIDIA Tesla V100 PCIe 32GB CPU. Simulation training was performed using the PPO algorithm in an Isaac Gym simulator. This application allows for parallel training with multiple (e.g., 4096) simulated robots. The experimental terrain was set to flat ground, with 1000 time steps per round, a total of 2000 rounds of training, and a learning rate of 1e-3. Other key experimental configuration parameters are exemplified below:

[0163] Table 1

[0164] Discount factor γ 0.99 Exponential weighting average λ 0.95 Clipping parameter 0.2 Value loss coefficient 1.0 Policy entropy loss coefficient 0.01 Expected KL divergence 0.01 Gradient descent mini batch 4

[0165] When deployed on a real machine, only the policy network (Actor) is deployed, and the target humanoid robot performs actions using only its own perception information. Because domain randomization and observation noise are added during simulation training, it can still maintain good performance when ported from simulation to the real world, and can achieve stable standing during turning and squatting.

[0166] The robot turning and squatting stable standing control device provided in the embodiments of this application is described below. The robot turning and squatting stable standing control device described below can be referred to in correspondence with the robot turning and squatting stable standing control method described above.

[0167] See Figure 4 , Figure 4 This is a schematic diagram of a robot turning and squatting stable standing control device disclosed in an embodiment of this application.

[0168] like Figure 4 As shown, the device may include:

[0169] The body perception information acquisition unit 11 is used to acquire the body perception information of the target humanoid robot. The body perception information includes the target rotation angle when performing the turning task and the target pelvic height when performing the squatting task.

[0170] The strategy network computing unit 12 is used to input the ontology perception information into the strategy network trained in the simulator to obtain the n-dimensional action output by the strategy network, where n is equal to the number of joint degrees of freedom of the target humanoid robot. The strategy network is trained in the simulator using a reinforcement learning training method. The input of the strategy network includes the ontology perception information, and the reward function includes: a first function that encourages the robot to squat down to the target pelvic height, a second function that encourages the robot to squat down with the pelvic movement speed in the z-axis direction close to the set desired squat speed, and a third function that encourages the robot to turn to the target rotation angle.

[0171] The motion execution unit 13 is used to convert the n-dimensional motion into joint torques through a proportional-derivative PD controller, and drive the corresponding joint motors according to the joint torques to perform the turning task and the squatting task.

[0172] Optionally, the apparatus of this application may further include:

[0173] A simulation training unit is used to perform reinforcement learning training on the policy network in a simulator, the process of which may include:

[0174] Import the target humanoid robot's universal robot description format urdf file into the simulator;

[0175] Configure a policy network and an evaluation network for reinforcement learning. Set the input of the policy network to include the ontology-aware information, and set the input of the evaluation network to include the ontology-aware information and privileged information, wherein the privileged information is other observable information different from the ontology-aware information.

[0176] The reward function for reinforcement learning includes: the first function, the second function, and the third function.

[0177] Optionally, the process of the simulation training unit performing reinforcement learning training on the policy network in the simulator further includes:

[0178] The dynamic parameters of a given type are randomized using a domain randomization method;

[0179] And / or,

[0180] A thrust is applied to a random position of the target humanoid robot, and a random thrust is applied to the torso position of the target humanoid robot;

[0181] And / or,

[0182] Noise is added to the ontological perception information.

[0183] The robot turning and crouching stable standing control device provided in this application embodiment can be applied to robot turning and crouching stable standing control equipment. Optionally, Figure 5 The hardware structure block diagram of the robot's turning, squatting, and stable standing control device is shown. (Refer to...) Figure 5 The hardware structure of the device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;

[0184] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;

[0185] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0186] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0187] The memory stores a program, which the processor can call. The program is used for:

[0188] Acquire the body perception information of the target humanoid robot, including the target rotation angle when performing a turning task and the target pelvic height when performing a squatting task;

[0189] The proprioceptive information is input into the policy network trained in the simulator to obtain the n-dimensional action output by the policy network, where n is equal to the number of joint degrees of freedom of the target humanoid robot. The policy network is trained in the simulator using a reinforcement learning method. The input of the policy network includes the proprioceptive information, and the reward function includes: a first function that encourages the robot to squat to a specified pelvic height, a second function that encourages the robot to squat with the pelvis moving at a speed close to the set desired squatting speed in the z-axis direction, and a third function that encourages the robot to turn to a specified rotation angle.

[0190] The n-dimensional motion is converted into joint torques by a proportional-derivative (PD) controller, and the corresponding joint motors are driven according to the joint torques to perform the turning and squatting tasks.

[0191] Optionally, the refined and extended functions of the program can be found in the description above.

[0192] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:

[0193] Acquire the body perception information of the target humanoid robot, including the target rotation angle when performing a turning task and the target pelvic height when performing a squatting task;

[0194] The proprioceptive information is input into the policy network trained in the simulator to obtain the n-dimensional action output by the policy network, where n is equal to the number of joint degrees of freedom of the target humanoid robot. The policy network is trained in the simulator using a reinforcement learning method. The input of the policy network includes the proprioceptive information, and the reward function includes: a first function that encourages the robot to squat to a specified pelvic height, a second function that encourages the robot to squat with the pelvis moving at a speed close to the set desired squatting speed in the z-axis direction, and a third function that encourages the robot to turn to a specified rotation angle.

[0195] The n-dimensional motion is converted into joint torques by a proportional-derivative (PD) controller, and the corresponding joint motors are driven according to the joint torques to perform the turning and squatting tasks.

[0196] Optionally, the refined and extended functions of the program can be found in the description above.

[0197] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0198] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0199] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for controlling a robot's turning, squatting, and stable standing, characterized in that, include: Acquire the body perception information of the target humanoid robot, including the target rotation angle when performing a turning task and the target pelvic height when performing a squatting task; The proprioceptive information is input into a policy network trained by reinforcement learning in a simulator to obtain an n-dimensional action output by the policy network, where n equals the number of joint degrees of freedom of the target humanoid robot. The policy network is trained using reinforcement learning in the simulator, with a dedicated reward function. The input to the policy network includes the proprioceptive information, and the reward function includes: a first function to encourage the robot to squat to a specified pelvic height; a second function to encourage the robot to squat with the pelvis moving at a speed close to the desired squatting speed along the z-axis; and a third function to encourage the robot to turn to a specified rotation angle. The reward function also... The reward function includes one or more of the following: a fourth function to encourage the robot to maintain a given standing posture during a turn; a fifth function to encourage the robot to maintain the same posture for both feet; a sixth function to encourage the robot to bring its pelvis close to its feet in the x-direction; a ninth function to encourage the robot to choose the action for the next time step with less movement difference; a tenth function to encourage the robot to maintain a stable joint velocity; an eleventh function to encourage the robot to remain stable and still after completing the turn and squat tasks; a fourteenth function to encourage the robot to keep its feet on the ground; and a sixteenth function to encourage the robot to remain standing stably until the end of the round. The n-dimensional motion is converted into joint torques by a proportional-derivative (PD) controller, and the corresponding joint motors are driven according to the joint torques to perform the turning and squatting tasks and maintain a stable standing position throughout the process.

2. The method according to claim 1, characterized in that, The training process of the policy network in the simulator includes: Import the target humanoid robot's universal robot description format urdf file into the simulator; Configure a policy network and an evaluation network for reinforcement learning. Set the input of the policy network to include the ontology-aware information, and set the input of the evaluation network to include the ontology-aware information and privileged information, wherein the privileged information is other observable information different from the ontology-aware information. The reward function for reinforcement learning includes: the first function, the second function, and the third function.

3. The method according to claim 2, characterized in that, The privileged information includes terrain height information.

4. The method according to claim 1, characterized in that, The ontology-aware information includes: The linear velocity of the robot body in the x, y, and z directions, the angular velocity of the robot body in the x, y, and z directions, the three-dimensional gravity vector, the linear velocity command of the robot body in the x, y, and z directions, the difference between the current position and the default joint position of each joint of the target humanoid robot, the velocity of each joint of the target humanoid robot, and the predicted position of each joint output at the previous time step.

5. The method according to claim 1, characterized in that, The n-dimensional actions output by the policy network include the relative position changes of each of the n joints relative to the default joint position.

6. The method according to claim 2, characterized in that, The reward function for reinforcement learning also includes any one or more of the following reward functions, which are weighted and summed to form the total reward function: The seventh function is used to encourage the robot to maintain a set linear velocity in the x and y directions. The eighth function is used to encourage the robot to approach the set angular velocity command of the robot body; The twelfth function is used to encourage the robot to move within the set joint limit positions; The thirteenth function is used to encourage robots to perform actions with less joint torque; The fifteenth function is used to encourage the robot to reduce its speed along the z-axis.

7. The method according to claim 2, characterized in that, The training process is simulated in the simulator, and the dynamic parameters of a set type are randomized using a domain randomization method. And / or, A thrust is applied to a random position of the target humanoid robot, and a random thrust is applied to the torso position of the target humanoid robot; And / or, Noise is added to the ontological perception information.

8. The method according to claim 7, characterized in that, The dynamic parameters of the set type include any one or more of the following: ground friction coefficient, joint friction coefficient, target humanoid robot torso mass, torso center of mass position, and motor strength.

9. A robot turning, squatting, and stable standing control device, characterized in that, include: The body perception information acquisition unit is used to acquire the body perception information of the target humanoid robot. The body perception information includes the target rotation angle when performing a turning task and the target pelvic height when performing a squatting task. The policy network computation unit is used to input the ontology perception information into a policy network trained by reinforcement learning in a simulator, and obtain the n-dimensional action output by the policy network, where n is equal to the number of joint degrees of freedom of the target humanoid robot. The policy network is trained using reinforcement learning in the simulator, and a dedicated reward function is set. The input to the policy network includes the ontology perception information, and the reward function includes: a first function to encourage the robot to squat to the target pelvic height; a second function to encourage the robot to squat with the pelvis moving at a speed close to the set desired squatting speed along the z-axis; and a third function to encourage the robot to turn to the target rotation angle. The reward function also includes any one or more of the following reward functions: a fourth function to encourage the robot to maintain a given standing posture during the turning process; a fifth function to encourage the robot to maintain the same posture for both feet; a sixth function to encourage the robot to keep its pelvis close to its feet in the x-direction; a ninth function to encourage the robot to choose the action for the next time step with a smaller difference in movement; a tenth function to encourage the robot to maintain a stable joint velocity; an eleventh function to encourage the robot to remain stable and still after completing the turning and squatting tasks; a fourteenth function to encourage the robot to keep its feet on the ground; and a sixteenth function to encourage the robot to stand stably until the end of the round. The motion execution unit is used to convert the n-dimensional motion into joint torques through a proportional-derivative (PD) controller, and drive the corresponding joint motors according to the joint torques to perform the turning task and the squatting task while maintaining a stable standing position throughout the process.

10. A robot turning, squatting, and stable standing control device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the robot turning and squatting stable standing control method as described in any one of claims 1 to 8.

11. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the robot turning and crouching stable standing control method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Biped robot gait control method and control device

    CN113467235A

  • Robot motion control method and system and electronic equipment

    CN116619382A