Method for controlling anthropomorphic running action of humanoid robot
Through contact mode decomposition and gait cycle modeling, anthropomorphic running reference trajectory is generated, and combined with mirror processing and reinforced imitation learning of layered reward structures, the symmetry and anti-interference problems of humanoid robot running movements are solved, and the efficient control of the robot in complex environments is achieved.
Patent Information
- Application Number
- CN202510744090.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to effectively construct accurate reference trajectories, express contact patterns, and design stable learning structures, which leads to the inability to achieve natural symmetry and energy-saving characteristics of humanoid robots, which limits the maneuverability and robustness of robots in complex environments.
Through contact mode decomposition and gait cycle modeling, anthropomorphic running reference trajectory is generated, mirror keyframe processing and enhanced imitation learning of layered reward structures are adopted, and combined with PVT-PD joint motor control, the anti-interference ability and dynamic rationality of the robot's anthropomorphic running movement are realized.
It realizes the high symmetry and anti-interference ability of the robot's running movement, improves the maneuverability and robustness of the robot in complex environments, and supports online running control.
Smart Images

Figure CN120382498A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and specifically to a control method for anthropomorphic running actions of a humanoid robot. Background Art
[0002] The design and control of the running actions of humanoid robots have always been important topics in the field of robot control. Compared with walking, running has more complex dynamic characteristics, including asymmetric contact patterns, short flight phases, rapid center of gravity conversion, etc. Traditional methods mostly combine finite state machines with simple gait generators, and cannot fully reproduce the energy-saving, symmetric and stable characteristics in human running, which limits the mobility and robustness of the robot in complex environments.
[0003] In recent years, the rise of imitation learning and reinforcement learning (RL) has brought new opportunities for the natural motion control of robots. However, how to construct accurate reference trajectories, effectively express contact patterns, and design stable learning structures are still the key difficulties in the research of anthropomorphic running control.
[0004] Therefore, how to extract high-quality and generalizable running reference trajectories from human actions, and on this basis, design a set of stable and effective control strategies to enable humanoid robots to achieve running actions with natural symmetry and energy-saving characteristics is a technical problem that needs to be solved urgently at present, specifically as follows:
[0005] 1. How to model and express the complex contact states during running.
[0006] 2. How to perform spatio-temporal symmetrization processing on the asymmetric running trajectories.
[0007] 3. How to fuse imitation goals and machine state constraints in reinforcement learning to ensure training stability.
[0008] 4. How to implement implicit learning and generalization of contact patterns in the control strategy. Summary of the Invention
[0009] The present invention proposes a control method for anthropomorphic running actions of a humanoid robot, aiming to control the robot to perform anthropomorphic running actions with strong anti-interference ability, high dynamic rationality and close to the optimal solution in multiple dimensions while imitating human running actions.
[0010] To achieve the above object, the technical solution adopted by the present invention is:
[0011] A control method for anthropomorphic running actions of a humanoid robot, characterized in that it includes the following steps:
[0012] S1. Human expert running motion analysis
[0013] S2. Generation of anthropomorphic running reference trajectory
[0014] S21. Modeling of anthropomorphic running motion of robot and key frame design
[0015] S22. Symmetrization processing of key frame sequence, interpolation of position, velocity and torque trajectories of key frame sequence, and interpolation of attitude trajectory of key frame sequence
[0016] S23. Obtain the anthropomorphic running reference trajectory
[0017] S24. Generate an anthropomorphic running trajectory library by instruction extension
[0018] S3. Reinforcement imitation learning based on reference trajectory library
[0019] S31. Design of reinforcement imitation learning framework, including design of asymmetric AC network mechanism, design of reinforcement imitation learning reward function based on hierarchical progression, and reinforcement learning optimization method of proximal policy optimization
[0020] S32. Obtain the anthropomorphic running policy network of the robot
[0021] S4. Physical machine deployment of anthropomorphic running control strategy
[0022] S41. Connect the input and output of the policy network to the real machine program, receive sensing information, and generate control instructions
[0023] S42. PVT-PD joint motor control
[0024] S5. Realize the anthropomorphic running motion control of the robot
[0025] As a preferred technical solution of the present invention: Step S21 is specifically as follows
[0026] Annotate the contact state between the sole of the foot and the ground frame by frame for the running motion video of the human expert, and define the basic contact types
[0027] Each frame constitutes a left and right foot contact combination
[0028] C t =(Left, Right)
[0029] There are 16 states in total, and extract the periodic template sequence
[0030] {lfc-rff, ltc-rff, lff-rff, lff-rhc, lff-rfc, lff-rtc, lff-rff, lhc-rff}
[0031] Assign a time stamp t to each frame ii With state K i =(p i , q i , C i ).
[0032] As a preferred technical solution of the present invention: In step S22, the key-frame sequence symmetrization process is specifically as follows:
[0033] Let the original key-frame sequence be
[0034] {K1, K2,..., K N}
[0035] The mirror transformation is defined as:
[0036]
[0037] Match the contact states of all mirror key-frames and the original key-frames, and perform equal-weight averaging on the key-frame data of the key-frames with the same contact state in the original key-frames and the mirror key-frames to obtain the key-frame sequence after symmetrization processing:
[0038]
[0039] As a preferred technical solution of the present invention: In step S22, the key-frame sequence position-velocity-torque trajectory interpolation is specifically as follows:
[0040] For the obtained key-frame sequence Construct a continuous reference trajectory T ref (t) through time interpolation. Each key-frame includes the position p i ∈R 3 of the floating base, the attitude quaternion q i ∈S 3 and the joint angle information. In order to generate a temporally continuous and smooth reference trajectory, different interpolation methods are used for different types of variables.
[0041] The position information in the key-frame includes the position of the floating base and the joint positions. There is no correlation between the data of each dimension among these information. Therefore, cubic splines are used to interpolate each dimension.
[0042] As a preferred technical solution of the present invention: In step S22, the key-frame sequence attitude trajectory interpolation is specifically as follows:
[0043] Since the attitude is represented by the unit quaternion q i ∈S 3 during interpolation, it must be kept on the unit sphere. Therefore, spherical linear interpolation is used for processing, specifically as follows:
[0044] Between two adjacent quaternions q i and q i +1, the Slerp interpolation formula is as follows:
[0045]
[0046] Where:
[0047] θ = cos -1 (q i ·q i +1) represents the angle between the two quaternions;
[0048] α ∈ [0, 1] is the normalized time ratio;
[0049] The interpolation result q(t) always remains in the S 3 unit quaternion space, ensuring smooth and continuous rotation.
[0050] As a preferred technical solution of the present invention: In step S23, by interpolating the translation and rotation parts of all key frames respectively, a continuous reference trajectory is defined:
[0051]
[0052] The reference trajectories obtained under different instructions in step S23 are extended to obtain a reference running trajectory library
[0053]
[0054] As a preferred technical solution of the present invention: In step S31, the design of the asymmetric AC network mechanism is specifically as follows:
[0055] An asymmetric Actor-Critic network structure is adopted, where:
[0056] The policy network π θ (a ∨ s) outputs each step action
[0057] The value network evaluates the return of the state
[0058] Input state:
[0059]
[0060] Wherein, is the one-hot encoding of the contact mode,
[0061] The goal is to maximize the expected return:
[0062]
[0063] As a preferred technical solution of the present invention: in step S31, the design of the hierarchical progressive reinforcement imitation learning reward function is specifically as follows:
[0064] Hierarchical Reward function design
[0065] The Reward is designed as a four-layer nested structure:
[0066] Safety Reward r safe :
[0067] r safe = w1·exp(-||τ t || 2 ) + w2·1 not falling
[0068] Regularization Reward r regular :
[0069]
[0070] Instruction following Reward r cmd :
[0071]
[0072] Imitation Reward r mimic :
[0073]
[0074] Final combined Reward:
[0075] r total = r safe + σ(r safe )·(r regular + σ(r regular )·(r cmd + σ(r cmd )·r mimic ))
[0076] Among them, is the sigmoid function.
[0077] As a preferred technical solution of the present invention: in step S31, the optimization method of proximal policy optimization reinforcement learning is specifically as follows:
[0078] The Proximal Policy Optimization algorithm is used for policy optimization:
[0079] Policy objective function:
[0080]
[0081] Wherein:
[0082]
[0083] is the generalized advantage estimation
[0084] After training, the model πθ is exported and deployed to the real robot controller, and humanoid running can be realized.
[0085] As a preferred technical solution of the present invention: in step S41, the PVT-PD joint motor control is specifically as follows:
[0086] The bottom motor control adopts PD control, and the output is pure torque control. The torque command is directly sent to the motor for execution.
[0087]
[0088] where Kp is the proportional feedback coefficient, q mea is the measured joint position, q des is the desired joint position, K d is the differential feedback coefficient, is the measured joint speed.
[0089] Compared with the prior art, the beneficial effects of the present invention are:
[0090] 1. Introduce contact mode and mirror processing to achieve highly symmetric humanoid running;
[0091] 2. The key frame generation mechanism and interpolation method ensure the spatio-temporal continuity of the trajectory;
[0092] 3. The reward hierarchy improves training stability and multi-objective balance;
[0093] 4. The PPO+AC reinforcement imitation learning greatly improves the anti-interference ability of the robot compared with the traditional model-based control, enables the robot to overcome the problem of insufficient model accuracy, and is applicable to real machine deployment;
[0094] 5. Support input of any speed command to achieve online running control. Description of the Drawings
[0095] Figure 1 is the overall framework diagram of the present invention;
[0096] Figure 2 is the definition diagram of the contact type between the sole of the foot and the ground in the present invention;
[0097] Figure 3 is Figure 2 the mirror transformation diagram of. Detailed Embodiments
[0098] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments:
[0099] As Figure 1 shown, a control method for the anthropomorphic running action of a humanoid robot proposed by the present invention includes the following steps:
[0100] S1. Analysis of the running action of a human expert
[0101] S2. Generation of an anthropomorphic running reference trajectory
[0102] S21. Modeling of the anthropomorphic running action of the robot and design of key frames,
[0103] S22. Symmetrization processing of the key frame sequence, interpolation of the position, speed, and torque trajectories of the key frame sequence, and interpolation of the attitude trajectory of the key frame sequence,
[0104] S23. Obtain the anthropomorphic running reference trajectory,
[0105] S24. Generate an anthropomorphic running trajectory library by command extension;
[0106] S3. Reinforcement imitation learning based on the reference trajectory library
[0107] S31. Design of the reinforcement imitation learning framework, including the design of an asymmetric AC network mechanism, the design of a reinforcement imitation learning reward function based on hierarchical progression, and a reinforcement learning optimization method for proximal policy optimization,
[0108] S32. Obtain the anthropomorphic running policy network of the robot;
[0109] S4. Physical deployment of the anthropomorphic running control strategy
[0110] S41. Connect the input and output of the policy network to the real machine program, receive sensing information, and generate control commands,
[0111] S42. PVT-PD joint motor control;
[0112] S5. Implement the anthropomorphic running action control of the robot.
[0113] As Figure 2 shown, contact mode template and key frame extraction,
[0114] Step S21 is specifically as follows: Frame by frame, annotate the contact state between the sole of the foot and the ground in the running action video of the human expert, and define the basic contact types,
[0115] Each frame constitutes a left and right foot contact combination:
[0116] C t =(Left, Right)
[0117] A total of 16 states, extract the periodic template sequence:
[0118] {lfc-rff, ltc-rff.lff-rff, lff-rhc, lff-rfc, lff-rtc, lff-rff, lhc-rff}
[0119] Assign a timestamp t to each frame i i With state K i =(p i , q i , C i ).
[0120] As Figure 3 shown, in step S22, the symmetrization process of the key frame sequence is as follows:
[0121] Let the original key frame sequence be
[0122] {K1, K2,..., K N}
[0123] The mirror transformation is defined as:
[0124]
[0125] Perform contact state matching on all mirror key frames and the original key frames, and perform equal-weight averaging on the key frame data with the same contact state in the original key frames and the mirror key frames to obtain the symmetrized key frame sequence:
[0126]
[0127] In step S22, the position, velocity, and torque trajectory interpolation of the key frame sequence is as follows:
[0128] For the obtained key frame sequence Construct a continuous reference trajectory T ref (t) through time interpolation. Each key frame contains the position p of the floating base i ∈R 3 , the attitude quaternion q i ∈S 3 and joint angle information. In order to generate a temporally continuous and smooth reference trajectory, different interpolation methods are used for different types of variables.
[0129] The position information in the key frame includes the position of the floating base and the joint positions. There is no correlation between the data in each dimension of these information. Therefore, cubic splines are used to interpolate each dimension.
[0130] Taking the position of the floating base as an example, given two adjacent frames p i and p i +1 and their corresponding times t i and t i+1 , we use the spline function s(t) to satisfy the following conditions:
[0131] s(t i ) = p i , s(t i+1 ) = p i+1 , s′(t) is continuous, s″(t) is continuous
[0132] Cubic Spline Interpolation is a commonly used interpolation method that extends key-frame data with piecewise cubic polynomials:
[0133] p(t) = [s x (t), s y (t), s z (t)] 1 , t ∈ [t i , t i+1
[0134] where S i (t) = a i + b i (t - t i ) + c i (t - t i ) 2 + d i (t - t i ) 3 , which can ensure the continuity of the first and second derivatives of the function at the joints of each segment.
[0135] This interpolation can ensure that the trajectory is continuous and smooth in space, which is beneficial to the tracking and stability of the control system.
[0136] In step S22, the interpolation of the key-frame sequence attitude trajectory is as follows:
[0137] Since the attitude is represented by the unit quaternion q i ∈ S 3 , it must be kept on the unit sphere during interpolation. Therefore, spherical linear interpolation is used for processing, as follows:
[0138] Between two adjacent quaternions q i and q i +1, the Slerp interpolation formula is as follows:
[0139]
[0140] where:
[0141] θ = cos -1 (q i ·q i + 1) represents the angle between two quaternions;
[0142] α ∈ [0, 1] is the normalized time ratio;
[0143] The interpolation result q(t) always remains in the S 3 unit quaternion space, ensuring smooth continuous rotation.
[0144] In step S23, by interpolating the translation and rotation parts of all key frames respectively, a continuous reference trajectory is defined:
[0145]
[0146] Expand the reference trajectories under different instructions obtained in step S23 to obtain a reference running trajectory library
[0147]
[0148] This continuous trajectory can be used for the motion controller to perform tracking control or as the target state sequence in imitation learning.
[0149] In step S31, the design of the asymmetric AC network mechanism is as follows:
[0150] Adopt an asymmetric Actor-Critic network structure, where:
[0151] Policy network π θ (a ∨ s) outputs the action for each step
[0152] Value network Evaluates the return of the state
[0153] Input state:
[0154]
[0155] Among them, is the one-hot encoding of the contact mode,
[0156] The goal is to maximize the expected return:
[0157]
[0158] In step S31, the design of the reward function based on hierarchical progressive reinforcement imitation learning is as follows:
[0159] Hierarchical Reward function design
[0160] Reward is designed as a four - layer nested structure:
[0161] Safety Rewardr safe :
[0162] r safe = w1·exp(-||τ t || 2 )+ w2·1 not falling
[0163] Regularized Rewardr regular :
[0164]
[0165] Instruction - following Reward r cmd :
[0166]
[0167] Imitation Reward r mimic :
[0168]
[0169] Final combined Reward:
[0170] r total = r safe + σ(r safe )·(r regular + σ(r regular )·(r cmd + σ(r cmd )·r mimic ))
[0171] Wherein, is the sigmoid function.
[0172] In step S31, the reinforcement learning optimization method of proximal policy optimization is specifically as follows:
[0173] Use the Proximal Policy Optimization algorithm for policy optimization:
[0174] Policy objective function:
[0175]
[0176] Wherein:
[0177]
[0178] is the generalized advantage estimation
[0179] After training, the model πθ is exported and deployed to the real robot controller, and humanoid running can be achieved.
[0180] In step S41, the PVT-PD joint motor control is specifically as follows:
[0181] The underlying motor control uses PD control, and the output is pure torque control. The torque command is directly sent to the motor for execution.
[0182]
[0183] where Kp is the proportional feedback coefficient, q mea is the measured joint position, q des is the desired joint position, K d is the differential feedback coefficient, is the measured joint speed.
[0184] The core of the present invention lies in solving the problem that the humanoid robot tilts or even falls when performing single-leg action switching. The present invention models the human expert demonstration action as key frame data through contact mode decomposition and gait cycle modeling; through the spatio-temporal symmetry processing of the mirror key frames, it ensures the rationality and simplicity of the robot imitating the reference trajectory, and avoids the robot learning bad actions; through the hierarchical reward structure of reinforcement imitation learning, it realizes the anti-interference and dynamic optimality of the running action.
[0185] The above is only a preferred embodiment of the present invention, and it is not any other form of limitation to the present invention. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.
Claims
1. A control method for a humanoid robot's anthropomorphic running motion, characterized in that: It includes the following steps: S1. Analysis of the running actions of human experts S2. Generation of a reference trajectory for anthropomorphic running S21. Modeling of the anthropomorphic running actions of the robot and design of key frames S22. Symmetrization processing of the key frame sequence, interpolation of the position, speed, and torque trajectories of the key frame sequence, and interpolation of the pose trajectory of the key frame sequence S23. Obtain a reference trajectory for anthropomorphic running S24. Generate an anthropomorphic running trajectory library by expanding the instructions S3. Reinforcement imitation learning based on the reference trajectory library S31. Design of the reinforcement imitation learning framework, including the design of an asymmetric AC network structure, the design of a reinforcement imitation learning reward function based on hierarchical progression, and a reinforcement learning optimization method for proximal policy optimization S32. Obtain a robot anthropomorphic running policy network S4. Actual deployment of the anthropomorphic running control strategy S41. Connect the input and output of the policy network to the real machine program, receive sensing information, and generate control instructions S42. PVT-PD joint motor control S5. Realize the control of the anthropomorphic running actions of the robot 2. The control method for the anthropomorphic running action of a humanoid robot according to claim 1, characterized in that: Step S21 is specifically as follows: Frame by frame, annotate the contact state between the sole of the foot and the ground in the running action video of the human expert, and define the basic contact types Each frame forms a left and right foot contact combination: C t =(Left, Right) There are a total of 16 states, and a periodic template sequence is extracted: Plfc-rff, ltc-rff, lff-rff, lff-rhc, lff-rfc, lff-rtc, lff-rff, lhc-rff Assign a timestamp t to each frame i i with state K i =(p i , q i , C i ).
3. The control method for the anthropomorphic running motion of a humanoid robot according to claim 1, characterized in that: In step S22, the symmetrization processing of the key frame sequence is specifically as follows: Let the original key-frame sequence be {K1, K2,..., K N} The mirror transformation is defined as: Match the contact states of all mirror key frames with the original key frames, and perform equal-weight averaging on the key frame data of the key frames with the same contact state in the original key frames and the mirror key frames to obtain a symmetrized key frame sequence 4. A control method for the anthropomorphic running motion of a humanoid robot according to claim 1, characterized in that: In step S22, the interpolation of the position, speed, and torque trajectories of the key frame sequence is specifically as follows: For the obtained key frame sequence Construct a continuous reference trajectory T ref (t) through time interpolation. Each key frame contains the position p i ∈R 3 of the floating base, the attitude quaternion q i ∈S 3 and joint angle information. To generate a temporally continuous and smooth reference trajectory, different interpolation methods are used for different types of variables The position information in the key frames includes the position of the floating base and the joint positions. Since there is no correlation between the data of each dimension among these information, cubic splines are used for interpolation in each dimension 5. The control method for the anthropomorphic running motion of a humanoid robot according to claim 1 or 4, characterized in that: In step S22, the interpolation of the pose trajectory of the key frame sequence is specifically as follows: Since the attitude is represented by the unit quaternion q i ∈S 3 When interpolating, it must be kept on the unit sphere. Therefore, spherical linear interpolation is used for processing, as follows: Between two adjacent quaternions qi and qi+1, the Slerp interpolation formula is as follows: Where: θ = cos -1 (qi·qi + 1) represents the angle between two quaternions; α ∈ [0, 1] is the normalized time ratio The interpolation result q(t) always remains within S 3 On the unit quaternion space, it ensures smooth and continuous rotation.
6. The control method of the anthropomorphic running action of a humanoid robot according to claim 1, characterized in that: In step S23, by interpolating the translation and rotation parts of all key frames respectively, a continuous reference trajectory is defined Expand the reference trajectories under different instructions obtained in step S23 to obtain a reference running trajectory library 7. A control method for an anthropomorphic running motion of a humanoid robot according to claim 1, characterized in that: In step S31, the design of the asymmetric AC network structure is specifically as follows: Adopt an asymmetric Actor-Critic network structure, where: Policy network π θ (a ∨ s) outputs the action at each step Value network Return on assessment status Input state: Among them, is the one-hot encoding of the contact mode, The goal is to maximize the expected return:
8. The control method for the anthropomorphic running motion of a humanoid robot according to claim 1, characterized in that: In step S31, the design of the reinforcement imitation learning reward function based on hierarchical progression is specifically as follows: Hierarchical Reward function design The Reward is designed as a four-layer nested structure: Safety Reward r safe : r safe = w1·exp(-||τ t || 2 ) + w2·1 not falling Regularized Reward r regular : Instruction Follow Reward r cmd : Imitate Reward r mimic : Final combined Reward: r total = r safe + σ(r safe )·(r regular + σ(r regular )·(r cmd + σ(r cmd )·r mimic )) where, is the sigmoid function.
9. A control method for the anthropomorphic running actions of a humanoid robot according to claim 1, characterized in that: In step S31, the reinforcement learning optimization method for proximal policy optimization is specifically as follows: Use the Proximal Policy Optimization algorithm for policy optimization: Policy objective function: Where: For generalized advantage estimation Export the trained model π θ , and deploy it to the real robot controller to achieve humanoid running.
10. A control method for anthropomorphic running actions of a humanoid robot according to claim 1, characterized in that: In step S41, the PVT-PD joint motor control is specifically as follows: The underlying motor control uses PD control, and the output is pure torque control. The torque command is directly sent to the motor for execution. where $K_p$ is the proportional feedback coefficient, $q$ mea is the measured joint position, $q$ des is the desired joint position, $K$ d is the derivative feedback coefficient, and is the measured joint velocity.