Motion control method for three-foot walking gait of quadruped robot
Through the framework of the teacher strategy network and the student strategy network, combined with prior knowledge of mechanics and reinforcement learning, we independently identify damaged legs and generate stable three-legged walking gait, which solves the stability problem of the four-legged robot under single leg damage and achieves stable three-legged movement under special operating conditions.
Patent Information
- Application Number
- CN202510443879.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Existing four-legged robots are difficult to achieve stable three-legged movements under single-leg damage. The existing methods have high complexity and poor adaptability in modeling. The learning-based methods rely on a large amount of data and do not fully consider the principles of mechanics, resulting in limited stability and efficiency.
The framework of teacher strategy network and student strategy network is adopted, combined with mechanical prior knowledge and reinforcement learning, and the damaged legs are automatically identified. By generating reference foot end positions and real-time optimization, a stable three-leg walking gait is generated using a recurrent neural network, and the optimization algorithm and reward function are combined to ensure the stability of the robot.
It realizes stable three-legged movement of the four-legged robot under any single leg damage, reduces data dependence, improves the stability and adaptability of three-legged movement, and provides key technical support for control under special operating conditions.
Smart Images

Figure CN120363183A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot control, and particularly relates to a motion generation method for a quadruped robot that can autonomously identify a damaged leg and achieve stable three-legged motion. Background Art
[0002] Quadruped robots have received extensive attention due to their excellent terrain adaptability. However, quadruped robots require special walking gaits under special working conditions. For example, quadruped robots may face the situation of a single leg damage in the military field and civil special working conditions, and manual intervention and repair may not be carried out in a timely and effective manner. At this time, the quadruped robot needs to use the remaining three legs to complete the task. Therefore, the three-legged motion of quadruped robots is an important challenge in the field of quadruped robot motion control.
[0003] Existing quadruped robot motion control methods mainly include model-based methods and learning-based methods. In the case of a single leg damage, model-based methods usually rely on separately modeling different damage situations to achieve three-legged motion control. This method not only has a high modeling complexity but also poor adaptability and is difficult to be extended to multiple damage scenarios. On the other hand, learning-based methods train control strategies in a data-driven manner and have been proposed in some patents. For example, patents CN202410971117.8 and CN202411048913.0 disclose learning-based three-legged motion control methods. Existing learning methods mostly rely on a large amount of data for training and do not fully consider the mechanical principles in three-legged motion, resulting in a training process that depends on a large amount of data and the stability and efficiency are limited in practical applications. Therefore, there is an urgent need for a quadruped robot control method that can reduce data dependence while integrating mechanical prior knowledge and adaptively achieve stable three-legged motion. Summary of the Invention
[0004] The present invention aims to solve the challenge of achieving stable three-legged motion of a quadruped robot in the case of a single leg damage in the prior art, and provides a method that integrates mechanical prior knowledge and learning, and can autonomously identify a damaged leg and achieve a motion generation method for a quadruped robot with stable three-legged motion.
[0005] The object of the present invention can be achieved by the following technical solutions:
[0006] A motion control method for a quadruped robot with a three-legged walking gait, the steps include:
[0007] Obtain the gait information of the quadruped robot walking on three legs in the case of a single leg damage;
[0008] Generate a reference foot end position through a foot end reference position generator, and perform real-time optimization on the reference foot end position through a teacher policy network;
[0009] Control the quadruped robot based on the optimized foot-end reference position, and use the motion data of the quadruped robot in different three-legged walking gaits to interactively train the teacher policy network through a reinforcement learning algorithm;
[0010] Train the student policy network using supervised learning to approximate the teacher policy network;
[0011] Deploy the trained student policy network on the physical robot, collect the historical state information of the quadruped robot, input it into the student policy network, and output the corresponding joint control commands of the quadruped robot.
[0012] As a preferred technical solution, the input of the teacher policy network is the state information of the robot at multiple moments, and the output is the foot-end reference position compensation command; the input of the student policy network is the state information of the robot at multiple moments, and the output is the position command of the joint motor.
[0013] As a preferred technical solution, the state information of the robot at each moment input into the teacher policy network and the student policy network includes: the robot speed, the pitch, roll, and yaw angles and angular velocities of the fuselage, the contact states of the four foot-ends with the ground, and the angles and angular accelerations of the various joint motors of the robot.
[0014] As a preferred technical solution, the teacher policy network and the student policy network adopt a recurrent neural network with the same structure. The recurrent neural network adopts a combined network structure of GRU plus MLP. The GRU network contains a hidden layer for extracting the temporal features in the observation sequence; the MLP part consists of three fully connected networks to output the control commands of the joints of the quadruped robot.
[0015] As a preferred technical solution, the specific process of the interactive training of the teacher policy network of the reinforcement learning is as follows:
[0016] Obtain the gait information of the three-legged walking of the quadruped robot in the case of a single leg damage;
[0017] Input the joint angle information of the quadruped robot in the three-legged gait at the start moment of each phase into the foot-end reference position generation module, and output the foot-end reference position within the future phase time through the foot-end reference position generation formulas of the stance phase and the swing phase;
[0018] Input the state information of the robot at multiple moments into the teacher policy network, and output the foot-end reference position compensation command;
[0019] Use the output of the teacher policy network to compensate the foot-end reference position, and control the three-legged movement of the quadruped robot based on the compensated foot-end coordinates;
[0020] Interactively train the teacher policy network using the motion data of the quadruped robot under different terrains and different three-legged walking gaits.
[0021] As a preferred technical solution, the specific process of the foot-end reference position generation module generating the foot-end reference position within the future phase time is as follows:
[0022] For the three supporting legs of the quadruped robot in the three-legged gait, at the beginning of each phase time, it is set that within the future phase time, two supporting legs are in the support phase and one supporting leg is in the swing phase; at the end of the future phase time, a support triangle is formed by the projection points of the feet of the three supporting legs on the horizontal plane;
[0023] It is set that the movement of the supporting leg in the support phase is the foot moving backward, and the coordinates of the projection points of the foot-end reference positions of the two supporting legs in the support phase at the end of the future phase time are set as fixed values according to engineering experience;
[0024] With the goal of maximizing the minimum distance from the projection point of the center of gravity of the quadruped robot on the horizontal plane to the three sides of the support triangle, solve for the projection point of the reference position of the foot of the supporting leg in the swing phase at the end of the phase time, that is, the best foothold. The specific objective function is:
[0025]
[0026] In the formula, d1, d2, and d3 are the distances from the projection point of the center of gravity at the end of the future phase time to the three sides of the support triangle respectively;
[0027] Use the optimization algorithm to solve for the foot-end reference positions of the three supporting legs at the end of the future phase time;
[0028] Generate the trajectory within the corresponding phase time according to the foot-end reference position at the end of the future phase time, where the foot-end of the trajectory of the supporting leg in the support phase is a straight line, and the foot-end trajectory of the supporting leg in the swing phase is a sine trajectory; take points on the reference trajectory at a set frequency to obtain a series of optimal foot-end reference positions in the future phase time.
[0029] As a preferred technical solution, when solving for the best foothold of the supporting leg in the swing phase, the following constraint conditions need to be satisfied:
[0030] The projection point (x g , y g ) of the center of gravity of the quadruped robot is within the support triangle formed by the projection points (x1, y1), (x2, y2), (x3, y3) of the feet of the three supporting legs:
[0031]
[0032] ω = 1 - u - v
[0033] u > 0, v > 0, ω > 0
[0034] According to the mechanical structure parameters of the quadruped robot, the value range of the optimal foothold of the leg in the swing phase at the end of the phase time is defined.
[0035] As a preferred technical solution, during the training process of the teacher policy network: when resetting the robot state in each training episode, the quadruped robot simulates the walking motion of the quadruped robot in a three-legged gait by lifting any one leg above a specified height; taking the damaged leg and the effective support legs as known prior information, determining the joint data of the three support legs and the joint data of the damaged leg in the input joint position information, and generating the foot end reference position with the joint data of the effective support legs.
[0036] As a preferred technical solution, the proximal policy optimization (PPO) algorithm is adopted for the reinforcement learning algorithm, and the reward function is set to include: the forward item reward of the quadruped robot, the stability item reward of the quadruped robot, the reward for the quadruped robot to maintain the moving direction, and the stability reward for three-legged walking.
[0037] As a preferred technical solution, the setting of the reward function is specifically as follows:
[0038]
[0039] In the formula, ω1, ω2, ω3, ω4, ω5, ω6, ω7 are the coefficients of each reward item, ω1, ω2, ω7 are positive, and ω3, ω4, ω5, ω6 are negative; v x is the forward speed of the robot, v y is the lateral speed of the robot; ω z is the yaw angular velocity of the robot; v z is the vibration speed of the body up and down; h is the body height, g z is the gravity vector projected onto the robot body coordinate system; ω x is the roll angular velocity of the robot, ω y is the pitch angular velocity of the robot, d min is the minimum distance from the center of gravity projection point to the three sides of the support triangle. When the center of gravity projection point is inside the support triangle, d min is positive. When the center of gravity projection point is outside the support triangle, d min is negative, is for v x , v y , ω z , v z , h are the expected values.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] 1) In the present invention, a reference foot-end position generator is designed in combination with the mechanical principle of three-legged stable movement to generate a series of reference foot-end positions. Reinforcement learning is introduced into the control framework to optimize the reference foot-end positions in real time. The optimized foot-end positions are converted into joint motor angles through inverse kinematics to achieve the control of the quadruped robot. Based on the teacher and student policy network structures of the recurrent neural network, the quadruped robot can implicitly and autonomously identify the damaged leg and implement the corresponding control of the quadruped robot. It can enable the robot to achieve stable three-legged movement in the case of any single leg damage, providing key technical support for the three-legged movement control of the quadruped robot under special working conditions.
[0042] 2) The present invention combines the mechanical principle of three-legged stable movement and the optimization algorithm to design a formula for generating the reference foot positions of the three-legged movement of a quadruped robot. At the beginning of each phase time, in the future phase time, two legs are in the support phase and one leg is in the swing phase. By setting the reference foot-end positions in the support phase and optimizing to solve the optimal best foothold of the support leg in the swing phase, and considering the mechanical principle of three-legged stable movement during the solution, the center of gravity projection point of the robot is within the support triangle, and the minimum distance from the center of gravity projection point to the three sides is the largest. A set of foot-end reference positions that fully consider three-legged stability can be obtained, providing a good prior trajectory for the entire training framework and reducing the training time.
[0043] 3) The reinforcement learning reward function set in the present invention includes the forward item reward, the reward for maintaining the movement direction, and the reward for maintaining the movement direction of the quadruped robot, and a stability reward item for the three-legged walking gait of the quadruped robot is designed in combination with the mechanical principle of three-legged stable movement, ensuring that the center of gravity projection point is within the support triangle formed by the three foot-end projection points on the horizontal plane, walking forward while ensuring the stability of the fuselage, and improving the stability of the three-legged movement. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is the overall flowchart of a method for generating the three-legged walking gait movement of a quadruped robot provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0045] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manner and specific operation process are given, but the protection scope of the present invention is not limited to the following embodiments.
[0046] Embodiment 1
[0047] In view of the situation where a single leg of a quadruped robot is damaged, the present invention proposes a motion control method for realizing stable three-legged walking of a quadruped robot. The present invention combines the mechanical principle of stable three-legged motion and an optimization algorithm to design a foot-end reference position generator considering the stability of three-legged motion to generate a series of foot-end reference positions for the three-legged walking gait, and introduces reinforcement learning into the control framework to optimize the foot-end reference position in real time. The optimized foot-end position is converted into joint motor angles through inverse kinematics to realize the control of the quadruped robot. The present invention adopts a teacher-student strategy framework, and the student strategy approximates the teacher strategy network through supervised learning to realize the deployment of the strategy network on a real quadruped robot. The teacher and student strategies adopt a strategy network based on a gated recurrent unit (GRU), and the damaged leg is implicitly and autonomously identified through the recurrent neural network to realize stable three-legged motion in the case of any damaged leg on a real quadruped robot. The present invention can enable the robot to achieve stable three-legged motion in the case of any single damaged leg, providing key technical support for the three-legged motion control of quadruped robots under special working conditions.
[0048] As Figure 1 shown, the method proposed by the present invention specifically includes the following steps:
[0049] S1. Build a quadruped robot model and a terrain environment in a simulation environment.
[0050] Specifically, in this embodiment, the simulation environment selects Pybullet, and the quadruped robot model selects Unitree A1. This robot has a total of 12 joint motors, and each leg has three joint motors, corresponding to the body joint, hip joint, and knee joint respectively. To improve the adaptability of the robot on real ground, the terrain environment settings mainly include flat ground and slightly rough terrain.
[0051] S2. Design a simulation mechanism for the three-legged gait of a quadruped robot in the case of a single damaged leg: Lift any leg of the quadruped robot above a specified height to simulate the three-legged gait walking motion of the quadruped robot, and use this mechanism in the simulation environment to train the three-legged motion of the quadruped robot.
[0052] Specifically, in this embodiment, when resetting the robot state in each episode, randomly lift a leg above a height of 8 cm to simulate the three-legged motion in the case of a single damaged leg. The initial standing height of the robot is 28 cm. Lifting any leg in the simulation is known prior information, so the three effective supporting legs are known during the training of the teacher strategy network. During the training of the teacher strategy network, the damaged leg and the effective supporting legs are used as known prior information to determine the joint data of the three supporting legs and the joint data of the damaged leg in the input 12-dimensional joint position information. The 9 joint data of the effective supporting legs are used to generate the foot-end reference position.
[0053] S3. Design a control framework for the three - legged walking gait motion of a quadruped robot based on a reinforcement learning algorithm. The control framework mainly includes the following modules: the foot - end reference position generation module, the teacher policy based on a recurrent neural network, and the student policy based on a recurrent neural network. The input of the foot - end reference position generation module is the information of the 12 joint angles of the quadruped robot at the start of each phase, and the output is the foot - end reference position within the time of the next phase. The teacher policy based on a recurrent neural network takes the state information of the robot at 10 moments as input and outputs a 12 - dimensional foot - end reference position compensation command. The student policy based on a recurrent neural network takes the state information of the robot at 10 moments as input and outputs the position commands of the 12 joint motors. The state information of the robot at each moment includes: the robot speed, the pitch, roll, and yaw angles and angular velocities of the fuselage, the contact states of the four foot - ends with the ground, and the angles and angular accelerations of the 12 joint motors of the robot.
[0054] Among them, the recurrent neural network adopts a combined network structure of GRU plus MLP. The GRU network contains 1 hidden layer with 128 hidden units, which is used to extract the temporal features in the observation sequence. The MLP part consists of three fully - connected networks with an inter - layer dimension of 128*128*128, and the output is the position commands of the 12 joint motors of the four legs of the robot.
[0055] S4. Design the foot - end reference position generation formulas for the support phase and the swing phase in the foot - end reference position generator, and solve the future foot - end reference position by the gradient descent method.
[0056] Specifically: The three - legged gait in the present invention adopts a walking gait. At the start of each phase time, in the future phase time, two legs are in the support phase and one leg is in the swing phase. Project the reference positions of the foot - ends of the three support legs (three support points) after the end of the next phase time onto the horizontal plane to obtain three points, and the coordinates of the three points are (x1,y1), (x2,y3), (x3,y3). The coordinates of the projection point of the center of gravity onto this horizontal plane are (x g ,y g ). The motion in the support phase is the backward movement of the foot - end. After the end of the phase time, the reference positions (x1,y1), (x2,y2) of the foot - ends of the two support legs can be set as fixed values according to engineering experience, and solve the optimal foothold (x3,y3) of the leg in the swing phase at the end of the phase time. The solution of this foothold position needs to meet the following conditions: The center - of - gravity projection point (x g ,y g ) is within the support triangle formed by the three foot - end projection points (x1,y1), (x2,y2), (x3,y3), and the minimum distance from the center - of - gravity projection point to the three sides of the support triangle is the largest. Use an optimization algorithm to solve it.
[0057] The objective function is:
[0058]
[0059] Solve the coordinates of the center-of-gravity projection point through the reference robot state (robot state ) after the end of the phase time. The robot state (robot state ) can be solved through the positions of the four foot ends. The coordinate solution formula for the center-of-gravity projection point is as follows:
[0060] (x g , y g ) = f(robot state )
[0061] Calculate the distances from the center-of-gravity projection point (x g , y g ) to the three sides of the support triangle. The formula is as follows:
[0062]
[0063] The constraint conditions are as follows:
[0064] Constraint condition 1: Ensure that the center-of-gravity projection point 9x g , y g ) is within the support triangle formed by the three foot-end projection points (x1, y1), (x2, y2), and (x3, y3). The formula is as follows:
[0065]
[0066] ω = 1 - u - v
[0067] u > 0, v > 0, ω > 0
[0068] Constraint condition 2: Limit the value range of (x3, y3) according to the mechanical structure parameters of the quadruped robot. The formula is as follows:
[0069] x low ≤ x3 ≤ x high
[0070] y low ≤ y3 ≤ y high
[0071] According to the above formula, the optimal algorithm can be used to solve the reference positions of the three support leg foot ends at the end of the future phase time. Generate the trajectory during this phase time based on the reference positions of the foot ends. Among them, the support phase trajectory is a straight line, and the swing phase trajectory is a sine trajectory. The height of the sine trajectory is set to 3 cm. Take points on the reference trajectory at a frequency of 50 Hz to obtain a series of optimal reference positions of the foot ends in the future phase time.
[0072] S5. Train the teacher policy network through reinforcement learning. The output of the teacher policy network compensates for the foot-end reference position. The sum of the output of the teacher policy network and the foot-end reference position coordinates is the final foot-end coordinate position. The final foot-end coordinate position is converted into the angles of the joint motors through inverse kinematics. Finally, a PD controller is used to achieve the three-legged movement of the quadruped robot. Lifting different legs of the quadruped robot will result in different three-legged gaits. Collect the motion data of the quadruped robot under different terrains and different three-legged walking gaits for interactive training to improve the stability of the three-legged gait movement of the quadruped robot. Design a reward function for the three-legged movement, strengthen the robot stability term in the reward function, and the reward setting should ensure that the robot walks forward while ensuring the stability of the fuselage. The setting of the reward function is as follows:
[0073]
[0074] In the formula, r t is the total reward term, is the forward term reward of the quadruped robot, is the stability term reward of the quadruped robot, is the reward for the quadruped robot to maintain the moving direction, is the three-legged walking stability reward. Through the reward term, it is ensured that on the horizontal plane, the center of gravity projection point (x g , y g ) is within the support triangle formed by the three foot-end projection points (x1, y1), (x2, y2), (x3, y3), and the minimum distance from the center of gravity projection point to the three sides of the support triangle is the largest.
[0075]
[0076] In the formula, ω1, ω2, ω3, ω4, ω5, ω6, ω7 are the coefficients of each reward term. ω1, ω2, ω7 are positive, while ω3, ω4, ω5, ω6 are negative. v x is the forward speed of the robot, v y is the lateral speed of the robot, ω z is the yaw angular velocity of the robot, v z is the vibration speed of the fuselage up and down, h is the height of the fuselage, g z is the gravity vector projected onto the robot's body coordinate system, ω x is the roll angular velocity of the robot, ω y is the pitch angular velocity of the robot, d min is the minimum distance from the center of gravity projection point to the three sides of the support triangle. When the center of gravity projection point is inside the support triangle, d min is positive. When the center of gravity projection point is outside the support triangle, d min is negative. are respectively vx , v y , ω z , v z , the expected value of h
[0077] During the training process, domain randomization is performed on the observation space and some physical parameters. The trained network is the teacher policy network. The reinforcement learning algorithm uses the Proximal Policy Optimization (PPO) algorithm. To reduce the error between the simulation environment and the ground in the real environment, parameter randomization of the friction coefficient is performed in the simulation. To reduce the error between the robot model in the simulation and the real robot, parameter randomization of the center of gravity position of the robot is performed in the simulation. To reduce the error caused by the sensor noise of the robot, parameter randomization of the observation space of the policy network is performed in the simulation.
[0078] S6. Use the method of supervised learning to train the student network to approximate the teacher policy network. The trained student network is the final network deployed on the robot. The present invention uses the method of imitation learning to train the student policy network. The GRU network of the student policy can autonomously identify the damaged leg according to the temporal information and realize the corresponding three-legged movement.
[0079] S7. Deploy the student policy network on the physical robot, collect the historical state information of the quadruped robot, and input it into the student policy. According to the historical state information, the student policy network can autonomously identify the faulty leg and output the joint control commands of the corresponding quadruped robot to realize the stable control of the three-legged gait of the quadruped robot.
[0080] The present invention aims at the situation of a single damaged leg of a quadruped robot and proposes a motion control method for realizing the stable three-legged walking of the quadruped robot. The present invention combines the mechanical principle of stable three-legged movement and the optimization algorithm to design a foot-end reference position generator considering the stability of three-legged movement to generate a series of foot-end reference positions, and introduces reinforcement learning into the control framework to optimize the foot-end reference positions in real time. The optimized foot-end positions are converted into joint motor angles through inverse kinematics to realize the control of the quadruped robot. The present invention adopts a teacher-student policy framework. The student policy network approximates the teacher policy network through the method of supervised learning to realize the deployment of the policy network on the real quadruped robot. The teacher and student policies adopt a policy network based on the Recurrent Neural Network (GRU). The damaged leg is implicitly and autonomously identified through the recurrent neural network to realize the stable three-legged movement in the case of any damaged leg on the real quadruped robot. The present invention can enable the robot to realize stable three-legged movement in the case of any single-leg failure, providing key technical support for the three-legged movement control of the quadruped robot under special working conditions.
[0081] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A three - legged walking gait motion control method for a quadruped robot, characterized in that the steps Including: Obtain the gait information of a quadruped robot walking on three legs when one leg is damaged. Generate a reference foot end position through a foot end reference position generator, and perform real-time optimization on the reference foot end position through a teacher policy network. Control the quadruped robot based on the optimized reference foot end position, and use the motion data of the quadruped robot in different three-legged walking gaits to perform interactive training on the teacher policy network through a reinforcement learning algorithm. Train a student policy network using supervised learning to approximate the teacher policy network. Deploy the trained student policy network on a physical robot, collect the historical state information of the quadruped robot, input it into the student policy network, and output the corresponding joint control commands of the quadruped robot.
2. The three - legged walking gait motion control method of a quadruped robot according to claim 1, wherein, The input of the teacher policy network is the state information of the robot at multiple moments, and the output is the foot end reference position compensation command; the input of the student policy network is the state information of the robot at multiple moments, and the output is the position command of the joint motor.
3. A method for controlling the three-legged walking gait movement of a quadruped robot according to claim 2, characterized in that, The state information of the robot at each moment input into the teacher policy network and the student policy network includes: the robot speed, the pitch, roll, and yaw angles and angular velocities of the fuselage, the contact states of the four foot ends with the ground, and the angles and angular accelerations of the joints of the robot.
4. A quadruped robot three-legged walking gait motion control method according to claim 2, characterized in that, The teacher policy network and the student policy network adopt a recurrent neural network with the same structure. The recurrent neural network adopts a combined network structure of GRU plus MLP. The GRU network contains a hidden layer for extracting temporal features in the observation sequence; the MLP part consists of three fully connected networks, and outputs the control commands of the joints of the quadruped robot.
5. A three-legged walking gait motion control method for a quadruped robot according to claim 1, characterized in that The specific process of performing interactive training on the teacher policy network of the reinforcement learning is as follows: Obtain the gait information of a quadruped robot walking on three legs when one leg is damaged. Input the joint angle information of the quadruped robot in the three-legged gait at the start of each phase into the foot end reference position generation module, and output the foot end reference position within the future phase time through the foot end reference position generation formulas for the stance phase and the swing phase. Input the state information of the robot at multiple moments into the teacher policy network, and output the foot end reference position compensation command. Use the output of the teacher policy network to compensate the foot end reference position, and control the three-legged movement of the quadruped robot based on the compensated foot end coordinates. Perform interactive training on the teacher policy network using the motion data of the quadruped robot on different terrains and in different three-legged walking gaits.
6. A three-legged walking gait motion control method for a quadruped robot according to claim 5, characterized in that, The specific process of the foot end reference position generation module generating the foot end reference position within the future phase time is as follows: For the three supporting legs of the quadruped robot in the three-legged gait, at the start of each phase time, it is set that two supporting legs are in the stance phase and one supporting leg is in the swing phase within the future phase time; at the end of the future phase time, a support triangle is formed by the projection points of the foot ends of the three supporting legs on the horizontal plane. It is set that the movement of the supporting leg in the stance phase is the foot end moving backward, and the coordinates of the projection points of the foot end reference positions of the two supporting legs in the stance phase at the end of the future phase time are set as fixed values according to engineering experience. Taking the maximum of the minimum distances from the projection point of the center of gravity of the quadruped robot on the horizontal plane to the three sides of the support triangle as the goal, solve the projection point of the reference position of the foot end of the supporting leg in the swing phase at the end of the phase time on the horizontal plane, that is, the optimal foothold. The specific objective function is as follows: In the formula, d1, d2, and d3 are the distances from the projection point of the center of gravity at the end of the future phase time to the three sides of the support triangle respectively; Use the optimization algorithm to solve the reference positions of the foot ends of the three supporting legs at the end of the future phase time; Generate the trajectory within the corresponding phase time according to the reference position of the foot end at the end of the future phase time. Among them, the foot end of the supporting leg trajectory in the support phase is a straight line, and the foot end trajectory of the supporting leg trajectory in the swing phase is a sine trajectory; Take points on the reference trajectory at a set frequency to obtain a series of optimal reference positions of the foot end in the future phase time.
7. A three-legged walking gait motion control method for a quadruped robot according to claim 6, characterized in that When solving the optimal foothold of the supporting leg in the swing phase, the following constraints need to be satisfied: The center of gravity projection point (x g , y g ) of the quadruped robot is within the support triangle formed by the projection points (x1, y1), (x2, y2), and (x3, y3) of the three support leg feet: ω = 1 - u - v u > 0, v > 0, ω > 0 According to the mechanical structure parameters of the quadruped robot, limit the value range of the optimal foothold of the leg in the swing phase at the end of the phase time.
8. A method for controlling the gait movement of a quadruped robot walking on three legs according to claim 5, characterized in that, During the training process of the teacher policy network: When resetting the robot state in each training episode, the quadruped robot simulates the walking motion of the quadruped robot's three-legged gait by lifting any one leg above a specified height; Take the damaged leg and the effective supporting leg as known prior information, determine the joint data of the three supporting legs and the joint data of the damaged leg in the input joint position information, and generate the reference position of the foot end with the joint data of the effective supporting leg.
9. A method for controlling the three-legged walking gait motion of a quadruped robot according to claim 5, characterized in that, The reinforcement learning algorithm adopts the proximal policy optimization (PPO) algorithm, and the reward function is set to include: the forward item reward of the quadruped robot, the stability item reward of the quadruped robot, the reward for the quadruped robot to maintain the moving direction, and the stability reward for three-legged walking.
10. A three-legged walking gait motion control method for a quadruped robot according to claim 9, characterized in that, The setting of the reward function is specifically as follows: where ω1, ω2, ω3, ω4, ω5, ω6, ω7 are the coefficients of each reward item, ω1, ω2, ω7 are positive, and ω3, ω4, ω5, ω6 are negative; v x is the forward speed of the robot, v y is the lateral speed of the robot; ω z is the yaw angular velocity of the robot; v z is the vibration speed of the robot up and down; h is the body height, g z is the gravity vector projected onto the robot's body coordinate system; ω x is the roll angular velocity of the robot, ω y is the pitch angular velocity of the robot, d min is the minimum distance from the center of gravity projection point to the three sides of the support triangle. When the center of gravity projection point is inside the support triangle, d min is positive, and when the center of gravity projection point is outside the support triangle, d min is negative, is v x , v y , ω z , v z , h's expected value.
Citation Information
Patent Citations
Quadruped robot single motor fault tolerance control method based on reinforcement learning
CN119065240A
Blind hexapod robot motion strategy training method
CN117340876A
Quadruped robot visual motion control method based on forward kinematics
CN118838388A
Motion control method for quadruped robot with damaged legs
CN118915802A
Multi-agent reinforcement learning task planning and control method under space-time task driving
CN119472783A