Quadruped robot motion control method and system based on reinforcement learning
By using the CVaR-PPO algorithm and risk assessment module to train the four-legged robot control strategy in the ISAAC GYM simulation environment, the problem of insufficient motion stability and safety of the four-legged robot in complex terrain is solved, and stronger environmental adaptability and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510337875.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
The existing reinforcement learning control methods have problems in the application of four-legged robots, lack of consideration of extreme risk scenarios, and insufficient generalization ability on complex terrain.
By building a four-legged robot simulation environment in the ISAAC GYM simulation environment, the control strategy is trained using the CVaR-PPO algorithm, and combining the risk assessment module and risk sensitivity factor, the control strategy is dynamically adjusted to ensure the safety and stability of the robot in complex terrain.
It improves the motion stability and safety of four-legged robots in complex terrain, enhances the environmental adaptability and generalization capabilities of control strategies, and reduces the cost and loss of hardware experiments.
Smart Images

Figure CN120178683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of motion control, and in particular, to a motion control method and system for a quadruped robot based on reinforcement learning. Background Art
[0002] As a kind of bionic robot, quadruped robots have received extensive attention in recent years in the fields of autonomous walking in complex environments, rescue operations, military reconnaissance, etc. Traditional control methods for quadruped robots mainly rely on preset motion patterns and control strategies based on dynamic modeling, such as gait control methods based on reflex arcs, model predictive control (MPC) methods, and optimization control methods. These methods perform well in specific terrains and known environments, but for complex, variable, and unknown terrain environments, their adaptability and robustness are relatively weak. In the field of quadruped robots, artificial intelligence technologies represented by deep reinforcement learning can break through the limitations of traditional robotics theories, avoiding problems such as the need for precise dynamic and kinematic modeling, professional knowledge in electromechanics, and complex and cumbersome manual tuning in traditional motion control, enabling quadruped robots to learn optimal control strategies. However, standard reinforcement learning algorithms still face many challenges in practical applications, such as low training efficiency, lack of consideration for extreme risk scenarios, and insufficient generalization ability on complex terrains.
[0003] The existing reinforcement learning control methods have the following deficiencies in the application of quadruped robots: (1) On complex terrains, traditional reinforcement learning methods mainly optimize the average return and are difficult to handle high-risk terrains, such as steep slopes and gravel ground, resulting in poor stability of the robot in extreme situations; (2) The reinforcement learning method based on the standard PPO algorithm lacks the ability to model risks, making the robot prone to high-risk behaviors during training, such as unstable gaits and falls, affecting actual deployment; (3) Most of the existing simulation training environments use simple terrains and do not fully cover the complex environments that quadruped robots may encounter in the real world, resulting in limited generalization ability of the trained strategies in the real environment. Summary of the Invention
[0004] The purpose of the present invention is to provide a motion control method and system for a quadruped robot based on reinforcement learning to solve the problems raised in the above background art.
[0005] To solve the above technical problems, the present invention provides the following technical solutions:
[0006] A motion control method for a quadruped robot based on reinforcement learning, the method comprising the following steps: Step S1: Use the reinforcement learning environment ISAAC GYM to build a simulation environment for the quadruped robot; Step S2: In the simulation environment, simulate the interaction process between the quadruped robot and the environment, obtain the action feedback of the robot on different terrains through multiple simulations, and record the interaction data with the environment; Step S3: Set the control strategy, and use the CVaR-PPO algorithm to train the control strategy in the simulation environment; Step S4: Verify the trained control strategy in the simulation environment, and after passing the verification, deploy the control strategy on the quadruped robot.
[0007] As a preferred solution of the motion control method for a quadruped robot based on reinforcement learning according to the present invention, the simulation environment includes different types of terrain scenarios, the terrain scenarios include slopes, steps and uneven ground, and the simulation environment is used to verify the stability and robustness of the control strategy.
[0008] In the present invention, the Isaac Gym framework of NVIDIA is used to build a simulation environment for the quadruped robot, and the quadruped robot model is imported by defining the robot URDF model. Define the characteristics of the terrain in the simulation environment, such as flat ground, slopes, obstacles, etc. Set the physical properties of the robot, such as the KP and KD parameters of joint control, torque limit, kinematic parameters, etc.
[0009] As a preferred solution of the motion control method for a quadruped robot based on reinforcement learning according to the present invention, the quadruped robot obtains terrain feature data, robot attitude data and joint state data in real time through the sensor module, and inputs the terrain feature data, robot attitude data and joint state data into the CVaR-PPO algorithm to train the control strategy.
[0010] In the present invention, in the constructed simulation environment, simulate the interaction process between the quadruped robot and the surrounding environment, obtain the action feedback of the robot on different terrains through multiple simulations, and record the interaction data with the environment. Use these data to model the motion characteristics of the robot in different terrains and formulate a motion control strategy adapted to different environments.
[0011] It should be noted that after the training is completed, the obtained control strategy is first verified multiple times in the simulation environment. After passing the simulation verification, it is finally deployed on the quadruped robot for testing to achieve the goal of automatic gait generation and posture adjustment.
[0012] As a preferred solution of the quadruped robot motion control method based on reinforcement learning according to the present invention, the CVaR-PPO algorithm is built-in with a risk assessment module, and the risk assessment module is used to evaluate the terrain risk in real time and dynamically adjust the control strategy according to the evaluation result.
[0013] It should be noted that in the simulation environment, the PPO reinforcement learning method improved based on CVaR is used for training. By introducing a risk assessment module, the risks of different terrains are evaluated in real time, and the control strategy is dynamically adjusted to ensure the safety and stability of the robot in complex terrains.
[0014] As a preferred solution of the quadruped robot motion control method based on reinforcement learning according to the present invention, the CVaR-PPO algorithm is provided with a risk sensitivity factor, and the risk sensitivity factor is used to make the quadruped robot have high control accuracy in complex terrains.
[0015] As a preferred solution of the quadruped robot motion control method based on reinforcement learning according to the present invention, based on the CVaR theory, the index CVaR is calculated α , and the calculation formula is as follows:
[0016]
[0017] Among them, α represents the confidence level, and its value is in (0, 1), which determines the proportion of tail samples used to calculate CVaR and reflects the degree of attention to risks. N represents the total number of sampled trajectories, which represents the number of trajectories collected from the simulation environment for analysis. represents the floor function to obtain the number of tail samples, which determines the number of samples of the worst return value selected when calculating CVaR. G i represents the worst return values, which are part of the return values of all sampled trajectories sorted from small to large.
[0018] As a preferred solution of the quadruped robot motion control method based on reinforcement learning according to the present invention, based on the PPO loss function, a risk-sensitive loss function is constructed by combining the CVaR risk measure, and the calculation formula of the risk-sensitive loss function is as follows:
[0019]
[0020] Among them, represents the risk-sensitive surrogate loss function, which is used to optimize the policy network parameter θ. θ represents the policy network parameter. represents the expectation operator, which calculates the expectation based on the sampled data at time step t. r t (θ) represents the importance sampling ratio, and πθ (a t |s t ) is the probability that the current policy θ takes action a t under state s t . is the probability that the old policy θ old takes the same action under the same state, represents the advantage function, ε is the clipping threshold of PPO, and β is the penalty coefficient.
[0021] As a preferred solution of the method for controlling the movement of a quadruped robot based on reinforcement learning according to the present invention, an Actor network and a Critic network are provided in the CVaR-PPO algorithm. The Actor network is responsible for generating action policies, and the Critic network is responsible for evaluating state values. When the CVaR-PPO algorithm trains the control policy, the parameters of the Actor network and the Critic network are trained by backpropagation using stochastic gradient descent.
[0022] A quadruped robot movement control system based on reinforcement learning. This system includes: a simulation environment construction module for building a quadruped robot simulation environment using the reinforcement learning environment ISAAC GYM; an interaction data collection module for simulating the interaction process between the quadruped robot and the environment in the simulation environment, obtaining the action feedback of the robot on different terrains through multiple simulations, and recording the interaction data with the environment; a policy training and optimization module for setting control policies and training the control policies using the CVaR-PPO algorithm in the simulation environment; a policy verification and deployment module for verifying the trained control policy in the simulation environment, and after passing the verification, deploying the control policy on the quadruped robot.
[0023] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: In a motion control method and system for a quadruped robot based on reinforcement learning provided by the present invention, a simulation environment for the quadruped robot is built by using NVIDIA's Isaac Gym framework, and training scenarios including various complex terrains such as flat ground, slopes, and steps are created, providing an efficient and realistic virtual training space for the robot, and reducing the cost and loss of hardware experiments. In the simulation environment, by simulating the interaction process between the robot and the environment multiple times, detailed action feedback, terrain features, and attitude state data are collected to construct a comprehensive data set, providing diverse and realistic data support for training, and ensuring that the control strategy has stronger environmental adaptability and generalization ability. Subsequently, the improved CVaR-PPO algorithm is used to train the control strategy. Combining the built-in risk assessment module and risk sensitivity factor, not only optimizes the policy network parameters, but also evaluates the terrain risk in real time and dynamically adjusts the control strategy to ensure the motion stability and safety of the robot in complex terrains. Finally, the trained control strategy is verified in the simulation environment and deployed to the actual robot after passing, achieving the goal of autonomous gait generation and attitude adjustment, making the motion control of the quadruped robot in unknown complex terrains more robust and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.
[0025] Figure 1 It is a schematic diagram of the steps of a motion control method for a quadruped robot based on reinforcement learning according to the present invention;
[0026] Figure 2 It is the specific framework of the algorithm of the present invention;
[0027] Figure 3 It is the result diagram of the algorithm training of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] Please refer to Figure 1 , in the first embodiment: A motion control method for a quadruped robot based on reinforcement learning is provided, and the method includes the following steps:
[0030] Step S1: Build a simulation environment for a quadruped robot using the reinforcement learning environment Isaac Gym.
[0031] Preferably, build a simulation environment, use Isaac Gym to construct a quadruped robot model and its motion simulation environment, and simulate complex terrains in the real world by simulating different terrain scenarios, such as slopes, steps, and uneven ground.
[0032] The quadruped robot obtains terrain feature data, robot attitude data, and joint state data in real time through a sensor module, and inputs the terrain feature data, robot attitude data, and joint state data into the CVaR-PPO algorithm to train the control strategy; the physical properties of the quadruped robot include the KP and KD parameters of joint control, torque limits, and kinematic parameters.
[0033] Step S2: In the simulation environment, simulate the interaction process between the quadruped robot and the environment, obtain the action feedback of the robot on different terrains through multiple simulations, and record the interaction data with the environment.
[0034] Preferably, define the simulation motion rules, define the motion rules for the simulation model, and set the state space, action space, and reward function of the multi-legged robot.
[0035] Step S3: Train the CVaR-PPO algorithm. In the simulation environment, use the PPO reinforcement learning method improved based on CVaR for training. By introducing a risk assessment module, the risks of different terrains are evaluated in real time, and the control strategy is dynamically adjusted to ensure the safety and stability of the robot in complex terrains.
[0036] The specific construction process of the CVaR-PPO algorithm includes the following steps:
[0037] Initialize the parameters θ of the Actor network and the parameters η of the Critic network, and create an experience buffer for storing data during the interaction with the environment.
[0038] Use the Actor network to interact with the simulation environment and collect trajectory data. The trajectory data mainly includes the state s t , action a t , reward r t , and the next state s t+1 . In this process, a risk sensitivity factor is used to dynamically evaluate the risk levels of different terrains, so as to obtain risk information about the motion process.
[0039] Calculate the discounted return {G t} for each sequence in the sampled trajectory, and all {G tSort them from small to large, and take the worst α% of the data as the tail return set to measure extremely adverse situations.
[0040] Calculate CvaR. Based on the worst α% return set, calculate its average value to obtain CVaR. α 。
[0041] Construct the objective function. Based on the objective function of the standard PPO, introduce the CVaR penalty term to obtain a risk-sensitive loss function:
[0042]
[0043] Among them, represents the risk-sensitive surrogate loss function, which is used to optimize the policy network parameter θ. θ represents the policy network parameter. represents the expectation operator, which calculates the expectation based on the sampled data at time step t. r t (θ) represents the importance sampling ratio, and π θ (a t |s t ) is the probability that the current policy θ takes action a t under the state s t . is the probability that the old policy θ old takes the same action under the same state. represents the advantage function. ε is the clipping threshold of PPO, and β is the penalty coefficient.
[0044] Preferably, update the Critic network parameter η, minimize the error between the predicted value and the true return value, and the optimization objective is:
[0045]
[0046] Among them, L value (η) represents the value loss function.
[0047] Update the Actor network parameter. Adopt the stochastic gradient descent optimization algorithm to minimize Update the Actor network parameter θ.
[0048] Repeat the above steps, continuously sample and optimize the policy until the training converges or reaches the preset iteration upper limit.
[0049] Step S4: Simulation verification. Use the trained CVaR-PPO control strategy to conduct multiple simulation verifications on the quadruped robot to test its motion performance under different terrain conditions. The simulation verification includes tests on the motion stability and adaptability in complex terrains such as flat terrain, obstacle terrain, and slopes, ensuring that the control strategy has high robustness in diverse terrains.
[0050] Furthermore, due to the characteristics of model-free reinforcement learning, training requires a large amount of computing power support. Therefore, parallel computing is used to reduce the training time.
[0051] This embodiment also provides a motion control system for a quadruped robot based on reinforcement learning. The system includes: a simulation environment construction module for building a simulation environment for the quadruped robot using the reinforcement learning environment ISAAC GYM; an interaction data acquisition module for simulating the interaction process between the quadruped robot and the environment in the simulation environment, obtaining the action feedback of the robot on different terrains through multiple simulations, and recording the interaction data with the environment; a policy training and optimization module for setting the control strategy and training the control strategy using the CVaR-PPO algorithm in the simulation environment; a policy verification and deployment module for verifying the trained control strategy in the simulation environment and, after passing the verification, deploying the control strategy on the quadruped robot.
[0052] Please refer to Figures 2 - 3 , in the second embodiment: A motion control method for a quadruped robot based on reinforcement learning is provided, including the following:
[0053] S1. Use solid works to establish a quadruped robot model and convert it into a urdf file for easy simulation and algorithm training.
[0054] S2. Create a simulation environment under the Isaac gym platform and add the quadruped robot model for algorithm training and demonstration.
[0055] S3. Set various parameters required for algorithm training, including the state space, action space, and reward function.
[0056] S4. Build the CvaR-PPO algorithm, including the construction of the actor network and critic network, and the establishment of the CvaR module.
[0057] S5. Use the CvaR-PPO algorithm to train the quadruped robot in Isaac gym.
[0058] Specifically, in S2, the Isaac Gym software is used to create a simulation environment required for parallel training of the robot and load the URDF model of the quadruped robot, and corresponding physical engine parameters, ground friction coefficient, gravitational acceleration, and possible environmental disturbance information are set.
[0059] Specifically, in S3, according to the running objectives and safety requirements of the robot, key parameters required for reinforcement learning are set, including the state space, action space, and reward function. The setting of the main reward function is shown in Table 1.
[0060] Table 1
[0061]
[0062] Specifically, in S4, the construction of the CvaR-PPO algorithm is as Figure 2 shown, which includes the construction of the Actor network and the Critic network, the introduction of the CVaR risk measurement module, the construction of the risk-sensitive objective function, and a simple training process.
[0063] The Actor network trains the robot by inputting the state s t of the quadruped robot and outputting the action probability distribution π θ (a t , s t ). It includes three hidden layers, each with 225 dimensions.
[0064] The CVaR risk measurement module adds the CVaR α term to the original objective function of PPO and balances the trade-off between the average return and the extreme loss through the weight coefficient β. Its main process is to sort all the returns sampled during the training process, then select the worst tail returns and calculate the mean value to obtain the CVaR α .
[0065] The specific training process of the CvaR-PPO algorithm is shown in Table 2:
[0066] Table 2
[0067]
[0068]
[0069] Specifically, in S5, parallel training is used, and 1024 robots are used for training simultaneously, which can quickly collect a large number of diverse training samples and greatly reduce the "exploration blind area" of the policy in complex terrains or random disturbances. After 6000 times of training, the results are as Figure 3 shown.
[0070] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0071] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A quadruped robot motion control method based on reinforcement learning, characterized in that: The method comprises the following steps: Step S1: Use the reinforcement learning environment ISAAC GYM to build a quadruped robot simulation environment; Step S2: In the simulation environment, simulating the interaction process between the quadruped robot and the environment, obtaining the action feedback of the robot on different terrains through multiple simulations, and recording the interaction data with the environment; Step S3: setting a control strategy, and training the control strategy using the CVaR-PPO algorithm in the simulation environment; Step S4: After the training is completed, the control strategy is verified in a simulation environment. After the verification is passed, the control strategy is deployed on the quadruped robot.
2. A quadruped robot motion control method based on reinforcement learning according to claim 1, characterized in that: The simulation environment includes different types of terrain scenes, including slopes, steps and uneven ground, and the simulation environment is used to verify the stability and robustness of the control strategy.
3. A quadruped robot motion control method based on reinforcement learning according to claim 1, characterized in that: The quadruped robot acquires terrain feature data, robot posture data and joint status data in real time through a sensor module, and inputs the terrain feature data, robot posture data and joint status data into the CVaR-PPO algorithm to train the control strategy; the physical properties of the quadruped robot include KP and KD parameters of joint control, torque limit and kinematic parameters.
4. A quadruped robot motion control method based on reinforcement learning according to claim 1, characterized in that: The CVaR-PPO algorithm has a built-in risk assessment module, which is used to assess terrain risks in real time and dynamically adjust the control strategy according to the assessment results.
5. The method for controlling quadruped robot motion based on reinforcement learning according to claim 1, characterized in that: The CVaR-PPO algorithm is provided with a risk sensitivity factor, and the risk sensitivity factor is used to enable the quadruped robot to have high control accuracy in complex terrain.
6. A quadruped robot motion control method based on reinforcement learning according to claim 1, characterized in that: Based on the CVaR theory, the CVaR indicator is calculated. α , the calculation formula is as follows: Where α represents the confidence level, N represents the total number of sampled trajectories, Indicates the number of tail samples rounded down, G i Indicates the worst A return value.
7. The method for controlling quadruped robot motion based on reinforcement learning according to claim 1, characterized in that: Based on the PPO loss function, a risk-sensitive loss function is constructed in combination with the CVaR risk metric. The calculation formula of the risk-sensitive loss function is as follows: in, represents the risk-sensitive surrogate loss function, θ represents the policy network parameters, represents the expectation operator, r t (θ) represents the importance sampling ratio, and π θ (a t |s t ) is the current policy θ in state s t Take action a t The probability of is the old strategy θ old The probability of taking the same action in the same state, represents the advantage function, ε is the clipping threshold of PPO, and β is the penalty coefficient.
8. The method for controlling quadruped robot motion based on reinforcement learning according to claim 4, characterized in that: The CVaR-PPO algorithm is provided with an Actor network and a Critic network, wherein the Actor network is responsible for generating action strategies, and the Critic network is responsible for evaluating state values; when the CVaR-PPO algorithm trains the control strategy, the parameters of the Actor network and the Critic network are back-propagated using stochastic gradient descent.
9. A quadruped robot motion control system based on reinforcement learning according to claim 8, characterized in that: The simulation environment building module is used to build a quadruped robot simulation environment using the reinforcement learning environment ISAAC GYM; An interactive data acquisition module is used to simulate the interaction process between the quadruped robot and the environment in the simulation environment, obtain the action feedback of the robot on different terrains through multiple simulations, and record the interactive data with the environment; A strategy training optimization module is used to set a control strategy, and train the control strategy using a CVaR-PPO algorithm in the simulation environment; Strategy verification and deployment module: used to verify the control strategy after training in a simulation environment. After verification, the control strategy is deployed on the quadruped robot.
Citation Information
Cited By
Fresh medicine pulp quality dynamic monitoring and control method based on big data
CN121806517A