A quadruped robot motion control method and system based on deep reinforcement learning based on constraint rewards
By simulating real environment parameters in the simulation training environment and tuning model in the inference testing environment, the deep reinforcement learning method based on constraint rewards solves the problem of insufficient stability in the real environment, and the long-term and stable operation of the four-legged robot in complex terrain is achieved.
Patent Information
- Application Number
- CN202510082218.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Four-legged robots have insufficient long-term stability in real environments, and the existing deep reinforcement learning motion control methods have poor stability when migrating from simulation environments to real environments.
A deep reinforcement learning four-legged robot motion control method based on constraint reward is proposed. By using domain randomized parameters in the simulation training environment, model inference test tuning is performed in the inference test environment, the difference between simulation and real environment is reduced.
Through the target strategy network model, the four-legged robot can achieve long-term and stable operation in the real environment, improving the stability and adaptability of the four-legged robot in complex terrain.
Smart Images

Figure CN119512184B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent motion technology, and in particular to a deep reinforcement learning quadruped robot motion control method and system based on constraint reward. Background Art
[0002] With the continuous development of mobile robot control technology, quadruped robots have been widely used in operations on complex terrain due to their high degrees of freedom and discrete footholds.
[0003] Among them, the mechanical structure of the quadruped robot is complex, and the traditional motion control method requires the construction of an accurate dynamic model, which has insufficient adaptability; the existing deep reinforcement learning motion control method has end-to-end advantages, but the control model has poor stability when migrating from the simulation environment to the real environment, which makes the long-term stability of the quadruped robot insufficient in the real environment. Summary of the invention
[0004] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0005] To this end, an object of the present invention is to propose a deep reinforcement learning quadruped robot motion control method based on constrained rewards. This method simulates the parameters of the real environment through domain randomization parameters in a simulation training environment, and performs model reasoning test tuning through an inference test environment, thereby reducing the difference between the simulation training environment and the real environment, thereby enabling the quadruped robot to operate stably and for a long time in the real environment through the target strategy network model.
[0006] Another object of the present invention is to propose a deep reinforcement learning quadruped robot motion control system based on constraint rewards.
[0007] To achieve the above object, an embodiment of the present invention proposes a deep reinforcement learning quadruped robot motion control method based on constraint reward, comprising:
[0008] Establishing a simulation training environment for deep reinforcement learning of a quadruped robot, wherein the simulation training environment includes robot information of the quadruped robot and first simulation environment information;
[0009] Determining a reward function, a domain randomization parameter, and a cost constraint function of the simulation training environment;
[0010] Based on the robot information and the first simulation environment information, the initial policy network model is trained in the simulation training environment by using the reward function and the cost constraint function to obtain a trained policy network model;
[0011] Establishing an inference test environment, and deploying the trained policy network model to the inference test environment for model inference test tuning to obtain a target policy network model;
[0012] The target strategy network model is deployed into the quadruped robot to perform motion control on the quadruped robot.
[0013] The constraint-reward-based deep reinforcement learning quadruped robot motion control method of the embodiment of the present invention may also have the following additional technical features:
[0014] Furthermore, the establishment of a simulation training environment for deep reinforcement learning of a quadruped robot includes:
[0015] Loading a URDF file containing robot information of the quadruped robot in the Isaac Gym simulation training environment, wherein the URDF file includes detailed information on the geometric shape, mass, inertia matrix, joint type, and connection relationship of the quadruped robot;
[0016] A first environment model file containing the first simulated environment information is loaded into the Isaac Gym simulation training environment, wherein the first environment model file includes information on flat ground, ascending stairs, descending stairs, uphill, downhill and rugged terrain, and mixed terrain information.
[0017] Further, determining the reward function, domain randomization parameter and cost constraint function of the simulation training environment includes:
[0018] Designing a reward function of the simulation training environment according to a Markov decision process;
[0019] Determining domain randomization parameters of the simulation training environment by a random function;
[0020] The cost constraint function of the simulation training environment is designed according to the actual environment requirements.
[0021] Furthermore, the reward function includes:
[0022] =λ w1 r tlv +λ w2 r tav +λ w3 r fat +λ w4 r cc +λ w5 r bh +λ w6 r da +λ w7 r so
[0023] Among them, the r tlv To calculate forward velocity tracking, the r tav To calculate the angular velocity tracking, the r fat To calculate the foot end air time, the r cc To detect whether a collision occurs, the r bh is the difference between the calculated height of the body and the target height, the r da is the joint acceleration penalty, the r so To calculate the difference between the joint position and the soft limit, the λ wi is the scaling factor of the i-th reward function value.
[0024] Furthermore, the cost constraint function includes:
[0025] C = c p + c t + c d
[0026] Among them, the c p To calculate the difference between the joint position and the first limiting threshold, the c t To calculate the difference between the joint torque and the second limiting threshold, c d To calculate the difference between the joint velocity and the third limiting threshold, , , and are the corresponding proportional coefficients.
[0027] Furthermore, the initial policy network model includes a first policy network and a second policy network, wherein the first policy network is used to generate an action strategy, and the second policy network is used to evaluate the generated action strategy to obtain reward information; based on the robot information and the first simulation environment information, the initial policy network model is trained in the simulation training environment by the reward function and the cost constraint function to obtain a trained policy network model, including:
[0028] configuring an initial policy network and a corresponding teacher network model in the simulation training environment, wherein the teacher network model is trained, and the first policy network, the second policy network and the teacher network model have the same network structure;
[0029] Based on the robot information and the first simulation environment information, according to the teacher network model, the reward function and the cost constraint function, the initial policy network is trained in the simulation training model through the proximal strategy optimization algorithm of multiple loss functions, and the weight parameters corresponding to the reward function and the cost constraint function are adjusted until the initial policy network model converges to obtain a trained policy network model.
[0030] Furthermore, the establishing of the reasoning test environment includes:
[0031] Loading a second environment model file containing second simulation environment information in the IDE, wherein the second environment model file includes a simulation environment world model;
[0032] The second environment model file is run in the IDE to obtain the Gazebo reasoning test environment.
[0033] To achieve the above object, another embodiment of the present invention proposes a deep reinforcement learning quadruped robot motion control system based on constraint rewards, the system comprising:
[0034] A first establishment module is used to establish a simulation training environment for deep reinforcement learning of a quadruped robot, wherein the simulation training environment includes robot information of the quadruped robot and first simulation environment information;
[0035] A determination module, used to determine a reward function, a domain randomization parameter, and a cost constraint function of the simulation training environment;
[0036] A training module, used to train the initial policy network model through the reward function and the cost constraint function in the simulation training environment based on the robot information and the first simulation environment information, to obtain a trained policy network model;
[0037] The second establishment module is used to establish an inference test environment, and deploy the trained policy network model to the inference test environment to perform model inference test tuning to obtain a target policy network model;
[0038] A control module is used to deploy the target strategy network model into the quadruped robot to perform motion control on the quadruped robot.
[0039] The constrained reward-based deep reinforcement learning quadruped robot motion control method and system proposed in the present invention simulates the parameters of the real environment through domain randomization parameters in a simulation training environment, and performs model reasoning test tuning through an inference test environment, thereby reducing the difference between the simulation training environment and the real environment, thereby enabling the quadruped robot to operate stably and for a long time in a real environment through the target strategy network model.
[0040] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0042] Figure 1 A flowchart of a quadruped robot motion control method based on deep reinforcement learning with constraint rewards according to an embodiment of the present invention;
[0043] Figure 2 A schematic diagram of a simulation training environment according to an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of an initial strategy network model according to an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of a Gazebo reasoning test environment according to an embodiment of the present invention;
[0046] Figure 5 is a schematic diagram of a motion effect in a real environment according to an embodiment of the present invention;
[0047] Figure 6 Schematic diagram of the structure of a quadruped robot motion control system based on deep reinforcement learning and constraint rewards according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0049] The following describes a method and system for controlling a quadruped robot using deep reinforcement learning based on constraint rewards according to an embodiment of the present invention with reference to the accompanying drawings.
[0050] First, the constraint-reward-based deep reinforcement learning quadruped robot motion control method proposed in accordance with an embodiment of the present invention will be described with reference to the accompanying drawings.
[0051] Figure 1 The present invention is a flowchart of a method for controlling a quadruped robot motion using deep reinforcement learning based on constraint rewards according to an embodiment of the present invention.
[0052] like Figure 1 As shown, the constraint reward-based deep reinforcement learning quadruped robot motion control method includes the following steps:
[0053] Step S1, establishing a simulation training environment for deep reinforcement learning of a quadruped robot, wherein the simulation training environment includes robot information of the quadruped robot and first simulation environment information.
[0054] In one embodiment of the present invention, the above-mentioned method of establishing a simulation training environment for deep reinforcement learning of a quadruped robot may include: loading a URDF file containing robot information of a quadruped robot in the Isaac Gym simulation training environment, and loading a first environment model file containing first simulation environment information in the Isaac Gym simulation training environment.
[0055] Among them, in one embodiment of the present invention, the above-mentioned URDF file may include detailed information on the geometry, mass, inertia matrix, joint type and connection relationship of the quadruped robot; the above-mentioned first environment model file may include flat ground, going up stairs, going down stairs, uphill, downhill and rugged terrain information, as well as mixed terrain information.
[0056] And, in one embodiment of the present invention, Figure 2 A schematic diagram of a simulation training environment proposed in an embodiment of the present invention.
[0057] Step S2, determining a reward function, a domain randomization parameter, and a cost constraint function of a simulation training environment;
[0058] Among them, in one embodiment of the present invention, after obtaining the simulation training environment through the above steps, it is also necessary to determine the reward function, domain randomization parameters and cost constraint function of the simulation training environment, so that model training can be performed in the simulation training environment later.
[0059] Specifically, in one embodiment of the present invention, the method for determining the reward function, domain randomization parameters and cost constraint function of the simulation training environment may include the following steps:
[0060] Step S21, designing a reward function of a simulation training environment according to a Markov decision process;
[0061] Step S22, determining a domain randomization parameter of a simulation training environment through a random function;
[0062] Step S23, designing a cost constraint function of the simulation training environment according to the actual environment requirements.
[0063] In one embodiment of the present invention, a reward function of a simulation training environment can be designed according to a Markov decision process through artificial experience, wherein the reward function is:
[0064] =λ w1 r tlv +λ w2 r tav +λ w3 r fat +λ w4 r cc +λ w5 r bh +λ w6 r da +λ w7 r so
[0065] Among them, r tlv To calculate the forward velocity tracking, r tav To calculate the angular velocity tracking, r fat To calculate the foot end air time, r cc To detect whether a collision occurs, r bh is the difference between the calculated body height and the target height, r da is the joint acceleration penalty, r so To calculate the difference between the joint position and the soft limit, λ wi is the scaling factor of the i-th reward function value, which is convenient for debugging parameters in the simulation environment.
[0066] Specifically, in one embodiment of the present invention, the above ,in, =1.0; above ,in, =0.5; above ,in, =1.0; above ,in, =-1.0; above ,in, =-0.5; above ,in, =-2.5× ; above ,in, =1.0.
[0067] Among them, in one embodiment of the present invention, is the actual simulation speed, is the target speed; is the simulated angular velocity, Reduce to the target angular velocity; is the simulated leg-lifting time, is the target time; r ccTo judge whether there is a collision; is the height in the simulation, is the target height; is the acceleration of 12 joints; is the difference between the current position and the target position, and p is the limit value configured in the robot information.
[0068] Furthermore, in one embodiment of the present invention, the method for determining the domain randomization parameters of the simulation training environment through a random function may include: determining a threshold range of the domain randomization parameters, and randomly determining the corresponding domain randomization parameters from the threshold range through a random function.
[0069] In one embodiment of the present invention, the domain randomization parameters may include a friction coefficient , elastic modulus , Loading mass , body center of mass and robot thrust randomization .
[0070] And, in one embodiment of the present invention, the friction coefficient , used to simulate the friction coefficient of different terrains in the training environment; the above elastic coefficient , elastic coefficient for different terrains in simulation training environment; load mass , used to simulate the effects of different loads on the quadruped robot in the training environment; the center of mass of the body , used to simulate the effects of different postures of quadruped robots in training environments; randomization of robot thrust , which is used to simulate the effect of a quadruped robot under external disturbances in a training environment.
[0071] Furthermore, in one embodiment of the present invention, the threshold range corresponding to the domain randomization parameter is: The corresponding first threshold range is [0.2, 1.2]; elastic coefficient The corresponding second threshold range is [0.0, 0.7]; load mass The corresponding third threshold range is [-1, 2]; the center of mass of the fuselage The corresponding fourth threshold range is [-0.05, 0.06]; the robot thrust is randomized The corresponding fifth threshold is 15.
[0072] And, in one embodiment of the present invention, the parameter value corresponding to the domain randomization parameter can be randomly determined from the corresponding threshold range by the domain_rand function. In one embodiment of the present invention, through the above-mentioned domain randomization parameter, physical properties can be simulated and added in the simulation training environment, so that the model can learn more realistic features in the simulation environment, and achieve the stability of the model migration in the real environment.
[0073] In one embodiment of the present invention, a cost constraint function of a simulation training environment can be designed based on real environment requirements through artificial experience, wherein the cost constraint function is:
[0074] C = c p + c t + c d
[0075] Among them, c p To calculate the difference between the joint position and the first limiting threshold, c t To calculate the difference between the joint torque and the second limiting threshold, c d To calculate the difference between the joint velocity and the third limiting threshold, , and is the corresponding proportionality coefficient.
[0076] In one embodiment of the present invention, the first limit threshold may be a joint limit threshold, the second limit threshold may be a joint torque limit threshold, and the third limit threshold may be a joint speed limit threshold. The joint limit threshold, the joint torque limit threshold, and the joint speed limit threshold may be thresholds configured in the robot information.
[0077] Among them, in one embodiment of the present invention, the above , and It can be set experimentally or manually, for example , and =0.3. In one embodiment of the present invention, according to the above cost constraint function, the direction of subsequent model learning can be guided by minimizing the cost constraint function, so that the model can quickly find an effective strategy, thereby improving the efficiency of model training.
[0078] Step S3, based on the robot information and the first simulation environment information, the initial policy network model is trained in a simulation training environment by using a reward function and a cost constraint function to obtain a trained policy network model;
[0079] In one embodiment of the present invention, the initial policy network model includes a first policy network and a second policy network, wherein the first policy network is used to generate an action policy, and the second policy network is used to evaluate the generated action policy to obtain reward information. In one embodiment of the present invention, the first policy network may be an Actor network, and the second policy network may be a Critic network (e.g. Figure 3 as shown).
[0080] And, in one embodiment of the present invention, based on the robot information and the first simulation environment information, the initial policy network model is trained by the reward function and the cost constraint function in the simulation training environment to obtain the trained policy network model. The method may include the following steps:
[0081] Step S31, configuring an initial policy network and a corresponding teacher network model in a simulation training environment, wherein the teacher network model is trained, and the first policy network, the second policy network and the teacher network model have the same network structure;
[0082] Step S32, based on the robot information and the first simulation environment information, according to the teacher network model and the reward function and the cost constraint function, the initial policy network is trained in the simulation training model through the proximal policy optimization algorithm of multiple loss functions, and the weight parameters corresponding to the reward function and the cost constraint function are adjusted until the initial policy network model converges to obtain a trained policy network model.
[0083] Among them, in one embodiment of the present invention, the first policy network and the second policy network in the above-mentioned initial policy network have the same network structure as the teacher network model. In one embodiment of the present invention, the policy network model composed of the teacher network model and the first policy network and the second policy network can be a multi-layer perceptron network. And, in one embodiment of the present invention, the teacher network model is a multi-layer perceptron network completed based on course learning training.
[0084] Furthermore, in one embodiment of the present invention, based on the robot information and the first simulation environment information, the course tasks can be designed from easy to difficult in the simulation training model, and the linear speed tracking method is adopted. When the quadruped robot can stand stably and the movement speed on the flat ground reaches 0.5m / s, it enters the next course, and the terrain information is upgraded according to the course content. The collected terrain information and proprioceptive perception information are used as observation data. After encoding the observation data, it is returned to the initial policy network for policy generation and policy evaluation, and the policy generation and policy evaluation in the initial policy network are adjusted according to the reward function value of the reward function and the constraint cost function value of the cost constraint function through the proximal policy optimization algorithm of multiple loss functions, and the weight parameters corresponding to the reward function and the cost constraint function are adjusted; the above process is continuously repeated until the initial policy network model training converges, all terrain tasks can be completed, and a trained policy network model is obtained.
[0085] Among them, in one embodiment of the present invention, the formula corresponding to the proximal strategy optimization algorithm of the above-mentioned multi-loss function is:
[0086]
[0087] in, is the gradient loss function, Designed for advantage, is entropy regularization, , is a hyperparameter.
[0088] Furthermore, in one embodiment of the present invention, after the initial policy network model is trained in a simulation training environment, the reward function value of the quadruped robot's reward function and the cost function value of the cost constraint function can be monitored in real time.
[0089] In one embodiment of the present invention, curriculum learning and teacher-student methods are introduced into the proximal strategy optimization algorithm of multiple loss functions, and the training efficiency of the model is improved by designing different levels of curriculum learning stages from easy to difficult. At the same time, the teacher model is trained using a multi-layer perceptron network, and the teacher model uses the knowledge distillation method to guide the student model training, which improves the model training efficiency and realizes model compression.
[0090] Step S4, establish an inference test environment, and deploy the trained policy network model to the inference test environment for model inference test tuning to obtain the target policy network model;
[0091] Among them, in one embodiment of the present invention, after obtaining the trained policy network model through the above steps, an inference test environment can be established, and the trained policy network model can be deployed to the inference test environment for model inference test tuning to obtain the target policy network model.
[0092] Specifically, in one embodiment of the present invention, the method for establishing the inference test environment may include: loading a second environment model file containing second simulation environment information in an IDE, running the second environment model file in the IDE, and obtaining a Gazebo inference test environment (such as Figure 4 As shown). In one embodiment of the present invention, the second environment model file includes a simulated environment world model (such as a flat ground, stairs, and rugged terrain world model, so as to make the reasoning test environment more realistic.
[0093] Furthermore, in one embodiment of the present invention, after the trained policy network model is deployed to the inference test environment, the keyboard can be used to control the movement of the quadruped robot in the inference test environment, and the motion control output of the trained policy network model in the inference test environment can be observed. The trained policy network model and the weight parameters of the reward function and the cost constraint function can be adjusted according to the motion output until the motion control output completes all terrain tasks and the target policy network model is obtained.
[0094] Step S5, deploying the target strategy network model into the quadruped robot to perform motion control on the quadruped robot.
[0095] Among them, in one embodiment of the present invention, after the target policy network model is obtained through the above steps, the target policy network model can be deployed in a quadruped robot to perform motion control on the quadruped robot.
[0096] Specifically, in one embodiment of the present invention, the reinforcement learning model corresponding to the target policy network model is converted into a jit format, and the deployment framework is used to deploy the jit format model to an embedded device for reasoning testing, and the high-level control system of the quadruped robot is switched to the low-level control system, and the state machine is used to switch to the reinforcement learning model control state, and the control system of the quadruped robot is connected using a joystick, so that the quadruped robot can be controlled to pass through flat ground, stairs and rugged terrain, and the motion state output of the quadruped robot is observed, and the motion effect is as follows: Figure 5 shown.
[0097] According to the constrained reward-based deep reinforcement learning quadruped robot motion control method proposed in an embodiment of the present invention, the parameters of the real environment are simulated through domain randomization parameters in the simulation training environment, and the model reasoning test tuning is performed through the reasoning test environment, thereby reducing the difference between the simulation training environment and the real environment, thereby enabling the quadruped robot to operate stably and for a long time in the real environment through the target strategy network model.
[0098] Next, a deep reinforcement learning quadruped robot motion control system based on constraint rewards proposed according to an embodiment of the present invention is described with reference to the accompanying drawings.
[0099] Figure 6 Schematic diagram of the structure of a quadruped robot motion control system based on deep reinforcement learning and constraint rewards according to an embodiment of the present invention.
[0100] like Figure 6 As shown, the deep reinforcement learning quadruped robot motion control system 10 based on constraint reward includes: a first establishment module 601, a determination module 602, a training module 603, a second establishment module 604 and a control module 605, wherein
[0101] A first establishing module 601 is used to establish a simulation training environment for deep reinforcement learning of a quadruped robot, wherein the simulation training environment includes robot information of the quadruped robot and first simulation environment information;
[0102] A determination module 602 is used to determine a reward function, a domain randomization parameter, and a cost constraint function of a simulation training environment;
[0103] A training module 603 is used to train the initial policy network model through a reward function and a cost constraint function in a simulation training environment based on the robot information and the first simulation environment information to obtain a trained policy network model;
[0104] The second establishment module 604 is used to establish an inference test environment and deploy the trained policy network model to the inference test environment to perform model inference test optimization to obtain a target policy network model;
[0105] The control module 605 is used to deploy the target strategy network model into the quadruped robot to perform motion control on the quadruped robot.
[0106] Furthermore, the first establishing module 601 is specifically used for:
[0107] Load the URDF file containing the robot information of the quadruped robot in the Isaac Gym simulation training environment, wherein the URDF file includes the geometric shape, mass, inertia matrix, joint type and connection relationship detailed information of the quadruped robot;
[0108] A first environment model file containing first simulation environment information is loaded into the Isaac Gym simulation training environment, wherein the first environment model file includes information on flat ground, ascending stairs, descending stairs, uphill, downhill and rugged terrain, and mixed terrain information.
[0109] Furthermore, the determination module 602 is further configured to:
[0110] Design the reward function of the simulation training environment based on the Markov decision process;
[0111] Determine the domain randomization parameters of the simulation training environment through a random function;
[0112] Design the cost constraint function of the simulation training environment according to the actual environment requirements.
[0113] Furthermore, the above reward function includes:
[0114] =λ w1 r tlv +λ w2 r tav +λ w3 r fat +λ w4 r cc +λ w5 r bh +λ w6 r da +λ w7 r so
[0115] Among them, r tlv To calculate the forward velocity tracking, r tav To calculate the angular velocity tracking, r fat To calculate the foot end air time, r cc To detect whether a collision occurs, r bh is the difference between the calculated body height and the target height, r da is the joint acceleration penalty, r so To calculate the difference between the joint position and the soft limit, λ wi is the scaling factor of the i-th reward function value.
[0116] Furthermore, the above cost constraint function includes:
[0117] C = c p + c t + c d
[0118] Among them, c p To calculate the difference between the joint position and the first limiting threshold, c t To calculate the difference between the joint torque and the second limiting threshold, c d To calculate the difference between the joint velocity and the third limiting threshold, , and is the corresponding proportionality coefficient.
[0119] Furthermore, the initial strategy network model includes a first strategy network and a second strategy network, wherein the first strategy network is used to generate an action strategy, and the second strategy network is used to evaluate the generated action strategy to obtain reward information; the training module 603 is specifically used to:
[0120] In a simulation training environment, an initial policy network and a corresponding teacher network model are configured, wherein the teacher network model is trained, and the first policy network, the second policy network and the teacher network model have the same network structure;
[0121] Based on the robot information and the first simulation environment information, according to the teacher network model, the reward function and the cost constraint function, the initial policy network is trained in the simulation training model through the proximal policy optimization algorithm of multiple loss functions, and the weight parameters corresponding to the reward function and the cost constraint function are adjusted until the initial policy network model converges to obtain a trained policy network model.
[0122] Furthermore, the second establishing module 604 is specifically used for:
[0123] Loading a second environment model file containing second simulation environment information in the IDE, wherein the second environment model file includes a simulation environment world model;
[0124] Run the second environment model file in the IDE to get the Gazebo reasoning test environment.
[0125] According to the constrained reward-based deep reinforcement learning quadruped robot motion control system proposed in an embodiment of the present invention, the parameters of the real environment are simulated through domain randomization parameters in the simulation training environment, and the model reasoning test tuning is performed through the reasoning test environment, thereby reducing the difference between the simulation training environment and the real environment, thereby enabling the quadruped robot to operate stably and for a long time in the real environment through the target strategy network model.
[0126] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0127] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0128] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A deep reinforcement learning quadruped robot motion control method based on constrained rewards, characterized in that: The method comprises: Establishing a simulation training environment for deep reinforcement learning of a quadruped robot, wherein the simulation training environment includes robot information of the quadruped robot and first simulation environment information; Determining a reward function, a domain randomization parameter, and a cost constraint function of the simulation training environment; Based on the robot information and the first simulation environment information, the initial policy network model is trained in the simulation training environment by using the reward function and the cost constraint function to obtain a trained policy network model; Establishing an inference test environment, and deploying the trained policy network model to the inference test environment for model inference test tuning to obtain a target policy network model; Deploying the target strategy network model into the quadruped robot to perform motion control on the quadruped robot; The reward function of the simulation training environment is designed according to the Markov decision process, and the reward function includes: =λ w1 r tlv +λ w2 r tav +λ w3 r fat +λ w4 r cc +λ w5 r bh +λ w6 r da +λ w7 r so Among them, the r tlv To calculate forward velocity tracking, the r tav To calculate the angular velocity tracking, the r fat To calculate the foot end air time, the r cc To detect whether a collision occurs, the r bh is the difference between the calculated height of the body and the target height, the r da is the joint acceleration penalty, the r so To calculate the difference between the joint position and the soft limit, the λ wi is the scaling factor of the i-th reward function value; in, ; ; ; ; ; ; ; is the actual simulation speed, is the target speed; is the simulated angular velocity, Reduce to the target angular velocity; is the simulated leg-lifting time, is the target time; r cc To judge whether there is a collision; is the height in the simulation, is the target height; is the acceleration of 12 joints; is the difference between the current position and the target position, and p is the limit value configured in the robot information; The cost constraint function includes: C = c p + c t + c d Among them, the c p To calculate the difference between the joint position and the first limiting threshold, the c t To calculate the difference between the joint torque and the second limit threshold, the c d To calculate the difference between the joint velocity and the third limiting threshold, , and is the corresponding proportionality coefficient.
2. The method according to claim 1, characterized in that The establishment of a simulation training environment for deep reinforcement learning of a quadruped robot includes: Loading a URDF file containing robot information of the quadruped robot in the Isaac Gym simulation training environment, wherein the URDF file includes detailed information on the geometric shape, mass, inertia matrix, joint type, and connection relationship of the quadruped robot; A first environment model file containing the first simulated environment information is loaded into the Isaac Gym simulation training environment, wherein the first environment model file includes information on flat ground, ascending stairs, descending stairs, uphill, downhill and rugged terrain, and mixed terrain information.
3. The method according to claim 1, characterized in that Determining a domain randomization parameter and a cost constraint function of the simulation training environment includes: Determining domain randomization parameters of the simulation training environment by a random function; The cost constraint function of the simulation training environment is designed according to the actual environment requirements.
4. The method according to claim 1, characterized in that: The initial policy network model includes a first policy network and a second policy network, wherein the first policy network is used to generate an action strategy, and the second policy network is used to evaluate the generated action strategy to obtain reward information; based on the robot information and the first simulation environment information, the initial policy network model is trained in the simulation training environment by the reward function and the cost constraint function to obtain a trained policy network model, including: configuring an initial policy network and a corresponding teacher network model in the simulation training environment, wherein the teacher network model is trained, and the first policy network, the second policy network and the teacher network model have the same network structure; Based on the robot information and the first simulation environment information, according to the teacher network model, the reward function and the cost constraint function, the initial policy network is trained in the simulation training model through the proximal strategy optimization algorithm of multiple loss functions, and the weight parameters corresponding to the reward function and the cost constraint function are adjusted until the initial policy network model converges to obtain a trained policy network model.
5. The method according to claim 1, characterized in that The establishing of the reasoning test environment comprises: Loading a second environment model file containing second simulation environment information in the IDE, wherein the second environment model file includes a simulation environment world model; The second environment model file is run in the IDE to obtain the Gazebo reasoning test environment.
6. A deep reinforcement learning quadruped robot motion control system based on constraint rewards, characterized in that: The system comprises: A first establishment module is used to establish a simulation training environment for deep reinforcement learning of a quadruped robot, wherein the simulation training environment includes robot information of the quadruped robot and first simulation environment information; A determination module, used to determine a reward function, a domain randomization parameter, and a cost constraint function of the simulation training environment; A training module, used to train the initial policy network model through the reward function and the cost constraint function in the simulation training environment based on the robot information and the first simulation environment information, to obtain a trained policy network model; The second establishment module is used to establish an inference test environment, and deploy the trained policy network model to the inference test environment to perform model inference test tuning to obtain a target policy network model; A control module, used for deploying the target strategy network model into the quadruped robot to perform motion control on the quadruped robot; The reward function of the simulation training environment is designed according to the Markov decision process, and the reward function includes: =λ w1 r tlv +λ w2 r tav +λ w3 r fat +λ w4 r cc +λ w5 r bh +λ w6 r da +λ w7 r so Among them, the r tlv To calculate forward velocity tracking, the r tav To calculate the angular velocity tracking, the r fat To calculate the foot end air time, the r cc To detect whether a collision occurs, the r bh is the difference between the calculated height of the body and the target height, the r da is the joint acceleration penalty, the r so To calculate the difference between the joint position and the soft limit, the λ wi is the scaling factor of the i-th reward function value; in, ; ; ; ; ; ; ; is the actual simulation speed, is the target speed; is the simulated angular velocity, Reduce to the target angular velocity; is the simulated leg-lifting time, is the target time; r cc To judge whether there is a collision; is the height in the simulation, is the target height; is the acceleration of 12 joints; is the difference between the current position and the target position, and p is the limit value configured in the robot information; The cost constraint function includes: C = c p + c t + c d Among them, the c p To calculate the difference between the joint position and the first limiting threshold, the c t To calculate the difference between the joint torque and the second limit threshold, the c d To calculate the difference between the joint velocity and the third limiting threshold, , and is the corresponding proportionality coefficient.
7. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Quadruped robot fall-down self-reset control method based on deep reinforcement learning
CN110861084A
Motion control method and system for quadruped robot under terrain subareas
CN118192254A
Cited By
Humanoid robot training system and humanoid robot control system
CN121680167A
Four-wheeled robot reinforcement learning control method and system based on safety barrier function and multi-modal observation
CN122411016A