Biped robot gait network training method

Through the dual-channel deep reinforcement learning architecture and curriculum learning method, the problems of low sample efficiency and poor stability in the gait training of bipedal robots were solved, the robot's adaptability and strategy robustness in complex environments were improved, and efficient gait training effects were achieved.

CN120722767AActive Publication Date: 2025-09-30HUNAN UNIV

Patent Information

Application Number
CN202511238850.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-09-30
Estimated Expiration
2045-09-01

AI Technical Summary

Technical Problem

In the existing technology, bipedal robot gait training has problems such as low sample efficiency, insufficient exploration, poor stability and hyperparameter sensitivity. It is difficult to perform stably in complex environments, and the gap between the simulation environment and the real environment is large, making it difficult to transfer training results.

Method used

A dual-channel deep reinforcement learning architecture is adopted, including a main network and an opponent network. Parameters are optimized through KL divergence and compound loss function, and a perturbation mechanism and hyperparameter memory curve adjustment are introduced. Combined with course learning, the robot's adaptability in complex environments is gradually improved.

Benefits of technology

It improves sample collection efficiency and training throughput, enhances the robustness and generalization ability of the strategy, reduces the complexity of the state space, achieves the stability and flexibility of the training process, and can effectively cope with complex environmental challenges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120722767A_ABST
    Figure CN120722767A_ABST
Patent Text Reader

Abstract

The invention discloses a biped robot gait network training method. The method comprises the following steps: constructing a dual-channel deep reinforcement learning architecture; the running states of the X biped robots in the simple terrain of the simulation environment are collected; obtaining a current reward according to the current running state, merging each piece of information into a Markov decision process, and storing the information into an experience playback area; when the number of the Markov decision-making processes in the experience playback area is greater than a preset threshold value n, randomly taking a preset number of Markov decision-making processes from the Markov decision-making processes, and updating parameters of the main network and the opponent network; disturbance is carried out on main network parameters, and a human memory curve is simulated to carry out continuous adjustment on main network hyper-parameters clip; the biped robot with the stable walking duration reaching the preset duration is moved to the terrain with the higher difficulty level, and the network parameter updating process is repeated; and course learning is continuously carried out until the accumulated reward information and the stable walking duration of all the biped robots reach preset values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot control, and in particular to a gait network training method for a biped robot. Background Art

[0002] In recent years, with the rapid development of deep reinforcement learning technology, using reinforcement learning to train the gait and motion control of bipedal robots has become a research hotspot. Traditional methods usually use a single main network to make decisions and use algorithms such as deep deterministic policy gradient (DDPG) and proximal policy optimization (PPO) to guide the robot to train in a simulation environment. The main process of these methods usually includes: (1) deploying a single or a small number of robot instances in the simulation platform and collecting their respective state data; (2) using a deep neural network (usually including actors and critics) to evaluate the robot's state and output action instructions; (3) after the robot executes the action, the reward information fed back by the environment (such as motion tracking, stability, energy efficiency, etc.) is used to adjust the network parameters and gradually optimize the control strategy.

[0003] While existing training methods have advanced bipedal robot gait control technology to some extent, several challenges remain in practical applications. Traditional methods often employ a single-channel main network architecture. Due to the high dimensionality of both the state and action spaces, sample utilization is low, training requires a large amount of simulation data, and convergence is slow. Relying solely on a single main network, training is prone to falling into local optima and lacks sufficient exploration. This is particularly true when faced with complex dynamic balance and changing terrain, where a single-channel network struggles to provide comprehensive exploration guidance, thus limiting the robot's motion flexibility and robustness. Current reinforcement learning algorithms exhibit instability when dealing with complex environmental disturbances, such as performance fluctuations in gait switching, impact absorption, and balance control. Furthermore, the "simulation-to-physical gap" between simulation and the real world makes it difficult to directly transfer training results to practical applications. Traditional algorithms are sensitive to the settings of hyperparameters (such as the clip parameter). Improper parameter adjustment can lead to training instability or performance degradation. Furthermore, there is a lack of effective control over the exploration strategy during the initial stages of training. Summary of the Invention

[0004] In view of this, the present invention provides a biped robot gait network training method to at least solve the problems of low sample efficiency, insufficient exploration, poor stability and hyperparameter sensitivity in the prior art.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions: A bipedal robot gait network training method comprises the following steps: S1. Based on the same-policy deep reinforcement learning PPO algorithm, a dual-channel deep reinforcement learning architecture is constructed as the bipedal robot gait network, including a main network and an opponent network with the same network structure but different initial parameters; S2. Collect the running status of X bipedal robots in a simple terrain in the simulation environment, and train the main network and the opponent network based on the running status; Construct a Markov decision process, obtain N samples from it, input the samples into the main network and the opponent network, calculate the KL divergence between the main network and the opponent network, and construct a composite loss function for the main network based on the mean squared error between the output action value and the target action value and the KL divergence; update the parameters of the main network according to the composite loss function; construct the loss function of the opponent network based on the KL divergence to update the parameters of the opponent network; apply perturbations to the main network, and update the hyperparameter clip value in the main network according to the forgetting curve; S3. When the preset performance indicators are met, the simple terrain is switched to complex terrain, and the main network and the opponent network are further trained until the cumulative reward information and stable walking time of all bipedal robots reach the preset network performance indicators, and a trained bipedal robot gait network is obtained.

[0006] Preferably, the specific content of the training includes:

[0007] S21. The collected current running state is input into the main network. The main network outputs an action command and converts it into a motor control vector, which is then sent back to the bipedal robot for execution. The next running state after the bipedal robot executes the action is further collected, and the current reward value is calculated based on the next running state and the preset target state. The current running state, action command, next running state, and current reward value are combined into a Markov decision process expressed as a tuple and stored in the experience replay area. S22. Repeat the contents of S21 until the number of Markov decision processes in the experience replay area is greater than the preset threshold n; S23. Randomly extract a preset number N of samples from the experience replay area, input the samples into the main network, calculate the output action value based on each set of action instructions and the current operating state, and calculate the target action value based on the current reward value and the action instruction output by the main network in the next operating state; calculate the KL divergence between the action probability distributions output by the main network and the opponent network under the same input, and construct a composite loss function for the main network based on the mean squared error between the output action value and the target action value and the KL divergence; update the parameters of the main network according to the composite loss function; and construct a loss function for the opponent network based on the KL divergence to update the parameters of the opponent network; S24. Apply perturbations to the main network and update the hyperparameter clip values ​​in the main network according to the forgetting curve. After completing the parameter perturbation and hyperparameter update, use the updated main network to regenerate the motor control vector and return it to the bipedal robot for execution.

[0008] Preferably, the network structures of the main network and the opponent network both include a strategy sub-network and a value sub-network, wherein the strategy sub-network and the value sub-network of the main network are respectively recorded as the first strategy network and the first value network, and the strategy sub-network and the value sub-network of the opponent network are respectively recorded as the second strategy network and the second value network; The first policy network is used to interact with the parallel simulation environment, receive the reduced and normalized operating state, output the probability distribution of the motor control action based on the current operating state, and generate control instructions based on the probability distribution; the first value network is used to estimate the expected return of the operating state, participate in the calculation of the advantage function and the construction of the PPO loss, thereby realizing the optimization of the main network strategy; Without directly interacting with the simulation environment, the second policy network receives the same operating state as the first policy network, and independently generates an action probability distribution. By comparing the probability distributions output by the first and second policy networks, the KL divergence between them is calculated, and the KL divergence is used as a loss term to update the parameters of the second policy network. During the optimization process, the first policy network and the first value network simultaneously integrate the PPO algorithm loss and KL divergence for joint gradient update.

[0009] Preferably, the running state of the bipedal robot includes the position and speed information of the freely movable joints of the bipedal robot; the reward information includes the bipedal robot's motion tracking performance reward, stability and balance reward, gait and contact reward, safety and hardware protection reward, motion efficiency and smoothness reward, command response and behavior reward.

[0010] Preferably, each strategy sub-network and value sub-network in the main network and the opponent network adopts a multi-layer perceptron structure, including an input layer, one or more hidden layers, and an output layer; The input layer is used to receive the current operating state and historical action instructions of the bipedal robot after dimensionality reduction processing, and preliminarily encode the input data before passing it to the hidden layer; The hidden layer consists of several fully connected neurons, which use activation functions to perform nonlinear transformations on the input signal to extract features of the complex correlation between state and action. The output layer is used to generate the policy distribution, that is, output the estimated value or probability distribution of the corresponding action based on the feature information extracted by the hidden layer, which is used to guide the robot to generate motor control instructions; The dimension of the input layer is equal to the state dimension of the biped robot after dimensionality reduction, and the dimension of the output layer is equal to the number of joints controlled by the robot. The input layer dimension is consistent with the policy network, and the output layer dimension is 1, which is used to estimate the state value corresponding to the current state.

[0011] Preferably, the composite loss function of the main network is: ; ; ; ; in, represents the total loss, represents the clipping constraint loss updated by the first policy network, represents the first value network loss, represents the first value network loss weight, represents the KL divergence, represents the KL divergence loss weight, Represents the time step The expected value of a sample of the Markov decision process on , Represents the probability ratio of the new and old strategies, where the new strategy refers to the updated strategy network in the current training process, and the old strategy refers to the strategy network used to generate the current training data, that is, the strategy before the update. represents the generalized advantage estimate, clip represents the pruning hyperparameter in the PPO algorithm, Indicates that at time step The clipping threshold when Indicates the operating status of the first value network The predicted value of represents the target value estimated by discounted reward and generalized advantage, represents the standard deviation of the action probability distribution output by the second policy network, represents the standard deviation of the action probability distribution output by the first policy network, represents the mean of the action probability distribution output by the second policy network, Represents the mean of the action probability distribution output by the first policy network.

[0012] Preferably, the loss function of the adversary network is: ; in, represents the KL divergence, represents the standard deviation of the action probability distribution output by the second policy network, represents the standard deviation of the action probability distribution output by the first policy network, represents the mean of the action probability distribution output by the second policy network, Represents the mean of the action probability distribution output by the first policy network.

[0013] Preferably, the method for perturbing the main network parameters in S24 is: ; ; in, represents the updated network parameters, Indicates the network parameters before the update, represents the update parameter discount, represents the noise term that conforms to the Gaussian distribution, represents a Gaussian distribution, Indicates the noise disturbance intensity.

[0014] Preferably, the specific content of continuously adjusting the main network hyperparameter clip by simulating the human memory curve in S24 includes: The hyperparameter clip of the main network is the clipping threshold in the PPO algorithm The hyperparameter clip is used to control the amplitude of the policy update to balance the stability and efficiency of training. The method for adjusting the hyperparameter clip by simulating the human memory curve is as follows: ; in, is the updated clip value at time t, is the initial clip value, is the decay period, is the decay rate.

[0015] The preferred course learning method is: ; ; in, Bipedal robot Whether the target distance indication of the current difficulty terrain has been completed, Bipedal robot The maximum distance traveled on the current terrain difficulty. The distance threshold set for the current difficulty terrain. Bipedal robot The difficulty level of the current terrain, The degree to which the terrain difficulty increases.

[0016] Preferably, in S3, when the preset performance indicators are met, the simple terrain is switched to the complex terrain, and the main network and the opponent network are further trained: Training a bipedal robot in a simulation environment to achieve a preset performance indicator of stable walking under the current terrain difficulty level. That is, during continuous training, the simulated walking time continuously exceeds the preset stable duration threshold. When the condition is met, the simulation environment is automatically switched to complex terrain. In complex terrain, further training of the main network and the adversary network continues.

[0017] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a biped robot gait network training method, which has the following beneficial effects: The present invention is applied to a bipedal robot gait training system based on deep reinforcement learning. By deploying multiple bipedal robots for parallel training in a simulation environment, the sample collection efficiency and training throughput are greatly improved, and the time required for policy convergence is significantly shortened. The dual-channel architecture consisting of the main network and the opponent network not only maintains the consistency of policy learning, but also introduces an enhanced exploration mechanism to effectively avoid falling into local optimality, thereby improving the robustness and generalization ability of the policy. The dimensionality reduction and normalization of the running state reduce the complexity of the state space and improve the efficiency of network learning. At the same time, the refined multi-dimensional reward function comprehensively covers the key performance indicators in gait control, making the training process more directional and accurate. The introduction of a perturbation mechanism and dynamic adjustment of the memory curve of the clip hyperparameter during the main network parameter update process further enhances the stability and flexibility of the training. Through the course learning mechanism, robots that already have stable walking capabilities are gradually transferred to more difficult terrain for continuous training, achieving layered and progressive capability improvement and effectively responding to the challenges of environmental complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 A schematic diagram of the overall model structure of a biped robot gait network training method provided by the present invention; Figure 2 A flowchart of a bipedal robot gait network training method is provided; Figure 3 Schematic diagram of the structure of the primary network and the opponent network and the perturbation method in an embodiment of the present invention; Figure 4Schematic diagram of a hyperparameter memory curve adjustment method in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] The present invention provides a biped robot gait network training method, such as Figure 1-2 As shown, the following steps are included: S1. Based on the same-policy deep reinforcement learning PPO algorithm, a dual-channel deep reinforcement learning architecture is constructed as the bipedal robot gait network, including a main network and an opponent network with the same network structure but different initial parameters; S2. Collect the running status of X bipedal robots in a simple terrain in the simulation environment as the environmental interaction data in the initial training phase, and train the main network and the opponent network based on the running status; Construct a Markov decision process, obtain N samples from it, input the samples into the main network and the opponent network, calculate the KL divergence between the main network and the opponent network, and construct a composite loss function for the main network based on the mean squared error between the output action value and the target action value and the KL divergence; update the parameters of the main network according to the composite loss function; construct the loss function of the opponent network based on the KL divergence to update the parameters of the opponent network; apply perturbations to the main network, and update the hyperparameter clip value in the main network according to the forgetting curve; S3. When the preset performance indicators are met, the simple terrain is switched to complex terrain, and the main network and the opponent network are further trained until the cumulative reward information and stable walking time of all bipedal robots reach the preset network performance indicators, and a trained bipedal robot gait network is obtained.

[0022] In order to further implement the above technical solutions, the specific content of the training includes: S21. The collected current running state is input into the main network. The main network outputs an action command and converts it into a motor control vector, which is then sent back to the bipedal robot for execution. The next running state after the bipedal robot executes the action is further collected. The current reward value is calculated based on indicators such as the deviation between the next running state and the preset target state, motion stability, energy consumption, or gait continuity. The current running state, action command, next running state, and current reward value are combined into a Markov decision process expressed as a tuple and stored in the experience replay area. S22. Repeat the contents of S21 until the number of Markov decision processes in the experience replay area is greater than the preset threshold n; S23. Randomly extract a preset number N of samples from the experience replay area, input the samples into the main network, calculate the output action value based on each set of action instructions and the current operating state, and calculate the target action value based on the current reward value and the action instruction output by the main network in the next operating state; calculate the KL divergence between the action probability distributions output by the main network and the opponent network under the same input, and construct a composite loss function for the main network based on the mean squared error between the output action value and the target action value and the KL divergence; update the parameters of the main network according to the composite loss function; and construct a loss function for the opponent network based on the KL divergence to update the parameters of the opponent network; S24. During the main network parameter update process, apply a small perturbation to the main network to enhance the policy's exploration capabilities. This perturbation acts on the main network's parameters, giving it stronger local exploration capabilities in the policy space. Apply the perturbation to the main network and update the value of the hyperparameter clip in the main network according to the forgetting curve, causing this parameter to gradually decrease during training. This tightens the clipping range during the policy update process, limits the policy update amplitude, and controls policy fluctuations. After completing the parameter perturbation and hyperparameter update, use the updated main network to regenerate the motor control vector and return it to the bipedal robot for execution.

[0023] It should be noted that the execution subject of the above method may be a computer device.

[0024] In this embodiment, simple terrain refers to flat ground without obstacles, and complex terrain refers to adding terrain elements such as stones, ups and downs, and slopes to the basic flat ground. In actual use, the complex terrain can be used to train the bipedal robot gait network step by step according to the level of difficulty.

[0025] The bipedal robot's simulation environment was constructed using Isaac Gym, which can be configured to accommodate terrains of varying difficulty. The bipedal robot's simulation model was constructed using the Unified Robot Description Format (URDF), which includes a kinematic and dynamic description of the bipedal robot model, its geometric representation, and its collision model. A dual-channel deep reinforcement learning network architecture based on the PPO algorithm was constructed, in which the policy and value subnetworks of both the main and opponent networks are five-layer multilayer perceptrons.

[0026] The simulation model of the biped robot interacts in the simulation environment and can sample the running state of the biped robot. The running state is the position and velocity information of each joint of the biped robot. After all the information is reduced in dimension and normalized, the running state space is obtained. ,Will Input to the constructed main network Actor layer to get the action output , and convert it into a motor control vector and feed it back to the biped robot. The biped robot interacts with the environment through the control vector to obtain the next operating state space And get multi-dimensional reward information , construct all the above information into a Markov decision process tuple , and store the tuple in the experience replay pool.

[0027] The course learning mechanism is continuously implemented. Through multiple rounds of simulation training and gradually increasing terrain difficulty, the gait stability and adaptability of the bipedal robot in complex environments are continuously improved until the cumulative reward information and stable walking time of all bipedal robots reach the preset performance indicators, completing the entire training process.

[0028] In order to further implement the above technical solution, the network structures of the main network and the opponent network both include a strategy sub-network and a value sub-network, wherein the strategy sub-network and the value sub-network of the main network are respectively recorded as the first strategy network and the first value network, and the strategy sub-network and the value sub-network of the opponent network are respectively recorded as the second strategy network and the second value network; The first policy network is used to interact with the parallel simulation environment, receive the reduced and normalized operating state, output the probability distribution of the motor control action based on the current operating state, and generate control instructions based on the probability distribution; the first value network is used to estimate the expected return of the operating state, participate in the calculation of the advantage function and the construction of the PPO loss, thereby realizing the optimization of the main network strategy; Without directly interacting with the simulation environment, the second policy network receives the same operating state as the first policy network, and independently generates an action probability distribution. By comparing the probability distributions output by the first and second policy networks, the KL divergence between them is calculated, and the KL divergence is used as a loss term to update the parameters of the second policy network. During the optimization process, the first policy network and the first value network simultaneously integrate the PPO algorithm loss and KL divergence for joint gradient update.

[0029] In order to further implement the above technical solution, the operating state of the bipedal robot includes the position and speed information of the freely movable joints of the bipedal robot; the reward information includes the bipedal robot's motion tracking performance reward, stability and balance reward, gait and contact reward, safety and hardware protection reward, motion efficiency and smoothness reward, command response and behavior reward.

[0030] To further implement the above technical solution, each strategy sub-network and value sub-network in the main network and the opponent network adopts a multi-layer perceptron structure, including an input layer, one or more hidden layers, and an output layer; The input layer is used to receive the current operating state and historical action instructions of the bipedal robot after dimensionality reduction processing, and preliminarily encode the input data before passing it to the hidden layer; The hidden layer consists of several fully connected neurons, which use activation functions to perform nonlinear transformations on the input signal to extract features of the complex correlation between state and action. The output layer is used to generate the policy distribution, that is, output the estimated value or probability distribution of the corresponding action based on the feature information extracted by the hidden layer, which is used to guide the robot to generate motor control instructions; The dimension of the input layer is equal to the state dimension of the biped robot after dimensionality reduction, and the dimension of the output layer is equal to the number of joints controlled by the robot. The input layer dimension is consistent with the policy network, and the output layer dimension is 1, which is used to estimate the state value corresponding to the current state.

[0031] To further implement the above technical solution, the composite loss function of the main network is: ; ; ; ; in, represents the total loss, represents the clipping constraint loss updated by the first policy network, represents the first value network loss, represents the first value network loss weight, represents the KL divergence, represents the KL divergence loss weight, Represents the time step The expected value of a sample of the Markov decision process on , Represents the probability ratio of the new and old strategies, where the new strategy refers to the updated strategy network in the current training process, and the old strategy refers to the strategy network used to generate the current training data, that is, the strategy before the update. represents the generalized advantage estimate, clip represents the pruning hyperparameter in the PPO algorithm, Indicates that at time step The clipping threshold when Indicates the operating status of the first value network The predicted value of represents the target value estimated by discounted reward and generalized advantage, represents the standard deviation of the action probability distribution output by the second policy network, represents the standard deviation of the action probability distribution output by the first policy network, represents the mean of the action probability distribution output by the second policy network, Represents the mean of the action probability distribution output by the first policy network.

[0032] To further implement the above technical solution, the loss function of the adversary network is: ; in, represents the KL divergence, represents the standard deviation of the action probability distribution output by the second policy network, represents the standard deviation of the action probability distribution output by the first policy network, represents the mean of the action probability distribution output by the second policy network, Represents the mean of the action probability distribution output by the first policy network.

[0033] In order to further implement the above technical solution, the method of perturbing the main network parameters in S24 is: ; ; in, represents the updated network parameters, Indicates the network parameters before the update, represents the update parameter discount, represents the noise term that conforms to the Gaussian distribution, represents a Gaussian distribution, Indicates the noise disturbance intensity.

[0034] To further implement the above technical solution, the specific contents of S24 to simulate the human memory curve to continuously adjust the main network hyperparameter clip include: The hyperparameter clip of the main network is the clipping threshold in the PPO algorithm The hyperparameter clip is used to control the amplitude of the policy update to balance the stability and efficiency of training. The method for adjusting the hyperparameter clip by simulating the human memory curve is as follows: ; in, is the updated clip value at time t, is the initial clip value, is the decay period, is the decay rate.

[0035] It should be noted that: During the main network parameter update process, a perturbation mechanism is introduced to moderately intervene in its parameters to enhance the policy's exploration capabilities. By simulating the human memory forgetting curve, the key hyperparameter in the main network, the clip parameter, is dynamically adjusted to achieve a parameter optimization process that is more consistent with learning principles. This mechanism helps the model maintain sensitivity to key policy changes during long-term training and avoid falling into local optimal solutions.

[0036] Based on this, a curriculum learning approach was employed to gradually migrate a bipedal robot, already capable of stably achieving a preset walking duration in a simulation environment, to terrain environments with higher difficulty levels. During this migration, the parameter update process for both the primary and adversary networks was repeated, ensuring the model's adaptability and generalization capabilities for increasingly complex tasks.

[0037] In order to further implement the above technical solutions, the course learning method is: ; ; in, Bipedal robot Whether the target distance indication of the current difficulty terrain has been completed, Bipedal robot The maximum distance traveled on the current terrain difficulty. The distance threshold set for the current difficulty terrain. Bipedal robot The difficulty level of the current terrain, The degree to which the terrain difficulty increases.

[0038] It should be noted that: This course-based training process will continue until all bipedal robot individuals can reach the established cumulative reward threshold and stable walking time targets in all terrain levels, thereby achieving the full achievement of training goals and improving the robustness and environmental adaptability of the overall system.

[0039] like Figure 3As shown, the Actor and Critic subnetworks of the main and opponent networks are five-layer multilayer perceptrons, consisting of one input layer, three hidden layers, and one output layer. The dimension of the Actor network's input layer is the number of operational states obtained after dimensionality reduction, and the dimension of the output layer is the number of joints of the bipedal robot. The dimension of the Critic network's input layer is the number of operational states obtained after dimensionality reduction, and the dimension of the output layer is 1. The hidden layer dimensions of the Actor and Critic networks are 512, 256, and 128, respectively. After each round of main network parameter updates, a parameter perturbation operation is introduced to enhance the strategy's adaptability to plasticity loss caused by terrain changes.

[0040] like Figure 4 As shown, in the PPO (Proximal Policy Optimization) algorithm, the clip hyperparameter is one of its core designs. It is used to construct the clipped surrogate objective to limit the amplitude of policy updates, thereby ensuring the stability and reliability of the training process. This embodiment provides a method for adjusting hyperparameters based on the human memory curve. At the beginning, the hyperparameter clip will be set relatively large to encourage the policy to conduct more extensive exploration in the early stages of training. As the training progresses, the clip value will gradually decrease according to the decay function that simulates the forgetting law of human memory, thereby strengthening the stability and fine optimization of the learned policy in the later stages, and improving the convergence performance and generalization ability of the policy in complex environments.

[0041] To further implement the above technical solution, in S3, when the preset performance indicators are met, the simple terrain is switched to complex terrain to further train the main network and the opponent network: Training a bipedal robot in a simulation environment to achieve a preset performance indicator of stable walking under the current terrain difficulty level. That is, during continuous training, the simulated walking time continuously exceeds the preset stable duration threshold. When the condition is met, the simulation environment is automatically switched to complex terrain. In complex terrain, further training of the main network and the adversary network continues.

[0042] The present invention is applied to a bipedal robot gait training system based on deep reinforcement learning. By deploying multiple bipedal robots for parallel training in a simulation environment, the sample collection efficiency and training throughput are greatly improved, and the time required for policy convergence is significantly shortened. The dual-channel architecture consisting of the main network and the opponent network not only maintains the consistency of policy learning, but also introduces an enhanced exploration mechanism to effectively avoid falling into local optimality, thereby improving the robustness and generalization ability of the policy. The dimensionality reduction and normalization of the running state reduce the complexity of the state space and improve the efficiency of network learning. At the same time, the refined multi-dimensional reward function comprehensively covers the key performance indicators in gait control, making the training process more directional and accurate. The introduction of a perturbation mechanism and dynamic adjustment of the memory curve of the clip hyperparameter during the main network parameter update process further enhances the stability and flexibility of the training. Through the course learning mechanism, robots that already have stable walking capabilities are gradually transferred to more difficult terrain for continuous training, achieving layered and progressive capability improvement and effectively responding to the challenges of environmental complexity.

[0043] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A bipedal robot gait network training method, characterized in that: The following steps are involved: S1. Based on the same-policy deep reinforcement learning PPO algorithm, a dual-channel deep reinforcement learning architecture is constructed as the bipedal robot gait network, including a main network and an opponent network with the same network structure but different initial parameters; S2. Collect the running status of X bipedal robots in a simple terrain in the simulation environment, and train the main network and the opponent network based on the running status; The specific content of the training includes: Construct a Markov decision process, obtain N samples from it, input the samples into the main network and the opponent network, calculate the KL divergence between the main network and the opponent network, and construct a composite loss function for the main network based on the mean squared error between the output action value and the target action value and the KL divergence; update the parameters of the main network according to the composite loss function; construct the loss function of the opponent network based on the KL divergence to update the parameters of the opponent network; apply perturbations to the main network, and update the hyperparameter clip value in the main network according to the forgetting curve; S3. When the preset performance indicators are met, the simple terrain is switched to complex terrain, and the main network and the opponent network are further trained until the cumulative reward information and stable walking time of all bipedal robots reach the preset network performance indicators, and a trained bipedal robot gait network is obtained.

2. A biped robot gait network training method according to claim 1, characterized in that: The specific content of training in S2 includes: S21. The collected current running state is input into the main network. The main network outputs an action command and converts it into a motor control vector, which is then sent back to the bipedal robot for execution. The next running state after the bipedal robot executes the action is further collected, and the current reward value is calculated based on the next running state and the preset target state. The current running state, action command, next running state, and current reward value are combined into a Markov decision process expressed as a tuple and stored in the experience replay area. S22. Repeat the contents of S21 until the number of Markov decision processes in the experience replay area is greater than the preset threshold n; S23. Randomly extract a preset number N of samples from the experience replay area, input the samples into the main network, calculate the output action value based on each set of action instructions and the current operating state, and calculate the target action value based on the current reward value and the action instruction output by the main network in the next operating state; calculate the KL divergence between the action probability distributions output by the main network and the opponent network under the same input, and construct a composite loss function for the main network based on the mean squared error between the output action value and the target action value and the KL divergence; update the parameters of the main network according to the composite loss function; and construct a loss function for the opponent network based on the KL divergence to update the parameters of the opponent network; S24. Apply perturbations to the main network and update the hyperparameter clip values ​​in the main network according to the forgetting curve. After completing the parameter perturbation and hyperparameter update, use the updated main network to regenerate the motor control vector and return it to the bipedal robot for execution.

3. A biped robot gait network training method according to claim 1, characterized in that: The network structures of the main network and the opponent network both include a strategy sub-network and a value sub-network. The strategy sub-network and value sub-network of the main network are respectively recorded as the first strategy network and the first value network, and the strategy sub-network and value sub-network of the opponent network are respectively recorded as the second strategy network and the second value network. The first policy network is used to interact with the parallel simulation environment, receive the reduced and normalized operating state, output the probability distribution of the motor control action based on the current operating state, and generate control instructions based on the probability distribution; the first value network is used to estimate the expected return of the operating state, participate in the calculation of the advantage function and the construction of the PPO loss, thereby realizing the optimization of the main network strategy; Without directly interacting with the simulation environment, the second policy network receives the same operating state as the first policy network, and independently generates an action probability distribution. By comparing the probability distributions output by the first and second policy networks, the KL divergence between them is calculated, and the KL divergence is used as a loss term to update the parameters of the second policy network. During the optimization process, the first policy network and the first value network simultaneously integrate the PPO algorithm loss and KL divergence for joint gradient update.

4. A biped robot gait network training method according to claim 2, characterized in that: The running status of the bipedal robot includes the position and speed information of the freely movable joints of the bipedal robot; the reward information includes the bipedal robot's motion tracking performance reward, stability and balance reward, gait and contact reward, safety and hardware protection reward, motion efficiency and smoothness reward, command response and behavior reward.

5. A biped robot gait network training method according to claim 3, characterized in that: Each strategy sub-network and value sub-network in the main network and the opponent network adopts a multi-layer perceptron structure, including an input layer, one or more hidden layers, and an output layer; The input layer is used to receive the current operating state and historical action instructions of the bipedal robot after dimensionality reduction processing, and preliminarily encode the input data before passing it to the hidden layer; The hidden layer consists of several fully connected neurons, which use activation functions to perform nonlinear transformations on the input signal to extract features of the complex correlation between state and action. The output layer is used to generate the policy distribution, that is, output the estimated value or probability distribution of the corresponding action based on the feature information extracted by the hidden layer, which is used to guide the robot to generate motor control instructions; The dimension of the input layer is equal to the state dimension of the biped robot after dimensionality reduction, and the dimension of the output layer is equal to the number of joints controlled by the robot. The input layer dimension is consistent with the policy network, and the output layer dimension is 1, which is used to estimate the state value corresponding to the current state.

6. A biped robot gait network training method according to claim 3, characterized in that: The composite loss function of the main network is: ; ; ; ; in, represents the total loss, represents the clipping constraint loss updated by the first policy network, represents the first value network loss, represents the first value network loss weight, represents the KL divergence, represents the KL divergence loss weight, Represents the time step The expected value of a sample of the Markov decision process on , Represents the probability ratio of the new and old strategies, where the new strategy refers to the updated strategy network in the current training process, and the old strategy refers to the strategy network used to generate the current training data, that is, the strategy before the update. represents the generalized advantage estimate, clip represents the pruning hyperparameter in the PPO algorithm, Indicates that at time step The clipping threshold when Indicates the operating status of the first value network The predicted value of represents the target value estimated by discounted reward and generalized advantage, represents the standard deviation of the action probability distribution output by the second policy network, represents the standard deviation of the action probability distribution output by the first policy network, represents the mean of the action probability distribution output by the second policy network, Represents the mean of the action probability distribution output by the first policy network.

7. A biped robot gait network training method according to claim 3, characterized in that: The loss function of the adversary network is: ; in, represents the KL divergence, represents the standard deviation of the action probability distribution output by the second policy network, represents the standard deviation of the action probability distribution output by the first policy network, represents the mean of the action probability distribution output by the second policy network, Represents the mean of the action probability distribution output by the first policy network.

8. A biped robot gait network training method according to claim 2, characterized in that: The method of perturbing the main network parameters is: ; ; in, represents the updated network parameters, Indicates the network parameters before the update, represents the update parameter discount, represents the noise term that conforms to the Gaussian distribution, represents a Gaussian distribution, represents the noise disturbance intensity; The specific contents of simulating the human memory curve to continuously adjust the main network hyperparameter clip include: The hyperparameter clip of the main network is the clipping threshold in the PPO algorithm The hyperparameter clip is used to control the amplitude of the policy update to balance the stability and efficiency of training. The method for adjusting the hyperparameter clip by simulating the human memory curve is as follows: ; in, is the updated clip value at time t, is the initial clip value, is the decay period, is the decay rate.

9. A biped robot gait network training method according to claim 1, characterized in that: The course learning method is: ; ; in, Bipedal robot Whether the target distance indication of the current difficulty terrain has been completed, Bipedal robot The maximum distance traveled on the current terrain difficulty. The distance threshold set for the current difficulty terrain. Bipedal robot The difficulty level of the current terrain, The degree to which the terrain difficulty increases.

10. A biped robot gait network training method according to claim 1, characterized in that: In S3, when the preset performance indicators are met, the simple terrain is switched to complex terrain to further train the main network and the opponent network: Training a bipedal robot in a simulation environment to achieve a preset performance indicator of stable walking under the current terrain difficulty level. That is, during continuous training, the simulated walking time continuously exceeds the preset stable duration threshold. When the condition is met, the simulation environment is automatically switched to complex terrain. In complex terrain, further training of the main network and the adversary network continues.

Citation Information

Patent Citations

  • Rapid path planning method based on variant dual DQNs (deep Q-networks) and mobile robot

    CN108375379A

  • Biped robot walking stability optimization method based on improved PPO algorithm

    CN114839878A

  • Strategy network training method and humanoid biped robot gait control method

    CN117555339A

  • Vector propeller control system and method based on flexible shaft

    CN117666355A

  • Multi-mobile robot autonomous obstacle avoidance method based on deep reinforcement learning

    CN117873116A

Cited By

  • Robot reinforcement learning motion control method and system

    CN121680293A

  • Control strategy learning method and system of biped wheeled robot and medium

    CN121704197A