Robot gait training method and system based on reinforcement learning

By optimizing robot gait training through imitation learning initialization, dynamic reward function and cognitive feedback mechanism, the problems of unstable strategy and poor human acceptance in existing methods are solved, and efficient, stable gait control and natural behavior are achieved.

CN120630670APending Publication Date: 2025-09-12NANJING KANGLONGWEI TECH IND CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510581188.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing robot gait training methods have problems such as unstable initialization strategies, rough reward function design, lack of scheduling mechanism in the training process, and non-cognitive-friendly modeling, resulting in low training efficiency, unstable strategy performance, and poor human acceptance.

Method used

Imitation learning initialization, dynamic reward function, adaptive environment scheduling and cognitive feedback mechanism are adopted. Pre-training is carried out by collecting expert motion trajectories, constructing multi-factor reward function and comprehensive performance indicators, combining with PPO algorithm for strategy optimization, and introducing cognitive load indicator to adjust gait strategy.

Benefits of technology

It significantly improves the convergence speed and stability of the gait strategy, enhances the quality and practicality of the gait strategy, enhances the robustness and human acceptance in complex environments, and is suitable for multi-degree-of-freedom robot platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120630670A_ABST
    Figure CN120630670A_ABST
Patent Text Reader

Abstract

The invention discloses a robot gait training method and system based on reinforcement learning, and relates to the technical field of electric vehicle charging, and the method comprises the steps: carrying out the simulation learning initialization of a strategy based on a reference video; selecting a plurality of evaluation indexes to design a dynamic reward function so as to guide the initial strategy network to optimize the gait performance; through an environment difficulty scheduling mechanism, the training environment difficulty is automatically adjusted according to strategy performance; performing optimization training on the strategy network by using a PPO algorithm; and judging the optimized gait strategy by constructing a cognitive load index, and outputting a final gait strategy. According to the method, imitation learning, self-adaptive reward modeling, multi-stage scheduling, reinforcement learning optimization and cognitive feedback are organically fused, and a robot gait training framework with generalization ability, stability and social adaptability is constructed. According to the method, the naturalness, the stability and the man-machine friendliness of the gait of the robot can be remarkably improved under complex terrains and interaction scenes, and the method has wide application prospects and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot gait training, and in particular to a robot gait training method and system based on reinforcement learning. Background Art

[0002] In recent years, with the continuous development of mobile robot technology, its application in a variety of complex environments, such as inspection, logistics, rehabilitation, and rescue, has gradually expanded. As one of the core motion control technologies, robot gait control is crucial for achieving smooth walking, flexible obstacle avoidance, and adapting to environmental changes. For legged robots in particular, optimizing their gait strategy not only affects mobility performance but also directly impacts the stability and safety of their mission execution. Existing gait control methods primarily include trajectory planning methods based on dynamic models, rule-driven finite state control methods, and the recently widely used policy optimization methods based on deep reinforcement learning. Trajectory planning and model control methods, among others, achieve controllable regulation of the robot's trajectory by establishing a mechanical model of the system. These methods, with their clear structure and easy explanation, have been successfully applied in many scenarios with regular structures and well-defined environments. Furthermore, reinforcement learning methods, by obtaining feedback and rewards from continuous interaction with the environment, directly learn policies from perception to action decisions. These methods possess strong environmental adaptability and can demonstrate flexibility and self-organization in irregular terrain and complex tasks. Based on reinforcement learning, researchers have proposed a variety of policy optimization methods, such as DDPG, PPO, and SAC, and combined them with imitation learning, reward function design, and policy stability constraints to further improve training effectiveness. In simulations and some actual platforms, these methods have achieved positive results in improving gait stability and reducing energy consumption.

[0003] However, although reinforcement learning has become an important trend in robot gait training, existing methods still face the following technical challenges: First, most reinforcement learning methods start with optimization from random strategies. In the early stages, due to sparse rewards, robots are prone to frequent falls, which may lead to low training efficiency. In addition, most existing methods use a fixed-structure reward function, which makes it difficult to dynamically adjust the focus at different training stages or with different target preferences, which may affect the continuous improvement of strategy performance. Second, the current training process has a relatively static scheduling of environmental complexity and lacks an automatic switching mechanism based on strategy quality assessment, which is not conducive to the robust generalization of strategies under multi-task or multi-scenario conditions. Third, when applied to public or human-machine integration scenarios, the visual naturalness of the robot's gait and the acceptability of its behavior have gradually become key, but the current training process has not fully considered the subjective evaluation factors of human observers.

[0004] Therefore, there is an urgent need for a gait training method that can integrate imitation learning initialization, dynamic reward adjustment, adaptive environment scheduling and cognitive feedback mechanism to improve the gait stability, behavioral rationality and social acceptance of robots in complex environments. Summary of the Invention

[0005] In view of the fact that the existing training methods for robot gait training have problems such as unstable initialization strategy, rough reward function design, lack of scheduling mechanism in the training process, and no cognitive-friendly modeling, the present invention is proposed.

[0006] Therefore, the problem to be solved by the present invention is how to provide a robot gait training method and system based on reinforcement learning.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0008] In the first aspect, an embodiment of the present invention provides a robot gait training method based on reinforcement learning, including: collecting expert motion trajectories and pre-training the initial policy network using behavioral cloning technology; selecting multiple evaluation indicators to design a dynamic reward function to guide the initial policy network to optimize gait performance; constructing an environmental difficulty index and a comprehensive strategy performance index to evaluate the performance of the current strategy in a specific environment, and adaptively adjusting the increase or decrease of the environmental difficulty based on the comprehensive performance index; combining the dynamic reward function with the adaptive adjustment mechanism, iteratively training the policy network based on the optimization objective function of the PPO algorithm to obtain an optimized gait strategy; judging the optimized gait strategy by constructing a cognitive load index, and outputting the final gait strategy.

[0009] As a preferred solution of the robot gait training method based on reinforcement learning described in the present invention, the training initial strategy network includes:

[0010] Collect reference videos containing natural gait, use a pose estimation algorithm to extract body key points in each image frame, and convert the key point trajectories into a state-action pair dataset;

[0011] The policy network is pre-trained using the behavior cloning method so that it can predict actions consistent with the reference trajectory based on the state.

[0012] As a preferred solution of the robot gait training method based on reinforcement learning described in the present invention, the construction of the comprehensive reward function includes:

[0013] Select multiple evaluation indicators covering stability, grounding status, energy efficiency, and speed deviation;

[0014] A comprehensive reward function is obtained by integrating multiple evaluation indicators such as stability, grounding status, energy efficiency, speed deviation, etc. in a weighted manner.

[0015] As a preferred solution of the robot gait training method based on reinforcement learning described in the present invention, the construction of comprehensive performance indicators includes:

[0016] Select multiple performance evaluation indicators including average round reward, fall frequency, task completion success rate, and action entropy;

[0017] Multiple performance evaluation indicators such as round average reward, fall frequency, task completion success rate, action entropy, etc. are integrated in a weighted manner to obtain a comprehensive performance indicator.

[0018] As a preferred solution of the robot gait training method based on reinforcement learning described in the present invention, the construction of the optimization objective function includes:

[0019] Construct a cutting function to constrain the degree of policy deviation at each step during policy update;

[0020] Based on the probability ratio of the current policy to the old policy, an optimization objective function is constructed for the reinforcement learning training process.

[0021] As a preferred solution of the robot gait training method based on reinforcement learning described in the present invention, the cognitive load index includes:

[0022] Collect and analyze cognitive signals from external observers’ images, voice, and facial expressions during gait training;

[0023] Extract multiple characteristic variables reflecting human cognitive responses, including avoidance behavior frequency, average gaze duration, negative speech, etc.

[0024] The avoidance behavior frequency, average gaze duration, and negative speech feature variables were integrated in a weighted manner to obtain the cognitive load index.

[0025] As a preferred embodiment of the robot gait training method based on reinforcement learning of the present invention, the determination of the optimized gait strategy includes:

[0026] comparing the cognitive load index with a cognitive threshold;

[0027] According to the combination of the comparison results, the optimized gait strategy is judged as having good human acceptability or poor human acceptability.

[0028] On the second aspect, in order to further solve the problems existing in robot gait training, the embodiment of the present invention provides a robot gait training system based on reinforcement learning, which includes: an imitation learning initialization module for extracting state-action pairs based on the input reference gait video; a dynamic reward construction module for constructing a composite reward function including multiple factors such as stability, energy consumption, grounding safety and speed deviation; an environment scheduling and performance evaluation module for dynamically adjusting the difficulty level of the training environment according to the training performance of the robot's current strategy; a strategy training module for iteratively optimizing the current strategy network using the PPO algorithm; and a cognitive feedback module for collecting cognitive signals and constructing cognitive feedback reward items during the human-computer interaction stage.

[0029] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the robot gait training method based on reinforcement learning as described in the first aspect of the present invention is implemented.

[0030] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the robot gait training method based on reinforcement learning as described in the first aspect of the present invention.

[0031] The beneficial effects of the present invention are:

[0032] 1. By introducing an imitation learning mechanism based on reference videos, this paper provides the policy network with reasonable action priors in the initial training stage. This avoids the problems of frequent falls and sparse rewards caused by traditional reinforcement learning that relies entirely on random exploration. It significantly accelerates the convergence of gait strategies and reduces energy consumption and risk in the early stages of training.

[0033] 2. The reward function designed in this paper adopts a multi-factor structure, covering multiple key indicators such as gait stability, energy consumption, ground contact status, and desired speed deviation. Through the adaptive adjustment mechanism of reward weights, the strategy can focus on the current weak performance indicators at different training stages, thereby improving the overall quality and practicality of the gait strategy.

[0034] 3. This invention introduces a "training stage scheduling mechanism based on environmental difficulty". By dynamically evaluating and regulating the training environment, the training difficulty is automatically regressed when the gait strategy performance does not meet the standard or falls into a local optimum. This ensures the consistency and robustness of the training process and effectively prevents strategy forgetting or performance degradation in complex environments.

[0035] 4. This invention introduces the PPO algorithm to train the strategy, combines the probability ratio constraint with the advantage function estimation, and ensures that the gait strategy optimization process maintains a stable improvement in each update step, avoiding the problems of strategy jitter or training non-convergence in traditional methods. It has a complete theoretical basis and excellent practical results.

[0036] 5. This invention innovatively introduces cognitive feedback rewards, collecting human cognitive signals such as avoidance behavior and voice reactions to the robot's gait, and quantifying them into negative reward feedback for strategy correction, guiding the robot's gait to evolve in a direction that is more natural and more in line with human cognitive preferences. This is suitable for application scenarios such as service robots and public space collaboration.

[0037] 6. This application defines comprehensive performance indicators for the strategy, including reward mean, fall rate, task completion rate, and action entropy, as the basis for judging stage evolution and regression. This gives the strategy training process clear evolution criteria and self-correction capabilities, improving the intelligence and adaptability of the entire system.

[0038] 7. Each sub-module proposed in this application (imitation learning, reward control, environment scheduling, PPO training, cognitive feedback) is modularly designed, which is convenient for engineering deployment in combination with different types of robot hardware platforms. It supports multiple robot forms such as multi-degree-of-freedom bipeds, quadrupeds, and wheel-leg fusion, and has broad engineering application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:

[0040] Figure 1 This is a flow chart for implementing the present invention in Example 1. DETAILED DESCRIPTION

[0041] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0042] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0043] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0044] Example 1

[0045] Reference Figure 1 This is the first embodiment of the present invention. This embodiment provides a robot gait training method based on reinforcement learning. It combines imitation learning initialization, dynamic reward function, adaptive environment scheduling, PPO training and cognitive feedback mechanism to achieve robust optimization and human-machine friendly adaptation of gait behavior in complex terrain environments. The specific steps are as follows:

[0046] S1: Imitation learning initialization strategy based on reference video: By collecting expert action trajectories, the initial policy network is trained using behavior cloning technology, including the following sub-steps A1-A5:

[0047] A1: Collect reference videos containing natural gait and convert them into image frame sequences;

[0048] A2: Use a posture estimation algorithm to extract body key points in each image frame, output 3D joint coordinates, and form a key point trajectory sequence describing gait behavior according to the time sequence;

[0049] A3: Convert the key point trajectory into a state-action pair dataset The state includes joint position, velocity, and trunk orientation, and the action is the desired control target at the next moment;

[0050] A4: Use the behavior cloning method to pre-train the policy network so that it can predict actions consistent with the reference trajectory based on the state;

[0051] A5: Use the trained policy network parameters as the initial policy parameters for the reinforcement learning phase.

[0052] Specifically, the pre-training of the policy network using the behavior cloning method in A1-A5 can be expressed by the following formula:

[0053]

[0054] Where L BC represents the mean square error loss function, s * Indicates the status obtained by video analysis; a * represents the corresponding joint angle or motion target; π θ Represents the policy network to be trained.

[0055] S2: Construct a multi-factor dynamic reward function: Design a dynamic reward function to guide the initial policy network to optimize gait performance, including the following sub-steps B1-B5:

[0056] B1: Select multiple evaluation indicators covering stability (such as whether there is a fall), ground contact (the contact state between the foot and the ground), energy efficiency (control input power per unit time), and speed deviation (the difference between the execution speed and the target speed);

[0057] B2: Construct a comprehensive reward function based on the above multiple evaluation indicators;

[0058] B3: Adjust the weights of various reward items in real time based on the performance of the current initial strategy network. In the early stage, the proportion of stability rewards can be increased, and in the later stage, the incentives for grounding status or energy efficiency can be increased.

[0059] B4: Set negative penalties for behaviors such as falling, violent shaking, and large movements to force the strategy to avoid dangerous or non-physiological movements;

[0060] B5: Outputs a reward value that changes dynamically over time, which is used for strategy evaluation during reinforcement learning.

[0061] Specifically, the comprehensive reward function constructed in B1-B5 is expressed by the following formula:

[0062] r t =w1(t)·R stability +w2(t)·R EC +w3(t)·R GS +w4(t)·R SD ;

[0063] Where r t Represents the reward value of the current time step, which determines whether the current gait is good or not, and the goal is to maximize the reward; R stability Indicates the stability score of the current gait, which evaluates whether the robot maintains balance during the gait execution; R EC represents the energy consumption score of the front gait, which measures the energy consumed during gait; R GS To evaluate the contact status between the robot's foot and the ground, if the robot loses contact or the contact is unstable, the value is low; R SD It represents the difference between the speed during gait execution and the target speed, the smaller the better; w i (t) represents the dynamic weight of the reward item, which changes over time or training stage, indicating the importance of different tasks or goals. w1(t), w2(t), w3(t), and w4(t) are the weight coefficients of stability, energy consumption, foot contact state, and speed deviation, respectively;

[0064] It should be noted that this formula is the reward function for robot gait training, and its purpose is to adjust the output of the policy network according to various aspects of gait (stability, energy consumption, safety, and speed) to optimize the robot gait.

[0065] Among them, in order to adapt to the stage changes of the training process, a weight adaptation mechanism is introduced:

[0066]

[0067] Where, Δ i (t) represents the degree of convergence or volatility of the current reward item i, γ i To regulate the sensitivity of different reward items;

[0068] For example:

[0069] Initial stage: The stability term has poor convergence, Δ 稳定性 If it is larger, the system will automatically increase w1 and strengthen stability training;

[0070] Later stage: The robot can now walk stably, and the system gradually increases its attention to energy consumption and cognitive friendliness.

[0071] For example, the above parameters can be specifically expressed as:

[0072] R stability =-||c COM (t)-c ZMP (t)|| 2 ;

[0073] Where c COM (t) represents the current center of mass position; c ZMP (t) represents the zero moment point in the support surface;

[0074] It should be noted that the goal of this step is to reward the center of mass to fall stably within the support surface during gait, to prevent the robot from falling due to center of gravity shift, and thus to encourage the robot to learn to maintain center of gravity control and posture balance, and to improve movement stability.

[0075]

[0076] Where, represents the torque applied by the i-th joint; Indicates the angular velocity of the joint;

[0077] It should be noted that the goal of this step is to reward low-energy gaits, prevent the strategy from using violent and large-scale movements, and guide the gait to evolve in a smoother and more energy-efficient direction, thereby extending the endurance time and reducing the heat dissipation requirements in actual deployment.

[0078]

[0079] Where, δ j (t) indicates whether the jth foot end is in a suspended state (1 is suspended, 0 is grounded); D j (t)

[0080] Indicates the distance from the foot end to the ground when suspended;

[0081] It should be noted that the goal of this step is to reward the reasonable landing of the foot during gait, prevent problems such as slipping, accidental touch, and falling, and constrain the robot to make clear contact with the sole of the foot and form a support surface.

[0082] This improves gait stability in complex terrain.

[0083] R SD =-||v t -v targer || 2 ;

[0084] Where, v t Indicates the current speed of the robot; v targer Indicates the expected forward speed;

[0085] It should be noted that the goal of this step is to reward gaits that are close to the target speed, improve speed control ability, suppress unexpected acceleration or deceleration, and enhance path tracking ability, thereby maintaining consistency with the navigation system.

[0086] S3: Training phase scheduling mechanism based on environment difficulty: Adaptively adjust the training phase according to the environment complexity and gait performance indicators, including the following sub-steps C1-C5:

[0087] C1: Pre-build training environments with multiple difficulty levels, such as ground friction changes, obstacle interference, wind disturbances, etc., divided into three stages: low difficulty, medium difficulty, and high difficulty;

[0088] C2: Build comprehensive performance indicators for the current strategy (such as continuous stable walking time and fall frequency) and use them as the basis for judgment. When the comprehensive performance indicators reach the threshold, it automatically switches to the next stage;

[0089] C3: retains the current policy parameters during switching and maintains some of the previous stage disturbance to achieve a smooth transition;

[0090] C4: If continuous training failures occur in a more difficult environment, you can return to the previous stage to continue optimization;

[0091] C5: Decide whether to stay in the current environment or switch to the next training environment based on the performance of the strategy, and then schedule PPO to train and optimize the strategy network.

[0092] Specifically, the comprehensive performance indicators constructed in C1-C5 are reflected by the following formula:

[0093]

[0094] Where, Represents the average reward value of the last N rounds, which is obtained by r in step S2 t Get the average value; F t Expressed as the frequency of falls; S t Indicates the success rate of task completion; A t represents action entropy, which measures strategy diversity; η1, η2, η3, and η4 represent the average reward value, fall frequency, task completion success rate, and weight coefficient of action entropy, respectively, which are set by the staff.

[0095] For example, the above parameters can be specifically expressed as:

[0096]

[0097] Where A represents the set of all possible actions; π(a|s t ) represents a given state s t When the strategy selects action a, the probability of logπ(a|s t ) represents the logarithm of the action selection probability and is used to calculate the amount of information;

[0098] Environmental Difficulty Index:

[0099] E d =α1C dx +α2F rd +α3V bhl ;

[0100] This formula is used to calculate the difficulty index of the environment and determine the difficulty level of training; where E d Indicates the environment difficulty index, which reflects the complexity of the current training environment. The higher the value, the more complex the environment; C dx Indicates the complexity of the terrain. The more complex the terrain, the greater the difficulty of training. rd Indicates the disturbance factors in the environment, which may be external interference or the instability of the robot itself; V bhl It represents the rate of environmental change, indicating the speed at which environmental conditions change. A rapidly changing environment is more difficult to adapt to.

[0101] For example, in C2-C4, the judgment mechanism is as follows:

[0102] If the current difficulty Down It will automatically switch to the next stage;

[0103] If the current difficulty Down Then keep the current stage;

[0104] If the current difficulty Down Drops by more than 20% or action entropy remains extremely low for a long time (such as A t <0.01), it will automatically return to the previous difficulty level Reset strategy fine-tuning;

[0105] It should be noted that this mechanism forms a dynamic closed-loop "course cycle", which solves the problem of strategy collapse during training in a high-difficulty environment and enhances adaptability and robustness.

[0106] S4: Policy network optimization based on reinforcement learning: Combining the aforementioned dynamic reward function with the stage scheduling mechanism, the policy network is iteratively trained using the PPO algorithm to obtain the optimized gait strategy, including the following sub-steps D1-D5:

[0107] D1: Use the policy network initialized in step S1 as the initial PPO parameters, and combine the comprehensive reward function in step S2 with the stage adjustment mechanism in step S3;

[0108] D2: Sample the state-action-reward trajectory under the above strategy and calculate the advantage function at each step to measure the pros and cons of the current action relative to the baseline;

[0109] D3: Construct an optimization objective function based on the PPO algorithm and optimize the current policy network;

[0110] D4: Use the gradient descent algorithm to update the policy network and value network parameters while ensuring that the constraints are not violated;

[0111] D5: The iteratively updated optimized gait policy network is used in the cognitive feedback optimization process.

[0112] Specifically, the optimization objective function calculated in D1-D5 can be expressed by the following formula:

[0113]

[0114] Where r t (θ) is expressed as a probability ratio; Represents the advantage function estimate; ∈ is the cutoff range, set by the staff; clip(θ), 1-∈, 1+∈) means limiting the strategy update to a reasonable interval;

[0115] It should be noted that the optimization principle of this step is:

[0116] When r tWhen (θ) is far away from 1 (i.e., the difference between the new and old strategies is large), it is truncated by the clip function to prevent the network from "running wild";

[0117] When the advantage function When it is positive, the strategy is encouraged to increase the probability of the action; when the advantage function When it is negative, the probability of the action is suppressed to prevent oscillation in the training process;

[0118] For example, the above parameters can be specifically expressed as:

[0119]

[0120] Where, π θ (a t |s t ) represents the current policy network; Represents the old strategy, which is the strategy parameters saved before the start of the current round of training. It serves as a "benchmark reference" to prevent the new strategy from deviating too far.

[0121]

[0122] Where, Q(s t ,a t ) means in state s t Take action a t After that, the expected total reward (action-value function) that can be obtained in the future; V(s t ) means in state s t The average return expected by the current strategy under (state-value function); and If it is greater than zero, it means that the current action is better than the average strategy. Otherwise, it means that the current action is worse than the average strategy. The specific value indicates the degree of good or bad.

[0123] The value function loss and entropy regularization term are introduced simultaneously during the optimization process:

[0124]

[0125] Where, L total Represents the total loss, which includes PPO target loss, value loss and entropy regularization; The weight representing the loss of the value function; represents the weight of entropy regularization, which controls the degree of strategy exploration; L value Represents the value function loss, which represents the difference between the state value of the strategy and the actual return; Represents the entropy of the strategy, which is used to increase the exploratory nature of the strategy and prevent overfitting;

[0126] It should be noted that the total loss function L totalThis ensures that during PPO training, the strategy can not only improve the reward, but also maintain a certain degree of exploration and avoid local optimality.

[0127] S5: Introducing cognitive feedback to optimize gait strategy: In the later stages of training or deployment, cognitive feedback signals are further introduced to fine-tune the strategy, including the following sub-steps E1-E4:

[0128] E1: Uses multi-channel sensing methods such as image channel (to identify human avoidance behavior or sudden movement), audio channel (to detect negative tones or exclamations in speech), and expression channel (to detect observers' fear or discomfort expressions) to collect cognitive signals;

[0129] E2: Construct a cognitive load index based on the above cognitive signals and determine human acceptability based on cognitive thresholds;

[0130] E3: When human acceptability is poor, adjust the total reward function based on the above cognitive load indicators;

[0131] E4: Based on the adjusted total reward function, further fine-tune the current gait strategy through the PPO algorithm in step S4 until the final gait strategy is output.

[0132] Specifically, the cognitive load index constructed in E1-E4 can be reflected by the following formula:

[0133] R cog =β1f br +β2t zs +β3P fmyy ;

[0134] Where R cog represents the cognitive load reward, which evaluates whether the robot has better behavioral performance when interacting with humans; f br Indicates the frequency of human avoidance. The lower the avoidance frequency, the higher the reward. zs It indicates the time that humans look at the robot when it passes by. The longer the look time, the more unnatural the robot’s behavior. fmyy Indicates whether there is negative speech feedback. If there is negative speech, the reward will be reduced. β1, β2, and β3 represent the human avoidance frequency, gaze time, and weight coefficients of negative speech, which are set by the staff.

[0135] It should be noted that, through this cognitive feedback, the robot can adjust its gait in time according to human response, thus optimizing the comfort of human-robot interaction;

[0136] For example, the comparison between the cognitive load reward index and the cognitive threshold is expressed as:

[0137] If R cogIf it is less than the cognitive threshold, it means that human acceptability is good, and the optimized gait strategy at this time is the final gait strategy;

[0138] If R cog If it is greater than or equal to the cognitive threshold, it means that human acceptability is poor and the current optimization strategy needs to be further fine-tuned.

[0139] Specifically, when human acceptance is poor, the total reward function is adjusted as follows:

[0140] r` t =r t -λ(t)·R cog ;

[0141] Where r t is the performance reward based on the action (such as stability, energy consumption, ground contact quality, etc.); λ(t) represents the dynamic introduction coefficient of cognitive feedback, which controls the influence of cognitive factors in the early and late stages of training.

[0142] For example, the above parameters can be specifically expressed as:

[0143] λ(t)=min(λ max ,λ0+k·log(1+t));

[0144] In the formula, λ0 represents the initial weight; k represents the growth rate; t is the current training step number; λ max Indicates the maximum limit of cognitive weight to avoid suppressing the main performance indicators.

[0145] In summary, the reward function designed by the present invention adopts a multi-factor structure, covering multiple key indicators such as gait stability, energy consumption, ground contact status and expected speed deviation, and through the adaptive adjustment mechanism of reward weights, the strategy can focus on the current weak performance indicators at different training stages, thereby improving the overall gait strategy quality and practicality; the introduction of the "training stage scheduling mechanism based on environmental difficulty" automatically regresses the training difficulty when the gait strategy performance does not meet the standard or falls into the local optimum through dynamic evaluation and regulation of the training environment, ensuring the consistency and robustness of the training process, and effectively avoiding the strategy from being forgotten or lost in complex environments. Performance degradation; by introducing the PPO algorithm to train the strategy, combining probability ratio constraints with advantage function estimation, the gait strategy optimization process maintains a stable improvement in each update step, avoiding the problems of strategy jitter or training non-convergence in traditional methods. The theoretical basis is complete and the practical effect is excellent; by innovatively introducing cognitive feedback reward items, human cognitive signals such as avoidance behavior and voice response to the robot's gait are collected, and quantified into negative reward feedback for strategy correction, guiding the robot's gait to evolve in a more natural direction that is more in line with human cognitive preferences. It is suitable for application scenarios such as service robots and public space collaboration.

[0146] Example 2

[0147] Embodiment 2 is the second embodiment of the present invention. This embodiment differs from the first embodiment in that it provides a robot gait training system based on reinforcement learning, including:

[0148] The imitation learning initialization module is used to extract state-action pairs based on the input reference gait video, generate joint trajectory data using the posture estimation method, and initialize the policy network based on the behavior cloning algorithm;

[0149] A dynamic reward building module is used to construct a composite reward function that includes multiple factors such as stability, energy consumption, ground safety, and speed deviation. The weight of each factor is dynamically adjusted through a reward weight adaptation mechanism to guide the strategy to focus on the current performance bottleneck at different training stages.

[0150] An environment scheduling and performance evaluation module, which dynamically adjusts the difficulty level of the robot's training environment based on the training performance of the robot's current strategy, which is based on comprehensive strategy performance indicators including average reward, fall frequency, success rate, and action entropy, and rewinds to the previous training environment for repair training when performance degrades;

[0151] The strategy training module is used to iteratively optimize the initialized strategy network using the PPO algorithm. The optimization process includes the clipping probability ratio, advantage function estimation, value function loss term and strategy entropy regularization term.

[0152] The cognitive feedback module is used to collect cognitive signals and construct cognitive feedback rewards during the human-computer interaction phase. The cognitive signals include avoidance behavior, gaze duration, and voice feedback signals. The cognitive feedback rewards are added to the total reward in a time-gated manner to guide the strategy to evolve towards a more human-acceptable gait behavior.

[0153] This embodiment also provides a computer device, which is suitable for a robot gait training method based on reinforcement learning, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement a robot gait training method based on reinforcement learning proposed in the above embodiment.

[0154] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.

[0155] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements a robot gait training method based on reinforcement learning as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0156] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A robot gait training method based on reinforcement learning, characterized by: include: S1, collects expert action trajectories and uses behavior cloning technology to pre-train the initial policy network; S2, select multiple evaluation indicators to design a dynamic reward function to guide the initial policy network to optimize gait performance; S3, constructs an environment difficulty index and a comprehensive strategy performance index to evaluate the performance of the current strategy in a specific environment, and adaptively adjusts the increase or decrease of the environment difficulty based on the comprehensive performance index; S4, combining the dynamic reward function with the adaptive adjustment mechanism, iteratively training the policy network through the optimization objective function based on the PPO algorithm to obtain the optimized gait strategy. S5, determines the optimized gait strategy by constructing a cognitive load index and outputs the final gait strategy.

2. The robot gait training method based on reinforcement learning according to claim 1, characterized in that: The training initial strategy network includes: Collect reference videos containing natural gait, use a pose estimation algorithm to extract body key points in each image frame, and convert the key point trajectories into a state-action pair dataset; The policy network is pre-trained using the behavior cloning method so that it can predict actions consistent with the reference trajectory based on the state.

3. The robot gait training method based on reinforcement learning according to claim 2, characterized in that: The construction of the comprehensive reward function includes: Select multiple evaluation indicators covering stability, grounding status, energy efficiency, and speed deviation; A comprehensive reward function is obtained by integrating multiple evaluation indicators such as stability, grounding status, energy efficiency, speed deviation, etc. in a weighted manner.

4. The robot gait training method based on reinforcement learning according to claim 3, characterized in that: The comprehensive performance indicators of the construction include: Select multiple performance evaluation indicators including average round reward, fall frequency, task completion success rate, and action entropy; Multiple performance evaluation indicators such as round average reward, fall frequency, task completion success rate, action entropy, etc. are weighted and integrated according to the set weight coefficients to obtain a comprehensive performance indicator.

5. The robot gait training method based on reinforcement learning according to claim 4, characterized in that: The construction of the optimization objective function includes: Construct a cutting function to constrain the degree of policy deviation at each step during policy update; Based on the probability ratio of the current policy to the old policy, an optimization objective function is constructed for the reinforcement learning training process.

6. The robot gait training method based on reinforcement learning according to claim 5, characterized in that: The cognitive load indicators include: Collect and analyze cognitive signals from external observers’ images, voice, and facial expressions during gait training; Extract multiple characteristic variables reflecting human cognitive responses, including avoidance behavior frequency, average gaze duration, negative speech, etc. The avoidance behavior frequency, average gaze duration, and negative speech feature variables were integrated in a weighted manner to obtain the cognitive load index.

7. The robot gait training method based on reinforcement learning according to claim 6, characterized in that: Determining the optimized gait strategy includes: comparing the cognitive load index with a cognitive threshold; According to the combination of the comparison results, the optimized gait strategy is judged as having good human acceptability or poor human acceptability.

8. A robot gait training system based on reinforcement learning, based on the robot gait training method based on reinforcement learning according to any one of claims 1 to 7, characterized in that: include, Imitation learning initialization module for extracting state-action pairs based on the input reference gait video; Dynamic reward building module, used to construct a compound reward function that includes multiple factors such as stability, energy consumption, ground safety, and speed deviation; The environment scheduling and performance evaluation module is used to dynamically adjust the difficulty level of the training environment based on the training performance of the robot's current strategy; Strategy training module, used to iteratively optimize the current strategy network using the PPO algorithm; The cognitive feedback module is used to collect cognitive signals and construct cognitive feedback reward items during the human-computer interaction stage.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the robot gait training method based on reinforcement learning are implemented in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the robot gait training method based on reinforcement learning according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Self-adaptive strategy optimization method and system for robot with body based on interactive feedback

    CN121223793A

  • Quadruped robot robust adaptive multi-skill learning method based on key frame guidance

    CN121523059A

  • A robust self-adaptive multi-skill learning method for quadruped robots based on key frame guidance

    CN121523059B