Autonomous gait control method and system for a multi-legged robot
By combining low-dimensional latent variable encoding and hysteresis weights, the problem of unstable gait switching in multi-legged robots driven by energy consumption is solved, achieving stable and adaptive gait control and reducing ineffective switching and fluctuations in joint control quantities.
Patent Information
- Application Number
- CN202610781582.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-25
AI Technical Summary
In the process of energy-driven autonomous gait control, existing multi-legged robots suffer from unstable gait switching boundaries, resulting in fluctuations in joint control quantities and a large number of invalid switching.
A method combining low-dimensional latent variable encoding and policy network is adopted. By acquiring a fixed-length historical sequence of the proprioceptive state of a multi-legged robot, low-dimensional latent variables are generated. Then, a latent variable hysteresis reward term is constructed using hysteresis weights and energy consumption references to suppress ineffective transitions of low-dimensional latent variables under similar energy consumption states and stabilize gait switching.
This study achieved stability and adaptability in gait switching during autonomous gait control of a multi-legged robot driven by energy consumption, reduced ineffective switching and fluctuations in joint control quantities, and improved the robot's motion stability under different speeds and terrains.
Smart Images

Figure CN122632604A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of legged robot motion control technology, and in particular to an autonomous gait control method and system for multi-legged robots. Background Technology
[0002] In existing multi-legged robot motion control methods, some schemes use proprioceptive historical sequences and speed commands to train the robot's control strategy, while others guide the robot to form different gait states at different speeds through energy-intensive rewards, thereby reducing reliance on manually preset gait states. CN119065240A discloses a fault-tolerant control method for a single motor of a quadruped robot based on reinforcement learning. Specifically, it involves a reinforcement learning control method for quadruped robots that combines teacher and student networks. The student network receives proprioceptive historical sequences and desired speed commands and outputs joint position control instructions. Related research on Adaptive Energy Regularization discloses a quadruped robot that autonomously forms and switches between different gait states through an energy center reward strategy.
[0003] However, in energy-driven autonomous gait control, the internal state of the strategy may still undergo repeated changes under similar energy consumption conditions, leading to instability in gait switching boundaries and fluctuations in joint control quantities. Therefore, it is necessary to provide an autonomous gait control method for multi-legged robots that can improve the stability of autonomous gait switching. Summary of the Invention
[0004] The purpose of this invention is to provide an autonomous gait control method and system for multi-legged robots, addressing the shortcomings of existing technologies. The technical problem it aims to solve is: how to suppress ineffective transitions of low-dimensional latent variables under similar energy consumption states during energy-intensive autonomous gait control of multi-legged robots, while retaining the autonomous gait switching capability under improved energy consumption states. Specifically, this invention adopts the following solution: An autonomous gait control method for a multi-legged robot includes the following steps: Historical encoding steps: Obtain a fixed-length historical sequence of the multi-legged robot's body perception state and the target motion speed command; input the fixed-length historical sequence of the body perception state into the latent variable encoder to obtain low-dimensional latent variables; Strategy control steps: Input the low-dimensional latent variables and the target motion speed command into the strategy network to obtain the joint target control quantity used to control the joint motion of the multi-legged robot; Energy consumption comparison steps: Calculate the current energy consumption index based on the joint torque, joint speed and body movement speed of the multi-legged robot at the current sampling moment; determine the energy consumption reference amount based on the historical energy consumption data corresponding to the target movement speed command; and calculate the energy consumption difference between the current energy consumption index and the energy consumption reference amount. Hysteresis constraint steps: Generate hysteresis weights based on the energy consumption difference, calculate the change in latent variables between the low-dimensional latent variables at the current sampling time and the low-dimensional latent variables at the previous sampling time, and construct a latent variable hysteresis reward term based on the hysteresis weights and the change in latent variables, wherein the hysteresis weights increase when the current energy consumption index is not lower than the energy consumption reference value, and decrease when the current energy consumption index is lower than the energy consumption reference value; Training update steps: Based on the total reward, which includes command tracking reward, energy consumption penalty, and latent variable hysteresis reward, perform reinforcement learning training on the latent variable encoder and the policy network; Deployment control steps: After training, the latent variable encoder generates the low-dimensional latent variable based on the fixed-length historical sequence of the actual collected body perception state, and the policy network generates the joint target control quantity based on the low-dimensional latent variable and the actual input target motion speed command to control the movement of the multi-legged robot.
[0005] Specifically, the core of the above scheme is not simply to reduce robot motion energy consumption through energy rewards, nor is it simply to smooth control actions. Instead, it sets energy consumption improvement criteria at the low-dimensional latent variable level upon which the policy network depends. Low-dimensional latent variables can be understood as internal motion state representations extracted by the latent variable encoder from ontological perception states such as joint position, joint velocity, body posture, and historical actions. The policy network generates target joint control quantities based on these internal motion state representations and the target motion speed command. During training, the current energy consumption index reflects the motion cost at the current sampling moment, while the energy consumption reference value reflects the energy consumption level already achieved or recently performed well under the speed conditions corresponding to the target motion speed command. The energy consumption difference between the two is used to determine whether the current latent variable change has energy consumption improvement significance. When the current energy consumption index is not lower than the energy consumption reference value, it indicates that significant changes in low-dimensional latent variables have not led to energy consumption improvement. In this case, increasing the hysteresis weight can suppress repeated jumps in low-dimensional latent variables under similar energy consumption conditions. When the current energy consumption index is lower than the energy consumption reference value, it indicates that changes in low-dimensional latent variables may correspond to a lower energy consumption movement mode. In this case, decreasing the hysteresis weight can allow changes in low-dimensional latent variables and promote the formation of a better gait. In this way, the robot can autonomously form and stably switch gaits without explicit gait commands, while reducing invalid jitter at gait switching boundaries.
[0006] Furthermore, the current energy consumption index is a transportation cost index, which is determined based on the sum of the positive mechanical power of each joint of the multi-legged robot at the current sampling moment, the robot's mass, gravitational acceleration, and body movement speed. Specifically, the transportation cost index is used to link the mechanical power consumed by the robot at the current sampling moment with the robot's own weight and movement speed, avoiding interference from factors such as robot mass and speed when evaluating energy consumption solely based on joint torque or mechanical power. The sum of the positive mechanical power of each joint can be obtained by multiplying the joint torque and joint speed of each joint, taking the positive power component, and then summing them. Using positive mechanical power helps to reflect the main part of the active energy output by the motor, reducing the interference of negative power feedback, braking, or numerical oscillations on training rewards. The multi-legged robot's mass and gravitational acceleration are used to form weight normalization, and the body movement speed is used to form speed normalization, thereby making the energy consumption comparison under different speed conditions more stable. This setting allows the current energy consumption index to truly reflect the energy required per unit weight and unit distance of movement, facilitating subsequent comparison with energy consumption reference values.
[0007] Furthermore, in the energy consumption comparison step, the target motion speed command is mapped to a speed range, historical energy consumption data corresponding to the speed range is obtained, and the energy consumption reference value is determined based on the quantile estimate of the historical energy consumption data. Specifically, different target motion speeds require different suitable motion rhythms and gait patterns for the robot; the reasonable energy consumption levels corresponding to low-speed walking, medium-speed trotting, and high-speed running are not the same. If a uniform fixed energy consumption threshold is used, it is easy to lead to an overly wide threshold in low-speed states and an overly strict threshold in high-speed states, thereby affecting the accuracy of latent variable hysteresis constraints. By mapping the target motion speed command to the corresponding speed range, historical energy consumption data can be maintained separately for different speed ranges, so that the energy consumption reference value matches the speed conditions. The quantile estimate can prevent a single abnormally low-energy consumption sample from becoming an overly strong constraint, and can also prevent the average value from being pulled up by high-energy-consumption failure samples. Through this processing, the hysteresis weight generation process can make judgments based on energy consumption performance under the same or similar speed conditions, thereby more accurately distinguishing meaningful latent variable changes from invalid latent variable jumps.
[0008] Furthermore, the energy consumption reference value is updated using an exponential sliding update method. This method determines the energy consumption reference value at the current update time based on the energy consumption reference value at the previous update time and the quantile estimate at the current update time. Specifically, the energy consumption reference value is not a fixed threshold that remains unchanged after being set once, but a dynamic reference value that is gradually updated as training progresses or running data is collected. The exponential sliding update method can preserve the continuity of historical energy consumption levels while incorporating new energy consumption performance at the current stage, ensuring that the energy consumption reference value is neither overly affected by a single abnormal data point nor remains at the poor energy consumption level of the early training stage for an extended period. In practice, the energy consumption reference value at the previous update time can be multiplied by a smoothing coefficient, and the quantile estimate at the current update time can be multiplied by the remaining weights and then summed to obtain the energy consumption reference value at the current update time. This allows the hysteresis criterion to gradually become stricter as the policy capability improves, prompting the robot to continue searching for lower energy consumption and more stable gait patterns in subsequent training.
[0009] Further, in the hysteresis constraint step, the energy consumption difference is input into a monotonic smooth mapping function to obtain the hysteresis weight, and the hysteresis weight is multiplied by the norm of the latent variable change and then inverted to obtain the latent variable hysteresis reward term. Specifically, the monotonic smooth mapping function is used to convert the energy consumption difference into a continuously changing hysteresis weight, avoiding sudden jumps in the hysteresis weight when the energy consumption difference approaches a critical position. The monotonic relationship ensures that the more unfavorable the current energy consumption index is relative to the energy consumption reference value, the larger the hysteresis weight; the more favorable the current energy consumption index is relative to the energy consumption reference value, the smaller the hysteresis weight. The norm of the latent variable change is used to measure the magnitude of change of the low-dimensional latent variable at the current sampling time relative to the low-dimensional latent variable at the previous sampling time. The larger the norm, the more obvious the change in the internal state of the policy. Multiplying the hysteresis weight by the norm of the latent variable change and inverting the sign before adding it to the total reward allows reinforcement learning training to penalize large changes in latent variables when energy consumption does not improve, and relax the restrictions on changes in latent variables when energy consumption improves. This technique differs from ordinary fixed smoothing penalties, which do not distinguish whether changes in latent variables lead to energy consumption improvement. This approach, however, adaptively determines the constraint strength based on the energy consumption improvement status.
[0010] Furthermore, the change in the latent variable is the first-order difference between the low-dimensional latent variable at the current sampling time and the low-dimensional latent variable at the previous sampling time, and the latent variable hysteresis reward term is determined based on the L2 norm of the first-order difference. Specifically, the first-order difference can directly reflect the direction and magnitude of change of the low-dimensional latent variable at adjacent sampling times, and is suitable for judging whether a rapid jump occurs in the motion state representation within the strategy. The L2 norm can synthesize the changes in each dimension of the low-dimensional latent variable into a single magnitude indicator, which is convenient to form a calculable reward term together with the hysteresis weight. Using the first-order difference between adjacent sampling times, instead of using the cumulative difference over a long window, allows the training process to promptly detect instantaneous jumps in the latent variable, thereby suppressing invalid changes at gait switching boundaries. This processing method is simple to implement and can be directly calculated in the reinforcement learning training loop based on the current low-dimensional latent variable and the previous low-dimensional latent variable, without the need for additional sensors or external annotations.
[0011] Furthermore, a representation constraint step is included before the training update step. This step constructs a privileged encoder and a latent variable encoder. The privileged encoder receives privileged environmental information from the training environment and outputs privileged latent variables. The latent variable encoder receives a fixed-length history sequence of the ontology-aware state and outputs a low-dimensional latent variable of the same dimension as the privileged latent variable. A representation constraint loss is constructed based on the synchronously acquired privileged latent variables and the low-dimensional latent variables. The training update step updates the latent variable encoder based on the total reward and the representation constraint loss. Specifically, privileged environmental information is information easily obtained in the training environment but inconvenient to obtain directly during the deployment phase. This information may include terrain height, contact state, friction conditions, external disturbances, or simulation environment parameters. The privileged encoder generates privileged latent variables based on the privileged environmental information, while the latent variable encoder generates low-dimensional latent variables only based on the fixed-length history sequence of the ontology-aware state. By constructing a representation constraint loss using synchronously acquired privileged latent variables and low-dimensional latent variables, the low-dimensional latent variables can be guided to learn implicit representations related to terrain, contact, and motion states without directly receiving privileged environmental information. The training update step simultaneously utilizes the total reward and representation constraint loss to update the latent variable encoder, enabling the low-dimensional latent variables to serve both policy control and energy hysteresis constraints, while also carrying environmental state-related information. Thus, even without terrain sensors or privileged environmental information during the deployment phase, the latent variable encoder can infer state changes related to gait selection from the robot's own motion history.
[0012] Furthermore, the representation constraint loss is a contrastive learning loss, which uses the privileged latent variables and the low-dimensional latent variables acquired synchronously as positive sample pairs, and the privileged latent variables acquired asynchronously within a batch as negative samples. Specifically, the contrastive learning loss is used to narrow the representation distance between the synchronously acquired privileged latent variables and the low-dimensional latent variables, and to widen the representation distance between different samples. Synchronous acquisition means that the privileged latent variables and the low-dimensional latent variables come from the same training environment, the same time step, or the same motion state; therefore, they should express corresponding terrain and motion state information. The privileged latent variables acquired asynchronously within a batch come from other environments or other time steps and can be used as negative samples to enable the low-dimensional latent variables to distinguish between different terrain states and different motion states. Compared to directly using mean squared error to force the low-dimensional latent variables to completely fit the privileged latent variables, the contrastive learning loss emphasizes the relative representation relationship, which helps to avoid the low-dimensional latent variables being over-constrained and maintains the space for subsequent adjustment through energy consumption hysteresis to form autonomous gait changes.
[0013] Furthermore, the policy network does not use the privileged latent variables as action output conditions during training, and it does not receive the environmental privileged information during deployment. Specifically, the privileged latent variables are only used to represent and constrain low-dimensional latent variables during the training phase, and are not used as direct conditions for the policy network to output actions. This avoids the policy network from over-relying on environmental privileged information during the training phase, which could lead to performance degradation when environmental privileged information is lacking during the deployment phase. During training, the action output conditions remain the low-dimensional latent variables and the target motion velocity command, and the same input structure is used during deployment. This design keeps the policy input consistent between the training and deployment ends, which helps reduce the input distribution deviation between simulation training and actual deployment. At the same time, since environmental privileged information does not enter the deployment control link, the actual robot does not need to be equipped with depth cameras, LiDAR, or additional sensors for directly measuring terrain parameters, thus reducing system implementation costs and deployment complexity.
[0014] Furthermore, the training update steps include a warm-up training phase and an energy-dominated training phase. The weights of the energy consumption penalty term and the latent variable hysteresis reward term are lower in the warm-up training phase than in the energy-dominated training phase. Specifically, the warm-up training phase is mainly used to enable the multi-legged robot to learn basic standing, velocity tracking, and stable movement capabilities. If strong energy consumption penalties and strong latent variable hysteresis constraints are imposed in the early stages of training, the latent variable encoder may be prematurely restricted before it has formed effective motion representations, leading to insufficient policy exploration or difficulty in convergence. The energy-dominated training phase increases the influence of the energy consumption penalty term and the latent variable hysteresis reward term after the robot has acquired basic motion capabilities, allowing the policy to be further optimized towards low energy consumption and stable gait switching. Through phased settings, early exploration and later constraints can be balanced, avoiding locking low-dimensional latent variables at the beginning of training and also avoiding frequent jumps in low-dimensional latent variables under similar energy consumption states in the later stages of training.
[0015] Furthermore, in the deployment control step, the external motion command received by the policy network is the target motion speed command. The policy network does not receive gait category commands, phase offset commands, duty cycle commands, or gait frequency commands. Specifically, during the deployment phase, the operator or upper-level controller only needs to input the target motion speed command; the policy network does not need to additionally receive explicit gait parameters such as gait category, phase offset, duty cycle, or gait frequency. Gait formation and gait switching are jointly determined by low-dimensional latent variables, the target motion speed command, and the energy consumption hysteresis mechanism formed during training. This reduces the operational complexity caused by manually setting gait parameters and avoids the problem of insufficient adaptability of fixed gait parameters under complex terrain or speed variation conditions. This setting is particularly suitable for scenarios where robots need to move continuously between different speeds and terrains, enabling the robot to autonomously adjust its internal motion state based on its proprioceptive history and energy consumption improvement state.
[0016] An autonomous gait control system for a multi-legged robot, comprising: The body perception history acquisition module is used to acquire a fixed-length history sequence of the body perception state of the multi-legged robot and the target motion speed command. The latent variable encoding module is connected to the ontology perception history acquisition module. The latent variable encoding module is used to encode the fixed-length history sequence of the ontology perception state into low-dimensional latent variables. A strategy control module is connected to the latent variable encoding module. The strategy control module is used to receive the low-dimensional latent variable and the target motion speed command, and output the joint target control quantity for controlling the joint motion of the multi-legged robot. The energy consumption comparison module is used to calculate the current energy consumption index based on the joint torque, joint speed and body movement speed of the multi-legged robot at the current sampling time, determine the energy consumption reference amount based on the historical energy consumption data corresponding to the target movement speed command, and calculate the energy consumption difference between the current energy consumption index and the energy consumption reference amount. The hysteresis constraint module is connected to the latent variable encoding module and the energy consumption comparison module. The hysteresis constraint module is used to generate hysteresis weights based on the energy consumption difference, calculate the change in latent variables between the low-dimensional latent variables at the current sampling time and the low-dimensional latent variables at the previous sampling time, and construct a latent variable hysteresis reward term based on the hysteresis weights and the change in latent variables. The hysteresis weights increase when the current energy consumption index is not lower than the energy consumption reference value, and decrease when the current energy consumption index is lower than the energy consumption reference value. The training update module is connected to the latent variable encoding module, the policy control module, and the hysteresis constraint module. The training update module is used to perform reinforcement learning training on the latent variable encoding module and the policy control module based on the total reward, which includes command tracking reward, energy consumption penalty, and latent variable hysteresis reward. A deployment control module is connected to the ontology perception history acquisition module, the latent variable encoding module, and the strategy control module. After training is completed, the deployment control module enables the latent variable encoding module to generate the low-dimensional latent variables based on the fixed-length history sequence of the actual acquired ontology perception states, and enables the strategy control module to generate joint target control quantities based on the low-dimensional latent variables and the actual input target motion speed command, so as to control the movement of the multi-legged robot.
[0017] Furthermore, the energy consumption comparison module includes a transportation cost calculation unit, which determines the transportation cost index as the current energy consumption index based on the sum of the positive mechanical power of each joint of the multi-legged robot at the current sampling time, the mass of the multi-legged robot, the gravitational acceleration, and the body movement speed. The energy consumption comparison module also includes a speed interval mapping unit and an energy consumption reference quantity determination unit. The speed interval mapping unit maps the target movement speed command to a speed interval, and the energy consumption reference quantity determination unit acquires historical energy consumption data corresponding to the speed interval and determines the energy consumption reference quantity based on the quantile estimates of the historical energy consumption data. The energy consumption comparison module further includes an energy consumption reference quantity update unit, which determines the energy consumption reference quantity at the current update time using an exponential sliding update method based on the energy consumption reference quantity at the previous update time and the quantile estimate at the current update time. The hysteresis constraint module includes a hysteresis weight generation unit and a latent variable hysteresis reward term generation unit. The hysteresis weight generation unit is used to input the energy consumption difference into a monotonically smooth mapping function to obtain the hysteresis weight. The latent variable hysteresis reward term generation unit is used to multiply the hysteresis weight by the norm of the latent variable change and then invert the sign to obtain the latent variable hysteresis reward term. It also includes a representation constraint module, which includes a privileged encoder and a representation constraint loss construction unit. The privileged encoder is used to receive privileged environmental information in the training environment and output privileged latent variables. The latent variable encoding module is used to receive a fixed-length historical sequence of the ontology-aware state and output a low-dimensional latent variable of the same dimension as the privileged latent variable. The representation constraint loss construction unit is used to construct a representation constraint loss based on the synchronously acquired privileged latent variable and the low-dimensional latent variable. The training update module is also used to update the latent variable encoding module together with the total reward and the representation constraint loss. The representation constraint loss construction unit is used to construct a contrastive learning loss by using the privileged latent variables and the low-dimensional latent variables collected synchronously as positive sample pairs and the privileged latent variables collected asynchronously within a batch as negative samples.
[0018] The beneficial effects of this invention are: By employing the aforementioned autonomous gait control method for multi-legged robots, the fixed-length historical sequence of the body's perceived state is encoded as a low-dimensional latent variable. The policy network then generates joint target control variables based on these low-dimensional latent variables and the target motion speed command. This eliminates the gait control of the multi-legged robot from relying on explicit gait commands such as gait type, phase offset, duty cycle, or gait frequency. Furthermore, a hysteresis weight is generated based on the energy consumption difference between the current energy consumption index and the energy consumption reference value corresponding to the target motion speed command. This hysteresis weight is then used to constrain the changes in low-dimensional latent variables at adjacent sampling times. This ensures that the changes in low-dimensional latent variables are no longer solely determined by the instantaneous output of the policy network but are regulated by the energy consumption improvement state. When the current energy consumption index is not lower than the energy consumption reference value, the hysteresis weight increases, creating stronger constraints on the transitions of the low-dimensional latent variables, thereby reducing ineffective gait switching under similar energy consumption states. When the current energy consumption index is lower than the energy consumption reference value, the hysteresis weight decreases, allowing the low-dimensional latent variables to change with energy consumption improvement, thus preserving the ability to autonomously form lower-energy-consumption gaits. Therefore, in the process of energy-driven autonomous gait control, multi-legged robots can simultaneously take into account gait switching stability and gait adaptation capability, reduce joint control fluctuations caused by repeated changes in the internal state of the strategy, and solve the problems of unstable gait switching boundaries, motion jitter and many invalid switching in existing energy-driven gait control. Attached Figure Description
[0019] Figure 1 This is a schematic flowchart of the method steps of the present invention.
[0020] Figure 2 This is a schematic block diagram of the overall system structure of the present invention. Detailed Implementation
[0021] The technical solutions in the embodiments of the present invention will now be clearly and completely described in conjunction with the accompanying drawings.
[0022] This embodiment provides an autonomous gait control method and system for multi-legged robots, applicable to multi-legged robots with a main body and multiple legs, including quadruped robots, hexapod robots, and other robots with multiple leg actuators. The autonomous gait control method uses reinforcement learning during the training phase to acquire a motion strategy capable of traversing speeds and terrains. During the deployment phase, it relies solely on a fixed-length historical sequence of the body's perceived state and target motion speed commands to generate joint target control quantities, thereby controlling the multi-legged robot to complete autonomous gait formation and switching. The body's perceived state refers to motion state information that the multi-legged robot can collect or calculate itself. This state includes joint positions, joint velocities, body posture observations, and joint target control quantities from the previous sampling time. The body's perceived state does not include terrain height maps, external depth maps, or lidar point clouds directly collected by external terrain sensors. The target motion speed command represents the forward motion speed that the multi-legged robot expects to achieve. In one embodiment, the target motion speed command is a one-dimensional speed command.
[0023] Combination Figure 1 As shown, the autonomous gait control method for multi-legged robots includes a history encoding step, a strategy control step, an energy consumption comparison step, a hysteresis constraint step, a training and update step, and a deployment control step. Figure 1 Used to indicate the signal flow direction between the above steps Figure 1 The diagram, from top to bottom, shows the fixed-length historical sequence input latent variable encoder for ontological perception, the low-dimensional latent variable input policy network, the energy consumption difference formed by the current energy consumption index and the energy consumption reference, the adjustment of the low-dimensional latent variable change by the energy consumption difference, the addition of the latent variable hysteresis reward term to the total reward, and the execution chain for deployment control after training. Figure 1 The execution order shown can unify the reward construction in the training phase, the hysteresis constraints of latent variables, and the control output in the deployment phase into the same computational chain.
[0024] In the historical encoding step, a fixed-length historical sequence of the multi-legged robot's body perception state and the target motion velocity command are first acquired. The fixed-length historical sequence of the body perception state can be represented as a time series composed of body perception states from multiple consecutive sampling moments. The body perception state at each sampling moment includes the joint positions of each joint of the multi-legged robot, the joint velocities of each joint, body posture observations, and the joint target control variables from the previous sampling moment. The body posture observations can be represented by the gravity direction projection in the body coordinate system or other posture-related observations. The joint target control variables from the previous sampling moment are used to reflect the action state output by the policy network at the previous sampling moment, enabling the latent variable encoder to infer the robot's current motion rhythm or contact state based on historical actions and the current motion state.
[0025] In a preferred embodiment, the fixed-length history sequence of the proprioceptive state has a history length of 50 frames, with each frame having an input dimension of 48 dimensions. In this case, the fixed-length history sequence of the proprioceptive state can be represented as time-series data with a size of 50*48. The history length is not limited to 50 frames; it can be adjusted between 20 and 100 frames. A shorter history length allows the latent variable encoder to respond more quickly to changes in the current state; a longer history length allows the latent variable encoder to obtain more comprehensive motion cycle information and contact change information. In practical implementation, the history length can be selected based on the multi-legged robot's control frequency, leg motion cycle, and the computing power of the computing platform.
[0026] After obtaining a fixed-length historical sequence of the proprioceptive state, the sequence is input into a latent variable encoder to obtain low-dimensional latent variables. The latent variable encoder compresses high-dimensional temporal states into low-dimensional latent variables that can characterize the current motion state, terrain adaptation state, and gait switching state. The dimension of the low-dimensional latent variables can be less than or equal to 16, or in alternative embodiments, it can be relaxed to no more than 32. The purpose of setting low-dimensional latent variables is to force the latent variable encoder to compress redundant information unrelated to gait selection and terrain adaptation, allowing the low-dimensional latent variables to more centrally express the motion state information affecting gait formation and gait switching.
[0027] In a preferred embodiment, the latent variable encoder employs a cascaded structure of causal one-dimensional convolution, recurrent units, and fully connected mapping layers. The causal one-dimensional convolution is used to extract local temporal features from a fixed-length historical sequence of ontology-aware states. The recurrent units are used to accumulate motion state information across time, and the fully connected mapping layers are used to output low-dimensional latent variables. The causal one-dimensional convolution only accesses historical data up to the current sampling time, excluding data from future sampling times, thus ensuring real-time operation during deployment. The recurrent units can be gated recurrent units, or alternatively, long short-term memory networks, simple recurrent units, state-space models, or other temporal models with causal temporal modeling capabilities. The causal one-dimensional convolution can also be replaced by temporal convolutional networks, causal attention structures, or dilated convolutional structures.
[0028] In a specific architecture, the latent variable encoder takes a fixed-length historical sequence of ontological perception states as input, with an input size of 50*48. The first layer is a causal one-dimensional convolutional layer with a kernel size of 5, 32 channels, a stride of 1, and an output size of 50*32. The second layer is also a causal one-dimensional convolutional layer with a kernel size of 5, 64 channels, a stride of 2, and an output size of 25*64. The third layer is another causal one-dimensional convolutional layer with a kernel size of 3, 64 channels, a stride of 2, and an output size of 13*64. The fourth layer is a gated recurrent unit with 128 hidden units, taking the hidden state of the last time step as the temporal aggregation result. The fifth layer is a multilayer perceptron, mapping from 128 to 32 and then to the low-dimensional latent variable dimension, outputting the low-dimensional latent variable. The causal one-dimensional convolutional layers achieve causality through left-side zero-padding, the gated recurrent unit preserves the motion phase, contact state, and action change trends across time, and the fully connected mapping layer forms a low-dimensional bottleneck representation.
[0029] In the strategy control step, low-dimensional latent variables and target motion velocity commands are input into the strategy network to obtain joint target control quantities for controlling the joint motion of the multi-legged robot. The strategy network can adopt a multilayer perceptron structure. In one embodiment, the strategy network includes an input layer, three hidden layers, and an output layer. The number of neurons in the hidden layers is 512, 256, and 128, respectively. The inputs to the strategy network include low-dimensional latent variables, target motion velocity commands, and statistics of the previous target motion velocity command or the previous velocity command. The output of the strategy network is the joint target position of each joint. The joint target position is input to a position-type proportional-differential controller (PDC), which generates motor control quantities to drive the multi-legged robot joint motion. In a preferred embodiment, the proportional gain is 30, and the differential gain is 0.65. The above proportional gain and differential gain can be adjusted according to the characteristics of the robot joint motors.
[0030] The joint target control quantities output by the policy network are not preset gait parameters, nor are they manually specified phase offsets, duty cycles, gait frequencies, or gait categories. The policy network directly generates the joint target control quantities based on low-dimensional latent variables and target motion velocity commands. The low-dimensional latent variables are generated by the latent variable encoder based on a fixed-length historical sequence of the robot's perceived state. Therefore, the policy network can implicitly adjust the joint target control quantities according to the robot's own motion state and target motion velocity commands, enabling the robot to form different motion patterns under different speeds and terrain conditions.
[0031] In the energy consumption comparison step, the current energy consumption index is calculated based on the joint torque, joint velocity, and body movement speed of the multi-legged robot at the current sampling time. The preferred current energy consumption index is the transportation cost index. To reduce the disturbance of negative work to the training gradient, only the positive mechanical power of each joint can be counted. The sum of the positive mechanical power of each joint at the current sampling time can be determined as follows.
[0032] Satisfying formula (1): ; Instantaneous mechanical power This represents the sum of the positive mechanical power of each joint at the current sampling time, and the number of joints. This represents the total number of joints involved in the calculation of a multi-legged robot, and the joint torque. Indicates the first The output torque and joint velocity of each joint at the current sampling moment. Indicates the first The joint angular velocity of each joint at the current sampling moment.
[0033] Obtaining instantaneous mechanical power Then, the transportation cost index can be determined based on the multi-legged robot's mass, gravitational acceleration, and body movement speed. The transportation cost index can be determined as follows.
[0034] Satisfying formula (2): ; Among them, transportation cost indicators This represents the normalized energy consumption index at the current sampling moment, and the instantaneous mechanical power. The mass of the multi-legged robot represents the sum of the positive mechanical power of each joint at the current sampling moment. This represents the robot's mass and gravitational acceleration. Represents the gravitational constant and the velocity of the body. This represents the forward velocity of the organism at the current sampling moment, with a lower bound. A numerical stability constant used to prevent division by zero or excessively large values when the body's velocity approaches zero. Used to improve computational stability.
[0035] In the energy consumption comparison step, an energy consumption reference value is determined based on historical energy consumption data corresponding to the target motion speed command, and the energy consumption difference between the current energy consumption index and the energy consumption reference value is calculated. Since reasonable energy consumption levels differ for different target motion speeds, the energy consumption reference value does not use a uniform fixed threshold but corresponds to the target motion speed command. Specifically, the target motion speed command is mapped to a speed range, and energy consumption data from the most recent sampling moments or the most recent training rounds are extracted from the historical energy consumption data corresponding to that speed range. The energy consumption reference value is then determined based on quantile estimates. The quantile estimates can be the lower quantile, median, or other statistics that represent the optimal energy consumption performance of the current speed range. Using quantile estimates avoids the excessive influence of a single abnormally low-energy-consumption sample on the energy consumption reference value, and also avoids excessive increases in the energy consumption reference value due to high-energy-consumption failure samples.
[0036] In one implementation, the energy consumption reference value can be updated using an exponential sliding update method. The exponential sliding update method determines the energy consumption reference value at the current update time based on the energy consumption reference value at the previous update time and the quantile estimate at the current update time. The update of the energy consumption reference value can be represented as follows.
[0037] Satisfying formula (3): ; Among them, the energy consumption reference quantity corresponding to the speed command Indicates the target movement speed command Historical energy consumption reference value for the speed range, smoothing coefficient The quantile function represents the proportion of the energy consumption reference value from the previous update time that is retained in the current update. This indicates taking the quantiles of the most recent energy consumption data set. This represents the set of transportation cost indicators from the most recent sampling times or the most recent training rounds within the corresponding speed range, with quantile parameters. This represents the quantile used to determine the reference energy consumption level.
[0038] In the hysteresis constraint step, hysteresis weights are generated based on the energy consumption difference. The energy consumption difference represents the degree of deviation of the current energy consumption index from the energy consumption reference value. When the current energy consumption index is not lower than the energy consumption reference value, it indicates that changes in low-dimensional latent variables under the current motion state have not led to energy consumption improvement, and the hysteresis weight increases. When the current energy consumption index is lower than the energy consumption reference value, it indicates that the current motion state may correspond to a lower energy consumption motion mode, and the hysteresis weight decreases. In this way, low-dimensional latent variables are not fixedly forced to smooth, but are subject to stronger constraints when energy consumption does not improve, and are allowed to change when energy consumption improves.
[0039] Specifically, we first calculate the change in latent variables between the low-dimensional latent variables at the current sampling time and the low-dimensional latent variables at the previous sampling time. The change in latent variables can be represented using first-order differences.
[0040] Satisfying formula (4): ; Among them, the change in latent variables This represents the change in low-dimensional latent variables between the current sampling time and the previous sampling time. The current low-dimensional latent variable... This represents the low-dimensional latent variable output by the latent variable encoder at the current sampling time, and the previous low-dimensional latent variable. This represents the low-dimensional latent variable output by the latent variable encoder at the previous sampling time.
[0041] Subsequently, the energy consumption difference is input into a monotonically smooth mapping function to obtain hysteresis weights, and a latent variable hysteresis reward term is constructed based on the hysteresis weights and the change in latent variables. In one embodiment, the latent variable hysteresis reward term can be determined as follows.
[0042] Satisfying formula (5): ; Among them, the latent variable is the delayed reward item. This represents the time-varying constraint term used to incorporate the latent variable into the total reward, and the hysteresis adjustment coefficient. This represents the overall strength of the latent variable, the delayed reward term, and the change in the latent variable. The L2 norm represents the change in low-dimensional latent variables between the current sampling time and the previous sampling time. The monotonic smooth mapping function represents the magnitude of change of low-dimensional latent variables. This represents a function that maps energy consumption differences to hysteresis weights, with the mapping steepness parameter... This indicates the strength of the impact of energy consumption difference on changes in hysteresis weight, a transportation cost indicator. This indicates the current energy consumption index and the energy consumption reference value corresponding to the speed command. This indicates the historical energy consumption reference value corresponding to the target movement speed command.
[0043] Current transportation cost indicators Significantly lower than the energy consumption reference value corresponding to the speed command At this time, the monotonic smooth mapping function outputs a smaller hysteresis weight, and the latent variable hysteresis reward term has a weaker penalty for changes in low-dimensional latent variables, thus allowing changes in low-dimensional latent variables and giving the policy network the opportunity to form a lower-energy-consumption movement pattern. When the current transportation cost index... Energy consumption reference value that is close to or higher than the speed command At this time, the monotonic smooth mapping function outputs a large hysteresis weight, and the latent variable hysteresis reward term strongly penalizes changes in low-dimensional latent variables, thus suppressing meaningless jumps in low-dimensional latent variables. Therefore, changes in low-dimensional latent variables have an energy consumption improvement threshold, and gait switching no longer depends entirely on the instantaneous policy output, but is jointly regulated by the energy consumption reference value and the energy consumption difference.
[0044] The monotonically smooth mapping function can be the Sigmoid function, or it can be the hyperbolic tangent function, piecewise linear function, or other monotonically smooth functions. The latent variable change can be represented by first-order differencing, higher-order differencing, exponential moving differencing, or distance relative to a reference low-dimensional latent variable. The energy consumption reference can be binned according to the target motion speed command, or by combined speed and attitude, terrain category, or time period. The energy consumption reference can be the historical best value, historical median, historical low quantile, or a distribution reference value obtained based on online estimation. All of these alternative methods serve the same purpose: adjusting the constraint strength of low-dimensional latent variable changes according to the energy consumption improvement status.
[0045] In the training update step, reinforcement learning is performed on the latent variable encoder and policy network based on the total reward, which includes a command tracking reward, an energy consumption penalty, and a latent variable hysteresis reward. The total reward can be expressed in the following form.
[0046] Satisfying formula (6): ; Among them, the total reward This represents the reward used for reinforcement learning updates at the current sampling time, tracking the weights. This indicates the weight of the command tracking reward item. This indicates the reward items corresponding to target motion speed command tracking, heading maintenance, or torso height maintenance, with energy consumption weights. The weight of the energy consumption penalty term, representing the change in the metric as the training progresses, and the transportation cost metric. This represents the current energy consumption index, with the latent variable being the delayed reward item. This represents the hysteresis constraint term determined based on the energy consumption difference and the change in low-dimensional latent variables, with auxiliary weights. Indicates the weight of auxiliary reward items, auxiliary reward items This indicates auxiliary constraints such as posture stability, joint limit, torque smoothing, and joint velocity smoothing.
[0047] Command tracking reward items can be constructed using the error between the target motion speed command and the machine's motion speed. For example, the command tracking reward item can include a speed tracking error item, which can be based on the machine's motion speed. Target movement speed command The squared difference is penalized. Auxiliary reward items can include posture maintenance, trunk height maintenance, joint limitation, joint velocity smoothing, joint torque smoothing, and movement variation smoothing. Auxiliary reward items are used to ensure movement safety and stability during training.
[0048] To enable the multi-legged robot to first learn basic movement capabilities and then gradually optimize towards low-energy gait and stable gait switching, energy consumption weights are used. The training progress metrics can change monotonically without decreasing. These metrics can include training steps, cumulative epochs, current tracking error quantile, and policy entropy reduction percentage. Energy consumption weights can be set as follows.
[0049] Satisfying formula (7): ; Among them, energy consumption weight Indicators of training progress The corresponding energy consumption penalty weight, minimum energy consumption weight This represents the lower bound of the energy consumption penalty term weights in the early stages of training, and the maximum energy consumption weights. This represents the upper limit of the energy consumption penalty term weight in the later stages of training, and is a training progress indicator. Indicates training steps or other training progress; warm-up threshold. This indicates the training progress position where the weight of the energy consumption penalty term begins to increase significantly, representing the transition steepness. The monotonic smooth mapping function represents the smoothness of the change in energy consumption weights from low to high. This represents the Sigmoid function or other monotonic smooth functions.
[0050] During the warm-up training phase, energy consumption weighting Approaching minimum energy consumption weight The policy network primarily learns standing, velocity tracking, and basic stable motion capabilities. Warm-up threshold. It can be preferably set to approximately The number of training steps. During the energy-dominated training phase, energy consumption weights... Gradually approaching the maximum energy consumption weight The policy network further reduces transportation cost metrics while maintaining speed tracking capabilities. The mapping function for energy consumption weights is not limited to the Sigmoid function; it can also employ piecewise linear functions, hyperbolic tangent functions, cosine annealing inverse functions, or event-triggered functions based on tracking error thresholds. Energy consumption is not limited to transportation cost metrics; it can also use instantaneous mechanical energy time integrals, electrical energy, motor copper loss terms, the first or second norm of mechanical power, or selective summation for joint groups such as the hip and knee joints. Positive mechanical power can also be extended to a weighted sum of positive work and small-weighted negative work.
[0051] The training and update steps can employ a policy gradient algorithm. Preferably, a proximal policy optimization algorithm is used to update the policy network and latent variable encoder. Training can be performed in a parallel physical simulation environment, with the number of parallel environments preferably between 2048 and 4096. The physical simulation environment can employ procedural mixed terrain, including flat ground, rough terrain, slopes, random steps, and soft ground. The round length is preferably 20 seconds, and at the start of each round, target motion velocity commands and terrain parameters are randomly sampled. Target motion velocity commands can be uniformly sampled from a preset speed range. The total number of training samples is preferably... to The magnitude of the randomization can be adjusted during training. Domain randomization can include randomization of fuselage mass, friction coefficient, attitude observation noise, joint position noise, and joint velocity noise. In a preferred embodiment, the randomized fuselage mass ranges from ±2.0 kg, and the randomized friction coefficient ranges from 0.2 to 2.0 times the default friction coefficient.
[0052] An asymmetric evaluation structure can be used during the training phase. The policy network, acting as the action output network, receives low-dimensional latent variables and target motion velocity commands as input, and outputs joint target control variables. During training, the value network can additionally receive privileged latent variables or other statistics to obtain more accurate value estimates. The value network can employ a multilayer perceptron structure, with the number of hidden layer neurons potentially reaching 512, 256, or 128. After training, the value network does not participate in deployment control.
[0053] Prior to the training and update step, a representation constraint step can be performed. This step constructs a privileged encoder and a latent variable encoder. The privileged encoder receives privileged environmental information from the training environment and outputs privileged latent variables. The latent variable encoder receives a fixed-length history sequence of the ontology-aware state and outputs low-dimensional latent variables of the same dimension as the privileged latent variables. The privileged environmental information is information available in the training environment but not used during deployment. This information may include terrain height grids, contact forces, friction coefficients, external force disturbances, subsystem fault flags, wind speed, or other simulation environment descriptions. The privileged latent variables and low-dimensional latent variables have the same dimension, preferably no greater than 16.
[0054] In one implementation, the privileged encoder employs a multilayer perceptron structure. The inputs include a terrain height grid, contact force, friction coefficient, and perturbation information. The number of neurons in the hidden layer can be 256 or 128, and the output is a privileged latent variable. The privileged latent variable output by the privileged encoder is not used as the action output condition of the policy network, but rather to construct a representation constraint loss. During the training phase, the policy network only receives low-dimensional latent variables and target motion velocity commands, and during the deployment phase, it also only receives low-dimensional latent variables and target motion velocity commands. In this way, the policy inputs at the training and deployment ends remain consistent, avoiding performance degradation during deployment caused by the policy network's reliance on environmental privileged information during the training phase.
[0055] The constraint representation loss can employ contrastive learning loss. Contrastive learning loss uses synchronously acquired privileged latent variables and low-dimensional latent variables as positive sample pairs, and intra-batch asynchronously acquired privileged latent variables as negative samples. Synchronous acquisition means that the privileged latent variables and low-dimensional latent variables come from the same parallel simulation environment, the same sampling time, or the same motion state. Intra-batch asynchronous acquisition means that negative samples come from other parallel simulation environments or other sampling times. Contrastive learning loss can take the following form.
[0056] Satisfying formula (8): ; Among them, contrastive learning loss The loss function used to characterize the constraints, and the batch size. Indicates the number of samples in the same batch, sample index. Indicates the current sample number, negative sample index. Indicates the privileged latent variable number within the batch, a low-dimensional latent variable. Indicates the first In a parallel simulation environment, the latent variables output by the latent variable encoder at the current sampling time are privileged latent variables. Indicates the first The latent variables and similarity functions output by the privileged encoder at the current sampling time in a parallel simulation environment. Represents cosine similarity, dot product similarity, Euclidean negative distance, or bilinear similarity, and temperature coefficient. This represents the temperature parameter in the contrastive learning loss.
[0057] Contrastive learning loss can be replaced by NT-Xent loss, BarlowTwins loss, VICReg loss, SimSiam-style noncontrastive loss, or a mean squared error auxiliary term with a smaller weight can be superimposed on the contrastive learning loss. The definition of positive samples is not limited to the same environment and time period; it can also be extended to samples of the same terrain category or within the same time window. The privileged encoder can also be replaced by the hidden layer representations of the value network. All of these alternatives are used to enable low-dimensional latent variables to carry environmental and motion state information even when derived solely from ontological perception.
[0058] During the training phase, a cost quantile adaptive terrain curriculum can also be employed. Terrain parameters can be represented as roughness, slope, friction coefficient, step height, and ground compliance. Each terrain parameter can be divided into different intervals according to terrain difficulty level. Each parallel simulation environment maintains its own terrain difficulty level. After each training round, the relative quantile ranking of the energy consumption index of that parallel simulation environment in this round is calculated within the set of energy consumption indices of all parallel simulation environments in the most recent rounds, and the terrain difficulty level is adjusted based on the relative quantile ranking. The terrain difficulty level update can be represented as follows.
[0059] Satisfying formula (9): ; Among them, the difficulty level of the parallel simulation environment Indicates the first The terrain difficulty level corresponding to each parallel simulation environment, with the maximum difficulty level. This indicates the maximum level that the terrain course can reach, and the round transport cost index. Indicates the first Energy consumption metrics of each parallel simulation environment in the current training round, recent quantile ranking. This represents the relative quantile of the current round's energy consumption metric within the set of recent energy consumption metrics across all parallel simulation environments, and is the difficulty increase threshold. This represents the quantile threshold required to increase the terrain difficulty level, and the threshold for decreasing difficulty. This represents the quantile threshold required to reduce the terrain difficulty level.
[0060] In a preferred embodiment, the difficulty increase threshold The threshold for difficulty is 0.3. The relative quantile ranking is 0.7. A lower recent quantile ranking indicates better energy efficiency. When a parallel simulation environment performs better in terms of energy efficiency than the better performers in the group under the current terrain, the terrain difficulty level of that parallel simulation environment is increased; when a parallel simulation environment performs poorly in terms of energy efficiency, the terrain difficulty level is decreased; otherwise, the terrain difficulty level remains unchanged. To avoid degradation by sacrificing speed for lower energy efficiency, speed tracking error or the energy consumption decreasing trend within rounds can be used as secondary criteria. The relative quantile ranking can also be replaced by an absolute energy consumption threshold, an individual's historical ranking, or a standard score relative to the group average. Energy consumption distribution estimation for multiple parallel simulation environments can use reservoir sampling, sliding window averaging, or asynchronous aggregation. The step size for terrain difficulty level changes can be a fixed step size or adaptively changed based on the quantile ranking. A minimum dwell time constraint can also be added to avoid frequent increases and decreases in terrain difficulty level.
[0061] The training process can be divided into a warm-up training phase and an energy-dominated training phase. In the warm-up training phase, the weights of the energy penalty term and the latent variable lag reward term are relatively low, while the weight of the representation constraint loss can be relatively high to accelerate the alignment of low-dimensional latent variables with privileged latent variables. Adaptive terrain courses can be enabled, but the difficulty increase conditions can be set more strictly to allow the policy network to first achieve stable tracking capabilities. In the energy-dominated training phase, the weights of the energy penalty term enter a higher range, the weights of the latent variable lag reward term increase, and the weights of the representation constraint loss can be maintained or gradually annealed. The adaptive terrain course progresses according to the normal threshold. This two-phase training avoids over-constraining of low-dimensional latent variables in the early stages of training and allows for focused optimization of low-energy gait and gait switching stability in the later stages of training.
[0062] In the deployment control step, after training, only the latent variable encoder, policy network, and downstream joint controller are retained; the privileged encoder, value network, terrain privileged information input, and environmental privileged information input do not participate in deployment control. During deployment, the latent variable encoder parameters and policy network parameters are loaded first, and the fixed-length history buffer of the body-aware state is initialized. If the history buffer is insufficient to the preset length upon power-up, it can be padded with zero padding. Then, the following processes are executed cyclically at a 50 Hz control frequency: reading joint position, joint velocity, and body attitude observations; reading the joint target control quantity from the previous sampling moment; and concatenating them to form a fixed-length history sequence of the body-aware state; inputting the fixed-length history sequence of the body-aware state into the latent variable encoder to obtain low-dimensional latent variables; reading the target motion velocity command; inputting the low-dimensional latent variables and the target motion velocity command into the policy network to obtain the joint target control quantity at the current sampling moment; distributing the joint target control quantity to the position-type proportional-derivative controller; and updating the fixed-length history buffer of the body-aware state. During the deployment phase, gait category instructions, phase offset instructions, duty cycle instructions, and gait frequency instructions are not received. Gait formation and gait switching are jointly determined by low-dimensional latent variables, target motion speed instructions, and the trained policy network.
[0063] Combination Figure 2As shown, the multi-legged robot autonomous gait control system includes a proprioception history acquisition module, a latent variable encoding module, a strategy control module, an energy consumption comparison module, a hysteresis constraint module, a training update module, a deployment control module, and a representation constraint module. The proprioception history acquisition module acquires a fixed-length historical sequence of the multi-legged robot's proprioception state and a target motion velocity command. This module outputs the fixed-length historical sequence of the proprioception state to the latent variable encoding module and provides the target motion velocity command to the strategy control module. The latent variable encoding module generates low-dimensional latent variables based on the fixed-length historical sequence of the proprioception state and outputs these low-dimensional latent variables to the strategy control module, the hysteresis constraint module, and the representation constraint module. The strategy control module generates joint target control quantities based on the low-dimensional latent variables and the target motion velocity command.
[0064] The energy consumption comparison module receives historical energy consumption data corresponding to joint torque, joint speed, body movement speed, and target movement speed commands, and generates current energy consumption indicators, energy consumption reference values, and energy consumption differences. The energy consumption comparison module may include a transportation cost calculation unit, a speed interval mapping unit, an energy consumption reference value determination unit, and an energy consumption reference value update unit. The transportation cost calculation unit determines the transportation cost indicator based on the sum of the positive mechanical power of each joint, the mass of the multi-legged robot, gravitational acceleration, and body movement speed. The speed interval mapping unit maps the target movement speed command to the corresponding speed interval. The energy consumption reference value determination unit obtains historical energy consumption data from the corresponding speed interval and determines the energy consumption reference value based on quantile estimates. The energy consumption reference value update unit updates the energy consumption reference value using an exponential sliding update method.
[0065] The hysteresis constraint module is connected to the latent variable encoding module and the energy consumption comparison module. The hysteresis constraint module generates hysteresis weights based on the energy consumption difference, calculates the change in latent variables based on the low-dimensional latent variables at the current sampling time and the low-dimensional latent variables at the previous sampling time, and constructs a latent variable hysteresis reward term based on the hysteresis weights and the change in latent variables. The hysteresis constraint module may include a hysteresis weight generation unit and a latent variable hysteresis reward term generation unit. The hysteresis weight generation unit is used to input the energy consumption difference into a monotonically smooth mapping function to obtain the hysteresis weights. The latent variable hysteresis reward term generation unit is used to multiply the hysteresis weights by the norm of the latent variable change and then invert the sign to obtain the latent variable hysteresis reward term.
[0066] The training and update module is connected to the latent variable encoding module, policy control module, hysteresis constraint module, and representation constraint module. The training and update module performs reinforcement learning training on the latent variable encoding module and policy control module based on the total reward, which includes command tracking reward, energy consumption penalty, and latent variable hysteresis reward. The training and update module can also jointly update the latent variable encoding module based on the representation constraint loss output by the representation constraint module. The training and update module can run in a parallel physics simulation environment, using proximal policy optimization algorithms or other policy gradient algorithms to update network parameters.
[0067] The representation constraint module includes a privileged encoder and a representation constraint loss construction unit. The privileged encoder receives privileged environmental information from the training environment and outputs privileged latent variables. The representation constraint loss construction unit constructs the representation constraint loss based on synchronously acquired privileged latent variables and low-dimensional latent variables. This unit can construct a contrastive learning loss using synchronously acquired privileged latent variables and low-dimensional latent variables as positive sample pairs and intra-batch asynchronously acquired privileged latent variables as negative samples. The representation constraint loss is output to the training update module for jointly updating the latent variable encoding module.
[0068] The deployment control module is used to enable the ontology perception history acquisition module, latent variable encoding module, and policy control module to participate in deployment control after training is completed. The deployment control module does not directly receive isolated low-dimensional latent variables as input; instead, it invokes the ontology perception history acquisition module, latent variable encoding module, and policy control module retained during the deployment phase to form a complete deployment chain. The signal flow during the deployment phase is as follows: the ontology perception history acquisition module acquires a fixed-length historical sequence of the actually acquired ontology perception states; the latent variable encoding module generates low-dimensional latent variables based on the fixed-length historical sequence of the actually acquired ontology perception states; and the policy control module generates joint target control quantities based on the low-dimensional latent variables and the actual input target motion velocity command. Therefore, the deployment control module represents the deployment execution control relationship after training is completed, rather than representing the direct input of low-dimensional latent variables into the deployment control module.
[0069] In a preferred hardware platform, the multi-legged robot is a 12-DOF quadruped robot, with each leg including a hip joint, thigh joint, and knee joint. The robot body can be a quadruped robot of the Unitree Go2 type. The robot's joint layers employ position-based proportional-derivative (PDD) control. The host computer can be a computing device equipped with a graphics processing unit (GPU), such as one equipped with an NVIDIA RTX 4090 GPU. The host computer can communicate with the robot body via a wired connection, using a lightweight communication and serialization protocol. The operator can input target motion speed commands via a joystick; no gait selection button is provided.
[0070] The simulation and training platform can utilize a parallel physics simulation environment similar to Isaac Lab. The terrain generator can generate procedurally mixed terrain, including flat land, rough terrain, slopes, random steps, and soft ground. The policy optimization algorithm preferably employs a proximal policy optimization algorithm. The policy network, latent variable encoder, privileged encoder, and value network can be updated within the same training loop. The contrastive learning loss and policy gradient loss can be backpropagated simultaneously to the latent variable encoder parameters. The privileged encoder can be updated along with the value network gradients or updated independently with a smaller learning rate.
[0071] To verify the effectiveness of the above scheme, multiple baseline schemes can be set for comparison. The first baseline scheme sets the energy consumption weight to a fixed value and removes the energy consumption-dominant curriculum. The second baseline scheme replaces the contrastive learning loss with mean squared error distillation loss. The third baseline scheme removes the latent variable hysteresis adjustment mechanism. The fourth baseline scheme replaces the cost quantile adaptive terrain curriculum with a fixed curriculum that progresses linearly by the number of training steps. The fifth baseline scheme uses phase bias commands and symmetric rewards as the gait control baseline. Evaluation metrics can include the average cross-terrain transportation cost index, the 95th quantile of the cross-terrain transportation cost index, the normalized root mean square error of velocity tracking, terrain passability, and gait switching frequency. The gait switching frequency can be statistically analyzed by counting the number of times the first difference of the low-dimensional latent variable exceeds a threshold per unit time. The complete scheme, while maintaining similar velocity tracking accuracy, can reduce the average cross-terrain transportation cost index, reduce ineffective jumps of low-dimensional latent variables at velocity or terrain boundaries, and reduce the peak variance of joint moments.
[0072] In terms of specific operational effects, through energy-dominated curriculum rewards, the policy network initially acquires basic velocity tracking capabilities in the early stages of training, gradually becoming dominated by energy-consumption penalties in the later stages, thus tending to form low-energy gait. Through representational constraint loss between the privileged encoder and the latent variable encoder, low-dimensional latent variables can carry terrain and contact state information even when derived solely from fixed-length historical sequences of ontology-perceived states. Through latent variable hysteresis rewards, low-dimensional latent variables are less prone to significant jumps when energy consumption is not improved, but can change when energy consumption improves. Through cost quantile adaptive terrain curriculum, the terrain difficulty of the parallel simulation environment can automatically adjust as the policy capability improves, concentrating training data in terrain difficulty regions slightly higher than the current policy capability. These mechanisms work together to enable the multi-legged robot to achieve autonomous gait control across flat ground, rough terrain, slopes, steps, and soft ground. During deployment, motion control is completed solely through target motion velocity commands, without requiring explicit gait selection commands or privileged information from the deployment environment.
[0073] In alternative implementations, the multi-legged robot can be a quadruped, a hexapod, or other multi-legged robot with a main body. The target motion velocity command can be a forward velocity command, or it can be extended to include velocity commands containing direction or other motion commands, depending on the control task. The latent variable encoder can employ a combination of causal one-dimensional convolution and recurrent units, or it can employ temporal convolutional networks, Transformer causal attention, dilated convolution, long short-term memory networks, gated recurrent units, state-space models, or other causal temporal coding models. Energy consumption indicators can include transportation cost indicators, instantaneous mechanical energy time integrals, electrical energy, motor copper loss terms, mechanical power L1 norm, mechanical power L2 norm, or selective joint power. Latent variable hysteresis constraints can be applied to low-dimensional latent variables, or, when the policy network has an implicit periodic structure, to internal state variables equivalent to gait phase. The terrain course can use a complete set of terrain parameters, or only roughness and slope dimensions, or it can be extended to a set of terrain parameters including elastic collision bodies, fluid resistance, or localized debris. The above alternatives do not change the basic technical idea of this scheme, which improves the constraint strength of low-dimensional latent variable changes by improving state regulation through energy consumption, and achieves autonomous gait control by relying only on the fixed-length historical sequence of the body's perceived state and the target's motion speed command during the deployment phase.
[0074] By employing the autonomous gait control method for multi-legged robots in this case, the fixed-length historical sequence of the body's perceived state is encoded as low-dimensional latent variables. The policy network then generates joint target control quantities based on these low-dimensional latent variables and the target motion speed command. This eliminates the gait control of the multi-legged robot from relying on explicit gait commands such as gait type, phase offset, duty cycle, or gait frequency. Furthermore, a hysteresis weight is generated based on the energy consumption difference between the current energy consumption index and the energy consumption reference value corresponding to the target motion speed command. This hysteresis weight is then used to constrain the changes in low-dimensional latent variables at adjacent sampling times. This ensures that the changes in low-dimensional latent variables are no longer solely determined by the instantaneous output of the policy network but are regulated by the energy consumption improvement state. When the current energy consumption index is not lower than the energy consumption reference value, the hysteresis weight increases, creating stronger constraints on the jumps in low-dimensional latent variables, thereby reducing ineffective gait switching under similar energy consumption states. When the current energy consumption index is lower than the energy consumption reference value, the hysteresis weight decreases, allowing the low-dimensional latent variables to change with energy consumption improvement, thus retaining the ability to autonomously form lower-energy-consumption gaits. Therefore, in the process of energy-driven autonomous gait control, multi-legged robots can simultaneously take into account gait switching stability and gait adaptation capability, reduce joint control fluctuations caused by repeated changes in the internal state of the strategy, and solve the problems of unstable gait switching boundaries, motion jitter and many invalid switching in existing energy-driven gait control.
[0075] Furthermore, based on the low-dimensional latent variable hysteresis adjustment mechanism, the entire technical solution introduces historical energy consumption data and quantile estimates corresponding to the target motion speed command to determine the energy consumption reference value. This allows energy consumption comparison to no longer rely on a fixed threshold but adapt to historical motion performance in different speed ranges. The energy consumption reference value is continuously updated through an exponential sliding update method, enabling the hysteresis criterion to adjust according to changes in energy consumption performance during training or operation, avoiding the insufficient adaptability of fixed thresholds under different speeds, terrains, or training stages. Simultaneously, a privileged encoder receives privileged environmental information and outputs privileged latent variables during the training phase. The latent variable encoder is then updated using the representational constraint loss between the privileged latent variables and the low-dimensional latent variables. This allows the low-dimensional latent variables to carry implicit information related to the environmental and motion states, derived solely from the ontological perception state. After training, the policy network can be deployed without receiving privileged environmental information, reducing reliance on external terrain sensors or additional environmental observation channels. By setting phased training stages—warm-up and energy-driven—the energy consumption penalty and latent variable hysteresis reward terms maintain low influence in the early stages. Energy consumption and hysteresis constraints are then strengthened after the multi-legged robot possesses basic tracking capabilities, preventing latent variables from being locked prematurely in the early training phase. Therefore, the overall technical solution not only improves the stability of energy-driven autonomous gait switching but also addresses the problems in existing technologies, such as the difficulty of stable deployment of body perception under terrain conditions, insufficient adaptability of fixed energy consumption thresholds, and the difficulty of balancing energy consumption optimization and gait switching stability. Furthermore, it simplifies deployment commands, reduces dependence on external environmental sensing, minimizes joint control fluctuations, and improves the stability of autonomous gait control across speed states.
[0076] Finally, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for autonomous gait control of a multi-legged robot, characterized in that, The steps include the following: Historical encoding steps: Obtain a fixed-length historical sequence of the multi-legged robot's body perception state and the target motion speed command; input the fixed-length historical sequence of the body perception state into the latent variable encoder to obtain low-dimensional latent variables; Strategy control steps: Input the low-dimensional latent variables and the target motion speed command into the strategy network to obtain the joint target control quantity used to control the joint motion of the multi-legged robot; Energy consumption comparison steps: Calculate the current energy consumption index based on the joint torque, joint speed and body movement speed of the multi-legged robot at the current sampling moment; determine the energy consumption reference amount based on the historical energy consumption data corresponding to the target movement speed command; and calculate the energy consumption difference between the current energy consumption index and the energy consumption reference amount. Hysteresis constraint steps: Generate hysteresis weights based on the energy consumption difference, calculate the change in latent variables between the low-dimensional latent variables at the current sampling time and the low-dimensional latent variables at the previous sampling time, and construct a latent variable hysteresis reward term based on the hysteresis weights and the change in latent variables. The hysteresis weights increase when the current energy consumption index is not lower than the energy consumption reference value, and decrease when the current energy consumption index is lower than the energy consumption reference value.
2. The autonomous gait control method for a multi-legged robot according to claim 1, characterized in that, The hysteresis constraint step also includes: Training update steps: Based on the total reward, which includes command tracking reward, energy consumption penalty, and latent variable hysteresis reward, perform reinforcement learning training on the latent variable encoder and the policy network; Deployment control steps: After training, the latent variable encoder generates the low-dimensional latent variable based on the fixed-length historical sequence of the actual collected body perception state, and the policy network generates the joint target control quantity based on the low-dimensional latent variable and the actual input target motion speed command to control the movement of the multi-legged robot. The current energy consumption index is a transportation cost index, which is determined based on the sum of the positive mechanical power of each joint of the multi-legged robot at the current sampling time, the mass of the multi-legged robot, the gravitational acceleration, and the body movement speed.
3. The autonomous gait control method for a multi-legged robot according to claim 2, characterized in that, In the energy consumption comparison step, the target motion speed command is mapped to a speed range, historical energy consumption data corresponding to the speed range is obtained, and the energy consumption reference amount is determined based on the quantile estimate of the historical energy consumption data. The energy consumption reference value is updated through an exponential sliding update method, which determines the energy consumption reference value at the current update time based on the energy consumption reference value at the previous update time and the quantile estimate at the current update time.
4. The autonomous gait control method for a multi-legged robot according to claim 3, characterized in that, In the hysteresis constraint step, the energy consumption difference is input into a monotonically smooth mapping function to obtain the hysteresis weight, and the hysteresis weight is multiplied by the norm of the latent variable change and then the sign is reversed to obtain the latent variable hysteresis reward term. The change in the latent variable is the first-order difference between the low-dimensional latent variable at the current sampling time and the low-dimensional latent variable at the previous sampling time, and the latent variable lag reward term is determined based on the L2 norm of the first-order difference.
5. The autonomous gait control method for a multi-legged robot according to claim 4, characterized in that, The training update step is preceded by: The representation constraint step involves constructing a privileged encoder and a latent variable encoder. The privileged encoder receives privileged environmental information from the training environment and outputs privileged latent variables. The latent variable encoder receives a fixed-length historical sequence of the ontology-aware state and outputs a low-dimensional latent variable of the same dimension as the privileged latent variable. A representation constraint loss is constructed based on the synchronously acquired privileged latent variable and the low-dimensional latent variable. The training update step updates the latent variable encoder based on the total reward and the representation constraint loss.
6. The autonomous gait control method for a multi-legged robot according to claim 5, characterized in that, The representation constraint loss is a contrastive learning loss, which uses the privileged latent variables and the low-dimensional latent variables collected synchronously as positive sample pairs, and the privileged latent variables collected asynchronously within the batch as negative samples. The policy network does not use the privileged latent variables as action output conditions during training, and the policy network does not receive the environmental privilege information during deployment.
7. The autonomous gait control method for a multi-legged robot according to claim 6, characterized in that, The training update steps include a warm-up training phase and an energy consumption-dominated training phase. The weights of the energy consumption penalty term and the latent variable delayed reward term are lower in the warm-up training phase than in the energy consumption-dominated training phase. In the deployment control step, the external motion command received by the policy network is the target motion speed command, and the policy network does not receive gait category command, phase offset command, duty cycle command, and gait frequency command.
8. An autonomous gait control system for a multi-legged robot, characterized in that, include: The body perception history acquisition module is used to acquire a fixed-length history sequence of the body perception state of the multi-legged robot and the target motion speed command. The latent variable encoding module is connected to the ontology perception history acquisition module. The latent variable encoding module is used to encode the fixed-length history sequence of the ontology perception state into low-dimensional latent variables. A strategy control module is connected to the latent variable encoding module. The strategy control module is used to receive the low-dimensional latent variable and the target motion speed command, and output the joint target control quantity for controlling the joint motion of the multi-legged robot. The energy consumption comparison module is used to calculate the current energy consumption index based on the joint torque, joint speed and body movement speed of the multi-legged robot at the current sampling time, determine the energy consumption reference amount based on the historical energy consumption data corresponding to the target movement speed command, and calculate the energy consumption difference between the current energy consumption index and the energy consumption reference amount. The hysteresis constraint module is connected to the latent variable encoding module and the energy consumption comparison module. The hysteresis constraint module is used to generate hysteresis weights based on the energy consumption difference, calculate the change in the latent variable between the low-dimensional latent variable at the current sampling time and the low-dimensional latent variable at the previous sampling time, and construct a latent variable hysteresis reward term based on the hysteresis weights and the change in the latent variable. The hysteresis weights increase when the current energy consumption index is not lower than the energy consumption reference value, and decrease when the current energy consumption index is lower than the energy consumption reference value.
9. The multi-legged robot autonomous gait control system according to claim 8, characterized in that, Also includes: The training update module is connected to the latent variable encoding module, the policy control module, and the hysteresis constraint module. The training update module is used to perform reinforcement learning training on the latent variable encoding module and the policy control module based on the total reward, which includes command tracking reward, energy consumption penalty, and latent variable hysteresis reward. A deployment control module is connected to the ontology perception history acquisition module, the latent variable encoding module, and the strategy control module. After training is completed, the deployment control module enables the latent variable encoding module to generate the low-dimensional latent variables based on the fixed-length history sequence of the actual acquired ontology perception states, and enables the strategy control module to generate joint target control quantities based on the low-dimensional latent variables and the actual input target motion speed command, so as to control the movement of the multi-legged robot.
Citation Information
Patent Citations
Quadruped robot single motor fault tolerance control method based on reinforcement learning
CN119065240A