Robot walking control method and device, electronic equipment and storage medium

By introducing a straight-knee walking mode and a cushioning mechanism to optimize the gait control of the humanoid robot, the problems of excessive knee joint torque and high energy consumption were solved, achieving more natural and stable walking control.

CN122261150APending Publication Date: 2026-06-23INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF AUTOMATION CHINESE ACAD OF SCI
Filing Date
2026-04-07
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

In existing technologies, humanoid robots exhibit unnatural walking postures, excessive knee joint torque, high energy consumption, and are prone to motor overheating, resulting in severe hardware mechanical wear.

Method used

By acquiring the robot's proprioception data, the reference body height and target foot pitch angle are determined, a straight-knee walking mode is introduced, and gait control is implemented through a gait control model. The gait generation is optimized by combining a buffer mechanism and a symmetry consistency reward.

Benefits of technology

It reduces knee joint torque and energy consumption, avoids motor overheating, slows down hardware mechanical wear, and achieves a more natural and stable walking gait.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122261150A_ABST
    Figure CN122261150A_ABST
Patent Text Reader

Abstract

The application provides a robot walking control method and device, electronic equipment and storage medium, and belongs to the technical field of robot motion control. The method comprises the following steps: acquiring robot body perception data and relative displacement of a foot relative to a pelvis; determining a reference body height based on a fully extended leg height; the maximum value of the reference body height is less than the fully extended leg height; determining a target foot pitch angle according to the relative displacement; inputting the body perception data, the reference body height and the target foot pitch angle into a gait control model to obtain a control instruction; and controlling the robot to walk according to the control instruction. The application sets the reference body height slightly lower than the fully extended leg height, guides the robot to adopt a walking mode closer to straight knees to effectively shorten the force arm of the trunk weight to the knee joint axis center, and introduces a soft buffer mechanism through the linear mapping relationship between the target foot pitch angle and the relative displacement, so that the knee joint torque and energy consumption can be reduced, the motor overheating can be avoided, and the hardware mechanical loss can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot motion control technology, and in particular to a robot walking control method, device, electronic device, and storage medium. Background Technology

[0002] Gait generation and control for humanoid robots is a crucial aspect of robot motion control. With the advancement of robotics technology, the demands for natural movement, energy efficiency, and reliability in robots are constantly increasing.

[0003] Currently, humanoid robot gait planning often adopts a knee-bending walking strategy, in which the robot's thighs and lower legs remain bent during the support phase.

[0004] However, this bent-knee walking strategy is visually significantly different from human walking posture, and the bent knee joint results in a longer lever arm from the torso weight to the knee joint axis, requiring greater knee joint torque to perform the same task. This not only increases energy consumption but also leads to motor overheating, performance degradation, and accelerated hardware mechanical wear. Summary of the Invention

[0005] This invention provides a walking control method, device, electronic device, and storage medium for robots, which solves the problems of unnatural walking posture of humanoid robots in the prior art, as well as excessive knee joint torque, high energy consumption, easy overheating of motors, performance degradation, and accelerated mechanical wear of hardware.

[0006] This invention provides a method for controlling the walking of a robot, comprising the following steps: Acquire the robot's body perception data, as well as the relative displacement of the robot's feet relative to the pelvis; A reference body height is determined based on the robot's fully extended leg height; wherein the maximum value of the reference body height is less than the fully extended leg height. The target foot pitch angle is determined based on the relative displacement, and the target foot pitch angle has a linear mapping relationship with the relative displacement; The body perception data, the reference body height, and the target foot pitch angle are input into the gait control model to obtain the control commands output by the gait control model; the robot is then controlled to walk according to the control commands.

[0007] According to a walking control method for a robot provided by the present invention, determining a reference body height based on the robot's fully extended leg height includes: The maximum value of the reference body height is determined based on the difference between the fully extended leg height and the target height; Based on the maximum value of the reference body height, the range of values ​​for the reference body height is determined; The reference body height is determined based on the range of values.

[0008] According to a walking control method for a robot provided by the present invention, the step of determining the target foot pitch angle based on the relative displacement includes: Obtain the buffer ratio coefficient; The target foot pitch angle is obtained based on the relative displacement and the buffer ratio coefficient. Wherein, if the robot’s foot is located in front of the pelvis, the relative displacement is greater than zero, and the pitch angle of the target foot is greater than zero; If the robot's foot is located behind the pelvis, the relative displacement is less than zero, and the pitch angle of the target foot is less than zero.

[0009] According to the present invention, a walking control method for a robot is provided, wherein the gait control model is trained based on the following steps: Obtain the robot's raw observation data, input the raw observation data into the initial gait control model, and obtain the raw motion commands output by the initial gait control model; The original observation data is subjected to an observation mirror transformation to obtain mirrored observation data; The mirror observation data is input into the initial gait control model to obtain the intermediate motion commands output by the initial gait control model; Perform an action mirror transformation on the intermediate action command to obtain a mirror action command; Calculate the difference between the original action instruction and the mirror action instruction, and determine the symmetry consistency reward based on the difference; Based on the symmetry consistency reward, the network parameters of the initial gait control model are updated until the preset training termination condition is met, thus obtaining the gait control model.

[0010] According to a robot walking control method provided by the present invention, the step of performing a motion mirror transformation on the intermediate motion command to obtain a mirror motion command includes: Obtain the target position of the robot's left leg joint, the target position of the right leg joint, and the lateral component of the pelvic linear velocity from the intermediate motion command; The target positions of the left and right leg joints are interchanged, and the signs of the non-pitch joint target positions are inverted to obtain mirrored joint target positions. Invert the sign of the horizontal component to obtain the mirrored horizontal component; The mirror motion command is generated based on the target position of the mirror joint and the mirror lateral component.

[0011] According to a robot walking control method provided by the present invention, after determining the symmetry consistency reward based on the difference, the method further includes: Obtain the robot's actual body height and actual foot position; Calculate the height error between the actual body height and the reference body height; Calculate the pose error between the actual foot pose and the target foot pitch angle; The trajectory tracking reward is determined based on the height error and the pose error.

[0012] According to the present invention, a walking control method for a robot, wherein updating the network parameters of the initial gait control model based on the symmetry consistency reward includes: Obtain both task objective rewards and system constraint rewards; The target reward is obtained by weighted summing of the task objective reward, the system constraint reward, the trajectory tracking reward, and the symmetry consistency reward. The network parameters of the initial gait control model are updated based on the target reward.

[0013] The present invention also provides a walking control device for a robot, comprising the following modules: The data acquisition module is used to acquire the robot's body perception data and the relative displacement of the robot's feet relative to the pelvis. A height determination module is used to determine a reference body height based on the robot's fully extended leg height; wherein the maximum value of the reference body height is less than the fully extended leg height; The pitch angle determination module is used to determine the target foot pitch angle based on the relative displacement, wherein the target foot pitch angle and the relative displacement have a linear mapping relationship. The motion control module is used to input the body perception data, the reference body height, and the target foot pitch angle into the gait control model, obtain the control commands output by the gait control model, and control the robot to walk according to the control commands.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the walking control method of the robot as described above.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the walking control method of the robot as described above.

[0016] The robot walking control method, device, electronic device, and storage medium provided by this invention guide the robot to adopt a walking mode closer to straight knees by setting a reference body height slightly lower than the height of the fully extended legs. This effectively shortens the lever arm from the torso weight to the knee joint axis. At the same time, a compliant buffering mechanism is introduced through the linear mapping relationship between the target foot pitch angle and relative displacement, thereby reducing knee joint torque and energy consumption, avoiding motor overheating, and mitigating hardware mechanical wear. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the robot walking control method provided by the present invention.

[0019] Figure 2 This is a schematic diagram of the training process of the gait control model provided by the present invention.

[0020] Figure 3 This is a flowchart illustrating the process of determining the mirror action instruction provided by the present invention.

[0021] Figure 4 This is a schematic diagram of the process for determining the reward for trajectory tracking provided by the present invention.

[0022] Figure 5 This is a schematic diagram of the humanoid robot provided by the present invention.

[0023] Figure 6 This is a schematic diagram comparing the leg postures of the flexed-knee walking strategy and the anthropomorphic straight-knee walking strategy provided by this invention.

[0024] Figure 7 This is a schematic diagram of the heel-toe cushioning mechanism provided by the present invention.

[0025] Figure 8 This is a schematic diagram comparing the peak torque of the knee joint provided by the present invention.

[0026] Figure 9 This is a schematic diagram comparing the impact force of the robot's feet on the ground, provided by the present invention.

[0027] Figure 10 This is a schematic diagram of the walking control device for the robot provided by the present invention.

[0028] Figure 11 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0030] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0031] To facilitate a full understanding of the technical solution of this application, the following content is hereby introduced: With the development of robotics technology, the requirements for the naturalness of robot movement, energy efficiency, and reliability are increasing. Currently, the mainstream methods for generating robot gait can be divided into model-based planning methods and imitation learning-based methods.

[0032] Model-based planning methods achieve rapid computation by simplifying dynamics, but the resulting trajectories often exhibit dynamic inconsistencies. For example, to simplify calculations and ensure stability, traditional methods often employ a bent-knee walking strategy, where the robot's thighs and lower legs remain bent during the stance phase. This strategy visually differs significantly from human walking posture and appears unnatural. More importantly, the bent knee joint results in a longer lever arm from the torso weight to the knee joint axis, requiring greater knee joint torque to perform the same task. This not only increases energy consumption but also leads to motor overheating, performance degradation, and even accelerated hardware mechanical wear.

[0033] Imitation learning-based methods can learn natural motion patterns from human motion capture data, but acquiring high-quality and sufficient motion capture data is costly, and kinematic reorientation alone cannot guarantee the dynamic feasibility of transferring these patterns to robot hardware. Furthermore, the trajectories generated by these methods may not fully consider physical constraints such as joint torque limits.

[0034] Therefore, the present invention provides a walking control method, device, electronic device and storage medium for robots, which can reduce knee joint torque and energy consumption, avoid motor overheating and reduce hardware mechanical wear.

[0035] The following is combined with Figures 1-11 This invention describes the robot walking control method, apparatus, electronic device, and storage medium provided by the present invention.

[0036] Figure 1 This is a flowchart illustrating the robot walking control method provided by the present invention, as shown below. Figure 1 As shown, the robot walking control method provided by the present invention can be executed by various types of devices such as a robot control terminal, a robot walking control device, or a computer capable of executing the method of the present invention. Unless otherwise specified, the robot control terminal will be used as an example in the following embodiments.

[0037] As an optional embodiment, the robot's walking control method mainly includes, but is not limited to, the following steps: Step 110: Obtain the robot's body perception data and the relative displacement of the robot's feet relative to the pelvis.

[0038] Propriometry data refers to the robot's internal state and kinematic characteristics. For example, propriometry data can include joint positions, joint velocities, inertial measurement unit data, and actions taken at the previous moment. It can be acquired through hardware such as encoders or inertial sensors configured on the robot body.

[0039] The relative displacement of a robot's feet relative to its pelvis refers to the relative position information of the feet in the pelvic coordinate system. For example, it can be characterized by the displacement of the feet in the forward direction.

[0040] Step 120: Determine a reference body height based on the robot's fully extended leg height; wherein the maximum value of the reference body height is less than the fully extended leg height.

[0041] The height of a robot when its legs are fully extended refers to its body height when its thighs and calves are in a straight line. This height can be calculated, for example, using the robot's fixed physical dimensions and linkage structure parameters.

[0042] Reference body height refers to the desired height parameter set in walking gait control to guide the robot to maintain its position. For example, it can be obtained by setting a value within a reasonable range below the height of the fully extended legs.

[0043] Considering that the knee joint is usually close to fully extended during the middle of the stance phase when humans walk, and to avoid the robot's knee joint from frequently reciprocating near the kinematic singularity in the fully extended position, which would lead to control instability, this invention sets the maximum reference body height to be slightly lower than the height of the fully extended leg. This forces the robot to adopt a leg configuration that is closer to a straight knee during movement, effectively shortening the lever arm and significantly reducing the knee joint torque and energy consumption required to maintain posture.

[0044] Step 130: Determine the target foot pitch angle based on the relative displacement. The target foot pitch angle has a linear mapping relationship with the relative displacement.

[0045] The target foot pitch angle refers to the desired pitch angle required for the robot's foot to contact the ground during walking. For example, it can be calculated by obtaining the displacement of the foot relative to the pelvis in the forward direction and multiplying this relative displacement by an adjustable proportional constant used to control the cushioning strength.

[0046] Step 140: Input the body perception data, reference body height, and target foot pitch angle into the gait control model to obtain the control commands output by the gait control model; control the robot to walk according to the control commands.

[0047] Gait control models refer to network models used to generate action strategies based on the robot's current observed state and desired anthropomorphic gait characteristics. For example, a gait control model can be a neural network model trained using a reinforcement learning framework and a proximal policy optimization algorithm.

[0048] Control commands refer to the specific control signals that drive the robot's joints to perform the expected walking movements. For example, control commands may include proportional-derivative control target position signals instructing the movement of each leg joint, estimated pelvic velocity signals characterizing pelvic motion, and foot contact force signals reflecting the interaction between the foot and the ground. The proportional-derivative control target position refers to the desired joint position parameters input to the robot's underlying proportional-derivative controller. This controller calculates and outputs the actual driving torque based on the deviation between the joint's actual position and the target position parameters, thereby controlling the robot's joints to rotate to the desired angle.

[0049] The robot walking control method provided by this invention guides the robot to adopt a walking mode closer to straight knees by setting a reference body height slightly lower than the height of the legs when fully extended. This effectively shortens the lever arm from the torso weight to the knee joint axis. At the same time, a compliant buffering mechanism is introduced through the linear mapping relationship between the target foot pitch angle and relative displacement, thereby reducing knee joint torque and energy consumption, avoiding motor overheating, and mitigating hardware mechanical wear.

[0050] In another embodiment of the present invention, determining a reference body height based on the robot's fully extended leg height includes: determining a maximum value of the reference body height based on the difference between the fully extended leg height and the target height; determining a range of values ​​for the reference body height based on the maximum value of the reference body height; and determining the reference body height based on the range of values.

[0051] Specifically, to avoid instability caused by frequent reciprocating movements of the robot's knee joint near its fully extended position (kinematic singularity), this invention introduces a target height difference to safely constrain the upper limit of the robot's body height. The maximum reference body height can be calculated by subtracting this target height difference from the pre-acquired fully extended leg height. This maximum value is then used as the upper limit to define a reasonable range of motion as a value interval, within which the reference body height parameters required for the gait template are finally determined.

[0052] For example, let's assume that the height of the robot in its fully extended state, with its thigh and calf in a straight line, is denoted as the height of its fully extended legs ( H full ), and setting the target height difference to 0.02 meters, then the maximum value of the determined reference body height is ( H max The full leg extension height is reduced by 0.02 meters, i.e. H max = H full -0.02m. Then, the reference body height ( H The value range of ) is set to the maximum value of the reference body height ( H max Extending downwards by 0.1 meters, the reference body height value satisfies... Each time a gait template is generated, a specific reference body height is determined by sampling from this value range. This height range design forces the robot to adopt a leg configuration closer to straight knees in order to achieve a higher reference body height during movement, thereby guiding the control strategy to learn the straight-knee walking pattern from the reference trajectory level.

[0053] As an optional embodiment, when generating a gait template in gait control, the set of input gait cycle parameters can be represented as follows: ,in The duration of the gait cycle. To support spatiotemporal parameters such as phase ratio or step size, To determine the reference body height by sampling within the value range, Preset parameters such as maximum foot lift height or maximum foot contact force.

[0054] In addition, the gait template generator can calculate and output the reference angles of all joints and the reference posture of the foot (including pitch angle) online in real time within each control cycle (typically 0.01-0.05 seconds), thereby providing a precise and continuous reference benchmark for subsequent gait control.

[0055] The robot walking control method provided by this invention determines the maximum value of the reference body height based on the difference between the fully extended leg height and the target height, and sets a value range based on the maximum value to determine the specific reference body height. This can effectively avoid the control instability problem caused by the frequent reciprocating motion of the robot's knee joint near the kinematic singularity close to full extension. Thus, while ensuring the stability of system control, it further guides and forces the control strategy to reliably learn the straight-knee walking mode from the reference trajectory level.

[0056] In another embodiment of the present invention, determining the target foot pitch angle based on the relative displacement includes: obtaining a buffer ratio coefficient; obtaining the target foot pitch angle based on the relative displacement and the buffer ratio coefficient; wherein, if the robot's foot is located in front of the pelvis, the relative displacement is greater than zero, and the target foot pitch angle is greater than zero; if the robot's foot is located behind the pelvis, the relative displacement is less than zero, and the target foot pitch angle is less than zero.

[0057] The buffer ratio coefficient refers to an adjustable proportional constant used to control the cushioning intensity of the foot landing. For example, it can be dynamically adjusted according to the robot's actual hardware structure parameters and the physical characteristics of the ground to be walked on.

[0058] Specifically, this invention introduces a dynamic adjustment mechanism in gait control that relates the foot pitch angle to the foot's position relative to the pelvis, mimicking the human walking pattern of heel-to-toe strike, followed by foot-to-toe flattening, and finally toe-to-toe push-off. When the robot's foot is in front of the pelvis at the end of the swing phase, the relative displacement is greater than zero, resulting in a target foot pitch angle greater than zero, i.e., the toes are raised, thus simulating the posture before heel impact. When the foot is behind the pelvis at the beginning of the swing phase, the relative displacement is less than zero, resulting in a target foot pitch angle less than zero, i.e., the toes are pressed down and the heel is raised, thus simulating the posture of toe-toe push-off. When the foot is directly below the pelvis, the relative displacement is close to zero, and the target foot pitch angle is also close to zero, thus making the foot tend to be horizontal.

[0059] For example, the target foot pitch angle can be designed as a linear function of the foot's relative displacement in the forward direction. That is, the target foot pitch angle is equal to the product of the buffer ratio coefficient and the relative displacement. This mechanism introduces a compliant rotational degree of freedom into the robot's foot landing process, which helps to effectively disperse the impact force of the foot and make the walking gait softer.

[0060] The linear mapping relationship between the target's foot pitch angle and relative displacement can be specifically expressed by the following formula: ; in, The target foot pitch angle, This is the buffer ratio coefficient. This refers to the relative displacement of the robot's feet with respect to the pelvis in the forward direction.

[0061] The robot walking control method provided by this invention obtains a buffer ratio coefficient and accurately sets the positive and negative states of the target foot pitch angle according to the different positive and negative states of the relative displacement of the foot in front of or behind the pelvis. It can accurately simulate the dynamic adjustment mechanism of the human foot lifting the toes at the end of the swing phase and lifting the heel at the beginning of the swing phase, thereby introducing a smooth rotational degree of freedom for the foot landing process, further effectively dispersing the impact force of the foot and making the robot's gait more gentle.

[0062] Figure 2 This is a schematic diagram of the training process of the gait control model provided by the present invention, as shown below. Figure 2 As shown, as another optional embodiment provided by the present invention, the gait control model is trained based on the following steps: Step 210: Obtain the robot's raw observation data, input the raw observation data into the initial gait control model, and obtain the raw motion commands output by the initial gait control model.

[0063] Raw observation data refers to various sensor measurements and gait template information used to characterize the robot's current state. For example, raw observation data can be joint positions, joint velocities, inertial measurement unit data, and swing and support state indicators in the gait template segment at the current moment. It can be obtained through the robot's body perception hardware combined with the control system.

[0064] The initial gait control model refers to the basic policy neural network that has not yet completed reinforcement learning training or whose network parameters need to be optimized. For example, the initial gait control model can be an initial policy network built on the framework of the proximal policy optimization algorithm.

[0065] The original motion command refers to the motion feature vector directly output by the initial gait control model based on the currently input original observation data, which is used to control the robot. For example, the original motion command can be the output of continuous motion features such as the control target position of each leg joint.

[0066] Step 220: Perform an observation mirror transformation on the original observation data to obtain mirrored observation data.

[0067] Observation mirror transformation refers to data processing operations that interchange or reverse the direction of data features with left-right symmetry in the observation space. For example, the swing and support status indicators of the left and right feet in the original observation data can be interchanged.

[0068] Mirror observation data refers to virtual state input data obtained after the original observation data has undergone the above-mentioned physical symmetry transformation operation. For example, mirror observation data can be a simulated observation vector after the left and right leg support states are interchanged.

[0069] Step 230: Input the mirror observation data into the initial gait control model and obtain the intermediate motion instructions output by the initial gait control model.

[0070] Intermediate action instructions refer to the forward predicted action output made by the initial gait control model based on the input virtual mirror observation data. For example, intermediate action instructions can be a set of predicted action vectors with mirror features output by the model after receiving the left-right swap state.

[0071] Step 240: Perform motion mirror transformation on the intermediate motion instructions to obtain mirror motion instructions.

[0072] Action mirror transformation refers to the reverse restoration or interchange operation of the action vector output by the policy network in accordance with the symmetry law of physics and dynamics. For example, action mirror transformation can be the operation of interchange the relevant action control parameters of the left leg and the right leg in the intermediate action instruction.

[0073] Mirrored motion instructions refer to reference instruction data that is restored and mapped to the original motion space after intermediate motion instructions have undergone motion mirroring transformation. For example, a mirrored motion instruction can be a target reference motion vector that has been transformed and is used for consistency comparison with the original motion instructions.

[0074] Step 250: Calculate the difference between the original action instruction and the mirror action instruction, and determine the symmetry consistency reward based on the difference.

[0075] Specifically, the present invention achieves physical constraints by directly constraining the consistency between the original action command and the transformed mirror action command.

[0076] For example, the symmetric consistency reward can be represented by the following formula: ; in, For symmetric consistency rewards, To control the scaling factor of the consistency penalty strength, The original action command, This is a mirroring action command.

[0077] Step 260: Based on symmetric consistency rewards, update the network parameters of the initial gait control model until the preset training termination condition is met, and obtain the gait control model.

[0078] The preset training termination condition refers to the set standard used to determine whether the neural network has converged or reached the expected motion control performance. For example, the preset training termination condition may be reaching the set maximum number of reinforcement learning iterations, or the policy can stably and accurately track a reference trajectory with anthropomorphic features and exhibit a high degree of motion symmetry.

[0079] Specifically, within the reinforcement learning training framework, the loss gradient can be calculated and updated via backpropagation based on the reward signal, which includes the symmetry consistency reward. For example, based on the proximal policy optimization algorithm, the weight parameters of the policy network are iteratively adjusted according to the obtained symmetry consistency reward, thereby causing the motion generated by the policy learning to exhibit high dynamic symmetry between the left and right limbs, and outputting a well-trained gait control model after the termination condition is met.

[0080] As an alternative implementation, the reinforcement learning training process described above can be performed in physical simulation software (such as Isaac Gym or MuJoCo) that has established accurate robot models, environments, and physics engines. Furthermore, domain randomization training can be incorporated during the training process, such as introducing randomization factors like terrain friction and robot mass distribution, to perform large-scale training in a parallel simulation environment, thereby effectively improving the robustness of the policy network.

[0081] The robot walking control method provided by this invention performs mirror transformation on the original observation data and corresponding intermediate action commands during the training process of the gait control model, and calculates the difference between the original action commands and the mirror action commands to determine the symmetry consistency reward and update the network parameters accordingly. This enables the movement generated by policy learning to exhibit high dynamic symmetry between the robot's left and right limbs, thereby promoting high bilateral consistency and coordination of the robot's left and right leg movements, making the overall walking gait smoother, more stable, and more in line with the visual and dynamic characteristics of human walking.

[0082] Figure 3 This is a flowchart illustrating the process of determining the mirror action instruction provided by the present invention, as shown below. Figure 3 As shown, as another optional embodiment provided by the present invention, the intermediate action instruction is subjected to an action mirror transformation to obtain a mirror action instruction, including but not limited to the following steps: Step 310: Obtain the target position of the robot's left leg joint, the target position of the right leg joint, and the lateral component of the pelvic linear velocity from the intermediate motion command.

[0083] Specifically, after obtaining the intermediate motion commands output by the initial gait control model based on mirror observation data, it is necessary to extract the control parameters directly related to the robot's left and right limb movements for subsequent symmetry processing. For example, the proportional differential control target positions of each joint of the left and right legs can be extracted from the motion vector corresponding to the intermediate motion command. At the same time, the Y component, i.e., the lateral component, in the estimated pelvic linear velocity of the model can be extracted. In addition, features such as the estimated contact force of the left and right feet in the Z direction can be extracted as needed.

[0084] Step 320: Swap the target positions of the left and right leg joints, and invert the signs of the non-pitch joint target positions after the swap to obtain the mirrored joint target positions.

[0085] Specifically, in order to achieve a mirror mapping of motion commands in physical space, the target position parameters controlling the movement of the left and right legs need to be swapped first. Furthermore, since the rotation direction of some joints of the robot will be physically reversed when mirror symmetry is performed, the numerical signs of the target positions of other joints except for the pitch direction need to be reversed after the swap.

[0086] For example, after assigning the proportional differential control target position of the right leg joint to the corresponding joint of the left leg, if the joint is a non-pitch joint that controls the roll or yaw direction, then the sign of its value needs to be reversed, while the sign of the pitch joint that controls the forward and backward swing remains unchanged, thus forming the target position of the mirror joint.

[0087] Step 330: Invert the sign of the horizontal component to obtain the mirrored horizontal component.

[0088] Specifically, the lateral component of the pelvic linear velocity represents the robot's upward movement tendency to the left and right sides. In a perfectly mirror-symmetric gait, the robot's lateral movement velocities to the left or right should be equal in magnitude but opposite in direction. Therefore, this component needs to be inverted. For example, if the extracted estimated pelvic linear velocity Y component, i.e., the lateral component, is positive, it is inverted to become negative, thus obtaining the mirror lateral component, while the X and Z components of the pelvic linear velocity remain unchanged.

[0089] Step 340: Generate mirror motion commands based on the target position of the mirror joint and the mirror lateral component.

[0090] Specifically, after completing the symmetrical transformation of each key action feature, these transformed parameters are recombined into an action vector consistent with the original action instruction structure, thereby forming the final target reference instruction for consistency comparison.

[0091] For example, the target position of the mirror joint, the mirror lateral component, and other velocity components that remain unchanged can be recombined. If the estimated contact forces of the left and right feet in the Z direction are also obtained during extraction, they can be interchanged and incorporated into the process to generate a complete mirror motion command for reinforcement learning constraint comparison.

[0092] The robot walking control method provided by this invention obtains the target positions of the robot's left and right leg joints and the lateral component of the pelvic linear velocity from the intermediate motion command, and swaps the target positions of the left and right leg joints and inverts the signs of the joint target positions in non-pitch directions. At the same time, it inverts the signs of the lateral component of the pelvic linear velocity to generate a mirror motion command. This method can accurately construct a physical mirror reversal operation that conforms to the kinematics and dynamics laws of the robot, thereby providing an accurate and reliable reference benchmark for calculating motion differences and ensuring that the gait control model can effectively and reasonably learn highly coordinated bilateral symmetrical motion characteristics.

[0093] Figure 4 This is a schematic diagram of the process for determining the reward for trajectory tracking provided by the present invention, as follows: Figure 4 As shown, as another optional embodiment provided by the present invention, after determining the symmetry consistency reward based on the difference, the following steps are included but not limited to: Step 410: Obtain the robot's actual body height and actual foot pose.

[0094] Specifically, during each control cycle of reinforcement learning training, the control system collects the robot's body perception data in real time and obtains the robot's actual state information at the current moment through forward kinematics calculation or state estimation modules.

[0095] For example, the actual height of the pelvis or center of mass relative to the ground can be calculated based on the actual angles of each leg joint and the physical dimensions of the robot's links. At the same time, the actual position and posture of the feet in the world coordinate system can be obtained by combining the data from the inertial measurement unit. This actual posture includes the actual foot pitch angle.

[0096] Step 420: Calculate the height error between the actual body height and the reference body height.

[0097] Specifically, the obtained actual body height is compared with the reference body height output by the gait template generator at the current moment, and the difference between the two is used as the basis for the penalty term. For example, suppose the robot's actual body height is... H actual The reference body height set in the gait template is H ref Then the height error is H actual and Href The numerical difference between the two values, this height error intuitively reflects the degree to which the robot deviates from the target height required by the straight-knee strategy during walking.

[0098] Step 430: Calculate the pose error between the actual foot pose and the target foot pitch angle.

[0099] Specifically, the actual foot pose is compared with the expected reference foot pose, which includes the target foot pitch angle dynamically determined based on relative displacement. The pose error is obtained by calculating the deviation between the actual and expected values ​​in spatial position and rotational attitude. For example, the actual position of the foot in the world coordinate system can be calculated separately. P foot,actual Euclidean distance error between the reference position and the actual attitude P foot,actual The angular deviation between the actual pitch angle and the target foot pitch angle is included, thereby comprehensively quantifying the robot's attitude tracking accuracy when executing the heel-toe cushioning mechanism.

[0100] Step 440: Determine the trajectory tracking reward based on the height error and pose error.

[0101] Specifically, in order to encourage reinforcement learning strategies to accurately track reference trajectories with anthropomorphic characteristics, this invention uses an exponentially decaying function to convert the above-mentioned error into a reward signal; the smaller the error, the higher the reward value obtained.

[0102] For example, the trajectory tracking reward can be expressed as follows: ; in, To track rewards, This is an adjustable scaling factor. This refers to the tracking error, specifically the height error or pose error.

[0103] The robot walking control method provided by this invention obtains the robot's actual body height and actual foot pose, and calculates the height error and pose error between them and the reference body height and the target foot pitch angle to determine the trajectory tracking reward. This method can provide precise optimization signals to the policy network during reinforcement learning training, thereby strongly guiding the robot control strategy to accurately track the preset straight-knee strategy and heel-toe cushioning mechanism, ensuring that the final generated walking gait has highly human-like motion characteristics and meets physical constraints.

[0104] In another embodiment of the present invention, updating the network parameters of the initial gait control model based on the symmetry consistency reward includes: obtaining the task objective reward and the system constraint reward; weighting and summing the task objective reward, the system constraint reward, the trajectory tracking reward, and the symmetry consistency reward to obtain the target reward; and updating the network parameters of the initial gait control model based on the target reward.

[0105] Task objective reward refers to feedback signals used to encourage robots to complete specific high-level instructions or basic motion goals. For example, task objective reward can be a speed tracking reward for the robot during walking, which can be obtained by calculating the deviation between the robot's actual center of mass velocity and the expected input velocity.

[0106] System constraint rewards refer to punitive feedback signals used to limit robot movements to ensure that they operate within physical feasibility and protect the hardware. For example, system constraint rewards can be penalties for joint torque exceeding limits or foot slippage, which can be obtained by monitoring the actual output torque of the underlying motors and the friction cone state between the foot and the ground.

[0107] Specifically, within the reinforcement learning training framework, it is necessary to comprehensively consider the robot's task execution capability, physical safety, anthropomorphic features, and motion symmetry, and linearly combine the above rewards according to the set weight coefficients to form the overall target reward function.

[0108] For example, the target reward can be represented by the following formula: ; in, The target reward is obtained after weighted summation. Rewards for mission objectives. To constrain rewards in the system, As a reward for trajectory tracking, For symmetric consistency rewards, , and These represent the weight coefficients of system constraint reward, trajectory tracking reward, and symmetry consistency reward, respectively.

[0109] After obtaining the target reward, the gradient of the loss function can be calculated using the proximal policy optimization algorithm, and the network parameters of the initial gait control model can be iteratively updated through backpropagation until the policy network converges and outputs the final gait control model.

[0110] The robot walking control method provided by this invention obtains task target reward and system constraint reward, and then weights and sums the task target reward, system constraint reward, trajectory tracking reward and symmetry consistency reward to obtain the target reward to update the network parameters of the initial gait control model. It can comprehensively consider the robot's task execution capability, physical safety constraints, anthropomorphic motion characteristics and bilateral motion symmetry, thereby guiding the reinforcement learning strategy to learn a walking gait that has a natural appearance, high energy efficiency and high coordination under the premise of satisfying physical constraints and hardware protection.

[0111] Figure 5 This is a schematic diagram of the humanoid robot provided by the present invention, as shown below. Figure 5 As shown, the robot can be a humanoid robot platform with 12 active degrees of freedom in its legs (such as the Q5 robot). The robot has an overall bipedal anthropomorphic form. Its upper body includes a head, torso, and arm structures with a white protective shell, while the lower body mainly consists of metal mechanical links and exposed joint actuators. Between the robot's pelvis, thighs, lower legs, and feet, there are clearly visible actuators such as disc motors. These mechanisms correspond to the active degrees of freedom of the robot's hip, knee, and ankle joints, respectively, and are used to receive control commands and drive the movement of the lower limbs.

[0112] Figure 5 The image shows the robot's instantaneous posture while performing a walking task in a real outdoor environment. It can be clearly observed that the robot's left leg is in the support phase, with its left knee joint in a nearly fully extended straight knee configuration. This posture can effectively shorten the lever arm from the torso weight to the knee joint axis. At the same time, the robot's right leg is stepping forward at the end of the swing phase, with its right toes clearly raised, indicating a posture of preparing to land on the heel first. This intuitively demonstrates the heel-toe cushioning mechanism introduced in this invention.

[0113] It should be noted that when the trained gait control model (i.e., control strategy) is deployed to real robot hardware, it can be deployed using a zero-sample transfer method. By combining high-level instructions (such as desired speed) with gait parameters, an anthropomorphic reference trajectory can be generated, thereby driving the robot to achieve anthropomorphic walking with straight knees, heel-toe cushioning, and bilateral coordinated symmetry.

[0114] Figure 6 This is a schematic diagram comparing the leg postures of the flexed-knee walking strategy and the anthropomorphic straight-knee walking strategy provided by this invention, as shown in the diagram. Figure 6 As shown, the differences in the lower limb configuration and force characteristics of the robot under two different control strategies are illustrated.

[0115] Figure 6The left side illustrates a knee-bending walking strategy commonly used in traditional model-based planning methods. In this strategy, the top black and white circles representing the robot's pelvis or center of mass are positioned relatively low. During the support phase, the robot maintains a significant bend between its thigh and lower leg links, for example, an angle of approximately 135 degrees, while its feet are flat against the ground. This bent leg posture not only visually differs significantly from natural human walking and appears unnatural, but more importantly, the bent knee joints result in a longer horizontal distance—the lever arm—from the line of action of the torso's weight to the knee joint axis. This requires greater knee joint torque to maintain the posture when performing the same task, increasing energy consumption, causing motor overheating, and accelerating hardware wear.

[0116] Figure 6 The right side illustrates the anthropomorphic straight-knee walking strategy proposed in this invention. Under this strategy, the black and white alternating circles representing the center of mass are positioned higher, thanks to the optimized design in this invention that sets the reference body height slightly below the height of the fully extended legs. It can be clearly seen that the supporting leg of the robot on the right exhibits a near-straight-knee configuration, with the thigh and lower leg forming a straight knee. This straight-knee strategy effectively shortens the lever arm from the torso weight to the knee joint axis, thus significantly reducing the required knee joint torque when supporting the same torso weight. Furthermore, the robot's foot segments in the right model are no longer rigidly flat but exhibit significant pitch angle changes. The toes of the front foot are raised to simulate a heel-first landing posture, while the heel of the rear foot is raised to simulate a toe-off posture. This fully demonstrates the heel-to-toe cushioning mechanism introduced in this invention. This mechanism dynamically adjusts the foot pitch angle, introducing a smooth rotational degree of freedom into the foot landing process, making the robot's walking gait visually closer to that of a human, while also being more efficient and hardware-friendly at the physical level.

[0117] Figure 7 This is a schematic diagram of the heel-toe cushioning mechanism provided by the present invention, as shown below. Figure 7 As shown, three consecutive side screenshots of the robot walking indoors visually demonstrate the dynamic changes in the expected foot pitch angle relative to the pelvis at different forward and backward positions when the robot executes the control strategy of this invention.

[0118] Figure 7 The image shows a bipedal humanoid robot walking forward in an indoor space. Figure 7In the left-hand screenshot, the robot's right leg is stepping forward at the end of the swing phase, with the right foot positioned in front of the pelvis (i.e., relative displacement greater than zero). Based on the linear mapping relationship between the foot pitch angle and relative displacement according to this invention, the desired foot pitch angle output by the control system is greater than zero, causing the robot's right toes to lift significantly upwards, thus presenting a posture of preparing to land heel-first, effectively simulating the cushioning action of human walking. Figure 7 In the middle screenshot, as the robot moves forward, its right foot gradually transitions to directly below the pelvis. At this point, the relative displacement is close to zero, and the corresponding pitch angle of the target foot also approaches zero, causing the robot's right foot to become horizontal and smoothly and completely flat on the ground, entering a stable support phase. Figure 7 In the screenshot on the right, the robot's left leg is behind its body and about to leave the ground, corresponding to the initial stage of the swing phase. At this time, the left foot is located behind the pelvis (i.e., the relative displacement is less than zero). According to the linear mapping relationship, the expected pitch angle of the foot is less than zero, so that the robot's left foot toes press down and the heel is raised high, thus presenting a posture that simulates the toes pushing off the ground.

[0119] The posture changes in these three consecutive stages demonstrate that the heel-toe cushioning mechanism proposed in this invention successfully breaks the rigid pattern of traditional robots landing flat on their feet, introducing a smooth rotational degree of freedom into the foot landing process. This not only visually replicates the natural cushioning pattern of the human foot landing first, then flattening the foot, and finally pushing off with the toes, but also effectively disperses and reduces the impact force on the foot at the physical level, making the overall walking gait softer and more stable.

[0120] Figure 8 This is a schematic diagram comparing peak torque of the knee joint provided by the present invention, as shown below. Figure 8 As shown, the waveforms of the knee joint torque of the left and right legs (vertical axis is torque, unit Nm) during the robot's walking process of about 3 seconds (horizontal axis is time, unit is s) are displayed. The solid line (corresponding to Left in the figure) represents the left knee joint torque, and the dashed line (corresponding to Right shifted in the figure) represents the right knee joint torque after half a gait cycle of translation and alignment.

[0121] exist Figure 8 In the test results of the traditional knee flexion walking strategy on the left, the peak torque of the knee joint was very large, reaching about 150 Nm (the lowest point of the curve was close to negative 150 Nm). Furthermore, there was a significant deviation in the waveform amplitude between the solid line representing the left leg and the dashed line representing the right leg, indicating a significant asymmetry between the left and right sides. This confirms that the traditional knee flexion strategy, due to its long lever arm, leads to a large consumption of joint torque and poor coordination.

[0122] And in Figure 8In the test results on the right side, after adopting the anthropomorphic straight-knee walking strategy provided by this invention and introducing a mirror symmetry reward mechanism, the absolute value of the peak torque of the knee joint was significantly reduced to about 75 Nm (the lowest point of the curve is near -75 Nm). At the same time, the phase alignment of the solid line representing the left leg and the dashed line representing the right leg was significantly improved, the amplitude error was extremely small, and the two curves were highly overlapping.

[0123] This fully demonstrates that the straight-knee strategy of the present invention effectively shortens the lever arm from the torso weight to the knee joint axis, significantly reducing the required knee joint torque when supporting the same torso weight, thereby reducing motor energy consumption and heat generation, improving hardware friendliness and facilitating long-term stable operation of the hardware system; and the introduction of mirror symmetry reward successfully promotes highly coordinated left and right side movements of the robot, demonstrating greatly improved motion symmetry and stability.

[0124] Figure 9 This is a schematic diagram comparing the impact force of the robot's foot on the ground, as provided by the present invention. Figure 9 As shown, the waveforms of the ground impact force (vertical axis is the impact force, unit is N) experienced by the robot's left foot (corresponding to the left red line in the legend) and right foot (corresponding to the right blue line in the legend) during a 3-second walking process (the horizontal axis is time, in seconds) are displayed.

[0125] Figure 9 The left side shows the impact data when using the traditional walking strategy. It can be observed that at the moment the foot lands in each gait cycle, the impact curve shows an extremely sharp and high peak, with the highest peak reaching about 900N. This violent instantaneous impact is mainly caused by the flat landing of the foot under the traditional strategy, which will bring a huge mechanical burden to the robot's foot structure and the motors of each joint.

[0126] Figure 9 The right side shows the impact force data after applying the heel-toe cushioning mechanism provided by this invention. It can be clearly observed that, under the premise of performing the same walking task, the peak value of the foot impact force is significantly reduced and smoothly limited to about 600N. Moreover, the fluctuation of the curve in the support phase is more gentle, without the violent sudden peaks seen in the left figure.

[0127] This experimental result fully verifies that by establishing a linear mapping relationship between the foot pitch angle and the foot displacement relative to the pelvis, the present invention successfully simulates the human heel-first landing and then foot-flattening cushioning pattern, introducing a smooth rotational degree of freedom into the foot landing process, thereby effectively dispersing and reducing the impact force of the foot, significantly improving the smoothness of the robot's movement, and making the overall walking gait more natural and stable while protecting the hardware mechanical structure.

[0128] It should be noted that the robot walking control method provided by this invention also has the following significant advantages: On the one hand, this invention does not rely on motion capture data, but automatically generates human-like gait through parameterized templates and reinforcement learning, avoiding the expensive and time-consuming acquisition of human motion data and the complex kinematic retargeting process; on the other hand, this invention is compatible with advanced control architectures, and the human-like high-quality reference trajectory it generates can serve as an ideal input for subsequent hierarchical control architectures (such as motion planners and tracking controllers), thereby further improving the overall motion performance of the robot.

[0129] Figure 10 This is a schematic diagram of the walking control device for the robot provided by the present invention, as shown below. Figure 10 As shown, it mainly includes, but is not limited to: The data acquisition module 1010 is used to acquire the robot's body perception data and the relative displacement of the robot's feet relative to the pelvis.

[0130] The height determination module 1020 is used to determine a reference body height based on the fully extended leg height of the robot; wherein the maximum value of the reference body height is less than the fully extended leg height.

[0131] The pitch angle determination module 1030 is used to determine the target foot pitch angle based on the relative displacement, wherein the target foot pitch angle and the relative displacement are linearly mapped.

[0132] The motion control module 1040 is used to input the body perception data, the reference body height, and the target foot pitch angle into the gait control model, obtain the control commands output by the gait control model, and control the robot to walk according to the control commands.

[0133] It should be noted that the robot walking control device provided by the present invention can execute the robot walking control method described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.

[0134] The robot walking control device provided by the present invention guides the robot to adopt a walking mode closer to straight knees by setting a reference body height slightly lower than the height of the legs when fully extended. This effectively shortens the lever arm from the torso weight to the knee joint axis. At the same time, a compliant buffering mechanism is introduced through the linear mapping relationship between the target foot pitch angle and the relative displacement, thereby reducing knee joint torque and energy consumption, avoiding motor overheating and slowing down hardware mechanical wear.

[0135] Figure 11 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 11As shown, the electronic device may include a processor 1110, a communications interface 1120, a memory 1130, and a communication bus 1140, wherein the processor 1110, the communications interface 1120, and the memory 1130 communicate with each other via the communication bus 1140. The processor 1110 can call logic instructions in the memory 1130 to execute a robot walking control method, which includes: acquiring the robot's body perception data and the relative displacement of the robot's feet relative to the pelvis; determining a reference body height based on the robot's fully extended leg height, wherein the maximum value of the reference body height is less than the fully extended leg height; determining a target foot pitch angle based on the relative displacement, wherein the target foot pitch angle has a linear mapping relationship with the relative displacement; inputting the body perception data, the reference body height, and the target foot pitch angle into a gait control model, and acquiring control instructions output by the gait control model; and controlling the robot to walk according to the control instructions.

[0136] Furthermore, the logical instructions in the aforementioned memory 1130 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0137] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the robot walking control method provided by the above methods. The method includes: acquiring the robot's body perception data and the relative displacement of the robot's feet relative to the pelvis; determining a reference body height based on the robot's fully extended leg height, wherein the maximum value of the reference body height is less than the fully extended leg height; determining a target foot pitch angle based on the relative displacement, wherein the target foot pitch angle has a linear mapping relationship with the relative displacement; inputting the body perception data, the reference body height, and the target foot pitch angle into a gait control model to obtain control commands output by the gait control model; and controlling the robot to walk according to the control commands.

[0138] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a robot walking control method provided by the above methods. The method includes: acquiring the robot's body perception data and the relative displacement of the robot's feet relative to the pelvis; determining a reference body height based on the robot's fully extended leg height, wherein the maximum value of the reference body height is less than the fully extended leg height; determining a target foot pitch angle based on the relative displacement, wherein the target foot pitch angle has a linear mapping relationship with the relative displacement; inputting the body perception data, the reference body height, and the target foot pitch angle into a gait control model to obtain control commands output by the gait control model; and controlling the robot to walk according to the control commands.

[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0140] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A walking control method of a robot, characterized by, include: Acquire the robot's body perception data, as well as the relative displacement of the robot's feet relative to the pelvis; A reference body height is determined based on the robot's fully extended leg height; wherein the maximum value of the reference body height is less than the fully extended leg height. The target foot pitch angle is determined based on the relative displacement, and the target foot pitch angle has a linear mapping relationship with the relative displacement; The body perception data, the reference body height, and the target foot pitch angle are input into the gait control model to obtain the control commands output by the gait control model; the robot is then controlled to walk according to the control commands.

2. The walking control method of the robot according to claim 1, characterized by, Determining the reference body height based on the robot's fully extended leg height includes: The maximum value of the reference body height is determined based on the difference between the fully extended leg height and the target height; Based on the maximum value of the reference body height, the range of values ​​for the reference body height is determined; The reference body height is determined based on the range of values.

3. The walking control method of the robot according to claim 1, characterized by, Determining the target foot pitch angle based on the relative displacement includes: Obtain the buffer ratio coefficient; The target foot pitch angle is obtained based on the relative displacement and the buffer ratio coefficient. Wherein, if the robot’s foot is located in front of the pelvis, the relative displacement is greater than zero, and the pitch angle of the target foot is greater than zero; If the robot's foot is located behind the pelvis, the relative displacement is less than zero, and the pitch angle of the target foot is less than zero.

4. The walking control method of the robot according to claim 1, characterized by, The gait control model is trained based on the following steps: Obtain the robot's raw observation data, input the raw observation data into the initial gait control model, and obtain the raw motion commands output by the initial gait control model; The original observation data is subjected to an observation mirror transformation to obtain mirrored observation data; The mirror observation data is input into the initial gait control model to obtain the intermediate motion commands output by the initial gait control model; Perform an action mirror transformation on the intermediate action command to obtain a mirror action command; Calculate the difference between the original action instruction and the mirror action instruction, and determine the symmetry consistency reward based on the difference; Based on the symmetry consistency reward, the network parameters of the initial gait control model are updated until the preset training termination condition is met, thus obtaining the gait control model.

5. The walking control method of a robot according to claim 4, characterized by, The step of performing an action mirroring transformation on the intermediate action instruction to obtain a mirrored action instruction includes: Obtain the target position of the robot's left leg joint, the target position of the right leg joint, and the lateral component of the pelvic linear velocity from the intermediate motion command; The target positions of the left and right leg joints are interchanged, and the signs of the non-pitch joint target positions are inverted to obtain mirrored joint target positions. Invert the sign of the horizontal component to obtain the mirrored horizontal component; The mirror motion command is generated based on the target position of the mirror joint and the mirror lateral component.

6. The walking control method of the robot according to claim 4, characterized by, After determining the symmetry consistency reward based on the difference, the method further includes: Obtain the robot's actual body height and actual foot position; Calculate the height error between the actual body height and the reference body height; Calculate the pose error between the actual foot pose and the target foot pitch angle; The trajectory tracking reward is determined based on the height error and the pose error.

7. The walking control method of a robot according to claim 6, wherein The process of updating the network parameters of the initial gait control model based on the symmetric consistency reward includes: Obtain both task objective rewards and system constraint rewards; The target reward is obtained by weighted summing of the task objective reward, the system constraint reward, the trajectory tracking reward, and the symmetry consistency reward. The network parameters of the initial gait control model are updated based on the target reward.

8. A walking control device of a robot characterized by comprising: include: The data acquisition module is used to acquire the robot's body perception data and the relative displacement of the robot's feet relative to the pelvis. A height determination module is used to determine a reference body height based on the robot's fully extended leg height; wherein the maximum value of the reference body height is less than the fully extended leg height; The pitch angle determination module is used to determine the target foot pitch angle based on the relative displacement, wherein the target foot pitch angle and the relative displacement have a linear mapping relationship. The motion control module is used to input the body perception data, the reference body height, and the target foot pitch angle into the gait control model, obtain the control commands output by the gait control model, and control the robot to walk according to the control commands.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the walking control method for the robot as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by the processor, it implements the walking control method for the robot as described in any one of claims 1 to 7.