Method and device for training a motion control policy network of a legged robot with a floating base based on reinforcement learning

By training the baseline posture control model through reinforcement learning, calculating the changes in joint positions and constructing reward terms, the problem of motion incoordination caused by neglecting joint position optimization in existing technologies is solved, and stable control and efficient terrain adaptation of legged robots in complex environments are achieved.

CN121179441BActive Publication Date: 2026-03-31SHENZHEN ZHUJI POWER TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, only the end effector of legged robots is optimized while joint position optimization is ignored, which leads to uncoordinated movement and delayed posture response, affecting the stability and terrain adaptability of the robot in unstructured environments.

Method used

By using a reinforcement learning-based method, the changes in joint position are calculated and a reward term is constructed. The base posture control model is trained to achieve coordinated control of joint movements and base posture. The changes in joint position are handled by inverse kinematics and the pseudo-inverse matrix of the Jacobian matrix. Singularities are handled by damped least squares method, and joint cooperative motion is optimized.

Benefits of technology

It improves the stability and terrain adaptability of legged robots in unstructured environments, and enhances dynamic balance performance and the reliability of high-mobility actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121179441B_ABST
    Figure CN121179441B_ABST
Patent Text Reader

Abstract

The present disclosure provides a kind of movement control strategy network training method and device of legged robot with floating base based on reinforcement learning, and it is related to robot technical field.The method comprises: obtaining the position of each joint and the position of each foot of legged robot before and after executing target action;According to the position change of each foot before and after target action is executed, the position change of each joint is calculated;According to the position of each joint before target action is executed and the position change of each joint, the reference value of the position of each joint after target action is executed is calculated;According to the reference value and the position of each joint after target action is executed, reward item is constructed, and base posture control model is trained using reward item.The present disclosure can realize that the coordinated joint action that is more in line with kinematic constraint is generated when legged robot adjusts base posture by the base posture control model trained, and then the terrain adaptability, dynamic balance performance and execution reliability of high maneuvering action are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of robotics technology, and to a method and apparatus for training a motion control strategy network for a legged robot with a floating base based on reinforcement learning. Background Technology

[0002] In the motion control of legged robots, the base typically refers to the robot's reference coordinate system, which is mainly used to describe the robot's overall spatial posture. A typical base posture is characterized by three angles: Roll, Pitch, and Yaw, formed by rotation around the three axes (X-axis, Y-axis, and Z-axis) of the robot's coordinate system.

[0003] Unlike robots with a fixed base, legged robots do not have a fixed base; instead, they are in a state of free motion, often referred to as a floating base. This floating base typically uses the robot's center of gravity or torso region as a reference point, and relies on attitude information collected by sensors such as inertial measurement units (IMUs) for real-time estimation and correction during movement. Therefore, in engineering implementations, the location of the IMU is often set as the base's position to ensure consistency in attitude description.

[0004] In some related technologies, such as Chinese patent application CN118331052A, a robotic arm path planning method based on an improved DDPG algorithm is disclosed. In this method, the path planning model consists of a state space, an action space, and a reward function. The state space includes the three-dimensional coordinates of the robotic arm's end effector, the three-dimensional coordinates of the target point, and the attitude angles of the end effector. The action space includes the displacement vector of the end effector, the rotation attitude angles, and the opening and closing states of the gripper. The reward function is initially a sparse reward mechanism, providing a reward only when the gripper reaches the target position (e.g., within a 5cm tolerance range). To improve training efficiency, this method introduces HER (Hindsight Experience Replay) technology, relabeling the states where the target is achieved or not achieved as potential target states, and constructing new reward samples accordingly, thus achieving efficient training under sparse reward conditions.

[0005] For example, the paper "A DeepReinforcement-Learning Approach for Inverse Kinematics Solution of a HighDegree of Freedom Robotic Manipulator" (Aryslan Malik et al., 2022-04-02) discloses a method for solving the inverse kinematics of a high degree of freedom robotic arm based on deep reinforcement learning. Its main idea is to use a deep Q-network (DQN) to solve the inverse kinematics (IK) problem of a 7-degree-of-freedom robotic arm. In order to realize the automatic generation of joint space trajectory from the trajectory of the end effector, the method also constructs a forward kinematics model based on the Product Index (PoE) method and uses the DQN network as the control strategy, and uses the State-Action-Reward-Transition (SART) mechanism to learn the joint action.

[0006] While the aforementioned methods all involve deep reinforcement learning to train robot control strategies and employ end-effector (foot) pose state modeling, their purpose is not to serve "robot base posture control," but rather to optimize the trajectory of the end-effector itself. However, optimizing only the end-effector while neglecting the optimization of the legged robot's joint positions can lead to stiff joint movements, poor coordination, and even body swaying, loss of balance, falls, and inability to effectively conform to complex terrain when the legged robot attempts to adjust its body posture (such as lifting its head when going uphill or tilting sideways when crossing obstacles). These issues directly weaken the legged robot's ability to walk stably in unstructured environments, its terrain adaptability, and its reliability in performing highly dynamic actions (such as jumping and sharp turns). Summary of the Invention

[0007] This disclosure provides a method and apparatus for training a motion control strategy network for a legged robot with a floating base based on reinforcement learning. It aims to solve the problems of motion incoordination and lag in posture response caused by only optimizing the end effector and neglecting the optimization of the joint position of the legged robot in related technologies. The trained base posture control model enables the legged robot to generate more coordinated joint movements that conform to kinematic constraints when adjusting the base posture, thereby enabling the legged robot to achieve more stable body posture control in unstructured environments, significantly improving its terrain adaptability, dynamic balance performance and the reliability of high-mobility motion execution.

[0008] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure.

[0009] According to a first aspect of this disclosure, a method for training a motion control policy network for a legged robot with a floating basis based on reinforcement learning is provided, comprising:

[0010] The positions of each joint and each foot of the legged robot are obtained before and after performing the target action;

[0011] Calculate the positional change of each joint based on the positional changes of each foot before and after the execution of the target action;

[0012] Based on the positions of each joint before the target action is executed and the amount of position change of the joints, calculate the reference values ​​of the positions of each joint after the target action is executed;

[0013] A reward term is constructed based on the reference value and the position of each joint after the target action is executed, and the base posture control model is trained using the reward term.

[0014] In an exemplary embodiment of this disclosure, calculating the positional change of each joint based on the positional changes of each foot before and after the execution of the target action includes:

[0015] For each of the aforementioned soles, the positional change of each of the aforementioned soles is determined based on the position of each of the aforementioned soles before and after the execution of the target action;

[0016] The positional changes of each joint are mapped based on the positional changes of each foot sole.

[0017] In one exemplary embodiment of this disclosure, mapping the positional changes of each joint based on the positional changes of each of the foot soles includes:

[0018] The positional changes of each joint are mapped based on the positional changes of each foot using inverse kinematics.

[0019] In one exemplary embodiment of this disclosure, mapping the positional changes of each joint based on the positional changes of each foot using inverse kinematics includes:

[0020] Obtain the preset Jacobian matrix for each of the joints; the Jacobian matrix represents the linear relationship between the positional change of the foot and the positional change of the joint;

[0021] The positional changes of each joint are calculated based on the positional changes of each foot and the pseudo-inverse of the Jacobian matrix.

[0022] In one exemplary embodiment of this disclosure, calculating the positional change of each joint based on the positional change of each of the foot soles and the pseudo-inverse of the Jacobian matrix includes:

[0023] If the Jacobian matrix is ​​not close to a singularity, then the positional change of each joint is calculated based on the positional change of each foot and the pseudo-inverse of the Jacobian matrix.

[0024] If the Jacobian matrix is ​​close to the singular point, a damped least squares method is used to add a regularization term to the pseudo-inverse of the Jacobian matrix to obtain an updated pseudo-inverse matrix. Based on the positional changes of each foot and the updated pseudo-inverse matrix, the positional changes of the joint are calculated.

[0025] In one exemplary embodiment of this disclosure, mapping the positional changes of each joint based on the positional changes of each of the foot soles includes:

[0026] For each of the aforementioned changes in the position of the sole of the foot, the change in the position of the sole of the foot is decomposed into position change components in each direction under the coordinate system in which it is located;

[0027] Map the positional change components of each foot sole to the positional change components of each joint;

[0028] The position change components of each joint are used to construct the position change amount of the joint.

[0029] In one exemplary embodiment of this disclosure, constructing a reward item based on the reference value and the position of each joint after the target action is performed includes:

[0030] For each joint, determine the positional difference between a reference value for the joint's position and the actual value of the joint's position after the target action is performed;

[0031] The reward term is constructed based on the norm of the positional difference.

[0032] In one exemplary embodiment of this disclosure, the method further includes:

[0033] For each joint, at least one other gap generated by the joint during the execution of the target action is identified; the other gap is a gap corresponding to factors other than the positional gap.

[0034] The construction of the reward term based on the norm of the positional difference includes:

[0035] A reward term is constructed based on the norm of the location gap and the norm of at least one of the other gaps.

[0036] In one exemplary embodiment of this disclosure, determining, for each of the joints, at least one other gap generated by the joint during the execution of the target action includes:

[0037] For each joint, obtain the position of the joint at each moment during the execution of the target action;

[0038] Based on the positional change of the joint between adjacent moments during the execution of the target action, the rotational smoothing gap of the joint is determined; the rotational smoothing gap is positively correlated with the positional change, and the rotational smoothing gap belongs to at least one other gap.

[0039] In one exemplary embodiment of this disclosure, training the basis pose control model using the reward term includes:

[0040] Obtain the reward value output by the reward item;

[0041] The baseline attitude control model is trained using reinforcement learning based on the reward value until the reward value output by the reward item is less than the first reward threshold, thus obtaining the trained baseline attitude control model.

[0042] In one exemplary embodiment of this disclosure, the base attitude control model includes a motion control policy network;

[0043] The step of training the baseline attitude control model using reinforcement learning based on the reward value until the reward value output by the reward item is less than a first reward threshold, thereby obtaining the trained baseline attitude control model, includes:

[0044] Based on the reward value, the parameters of the motion control policy network in the base attitude control model are adjusted to obtain a new base attitude control model.

[0045] The new baseline attitude control model is used for reinforcement learning until the reward value output by the reward item is less than the first reward threshold, thus obtaining the trained baseline attitude control model.

[0046] In one exemplary embodiment of this disclosure, the step of performing reinforcement learning through the new basis pose control model until the reward value output by the reward term is less than a first reward threshold, thereby obtaining a trained basis pose control model, includes:

[0047] If the reward value is greater than or equal to the first reward threshold, then the training weight corresponding to the reward value is obtained; the training weight represents the degree of reinforcement learning training of the base posture control model, and the training weight and the reward value are positively correlated.

[0048] The baseline posture control model is trained using reinforcement learning based on the reward value and the training weights until the reward value output by the reward item is less than the first reward threshold, thus obtaining the trained baseline posture control model.

[0049] In one exemplary embodiment of this disclosure, the step of training the basis posture control model using reinforcement learning based on the reward value and the training weights until the reward value output by the reward item is less than the first reward threshold, thereby obtaining a trained basis posture control model, includes:

[0050] If the reward value is less than the second reward threshold, then the first training weight corresponding to the reward value is determined, and the base posture control model is trained by reinforcement learning based on the reward value and the first training weight until the reward value output by the reward item is less than the first reward threshold, and the trained base posture control model is obtained; the second reward threshold is greater than the first reward threshold.

[0051] If the reward value is greater than or equal to the second reward threshold, then the second training weight corresponding to the reward value is determined, and the base posture control model is trained by reinforcement learning based on the reward value and the second training weight until the reward value output by the reward item is less than the first reward threshold, and the trained base posture control model is obtained; the second training weight is greater than the first training weight.

[0052] In one exemplary embodiment of this disclosure, the step of training the basis posture control model using reinforcement learning based on the reward value until the reward value output by the reward item is less than a first reward threshold, thereby obtaining a trained basis posture control model, includes:

[0053] The training duration is statistically analyzed, and the learning factor corresponding to the training duration is obtained; the learning factor represents the difficulty of reinforcement learning training of the basis posture control model, and the learning factor is positively correlated with the training duration.

[0054] The baseline attitude control model is trained using reinforcement learning based on the reward value and the learning factor until the reward value output by the reward item is less than the first reward threshold, thus obtaining the trained baseline attitude control model.

[0055] In one exemplary embodiment of this disclosure, after calculating the positional change of each joint based on the positional changes of each foot before and after the execution of the target action, the method further includes:

[0056] For each of the joints, the physical constraint range corresponding to the joint is obtained;

[0057] If the position change of the joint is within the corresponding physical constraint range, then the step of calculating the reference value of the position of each joint after the target action is executed based on the position of each joint before the target action is executed and the position change of the joint is executed;

[0058] If the position change of the joint exceeds the corresponding physical constraint range, the position change of the joint is updated based on the preset joint executable strategy to obtain the updated position change. Then, the step of calculating the reference value of the position of each joint after the target action is executed is performed based on the position of each joint before the target action is executed and the position change of the joint.

[0059] In one exemplary embodiment of this disclosure, the physical constraint range includes at least one of the joint angle range and the motor torque range.

[0060] In one exemplary embodiment of this disclosure, the reward term is a quadratic norm.

[0061] According to a second aspect of this disclosure, a motion control method for a legged robot with a floating basis based on reinforcement learning is provided, comprising:

[0062] Obtain the current motion state information of the legged robot;

[0063] Based on the current motion state information and the trained baseline posture control model, an action strategy for controlling the legged robot's motion is output.

[0064] The trained basal attitude control model is obtained according to the reinforcement learning-based motion control strategy network training method for a legged robot with a floating basal base described in the above embodiment.

[0065] According to a third aspect of this disclosure, a training device for a motion control policy network of a legged robot with a floating basis based on reinforcement learning is provided, comprising:

[0066] The training position acquisition module is used to acquire the positions of each joint and each foot sole of the legged robot before and after performing the target action;

[0067] The joint position change calculation module is used to calculate the position change of each joint based on the position change of each foot before and after the target action is performed.

[0068] The joint position reference value calculation module is used to calculate the reference value of the position of each joint after the target action is executed, based on the position of each joint before the target action is executed and the change in the position of the joint.

[0069] The model training module is used to construct a reward term based on the reference value and the position of each joint after the target action is executed, and to train the base posture control model using the reward term.

[0070] According to a fourth aspect of this disclosure, a motion control device for a legged robot with a floating base based on reinforcement learning is provided, the device comprising:

[0071] The running status information acquisition module is used to acquire the current motion status information of the legged robot;

[0072] The motion strategy output module is used to output a motion strategy for controlling the movement of the legged robot based on the current motion state information and the trained base posture control model.

[0073] The trained basal attitude control model is obtained according to the reinforcement learning-based motion control strategy network training method for a legged robot with a floating basal base described in the above embodiment.

[0074] According to a fifth aspect of this disclosure, an electronic device is provided, comprising:

[0075] Processor; and

[0076] A memory that stores computer-readable instructions, which, when executed by a processor, implement the method as described in the above embodiments.

[0077] According to a sixth aspect of this disclosure, a robot is provided, comprising:

[0078] Processor; and

[0079] A memory that stores computer-readable instructions, which, when executed by a processor, implement the method as described in the above embodiments.

[0080] In one exemplary embodiment of this disclosure, the legged robot includes any one of a quadruped robot, a bipedal robot, a humanoid robot, and a mobile robot.

[0081] According to a seventh aspect of this disclosure, a computer-readable storage medium is provided that stores computer program code instructions, which, when invoked by a robot's processor, cause the robot to perform the method as described in the above embodiments.

[0082] As can be seen from the above technical solution, this disclosure possesses at least one of the following advantages and positive effects:

[0083] This disclosure provides a reinforcement learning-based method for training a motion control strategy network for a legged robot with a floating base. The method acquires the positions of each joint and foot before and after the robot performs a target action. Based on the changes in foot position, the positional changes of each joint can be calculated. Then, based on the positions of each joint before the target action and the changes in joint position, reference values ​​for each joint position are calculated. A reward term is constructed based on the reference values ​​of each joint under ideal conditions and the actual positions of each joint after the target action. This reward term is used to train the base posture control model. This allows the reinforcement learning strategy to explicitly learn the mapping relationship between base posture adjustment and joint coordinated motion during the training process. The trained base posture control model enables the legged robot to generate more kinematically consistent coordinated joint movements when adjusting its base posture. This effectively overcomes the problems of stiff joint movements and delayed trunk posture response caused by optimizing only the end effector trajectory in related technologies. Therefore, the trained base posture control model enables the legged robot to achieve more stable body posture control in unstructured environments, significantly improving its terrain adaptability, dynamic balance performance, and the reliability of high-maneuverability actions. Attached Figure Description

[0084] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0085] Figure 1 A system architecture diagram is shown that can be applied to the reinforcement learning-based motion control policy network training method for a legged robot with a floating base, as described in the embodiments of this disclosure.

[0086] Figure 2 The diagram illustrates a flowchart of a motion control policy network training method for a legged robot with a floating base based on reinforcement learning, according to an embodiment of this disclosure.

[0087] Figure 3 A schematic diagram of a legged robot performing a target action is shown in an embodiment of this disclosure.

[0088] Figure 4 A schematic diagram of another legged robot performing a target action is shown in an embodiment of this disclosure.

[0089] Figure 5 The diagram illustrates a process for calculating the positional change of each joint based on the positional changes of each foot before and after the execution of a target action, according to an embodiment of this disclosure.

[0090] Figure 6 A schematic diagram of a legged robot according to an embodiment of the present disclosure is shown.

[0091] Figure 7 This illustration shows a flowchart of a process for constructing a reward item based on a reference value and the position of each joint after the target action is performed, according to an embodiment of this disclosure.

[0092] Figure 8 A flowchart illustrating a process for training a basis posture control model using reward terms is shown in an embodiment of this disclosure.

[0093] Figure 9 A flowchart illustrating a motion control method for a legged robot with a floating base based on reinforcement learning, according to an embodiment of this disclosure, is shown.

[0094] Figure 10 A schematic diagram illustrating the principle of a motion control method for a legged robot with a floating base based on reinforcement learning, according to an embodiment of this disclosure, is shown.

[0095] Figure 11 A block diagram of a motion control policy network training device for a legged robot with a floating base based on reinforcement learning, according to an embodiment of this disclosure, is shown.

[0096] Figure 12 A block diagram of a motion control device for a legged robot with a floating base based on reinforcement learning, according to an embodiment of this disclosure, is shown.

[0097] Figure 13 A schematic diagram of a legged robot according to an embodiment of this disclosure is shown.

[0098] Figure 14 A schematic diagram of another legged robot according to an embodiment of this disclosure is shown.

[0099] Figure 15 A schematic diagram of yet another legged robot according to an embodiment of this disclosure is shown.

[0100] Figure 16 A schematic diagram of the structure of a computer system suitable for implementing the embodiments of the present disclosure is shown.

[0101] Figure 17 A schematic diagram of a computer-readable storage medium according to an embodiment of the present disclosure is shown. Detailed Implementation

[0102] In this disclosure, the terms "first" and "second" are used for description only and do not indicate relative importance or imply the number of technical features. Therefore, the features referred to as "first" or "second" may explicitly or implicitly include at least one of those features. "A plurality of" means at least two, unless otherwise expressly defined.

[0103] First, the relevant terms used in the exemplary embodiments of this disclosure will be explained:

[0104] Legged robots are biomimetic mobile robots with multiple mechanical legs, capable of walking, running, or traversing through alternating leg support and swinging. Legged robots interact with their environment via discrete foot contact points, maintaining high mobility and adaptability in unstructured, rugged, soft, or obstacle-rich terrain. They possess a floating base, meaning their body (base) has no fixed fulcrum in space; their posture changes freely with movement, requiring them to rely on their own perception and control to maintain dynamic balance.

[0105] Structurally, legged robots consist of a torso (base) and multiple mechanical legs with various degrees of freedom. Each leg contains one or more joints (such as hip, knee, and ankle joints), driven by motors, reducers, and actuators, and equipped with sensors such as angle encoders, force sensors, or tactile sensors to perceive its own state and contact information with the environment. Based on the number and configuration of the legs, legged robots can be classified into bipedal robots (such as humanoid robots), quadrupedal robots (such as canine-like robots), and hexapedal robots (such as insect-like robots), etc.

[0106] Floating base: In a legged robot, the body (i.e., the base) has no fixed connection points with the environment during movement. Its spatial pose (including position and orientation) can be determined by the interaction between the legged robot's own dynamics and the environment, allowing it to move freely. The overall motion freedom of a legged robot with a floating base can include the six-dimensional pose of the floating base (three translational degrees of freedom and three rotational degrees of freedom, namely Roll, Pitch, and Yaw). Its stability can be maintained by actively controlling the coordinated movement of the leg joints.

[0107] In some embodiments, the center of the legged robot's torso or the location of its inertial measurement unit (IMU) can be used as a reference point for the floating base to uniformly describe the geometric relationship between the body's posture and limb movement.

[0108] Reinforcement learning is a machine learning method whose basic principle is that an agent learns how to choose actions through trial and error in continuous interaction with the environment, aiming to minimize the long-term accumulated reward signal. In this framework, the agent perceives the state of the environment at each moment, selects an action based on the current policy, and executes it. The environment then transitions to a new state and provides a scalar reward. The agent's goal is to continuously optimize its policy to minimize the expected total reward obtained in the long run.

[0109] A motion control policy network (PCN) is a neural network model used to generate joint control commands for a legged robot. It is a concrete implementation of the policy network within a reinforcement learning framework. This PCN takes the current state of the legged robot (such as joint positions, foot position, and historical actions) as input and outputs control variables such as the target joint position, velocity, or torque for the next moment. In some embodiments, the PCN can be trained by minimizing a kinematic consistency-based reward signal. Its goal is to learn the mapping relationship from the intention to adjust the base posture (implicit in changes in foot position) to coordinated joint movements, thereby achieving active and smooth control of the floating base posture. After training, the PCN can be deployed on the robot body and respond to environmental changes in real time during runtime, autonomously generating control commands that conform to dynamic and kinematic constraints.

[0110] Target action: During reinforcement learning training, this refers to a set of control commands output by the policy network based on the current state, used to drive the robot to perform an exploratory or strategic movement. This target action can be manifested as the target position, velocity, or torque of each joint, and its purpose is to adjust the robot's base posture, such as raising its head when going uphill, tilting to the side when crossing obstacles, or coordinating trunk yaw when turning.

[0111] Joint position: refers to the actual pose of all actively driven joints of a legged robot in the body coordinate system. For rotary joints, this position is represented by joint angles; for linear joints, it is represented by linear displacements. These data are collected in real time by angle encoders, resolvers, or other position sensors installed at each joint, forming a joint space state vector describing the internal configuration of the legged robot. Specifically, joint position data may include, but is not limited to, the actual angle values ​​or displacements of each driven joint such as the hip, knee, and ankle joints.

[0112] Foot position: refers to the three-dimensional spatial coordinates of each foot (i.e., the end effector) in the body coordinate system. The body coordinate system has its origin at the center of the robot's torso or the location of the inertial measurement unit (IMU), with the X-axis pointing forward, the Y-axis pointing to the left, and the Z-axis pointing vertically upward. The foot position can be calculated from joint positions and link parameters through forward kinematics, or it can be estimated by combining IMU data, foot contact states, and kinematic models to characterize the contact relationship between the legged robot and its environment and its intended base posture changes. Specifically, the foot position data can include the precise three-dimensional coordinate values ​​of each foot in the body coordinate system.

[0113] Foot position change: refers to the difference in the position of each foot before and after the execution of the target action, that is, the actual spatial displacement of each foot in the body coordinate system. This change reflects the overall motion effect of the legged robot interacting with the environment during a single action, and also implies the intention to adjust the base posture, such as overall lifting, lateral displacement, or rotational tendency.

[0114] Foot position change component: refers to the overall position change of the sole of one leg of a legged robot during the execution of the target action, which is decomposed into the displacement in a single direction according to the orthogonal directions of a preset coordinate system (usually the body coordinate system) (such as X-axis: front-back direction, Y-axis: left-right direction, Z-axis: up-down direction).

[0115] The change in position of each joint refers to the amount of displacement that each joint should produce to achieve the observed change in foot position. The change in position of each joint is not the actual displacement of each joint after the robot actually performs the operation (i.e., the difference between the actual joint positions), but a set of reference joint adjustment amounts derived using inverse kinematics methods based on the current robot configuration and foot displacement requirements.

[0116] Joint position change component: refers to the angular or displacement adjustment that a joint should produce in response to the position change component of the foot in a specific direction. Since the movement of a single joint usually affects the foot in multiple directions, and conversely, the movement of the foot in a certain direction is often achieved by the coordination of multiple joints, the contribution of the foot position change component in each direction to the relevant joints is calculated separately during the mapping process.

[0117] Reference values ​​for the positions of each joint after the target action is executed: This refers to an idealized target joint position constructed during reinforcement learning training to evaluate the rationality of the actions generated by the policy network. Specifically, this reference value is obtained based on the actual positions and changes in joint positions of each joint before the target action is executed. It represents the theoretically optimal coordinated position that each joint should achieve in the current legged robot configuration to realize the observed change in foot position.

[0118] Reward Term: This refers to a quantitative metric used during the training of the baseline posture control model to measure the kinematic consistency of the actions generated by the policy. The reward term is constructed by comparing the deviation between the actual joint positions after the target action is executed and the ideal joint position reference values ​​derived from changes in foot position. It should be noted that the value of the reward term is calculated from indicators such as joint position tracking error. The smaller the output value of the reward term, the closer the actual movement of the legged robot joints matches the expected output of the policy network, and the better the control effect.

[0119] A basal attitude control model is a policy model used to generate joint control commands for legged robots to actively adjust the floating basal attitude (including roll, pitch, yaw angles, and center of mass position). This basal attitude control model can be represented as a neural network-like policy network. Its inputs include the robot's current state (such as joint position, foot position, and basal attitude), and its output is the target joint position or control variable for the next moment. This basal attitude control model can be trained end-to-end using reinforcement learning algorithms. The goal is to learn the mapping relationship from the intention to adjust the basal attitude (implicit in foot movement) to coordinated joint movements, thereby maintaining stable and flexible basal control capabilities in complex terrain or dynamic tasks.

[0120] Inverse kinematics (IK) refers to the process of determining the positions or angles of the corresponding joints based on the desired position or displacement of the robot's end effector (the foot in legged robots). Unlike forward kinematics (which calculates the end effector position from joint positions), inverse kinematics addresses the question of "how the joints should move to achieve a specific foot target." Inverse kinematics maps observed changes in foot position to ideal changes in joint position, thus providing a basis for constructing kinematically consistent reference values. Since legged robots have redundant degrees of freedom (more joints than required for foot pose), optimal solutions can be selected using optimization criteria (such as minimum joint movement and singularity avoidance).

[0121] The Jacobian matrix is ​​a mathematical tool used in the kinematics of legged robots to represent the instantaneous linear relationship between the velocity of the end effector and the joint velocities. This Jacobian matrix can be constructed under the current robot configuration before the target action is executed. The rows of the matrix correspond to the three spatial directions (X, Y, Z) of the foot in the body coordinate system, and the columns correspond to each active joint. Each element in the Jacobian matrix represents the velocity component of the foot in the corresponding direction when a unit velocity change occurs at a joint. The Jacobian matrix reflects the sensitivity of joint motion to foot motion under the current limb configuration and is a core foundation for calculating differential inverse kinematics.

[0122] The pseudo-inverse of the Jacobian matrix is ​​the one that minimizes the joint motion amplitude (i.e., minimizes the L2 norm) among all possible joint velocity solutions, thus mapping a given plantar velocity (or small displacement) to a set of the smoothest joint velocities (or position changes). By calculating the pseudo-inverse of the Jacobian matrix, the position changes of the plantar surface can be converted into position changes of each joint, achieving efficient and stable inverse kinematics solutions.

[0123] Singularity: refers to a decrease in the rank (i.e. effective degrees of freedom) of the Jacobian matrix, which causes the end effector (such as the sole of the foot) to lose its ability to move in one or more directions, or a state in which a small joint movement causes a large or even infinite movement of the end effector.

[0124] For example, when the legs of a legged robot are fully extended or fully folded, joint movements may only cause the foot to rotate around a point, rather than move in a straight line. In this case, the Jacobian matrix approaches singularity. Near singularities, standard pseudo-inverse calculations can amplify noise or cause abnormally drastic joint velocities / displacements, leading to control instability or actuator saturation.

[0125] Damped least squares is a numerical method used to improve the stability of inverse kinematics solutions, applicable when the Jacobian matrix is ​​close to singular. Its core idea is to introduce a regularization term (also called a damping factor) into the standard pseudo-inverse calculation, smoothing the singular values ​​of the Jacobian matrix and preventing small singular values ​​from being over-amplified.

[0126] Regularization term: An additional constraint term artificially introduced during the mathematical solution process to improve the stability or uniqueness of the solution. The regularization term manifests as a damping factor. square ( The regularization term is incorporated into the formula for calculating the pseudo-inverse of the Jacobian matrix. Its function is to limit the amplitude of joint solutions, preventing severe oscillations or divergence in solutions caused by ill-conditioned Jacobian matrices (such as near-singularities). The magnitude of the regularization term can be dynamically adjusted according to the robot's current configuration: a smaller value is taken when far from singular points to ensure accuracy, and an increased value is taken when approaching singular points to enhance stability. This mechanism significantly improves the robustness of inverse kinematics solutions without significantly sacrificing motion performance.

[0127] Figure 1 A system architecture diagram is shown that can be applied to the reinforcement learning-based motion control policy network training method for a legged robot with a floating base, as described in the embodiments of this disclosure.

[0128] like Figure 1As shown, the system architecture 100 may include a terminal device 101, a legged robot 102, a network 103, and a server 104. The terminal device 101 includes, but is not limited to, desktop computers, laptops, smartphones, and tablets, and is equipped with a graphical user interface for visually displaying the legged robot 102's base posture (Roll / Pitch / Yaw), real-time joint positions, foot contact status, and motion trajectory during posture adjustment. Users can use this interface to set target base postures, intervene in the training process, or evaluate posture control performance.

[0129] The legged robot 102 is a multi-legged robot with a floating base. It is equipped with an inertial measurement unit, joint angle encoder and foot force / position sensor to collect the position of each joint and the spatial position of each foot in the body coordinate system in real time before and after the execution of the action, so as to provide state input and feedback basis for base posture adjustment.

[0130] Server 104 is equipped with a training module for a basal posture control strategy. It receives state data from the legged robot 102, calculates the ideal position changes of each joint based on the changes in foot position before and after the action, and generates joint position reference values ​​by combining the joint state before the action. Then, it constructs a reward term based on the deviation between the actual joint position and this reference value, and uses reinforcement learning algorithms to train a basal posture control model capable of autonomously coordinating joint movements to achieve free adjustment of the basal posture. After training, this basal posture control model can be deployed to the legged robot 102, dynamically adjusting the torso posture in real time in response to terrain changes or task requirements during operation.

[0131] Network 103 provides a communication link between terminal device 101, legged robot 102, and server 104, supporting wired, wireless, or hybrid communication methods to ensure low latency and high reliability for status data upload, policy distribution, and human-machine interaction. It should be understood that... Figure 1 The number and type of each component are for illustrative purposes only; the actual system can be flexibly expanded according to deployment requirements.

[0132] Through the collaborative work of the modules in the above system architecture, this solution realizes a closed loop of posture control from foot displacement perception and joint collaborative reference generation to reinforcement learning-driven posture control. This enables the legged robot to actively and smoothly adjust its base posture in complex scenarios such as uphill, obstacle crossing, and sharp turns. It effectively avoids problems such as body swaying, imbalance and falls caused by stiff joint movements or insufficient coordination, and significantly improves its motion stability, terrain adaptability and high dynamic behavior execution capabilities in unstructured environments.

[0133] This disclosure provides an exemplary implementation of a method for training a motion control policy network for a legged robot with a floating base based on reinforcement learning, referencing... Figure 2As shown, the method may include the following steps S201 to S204:

[0134] Step S201: Obtain the position of each joint and the position of each foot of the legged robot before and after performing the target action;

[0135] Step S202: Calculate the positional change of each joint based on the positional changes of each foot before and after the target action is performed;

[0136] Step S203: Calculate the reference values ​​of the positions of each joint after the execution of the target action, based on the positions of each joint before the target action is executed and the amount of position change of the joints.

[0137] Step S204: Construct a reward term based on the reference value and the position of each joint after the target action is executed, and use the reward term to train the base posture control model.

[0138] The method for training a motion control strategy network for a legged robot with a floating base based on reinforcement learning, as provided in the exemplary embodiments of this disclosure, obtains the positions of each joint and each foot of the legged robot before and after performing a target action. The positional changes of each joint can be calculated by inversely based on the changes in the foot positions. Then, reference values ​​for the positions of each joint are calculated based on the positions of each joint before the target action and the changes in their positions. A reward term is constructed based on the reference values ​​of each joint under ideal conditions and the actual positions of each joint after the target action. This reward term is then used to train the base posture control model. This allows the reinforcement learning strategy to explicitly learn the mapping relationship between base posture adjustment and joint coordinated movement during the training process of the base posture control model. Thus, the trained base posture control model enables the legged robot to generate more kinematically consistent coordinated joint movements when adjusting its base posture. This effectively overcomes the problems of stiff joint movements and delayed trunk posture response caused by only optimizing the end-effector trajectory in related technologies. Therefore, the trained base posture control model enables the legged robot to achieve more stable body posture control in unstructured environments, significantly improving its terrain adaptability, dynamic balance performance, and the reliability of high-mobility movements.

[0139] The following will provide a detailed description of the motion control policy network training method for a legged robot with a floating base based on reinforcement learning in this example embodiment.

[0140] In step S201, the positions of each joint and the position of each foot of the legged robot are obtained before and after performing the target action.

[0141] Here, "before" and "after" refer to two discrete moments within a control cycle: before the action command is issued and after its execution. The state data of the legged robot before the target action is executed can include the positions of each joint and each foot; the state data after the target action is executed can also include the positions of each joint and each foot. These two sets of state data together constitute a training sample, used to subsequently calculate the change in foot position and derive the ideal changes in joint position based on differential inverse kinematics, providing a basis for constructing a reward function with kinematic consistency.

[0142] Understandably, to ensure data comparability under a unified reference system, all positional information is calculated based on the body frame with the legged robot's torso or IMU as the origin, thereby accurately reflecting the geometric relationship between the floating base posture and limb movement.

[0143] Specifically, during reinforcement learning training, whenever the motion control policy network outputs an action command and it is executed by the legged robot, the state data of the legged robot before and after the action execution can be recorded simultaneously: the joint positions before the action execution are recorded as joint_pos_pre and the foot positions are recorded as A1B1_pre; the joint positions after the action execution are recorded as joint_pos_post and the foot positions are recorded as A1B1_post.

[0144] For example, taking a quadruped robot as an example, each leg contains three active joints: hip, knee, and ankle. Therefore, a total of 12-dimensional joint position vectors need to be obtained. At the same time, the three-dimensional coordinates of the four feet in the body coordinate system constitute a 12-dimensional foot position vector. In each training time step, a complete state sampling can be completed, and the above-mentioned before-and-after state data can be cached for subsequent reference joint position generation based on the principle of differential inverse kinematics.

[0145] Understandably, the target action is not a pre-set fixed trajectory, but rather an exploratory or executive action of the reinforcement learning strategy in the current state. Its implicit goal is to drive the legged robot to actively adjust its base posture (including Roll angle, Pitch angle, and Yaw angle). Therefore, the acquired state data before and after execution essentially reflects the actual kinematic response of the joints and feet during the legged robot's attempt to adjust its trunk posture, providing a physical basis and data support for constructing a reward signal with kinematic consistency.

[0146] In actual training, the above data collection can be performed in a simulation environment or on a real robot platform. For example... Figure 3 The diagram illustrates a legged robot performing a target action in some embodiments. The legged robot, with a floating base 301, rotates along the X-axis (Roll angle) to perform the target action. Figure 4 The diagram shown is a schematic of a legged robot performing a target action in some other embodiments. The legged robot with a floating base 401 rotates along the Y-axis (Pitch angle) to perform the target action.

[0147] In some embodiments, in a real robotic motion environment, the positions of each joint and each foot of the legged robot are acquired before and after performing the target action.

[0148] Specifically, the positions of each joint before and after performing the target action can be collected by angle encoders or position sensors installed at each joint of the legged robot; the positions of each foot of the legged robot before and after performing the target action can be obtained by state estimation methods or optical motion capture systems (such as Vicon, OptiTrack, etc.).

[0149] In other embodiments, the positions of each joint and the sole of each foot of the legged robot before and after performing the target action can be obtained in a physical simulation environment.

[0150] Specifically, in a physical simulation environment (such as PyBullet, MuJoCo, or Isaac Gym), a dynamic model consistent with that of a real legged robot is constructed, including parameters such as mass distribution, joint constraints, friction coefficients, and actuation characteristics. During training, whenever the policy network outputs a command for a target action, the simulation engine drives the robot model to execute that action, recording the true positions of all joints and the precise coordinates of the foot in the body coordinate system before and after the action is executed. Because the physical simulation environment can directly access internal state variables without relying on sensor estimation, the data has high accuracy, low noise, and supports large-scale parallel sampling.

[0151] In other embodiments, high-precision encoders and tactile sensors can be pre-set at each joint and foot of the legged robot, and the positions of each joint and foot before and after the execution of the target action can be directly and synchronously collected in conjunction with the real-time kinematic model. This is not limited here.

[0152] Step S202: Calculate the positional change of each joint based on the positional changes of each foot before and after the target action is performed.

[0153] In some embodiments, the Differential Inverse Kinematics method can be used to calculate the positional changes of each joint based on the positional changes of each foot before and after the execution of the target action. This differential inverse kinematics method performs linearization processing near the current legged robot configuration, mapping foot displacements to velocity or displacement increments in joint space, and then integrating or directly solving for the positional changes of each joint.

[0154] In other embodiments, analytical inverse kinematics or numerical optimization methods can also be used to calculate the positional changes of each joint. For example, for a quadruped robot with a regular structure, an analytical inverse kinematics model can be independently established for each leg, and the ideal changes of the corresponding three joints can be directly solved based on the positional changes of the foot end of that leg; for more complex configurations, an optimization problem with the goal of minimizing the displacement error of the foot end can be constructed, and the positional changes of each joint can be obtained by iterative solution.

[0155] In some other embodiments, the calculation results can be corrected by combining the robot's motion feasibility constraints (such as joint limits, speed limits, contact stability, etc.) to ensure that the obtained joint position changes are not only kinematically reasonable, but also feasible in terms of dynamics and actual execution.

[0156] It is understood that in other embodiments, the positional changes of each joint can also be approximated by a pre-trained neural network model, a lookup table method, or a hybrid analytical-learning method, and this disclosure does not limit this.

[0157] Step S203: Calculate the reference values ​​of the positions of each joint after the execution of the target action, based on the positions of each joint before the target action is executed and the amount of position change of the joints.

[0158] Specifically, the positions of each joint before the target action is executed and the changes in joint positions can be superimposed to obtain reference values ​​for the positions of each joint after the target action is executed.

[0159] In some embodiments, the reference value can be calculated using direct vector addition. That is, for each joint, its pre-execution position value is added to the corresponding position change to obtain the reference position of that joint. For example, if a hip joint's pre-execution angle is 0.2 radians and the calculated position change is 0.15 radians, then its reference value is 0.35 radians. This method can calculate the reference value more efficiently.

[0160] In other embodiments, if the joints have physical limitations or motion coupling constraints, the positions of each joint before the target action is executed and the changes in joint positions can be superimposed to obtain candidate values. Then, the candidate values ​​are verified and corrected to obtain reference values.

[0161] For example, after superimposing the positions of each joint before the target action is executed with the changes in joint positions, if the calculated candidate value exceeds the joint travel range, it can be trimmed to the limit boundary, and the reference values ​​of other related joints can be adjusted accordingly to maintain the overall consistency of foot displacement; or through secondary optimization, the deviation from the original reference value can be minimized while satisfying the constraints.

[0162] In some other embodiments, the reference value can also be used as a supervisory signal to assist in the training of the policy network. For example, within the framework of behavior cloning or auxiliary loss, the policy network needs to make its output joint target position as close as possible to the reference value, thereby accelerating convergence and improving the kinematic rationality of the action.

[0163] It is understood that in other embodiments, the generation method of the reference value may also be dynamically adjusted in combination with the robot dynamics model, contact state, or task context. This disclosure does not limit the specific calculation and optimization method of the reference value. For example, when the foot is in a sliding state, the weight of the corresponding leg reference value can be reduced; or when performing a high-dynamic action (such as jumping), a velocity or acceleration term can be introduced to feedforward compensate the reference value.

[0164] Step S204: Construct a reward term based on the reference value and the position of each joint after the target action is executed, and use the reward term to train the base posture control model.

[0165] The reward term can be a quadratic norm. The reward term is calculated based on the error vector between the desired joint position and the actual executed position, and the sum of the squares of the elements of this error vector is taken as the final reward value. This reward value is a non-negative real number; the smaller the value, the smaller the deviation between the robot's actual movement and the desired movement, and the higher the control accuracy.

[0166] Specifically, a reward term for reinforcement learning training can be constructed by comparing the deviation between reference values ​​and the actual joint positions after the target movement is executed. This reward term measures the kinematic rationality of the movement generated by the strategy. For example, the closer the actual joint positions are to the reference values, the more the movement conforms to the kinematic laws deduced from foot displacement, and the smaller the reward value output by the reward term. Conversely, if the deviation is large, it indicates that the movement may have problems such as joint incoordination, energy waste, or inconsistency with the intention of adjusting the base posture, and the reward will be reduced or a penalty will be imposed accordingly.

[0167] In some embodiments, a weighted approach can be used to determine the reward based on the importance of the joints. For example, the hip joint, which has a greater impact on base posture, can be given a higher weight, while the distal ankle joint can be given a lower weight, in order to highlight the joint coordination that plays a key role in overall balance.

[0168] In other embodiments, it can also be used in combination with other auxiliary reward items to construct the reward item, such as a base attitude stability reward (e.g., the smaller the Roll / Pitch angle, the higher the reward) or an energy consumption penalty item (e.g., the smaller the sum of squares of joint velocity or torque, the better), which together constitute a complete reward function.

[0169] Understandably, by using a reinforcement learning framework, the system continuously interacts with the environment, collects state-action-reward samples, and optimizes the policy network parameters based on the reward. This reward serves as a feedback signal, guiding the basic posture control model to gradually reduce the deviation between the actual joint position and the reference value. This ensures that the policy output action not only meets the task objective (such as maintaining balance and adjusting orientation) but also conforms to kinematic laws, avoiding undesirable behaviors such as severe joint shaking, foot slippage, or energy waste.

[0170] In other embodiments, the base attitude control model may also be implemented using a hybrid architecture (such as combining model predictive control with neural networks), which is not limited here.

[0171] In some example implementations, references Figure 5 As shown, the process of calculating the positional change of each joint based on the positional changes of each foot before and after the target action can include the following steps S501 to S502:

[0172] Step S501: For each foot sole, determine the change in position of each foot sole based on the position of each foot sole before and after the target action is performed.

[0173] For each foot, the foot position before and after the target action can be obtained. By comparing the foot positions before and after the target action, the actual spatial displacement of the foot relative to the robot's torso during the action can be determined, i.e., the change in foot position. This change in foot position can include movement information in the forward, backward, left, right, and up / down directions, reflecting the leg's movement intention in this action.

[0174] Step S502: Map the positional changes of each joint based on the positional changes of each foot.

[0175] In some implementations, if the leg structure is a regular structure (such as the leg of a three-degree-of-freedom quadruped robot), the positional changes of each joint can be directly solved using analytical inverse kinematics.

[0176] Specifically, based on the target position of the foot in the hip joint coordinate system, the ideal changes in the hip joint yaw angle, hip joint pitch angle, and knee joint angle are derived sequentially, thereby obtaining the precise changes in joint position.

[0177] In other implementations, when multiple legs are displaced simultaneously, all changes in the position of the soles of the feet can be considered comprehensively. The changes in the position of each joint can be solved in a coordinated manner through a global optimization strategy (such as minimizing the range of motion of the joints or avoiding joint limit conflicts) to improve the smoothness and coordination of the overall movement.

[0178] In some other implementations, the positional changes of each joint are mapped based on the positional changes of each foot, including: mapping the positional changes of each joint based on the positional changes of each foot using inverse kinematics.

[0179] Understandably, since each leg of a legd robot typically consists of multiple cascaded joints (e.g., hip, knee, and ankle joints), there is a defined geometric relationship between foot movement and joint movement. Therefore, inverse kinematics can be used independently for each leg to convert the desired displacement of the foot into the angular adjustment that each joint should produce.

[0180] Based on the current limb configuration of the legged robot, a linear relationship between the small displacement of the foot and the small rotation of the joint can be established in advance. Then, the positional changes of each joint can be mapped according to this linear relationship.

[0181] Specifically, the Jacobian matrix of each joint is obtained; the Jacobian matrix represents the linear relationship between the positional change of the foot and the positional change of the joint; based on the positional change of each foot and the pseudo-inverse matrix of the Jacobian matrix, the positional change of each joint is calculated.

[0182] Specifically, the Jacobian matrix of each leg can be calculated based on the current configuration of the legged robot before the target action is executed. Specifically, the Jacobian matrix for each leg is constructed in real time based on the position of each joint before execution and the known link geometry parameters. Since legged robots typically have multiple independent legs, their local Jacobian matrix can be calculated separately, thus achieving modular processing.

[0183] After obtaining the Jacobian matrix for each leg, the pseudo-inverse matrix of the Jacobian matrix can be further calculated. In some embodiments, singular value decomposition can be used to calculate the pseudo-inverse matrix of the Jacobian matrix.

[0184] Understandably, the pseudo-inverse of the Jacobian matrix serves to deduce a set of optimal joint position changes when a desired change in foot position (i.e., foot displacement) is given, so that the actual foot movement is as close as possible to the desired displacement.

[0185] In some embodiments, the change in foot position can be multiplied by the pseudo-inverse of the Jacobian matrix to obtain the change in position of each joint of the corresponding leg.

[0186] Furthermore, if the Jacobian matrix is ​​not close to a singularity, the positional changes of each joint are calculated based on the positional changes of each foot and the pseudo-inverse of the Jacobian matrix; if the Jacobian matrix is ​​close to a singularity, a damped least squares method is used to add a regularization term to the pseudo-inverse of the Jacobian matrix to obtain an updated pseudo-inverse matrix, and the positional changes of the joints are calculated based on the positional changes of each foot and the updated pseudo-inverse matrix.

[0187] In some embodiments, it can be determined whether the Jacobian matrix is ​​close to a singular point by judging whether the minimum singular value of the Jacobian matrix is ​​lower than a preset threshold (e.g., 0.01). Specifically, if the minimum singular value is greater than or equal to the preset threshold, it is determined that the Jacobian matrix is ​​not close to a singular point; if the minimum singular value is less than the preset threshold, it is determined that the Jacobian matrix is ​​close to a singular point.

[0188] If the Jacobian matrix is ​​not close to a singularity, it indicates that the legged robot is currently in a conventional configuration with good kinematics. The positional changes of each joint can be directly calculated based on the positional changes of each foot and the pseudo-inverse of the Jacobian matrix. This method can accurately meet the movement requirements of the foot and has high computational efficiency.

[0189] If the Jacobian matrix approaches a singularity (e.g., when the leg is nearly fully extended, fully folded, or in other postures that degrade the degrees of freedom), the standard pseudo-inverse calculation is prone to producing excessively large joint adjustments due to matrix ill-conditioning, leading to control instability or actuator saturation. In such cases, damped least squares (DLS) can be used, introducing a regularization term (derived from the damping factor) into the pseudo-inverse calculation of the Jacobian matrix. The original pseudo-inverse matrix is ​​corrected by a control mechanism to obtain an updated pseudo-inverse matrix. This updated pseudo-inverse matrix effectively suppresses drastic fluctuations in joint solutions while maintaining the foot movement trend, ensuring that the output joint position changes are smooth, safe, and executable.

[0190] In some embodiments, the damping factor can be dynamically adjusted according to the degree of singularity. For example, a smaller value is used when far from the singularity to ensure tracking accuracy, while the damping factor is increased when approaching the singularity to enhance stability. This can significantly improve the robustness of the joint solution while ensuring foot motion tracking performance and avoid control failure due to configurational abrupt changes.

[0191] In other embodiments, other numerical stabilization strategies (such as singular value truncation, weighted pseudo-inverse, etc.) can also be used to deal with the ill-conditioned problem of the Jacobian matrix, which is not limited here.

[0192] Through the above mechanism, the foot tracking accuracy and joint motion stability can be automatically balanced under different motion states, ensuring that the inverse kinematic mapping results always have good physical rationality and control feasibility, providing reliable support for subsequent generation of reference values ​​and construction of kinematic consistency rewards.

[0193] In other embodiments, the positional changes of each joint are mapped based on the positional changes of each foot, including: decomposing the positional changes of each foot into positional change components in each direction under the coordinate system; mapping the positional change components of each foot to positional change components of each joint; and constructing the positional changes of the joints from the positional change components of each joint.

[0194] For the positional change of each foot, the position of the foot for each leg before and after the target action is obtained, and the positional change of that foot is calculated. This positional change is a three-dimensional vector representing the displacement of the foot from its initial position to its final position in the body coordinate system. This three-dimensional positional change can be decomposed along the three orthogonal axes of the body coordinate system (X-axis for forward / backward direction, Y-axis for left / right direction, and Z-axis for up / down direction) to obtain three independent positional change components, each corresponding to the positional change of the foot in its respective direction.

[0195] For each position change component in each direction, based on the kinematic structure of the leg of the legged robot, it can be mapped to the joint position change component contributed by the movement of the corresponding joint in that direction.

[0196] For example, changes in the foot's position along the Z-axis (height direction) are primarily achieved through the coordinated movement of the knee and hip flexion / extension joints, while changes along the X-axis (anteroposterior direction) may involve adjustments in both the hip flexion / extension joint and the ankle joint. This mapping process can be based on analytical kinematic models, empirical proportional allocation rules, or pre-calibrated local Jacobian relationships, ensuring that foot movements in different directions are appropriately allocated to the corresponding joint degrees of freedom.

[0197] Therefore, by superimposing the position change components of the same joint in different directions (e.g., through algebraic summation), the complete position change of that joint can be constructed. After integrating the position change components of all joints, a complete joint position change vector is formed, which is used for generating reference values ​​and constructing reward items in subsequent steps.

[0198] Understandably, this decomposition-mapping-reconstruction approach helps to explicitly decouple the contributions of multidimensional foot movements to each joint, improving the interpretability and controllability of the mapping process. Furthermore, while ensuring kinematic consistency, it can also reduce computational complexity, facilitating real-time deployment on resource-constrained embedded platforms.

[0199] In some embodiments, such as Figure 6 The diagram shows a legged robot. In the coordinate system (Body Frame) where the floating base 601 of the legged robot is located, O1 represents the origin, A1 and B1 are two key nodes of the leg of the legged robot, and A1B1 is the pose vector of the foot (end effector).

[0200] In the reinforcement learning training process, in order to evaluate the accuracy of action execution and construct reward items, the following steps need to be performed in the footed robot's base coordinate system:

[0201] Record the state data of the legged robot before it performs the target action, including the position information of each joint 602. and the location information of each foot sole 603 And record the state data of the legged robot after performing the target action, including the position information of each joint 602. and the location information of each foot sole 603 .

[0202] The positional change of the 603 point on the sole of the foot can be calculated using the following formula (1).

[0203] (1)

[0204] in, It indicates the amount of change in the position of the sole of the foot.

[0205] So, based on the change in the position of the foot It can map the positional changes of each joint. It also calculates the reference values ​​of the positions of each joint after the target action is executed, which are the expected changes in the joints.

[0206] The reward item is constructed using the following formula (2):

[0207] (2)

[0208] in, This is the reward value of the reward item. This reward value represents the magnitude of the deviation between the actual movement and the expected movement of the joint. The smaller the value, the higher the joint tracking accuracy and the better the control effect.

[0209] Through the above process, a quantitative assessment of the precision of joint motion execution is achieved and embedded into the reward function, enabling the reinforcement learning model to explicitly learn a high-precision, low-error coordinated motion strategy during training, thereby improving the robot's motion stability and control performance in complex environments.

[0210] In some example implementations, references Figure 7As shown, the process of constructing a reward item based on reference values ​​and the positions of each joint after the target action is performed may include the following steps S701 to S702:

[0211] Step S701: For each joint, determine the positional difference between the reference value of the joint's position and the actual value of the joint's position after the target action is performed.

[0212] In some implementations, the algebraic difference between the reference value and the actual value can be directly calculated as the positional difference. For example, for a linear translational joint, if the reference value is 0.15 meters and the actual value is 0.13 meters, then the positional difference is 0.02 meters.

[0213] In other implementations, for periodic rotary joints (such as hip yaw angle), the reference and actual values ​​can be normalized to the same angular range, and then the minimum difference in the orbital angle between the two can be calculated as the position difference. For example, when the reference value is 3.1 radians and the actual value is −3.1 radians, the difference is directly subtracted to get 6.2 radians, but considering the periodicity of the angle, the actual minimum difference is about 0.08 radians.

[0214] In some other implementations, for each joint, the position of the joint at each moment during the execution of the target action is obtained; based on the position change of the joint between adjacent moments during the execution of the target action, the rotational smoothing gap of the joint is determined; the rotational smoothing gap is positively correlated with the position change, and the rotational smoothing gap belongs to at least one other gap.

[0215] Rotational smoothness gap is a deviation index used to quantify the smoothness of motion of a legged robot joint during the execution of a target action. Its value is positively correlated with the amount of position change of the joint between adjacent control moments. Specifically, the rotational smoothness gap is constructed by analyzing the difference between consecutive frames in the joint position time series. The larger the difference, the more violent and discontinuous the joint movement, and the larger the corresponding rotational smoothness gap; conversely, if the joint position change is small and uniform, the rotational smoothness gap is small, indicating that the movement is more stable and smooth.

[0216] Understandably, rotational smoothing gap, as a regularization signal in the reinforcement learning reward function, is used to suppress the policy from generating high-frequency jitter, abrupt changes, or high-acceleration joint commands, thereby improving the naturalness of the movement, energy efficiency, and hardware safety. For example, in walking or jumping tasks, an excessively large rotational smoothing gap may mean that the leg joints will make unnecessary sudden stops or rebounds during the stance phase, which can easily cause the machine to wobble or even become unstable; by penalizing such gaps in the reward, the policy can be effectively guided to learn smooth and coordinated movement patterns.

[0217] Specifically, during the execution of the target action, position data of each joint is continuously collected in control cycles (e.g., every 10 milliseconds) to form a complete time series; this time series fully records the entire motion trajectory of the joint from the start to the end of the action.

[0218] Therefore, we can iterate through the time series and calculate the difference in joint position between any two adjacent moments. The cumulative measure of all adjacent position changes during the entire movement is used as the rotational smoothing difference. This difference reflects the instantaneous range of motion of the joint within the control step. Since the more intense the joint movement, the greater the change between adjacent positions, this difference directly reflects the local discontinuity or dynamic impact of the movement.

[0219] The rotational smoothness gap is positively correlated with the amount of position change at each moment. For example, the greater and more frequent the position change, the larger the rotational smoothness gap; conversely, if the joint movement is smooth and the change is uniform, the gap is smaller. This rotational smoothness gap can be used to evaluate the smoothness and coordination of the motion execution process.

[0220] For example, when a quadruped robot performs a "trotting" motion, if a hip joint experiences high-frequency, micro-amplitude tremors during the support phase due to unstable control, the positional differences between adjacent moments, though small, will accumulate and lead to a significant rotational smoothness gap. In ideal smooth motion, the positional changes of the same joint should be continuous, monotonous, or gradually changing, resulting in a significantly lower rotational smoothness gap. By incorporating this gap into the reward, reinforcement learning strategies will proactively avoid such unnecessary high-frequency movements during training, thereby improving motion efficiency, reducing energy consumption, and extending actuator lifespan.

[0221] It should be noted that the calculation of the rotational smoothing difference can also be carried out in a variety of accumulation methods, such as summing the absolute difference, summing the squares of the difference, or combining time weighting for integral approximation. The specific form can be flexibly selected according to the control frequency, task characteristics and optimization objectives, and is not limited here.

[0222] Step S702: Construct a reward item based on the norm of the positional gap.

[0223] In some implementations, the positional differences of all joints can be combined into an overall deviation vector, and the Euclidean norm (L2 norm) of this vector can be calculated as an indicator of the overall joint motion deviation. This indicator is then used as a reward. That is, the smaller the overall joint motion deviation, the smaller the reward; conversely, the larger the overall joint motion deviation, the larger the reward. This approach is simple in structure and has smooth gradients, effectively guiding the policy network to gradually reduce the overall difference between the actual action and the ideal kinematic solution, thereby learning to generate coordinated and reasonable joint motion sequences.

[0224] In other implementations, different weights can be assigned to the positional differences of different joints, and a comprehensive deviation norm can be obtained by weighted calculation based on the positional differences of each joint to construct the reward item.

[0225] For example, higher weights are assigned to the hip or knee joints, which directly affect trunk posture stability; while lower weights are assigned to the ankle joints or distal degrees of freedom, which have a smaller impact on overall balance.

[0226] In this way, the reward can be more focused on the motion accuracy of key joints, allowing the strategy to prioritize the optimization of degrees of freedom that play a dominant role in base posture control during training, thereby improving motion stability and task adaptability.

[0227] In addition, other implementation methods can be used to construct the reward term, such as nonlinear mapping of the deviation norm, for example, using a saturation function or exponential decay form, so as to provide high-sensitivity feedback when the deviation is small and avoid drastic fluctuations in the reward signal when the deviation is large, thereby enhancing the convergence stability in the early stage of training. No limitation is made here.

[0228] Furthermore, the above method also includes: for each joint, determining at least one other gap generated by the joint during the execution of the target action; the other gap is the gap corresponding to factors other than the position gap; constructing a reward term based on the norm of the position gap, including: constructing a reward term based on the norm of the position gap and the norm of at least one other gap.

[0229] Other gaps refer to deviation indicators other than positional gaps, used to measure the quality of movement, dynamic characteristics, or rationality of movement of legged robots during the execution of target actions.

[0230] Specifically, other discrepancies can be reflected in the drastic changes in joint position between consecutive control cycles, i.e., rotational smoothness discrepancies. For example, if a joint's angle changes abruptly between adjacent moments, it indicates jitter or shock in its movement, resulting in a larger rotational smoothness discrepancy; conversely, if the angle change is gradual, the discrepancy is smaller. This discrepancy can be obtained by calculating the absolute or squared value of the difference in joint position between adjacent moments and accumulating it over time, thus encouraging the strategy to generate continuous, low-frequency joint trajectories.

[0231] In addition, other discrepancies may include the deviation between joint velocity and ideal velocity. For example, in some tasks, it is expected that the joint will tend to come to rest at the end of the movement. If the actual velocity is still large, the velocity discrepancy is significant, reflecting insufficient control convergence. Similarly, acceleration can be obtained by differentiating the velocity again, and then the acceleration discrepancy can be constructed to measure the impact level during movement—high acceleration often means high inertial forces and foot impact risk, which is detrimental to balance maintenance and hardware safety.

[0232] In other embodiments, other discrepancies may extend to energy consumption dimensions, such as the deviation between the actual output torque and the theoretical minimum torque, or the degree of mismatch between the foot contact state and the expected support / swing timing (e.g., the supporting leg is unexpectedly lifted).

[0233] In some implementations, the norm of the positional difference can be weighted and summed with the norm of at least one other difference to obtain the reward term. For example, the L2 norm of all joint positional differences and the L2 norm of rotational smoothing differences can be calculated separately, and then multiplied by their respective positive weight coefficients (e.g., position weight is 1.0 and smoothing weight is 0.1) to obtain the final reward value.

[0234] In other implementations, a hierarchical or conditional fusion strategy can be used to construct the reward term. When the norm of the positional gap is below a preset threshold (indicating that the action is basically in place), the norms of other gaps are included in the reward calculation; when the norm of the positional gap is higher than or equal to the preset threshold, the reward term is constructed primarily based on the positional gap. For example, in the early stages of training, before the policy has learned the basic posture, positional accuracy is optimized first; once the policy can stably reach the target configuration, rotational smoothing gaps are introduced for fine-tuning. This approach can avoid training oscillations caused by multi-objective conflicts, improving convergence efficiency and final performance.

[0235] Other implementation methods can also be used to construct reward items, which are not limited here.

[0236] In some example implementations, references Figure 8 As shown, the process of training the basis posture control model using the reward term may include the following steps S801 to S802:

[0237] Step S801: Obtain the reward value output by the reward item.

[0238] The reward value is a scalar value output by the constructed reward term after each target action is executed by the policy. It is used to quantify the quality of the action generated by the current basis posture control model. This reward value serves as the core feedback signal of the reinforcement learning algorithm, directly guiding the direction and magnitude of model parameter updates.

[0239] Specifically, the reward value is calculated based on the norm of one or more kinematically related gaps (such as joint position gaps, rotational smoothness gaps, etc.), and can be configured as a positive correlation function of one or more gaps—that is, the smaller the gap, the smaller the reward value; the larger the gap, the larger the reward value.

[0240] For example, when the actual positions of each joint are very close to the reference values ​​and the movement process is smooth without abrupt changes, the difference norms are small, and the reward value will approach zero or be a small negative number, indicating that the movement quality is excellent. Conversely, if there is a significant positional deviation or violent shaking, the difference norms will increase and the reward value will decrease significantly, indicating that the movement is unreasonable.

[0241] It should be noted that the reward value is an evaluation metric oriented towards minimizing the loss. During training, the reinforcement learning algorithm continuously tries different control strategies to seek policy parameters that minimize the reward value (i.e., get closer to the ideal state).

[0242] Furthermore, the construction of the reward value integrates considerations of both endpoint accuracy and process quality, distinguishing it from sparse rewards that rely solely on task success or failure (such as whether a fall occurs). It provides dense, immediate, and kinematically meaningful feedback, significantly improving the learning efficiency and convergence stability of the policy in a high-dimensional continuous action space. Finally, when the reward value consistently falls below a preset first reward threshold, the model is considered to have learned to generate a basis posture control policy that meets the accuracy and smoothness requirements, and the training process can be terminated.

[0243] Step S802: Perform reinforcement learning training on the base attitude control model based on the reward value until the reward value output by the reward item is less than the first reward threshold, and obtain the trained base attitude control model.

[0244] The first reward threshold is a pre-set scalar value used to characterize the upper limit of acceptable motion quality in the basic attitude control task.

[0245] The first reward threshold is the upper limit of the reward value. For example, if the reward value remains below the first reward threshold during training, it indicates that the actions generated by the model have met the preset requirements in terms of endpoint accuracy and process smoothness, and further training is unnecessary. This threshold can be calibrated according to the robot's structural characteristics, task complexity, and actual control accuracy requirements, for example, by setting specific values ​​such as 0.1 or 0.05 through simulation trial and error or expert experience.

[0246] This model can autonomously output coordinated, stable, and kinematically sound joint control commands based on the robot's current state and target action. In practical deployment, it can effectively maintain or adjust the posture of the floating base, avoiding risks such as robot swaying, excessive energy consumption, or instability caused by unreasonable joint movements. This model can be directly used in subsequent whole-machine motion control modules or embedded as a sub-policy in higher-level hierarchical reinforcement learning architectures.

[0247] In some implementations, after calculating the reward value, reinforcement learning algorithms (such as PPO, SAC, etc.) can update the parameters of the basis attitude control model based on the reward value and its related trajectory data, and generate a new basis attitude control model.

[0248] This process is repeated cyclically, allowing continuous monitoring of the reward value's changing trend after each training iteration. When the average reward value for several consecutive rounds (e.g., 10 rounds) is less than the first reward threshold, or when the single reward value is consistently lower than the threshold without significant fluctuations, the training is determined to have converged, parameter updates are stopped, and the current model is saved as the trained base pose control model.

[0249] In other implementations, the baseline posture control model is trained using reinforcement learning based on reward values ​​until the reward value output by the reward item is less than a first reward threshold, thereby obtaining a trained baseline posture control model. This includes: calculating the training time and obtaining the learning factor corresponding to the training time; the learning factor characterizes the difficulty of reinforcement learning training of the baseline posture control model, and the learning factor is positively correlated with the training time; and the baseline posture control model is trained using reinforcement learning based on the reward values ​​and the learning factor until the reward value output by the reward item is less than the first reward threshold, thereby obtaining a trained baseline posture control model.

[0250] Training duration refers to the cumulative training progress of the base posture control model from the start to the current moment during reinforcement learning training, used to quantify the amount of training the model has received. Specifically, it can be expressed as the total number of steps in environmental interaction, the number of completed training episodes, or the actual running time, reflecting the scale of experience and optimization depth accumulated by the policy network.

[0251] The learning factor is a dynamic parameter positively correlated with training duration, used to characterize the intensity of the performance requirements of the policy or the difficulty of task optimization in the current training phase. As training duration increases, the learning factor gradually increases, indicating that the model should possess higher action accuracy and smoothness, and a lower tolerance for the same motion bias. The learning factor can be implemented in the form of linear, logarithmic, or piecewise functions, used to dynamically adjust the sensitivity of the reward signal or the weight of the loss function, thereby guiding the policy to focus on exploration in the early stages of training and on fine-tuning in the later stages, improving overall convergence efficiency and control performance.

[0252] The corresponding learning factor can be queried or calculated based on the training duration. In some embodiments, the learning factor may adopt a linear growth function. For example: Learning factor = α × training duration, where α is a preset growth coefficient. In other embodiments, the learning factor may adopt a piecewise stepwise or logarithmic growth form to adapt to the convergence characteristics of different training stages. For example, the learning factor may be kept at a low value in the first 1000 training rounds to encourage exploration, and then gradually increased to strengthen the requirements for accuracy and smoothness. In other embodiments, other methods may be used for calculation, which are not limited here.

[0253] The learning factor, combined with the original reward value, is used to adjust the objective or loss function of reinforcement learning. For example, the effective reward value can be defined as the original reward value multiplied by the learning factor, or the learning factor can be introduced as an importance weight in the policy gradient calculation, so that a stronger penalty is imposed on the same action bias in the later stages of training. Alternatively, the learning factor can be used to dynamically adjust the effective criterion of the first reward threshold, realizing an adaptive convergence mechanism that tightens performance requirements as training progresses.

[0254] Understandably, by introducing a learning factor, the reinforcement learning process can dynamically adjust the optimization intensity according to the training progress, avoiding training stagnation due to the inability to meet the fixed threshold in the early stages, and preventing convergence to a suboptimal strategy due to excessively low standards in the later stages. Finally, when the reward value modulated by the learning factor (or the original reward value guided by the learning factor) remains below the first reward threshold, it can be determined that the basis posture control model has been sufficiently learned, and the trained basis posture control model is output.

[0255] In other implementations, the base attitude control model may include a motion control policy network; the base attitude control model is trained by reinforcement learning based on the reward value until the reward value output by the reward term is less than a first reward threshold, thereby obtaining a trained base attitude control model, including: adjusting the parameters of the motion control policy network in the base attitude control model based on the reward value to obtain a new base attitude control model; and performing reinforcement learning on the new base attitude control model until the reward value output by the reward term is less than the first reward threshold, thereby obtaining a trained base attitude control model.

[0256] The motion control strategy network is a parameterized function approximator (typically a deep neural network) that takes the current robot state (such as joint position, velocity, base pose, and target action) as input and outputs target action commands for each joint (such as target position, torque, or velocity). This motion control strategy network can continuously optimize its internal parameters by interacting with the environment and receiving reward signals to learn and generate stable and efficient base pose control strategies that conform to kinematic laws. This motion control strategy network can employ a deep neural network structure (such as a multilayer perceptron, graph neural network, or temporal network), and its parameters are continuously updated during training to gradually improve motion quality.

[0257] Specifically, if the reward value is greater than or equal to the first reward threshold, the training weight corresponding to the reward value is obtained; the training weight represents the degree of reinforcement learning training of the base posture control model, and the training weight and the reward value are positively correlated; the base posture control model is trained by reinforcement learning based on the reward value and the training weight until the reward value output by the reward item is less than the first reward threshold, and the trained base posture control model is obtained.

[0258] Among them, the training weight is a dynamically adjusted scalar parameter used to reflect the degree of gap between the current policy performance and the ideal state, and indirectly characterizes the current reinforcement learning training stage or optimization difficulty of the model.

[0259] The training weights are positively correlated with the reward value; that is, the larger the reward value (the more severe the bias), the higher the training weight; the closer the reward value is to the first reward threshold (the smaller the bias), the lower the training weight. This reflects an adaptive training mechanism where the greater the bias, the higher the learning intensity.

[0260] Training weights are used to adjust the intensity of parameter updates during reinforcement learning. For example, in policy gradient calculation or loss function construction, the gradient or loss term is multiplied by the training weight, thereby applying a stronger correction signal when the bias is large, accelerating convergence; and reducing the update magnitude when the bias is small, avoiding overfitting or oscillation. By combining the reward value with the corresponding training weight for model updates, the system can dynamically adjust the learning intensity, balancing training efficiency and stability.

[0261] The reward term quantifies the degree of deviation in robot motion control. Its output reward value is a non-negative real number. The smaller the reward value, the closer the joint movement is to the desired trajectory, and the better the control performance. The first reward threshold is a preset convergence criterion (e.g., 0.1). When the reward value is lower than this threshold, the model training is considered complete.

[0262] During training, if the current reward value is greater than or equal to the first reward threshold, it indicates that the basic posture control model has not yet reached the expected control accuracy and needs further optimization. The corresponding training weights can be determined based on the reward value, and reinforcement learning training can be performed again. Repeat the above process to continuously train the basic posture control model using reinforcement learning until the reward value output by the reward item is less than the first reward threshold. At this point, it is determined that the model has fully converged, and the trained basic posture control model is output, which can be used for real-time basic posture control tasks of legged robots.

[0263] In some embodiments, training weights can also be used as scaling factors for policy gradients or loss functions to weight the parameter update step size. For example, when the training weights are greater than a preset weight threshold, the gradient update magnitude is amplified to accelerate the correction of severely biased actions; when the training weights are less than or equal to the preset weight threshold, the update step size is reduced to avoid oscillations caused by over-adjustment near convergence. In another implementation, training weights can also be used to adjust the sampling priority of samples in experience replay, so that trajectories with large deviations are replayed more frequently during training, thereby improving learning efficiency.

[0264] In some embodiments, if the reward value is less than a second reward threshold, a first training weight corresponding to the reward value is determined, and reinforcement learning training is performed on the base posture control model based on the reward value and the first training weight until the reward value output by the reward item is less than the first reward threshold, thus obtaining a trained base posture control model; the second reward threshold is greater than the first reward threshold; if the reward value is greater than or equal to the second reward threshold, a second training weight corresponding to the reward value is determined, and reinforcement learning training is performed on the base posture control model based on the reward value and the second training weight until the reward value output by the reward item is less than the first reward threshold, thus obtaining a trained base posture control model; the second training weight is greater than the first training weight.

[0265] The first reward threshold is the convergence criterion for model training completion (e.g., 0.1), and the second reward threshold is an intermediate performance threshold between the initial training phase and the convergence phase, and satisfies that the second reward threshold is greater than the first reward threshold (e.g., the first reward threshold is 0.1 and the second reward threshold is 0.5).

[0266] When the reward value is less than the second reward threshold, i.e. the deviation is small, it indicates that the base attitude control model has entered the fine-tuning stage, and a smaller first training weight is used. At this time, the update step size of the base attitude control model is small, which helps to maintain training stability when approaching convergence and avoid performance fluctuations caused by over-adjustment;

[0267] When the reward value is greater than or equal to the second reward threshold, i.e. the deviation is large, it indicates that the base attitude control model is still in the coarse-tuning stage. In this case, a larger second training weight is adopted, and the second training weight is greater than the first training weight.

[0268] In some embodiments, training weights can be used as scaling factors for the policy loss function or policy gradient. For example, in algorithms such as PPO and SAC, the advantage function or TD error is multiplied by the corresponding training weights before parameter updates. They can also be used to calculate the priority of samples in empirical replay, so that high-bias trajectories can be replayed more frequently.

[0269] Understandably, the aforementioned dual-threshold, dual-weight mechanism enables phased adaptive learning intensity adjustment: "relearning and rapid adjustment" when the deviation is large in the early stages of training, and "light adjustment and stable convergence" when the deviation is small in the later stages of training, effectively balancing convergence speed and final control accuracy. This training process is continuously executed until the reward value is less than the first reward threshold, at which point it can be determined that the basal attitude control model has met the performance requirements, and the trained basal attitude control model is output for high-precision basal attitude control of legged robots.

[0270] In some example implementations, after calculating the positional change of each joint based on the positional changes of each foot before and after the execution of the target action, the following steps may be included: for the positional change of each joint, obtain the physical constraint range corresponding to the joint; if the positional change of the joint is within the corresponding physical constraint range, then perform the step of calculating the reference value of the position of each joint after the execution of the target action based on the position of each joint before the execution of the target action and the positional change of the joint; if the positional change of the joint exceeds the corresponding physical constraint range, then update the positional change of the joint based on a preset joint executable strategy to obtain the updated positional change, and perform the step of calculating the reference value of the position of each joint after the execution of the target action based on the position of each joint before the execution of the target action and the positional change of the joint with the updated positional change.

[0271] The physical constraint range refers to the allowable range of position, velocity, acceleration, or position change that each joint of a legged robot can safely and reliably execute under the constraints of its actual physical structure and driving capabilities. This physical constraint range is jointly determined by the robot's mechanical design, actuator performance, and safe operation requirements. It is used to ensure that control commands are executable on the actual hardware and to avoid execution failures, performance degradation, or hardware damage caused by commands exceeding limits.

[0272] In some embodiments, the physical constraint range may include at least one of the joint angle range and the motor torque range.

[0273] The joint angle range refers to the interval between the minimum and maximum angles that a legged robot's joints can safely rotate within the limits allowed by its mechanical structure. This range can be determined by the joint's mechanical limits (such as hard limit blocks or flexible stops) or safe operating specifications. The motor torque range refers to the interval between the maximum positive torque and the minimum (or maximum negative) torque that the motor driving the joint movement can continuously or instantaneously output without overload, overheating, or triggering protection mechanisms.

[0274] In other embodiments, the physical constraint range may also include velocity constraint range, acceleration constraint range, and position change constraint range, etc., and is not limited thereto.

[0275] Joint executable strategies refer to a set of preset rules or algorithms used to correct or remap the change in position of a joint when it exceeds the physical constraints of the joint during the motion control of a legged robot. The goal is to generate an executable action that is as close as possible to the original command within the robot's physical capabilities, so as to ensure the safety, feasibility and motion coordination of the control command.

[0276] Specifically, if the change in joint position is within the corresponding physical constraint range, the change is considered physically executable. The change in position is directly used, and combined with the actual position of each joint before the target action is executed, the reference value of each joint position after the target action is executed is calculated. If the change in joint position exceeds the corresponding physical constraint range, it means that the change cannot be actually executed by the robot. At this time, the preset joint executable strategy is called to correct or update the change in position, resulting in an updated change in position that conforms to the physical constraints. The updated change in position is used to replace the original change in position, and combined with the joint position before the action is executed, the reference value of each joint position after the target action is executed is calculated.

[0277] In some embodiments, a joint-executable strategy may include at least one of the following: truncating the excess variation to an allowable boundary, scaling the variation of multiple joints as a whole to maintain limb coordination, or finding an alternative solution that is closest to the original instruction within the feasible domain based on an optimization method.

[0278] Through the above mechanism, it is possible to automatically filter or correct motion instructions that exceed the robot's physical capabilities during reinforcement learning training and control execution, ensuring that all generated reference trajectories are within the hardware's executable range, thereby effectively improving the practicality, safety, and deployment success rate of the strategy on real robots.

[0279] This disclosure also provides an exemplary implementation of a motion control method for a legged robot with a floating base based on reinforcement learning, referencing... Figure 9As shown, the method may include the following steps S901 to S902:

[0280] Step S901: Obtain the current motion state information of the legged robot.

[0281] Motion state information refers to multidimensional state data used to characterize the overall motion of a legged robot at a given moment. Motion state information can include the position and orientation of the legged robot's base, the linear and angular velocities of the base, the angles and angular velocities of each joint, the position, velocity, and contact state of each foot with the ground, as well as the acceleration and angular velocity output by the inertial measurement unit (IMU).

[0282] Specifically, the current motion state information can be obtained in real time through the body sensors or state estimation algorithms of the legged robot, and used as input to the base posture control model to determine the current motion situation and support subsequent action decisions.

[0283] Step S902: Output the motion strategy for controlling the legged robot's motion based on the current motion state information and the trained basal attitude control model; wherein, the trained basal attitude control model is obtained according to the above-mentioned reinforcement learning-based motion control strategy network training method for a legged robot with a floating basal.

[0284] The motion strategy refers to the control commands generated by the trained base posture control model based on the current motion state information, used to drive the robot to execute the next movement. The motion strategy can be expressed as the desired target for each joint, including but not limited to the target position, target velocity, or target torque of each joint, or it can be expressed as foot trajectory planning parameters or high-level gait commands. This motion strategy directly determines the robot's motion behavior and is the core output of the reinforcement learning model during the deployment phase. Its rationality and executability directly affect the stability, accuracy, and adaptability of the robot's motion.

[0285] Understandably, the trained baseline posture control model can be deployed in this legged robot.

[0286] In some implementations, the current motion state information (including base pose, joint angles, velocity, foot contact state, etc.) can be directly input into a trained base pose control model (such as a deep neural network), which then outputs the motion strategy directly in an end-to-end manner. This motion strategy may include low-level control commands such as target position, target velocity, or target torque for each joint.

[0287] In other implementations, current motion state information (including base pose, joint angles, velocity, foot contact state, etc.) can be input into a trained base pose control model. The trained base pose control model outputs a high-level motion intention based on the current motion state information. Combining this high-level motion intention with the robot's kinematics model (such as inverse kinematics or a whole-body dynamics optimizer), the high-level intention is transformed into a motion strategy. This motion strategy may include specific position or torque commands for each joint.

[0288] In robot control, high-level motion intention refers to the abstract, macroscopic motion target or command output by the policy model, which expresses how the robot as a whole should move. For example, high-level motion intention can be the expected base acceleration, the offset of the center of mass trajectory, or the foot swing target.

[0289] In other implementations, the motion strategy can be obtained in other ways, such as inputting the current motion state information into the trained base posture control model, which outputs intermediate representations such as foot trajectory parameters, gait timing, or center of mass motion target, and combines them with a real-time motion planner or a whole-body optimization controller to generate the final joint-level motion strategy, which is not limited here.

[0290] refer to Figure 10 The diagram illustrates the principle of a motion control method for a legged robot with a floating base based on reinforcement learning. Figure 10 In the process, after the trained base posture control model 1003 is deployed to the legged robot with a floating base 1001, the current motion state information 1002 of the legged robot is obtained, and the current motion state information 1002 is input into the trained base posture control model 1003. The action strategy 1004 for controlling the movement of the legged robot is output. The action strategy 1004 can instruct the legged robot to perform actions according to the instructions of the action strategy.

[0291] The motion control method for a legged robot with a floating base based on reinforcement learning, as provided in the exemplary embodiments of this disclosure, acquires the current motion state information of the legged robot and outputs a motion strategy to control the movement of the legged robot based on this information and a trained base posture control model. This trained base posture control model is obtained according to the aforementioned reinforcement learning-based motion control strategy network training method for a legged robot with a floating base. That is, during training, the model explicitly learns the intrinsic mapping relationship between base posture adjustment and multi-joint coordinated movement through a reward term constructed based on joint position reference values ​​and actual execution positions. Therefore, in actual motion control, this trained base posture control model can generate coordinated joint movements in real time based on the current state, which both conform to kinematic constraints and take into account trunk posture stability. This effectively overcomes the problems of stiff joint movements, delayed base posture response, and severe trunk swaying caused by relying solely on preset gait or end-effector trajectory tracking in related technologies. As a result, legged robots can autonomously maintain a stable floating base posture when performing actions such as walking, turning, and crossing, which significantly improves their dynamic balance ability, terrain adaptability, and reliability of high-mobility actions in unstructured environments.

[0292] It is understood that in steps S901 to S902 above, the legged robot trains the base posture control model based on steps S201 to S204 in another example embodiment.

[0293] In an exemplary embodiment of this disclosure, another reinforcement learning-based motion control policy network training device for a legged robot with a floating base is also provided, which is applied to a control terminal providing a graphical user interface. (Reference) Figure 11 As shown, the motion control policy network training device 1100 for a legged robot with a floating base based on reinforcement learning includes a training position acquisition module 1101, a joint position change calculation module 1102, a joint position reference value calculation module 1103, and a model training module 1104, wherein:

[0294] The training position acquisition module 1101 is used to acquire the position of each joint and the position of each foot of the legged robot before and after performing the target action;

[0295] The joint position change calculation module 1102 is used to calculate the position change of each joint based on the position changes of each foot before and after the target action is performed.

[0296] The reference value calculation module 1103 for joint positions is used to calculate the reference values ​​of the positions of each joint after the execution of the target action based on the positions of each joint before the target action is executed and the amount of change in the joint positions.

[0297] The model training module 1104 is used to construct a reward term based on the reference value and the position of each joint after the target action is executed, and to train the base posture control model using the reward term.

[0298] The specific details of each module in the above-mentioned reinforcement learning-based training device for the motion control strategy network of a legged robot with a floating base have been described in detail in the corresponding reinforcement learning-based training method for the motion control strategy network of a legged robot with a floating base, so they will not be repeated here.

[0299] In an exemplary embodiment of this disclosure, another motion control device for a legged robot with a floating base based on reinforcement learning is also provided, which is applied to a control terminal providing a graphical user interface. (Reference) Figure 12 As shown, the motion control device 1200 for a legged robot with a floating base based on reinforcement learning includes a running state information acquisition module 1201 and a motion strategy output module 1202, wherein:

[0300] The running status information acquisition module 1201 is used to acquire the current motion status information of the legged robot;

[0301] The motion strategy output module 1202 is used to output a motion strategy for controlling the movement of the legged robot based on the current motion state information and the trained basal attitude control model; wherein, the trained basal attitude control model is obtained according to the above-mentioned reinforcement learning-based motion control strategy network training method for a legged robot with a floating basal.

[0302] The specific details of each module in the motion control device for the aforementioned legged robot with a floating base based on reinforcement learning have been described in detail in the corresponding motion control method for the same legged robot based on reinforcement learning with a floating base, so they will not be repeated here.

[0303] In an exemplary embodiment of this disclosure, a legged robot is also provided. The legged robot includes a processor and a memory, the memory storing computer-readable instructions. When executed by the processor, the computer-readable instructions implement the aforementioned method. The legged robot includes any one of quadrupedal robots, bipedal robots, humanoid robots, and mobile robots. (Reference) Figure 13 A schematic diagram of a legged robot with a floating base 1301 is shown. (Reference) Figure 14 This diagram illustrates another legged robot with a floating base 1401. (Reference) Figure 15 This shows a schematic diagram of yet another legged robot with a floating base 1501.

[0304] refer to Figure 16As shown, an electronic device capable of implementing the above method is also provided. The electronic device 1600 includes a processor 1601 and a memory 1602. The memory 1602 stores computer-readable instructions, which, when executed by the processor 1601, implement the method of this disclosure.

[0305] In an exemplary embodiment of this disclosure, a computer-readable storage medium is also provided, having stored thereon computer program code instructions that, when invoked by a robot's processor, cause the robot to perform the method described in the embodiments.

[0306] refer to Figure 17 As shown, a program product 1700 for implementing the above-described method according to an embodiment of the present disclosure is described. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0307] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0308] Finally, the above preferred embodiments are only used to illustrate the technical solutions of this application and are not restrictive. Although this application has been described in detail, those skilled in the art should understand that changes in form and detail can be made without departing from the scope defined by the claims of this application. The dimensions in the drawings are not related to the specific physical object, and the physical object dimensions can be arbitrarily changed.

Claims

1. A method for training a motion control policy network of a legged robot with a floating base based on reinforcement learning, characterized in that, The method comprises: obtaining positions of each joint and positions of each foot sole of the foot robot before and after performing a target action; calculating a position change amount of each joint according to position changes of each foot sole before and after performing the target action; calculating a reference value of the position of each joint after performing the target action according to the position of each joint before performing the target action and the position change amount of each joint; constructing a reward term according to the reference value and the position of each joint after performing the target action, and training a base posture control model by using the reward term. 2.The method of claim 1, wherein, The method comprises: for each foot sole, determining a position change amount of the foot sole according to the position of each foot sole before and after performing the target action; mapping the position change amount of each joint based on the position change amount of each foot sole.

3. The method of claim 2, wherein the method further comprises: determining a reward value based on the state of the robot and the action; and updating the policy network based on the reward value. The method comprises: mapping the position change amount of each joint based on the position change amount of each foot sole by inverse kinematics.

4. The method of claim 3, wherein the method further comprises: The method comprises: obtaining a preset Jacobian matrix of each joint; the Jacobian matrix represents a linear relationship between the position change amount of the foot sole and the position change amount of the joint; calculating the position change amount of each joint based on the position change amount of each foot sole and a pseudo-inverse matrix of the Jacobian matrix.

5. The method of claim 4, wherein the method further comprises: determining a reward value based on the state of the robot and the action; and updating the policy network based on the reward value. The method comprises: if the Jacobian matrix is not close to a singular point, calculating the position change amount of each joint based on the position change amount of each foot sole and the pseudo-inverse matrix of the Jacobian matrix; if the Jacobian matrix is close to the singular point, adding a regularization term to the pseudo-inverse matrix of the Jacobian matrix by using a damped least square method to obtain an updated pseudo-inverse matrix, and calculating the position change amount of the joint based on the position change amount of each foot sole and the updated pseudo-inverse matrix.

6. The method of claim 2, wherein the method further comprises: The method comprises: for the position change amount of each foot sole, decomposing the position change amount of the foot sole into position change components in each direction in a coordinate system in which the foot sole is located; mapping the position change components of each foot sole into position change components of each joint; constructing the position change amount of each joint by using the position change components of each joint.

7. The method of claim 1, wherein the method further comprises: determining a reward value based on a result of the movement of the robot; and updating the policy network based on the reward value. The method comprises: for each joint, determining a position difference between the reference value of the position of the joint and an actual value of the position of the joint after performing the target action; constructing a reward term based on a norm of the position difference.

8. The method of claim 7, wherein the method further comprises: determining a reward value based on the state of the robot and the action; and updating the policy network based on the reward value. The method further comprises: For each joint, at least one other gap generated by the joint during the execution of the target action is identified; the other gap is a gap corresponding to factors other than the positional gap. The construction of the reward term based on the norm of the positional difference includes: A reward term is constructed based on the norm of the location gap and the norm of at least one of the other gaps.

9. The method of claim 8, wherein the method further comprises: determining a reward value based on the state of the robot and the action; and updating the policy network based on the reward value. For each of the joints, determining at least one other gap generated by the joint during the execution of the target action includes: For each joint, obtain the position of the joint at each moment during the execution of the target action; Based on the positional change of the joint between adjacent moments during the execution of the target action, the rotational smoothing gap of the joint is determined; the rotational smoothing gap is positively correlated with the positional change, and the rotational smoothing gap belongs to at least one other gap.

10. The method of claim 1, wherein, The process of training the basis posture control model using the reward term includes: Obtain the reward value output by the reward item; The baseline attitude control model is trained using reinforcement learning based on the reward value until the reward value output by the reward item is less than the first reward threshold, thus obtaining the trained baseline attitude control model.

11. The method of claim 10, wherein the method further comprises: determining a reward value based on the state of the robot and the action; and updating the policy network based on the reward value. The baseline attitude control model includes a motion control policy network; The step of training the baseline attitude control model using reinforcement learning based on the reward value until the reward value output by the reward item is less than a first reward threshold, thereby obtaining the trained baseline attitude control model, includes: Based on the reward value, the parameters of the motion control policy network in the base attitude control model are adjusted to obtain a new base attitude control model. The new baseline attitude control model is used for reinforcement learning until the reward value output by the reward item is less than the first reward threshold, thus obtaining the trained baseline attitude control model.

12. The method of claim 11, wherein the method further comprises: The step of performing reinforcement learning on the new basis posture control model until the reward value output by the reward item is less than a first reward threshold, to obtain the trained basis posture control model, includes: If the reward value is greater than or equal to the first reward threshold, then the training weight corresponding to the reward value is obtained; the training weight represents the degree of reinforcement learning training of the base posture control model, and the training weight and the reward value are positively correlated. The baseline posture control model is trained using reinforcement learning based on the reward value and the training weights until the reward value output by the reward item is less than the first reward threshold, thus obtaining the trained baseline posture control model.

13. The method of claim 12, wherein the method further comprises: determining a reward value based on the state of the robot and the action; and updating the policy network based on the reward value. The step of training the base posture control model using reinforcement learning based on the reward value and the training weights until the reward value output by the reward item is less than the first reward threshold, thereby obtaining the trained base posture control model, includes: If the reward value is less than the second reward threshold, then the first training weight corresponding to the reward value is determined, and the base posture control model is trained by reinforcement learning based on the reward value and the first training weight until the reward value output by the reward item is less than the first reward threshold, and the trained base posture control model is obtained; the second reward threshold is greater than the first reward threshold. If the reward value is greater than or equal to the second reward threshold, a second training weight corresponding to the reward value is determined, and reinforcement learning training is performed on the base posture control model based on the reward value and the second training weight until the reward value output by the reward item is less than the first reward threshold, so as to obtain the trained base posture control model; the second training weight is greater than the first training weight.

14. The method of claim 10, wherein the method further comprises: The reinforcement learning training of the base posture control model based on the reward value until the reward value output by the reward item is less than the first reward threshold, so as to obtain the trained base posture control model, comprises: a training duration is counted, and a learning factor corresponding to the training duration is obtained; the learning factor represents the difficulty of the reinforcement learning training of the base posture control model, and the learning factor and the training duration are positively correlated; the reinforcement learning training of the base posture control model based on the reward value and the learning factor until the reward value output by the reward item is less than the first reward threshold, so as to obtain the trained base posture control model.

15. The method of claim 1, wherein, The method further comprises, after the calculation of the position change amount of each joint according to the position change of each foot sole before and after the execution of the target action: for the position change amount of each joint, a corresponding physical constraint range of the joint is obtained; if the position change amount of the joint is within the corresponding physical constraint range, the step of calculating the reference value of the position of each joint after the execution of the target action according to the position of each joint before the execution of the target action and the position change amount of the joint is executed; if the position change amount of the joint exceeds the corresponding physical constraint range, the position change amount of the joint is updated based on a preset joint executable strategy to obtain an updated position change amount, and the step of calculating the reference value of the position of each joint after the execution of the target action according to the position of each joint before the execution of the target action and the position change amount of the joint is executed with the updated position change amount.

16. The method of claim 15, wherein the method further comprises: The physical constraint range comprises at least one of a joint angle range and a motor torque range.

17. The method of claim 1-16, wherein, The reward item is a quadratic norm.

18. A method for motion control of a legged robot with a floating base based on reinforcement learning, characterized by, The method comprises: obtaining current motion state information of the legged robot; outputting a motion strategy for controlling the motion of the legged robot based on the current motion state information and the trained base posture control model; The trained base posture control model is obtained according to the motion control strategy network training method of the legged robot with a floating base based on reinforcement learning in any one of claims 1-17.

19. A training device for a motion control policy network of a legged robot with a floating basis based on reinforcement learning, characterized in that, The device comprises: a training position acquisition module configured to acquire the positions of each joint and the positions of each foot sole of the legged robot before and after the execution of a target action; a joint position change amount calculation module configured to calculate the position change amount of each joint according to the position change of each foot sole before and after the execution of the target action; a reference value calculation module configured to calculate the reference value of the position of each joint after the execution of the target action according to the position of each joint before the execution of the target action and the position change amount of the joint. A model training module is configured to construct a reward term according to the reference value and the position of each joint after the target action is performed, and train a base posture control model by using the reward term.

20. A motion control device for a legged robot with a floating base based on reinforcement learning, characterized by, The device comprises: An operating state information acquisition module is configured to acquire current motion state information of the legged robot. An action policy output module is configured to output an action policy for controlling the motion of the legged robot based on the current motion state information and the trained base posture control model. The trained base posture control model is obtained according to the motion control strategy network training method for the legged robot with a floating base based on reinforcement learning in any one of claims 1-17.

21. An electronic device, comprising: It comprises: a processor; and a memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement the motion control strategy network training method for the legged robot with a floating base based on reinforcement learning in any one of claims 1-17, or the motion control method for the legged robot with a floating base based on reinforcement learning in claim 18.

22. A robot, characterized in that It comprises: a processor; and a memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement the motion control method for the legged robot with a floating base based on reinforcement learning in claim 18.

23. The robot of claim 22, wherein, The legged robot can be any one of a quadruped robot, a biped robot, a humanoid robot, and a mobile robot.

24. A computer-readable storage medium, characterized in that, The computer readable storage medium has computer program code instructions stored thereon, which, when invoked by the processor of the robot, causes the robot to perform the motion control strategy network training method for the legged robot with a floating base based on reinforcement learning in any one of claims 1-17, or the motion control method for the legged robot with a floating base based on reinforcement learning in claim 18.

Citation Information

Patent Citations

  • Visual robot control method based on time sequence prediction model

    CN118331052A

  • Master policy training method of hierarchical reinforcement learning with asymmetrical policy architecture

    US20230362196A1

  • Control method and apparatus for legged robot, and legged robot and medium

    WO2025011165A1