A Hybrid Balancing Control Method and Device for a Two-Wheeled Legged Robot Optimized by Reinforcement Learning
By combining model-based controllers and reinforcement learning strategies in the two-wheeled foot robot, optimizing the control torque, the problem of balance control of the two-wheeled foot robot is solved, and higher balance performance and robustness are achieved.
Patent Information
- Application Number
- CN202510333313.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Due to its instability and complexity, traditional model-based controllers are difficult to effectively solve, and directly applying reinforcement learning algorithms requires a large amount of data. It is easy to deviate from the expected state in the early stage of exploration and it is difficult to find an effective control strategy.
A hybrid balance control method based on reinforcement learning optimization is proposed. By defining the coordinate system and deriving the positive and inverse kinematics model, a model-based balance and attitude controller is established as a basic controller, and on it, and reinforcement learning strategies are trained. Through real-time online compensation of control torque, all joint motors are coupled and optimized to improve balance performance.
This method effectively reduces the difficulty of controller design, improves the balance performance of the double-wheeled foot robot, enhances the robustness and generalization capabilities of the system, and is suitable for generalized double-wheeled foot robots and different traditional controllers.
Smart Images

Figure CN119882406B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine control, and particularly to a hybrid balance control method and device for a two-wheeled legged robot optimized based on reinforcement learning. Background Art
[0002] The two-wheeled legged robot cleverly combines the advantages of high-speed movement of wheels and obstacle-crossing ability of legs, enabling it to exhibit excellent mobility and adaptability in complex and changing environments, and having broad application prospects in fields such as logistics, search and rescue, and surveying. However, at the same time, as an underactuated, non-linear and unstable system, the balance control of the two-wheeled legged robot has always been a research difficulty. At the same time, good balance performance is not only a prerequisite for realizing complex robot actions, but also the key to its adaptation to complex terrains and completion of diverse tasks. Traditional model-based controllers have many limitations, and their performance is severely restricted by the accuracy of the model. When the robot has a high degree of joint freedom, the controller design becomes particularly complex. In view of the limitations of traditional controllers, deep reinforcement learning provides a data-driven method. It directly interacts with the environment and learns the end-to-end mapping from perception to action without the need for an accurate system model, and can better handle complex and dynamic robot control tasks. However, the two-wheeled legged robot itself is unstable, which leads to the need for a large amount of data for training when directly applying standard reinforcement learning algorithms. The robot is prone to deviate from the desired state during the initial exploration and it is difficult to find an effective control strategy. To solve the above problems, there is an urgent need for a new balance control scheme that can not only reduce the design difficulty of the controller but also improve the balance performance. Summary of the Invention
[0003] In view of the above technical problems, the present invention proposes a hybrid balance control method and device for a two-wheeled legged robot optimized based on reinforcement learning, and the technical solution adopted is as follows:
[0004] According to a first aspect of the present invention, there is provided a hybrid balance control method for a two-wheeled legged robot optimized based on reinforcement learning, the method comprising the following steps:
[0005] S100, based on the mechanical structure of the two-wheeled legged robot, define a coordinate system, and sequentially deduce its forward kinematics and inverse kinematics models to describe the basic physical characteristics of the two-wheeled legged robot;
[0006] S200, simplify the two-wheeled legged robot into a second-order inverted pendulum model, and establish a model-based balance controller according to the target moving distance and target moving speed to obtain the control torque of the hub motor;
[0007] S300, based on the forward and inverse kinematics models, establish a model-based attitude controller according to the target center-of-gravity height and target attitude to obtain the control torques of the front swing motor and the knee-bending motor;
[0008] In S400 and S200, the balance controller and the attitude controller in S300 are combined as a model-based basic controller to initially ensure the stability of the system motion;
[0009] In S500, based on the S400 controller, a reinforcement learning strategy is trained. By compensating the control torque in real-time online, the coupling optimization of all joint motors is performed to improve the balance performance;
[0010] All the joint motors mentioned above include hub motors, front swing motors, and knee motors.
[0011] Furthermore, the process of coordinate system definition, forward kinematics, and inverse kinematics model establishment in step S100 of the present invention satisfies the following conditions:
[0012] First, the homogeneous transformation matrix from the base coordinate system {B} to each joint coordinate system is derived ; then the homogeneous transformation matrix from the source coordinate system {O} to the base coordinate system {B} is derived , and based on this, the homogeneous transformation matrix from the control coordinate system {C} to the base coordinate system {B} is derived ; according to the multiplication rule, the homogeneous transformation matrix from the control coordinate system {C} to each joint coordinate system is obtained . By virtue of this transformation, control commands are applied in the control system {C} to achieve precise control of each joint of the robot;
[0013] In the above derivation, the base coordinate system {B} is fixed on the head of the floating base of the robot. The origin of {B} coincides with the center of the inertial measurement unit, and the directions of its x, y, and z axes change following the rotation of the head. It is the starting point for connecting each joint coordinate system; the origin of the control coordinate system {C} is located at the midpoint of the connection line of the two wheels, the x-axis points in the forward direction of the robot, the z-axis is vertically upward, and the positive direction of the y-axis is obtained according to the right-hand rule; the origin of the source coordinate system {O} coincides with the base coordinate system {B}, and the positive directions of its x, y, and z axes are the same as those of the control coordinate system {C}. The left front swing motor coordinate system {1L}, the right front swing motor coordinate system {1R}, the left knee motor coordinate system {2L}, the right knee motor coordinate system {2R}, the left hub motor coordinate system {3L}, and the right hub motor coordinate system {3R} are the joint motor coordinate systems of the two-wheel foot robot;
[0014] Define the state vector , which includes the rotation angles of each joint motor , and the roll angle , pitch angle , and yaw angle of the floating base. Based on the conversion relationship between coordinate systems, the center-of-gravity position of the robot relative to the control coordinate system is obtained , and its expression is , where is the position of the centroid of the link in the control coordinate system {C}, represents the mass of the link. Further, the Jacobian matrix for inverse-solving the joint position from the centroid position is derived:
[0015] .
[0016] Further, the second-order inverted pendulum model establishment and balance controller design in step S200 of the present invention satisfy the following conditions:
[0017]
[0018] where , where is the wheel mass, is the floating base mass, is the wheel radius, is the moving displacement, is the first derivative of is the second derivative of represents the length of the pendulum, corresponding to the height of the robot, is the forward tilt angle of the vehicle body, is the first derivative of is the second derivative of is the wheel control torque. Define the current balanced state of the system as . Take the deviation between the current balanced state and the desired state as the feedback signal. After multiplying by the gain matrix , the optimal hub motor control torque input is obtained: .
[0019] Further, the attitude controller design in step S300 of the present invention satisfies the following conditions:
[0020]
[0021] where, represents the target centroid height tracking task, represents the target attitude angle task, represents that the wheel lateral slip is 0, represents the joint angle vector, represents the attitude vector, is the corresponding Jacobian matrix. Through inverse solution, the target speeds of the front swing motor and the knee bending motor are obtained by solving the task target:
[0022]
[0023] where the subscript represents the target value.
[0024] Furthermore, the model-based basic controller described in step S400 of the present invention satisfies the following conditions:
[0025] The balance controller that outputs the control torque of the hub motor and the attitude controller that outputs the control torque of the front swing motor and the knee bending motor are combined as the model-based basic controller to initially ensure the stable operation of the system.
[0026] Furthermore, the reinforcement learning optimization strategy described in step S500 of the present invention satisfies the following conditions:
[0027] The definition of the state s of the reinforcement learning optimization strategy at time t is where is the robot state vector, defined as , is the desired balance state and the desired pose vector, defined as , particularly, in order to enhance the robustness of the strategy under unknown terrains, the inaccurate desired pitch angle information is intentionally removed from , is the motion error and the history vector, defined as , where the error is calculated according to each component in the desired vector , is the control torque and the torque history vector, defined as , where represents the control torque of the basic controller, represents the optimized compensation torque of the reinforcement learning strategy, represents the total control torque finally applied to each joint.
[0028] According to the second aspect of the present invention, there is provided a hybrid balance control device for a two-wheeled legged robot based on reinforcement learning optimization. The device includes:
[0029] A dynamic model establishment module, configured to define a coordinate system according to the mechanical structure of the two-wheeled legged robot and deduce its forward kinematics and inverse kinematics models;
[0030] The model-based balance controller module is used to simplify the two-wheeled legged robot into a second-order inverted pendulum model. Based on the target moving distance and target moving speed, a model-based balance controller is established to obtain the control torque of the hub motor.
[0031] The model-based attitude controller module is used to establish a model-based attitude controller based on the forward and inverse kinematic models according to the target center-of-gravity height and target attitude, and obtain the control torques of the front swing motor and the knee joint motor.
[0032] The basic controller module is used to combine the balance controller and the attitude controller as a model-based basic controller to initially ensure the stable movement of the system.
[0033] The reinforcement learning optimization module trains the reinforcement learning policy, and through real-time online compensation of the control torque, performs coupled optimization on all joint motors to improve the balance performance.
[0034] All the joint motors mentioned above include the hub motor, the front swing motor, and the knee joint motor.
[0035] The present invention has at least the following beneficial effects:
[0036] A hybrid balance control method and device for a two-wheeled legged robot based on reinforcement learning optimization provided by an embodiment of the present invention fully combines the advantages of traditional controllers and reinforcement learning. On the one hand, the addition of a traditional controller can stabilize the training process and improve the sample efficiency of learning; on the other hand, reinforcement learning can perform online optimization on the traditional controller, thereby improving the control accuracy, enhancing the robustness and generalization ability of the system. The reinforcement learning policy directly compensates the control torque, and this framework has good applicability and scalability for general two-wheeled legged robots and different traditional controllers.
[0037] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0039] Figure 1 It is a flowchart of a hybrid balance control method for a two-wheeled legged robot based on reinforcement learning optimization provided by an embodiment of the present invention.
[0040] Figure 2 It is a diagram of the mechanical connection structure and coordinate system definition of the two-wheeled legged robot provided by the embodiment of the present invention;
[0041] Figure 3 It is a schematic diagram of the simplified second-order inverted pendulum structure of the two-wheeled legged robot provided by the embodiment of the present invention;
[0042] Figure 4 It is a curve graph of the balance task execution state of the two-wheeled legged robot provided by the embodiment of the present invention;
[0043] Figure 5 It is a curve graph of the position and pitch angle changes of the two-wheeled legged robot under external force interference provided by the example of the present invention. Detailed implementation manners
[0044] Next, in combination with the accompanying drawings of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0045] The embodiment of the present invention provides a hybrid balance control method for a two-wheeled legged robot optimized based on reinforcement learning, as Figure 1 shown, the method includes the following steps:
[0046] S100, based on the mechanical structure of the two-wheeled legged robot, define a coordinate system, and sequentially deduce its forward kinematics and inverse kinematics models to describe the basic physical characteristics of the two-wheeled legged robot.
[0047] In the embodiment of the present invention, the structure of the two-wheeled legged robot can be an existing structure, including a base, two hip joints, left and right legs, and hardware devices carried, etc. Among them, 3 joint motors are arranged on each hip joint, one motor is used to control the front and back swing of the leg, one motor is used to control the left and right swing of the leg, and one motor is used to control the angle of the knee joint. Hub motors and wheels are installed at the foot end of each leg, and the hub motors are used to control the movement of the wheels. The hardware devices carried can include an IMU and an on-board computing unit, etc.
[0048] In the example of the present invention, Figure 2Shows the front view of the mechanical structure of the two-wheeled legged robot and the distribution of the coordinate system. In the present invention, the optimization of the motor without considering the left and right swing is not considered. At the same time, the specific positions of the front swing motor, the knee bending motor and the hub motor are marked in the figure. The coordinate system is defined as follows: The base coordinate system {B} is fixed on the head of the floating base of the robot, the origin of {B} coincides with the center of the inertial measurement unit, and the directions of its x, y, and z axes change with the rotation of the head, which is the starting point for connecting each joint coordinate system; the origin of the control coordinate system {C} is located at the midpoint of the line connecting the two wheels, the x-axis points in the forward direction of the robot, the z-axis is vertically upward, and the positive direction of the y-axis is obtained according to the right-hand rule; the origin of the source coordinate system {O} coincides with the base coordinate system {B}, and the positive directions of its x, y, and z axes are the same as those of the control coordinate system {C}. The origin of the inertial coordinate system {I} coincides with the base coordinate system {B}, and the positive directions of its x, y, and z axes are the same as those of the world coordinate system. The left front swing motor coordinate system {1L}, the right front swing motor coordinate system {1R}, the left knee bending motor coordinate system {2L}, the right knee bending motor coordinate system {2R}, the left hub motor coordinate system {3L}, and the right hub motor coordinate system {3R} are the joint motor coordinate systems of the two-wheeled legged robot, and the rotation angles of each joint motor are respectively .
[0049] In the example of the present invention, the derivation process of the forward kinematics is as follows: First, derive the homogeneous transformation matrix from the base coordinate system {B} to each joint coordinate system ; Then, derive the homogeneous transformation matrix from the source coordinate system {O} to the base coordinate system {B} , and based on this, derive the homogeneous transformation matrix from the control coordinate system {C} to the base coordinate system {B} ; According to the multiplication rule, obtain the homogeneous transformation matrix from the control coordinate system {C} to each joint coordinate system . By virtue of this transformation, by applying control commands in the control system {C}, precise control of each joint of the robot can be achieved.
[0050] In the example of the present invention, the derivation of the position of the center of gravity of the robot is as follows: Define the state vector , which includes the rotation angles of each joint motor and the roll angle , pitch angle , and yaw angle . Based on the conversion relationship between coordinate systems, the position of the center of gravity of the robot relative to the control coordinate system can be obtained , and its expression is , where is the position of the center of mass of the th link in the control coordinate system {C}, represents the mass of the th link. Further, the Jacobian matrix for inverse solving the joint position from the center of gravity position can be derived:
[0051] 。
[0052] S200 simplifies the two-wheeled legged robot into a second-order inverted pendulum model. Based on the target moving distance and target moving speed, a model-based balance controller is established to obtain the control torque of the hub motor.
[0053] In the example of the present invention, the structure of the two-wheeled legged robot is simplified into a second-order inverted pendulum, as Figure 3 shown. According to the Lagrangian dynamics equation, the second-order dynamic model of the system is established as:
[0054]
[0055] where 。 is the mass of the wheel, is the mass of the floating base, is the radius of the wheel, is the moving displacement, is the first derivative of is the second derivative of represents the length of the pendulum, corresponding to the height of the robot. is the forward tilt angle of the vehicle body, is the first derivative of is the second derivative of is the wheel control torque.
[0056] In the example of the present invention, a linear quadratic regulator (LQR) is selected as the basic balance controller. It is an optimized linear feedback controller that minimizes a quadratic performance index, enabling the system to balance the system state and the cost of control input while reaching the desired state. Specifically, by adjusting the weight matrices Q and R, the Riccati equation is solved to obtain an optimal feedback gain matrix . The system balance state vector is defined as: , and the deviation between the current balance state and the desired state is used as the feedback signal. After multiplying by the gain matrix , the optimal control torque input of the hub motor is obtained: . At the same time, the adjustment of the yaw angle γ of the vehicle body is integrated in the LQR. The optimal parameters need to be specifically adjusted for different systems.
[0057] S300, based on the forward and inverse kinematic models, establish a model-based attitude controller according to the target center-of-gravity height and target attitude, and obtain the control torques of the front swing motor and the knee joint motor.
[0058] In the embodiment of the present invention, in order to achieve the dynamic control of the floating base pose of the robot, three specific control objectives are set: represents the center-of-gravity height tracking task, and is represented by represents the base roll angle and pitch angle adjustment tasks, and is represented by represents the task of ensuring that the wheels do not slip laterally. Define represents the joint angle vector, represents the attitude vector, and organize the control tasks into the following expression form:
[0059]
[0060] Through inverse kinematics, the target speeds of the front swing motor and the knee joint motor can be solved from the task objectives:
[0061]
[0062] where, is the corresponding Jacobian matrix, and the subscript represents the target value, and + represents the pseudo-inverse of the matrix.
[0063] In the embodiment of the present invention, a Proportion Differential (PD) controller is selected to achieve the dynamic control of the floating base pose of the robot. The PD controller measures the deviation between the actual output of the system and the set value, that is, the current error e(t) of the system and its change rate de(t) / dt, and calculates the control actions of the proportional term and the derivative term according to the proportional gain Kp and the derivative gain Kd, and adds the two to obtain the final control torque, which acts on the system to reduce the error and make the system output track the set value. In the present invention, different PD control parameters are used for the front swing motor and the knee joint motor respectively to obtain the corresponding control torques. Kp and Kd are adjustable parameters.
[0064] S400, combine the balance controller in S200 and the attitude controller in S300 as a model-based basic controller to initially ensure the stability of the system movement.
[0065] In the embodiment of the present invention, the balance controller that outputs the control torque of the hub motor is combined with the attitude controller that outputs the control torques of the front swing motor and the knee joint motor to obtain a model-based basic controller. This controller can initially ensure the stability of the system movement and provide a stable initial state for subsequent reinforcement learning training.
[0066] The S500, based on the S400 controller, trains a reinforcement learning policy. By compensating the control torque in real-time online, it couples and optimizes all joint motors (including hub motors, front swing motors, and knee motors) to improve the balance performance.
[0067] In the example of the present invention, the state space of the reinforcement learning optimization policy is defined as follows: The state s at time t is defined as . Where is the robot state vector, defined as . is the desired balance state and desired pose vector, defined as . In particular, to enhance the robustness of the policy under unknown terrains, the inaccurate desired pitch angle information is deliberately removed in . The aim is to stimulate the agent to learn from other state information, so as to make accurate decisions even when facing unknown angular terrains. is the motion error and history vector, defined as , where the error is calculated according to the components in the desired vector . is the control torque and torque history vector, defined as , where represents the control torque of the basic controller, represents the optimized compensation torque of the reinforcement learning policy, represents the total control torque finally applied to each joint.
[0068] In the example of the present invention, the reinforcement learning optimization policy directly outputs a 6-dimensional compensation torque . Based on the control torque , it couples and optimizes a total of 6 motors of the whole body. The total control torque finally applied to the motors is . The Soft Actor-Critic (SAC) algorithm is used to train the policy network, which is a reinforcement learning method based on a stochastic policy. By introducing the concept of entropy, SAC prompts the agent to explore more state-action pairs, thus avoiding getting stuck in local optimal solutions. The stochastic exploration policy can effectively balance exploration and exploitation, improve the learning efficiency, and finally obtain a better policy. At the same time, the reinforcement learning optimization policy in the present invention interacts with the basic controller only through torque information, and the design of the state space and action space has nothing to do with the specific structure inside the basic controller. This modular design enables this method to be flexibly integrated into other model-based control frameworks and has good generality.
[0069] In the example of the present invention, the reward function is designed as : The first part It represents the balance reward and the pose reward, optimizing the terminal task performance; the second part represents the joint limit reward, ensuring the motion safety of the robot. In the design of this reward function, only the maximum limit is taken for the state of each joint, and the joint state is not used to guide the task optimization. By measuring the deviation between the terminal state and the target to guide the optimization of performance, expressed as . The first term measures the balance state performance, encourages the robot to reduce the tracking error, and encourages the robot's floating base to be as stable as possible; the second term encourages the robot to maintain the desired center of gravity height ; the third term measures the pose performance of the robot . All the error terms measured in correspond to the terms in the error vector in order, and the function is used to quantify them.
[0070] In the example of the present invention, compared with the restriction on the movement of a single joint, the present invention adopts a more global reward function design, that is, directly evaluating the performance of the robot according to the completion of the terminal task. Guiding the policy update according to the actual balance performance and pose performance of the robot, enabling it to make decisions from the perspective of the overall task, rather than only focusing on the movement optimization of a certain local action or joint. This design gives the robot a greater degree of exploration freedom, and can prompt it to learn a better action strategy that comprehensively considers the global goal. In addition, the reward design based on the terminal task can also avoid the local optimum problem that may exist in the traditional method, and improve the adaptability of the robot in a diverse and dynamic environment.
[0071] In the example of the present invention, in order to improve the robustness of the model to environmental uncertainty and model parameter changes, a threshold randomization strategy is introduced during the training process. It includes the randomization of physical parameters, the randomization of sensor data, and the randomization of environmental parameters, etc.
[0072] In the example of the present invention, the policy network adopts a two-layer fully connected structure, with the number of neurons being 1024 and 512 respectively, and the activation function being tanh. The network input is a 15-dimensional state vector, and the output is a 6-dimensional action vector. The SAC algorithm is used for training, with the maximum number of iterations set to 50000 steps, the policy network updated every 1000 steps, and the batch size being 512. The learning rate is 1e-3, the discount factor is 0.99, the entropy regularization coefficient is initialized to 0.1, and the reward scaling coefficient is 10. The algorithm is trained on a computer equipped with a 12-core CPU and an NVIDIA 3060 GPU. The specific structure of the network and the training hyperparameters need to be adjusted specifically.
[0073] In the example of the present invention, two sets of training tasks are set up to verify the effectiveness of the method. Group 1: The height remains unchanged, and a moving target between 1 and 4 meters is randomly set, and slopes with three different gradients of 0 degrees, 5 degrees, and 10 degrees are considered. Group 2: A moving target between 1 and 4 meters on flat ground is randomly set, and a target for the change in the center of gravity height between 0.05 and 0.15 meters is randomly generated. The specific parameters in the tasks can all be adjusted.
[0074] A hybrid balance control method for a two-wheeled legged robot optimized based on reinforcement learning provided by an embodiment of the present invention. In an actual application scenario, the two-wheeled legged robot moves straight on flat ground while maintaining the stability of the center of gravity height and posture. The motion curve diagram is as Figure 4 shown. It can be seen from Figure 4 that the robot can balance various tasks and stably complete the target. The comparison method is before the reinforcement learning optimization, and there is only the effect of controlling based on the basic controller. On the basis of maintaining the stability of the front inclination angle and the posture angle, the tracking accuracy of the method of the present invention in terms of position and height is significantly improved. In particular, the method of the present invention almost coincides with the desired trajectory in terms of position. In addition, tests are carried out, and instantaneous external forces of 10N, 20N, and 30N are applied at 12s, 20s, and 32s respectively. The test results are as Figure 5 shown. Whether during the movement process or at rest, the proposed method can quickly return to the target trajectory after being disturbed by an external force in the opposite direction and effectively suppress the oscillation of the front inclination angle. While the comparison method shows obvious oscillations in both the front inclination angle and the position and it is difficult to return to stability in a short time.
[0075] A hybrid balance control method for a two-wheeled legged robot optimized based on reinforcement learning provided by an embodiment of the present invention fully combines the advantages of traditional controllers and reinforcement learning. The addition of a traditional controller can stabilize the training process and improve the sample efficiency of learning. At the same time, reinforcement learning can perform online optimization on the traditional controller, thereby improving the control accuracy, enhancing the robustness and generalization ability of the system. The reinforcement learning strategy directly compensates the control torque, and this framework has good applicability and scalability for general two-wheeled legged robots and different traditional controllers.
[0076] Based on the same inventive concept, an embodiment of the present invention further provides a hybrid balance control device for a two-wheeled legged robot optimized by reinforcement learning. The device includes:
[0077] A kinematic model establishment module, configured to define a coordinate system according to the mechanical structure of the two-wheeled legged robot, and deduce its forward kinematic and inverse kinematic models;
[0078] A model-based balance controller module, configured to simplify the two-wheeled legged robot into a second-order inverted pendulum model, and establish a model-based balance controller according to the target moving distance and target moving speed to obtain the control torque of the hub motor;
[0079] A model-based attitude controller module, configured to establish a model-based attitude controller according to the target center of gravity height and target attitude based on the forward and inverse kinematic models to obtain the control torques of the front swing motor and the knee motor;
[0080] A basic controller module, configured to combine the balance controller and the attitude controller as a model-based basic controller to initially ensure the stable movement of the system;
[0081] A reinforcement learning optimization module, which trains a reinforcement learning policy, and performs coupled optimization on all joint motors (including the hub motor, the front swing motor, and the knee motor) by compensating the control torque in real time online to improve the balance performance.
[0082] This device can be used to execute Figure 1 the method shown in the embodiments shown, therefore, for the functions that can be realized by each functional module of this device, reference can be made to Figure 1 the description of the embodiments shown, and details will not be repeated here.
[0083] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and the present invention is not limited herein.
[0084] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A hybrid balance control method for a two-wheeled legged robot based on reinforcement learning optimization, characterized in that: The method comprises the following steps: S100, based on the mechanical structure of the two-wheeled legged robot, define a coordinate system, and derive its forward kinematics and inverse kinematics models in sequence to describe the basic physical characteristics of the two-wheeled legged robot; S200, simplify the two-wheeled legged robot into a second-order inverted pendulum model, establish a model-based balance controller according to the target moving distance and target moving speed, and obtain the control torque of the wheel hub motor; S300, based on the forward and inverse kinematics model, according to the target center of gravity height and the target posture, establish a model-based posture controller to obtain the control torque of the forward swing motor and the knee flexion motor; S400 combines the balance controller in S200 with the attitude controller in S300 as a model-based basic controller to initially ensure the stability of the system motion; S500, based on the S400 controller, trains a reinforcement learning strategy to optimize the coupling of all joint motors through real-time online compensation of control torque to improve balance performance; The described total joint motors include a wheel hub motor, a front swing motor and a knee flexion motor.
2. The method according to claim 1, characterized in that: The coordinate system definition and the process of establishing the forward kinematics and inverse kinematics models in step S100 meet the following conditions: First, derive the homogeneous transformation matrix from the base coordinate system {B} to each joint coordinate system ; Then derive the homogeneous transformation matrix from the source coordinate system {O} to the base coordinate system {B} , and based on this, the homogeneous transformation matrix from the control coordinate system {C} to the base coordinate system {B} is derived ; According to the continuous multiplication rule, the homogeneous transformation matrix from the control coordinate system {C} to each joint coordinate system is obtained ,With this transformation, control instructions are applied under the control system {C} to achieve precise control of each joint of the robot; In the above derivation, the base coordinate system {B} is fixed on the head of the robot's floating base, the origin of {B} coincides with the center of the inertial measurement unit, and the directions of its x, y, and z axes change with the rotation of the head, which is the starting point for connecting the coordinate systems of each joint; the origin of the control coordinate system {C} is located at the midpoint of the line connecting the two wheels, the x axis points to the robot's forward direction, the z axis is vertically upward, and the positive direction of the y axis is obtained according to the right-hand rule; the origin of the source coordinate system {O} coincides with the base coordinate system {B}, and the positive directions of its x, y, and z axes are consistent with the control coordinate system {C}, the left front swing motor coordinate system {1L}, the right front swing motor coordinate system {1R}, the left knee flexion motor coordinate system {2L}, the right knee flexion motor coordinate system {2R}, the left hub motor coordinate system {3L}, and the right hub motor coordinate system {3R} are the coordinate systems of the joint motors of the two-wheeled robot; Define the state vector , which includes the rotation angle of each joint motor , and the rolling angle of the floating base , pitch angle , and the yaw angle , based on the conversion relationship between coordinate systems, the center of gravity position of the robot relative to the control coordinate system is obtained , whose expression is ,in For the The position of the center of mass of the connecting rod in the control coordinate system {C}, Representative The mass of the connecting rod, further, derives the Jacobian matrix of the joint position from the center of gravity position: 。 3. The method according to claim 1, characterized in that: The second-order inverted pendulum model establishment and balance controller design described in step S200 meet the following conditions: in ,in is the wheel mass, is the mass of the floating base, is the wheel radius, is the displacement, for The first derivative of for The second derivative of represents the length of the pendulum, corresponding to the height of the robot, is the forward tilt angle of the vehicle body, for The first derivative of for The second derivative of is the wheel control torque, defining the current equilibrium state of the system as , the current equilibrium state and expected state The deviation is used as the feedback signal and multiplied by the gain matrix After that, the optimal wheel hub motor control torque input is obtained. : .
4. The method according to claim 1, characterized in that: The posture controller design in step S300 satisfies the following conditions: in, represents the target center of gravity height tracking task, represents the target attitude angle task, Indicates that the wheel lateral slip is 0, represents the joint angle vector, represents the posture vector, is the corresponding Jacobian matrix. By inverse solution, the target speeds of the front swing motor and the knee flexion motor are obtained from the task objectives: The corner mark Indicates the target value.
5. The method according to claim 1, characterized in that The model-based basic controller in step S400 satisfies the following conditions: The balance controller that outputs the control torque of the hub motor and the posture controller that outputs the control torque of the front swing motor and the knee flexion motor are combined as the basic controller based on the model to initially ensure the stability of the system operation.
6. The method according to claim 1, characterized in that The reinforcement learning strategy described in step S500 satisfies the following conditions: The reinforcement learning policy state s at time t is defined as ,in is the robot state vector, defined as , is the desired equilibrium state and desired posture vector, defined as In order to enhance the robustness of the strategy in unknown terrain, Inaccurate expected pitch angles intentionally removed information, is the motion error and history vector, defined as , where the error According to the expected vector Calculate the components in is the control torque and torque history vector, defined as ,in represents the control torque of the basic controller, represents the optimal compensation torque of the reinforcement learning strategy, It represents the total control torque finally applied to each joint.
7. The method according to claim 1, characterized in that The reinforcement learning strategy described in step S500 satisfies the following conditions: Reinforcement learning strategy directly outputs 6-dimensional compensation torque , in the control torque Based on this, the coupling optimization is performed on the 6 motors in the whole body, and the total control torque applied to the motor is .
8. The method according to claim 1, characterized in that The reinforcement learning strategy described in step S500 satisfies the following conditions: The reward function is designed as : Part 1 It represents the balance reward and posture reward, and optimizes the terminal task performance; The second part represents the joint restriction reward, which ensures the robot's motion safety. In the design of this reward function, only the maximum restriction is taken for the state of each joint, and the joint state is not used to guide task optimization; By measuring the terminal status and the target The deviation is used to guide the optimization performance, expressed as , the first item Measuring the state of equilibrium performance, encouraging the robot to reduce tracking errors and encouraging the robot's floating base to be as stable as possible; the second Encourage the robot to maintain the desired center of gravity height ; Item 3 Measuring the robot’s posture performance , The order of all error terms measured in corresponds to the error vector Each item in the function To quantify.
9. A hybrid balance control device for a two-wheeled legged robot based on reinforcement learning optimization that implements any one of the methods of claims 1-8, characterized in that: The device comprises: The dynamic model building module is used to define the coordinate system according to the mechanical structure of the two-wheeled foot robot and derive its forward kinematics and inverse kinematics models; The model-based balance controller module is used to simplify the two-wheeled legged robot into a second-order inverted pendulum model. According to the target moving distance and target moving speed, a model-based balance controller is established to obtain the control torque of the wheel hub motor. A model-based attitude controller module is used to establish a model-based attitude controller based on the forward and inverse kinematics models according to the target center of gravity height and target attitude, and obtain the control torque of the forward swing motor and the knee flexion motor; The basic controller module is used to combine the balance controller and the attitude controller as a model-based basic controller to initially ensure the stability of the system motion; Reinforcement learning optimization module, which trains reinforcement learning strategies, optimizes the coupling of all joint motors through real-time online compensation of control torque, and improves balance performance; The described total joint motors include a wheel hub motor, a front swing motor and a knee flexion motor.
Citation Information
Patent Citations
Quadruped robot motion control method based on deep learning
CN116627041A
Biped robot complex terrain adaptive gait planning method and biped robot
CN119644704A