Sliding-mode control double-foot joint trajectory tracking optimization method based on reinforcement learning
By combining reinforcement learning and sliding mode control, a kinematic and dynamic model of bipedal robot is established and the parameters of sliding mode controller are updated in real time, which solves the problems of trajectory tracking accuracy and jitter of bipedal robots in complex environments, and high-precision adaptive control is achieved.
Patent Information
- Application Number
- CN202510333068.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to achieve accurate tracking of bipedal robot joint trajectory in a changing environment. Sliding mode control is prone to vibration during dynamic processes, and the controller parameters dynamic adjustment ability is insufficient.
Combining reinforcement learning and sliding mode control, by establishing kinematic and dynamic models of n-link mechanisms, a sliding mode controller is designed, and using reinforcement learning to update the controller parameters in real time, adaptive sliding mode control is realized.
It improves the trajectory tracking accuracy of bipedal robots in complex environments and the adaptability of the control system, reduces vibration, and enhances the robustness to external disturbances.
Smart Images

Figure CN120255555A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biped robot joint position control, and particularly relates to a sliding mode control biped joint trajectory tracking optimization method based on reinforcement learning. Background Technique
[0002] Compared with wheeled robots, biped robots have stronger terrain adaptability and a wider range of application scenarios. For example, they can cross obstacles, walk on complex terrains, and even perform actions such as running, which makes biped robots have broad application prospects in the fields of service, rescue, scientific research, etc. However, biped gait is an unstable motion mode. In the process of stable walking, the position of the center of mass and the poses of each joint need to satisfy the stability relationship, which requires ensuring the tracking accuracy of the joint space trajectory, meaning more stringent requirements are put forward for the control of the force and torque of the joint drive.
[0003] Compared with traditional control methods, such as PID control, model predictive control, etc., sliding mode control has stronger robustness. It shows excellent insensitivity to the model error, parameter change, and external disturbance of the controlled object. This is because its core idea is to design a sliding mode surface so that the system can reach the sliding mode surface from any position in the state space in a finite time and then slide along this surface to the desired state. And during the dynamic process, the sliding mode control can change purposefully according to the current state of the system, forcing the system to move along the state trajectory of the predetermined "sliding mode", which is especially suitable for dealing with nonlinear systems. However, when the state trajectory reaches the sliding mode surface, it is difficult to strictly slide along the sliding mode surface to the equilibrium point, but approaches the equilibrium point by crossing back and forth on both sides of it, resulting in chattering, which is a disadvantage of sliding mode control.
[0004] Selecting appropriate controller parameters can optimize the approaching speed of the system to the sliding mode surface, improve the ability to overcome perturbations / external disturbances, thereby improving the dynamic performance of the control system and reducing chattering. However, during the motion control process, the dynamic characteristics change, and the fixed controller parameters have poor dynamic adjustment ability. Therefore, it is very necessary to enhance the adaptability of the controller to the changing dynamic characteristics. Reinforcement learning learns by interacting with the environment and has the ability to make adaptive decisions according to the environment. Therefore, combining reinforcement learning with sliding mode control, using reinforcement learning to determine the controller parameters in different environments, can enhance the adaptability and control accuracy of the control system.
[0005] Precise control of the motion behavior of a legged robot can improve its adaptability to complex environments. In order to flexibly adjust the control strategy for a changing environment and make optimal or sub-optimal motion behaviors, the present invention combines the idea of reinforcement learning with the sliding mode control method, and proposes a sliding mode control biped joint trajectory tracking optimization method based on reinforcement learning. The proposed control strategy uses the method of reinforcement learning to select the parameters of the sliding mode controller, realizes the environmental adaptability of the control method, and is of great significance for finally realizing the precise control of trajectory tracking. Summary of the Invention
[0006] (1) Technical Problems to be Solved
[0007] The technical problem to be solved by the present invention is: how to propose a sliding mode control biped joint trajectory tracking optimization method based on reinforcement learning.
[0008] (2) Technical Solutions
[0009] To solve the above technical problems, the present invention provides a sliding mode control biped joint trajectory tracking optimization method based on reinforcement learning. The method includes the following steps:
[0010] Step 1: According to the simplified link structure of the leg of the legged robot, and connecting adjacent links with joints, establish the kinematic model of the n-link mechanism, so as to obtain the angular velocities of each joint and the end effector.
[0011] Step 2: Based on the angular velocities of each joint and the end effector obtained from the n-link mechanism, establish a dynamic model to obtain the relationship between the corresponding torque and motion parameters.
[0012] Step 3: Use the tracking error between the ideal trajectory of each joint and the actual trajectory of each corresponding joint as the control target, and design a sliding mode controller.
[0013] Step 4: Based on the method of reinforcement learning, update the parameters of the sliding mode controller in real time to realize adaptive sliding mode control.
[0014] Among them, in the above Step 1, establishing the kinematic model of the n-link mechanism is divided into trajectory planning and joint coordinate system transformation; for the establishment of trajectory planning, by giving the displacement-time function of the end spatial motion trajectory, and then through the kinematic relationship of each joint, calculate the change law of each joint angle with time, and obtain the trajectory planning.
[0015] Among them, in the above Step 1, for the joint coordinate system transformation, use the D-H method to establish the transformation relationship between each joint local coordinate system and the world coordinate system; if represents the homogeneous transformation matrix of the relative pose of link i with respect to link i-1, use the right multiplication method to continuously multiply 4 homogeneous transformation matrices, that is
[0016]
[0017] Among them, Rot represents the rotation matrix, Trans represents the translation matrix, x and z represent the x and z coordinate axes, θ i represents the z of the ith joint around the ith link i The rotation angle of the axis; d i represents the z of the ith joint along the ith link i Axis displacement, α i represents the torsion angle between the i-1th connecting rod and the i-th connecting rod; a i Represents the distance between the i-th joint and the i-1-th joint, that is, the connecting rod length.
[0018] Among them, in step 1, for a simpler mechanical structure, the posture relationship between the joints can be directly established based on the geometric relationship.
[0019] Wherein, in step 2, the Lagrangian method is used to establish the dynamic model of the n-link mechanism, and the expression is:
[0020]
[0021] in, q is the joint angular displacement vector, is the first-order derivative of q, q i is the angular displacement of connecting rod i, τ i is the torque of joint i;
[0022] The total kinetic energy and total potential energy of the n-link mechanism are and V(q); the difference between the total kinetic energy and the total potential energy is defined as the Lagrangian function Used to simplify the expression of formulas; since the total energy of the n-link mechanism is the sum of the energies of each joint, the total kinetic energy and total potential energy are expressed as:
[0023]
[0024]
[0025] Among them, T i 、V i are the kinetic energy and potential energy of link i and joint i respectively; ω i is the angular velocity of joint i, I i is the central inertia tensor of joint i, m i is the mass of connecting rod i, g is the acceleration due to gravity, r i is the coordinate of connecting rod i in the inertial coordinate system; For r i The first derivative of iy For ri The component on the y-axis of the inertial coordinate system; the superscript T in the upper right represents the transpose matrix;
[0026] The dynamic model of the n-link mechanism obtained from the Lagrange equation is arranged in the following form
[0027]
[0028] where, M(q) ∈ R n×n is the inertia matrix of the n-link mechanism, is the centrifugal force and Coriolis force matrix, G(q) ∈ R n×n is the gravity matrix, q ∈ R n×1 is the joint angular displacement vector, is the second derivative of q, τ ∈ R n×1 is the joint torque vector.
[0029] where, in the step 3, the design content of the sliding mode controller includes the design of the sliding mode surface, the design of the reaching law and the solution of the sliding mode controller.
[0030] where, in the step 3, first, the state space equation of the n-link mechanism model is expressed according to the transfer function of the n-link mechanism:
[0031]
[0032] where, x1 represents the difference between the desired joint angle and the actual joint angle, x2 is the first derivative of x1, is also the first derivative of x1, so and other parameters are deduced by analogy; A and B are coefficients that make the above expression hold;
[0033] Then the above n-link mechanism constitutes an n+1 order system, and at this time the sliding mode surface is designed as:
[0034]
[0035] where, K i is the design parameter of the sliding mode controller.
[0036] where, in the step 3, it is considered that at the equilibrium position, s(x) = 0; the purpose of the reaching law design is to make the designed sliding mode surface s(x) = 0, and thus, the designed reaching law is: constant speed reaching law, exponential reaching law or power reaching law;
[0037] Then, combining the state space equation, the sliding mode surface and the reaching law, u is expressed as a function of the desired joint angle and the actual joint angle, and their respective derivatives.
[0038] Among them, in the step 4, the method of reinforcement learning is adopted to select the parameters of the controller, so as to maintain the optimal or sub-optimal control effect of the system.
[0039] Among them, in the step 4, the algorithm of reinforcement learning is selected as Actor-Critic, Deep Q network or DPPG.
[0040] (III) Beneficial effects
[0041] Compared with the prior art, the beneficial effects produced by the method of the present invention are as follows:
[0042] (1) This method combines reinforcement learning with sliding mode control. In the control strategy, the method of reinforcement learning is used to select the parameters of the sliding mode controller, and the control strategy is flexibly adjusted for the changing environment to make optimal or sub-optimal motion behaviors, realizing environmental adaptability.
[0043] (2) Reinforcement learning is a parameter optimization idea, and various control methods all require parameter optimization. Therefore, this method is not limited to a single form of control method and can be applied to the parameter optimization of various forms of control such as robust control and PID control. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic diagram of a single-degree-of-freedom joint.
[0045] Figure 2 It is a schematic diagram of the design principle of a sliding mode controller.
[0046] Figure 3 It is a schematic diagram of the network structure of the DDPG method.
[0047] Figure 4a and Figure 4b They are respectively schematic diagrams of the network structures of the Actor network and the Critic network.
[0048] Figure 5 It is a schematic diagram of the training process flow. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To make the objectives, contents, and advantages of the present invention clearer, the following further describes in detail the specific embodiments of the present invention with reference to the drawings and embodiments.
[0050] To solve the above technical problems, the present invention provides a sliding mode control biped joint trajectory tracking optimization method based on reinforcement learning, aiming to propose a high-precision joint trajectory tracking method for adaptively controlling the parameters of the controller; the method includes the following steps:
[0051] Step 1: Based on the link structure simplified from the leg structure of the legged robot, with adjacent links connected by joints, establish the kinematic model of the n-link mechanism, so as to obtain the angular velocities of each joint and the end effector.
[0052] Step 2: Based on the angular velocities of each joint and the end effector obtained from the n-link mechanism, establish the dynamic model to obtain the relationship between the corresponding torque and motion parameters.
[0053] Step 3: Take the tracking error between the ideal trajectory and the actual trajectory of each joint as the control objective, and design a sliding mode controller.
[0054] Step 4: Based on the method of reinforcement learning, update the parameters of the sliding mode controller in real time to achieve adaptive sliding mode control.
[0055] Among them, in the above Step 1, establishing the kinematic model of the n-link mechanism is divided into trajectory planning and joint coordinate system transformation; for the establishment of trajectory planning, by giving the displacement-time function of the end spatial motion trajectory, and then through the kinematic relationship of each joint, calculate the variation law of each joint angle with time to obtain the trajectory planning.
[0056] Among them, in the above Step 1, for the joint coordinate system transformation, use the D-H method to establish the transformation relationship between each joint local coordinate system and the world coordinate system; if represents the homogeneous transformation matrix of the relative pose of link i with respect to link i-1, use the right multiplication method to continuously multiply 4 homogeneous transformation matrices, that is
[0057]
[0058] Among them, Rot represents the rotation matrix, Trans represents the translation matrix, x and z represent the x and z coordinate axes, θ i represents the rotation angle of the i-th joint around the z i axis of the i-th link; d i represents the displacement of the i-th joint along the z i axis of the i-th link, α i represents the twist angle between the (i-1)-th link and the i-th link; a i represents the distance between the i-th joint and the (i-1)-th joint, that is, the link length.
[0059] Among them, in the above Step 1, for a relatively simple mechanical structure, the pose relationship between joints can be directly established according to geometric relationships.
[0060] Among them, in the above Step 2, use the Lagrange method to establish the dynamic model of the n-link mechanism, and the expression is
[0061]
[0062] Among them, q is the joint angular displacement vector, is the first derivative of q, q i is the angular displacement of link i, τ i is the torque of joint i;
[0063] The total kinetic energy and total potential energy of the n-link mechanism are respectively and V(q); The difference between the total kinetic energy and the total potential energy is defined as the Lagrangian function for simplifying the formula expression form; Since the total energy of the n-link mechanism is the sum of the energies of each joint, then the total kinetic energy and the total potential energy are expressed as:
[0064]
[0065] Among them, T i , V i are respectively the kinetic energy and potential energy of link i and joint i; ω i is the angular velocity of joint i, I i is the central inertia tensor of joint i, m i is the mass of link i, g is the acceleration due to gravity, r i is the coordinate of link i in the inertial coordinate system; is the first derivative of r i , r iy is r i component on the y-axis of the inertial coordinate system; The superscript T in the upper right represents the transpose matrix;
[0066] The dynamic model of the n-link mechanism obtained from the Lagrangian equation is sorted into the following form
[0067]
[0068] Among them, M(q) ∈ R n×n is the inertia matrix of the n-link mechanism, is the centrifugal force and Coriolis force matrix, G(q) ∈ R n×n is the gravity matrix, q ∈ R n×1 is the joint angular displacement vector, is the second derivative of q, τ ∈ R n×1 is the joint torque vector.
[0069] Among them, in the said step 3, the design content of the sliding mode controller includes the design of the sliding mode surface, the design of the reaching law and the solution of the sliding mode controller.
[0070] Among them, in the said step 3, first, according to the transfer function of the n-link mechanism, the state space equation of the n-link mechanism model is expressed as:
[0071]
[0072] Among them, x1 represents the difference between the desired joint angle and the actual joint angle, and x2 is the first derivative of x1. It is also the first derivative of x1. Therefore The other parameters are deduced by analogy; A and B are coefficients that make the above expressions hold;
[0073] Then the above n-link mechanism constitutes an n+1 order system. At this time, the sliding mode surface is designed as:
[0074]
[0075] Among them, K i is the design parameter of the sliding mode controller.
[0076] Among them, in step 3, it is considered that at the equilibrium position, s(x)=0; the purpose of the reaching law design is to make the designed sliding mode surface s(x)=0. Therefore, the designed reaching law is: constant speed reaching law, exponential reaching law or power reaching law;
[0077] Then, combining the state space equation, the sliding mode surface and the reaching law, u is expressed as a function of the desired joint angle and the actual joint angle, and their respective derivatives.
[0078] Among them, in step 4, the method of reinforcement learning is used to select the parameters of the controller, so as to maintain the optimal or sub-optimal control effect of the system.
[0079] Among them, in step 4, the algorithm of reinforcement learning is selected as Actor-Critic, Deep Q network or DPPG.
[0080] Embodiment 1
[0081] This embodiment discloses a sliding mode control single-degree-of-freedom joint trajectory tracking optimization method based on reinforcement learning. What the embodiment describes is the motion tracking of a single-degree-of-freedom joint, not limited to a single-degree-of-freedom joint, not limited to sliding mode control, and can be extended to multi-degree-of-freedom joints and various control methods. The present invention will be further described in conjunction with the drawings and embodiments. The method includes the following steps:
[0082] Step (1): Figure 1 A schematic diagram of a single-degree-of-freedom joint is given. The a end is fixed and the b end can move freely. The a end and the b end are connected by a connecting rod. Based on this structure, the corresponding kinematic and dynamic models are established.
[0083] Assume that the length of the connecting rod is 2l. If the center of mass is considered to be at the center of the connecting rod, then at time t, the angle at the center of mass is q, and the real-time coordinates of the trajectory at the center of mass can be expressed as
[0084] x m = lcosθ, y m = lsinθ
[0085] Correspondingly, the linear velocity of the centroid can be expressed as
[0086]
[0087] If the mass of the connecting rod is m, then the kinetic energy and potential energy of the connecting rod are respectively expressed as
[0088]
[0089] Then, the difference between the total kinetic energy and the total potential energy of the single-degree-of-freedom joint can be expressed as
[0090]
[0091] According to Lagrange's equation, the dynamic model of the single-degree-of-freedom joint is expressed as
[0092]
[0093] Then, the driving torque can be further expressed as
[0094]
[0095] Rewrite the torque into the form of, and the corresponding parameters are obtained as follows
[0096] M = ml 2 , N = 0, G(θ) = mglcosθ.
[0097] Step (2): Take the tracking error between the ideal trajectory and the actual trajectory of the joint at the b end as the control objective, and design a sliding mode controller.
[0098] The trajectory tracking problem of the joint at the b end can be described as designing reasonable controller parameters and outputting a suitable control torque τ so that the actual joint angle and joint angular velocity can quickly track the desired joint angle and joint angular velocity, and keep the tracking error within a very small range.
[0099] 2.1) Since the parameters are up to the second derivative, the system composed of single-degree-of-freedom joints is of the second order, and the state space equation of the discrete system model is
[0100]
[0101] Then the tracking error can be expressed as
[0102]
[0103] Among them, and is the tracking signal, and x i is the current signal.
[0104] 2.2) At time k, the sliding surface is designed as
[0105] s(k) = Kx1(k) + x2(k)
[0106] If the control objective is the error, then the sliding surface can be written in the form of the error
[0107] s e (k) = Ke1(k) + e2(k)
[0108] Among them, K is the design parameter of the sliding mode controller.
[0109] Select the exponential reaching law as
[0110]
[0111] Then, the controller is expressed as
[0112]
[0113] Substitute it into the state equation expression, and get
[0114] u(k) = -sgn(s(k)) - s(k) - Ke1(k + 1) = -sgn(s(k)) - s(k) - K(x2 d (k + 1) - x2(k + 1))
[0115] Then, after calculating u(k) from the above formula, e i (k + 1) at time k + 1 can be obtained according to the state equation. The above process can be described as Figure 2 .
[0116] Step (3): Based on the DDPG reinforcement learning method, update the parameters of the sliding mode controller in real time to achieve adaptive sliding mode control.
[0117] The parameters of the sliding mode controller can be expressed as a(K1, K2,..., K j ), where j = 1. Assume that the difference between the ideal trajectory of each joint and the actual trajectory of each corresponding joint is the tracking error, and the control objective is to minimize the overall tracking error. Then this control problem can be expressed as
[0118]
[0119] Among them, is the desired joint angle, x i(a) is the actual joint angle. Link i is the i-th link, where i = 1.
[0120] It can be seen that the actual joint angle is affected by the controller parameters. If we want to reduce the tracking error of the joint angle, then the sliding mode controller parameters should be automatically adjusted according to this value to approximate the desired joint angle. Therefore, the method of reinforcement learning is introduced to train the agent. Here, the agent refers to a single-degree-of-freedom joint. Let the agent give reasonable control parameters according to the feedback of the environment, so as to improve the adaptability of the control system and maintain the optimal or sub-optimal control effect. The idea of reinforcement learning is to use the information interaction between the agent and the environment as the basis for updating the network weights, so that the reward is maximized when the agent reaches the final target action, that is, the optimal strategy is formulated. After the agent executes an action based on the current policy in the current state, the environment is updated to a new state and the corresponding reward value is fed back to the agent. The agent evaluates the quality of the current action according to the new state and the reward value, and then updates the corresponding policy. After multiple rounds of training, the agent finally completes the formulation of the optimal strategy.
[0121] 3.1) Determine the state vector and action vector according to the observed quantity and control quantity of the sliding mode controller.
[0122] As described in step (1), the angle change of the joint is continuous. For compliant movements, it means that the sliding mode controller parameters that determine the joint movement also change continuously. Therefore, a neural network is used to fit the relationship between the input vector and the control parameters. At the same time, a network also needs to be established for the strategy of the agent to select actions. Therefore, it can be said that deep reinforcement learning uses the advantage of neural networks to fit non-linear models to represent the continuous change laws of the agent, environment, state, reward, and action with neural networks, so as to serve as the strategy for decision-making behavior.
[0123] As can be seen from step (2), the system observes the movement of the joint, feeds the real-time movement parameters back to the sliding mode controller, and uses the error between the desired and actual movements as the input vector that needs to be controlled for a single-degree-of-freedom joint. Since the sliding mode controller directly outputs torque and the movement of a single-degree-of-freedom joint is derived from the dynamic model, the first network established should be the non-linear relationship between the movement quantity and torque of a single-degree-of-freedom joint. However, in step (2), the torque is the result of the control of the sliding mode controller. Therefore, in the first network, the state vector s(t) information should include the angles, angular velocities, angle tracking errors, and angular velocity tracking errors of each section, as well as the expected angles, angular velocities, and angular accelerations at the next moment, to reasonably represent the dynamic information of the single-degree-of-freedom joint's real-time and target trajectories; and the action vector a(t) is the control parameter of the sliding mode controller. In addition, the second network established should be used to evaluate the quality of the actions selected by the agent based on the policy.
[0124] 3.2) The network structure of reinforcement learning is constructed using the DDPG (Deep Deterministic Policy Gradient) method.
[0125] The DDPG algorithm includes both an Actor network for learning the relationship between states and actions, a Critic network for evaluating the Q-value of the current action, and two fixed networks, Actor_Target and Critic_Target. Among them, the input layer of the Actor network receives the current state vector s(t), and the output layer is the action vector a(t); the input layer of the Critic network receives the current state variable s(t) and the action vector a(t) passed in after being affected by the Actor network, and the output is the Q-value for evaluating the current action vector; the network structures of Actor and Actor_Target are the same, and the network structures of Critic and Critic_Target are the same. Moreover, Actor_Target and Critic_Target do not participate in the training. Their role is to be gradually updated according to the parameters of Actor and Critic, so as to calculate the Q_Target value and further calculate the TD_error value. This is because when updating the Critic network using the weighted gradient method, it is necessary to calculate the TD-error, which means that not only the Q-value corresponding to the state vector s(t) and the action vector a(t) is required, but also the Q-value of the predicted state vector s(t + 1).
[0126] The DPPG method uses the Actor_Target network to calculate a’(t + 1) at the s(t + 1) state and calculates the corresponding Q_Target value through the Critic_Target network. The benefit of this is that the Actor_Target network and the Critic_Target network with smoother parameter updates are used, which is equivalent to using a smaller step size when calculating the gradient and can be considered closer to the true gradient. Substitute the calculated Q_Target(s(t + 1), a’(t + 1)) value, the reward value r(t), and the reward discount rate γ into the Bellman equation to estimate the Q-value of state s(t + 1).
[0127] y = [r(t) + γ × Q_Target(s(t + 1), a'|θ i-1 )|s,a]
[0128] Then, the TD-error is calculated using the mean square error, and the expression is
[0129] TD-error = E s,a (y - q) 2
[0130] Among them, y is the Q(s(t+1)) value estimated by the Bellman equation, and q is the Q(s(t), a(t)) value of the current Critic network.
[0131] The weights of the Critic network update the network parameters θ using the weighted gradient method c , and the Actor's network updates θ using the method of gradient ascent a . The parameters θ of the Critic_Target network and the Actor_Target network ct and θ at are updated using the method of exponential moving average, and the expression is as follows
[0132]
[0133] Among them, ρ is the update rate, and θ a , θ at , θ c and θ ct represent the network parameters of Actor, Actor_Target, Critic, and Critic_Target respectively. The relationship between the four networks in DDPG is as Figure 3 shown.
[0134] The established Actor structure is as Figure 4a shown. The input of the network is the state vector s(t) of a single-degree-of-freedom joint. The middle two hidden layers both have 300 units, and the output layer is the parameter vector a(t) of the sliding mode controller. The established Critic network is as Figure 4b shown. The input includes the state vector s(t) and the action vector a(t). The middle two hidden layers have 400 units and 300 units respectively, and the output layer is the Q value of the action. In addition, to stably track the desired trajectory and keep the tracking error less than 0.02 rad for more than 25 time steps, an additional reward value can be obtained and set to 1.
[0135] 3.3) Testing.
[0136] During testing, given the desired motion trajectory of the b end, the corresponding variation law of the desired rotation angle of the single-degree-of-freedom joint with time is obtained based on the kinematic method. The desired rotation angle, angular velocity, and the real-time joint rotation angle and angular velocity together constitute the state vector s(t). With the help of the trained Actor network, the corresponding action vector a(t) is output, that is, the corresponding parameters of the sliding mode controller. Correspondingly, the sliding mode controller outputs the corresponding torque to complete the adjustment of the joint angle. The specific process is as Figure 5 shown.
[0137] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A sliding mode control biped joint trajectory tracking optimization method based on reinforcement learning, characterized in that The method includes the following steps: Step 1: According to the link structure simplified from the leg structure of the legged robot, with adjacent links connected by joints, establish the kinematic model of the n-link mechanism, so as to obtain the angular velocities of each joint and the end effector. Step 2: Based on the angular velocities of each joint and the end effector obtained from the n-link mechanism, establish the dynamic model to obtain the relationship between the corresponding torque and motion parameters. Step 3: Take the tracking error between the ideal trajectory of each joint and the actual trajectory of the corresponding joint as the control objective, and design a sliding mode controller. Step 4: Based on the method of reinforcement learning, update the parameters of the sliding mode controller in real time to achieve adaptive sliding mode control.
2. The optimized method for biped joint trajectory tracking based on reinforcement learning and sliding mode control according to claim 1, wherein In the said Step 1, establishing the kinematic model of the n-link mechanism is divided into trajectory planning and joint coordinate system transformation; for the establishment of trajectory planning, by giving the displacement-time function of the end spatial motion trajectory, and then through the kinematic relationship of each joint, calculate the variation law of each joint angle with time to obtain the trajectory planning.
3. The optimized biped joint trajectory tracking method based on reinforcement learning and sliding mode control according to claim 2, characterized in that In the above step 1, for the transformation of the joint coordinate system, the D-H method is used to establish the transformation relationships between the local coordinate systems of each joint and with the world coordinate system; if represents the homogeneous transformation matrix of the relative pose of link i with respect to link i-1, the four homogeneous transformation matrices are continuously multiplied using the right multiplication method, that is Among them, Rot represents the rotation matrix, Trans represents the translation matrix, x and z represent the x and z coordinate axes, and θ i represents the rotation angle of the i-th joint around the z-axis of the i-th link i ; d i represents the displacement of the i-th joint along the z-axis of the i-th link i ; α i represents the twist angle between the (i - 1)-th link and the i-th link; a i represents the distance between the i-th joint and the (i - 1)-th joint, that is, the link length.
4. The optimized method for biped joint trajectory tracking based on reinforcement learning and sliding mode control according to claim 3, characterized in that, In the said Step 1, for a relatively simple mechanical structure, the pose relationship between joints can be directly established according to geometric relationships.
5. The optimized method for biped joint trajectory tracking based on reinforcement learning and sliding mode control according to claim 3, characterized in that In the said Step 2, the Lagrangian method is used to establish the dynamic model of the n-link mechanism, and the expression is Among them, q is the joint angular displacement vector, is the first derivative of q, q i is the angular displacement of link i, τ i is the torque of joint i; The total kinetic energy and total potential energy of the n-link mechanism are respectively and V(q); the difference between the total kinetic energy and the total potential energy is defined as the Lagrangian function to simplify the formula expression; since the total energy of the n-link mechanism is the sum of the energies of each joint, the total kinetic energy and total potential energy are expressed as: where, T i and V i are the kinetic energy and potential energy of link i and joint i, respectively; ω i is the angular velocity of joint i, I i is the central inertia tensor of joint i, m i is the mass of link i, g is the acceleration due to gravity, and r i is the coordinate of link i in the inertial coordinate system; is the first derivative of r i , and r iy is the component of r i on the y-axis of the inertial coordinate system; the superscript T in the upper right represents the transpose matrix; The dynamic model of the n-link mechanism obtained from the Lagrangian equation is sorted into the following form where \(M(q)\in\mathbb{R}\) n×n is the inertia matrix of the \(n\)-link mechanism, is the centrifugal and Coriolis force matrix, \(G(q)\in\mathbb{R}\) n×n is the gravity matrix, \(q\in\mathbb{R}\) n×1 is the joint angular displacement vector, is the second derivative of \(q\), \(\tau\in\mathbb{R}\) n×1 is the joint torque vector.
6. The optimized biped joint trajectory tracking method based on reinforcement learning and sliding mode control according to claim 5, characterized in that In the said Step 3, the design content of the sliding mode controller includes the design of the sliding mode surface, the design of the reaching law and the solution of the sliding mode controller.
7. The optimized method for biped joint trajectory tracking based on reinforcement learning and sliding mode control according to claim 6, characterized in that In the said Step 3, first represent the state space equation of the n-link mechanism model according to the transfer function of the n-link mechanism: x = [x1, x2, …, x n+1 T , u = x n+2 ; where x1 represents the difference between the desired joint angle and the actual joint angle, and x2 is the first derivative of x1, which is also the first derivative of x1. Therefore, the other parameters follow the same analogy; A and B are coefficients that make the above expression hold; Then the above n-link mechanism constitutes an n+1 order system, and at this time the sliding mode surface is designed as: Among them, K i is the design parameter of the sliding mode controller.
8. The optimized method for biped joint trajectory tracking based on reinforcement learning and sliding mode control according to claim 7, wherein In the said Step 3, it is considered that at the equilibrium position, s(x)=0; the purpose of the reaching law design is to make the designed sliding mode surface s(x)=0. Thus, the designed reaching law is: constant speed reaching law, exponential reaching law or power reaching law; Then, combining the state space equation, the sliding mode surface and the reaching law, express u as a function of the desired joint angle and the actual joint angle, as well as their respective derivatives.
9. The optimized method for biped joint trajectory tracking based on reinforcement learning and sliding mode control according to claim 8, characterized in that In the said Step 4, the method of reinforcement learning is used to select the parameters of the controller, so as to maintain the optimal or sub-optimal control effect of the system.
10. The optimized biped joint trajectory tracking method based on reinforcement learning and sliding mode control according to claim 9, characterized in that, In the said Step 4, the algorithm of reinforcement learning is selected as Actor-Critic, Deep Q network or DPPG.