Reinforced learning motion control method and system for lower limbs of linear joint humanoid robot

The reinforcement learning-based motion control system for straight-line joint humanoid robots addresses the complexity of current control methods by utilizing linear actuator force control, enabling stable dynamic walking on diverse terrains.

CN120307298APending Publication Date: 2025-07-15HUBEI JINGCHU HUMANOID ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510665214.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, the motion control method of linear joint humanoid robots relies on complex modeling, resulting in poor anti-interference ability, limiting its application scenarios, especially its stable walking ability in complex ground environments.

Method used

The reinforcement learning algorithm is used combined with simulation algorithms to achieve stable dynamic walking control without complex modeling through the force control characteristics of linear electric cylinders. The actor-critician algorithm and partially observable Markov decision-making process optimization strategy are used, and the motion strategy of linear joint humanoid robots is generated by combining positive kinematics and closed-chain mechanism transformation.

Benefits of technology

The stable and dynamic walking of a linear joint humanoid robot on complex ground is realized, which improves its motion stability and robustness in complex environments, reduces costs and improves control accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120307298A_ABST
    Figure CN120307298A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of robot control, and discloses a reinforcement learning motion control method and system for lower limbs of a linear joint humanoid robot, and the control method comprises the steps: structurally separating the physical coupling constraint of a linear motor and a rotary joint in a training environment; a reinforcement learning method is adopted, an actor-commentator method and privilege observation variables are supplemented, and multiple agents are trained in parallel; a trained strategy is deployed to a real machine, input and output of the strategy obtained through training are rotating joint information, actual robot control and collected data are linear joint motor side information, and the difference between simulation and reality can be overcome through the mapping relation of force, position and speed. Compared with a traditional position control method, the force control characteristic of the linear joint is used, and the higher adaptability is achieved; and meanwhile, a reinforcement learning method is used for training the linear joint humanoid robot, accurate modeling is not needed, and a global and long-term motility simulation algorithm stage and a deployment algorithm stage can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot control, and particularly relates to a reinforcement learning motion control method and system for the lower limbs of a linear joint humanoid robot. Background Technique

[0002] A humanoid robot is based on the human bionic prototype and has an innate bionic advantage of legged locomotion compared with robots of other motion modes. At present, the configurations of humanoid robots can be divided into rotary joint drive and linear joint drive. The linear electric cylinder has advantages such as large load and high endurance due to its internal planetary roller screw structure. However, the current motion control method for linear joint humanoid robots is a model-based method, that is, physical parameters such as the center of mass of the robot are obtained through complex and precise modeling to maintain balanced motion. The modeling of the whole-body dynamics is complex and time-consuming. At the same time, the linear electric cylinder is for position control and has poor anti-interference ability, which greatly limits its application scenarios. Summary of the Invention

[0003] Aiming at the above defects or improvement requirements of the prior art, the present invention provides a reinforcement learning motion control method and system for the lower limbs of a linear joint humanoid robot, which can utilize the force control characteristics of the linear electric cylinder to realize the reinforcement learning motion control of the linear joint humanoid robot, so that the linear joint humanoid robot can achieve stable dynamic walking without complex modeling and can adapt to various complex terrains at the same time.

[0004] To achieve the above object, according to one aspect of the present invention, a reinforcement learning motion control method for the lower limbs of a linear joint humanoid robot is provided, including a simulation algorithm stage and a deployment algorithm stage:

[0005] S1 Simulation algorithm stage, including the following steps:

[0006] S11: The linear joint humanoid robot model is decoupled by a physical closed-chain method to obtain a rotary joint humanoid robot model and imported into the simulation environment for training;

[0007] S12: Use the reinforcement learning algorithm to train the rotary joint humanoid robot model obtained in step S11 to obtain a reinforcement learning model, establish a partially observable Markov decision process, use the proximal policy optimization algorithm to update the policy gradient and optimize the policy, and adopt the actor-critic algorithm to guide the generation of the critic and actor policy networks according to the asymmetry of the observation variables in the simulation environment and the real world;

[0008] S2 Deployment algorithm stage, including the following steps:

[0009] S21: Deploy the actor policy network obtained in step S12 to the physical machine controller. The input layer of this network receives the position, speed, torque, attitude sensor angle, and angular velocity data of the previous action of the rotating joint. After passing through the hidden layer network to the output layer, it outputs the next action, that is, the position increment of the rotating joint. Finally, joint torque data is generated through a PD controller.

[0010] S22: Input the joint torque data obtained in step S21 into the rotating joint torque data to the rotating motor controller; for the joint torque data obtained in step S21, obtain the motor torque through forward kinematics and the thrust-torque transformation formula and input it to the linear motor controller.

[0011] S23: On the basis of step S22, the rotating motor controller directly outputs the position, speed, and torque data of the rotating joint. The position and speed data of the linear motor controller generate the position and speed data of the rotating joint through forward kinematics. The torque data generates the rotating joint torque through torque-thrust transformation and closed-chain mechanism forward dynamics.

[0012] S24: Concatenate the rotating joint data, attitude sensor data, etc. generated in S23 into an observation vector and input it into the policy network again to achieve reinforcement learning closed-loop control, thereby formulating a complex environment dynamic motion strategy for the linear joint humanoid robot.

[0013] Preferably, in step S11, the physical decoupling closed-chain method is to cancel the physical constraints between the linear transmission joint and the rotating joint. There is no coupled motion between the two parts. At the same time, the linear transmission joint is no longer used as a power component, and the rotating joint becomes the power component.

[0014] Preferably, the partially observable Markov formula in step S12 is:

[0015] M = <S, A, T, O, R, γ>

[0016] Among them, S and A respectively represent the full state space and the rotating joint action space; T represents the state transition dynamic function, specifically expressed as T(s’ / s, a), and its meaning is the probability of reaching s’ after executing action a in state s; R is the reward function, specifically expressed as R(s, a), and its meaning is the reward value of executing action a in state s; γ ∈ [0, 1] is the discount factor; O is the partially observable state space.

[0017] By solving the partially observable Markov process, that is, generating a policy π(a|o≤t) from the partially observable state space to the rotating joint action space, maximizing the expected return J = E[R] = E[∑ t γ t r t ;

[0018] In step S12, the specific implementation process of the actor-critic algorithm includes: the actor policy function π θ (O), which is implemented by a multi-layer perceptron. The input is the partially observable state space O, and the output is the next action of the robot; the critic value function V π (s), which is implemented by a multi-layer perceptron, estimates the value function of the current policy, and evaluates the quality of the actor policy network.

[0019] Preferably, step S22 specifically includes: for the torque data of the rotary joint, directly input the torque data of the rotary joint into the rotary motor controller;

[0020] For the torque data of the linear transmission joint, after obtaining the adjustment data through the inverse dynamics of the closed-chain mechanism and the thrust-torque transformation equation, input it into the linear motor controller;

[0021] At the same time, for the motor position, speed, and torque data obtained by the linear motor controller, it needs to go through forward kinematics, the torque-thrust transformation equation, and the forward dynamics of the closed-chain mechanism before being input into the policy network.

[0022] Preferably, the equation of the inverse dynamics of the closed-chain mechanism is:

[0023] F = ID(τ 关节 )

[0024] where F represents the thrust at the end of the linear cylinder push rod, and τ 关节 represents the torque on the rotary joint side of the linear transmission closed-chain mechanism.

[0025] The thrust-torque transformation equation is:

[0026] τ 电机 = f(F)

[0027] where F represents the thrust at the end of the linear cylinder push rod, and τ 电机 represents the torque on the linear cylinder motor side.

[0028] Preferably, the position and speed data of the linear transmission motor need to go through forward kinematics, and its equation is:

[0029] [r, w] = FK(p, v)

[0030] where r and w are the angle and angular velocity on the rotary joint side of the linear transmission closed-chain mechanism respectively, and p and v are the feedback position and feedback speed on the motor side of the linear transmission closed-chain mechanism respectively;

[0031] The torque data needs to go through the torque-thrust transformation equation and the forward dynamics of the closed-chain mechanism. The torque-thrust transformation equation is:

[0032] F = f(τ 电机 )

[0033] Among them, F represents the thrust at the end of the linear electric cylinder push rod, and τ 电机 represents the torque on the motor side of the linear electric cylinder.

[0034] The forward dynamics equation of the closed-loop mechanism is:

[0035] τ 关节 = FD(F)

[0036] Among them, F represents the thrust at the end of the linear electric cylinder push rod, and τ 关节 represents the torque on the side of the rotating joint of the linear transmission closed-loop mechanism.

[0037] Preferably, the thrust-torque transformation and torque-thrust transformation need to be realized by calibrating the torque on the motor side of the linear electric cylinder and the thrust at the end of the linear electric cylinder push rod. The specific method includes the following steps:

[0038] Step S31: Set a series of torque command values sent to the motor side of the linear electric cylinder, and send the torque command values to the driver on the motor side of the linear electric cylinder at a period.

[0039] Step S32: Obtain the corresponding thrust magnitude of the linear electric cylinder push rod.

[0040] Step S33: Establish a linear regression equation between the torque command value on the motor side of the linear electric cylinder and the thrust of the linear electric cylinder push rod, and establish a motor torque-thrust model. Its equation is:

[0041] F = f(τ 电机 )

[0042] τ 电机 = f(F)

[0043] Among them, τ 电机 is the torque value on the motor side of the linear electric cylinder, and F is the thrust value of the linear electric cylinder push rod.

[0044] To achieve the above object, according to another aspect of the present invention, there is provided a reinforcement learning motion control system for the lower limbs of a linear joint humanoid robot, including:

[0045] An acquisition module for acquiring the state of the motor side of the linear electric cylinder, including current position, speed, and torque information;

[0046] A data processing module for converting the state data of the motor side into the state data of the rotating joint through the torque-thrust relationship;

[0047] A calculation module for acquiring the state data of the rotating joint and outputting the action data of the rotating joint;

[0048] A data processing module for converting the action data of the rotating joint into the command data of the motor side;

[0049] A control module for sending command data to a motor driver at a certain period.

[0050] Preferably, it further includes a processor, a bus, and a memory. The processor includes a central processing unit and a network processor; the bus includes an address bus, a data bus, and a control bus; the memory is a type of semiconductor memory.

[0051] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the following beneficial effects are achieved:

[0052] The present invention proposes a method for calibrating the torque on the motor side of a linear cylinder and the output force at the end of the push rod of the linear cylinder. This method can avoid installing a force sensor at the end of the linear cylinder, reduce costs, and at the same time has higher force control accuracy;

[0053] The present invention utilizes the force control characteristics of the linear cylinder to realize the reinforcement learning motion control of the lower limbs of a linear joint humanoid robot. Compared with the traditional position control scheme, the linear joint humanoid robot can achieve stable dynamic walking without complex modeling, and at the same time can adapt to various complex terrains, enhancing the stability and robustness of the lower limb motion system of the linear joint humanoid robot. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a structural diagram of the lower limbs of a linear joint humanoid robot provided by an embodiment of the present invention;

[0055] Figure 2 It is a flowchart of the reinforcement learning motion control system of the lower limbs of a linear joint humanoid robot provided by an embodiment of the present invention;

[0056] Figure 3 It is a schematic diagram of the closed-chain structure of the hip Pitch joint, knee joint, and two-degree-of-freedom ankle joint provided by an embodiment of the present invention;

[0057] Figure 4 It is a calibration flowchart of the torque on the motor side of the linear cylinder and the thrust on the push rod side of the linear cylinder provided by an embodiment of the present invention.

[0058] Figure 5 It is a structural block diagram of the control system of a linear joint humanoid robot provided by an embodiment of the present invention;

[0059] Figure 6 It is a schematic diagram of the structure of an electronic device in the control system provided by an embodiment of the present invention.

[0060] In the accompanying drawings: 11 - Hip Roll joint, 12 - Hip Yaw joint, 13 - Hip Pitch joint, 14 - Knee joint, 15 - Ankle joint 1, 16 - Ankle joint 2, 17 - Linear electric cylinder, 18 - Pusher side, 19 - Motor side, 50 - Control system, 51 - Acquisition module, 52 - Data processing module, 53 - Calculation module, 54 - Data processing module, 55 - Control module, 60 - Processor, 61 - Bus, 62 - Processor. Detailed implementation manner

[0061] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0062] Please refer to Figures 1-6 , a reinforcement learning motion control method and system for the lower limbs of a linear joint humanoid robot provided by the present invention, the method includes a simulation algorithm stage and a deployment algorithm stage:

[0063] S1 Simulation algorithm stage, including the following steps:

[0064] S11: The linear joint humanoid robot model is obtained through a physical decoupling closed-loop method to obtain a rotary joint humanoid robot model and imported into the simulation environment for training;

[0065] Among them, the physical decoupling closed-loop method is to cancel the physical constraints between the linear drive joint and the rotary joint, there is no coupled motion between the two parts, and at the same time, the linear drive joint is no longer used as a power component, and the rotary joint becomes a power component.

[0066] S12: Use the reinforcement learning algorithm to train the rotary joint humanoid robot model obtained in step S11 to obtain a reinforcement learning model, establish a partially observable Markov decision process (POMDP), use the proximal policy optimization algorithm (PPO) to update the policy gradient and optimize the policy, and use the actor-critic algorithm (Actor-Critic) to guide the generation of critic and actor policy networks according to the asymmetry of the observed variables in the simulation environment and the real world.

[0067] Among them, the partially observable Markov formula is:

[0068] M = <S, A, T, O, R, β>

[0069] Among them, S and A represent the full state space and the rotational joint action space respectively; T represents the state transition dynamic function, specifically expressed as T(s’ / s,a), which means the probability of reaching s’ after executing action a in state s; R is the reward function, specifically expressed as R(s,a), which means the reward value of executing action a in state s; γ∈[0,1] is the discount factor; O is the partially observable state space;

[0070] Solve the partially observable Markov process, that is, generate a policy π(a|o≤t) from the partially observable state space to the rotational joint action space, and maximize the expected return J = E[R] = E[∑ t γ t r t ;

[0071] In step S12, the specific implementation process of the Actor-Critic algorithm includes: the Actor policy function π θ (O), implemented using a multi-layer perceptron (MLP), takes the partially observable state space O as input and outputs the next action of the robot; the Critic value function V π (s), implemented using a multi-layer perceptron (MLP), estimates the value function of the current policy and evaluates the quality of the Actor policy network.

[0072] In the S2 deployment algorithm stage, it includes the following steps:

[0073] S21: Deploy the Actor policy network obtained in step S12 to the real machine controller. The input layer of this network receives data such as the position, speed, torque, attitude sensor angle, and angular velocity of the previous action of the rotational joint, passes through the hidden layer network to the output layer to output the next action, that is, the rotational joint position increment, and finally generates joint torque data through the PD controller;

[0074] The output layer of the network in this step is action, that is, the next action, which can be used as a position increment, speed increment, or torque increment, as well as any control parameter of the robot. When used as a speed increment or torque increment, the system control frequency requirement is higher, and using it as a position increment is the best choice.

[0075] S22: Input the joint torque data obtained in step S21 into the rotational joint torque data to the rotational motor controller; the joint torque data obtained in step S21 is used to obtain the motor torque through the forward kinematics and thrust-torque transformation formula and input it to the linear motor controller;

[0076] Step S22 specifically includes: for the torque data of the rotational joint, directly input the rotational joint torque data to the rotational motor controller;

[0077] For the torque data of the linear drive joint, after obtaining the adjustment data through the inverse dynamics of the closed-chain mechanism and the thrust-torque transformation equation, it is input into the linear motor controller;

[0078] Meanwhile, for the motor position, speed, and torque data obtained by the linear motor controller, they need to go through forward kinematics, the torque-thrust transformation equation, and the forward dynamics of the closed-chain mechanism before being input into the policy network.

[0079] The equation of the inverse dynamics of the closed-chain mechanism is:

[0080] F = ID(τ 关节 )

[0081] where F represents the thrust at the end of the push rod of the linear cylinder, and τ 关节 represents the torque on the rotary joint side of the linear drive closed-chain mechanism.

[0082] The thrust-torque transformation equation is:

[0083] τ 电机 = f(F)

[0084] where F represents the thrust at the end of the push rod of the linear cylinder, and τ 电机 represents the torque on the motor side of the linear cylinder. The position and speed data of the linear drive motor need to go through forward kinematics, and its equation is:

[0085] [r, w] = FK(p, v)

[0086] where r and w are the angle and angular velocity on the rotary joint side of the linear drive closed-chain mechanism respectively, and p and v are the feedback position and feedback speed on the motor side of the linear drive closed-chain mechanism respectively.

[0087] The torque data needs to go through the torque-thrust transformation equation and the forward dynamics of the closed-chain mechanism. The torque-thrust transformation equation is:

[0088] F = f(τ 电机 )

[0089] where F represents the thrust at the end of the push rod of the linear cylinder, and τ 电机 represents the torque on the motor side of the linear cylinder.

[0090] The forward dynamics equation of the closed-chain mechanism is:

[0091] τ 关节 = FD(F)

[0092] where F represents the thrust at the end of the push rod of the linear cylinder, and τ 关节 represents the torque on the rotary joint side of the linear drive closed-chain mechanism.

[0093] The thrust-moment transformation and moment-thrust transformation need to be achieved through the calibration of the moment on the motor side of the linear cylinder and the thrust at the end of the push rod of the linear cylinder. The specific method includes the following steps:

[0094] Step S31: Set a series of moment command values sent to the motor side of the linear cylinder, and send the moment command values to the driver on the motor side of the linear cylinder at regular intervals.

[0095] Step S32: Obtain the magnitude of the thrust of the push rod of the corresponding linear cylinder.

[0096] Step S33: Establish a linear regression equation between the moment command value on the motor side of the linear cylinder and the thrust of the push rod of the linear cylinder, and establish a motor moment-push rod thrust model. The equation is:

[0097] F = f(τ 电机 )

[0098] τ 电机 = f(F)

[0099] where τ 电机 is the moment value on the motor side of the linear cylinder, and F is the thrust value of the push rod of the linear cylinder.

[0100] S23: On the basis of S22, the rotation motor controller directly outputs the rotation joint position, speed, and moment data. The position and speed data of the linear motor controller generate the rotation joint position and speed data through forward kinematics, and the moment data generates the rotation joint moment through moment-thrust transformation and the forward dynamics of the closed-chain mechanism.

[0101] S24: Concatenate the rotation joint data, attitude sensor (IMU) data, etc. generated in S23 into an observation vector, and input it into the policy network again to achieve closed-loop reinforcement learning control, so as to formulate a complex environment dynamic motion strategy for the linear joint humanoid robot.

[0102] Specifically, as Figures 1-6 shown:

[0103] Figure 1 FIG. is a schematic diagram of the lower limb structure of a linear joint humanoid robot provided by an embodiment of the present invention. It is only an exemplary lower limb of a linear joint humanoid robot. The embodiments of the present application can be applied to the control system of any humanoid robot with a linear cylinder in the lower limb joint. The embodiments of the present application do not limit the specific application platform.

[0104] In Figure 1Among them, the lower limb joints of the straight-joint humanoid robot include the hip Roll joint 11, the hip Yaw joint 12, the hip Pitch joint 13, the knee joint 14, the ankle joint one 15, and the ankle joint two 16. Among them, the hip Roll joint 11 and the hip Yaw joint 12 are driven by rotary joint motors, and the remaining joints are linearly driven by a linear electric cylinder 17.

[0105] In Figure 1 the linear electric cylinder 17 has a motor side 19 and a push rod side 18 at both ends respectively. The motor side 19 generates a rotary motion to drive the push rod side 18 to perform a reciprocating linear motion.

[0106] Figure 2 is a flowchart of the reinforcement learning motion control method for the lower limbs of the straight-joint humanoid robot provided by the embodiment of the present invention, including a simulation algorithm stage and a deployment algorithm stage.

[0107] In the simulation algorithm stage, a physical decoupling closed-chain structure is used to obtain a rotary-joint humanoid robot model and import it into the simulation environment for training. Specifically as follows:

[0108] The lower limbs of the straight-joint humanoid robot adopted in the embodiment of the present application are as Figure 1 shown, where the closed-chain mechanism includes the hip Pitch joint 13, the knee joint 14, the ankle joint one 15, and the ankle joint two 16. For clearer representation, these three joints are drawn separately, as Figure 3 shown in (a), (b), and (c) in. The linear transmission parts a31, b31, and c31 are the power components at the joints of the actual robot, and the rotary joints a33, b33, and c33 are the passive traction motion components, and a32, b32, and c32 are the connection components between the linear transmission parts and the rotary joints.

[0109] The physical decoupling closed-chain mechanism method, that is, cancel the physical constraints between the linear transmission parts a31, b31, c31 and the rotary joints a33, b33, c33, there is no coupled motion between the two parts. At the same time, the linear transmission parts a31, b31, c31 are no longer used as power components, but are one of the conventional rigid components, and the rotary joints a33, b33, c33 become power components. At the same time, the hip Yaw joint 12 and the hip Pitch joint 13 remain unchanged.

[0110] In the simulation environment, the motion joints of the rotary-joint humanoid robot model are all rotary joint parts.

[0111] The rotating joint humanoid robot model is trained using a reinforcement learning algorithm. The reinforcement learning model is M = <S, A, T, O, R, γ>, where S and A represent the state space and action space of the rotating joint part 34 respectively. T represents the state transition dynamic function, specifically expressed as T(s’ / s, a), which means the probability of reaching s’ after executing action a in state s. R is the reward function, specifically expressed as R(s, a), which means the reward value of executing action a in state s. γ ∈ [0, 1] is the discount factor. O is the partially observable state space. By solving the partially observable Markov process, the proximal policy optimization algorithm is used to update the policy gradient and optimize the policy, and the Actor-Critic algorithm is used to guide the generation of the actor policy. The input of this network is the observed variables such as the current position and current speed of the rotating joint part, and the output is the action of the rotating joint, that is, to generate the policy π(a|o≤t) from the partially observable state space to the rotating joint action space, maximizing the expected return J = E[R] = E[∑ t γ t r t ;

[0112] This actor policy network (Actor) is what is needed in the deployment algorithm stage.

[0113] Then comes the deployment algorithm stage. The actor policy network is deployed to the real machine controller. The policy network receives the position, speed, and torque data of the rotating joint, outputs the position increment of the rotating joint, and generates the torque data of the rotating joint through a PD controller. For the rotating electric joint, the torque data of the rotating joint can be directly input into the motor controller. For the linear transmission joint, the motor torque input to the motor controller needs to be obtained through the inverse dynamics of the closed-chain mechanism and the thrust-torque transformation equation. Similarly, the motor position, speed, and torque data obtained by the linear transmission joint need to go through forward kinematics, torque-thrust transformation equation, and forward dynamics of the closed-chain mechanism before being input into the policy neural network. Specifically as follows:

[0114] The policy neural network is the actor policy network in the simulation algorithm stage, which receives control commands, IMU data, the position, speed, and torque data of the rotating joint part, etc., and outputs the torque data of the rotating joint. For the rotating joints 11 and 12, the torque of the rotating joint part can be directly written into the motor controller. For the linear transmission closed-chain mechanisms 13, 14, 15, and 16, it needs to go through the inverse dynamics of the closed-chain mechanism F = ID(τ 关节 ) and the thrust-torque transformation τ 电机 = f(F) algorithms to be written into the motor controller.

[0115] Similarly, for the linear drive closed-loop mechanisms 13, 14, 15, and 16, the motor position and speed need to go through the forward kinematics [r, w] = FK(p, v), and the torque needs to go through the torque-thrust transformation F = f(τ 电机 ) and the forward dynamics of the closed-loop mechanism τ 关节 = FD(F) before it can be input into the policy network.

[0116] The thrust-torque transformation and the torque-thrust transformation need to be achieved through the calibration of the torque on the motor side of the linear cylinder and the thrust at the end of the push rod of the linear cylinder. Refer to Figure 4 , including the following steps:

[0117] Step S31: Set a series of torque command values sent to the motor side 19 of the linear cylinder, and send the torque command values to the drive on the motor side of the linear cylinder at regular intervals.

[0118] Step S32: Obtain the corresponding thrust magnitude of the push rod 18 of the linear cylinder.

[0119] Step S33: Establish the linear regression equations τ 电机 = f(F) and F = f(τ 电机 ) between the torque command value on the motor side 19 of the linear cylinder and the thrust of the push rod 18 of the linear cylinder.

[0120] In the exemplary embodiment of the present invention, the input is the torque command value on the motor side 19 of the linear cylinder, and the collected signals are the torque value of the drive on the motor side 19 of the linear cylinder and the thrust magnitude of the push rod 18 of the linear cylinder. τ is the torque value on the motor side 19 of the linear cylinder, and F is the thrust value of the push rod 18 of the linear cylinder. Through the two conversion equations τ 电机 = f(F) and F = f(τ 电机 ), the conversion between the torque on the motor side 19 of the linear cylinder and the thrust of the push rod 18 of the linear cylinder can be achieved.

[0121] The schematic diagrams of the closed-loop mechanisms are shown in (a), (b), and (c) of Figure 3 , which represent the hip Pitch joint, the knee joint, and the two-degree-of-freedom ankle joint respectively. Among them, the motor side 31 and the push rod side 32 correspond one-to-one with Figure 1 the motor side 19 and the push rod side 18 in

[0122] In the closed-loop mechanism, the actively controlled object is the push rod side 32 of the linear drive part, and the passively moving part is the rotary joint 33. The forward and inverse kinematics and dynamics of the closed-loop mechanism establish the force-position-velocity mapping relationship between the push rod side 32 and the rotary joint 33.

[0123] The forward kinematics and dynamics are from the push rod side 32 of the linear drive part to the rotary joint 33, and the inverse kinematics and dynamics are the opposite.

[0124] Further, establish the forward kinematics and solve it. Through geometric relationships, the forward kinematics equation can be established as: [r, w] = FK(p, v);

[0125] Rewrite the equation to satisfy the iterative solution format, [r, w] - FK(p, v) = 0; Set the initial values of the angle r and angular velocity w, the maximum number of iterations, and the solution accuracy for numerical iterative solution; Through the forward kinematics, the equivalent conversion relationship between the angular velocity and angle of the rotary joint 33 and the linear velocity and linear displacement of the push rod side 32 of the linear transmission part can be obtained.

[0126] Further, establish and solve the forward and inverse dynamics. The forward dynamics and inverse dynamics are established as follows: τ 关节 = FD(F) and F = ID(τ 关节 ); where, the solution method of the forward dynamics equation τ 关节 = FD(F) is the same as that of the forward kinematics, which will not be elaborated here. The inverse dynamics can directly substitute the joint torque into the equation to obtain the thrust at the push rod 32 of the linear electric cylinder.

[0127] In summary, through the various embodiments of the present invention, the reinforcement learning training and deployment of the lower limbs of the linear joint humanoid robot can be realized, which can effectively save the training cost and improve the dynamic motion ability of the linear joint humanoid robot in complex environments at the same time.

[0128] Refer to Figure 5 , which is the structural block diagram of the control system 50 of the linear joint humanoid robot provided by the present invention. The control system 50 of the lower limbs of the linear joint humanoid robot includes:

[0129] An acquisition module 51, configured to acquire the states of the motor side 31 of the linear electric cylinder, including information such as the current position, speed, and torque.

[0130] A data processing module 52, configured to convert the state data of the motor side 31 into the state data of the rotary joint 33 through the torque-thrust relationship;

[0131] A calculation module 53, which acquires the state data of the rotary joint 33 and outputs the action data of the rotary joint 33;

[0132] A data processing module 54, configured to convert the action data of the rotary joint 33 into the command data of the motor side 31;

[0133] A control module 55, configured to send command data to the motor driver at a certain period.

[0134] The control system of the linear joint humanoid robot provided by the present invention can implement the control method of the above-mentioned linear joint humanoid robot. For details, refer to the above, which will not be elaborated here.

[0135] In addition, in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. They are merely used to distinguish different operations, and the serial numbers themselves do not represent any order of execution.

[0136] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention.

[0137] This control system further includes a processor 62, a bus 61, and a memory 60. The processor 62 can be an integrated circuit chip with signal processing capabilities. The processor 62 can also be a general-purpose processor, including a central processing unit, a network processor, etc. The bus 61 can be an address bus, a data bus, a control bus, etc. The memory 60 can be one of semiconductor memories.

[0138] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A reinforcement learning motion control method for the lower limbs of a straight-joint humanoid robot, characterized in that, It includes a simulation algorithm stage and a deployment algorithm stage: S1 Simulation algorithm stage, including the following steps: S11: The linear joint humanoid robot model is transformed into a rotary joint humanoid robot model through a physical decoupling closed-chain method and imported into the simulation environment for training; S12: Use the reinforcement learning algorithm to train the rotary joint humanoid robot model obtained in step S11 to obtain a reinforcement learning model, establish a partially observable Markov decision process, use the proximal policy optimization algorithm to update the policy gradient and optimize the policy, and adopt the actor-critic algorithm to guide the generation of the critic and actor policy networks according to the asymmetry of the observation variables in the simulation environment and the real world; S2 Deployment algorithm stage, including the following steps: S21: Deploy the actor policy network obtained in step S12 to the real machine controller. The input layer of this network receives the position, speed, torque, attitude sensor angle, and angular velocity data of the previous action of the rotary joint. After passing through the hidden layer network to the output layer, it outputs the next action, that is, the rotary joint position increment. Finally, the joint torque data is generated through the PD controller; S22: Input the joint torque data obtained in step S21 into the rotary joint torque data to the rotary motor controller; for the joint torque data obtained in step S21, obtain the motor torque through the forward kinematics and thrust-torque transformation formula and input it to the linear motor controller; S23: On the basis of step S22, the rotary motor controller directly outputs the rotary joint position, speed, and torque data. The position and speed data of the linear motor controller generate the rotary joint position and speed data after forward kinematics, and the torque data generates the rotary joint torque through the torque-thrust transformation and the forward dynamics of the closed-chain mechanism; S24: Concatenate the rotary joint data and attitude sensor data generated in S23 into an observation vector, and input it into the policy network again to achieve reinforcement learning closed-loop control, thereby formulating the dynamic motion strategy of the linear joint humanoid robot in a complex environment.

2. The reinforcement learning motion control method for the lower limbs of a linear joint humanoid robot according to claim 1, characterized in that, In step S11, the physical decoupling closed-chain method is to cancel the physical constraints between the linear transmission joint and the rotary joint. There is no coupled movement between the two parts. At the same time, the linear transmission joint is no longer used as a power component, and the rotary joint becomes the power component.

3. The reinforcement learning motion control method for the lower limbs of a straight-joint humanoid robot according to claim 1, characterized in that, The partially observable Markov formula in step S12 is: M = <S, A, T, O, R, γ> Among them, S and A represent the full state space and the rotary joint action space respectively; T represents the state transition dynamic function, specifically expressed as T(s’ / s, a), whose meaning is the probability of reaching s’ after executing action a in state s; R is the reward function, specifically expressed as R(s, a), whose meaning is the reward value of executing action a in state s; γ ∈ [0, 1] is the discount factor; O is the partially observable state space; By solving the partially observable Markov process, that is, generating a policy π(a|o≤t) from the partially observable state space to the rotational joint action space, maximizing the expected return J = E[R] = E[∑ t γ t r t ; In step S12, the specific implementation process of the actor-critic algorithm includes: the actor policy function π θ (O), which is implemented using a multi-layer perceptron. The input is the partially observable state space O, and the output is the next action of the robot; the critic value function V π (s), which is implemented using a multi-layer perceptron, estimates the value function of the current policy, and evaluates the quality of the actor policy network.

4. The reinforcement learning motion control method for the lower limbs of a linear joint humanoid robot according to claim 1, characterized in that, Step S22 specifically includes: For the torque data of the rotary joint, directly input the rotary joint torque data to the rotary motor controller; For the torque data of the linear transmission joint, obtain the adjustment data through the inverse dynamics of the closed-chain mechanism and the thrust-torque transformation equation and input it to the linear motor controller; Meanwhile, the motor position, speed, and torque data obtained by the linear motor controller need to go through forward kinematics, torque-thrust transformation equations, and forward dynamics of the closed-loop mechanism before being input into the policy network.

5. The reinforcement learning motion control method for the lower limbs of a straight-joint humanoid robot according to claim 4, wherein The equation for the inverse dynamics of the closed-loop mechanism is: F = ID(τ 关节 ) Among them, F represents the thrust at the end of the push rod of the linear electric cylinder, and τ 关节 represents the torque on the rotary joint side of the linear transmission closed-chain mechanism; The thrust-torque transformation equation is: τ 电机 = f(F) Among them, F represents the thrust at the end of the push rod of the linear electric cylinder, and τ 电机 represents the torque on the motor side of the linear electric cylinder.

6. The reinforcement learning motion control method for the lower limbs of a linear joint humanoid robot according to claim 4, characterized in that, The position and speed data of the linear transmission motor need to go through forward kinematics, and its equation is: [r, w] = FK(p, v) where r and w are the angles and angular velocities on the rotating joint side of the linear transmission closed-loop mechanism, respectively, and p and v are the feedback positions and feedback speeds on the motor side of the linear transmission closed-loop mechanism; The torque data needs to go through the torque-thrust transformation equation and the forward dynamics of the closed-loop mechanism. The torque-thrust transformation equation is: F = f(τ 电机 ) Among them, F represents the thrust at the end of the push rod of the linear electric cylinder, and τ 电机 represents the torque on the motor side of the linear electric cylinder. The forward dynamics equation of the closed-loop mechanism is: τ 关节 = FD(F) Among them, F represents the thrust at the end of the push rod of the linear electric cylinder, and τ 关节 represents the torque on the rotating joint side of the linear transmission closed chain mechanism.

7. A reinforcement learning motion control method for the lower limbs of a linear joint humanoid robot according to any one of claims 4-6, characterized in that, The thrust-torque transformation and torque-thrust transformation need to be realized through the calibration of the torque on the motor side of the linear cylinder and the thrust at the end of the linear cylinder push rod. The specific method includes the following steps: Step S31: Set a series of torque command values sent to the motor side of the linear cylinder, and send the torque command values to the driver on the motor side of the linear cylinder at a certain period. Step S32: Obtain the magnitude of the thrust of the linear cylinder push rod. Step S33: Establish a linear regression equation between the torque command value on the motor side of the linear cylinder and the thrust of the linear cylinder push rod, and establish a motor torque-push rod thrust model. Its equation is: F = f(τ 电机 ) τ 电机 = f(F) Among them, τ 电机 is the torque value on the motor side of the linear electric cylinder, and F is the thrust value of the push rod of the linear electric cylinder.

8. The reinforcement learning motion control system for the lower limbs of a straight-joint humanoid robot according to claim 7, characterized in that, Including: An acquisition module for acquiring the state of the motor side of the linear cylinder, including current position, speed, and torque information; A data processing module for converting the state data of the motor side into the state data of the rotating joint through the torque-thrust relationship; A calculation module for obtaining the state data of the rotating joint and outputting the action data of the rotating joint; A data processing module for converting the action data of the rotating joint into the command data of the motor side; A control module for sending command data to the motor driver at a certain period.

9. The reinforcement learning motion control system for the lower limbs of a linear joint humanoid robot according to claim 8, characterized in that, It also includes a processor, a bus, and a memory.

10. The reinforcement learning motion control system for the lower limbs of a linear joint humanoid robot according to claim 9, characterized in that, The processor includes a central processing unit and a network processor; The bus includes an address bus, a data bus, and a control bus; The memory is one of the semiconductor memories.