Reinforcement Learning Training Method, Device, Equipment and Storage Medium for Humanoid Robots
By analyzing robot models to establish inverse kinematics and dynamics for parallel kinematic ankle joints, the method addresses low efficiency and high sampling issues in personification robot training, enhancing computational efficiency and motion planning accuracy.
Patent Information
- Application Number
- CN202510493067.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-18
AI Technical Summary
In the prior art, the humanoid robot control algorithm has low computational efficiency and high parallel sampling efficiency during training, which leads to the long time for neural networks to find the optimal solution in the action space, making it difficult to efficiently carry out reinforcement learning.
By obtaining the model file of the humanoid robot, analyzing its positional relationship and component constraint information, establishing a benchmark coordinate system, building a parallel ankle joint mechanism model, determining inverse kinematics information, and conducting reinforcement learning strategy training based on this information, avoiding serial and parallel conversion, improving computing efficiency and reducing parallel sampling requirements.
It improves computing efficiency, simplifies the training process, accelerates the convergence speed of reinforcement learning strategies, improves the accuracy of motion planning, and reduces unstable factors in training.
Smart Images

Figure CN120023832B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of robot control, and more particularly, to a method, device, equipment and storage medium for training a humanoid robot using reinforcement learning. Background Art
[0002] A humanoid robot is an extremely complex form of robot. Generally, a humanoid robot has a complex bipedal and dual-arm mechanical structure, which includes numerous serial and parallel mechanisms. This complexity makes the control problem of a humanoid robot a typical non-linear, multi-variable and time-varying control problem. Traditional model-based control algorithms usually require accurate modeling of the robot and its operating environment, resulting in a highly complex model. In addition, when these models are generalized to different tasks and environments, they usually need to be readjusted. Therefore, more and more researchers are adopting deep reinforcement learning methods to solve the control problem of humanoid robots.
[0003] The humanoid robot control algorithms in the prior art usually adopt an open-loop topology during the training process and complete the serial-parallel conversion during the stage of migrating from the simulation environment to the real world (Simulation to Reality, abbreviated as sim2real).
[0004] However, this training method makes the neural network spend more time finding the optimal solution in the action space, resulting in a problem of low computing efficiency. At the same time, it also has high requirements for the parallel sampling efficiency of the reinforcement learning of the humanoid robot. Summary of the Invention
[0005] The purpose of the present application is to provide a method, device, equipment and storage medium for training a humanoid robot using reinforcement learning to solve the problems of low computing efficiency and high requirements for parallel sampling efficiency in the prior art.
[0006] To achieve the above object, the technical solutions adopted in the embodiments of the present application are as follows:
[0007] In a first aspect, an embodiment of the present application provides a method for training a humanoid robot using reinforcement learning, the method including:
[0008] Obtaining a model file of the humanoid robot, the model file being at least used to indicate the physical structure, sensor information and kinematic parameters of the humanoid robot;
[0009] Parsing the model file to obtain the positional relationship between the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establishing a reference coordinate system according to the positional relationship between the two feet;
[0010] Based on the reference coordinate system and the component constraint information, a parallel ankle joint mechanism model of the humanoid robot is constructed, and inverse kinematic information of the humanoid robot is determined according to the parallel ankle joint mechanism model, where the inverse kinematic information includes position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot;
[0011] Based on the inverse kinematic information and the component constraint information, the reinforcement learning strategy of the humanoid robot is trained.
[0012] Optionally, determining the inverse kinematic information according to the parallel ankle joint mechanism model includes:
[0013] Determining the position inverse solution information according to the parallel ankle joint mechanism model;
[0014] Determining the velocity inverse solution information according to the position inverse solution information;
[0015] Determining the acceleration inverse solution information according to the velocity inverse solution information.
[0016] Optionally, determining the position inverse solution information according to the parallel ankle joint mechanism model includes:
[0017] Determining the attitude of the foot plate, the length of the active arm, the length of the passive rod, the geometric offset, and the geometric constraint conditions according to the parallel ankle joint mechanism model;
[0018] Inputting the attitude of the foot plate, the length of the active arm, the length of the passive rod, and the geometric offset into a pre-constructed vector loop equation to solve for at least one initial value of the first active arm angle and at least one initial value of the second active arm angle;
[0019] Determining a target value of the first active arm angle and a target value of the second active arm angle according to each of the initial values of the first active arm angle, each of the initial values of the second active arm angle, and the geometric constraint conditions, and using the target value of the first active arm angle and the target value of the second active arm angle as the position inverse solution information.
[0020] Optionally, determining the velocity inverse solution information according to the position inverse solution information includes:
[0021] Determining an active arm vector and a passive rod vector according to the target value of the first active arm angle and the target value of the second active arm angle;
[0022] Calculating a rotation matrix of the end point according to the active arm vector and the passive rod vector;
[0023] Constructing a Jacobian matrix according to the active arm vector, the passive rod vector, and the rotation matrix of the end point;
[0024] Obtain the speed of the foot plate;
[0025] Calculate the product of the speed of the foot plate and the Jacobian matrix to obtain the angular velocity of the first active arm and the angular velocity of the second active arm, and use the Jacobian matrix, the angular velocity of the first active arm, and the angular velocity of the second active arm as the inverse kinematic information of the speed.
[0026] Optionally, determining the inverse acceleration information according to the inverse kinematic information of the speed includes:
[0027] Determine the acceleration of the foot plate;
[0028] Calculate the time derivative of the Jacobian matrix;
[0029] According to the Jacobian matrix, the speed of the foot plate, the acceleration of the foot plate, and the time derivative of the Jacobian matrix, calculate the angular acceleration of the first active arm and the angular acceleration of the second active arm, and use the angular acceleration of the first active arm and the angular acceleration of the second active arm as the inverse acceleration information.
[0030] Optionally, training the reinforcement learning strategy of the humanoid robot according to the inverse kinematic information and the component constraint information includes:
[0031] Initialize the humanoid robot model and the reinforcement learning environment. The humanoid robot model is a humanoid robot with a serial model, and the type of the reinforcement learning environment is a serial model;
[0032] In the current training round, obtain the current state from the reinforcement learning environment. The current state includes joint angles, speeds, and accelerations;
[0033] According to the current reinforcement learning strategy, the inverse position information, and the current state, determine the target parallel action. The target parallel action is used to indicate the target action when the humanoid robot model has a parallel ankle structure;
[0034] According to the target parallel action, the current state, the inverse kinematic information, and the component constraint information, determine the serial expected torque, and apply the serial expected torque to the reinforcement learning environment, and update the reinforcement learning strategy to train the reinforcement learning strategy of the humanoid robot.
[0035] Optionally, determining the serial expected torque according to the target parallel action, the current state, the inverse kinematic information, and the component constraint information includes:
[0036] Calculate a base torque based on the current angle in the current state and the angle of the target parallel motion;
[0037] Calculate a compensation torque based on the inverse kinematic information and the component constraint information;
[0038] Superimpose the base torque and the compensation torque to obtain the desired parallel torque;
[0039] Calculate a transposed Jacobian matrix according to the Jacobian matrix;
[0040] Calculate the product of the transposed Jacobian matrix and the desired parallel torque to obtain the desired serial torque.
[0041] In a second aspect, another embodiment of the present application provides a humanoid robot reinforcement learning training device, and the device includes:
[0042] An acquisition module, configured to acquire a model file of the humanoid robot, and the model file is at least used to indicate the physical structure, sensor information, and kinematic parameters of the humanoid robot;
[0043] An analysis module, configured to analyze the model file to obtain the positional relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establish a reference coordinate system according to the positional relationship of the two feet;
[0044] A determination module, configured to construct a parallel ankle joint mechanism model of the humanoid robot according to the reference coordinate system and the component constraint information, and determine inverse kinematic information according to the parallel ankle joint mechanism model, where the inverse kinematic information includes position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot;
[0045] A training module, configured to train the reinforcement learning strategy of the humanoid robot according to the inverse kinematic information and the component constraint information.
[0046] In a third aspect, another embodiment of the present application provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the method according to any one of the first aspects above.
[0047] In a fourth aspect, another embodiment of the present application provides a storage medium, on which a computer program is stored, and when the computer program is run by a processor, it performs the steps of the method according to any one of the first aspects above.
[0048] The beneficial effects of the present application are as follows: By obtaining the model file of the humanoid robot, parsing the model file to obtain the positional relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establishing a reference coordinate system based on the positional relationship of the two feet, it is possible to construct a parallel ankle joint mechanism model of the humanoid robot according to the reference coordinate system and the component constraint information, and determine the inverse kinematics information according to the parallel ankle joint mechanism model. Thus, it is possible to train the reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information, avoiding the series-parallel conversion through the calculation method during the training process, improving the operation efficiency. At the same time, it also reduces the requirement for the parallel sampling efficiency of reinforcement learning, simplifies the training process, and speeds up the convergence speed of the reinforcement learning strategy. In addition, training the reinforcement learning strategy based on the parallel ankle joint model and the component constraint information can also improve the accuracy of motion planning. Optimizing the strategy training through the inverse kinematics information and the constraint information can also reduce the unstable factors in the training. Brief Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.
[0050] Figure 1 A flowchart of a method for training the reinforcement learning of a humanoid robot provided by an embodiment of the present application;
[0051] Figure 2 A schematic diagram of a parallel ankle joint mechanism model provided by an embodiment of the present application;
[0052] Figure 3 A flowchart of a method for determining inverse kinematics information in the method for training the reinforcement learning of a humanoid robot provided by an embodiment of the present application;
[0053] Figure 4 A flowchart of a method for determining the position inverse solution information in the method for training the reinforcement learning of a humanoid robot provided by an embodiment of the present application;
[0054] Figure 5 A flowchart of a method for determining the velocity inverse solution information in the method for training the reinforcement learning of a humanoid robot provided by an embodiment of the present application;
[0055] Figure 6 A flowchart of a method for determining the acceleration inverse solution information in the method for training the reinforcement learning of a humanoid robot provided by an embodiment of the present application;
[0056] Figure 7 It is a schematic flowchart for training the reinforcement learning policy of a humanoid robot in the humanoid robot reinforcement learning training method provided by an embodiment of the present application;
[0057] Figure 8 It is a schematic flowchart for determining the series expected torque in the humanoid robot reinforcement learning training method provided by an embodiment of the present application;
[0058] Figure 9 It is a schematic diagram of a humanoid robot reinforcement learning training device provided by an embodiment of the present application;
[0059] Figure 10 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application only serve the purposes of illustration and description, and are not used to limit the protection scope of the present application. Additionally, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. Furthermore, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0061] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present application.
[0062] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated thereafter, but does not exclude adding other features.
[0063] The humanoid robot control algorithms in the prior art usually adopt an open-loop topology structure during the training process and complete the series-parallel conversion during the stage of migrating from the simulation environment to the real world (Simulation to Reality, abbreviated as sim2real).
[0064] However, this training method makes the neural network spend more time searching for the optimal solution in the action space, resulting in a problem of low computing efficiency. At the same time, it also has high requirements for the parallel sampling efficiency of the reinforcement learning of humanoid robots.
[0065] Based on the above problems, the embodiments of the present application propose a method for training the reinforcement learning of humanoid robots. By obtaining the model file of the humanoid robot and parsing the model file, the positional relationship of the feet of the humanoid robot and the component constraint information of the humanoid robot are obtained. And based on the positional relationship of the feet, a reference coordinate system is established. Then, according to the reference coordinate system and the component constraint information, a parallel ankle joint mechanism model of the humanoid robot can be constructed, and the inverse kinematics information can be determined according to the parallel ankle joint mechanism model. Thus, according to the inverse kinematics information and the component constraint information, the reinforcement learning strategy of the humanoid robot can be trained, avoiding the series-parallel conversion through the calculation method during the training process, improving the computing efficiency. At the same time, it also reduces the requirements for the parallel sampling efficiency of the reinforcement learning, simplifies the training process, and speeds up the convergence rate of the reinforcement learning strategy.
[0066] The following describes in detail the method for training the reinforcement learning of humanoid robots provided by the embodiments of the present application in combination with multiple embodiments.
[0067] Figure 1 FIG. is a schematic flowchart of a method for training the reinforcement learning of humanoid robots provided by an embodiment of the present application. Referring to Figure 1 as shown, the execution subject of this method can be any electronic device with processing capabilities. This method includes:
[0068] S101. Obtain the model file of the humanoid robot.
[0069] Among them, the model file is at least used to indicate the physical structure, sensor information, and kinematic parameters of the humanoid robot.
[0070] Optionally, the model file further includes visual appearance and collision model information.
[0071] Exemplarily, the physical structure may include the rigid components of a humanoid robot, such as: the chassis, sensors, joints of the robotic arm, geometric connection methods between linkages, the relationship between joints and motors, etc. The sensor information may include the basic information of the sensors on the humanoid robot (specifically including type, identifier, manufacturer, model, and version), physical attributes (specifically including size, shape, installation location, Euler angle position, electrical interface, operating temperature, waterproof level, and electromagnetic interference resistance), performance parameters (specifically including measurement range, resolution, accuracy, sampling rate, field of view angle, and sensitivity), communication interfaces (specifically including communication protocol, data format, data synchronization mechanism, bus configuration), calibration parameters (specifically including internal parameter matrix, external field matrix, offset compensation, and calibration timestamp), noise and error information (specifically including noise characteristics, systematic error, interference model), data output configuration (specifically including output mode, supported data compression algorithms, supported data filtering algorithms), driver dependency configuration (specifically including driver program, dependencies, configuration file path), and other information (specifically including cost, lifespan, certification, and user manual link, etc.).
[0072] Exemplarily, the kinematic parameters are used to indicate the joint motion relationship and motion limitations of the humanoid robot, such as joint variable range, speed and acceleration limitations, definition of singular configurations, geometric and inertial parameters, homogeneous transformation matrix, self-motion parameters, and joint coupling relationship, etc.
[0073] Exemplarily, the visual appearance includes color, texture, etc., and the collision model information includes mesh model, elastic coefficient, collision mask, hierarchy, trigger area, and component attribution.
[0074] By obtaining the model file of the humanoid robot, it is possible to ensure that the behavior of the humanoid robot conforms to the physical laws.
[0075] S102. Parse the model file to obtain the position relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establish a reference coordinate system based on the position relationship of the two feet.
[0076] Optionally, the model file can be parsed by a parsing tool to extract the position relationship of the two feet of the humanoid robot. Specifically, the parsing tool can be the urdf_parser of ROS or the SDFormat library. The position relationship of the two feet of the humanoid robot is used to indicate the linkage and joint information of the left and right feet of the humanoid robot.
[0077] Optionally, the kinematic parameters of each component, the geometric information and inertial information in the physical structure are read from the model file, and based on the kinematic parameters of each component, the geometric information and inertial information in the physical structure, the mass matrix, the Coriolis force and centrifugal force matrix, and the gravity term are calculated in combination with the dynamic model as the component constraint information.
[0078] Specifically, the mass matrix can be obtained from the mass and inertia tensor of the link, the Coriolis force and centrifugal force matrix can be obtained from the mass matrix and joint velocity, and the gravity term can be obtained from the instruction of the link, the position of the center of mass, and the gravitational acceleration.
[0079] Exemplarily, the model file can be parsed to obtain the left foot position, the right foot position, and the joint constraints, and the center point can be calculated to establish a reference coordinate system.
[0080] S103. According to the reference coordinate system and the component constraint information, a parallel ankle joint mechanism model of the humanoid robot is constructed, and the inverse kinematic information is determined according to the parallel ankle joint mechanism model.
[0081] Optionally, after obtaining the reference coordinate system, the attitude of the ankle joint end relative to the calf can be defined , that is, the attitude of the foot plate, where and respectively represent the rotation angles around the X-axis and Y-axis, and the rotation matrices RX( ) and RY( ) are used to calculate the position o′Pi of the parallel link hinge point Pi. Specifically, it can be calculated with reference to the following formula:
[0082] o′Pi = RY(θ)RX( ) o′Pi′
[0083] where o′Pi′ is the coordinate of the hinge point at the initial position.
[0084] Exemplarily, Figure 2 is a schematic diagram of a parallel ankle joint mechanism model provided by an embodiment of the present application. Referring to Figure 2 shown, after calculating the position o′Pi of the parallel link hinge point Pi, the length of the active arm can be determined as L1, and the length of the passive rod can be determined as L2 to obtain the parallel ankle joint mechanism model.
[0085] Among them, Pi is the hinged point of the parallel link corresponding to the i-th parallel joint in the parallel ankle joint mechanism model. Depending on the different structures of the humanoid robot, there can be different numbers of parallel joints in the parallel ankle joint mechanism model of the humanoid robot. Ai is the i-th parallel node in the Z-axis direction in the parallel ankle joint mechanism model, Bi is the i-th parallel node in the YO'Z plane in the parallel ankle joint mechanism model, and Ci is the i-th parallel node in the XYZ space in the parallel ankle joint mechanism model.
[0086] Optionally, after obtaining the parallel ankle joint mechanism model, the inverse kinematic information can be calculated based on the geometric relationships in the parallel ankle joint mechanism model.
[0087] Among them, the inverse kinematic information includes the position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot.
[0088] Exemplarily, the position inverse solution information includes the angles of the active arms of the humanoid robot, the velocity inverse solution information includes the angular velocities of the active arms of the humanoid robot, and the acceleration inverse solution information includes the angular accelerations of the active arms of the humanoid robot.
[0089] Through the angles of the active arms of the humanoid robot, the arm postures and the positions of the end effectors of the humanoid robot can be determined, thereby ensuring the accuracy of motion planning and inverse kinematic calculations. Through the angular velocities of the active arms of the humanoid robot, the arm movement speeds of the humanoid robot can be determined, thereby ensuring the accuracy of dynamic motion control and balance adjustment. Through the angular accelerations of the active arms of the humanoid robot, the accelerations of the arm movements can be determined, thereby ensuring the accuracy of dynamic modeling and force control.
[0090] S104. Train the reinforcement learning strategy of the humanoid robot according to the inverse kinematic information and the component constraint information.
[0091] Optionally, after obtaining the inverse kinematic information, the reinforcement learning strategy of the humanoid robot can be adjusted based on the inverse kinematic information and the component constraint information, and trained based on the adjusted reinforcement learning strategy.
[0092] In this embodiment, by obtaining the model file of the humanoid robot, parsing the model file to obtain the positional relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establishing a reference coordinate system according to the positional relationship of the two feet, it is possible to construct a parallel ankle mechanism model of the humanoid robot based on the reference coordinate system and the component constraint information, and determine the inverse kinematic information according to the parallel ankle mechanism model. Thus, it is possible to train the reinforcement learning strategy of the humanoid robot based on the inverse kinematic information and the component constraint information, avoiding the series-parallel conversion through the calculation method during the training process, improving the operation efficiency. At the same time, it also reduces the requirement for the parallel sampling efficiency of reinforcement learning, simplifies the training process, and accelerates the convergence speed of the reinforcement learning strategy. In addition, training the reinforcement learning strategy based on the parallel ankle model and the component constraint information can also improve the accuracy of motion planning, and optimizing the strategy training through the inverse kinematic information and the constraint information can also reduce the unstable factors in the training.
[0093] As a possible implementation Figure 3 is a schematic flow chart when determining the inverse kinematic information in the humanoid robot reinforcement learning training method provided by the embodiment of the present application. Referring to Figure 3 as shown, in step S103, determining the inverse kinematic information according to the parallel ankle mechanism model includes:
[0094] S301. Determine the inverse position solution information according to the parallel ankle mechanism model.
[0095] Optionally, geometric analysis can be performed on the parallel ankle mechanism model and solved to obtain the angle of the active arm of the humanoid robot as the inverse position solution information.
[0096] Exemplarily, the vector loop method can be used for solving, and the homogeneous coordinate transformation method or the screw theory method can also be used for solving, or equations can be directly established through geometric analysis for solving.
[0097] S302. Determine the inverse velocity solution information according to the inverse position solution information.
[0098] Optionally, after obtaining the inverse position solution information, the inverse velocity solution information can be calculated according to the inverse position solution information.
[0099] Exemplarily, the inverse velocity solution information can be calculated by the multi-joint Jacobian matrix method or the numerical integration method. Among them, the inverse velocity solution information can be the angular velocity of the active arm of the humanoid robot.
[0100] S303. Determine the inverse acceleration solution information according to the inverse velocity solution information.
[0101] Optionally, after obtaining the inverse kinematic velocity information, the inverse kinematic acceleration information can be calculated based on the inverse kinematic velocity information.
[0102] Exemplarily, the angular acceleration of the active arm of the humanoid robot can be calculated by any one of the Jacobian matrix derivative method, the Hessian matrix method, and the Lagrange equation method for the angular velocity of the active arm of the humanoid robot.
[0103] Through the parallel ankle mechanism model, first determine the inverse kinematic position information, then determine the inverse kinematic velocity information based on the inverse kinematic position information, and then determine the inverse kinematic acceleration information based on the inverse kinematic velocity information, so as to gradually obtain the inverse kinematic information of the humanoid robot in a hierarchical manner, enabling accurate positioning of the end effector through inverse kinematic position, optimizing dynamic response through inverse kinematic velocity, eliminating shocks through inverse kinematic acceleration, realizing smooth movement of the humanoid robot, avoiding singular configurations, reducing computational complexity, reducing energy consumption, and ensuring stability and real-time performance.
[0104] As a possible implementation Figure 4 FIG. is a schematic flowchart of a process for determining inverse kinematic position information in the humanoid robot reinforcement learning training method provided by an embodiment of the present application. Refer to Figure 4 As shown, according to the parallel ankle mechanism model, determining the inverse kinematic position information includes:
[0105] S401. According to the parallel ankle mechanism model, determine the posture of the foot plate, the length of the active arm, the length of the passive rod, the geometric offset, and the geometric constraint conditions.
[0106] Optionally, the kinematics of the parallel ankle joint can be determined from the parallel ankle mechanism model to obtain the posture χ of the foot plate, the length L1 of the active arm, the length L2 of the passive rod, the geometric offset r1, the hinge point Pi of the parallel link, and the geometric constraint conditions. Among them, the geometric constraint conditions can be the rod length constraint conditions and the joint angle range of the humanoid robot.
[0107] S402. Input the posture of the foot plate, the length of the active arm, the length of the passive rod, and the geometric offset into a pre-constructed vector loop equation, and solve to obtain at least one initial value of the first active arm angle and at least one initial value of the second active arm angle.
[0108] Optionally, the first active arm vector loop equation and the second active arm vector loop equation can be established in advance through geometric constraints, and the attitude χ of the foot plate, the length L1 of the active arm, the length L2 of the passive rod, the geometric offset r1, and the hinge point Pi of the parallel link are respectively input into the first active arm vector loop equation and the second active arm vector loop equation, and the first active arm vector loop equation and the second active arm vector loop equation are solved to obtain at least one initial value of the first active arm angle and at least one initial value of the second active arm angle.
[0109] Exemplarily, numerical methods such as the Newton iteration method can be used to solve the first active arm vector loop equation and the second active arm vector loop equation to obtain at least one initial value of the first active arm angle and at least one initial value of the second active arm angle.
[0110] S403. Determine the target value of the first active arm angle and the target value of the second active arm angle according to each initial value of the first active arm angle, each initial value of the second active arm angle, and the geometric constraint conditions, and use the target value of the first active arm angle and the target value of the second active arm angle as the inverse kinematic solution information.
[0111] Optionally, after obtaining at least one initial value of the first active arm angle and at least one initial value of the second active arm angle, the target value of the first active arm angle and the target value of the second active arm angle can be screened from each initial value of the first active arm angle and each initial value of the second active arm angle according to the joint angle range and the rod length constraint, and the target value of the first active arm angle and the target value of the second active arm angle are used as the inverse kinematic solution information.
[0112] By using the parallel ankle joint mechanism model, the attitude of the foot plate, the length of the active arm, the length of the passive rod, the geometric offset, and the geometric constraint conditions are determined, and the attitude of the foot plate, the length of the active arm, the length of the passive rod, and the geometric offset are input into the pre-constructed vector loop equation. At least one initial value of the first active arm angle and at least one initial value of the second active arm angle are solved, and according to each initial value of the first active arm angle, each initial value of the second active arm angle, and the geometric constraint conditions, the target value of the first active arm angle and the target value of the second active arm angle are determined, and the target value of the first active arm angle and the target value of the second active arm angle are used as the inverse kinematic solution information. It can directly model based on the geometric relationship of the parallel ankle joint mechanism model, avoid kinematic approximation errors, ensure that the calculation results strictly meet the actual physical constraints, and can also improve the motion efficiency of the humanoid robot. In addition, it helps to improve the fault tolerance ability of the humanoid robot to physical deviations and enhance control stability.
[0113] As a possible implementation method Figure 5This is a schematic flow diagram for determining velocity inverse solution information in the humanoid robot reinforcement learning training method provided by the embodiments of this application. Refer to Figure 5 As shown in
[0114] S501. Determine the active arm vector and the passive rod vector according to the first active arm angle target value and the second active arm angle target value.
[0115] Optionally, according to the first active arm angle target value and the second active arm angle target value in the position inverse solution information, determine the active arm vector BiCi and the passive rod vector CiPi.
[0116] S502. Calculate the rotation matrix of the end point according to the active arm vector and the passive rod vector.
[0117] Optionally, calculate the rotation matrix of the end point Pi in the reference coordinate system according to the active arm vector BiCi and the passive rod vector CiPi.
[0118] S503. Construct the Jacobian matrix according to the active arm vector, the passive rod vector, and the rotation matrix of the end point.
[0119] Optionally, vector operations can be performed according to the active arm vector BiCi, the passive rod vector CiPi, and the rotation matrix of the end point Pi to construct the Jacobian matrix.
[0120] Exemplarily, taking the i-th row of the Jacobian matrix as an example, the y component of the vector cross product of the active arm vector BiCi and the passive rod vector CiPi can be calculated, and the transposed vector of the passive rod vector CiPi can be calculated, and the ratio of the transposed vector of the passive rod vector CiPi to the y component can be calculated, and the product of the ratio of the transposed vector of the passive rod vector CiPi to the y component and the rotation matrix of the end point Pi in the reference coordinate system can be calculated to obtain the i-th row of the Jacobian matrix. By analogy, the Jacobian matrix can be obtained.
[0121] S504. Obtain the velocity of the foot plate.
[0122] Optionally, the velocity of the first foot plate and the velocity of the second foot plate, that is, the velocity of the end effector, can be obtained.
[0123] S505. Calculate the product of the velocity of the foot plate and the Jacobian matrix to obtain the first active arm angular velocity and the second active arm angular velocity, and use the Jacobian matrix, the first active arm angular velocity, and the second active arm angular velocity as the velocity inverse solution information.
[0124] Optionally, the product of the velocity of the first foot plate and the Jacobian matrix, and the product of the velocity of the second foot plate and the Jacobian matrix can be calculated respectively through matrix multiplication to obtain the angular velocity of the first active arm and the angular velocity of the second active arm, and the Jacobian matrix, the angular velocity of the first active arm, and the angular velocity of the second active arm are used as inverse kinematic velocity information.
[0125] Based on the target value of the angle of the first active arm and the target value of the angle of the second active arm, the active arm vector and the passive rod vector are determined. According to the active arm vector and the passive rod vector, the rotation matrix of the end point is calculated. According to the active arm vector, the passive rod vector, and the rotation matrix of the end point, the Jacobian matrix is constructed. The product of the velocity of the foot plate and the Jacobian matrix is calculated to obtain the angular velocity of the first active arm and the angular velocity of the second active arm. The Jacobian matrix, the angular velocity of the first active arm, and the angular velocity of the second active arm are used as inverse kinematic velocity information, which can avoid complex iterative calculations, adapt to high-frequency control loops, ensure real-time dynamic tracking, and also avoid step-by-step approximation errors, especially suitable for the rapid direction-changing requirements of high-stiffness ankle joints. In addition, by redistributing the velocity or correcting the trajectory, the singular region can be avoided in advance to prevent the mechanism from getting out of control due to sudden changes in joint velocity. In addition, an optimal balance can be achieved among accuracy, efficiency, and controllability.
[0126] As a possible implementation Figure 6 is a schematic flowchart of a process for determining inverse kinematic acceleration information in the humanoid robot reinforcement learning training method provided by an embodiment of the present application. Refer to Figure 6 shown, the above S303 determines inverse kinematic acceleration information according to inverse kinematic velocity information, including:
[0127] S601. Determine the acceleration of the foot plate.
[0128] Optionally, the acceleration of the first foot plate and the acceleration of the second foot plate, that is, the acceleration of the end effector, can be obtained from the sensor side.
[0129] Optionally, the velocity of the first foot plate and the velocity of the second foot plate can also be differentiated with respect to time to obtain the acceleration of the first foot plate and the acceleration of the second foot plate.
[0130] S602. Calculate the time derivative of the Jacobian matrix.
[0131] Optionally, the time derivative of the Jacobian matrix can be calculated through the chain rule.
[0132] Exemplarily, the partial derivative of the Jacobian matrix with respect to the joint angle is calculated, and the partial derivative is multiplied by the joint velocity to obtain the time derivative of the Jacobian matrix.
[0133] S603. Calculate the angular acceleration of the first active arm and the angular acceleration of the second active arm based on the Jacobian matrix, the velocity of the foot plate, the acceleration of the foot plate, and the time derivative of the Jacobian matrix, and use the angular acceleration of the first active arm and the angular acceleration of the second active arm as the inverse acceleration solution information.
[0134] Optionally, calculate the angular acceleration of the first active arm based on the Jacobian matrix, the velocity of the first foot plate, the acceleration of the first foot plate, and the time derivative of the Jacobian matrix.
[0135] Optionally, calculate the angular acceleration of the second active arm based on the Jacobian matrix, the velocity of the second foot plate, the acceleration of the second foot plate, and the time derivative of the Jacobian matrix.
[0136] Exemplarily, calculate the product of the Jacobian matrix and the acceleration of the first foot plate as the first product, calculate the product of the time derivative of the Jacobian matrix and the velocity of the first foot plate as the second product, and calculate the sum of the first product and the second product to obtain the angular acceleration of the first active arm.
[0137] Exemplarily, calculate the product of the Jacobian matrix and the acceleration of the second foot plate as the third product, calculate the product of the time derivative of the Jacobian matrix and the velocity of the second foot plate as the fourth product, and calculate the sum of the third product and the fourth product to obtain the angular acceleration of the second active arm.
[0138] By determining the acceleration of the foot plate, calculating the time derivative of the Jacobian matrix, and calculating the angular acceleration of the first active arm and the angular acceleration of the second active arm based on the Jacobian matrix, the velocity of the foot plate, the acceleration of the foot plate, and the time derivative of the Jacobian matrix, and using the angular acceleration of the first active arm and the angular acceleration of the second active arm as the inverse acceleration solution information, it is possible to eliminate the dynamic error in high-speed motion, achieve accurate tracking of the end acceleration, avoid mechanical vibration caused by sudden changes in joint velocity, and extend the fatigue life of the mechanism.
[0139] As a possible implementation manner, Figure 7 This is a schematic flowchart of a process for training the reinforcement learning strategy of a humanoid robot in the humanoid robot reinforcement learning training method provided by an embodiment of the present application. Refer to Figure 7 As shown, the above S104 trains the reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information, including:
[0140] S701. Initialize the humanoid robot model and the reinforcement learning environment.
[0141] Among them, the humanoid robot model is a humanoid robot with a serial model, and the type of the reinforcement learning environment is a serial model.
[0142] It can be understood that the serial model means that during the training process of the reinforcement learning policy, the humanoid robot is modeled as an open-loop series structure model, rather than the actual closed-loop parallel structure (such as the parallel mechanism of the ankle). That is, the robot joints are connected in series, and the movement of each joint only depends on the state of its previous joint, so as to simplify the dynamic calculation and make the training process more efficient.
[0143] S702. In the current training round, obtain the current state from the reinforcement learning environment.
[0144] The current state includes joint angles, velocities, and accelerations.
[0145] Optionally, determine whether the reinforcement learning policy converges. If it does not converge, in the current training round, obtain the current state from the reinforcement learning environment. The current state includes joint angles, velocities, and accelerations.
[0146] S703. Determine the target parallel action according to the current reinforcement learning policy, position inverse kinematics information, and the current state.
[0147] Optionally, the target parallel action, that is, the parallel mechanism action, can be generated through the current reinforcement learning policy based on the angles of the first active arm and the second active arm in the position inverse kinematics information. The target parallel action is used to indicate the target action when the humanoid robot model has a parallel ankle structure.
[0148] S704. Determine the series expected torque according to the target parallel action, the current state, the inverse kinematics information, and the component constraint information, apply the series expected torque to the reinforcement learning environment, and update the reinforcement learning policy to train the reinforcement learning policy of the humanoid robot.
[0149] Optionally, the target parallel action can be converted into the joint state of the series model through the inverse kinematics information and the component constraint information, calculate the series expected torque, apply the series expected torque to the reinforcement learning environment, observe the next state and the reward, and thus update the reinforcement learning policy to achieve the training of the reinforcement learning policy of the humanoid robot.
[0150] As a possible implementation Figure 8 is a schematic flow diagram for determining the series expected torque in the humanoid robot reinforcement learning training method provided by the embodiments of the present application. Refer to Figure 8 As shown, in S704 above, determining the series expected torque according to the target parallel action, the current state, the inverse kinematics information, and the component constraint information includes:
[0151] S801. Calculate the basic torque according to the current angle in the current state and the angle of the target parallel action.
[0152] Optionally, taking the angle as an example, the basic torque can be calculated based on the current angle in the current state and the angle of the target parallel action through a PD controller. Optionally, calculations can also be made taking the angular velocity or angular acceleration as examples, which will not be elaborated in this application.
[0153] Exemplarily, taking the first active arm as an example, the error between the current angle of the first active arm and the angle of the first active arm in the target parallel action can be calculated first, and the basic torque of the first active arm can be calculated according to the error and the PD control algorithm.
[0154] Exemplarily, taking the second active arm as an example, the error between the current angle of the second active arm and the angle of the second active arm in the target parallel action can be calculated first, and the basic torque of the second active arm can be calculated according to the error and the PD control algorithm.
[0155] S802. Calculate the compensation torque according to the inverse kinematic information and the component constraint information.
[0156] Optionally, taking the first active arm as an example, the compensation torque of the first active arm can be calculated according to the angle, angular velocity, and angular acceleration of the first active arm in the inverse kinematic information and the mass matrix, Coriolis force matrix, and gravity term in the component constraint information.
[0157] Exemplarily, the product of the mass matrix and the angular acceleration of the first active arm can be calculated as the fifth product, the product of the Coriolis force matrix and the angular velocity of the first active arm can be calculated as the sixth product, and the sum of the fifth product, the sixth product, and the gravity term can be calculated as the compensation torque of the first active arm.
[0158] Optionally, taking the second active arm as an example, the compensation torque of the second active arm can be calculated according to the angle, angular velocity, and angular acceleration of the second active arm in the inverse kinematic information and the mass matrix, Coriolis force matrix, and gravity term in the component constraint information.
[0159] Exemplarily, the product of the mass matrix and the angular acceleration of the second active arm can be calculated as the seventh product, the product of the Coriolis force matrix and the angular velocity of the second active arm can be calculated as the eighth product, and the sum of the seventh product, the eighth product, and the gravity term can be calculated as the compensation torque of the second active arm.
[0160] S803. Superimpose the basic torque and the compensation torque to obtain the parallel desired torque.
[0161] Optionally, taking the first active arm as an example, the base torque of the first active arm and the compensation torque of the first active arm can be superimposed to obtain the parallel desired torque of the first active arm.
[0162] Optionally, taking the second active arm as an example, the base torque of the second active arm and the compensation torque of the second active arm can be superimposed to obtain the parallel desired torque of the second active arm.
[0163] S804. Calculate the transposed Jacobian matrix according to the Jacobian matrix.
[0164] Optionally, calculate the transposed matrix of the Jacobian matrix to obtain the transposed Jacobian matrix.
[0165] S805. Calculate the product of the transposed Jacobian matrix and the parallel desired torque to obtain the series desired torque.
[0166] Optionally, calculate the product of the transposed Jacobian matrix and the parallel desired torque as the series desired torque.
[0167] Based on the current angle in the current state and the angle of the target parallel action, calculate the base torque, and according to the inverse kinematics information and the component constraint information, calculate the compensation torque, and superimpose the base torque and the compensation torque to obtain the parallel desired torque, and calculate the transposed Jacobian matrix according to the Jacobian matrix, and calculate the product of the transposed Jacobian matrix and the parallel desired torque to obtain the series desired torque, which realizes decoupling the multi-physics field coupling problem into hierarchical sub-problems through matrix operations, achieving global optimization among model accuracy, calculation efficiency and control stability, being able to avoid directly solving complex non-linear equations, improving the real-time calculation efficiency, reducing redundant energy consumption, and at the same time, enhancing the robustness of the system to external disturbances.
[0168] Based on the same inventive concept, an embodiment of the present application also provides a humanoid robot reinforcement learning training device corresponding to the humanoid robot reinforcement learning training method. Since the principle of solving problems by the device in the embodiment of the present application is similar to that of the above-mentioned humanoid robot reinforcement learning training method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0169] Refer to Figure 9 as shown in Figure 9 is a schematic diagram of a humanoid robot reinforcement learning training device provided by an embodiment of the present application. The device includes: an acquisition module 901, an analysis module 902, a determination module 903, and a training module 904;
[0170] The acquisition module 901 is used to acquire a model file of the humanoid robot, and the model file is at least used to indicate the physical structure, sensor information, and kinematic parameters of the humanoid robot;
[0171] A parsing module 902, configured to parse a model file to obtain the positional relationship of the feet of the humanoid robot and the component constraint information of the humanoid robot, and establish a reference coordinate system according to the positional relationship of the feet;
[0172] A determination module 903, configured to construct a parallel ankle mechanism model of the humanoid robot according to the reference coordinate system and the component constraint information, and determine inverse kinematic information according to the parallel ankle mechanism model, where the inverse kinematic information includes position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot;
[0173] A training module 904, configured to train the reinforcement learning policy of the humanoid robot according to the inverse kinematic information and the component constraint information.
[0174] Optionally, the determination module 903 is specifically configured to:
[0175] Determine the position inverse solution information according to the parallel ankle mechanism model;
[0176] Determine the velocity inverse solution information according to the position inverse solution information;
[0177] Determine the acceleration inverse solution information according to the velocity inverse solution information.
[0178] Optionally, the determination module 903 is specifically configured to:
[0179] Determine the attitude of the foot plate, the length of the active arm, the length of the passive rod, the geometric offset, and the geometric constraint conditions according to the parallel ankle mechanism model;
[0180] Input the attitude of the foot plate, the length of the active arm, the length of the passive rod, and the geometric offset into a pre-constructed vector loop equation, and solve to obtain at least one initial value of the first active arm angle and at least one initial value of the second active arm angle;
[0181] Determine the target value of the first active arm angle and the target value of the second active arm angle according to each initial value of the first active arm angle, each initial value of the second active arm angle, and the geometric constraint conditions, and use the target value of the first active arm angle and the target value of the second active arm angle as the position inverse solution information.
[0182] Optionally, the determination module 903 is specifically configured to:
[0183] Determine the active arm vector and the passive rod vector according to the target value of the first active arm angle and the target value of the second active arm angle;
[0184] Calculate the rotation matrix of the end point according to the active arm vector and the passive rod vector;
[0185] Construct the Jacobian matrix based on the active arm vector, passive rod vector, and rotation matrix of the end point;
[0186] Obtain the speed of the foot plate;
[0187] Calculate the product of the speed of the foot plate and the Jacobian matrix to obtain the first active arm angular velocity and the second active arm angular velocity, and use the Jacobian matrix, the first active arm angular velocity, and the second active arm angular velocity as the velocity inverse solution information.
[0188] Optionally, the determination module 903 is specifically used for:
[0189] Determine the acceleration of the foot plate;
[0190] Calculate the time derivative of the Jacobian matrix;
[0191] Calculate the first active arm angular acceleration and the second active arm angular acceleration based on the Jacobian matrix, the speed of the foot plate, the acceleration of the foot plate, and the time derivative of the Jacobian matrix, and use the first active arm angular acceleration and the second active arm angular acceleration as the acceleration inverse solution information.
[0192] Optionally, the training module 904 is specifically used for:
[0193] Initialize the humanoid robot model and the reinforcement learning environment. The humanoid robot model is a humanoid robot with a serial model, and the type of the reinforcement learning environment is a serial model;
[0194] In the current training round, obtain the current state from the reinforcement learning environment. The current state includes joint angles, speeds, and accelerations;
[0195] Determine the target parallel action according to the current reinforcement learning policy, the position inverse solution information, and the current state. The target parallel action is the target action when indicating that the humanoid robot model has a parallel ankle structure;
[0196] Determine the serial expected torque according to the target parallel action, the current state, the inverse kinematics information, and the component constraint information, apply the serial expected torque to the reinforcement learning environment, and update the reinforcement learning policy to train the reinforcement learning policy of the humanoid robot.
[0197] Optionally, the training module 904 is specifically used for:
[0198] Calculate the basic torque according to the current angle in the current state and the angle of the target parallel action;
[0199] Calculate the compensation torque according to the inverse kinematics information and the component constraint information;
[0200] Superimpose the base torque and the compensation torque to obtain the parallel expected torque;
[0201] Calculate the transposed Jacobian matrix according to the Jacobian matrix;
[0202] Calculate the product of the transposed Jacobian matrix and the parallel expected torque to obtain the series expected torque.
[0203] The description of the processing flow of each module in the device and the interaction flow between modules can refer to the relevant descriptions in the above method embodiments and will not be elaborated here.
[0204] The embodiment of the present application also provides an electronic device, as Figure 10 shown, Figure 10 is a schematic structural diagram of the electronic device provided by the embodiment of the present application, including: a processor 1001, a memory 1002. Optionally, a bus 1003 may also be included. The memory 1002 stores machine-readable instructions executable by the processor 1001 (for example, Figure 9 the execution instructions corresponding to the acquisition module 901, the parsing module 902, the determination module 903, and the training module 904 in the device in), when the electronic device runs, the processor 1001 communicates with the memory 1002 through the bus 1003, and when the machine-readable instructions are executed by the processor 1001, the steps of the above-mentioned humanoid robot reinforcement learning training method are executed.
[0205] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the steps of the above-mentioned humanoid robot reinforcement learning training method are executed.
[0206] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the method embodiments, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces, and the indirect coupling or communication connection of the device or module may be electrical, mechanical or other forms.
[0207] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0208] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application.
Claims
1. A method for training a humanoid robot using reinforcement learning, characterized in that, Including: Obtain a model file of a humanoid robot, where the model file is at least used to indicate the physical structure, sensor information, and kinematic parameters of the humanoid robot; Parse the model file to obtain the positional relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establish a reference coordinate system according to the positional relationship of the two feet; Construct a parallel ankle joint mechanism model of the humanoid robot according to the reference coordinate system and the component constraint information, and determine inverse kinematic information according to the parallel ankle joint mechanism model. The inverse kinematic information includes position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot. The parallel ankle joint mechanism model of the humanoid robot at least includes: a plurality of parallel joints, parallel link articulation points corresponding to each parallel joint, active arms corresponding to each parallel joint, and passive rods corresponding to each parallel joint; In the current training round, obtain the current state from the reinforcement learning environment; Determine a target parallel action according to the current reinforcement learning policy, the position inverse solution information, and the current state. The target parallel action is used to indicate the target action when the humanoid robot model has a parallel ankle structure; Calculate a basic torque according to the current angle in the current state and the angle of the target parallel action; calculate a compensation torque according to the inverse kinematic information and the component constraint information; superimpose the basic torque and the compensation torque to obtain a parallel expected torque; calculate a transposed Jacobian matrix according to the Jacobian matrix; calculate the product of the transposed Jacobian matrix and the parallel expected torque to obtain a series expected torque, and apply the series expected torque to the reinforcement learning environment and update the reinforcement learning policy to train the reinforcement learning policy of the humanoid robot.
2. The method for training the humanoid robot reinforcement learning according to claim 1, wherein, The determining the inverse kinematic information according to the parallel ankle joint mechanism model includes: Determine the position inverse solution information according to the parallel ankle joint mechanism model; Determine the velocity inverse solution information according to the position inverse solution information; Determine the acceleration inverse solution information according to the velocity inverse solution information.
3. The method for training a humanoid robot through reinforcement learning according to claim 2, wherein The determining the position inverse solution information according to the parallel ankle joint mechanism model includes: Determine the attitude of the foot plate, the length of the active arm, the length of the passive rod, the geometric offset, and the geometric constraint conditions according to the parallel ankle joint mechanism model; Input the attitude of the foot plate, the length of the active arm, the length of the passive rod, and the geometric offset into a pre-constructed vector loop equation to solve for at least one initial value of the first active arm angle and at least one initial value of the second active arm angle; Determine a target value of the first active arm angle and a target value of the second active arm angle according to each of the initial values of the first active arm angle, each of the initial values of the second active arm angle, and the geometric constraint conditions, and use the target value of the first active arm angle and the target value of the second active arm angle as the position inverse solution information.
4. The method for training a humanoid robot using reinforcement learning according to claim 3, wherein, The determining the velocity inverse solution information according to the position inverse solution information includes: Determine the active arm vector and the passive rod vector according to the first active arm angle target value and the second active arm angle target value; Calculate the rotation matrix of the end point based on the active arm vector and the passive rod vector; Construct the Jacobian matrix based on the active arm vector, the passive rod vector and the rotation matrix of the end point; Obtain the speed of the foot plate; Calculate the product of the speed of the foot plate and the Jacobian matrix to obtain the first active arm angular velocity and the second active arm angular velocity, and use the Jacobian matrix, the first active arm angular velocity and the second active arm angular velocity as the velocity inverse solution information.
5. The method for training a humanoid robot through reinforcement learning according to claim 4, wherein The determining the acceleration inverse solution information according to the velocity inverse solution information includes: Determine the acceleration of the foot plate; Calculate the time derivative of the Jacobian matrix; Calculate the first active arm angular acceleration and the second active arm angular acceleration based on the Jacobian matrix, the speed of the foot plate, the acceleration of the foot plate and the time derivative of the Jacobian matrix, and use the first active arm angular acceleration and the second active arm angular acceleration as the acceleration inverse solution information.
6. The method for training a humanoid robot through reinforcement learning according to claim 1, wherein The current state includes joint angles, speeds and accelerations; before obtaining the current state from the reinforcement learning environment, it further includes: Initialize the humanoid robot model and the reinforcement learning environment, where the humanoid robot model is a humanoid robot with a serial model, and the type of the reinforcement learning environment is a serial model.
7. A reinforcement learning training device for a humanoid robot, characterized in that, It includes: An acquisition module for acquiring the model file of the humanoid robot, where the model file is at least used to indicate the physical structure, sensor information and kinematic parameters of the humanoid robot; An analysis module for analyzing the model file to obtain the position relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establishing a reference coordinate system according to the position relationship of the two feet; A determination module for constructing the parallel ankle joint mechanism model of the humanoid robot according to the reference coordinate system and the component constraint information, and determining the inverse kinematic information according to the parallel ankle joint mechanism model, where the inverse kinematic information includes the position inverse solution information, velocity inverse solution information and acceleration inverse solution information of the humanoid robot, and the parallel ankle joint mechanism model of the humanoid robot at least includes: a plurality of parallel joints, the parallel link hinge points corresponding to each parallel joint, the active arms corresponding to each parallel joint, and the passive rods corresponding to each parallel joint; A training module, which is configured to obtain the current state from the reinforcement learning environment in the current training round; determine a target parallel action according to the current reinforcement learning policy, the position inverse solution information, and the current state, where the target parallel action is the target action when indicating that the humanoid robot model has a parallel ankle structure; calculate a basic torque according to the current angle in the current state and the angle of the target parallel action; calculate a compensation torque according to the inverse kinematics information and the component constraint information; superimpose the basic torque and the compensation torque to obtain a parallel desired torque; calculate a transposed Jacobian matrix according to the Jacobian matrix; calculate the product of the transposed Jacobian matrix and the parallel desired torque to obtain a series desired torque, and apply the series desired torque to the reinforcement learning environment, and update the reinforcement learning policy to train the reinforcement learning policy of the humanoid robot.
8. An electronic device, characterized in that, Comprising: A processor and a memory, where the memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor executes the machine-readable instructions to perform the steps of the humanoid robot reinforcement learning training method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, it executes the steps of the humanoid robot reinforcement learning training method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Gait training method and device of quadruped robot based on deep reinforcement learning, electronic equipment and medium
CN112596534A
Bionic biped robot balance control method and humanoid robot system
CN118579173A