Humanoid robot reinforcement learning training method, device and equipment and storage medium

By analyzing the humanoid robot model file, building a parallel ankle joint mechanism model and determining the inverse kinematic information, the problems of low training operation efficiency and high parallel sampling efficiency in the existing technology are solved, and a more efficient training process and faster strategy convergence are achieved.

CN120023832AActive Publication Date: 2025-05-23BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510493067.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-23
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The humanoid robot control algorithm in the prior art has low computational efficiency during training and has high requirements for parallel sampling efficiency of reinforcement learning.

Method used

By obtaining the model file of the humanoid robot, analyzing the model file to obtain the position relationship between the feet and component constraint information, establishing a benchmark coordinate system, constructing a parallel ankle joint mechanism model, determining the inverse kinematic information, and training the reinforcement learning strategy based on this information.

Benefits of technology

It improves computing efficiency, reduces the requirements for parallel sampling efficiency of reinforcement learning, simplifies the training process, speeds up the convergence speed of reinforcement learning strategies, and improves the accuracy of motion planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120023832A_ABST
    Figure CN120023832A_ABST
Patent Text Reader

Abstract

The invention provides a humanoid robot reinforcement learning training method, device and equipment and a storage medium, and the method comprises the steps: obtaining a model file of a humanoid robot; analyzing the model file to obtain a position relation of two feet of the humanoid robot and part constraint information of the humanoid robot, and establishing a reference coordinate system according to the position relation of the two feet; according to the reference coordinate system and the component constraint information, a parallel ankle joint mechanism model of the humanoid robot is constructed, and inverse kinematics information is determined according to the parallel ankle joint mechanism model; and training a reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information. According to the method, series-parallel connection conversion in a resolving mode in the training process is avoided, the operation efficiency is improved, meanwhile, the requirement for reinforcement learning parallel sampling efficiency is lowered, the training process is simplified, and the convergence speed of reinforcement learning strategies is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robot control technology, and in particular to a method, device, equipment and storage medium for reinforcement learning training of a humanoid robot. Background Art

[0002] Humanoid robots are an extremely complex form of robot. Typically, humanoid robots have complex bipedal and bimanual mechanical structures that contain numerous serial and parallel mechanisms. This complexity makes the control problem of humanoid robots a typical nonlinear, multivariable, and time-varying control problem. Traditional model-based control algorithms usually require accurate modeling of the robot and its operating environment, resulting in highly complex models. In addition, when these models are generalized to different tasks and environments, they usually need to be readjusted. Therefore, more and more researchers are using deep reinforcement learning methods to solve the control problem of humanoid robots.

[0003] The humanoid robot control algorithm in the prior art usually adopts an open-loop topology during the training process and completes the series-parallel conversion in the stage of migrating from the simulation environment to the real world (Simulation to Reality, sim2real for short).

[0004] However, this training method causes the neural network to spend more time searching for the optimal solution in the action space, resulting in low computational efficiency. At the same time, it also places high demands on the parallel sampling efficiency of reinforcement learning for humanoid robots. Summary of the invention

[0005] The purpose of this application is to provide a humanoid robot reinforcement learning training method, device, equipment and storage medium to address the deficiencies in the above-mentioned prior art, so as to solve the problems of low computing efficiency and high parallel sampling efficiency requirements in the prior art.

[0006] To achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows: In a first aspect, an embodiment of the present application provides a humanoid robot reinforcement learning training method, the method comprising: Acquire a model file of a humanoid robot, wherein the model file is used to indicate at least a physical structure, sensor information, and kinematic parameters of the humanoid robot; Parsing the model file to obtain the positional relationship of the two feet of the humanoid robot and component constraint information of the humanoid robot, and establishing a reference coordinate system according to the positional relationship of the two feet; According to the reference coordinate system and the component constraint information, a parallel ankle joint mechanism model of the humanoid robot is constructed, and inverse kinematics information is determined according to the parallel ankle joint mechanism model, wherein the inverse kinematics information includes position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot; The reinforcement learning strategy of the humanoid robot is trained according to the inverse kinematics information and the component constraint information.

[0007] Optionally, determining inverse kinematics information according to the parallel ankle joint mechanism model includes: Determining the position inverse solution information according to the parallel ankle joint mechanism model; Determining the velocity inverse solution information according to the position inverse solution information; The acceleration inverse solution information is determined according to the velocity inverse solution information.

[0008] Optionally, determining the position inverse solution information according to the parallel ankle joint mechanism model includes: According to the parallel ankle joint mechanism model, determining the posture of the foot plate, the length of the active arm, the length of the passive rod, the geometric offset and the geometric constraint conditions; Inputting the posture of the foot plate, the length of the active arm, the length of the passive rod, and the geometric offset into a pre-constructed vector loop equation, and solving to obtain at least one first active arm angle initial value and at least one second active arm angle initial value; According to each of the first active arm angle initial values, each of the second active arm angle initial values ​​and the geometric constraints, the first active arm angle target value and the second active arm angle target value are determined, and the first active arm angle target value and the second active arm angle target value are used as the position inverse solution information.

[0009] Optionally, determining the velocity inverse solution information according to the position inverse solution information includes: Determine an active arm vector and a passive rod vector according to the first active arm angle target value and the second active arm angle target value; According to the active arm vector and the passive rod vector, a rotation matrix of the end point is calculated; Constructing a Jacobian matrix according to the active arm vector, the passive rod vector and the rotation matrix of the end point; Get the speed of the foot; The product of the foot plate speed and the Jacobian matrix is ​​calculated to obtain a first active arm angular velocity and a second active arm angular velocity, and the Jacobian matrix, the first active arm angular velocity and the second active arm angular velocity are used as the velocity inverse solution information.

[0010] Optionally, determining the acceleration inverse solution information according to the velocity inverse solution information includes: Determine the acceleration of the foot; calculating the time derivative of the Jacobian matrix; The first active arm angular acceleration and the second active arm angular acceleration are calculated according to the Jacobian matrix, the speed of the foot, the acceleration of the foot and the time derivative of the Jacobian matrix, and the first active arm angular acceleration and the second active arm angular acceleration are used as the acceleration inverse solution information.

[0011] Optionally, the training of the reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information includes: Initializing a humanoid robot model and a reinforcement learning environment, wherein the humanoid robot model is a humanoid robot with a serial model, and the type of the reinforcement learning environment is a serial model; In a current training round, obtaining a current state from the reinforcement learning environment, wherein the current state includes a joint angle, a velocity, and an acceleration; Determine a target parallel action according to a current reinforcement learning strategy, the position inverse solution information, and the current state, wherein the target parallel action is used to indicate a target action when the humanoid robot model has a parallel ankle structure; According to the target parallel action, the current state, the inverse kinematics information and the component constraint information, the expected series torque is determined, the expected series torque is applied to the reinforcement learning environment, and the reinforcement learning strategy is updated to train the reinforcement learning strategy of the humanoid robot. Optionally, determining the expected series torque according to the target parallel action, the current state, the inverse kinematics information and the component constraint information includes: Calculating a basic torque according to a current angle in the current state and an angle of the target parallel action; Calculating a compensation torque according to the inverse kinematics information and the component constraint information; Superimposing the basic torque and the compensation torque to obtain a parallel desired torque; According to the Jacobian matrix, the transposed Jacobian matrix is ​​calculated; The product of the transposed Jacobian matrix and the parallel desired torque is calculated to obtain the series desired torque.

[0012] In a second aspect, another embodiment of the present application provides a humanoid robot reinforcement learning training device, the device comprising: An acquisition module, used for acquiring a model file of a humanoid robot, wherein the model file is used for indicating at least a physical structure, sensor information and kinematic parameters of the humanoid robot; A parsing module, used to parse the model file, obtain the position relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establish a reference coordinate system according to the position relationship of the two feet; a determination module, configured to construct a parallel ankle joint mechanism model of the humanoid robot according to the reference coordinate system and the component constraint information, and determine inverse kinematics information according to the parallel ankle joint mechanism model, wherein the inverse kinematics information includes position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot; A training module is used to train the reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information.

[0013] In the third aspect, another embodiment of the present application provides an electronic device, comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of any method described in the first aspect above.

[0014] In a fourth aspect, another embodiment of the present application provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any method described in the first aspect are executed.

[0015] The beneficial effects of the present application are: by obtaining the model file of the humanoid robot and parsing the model file, the positional relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot are obtained, and according to the positional relationship of the two feet, a reference coordinate system is established, and the parallel ankle joint mechanism model of the humanoid robot can be constructed according to the reference coordinate system and the component constraint information, and the inverse kinematics information is determined according to the parallel ankle joint mechanism model, so that the reinforcement learning strategy of the humanoid robot can be trained according to the inverse kinematics information and the component constraint information, avoiding the series-parallel conversion by solving in the training process, improving the operation efficiency, and at the same time, reducing the requirements for the parallel sampling efficiency of reinforcement learning, simplifying the training process, and accelerating the convergence speed of the reinforcement learning strategy. In addition, the training of the reinforcement learning strategy based on the parallel ankle joint model and the component constraint information can also improve the accuracy of motion planning, and the instability factors in training can be reduced by optimizing the strategy training through inverse kinematics information and constraint information. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 A schematic diagram of a flow chart of a humanoid robot reinforcement learning training method provided in an embodiment of the present application; Figure 2 A schematic diagram of a parallel ankle joint mechanism model provided in an embodiment of the present application; Figure 3 A schematic diagram of a process for determining inverse kinematics information in a humanoid robot reinforcement learning training method provided in an embodiment of the present application; Figure 4 A schematic diagram of a process for determining position inverse solution information in the humanoid robot reinforcement learning training method provided in an embodiment of the present application; Figure 5 A schematic diagram of a process for determining velocity inverse solution information in the humanoid robot reinforcement learning training method provided in an embodiment of the present application; Figure 6 A schematic diagram of a process for determining acceleration inverse solution information in the humanoid robot reinforcement learning training method provided in an embodiment of the present application; Figure 7 A schematic diagram of a flow chart of training a reinforcement learning strategy for a humanoid robot in the reinforcement learning training method for a humanoid robot provided in an embodiment of the present application; Figure 8 A schematic diagram of a process for determining a series desired torque in a humanoid robot reinforcement learning training method provided in an embodiment of the present application; Fig. 9 A schematic diagram of a humanoid robot reinforcement learning training device provided in an embodiment of the present application; Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0018] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.

[0019] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0020] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0021] The humanoid robot control algorithm in the prior art usually adopts an open-loop topology during the training process and completes the series-parallel conversion in the stage of migrating from the simulation environment to the real world (Simulation to Reality, sim2real for short).

[0022] However, this training method causes the neural network to spend more time searching for the optimal solution in the action space, resulting in low computational efficiency. At the same time, it also places high demands on the parallel sampling efficiency of reinforcement learning for humanoid robots.

[0023] Based on the above-mentioned problems, an embodiment of the present application proposes a reinforcement learning training method for a humanoid robot. By acquiring a model file of the humanoid robot and parsing the model file, the position relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot are obtained, and a reference coordinate system is established according to the position relationship of the two feet. A parallel ankle joint mechanism model of the humanoid robot can be constructed according to the reference coordinate system and the component constraint information, and inverse kinematics information can be determined according to the parallel ankle joint mechanism model. Therefore, the reinforcement learning strategy of the humanoid robot can be trained according to the inverse kinematics information and the component constraint information, thereby avoiding series-parallel conversion by solving the problem during the training process, improving the computational efficiency, and at the same time, reducing the requirements for parallel sampling efficiency of reinforcement learning, simplifying the training process, and accelerating the convergence speed of the reinforcement learning strategy.

[0024] The following is a detailed description of the humanoid robot reinforcement learning training method provided in the embodiments of the present application in combination with multiple embodiments.

[0025] Figure 1 A flowchart of a humanoid robot reinforcement learning training method provided in an embodiment of the present application, referring to Figure 1 As shown, the execution subject of the method can be any electronic device with processing capability, and the method includes: S101. Obtain a model file of a humanoid robot.

[0026] The model file is at least used to indicate the physical structure, sensor information and kinematic parameters of the humanoid robot.

[0027] Optionally, the model file also includes visual appearance and collision model information.

[0028] Exemplarily, the physical structure may include rigid components of the humanoid robot, such as: chassis, sensors, joints of the robotic arm, geometric connections between connecting rods, relationships between joints and motors, etc. The sensor information may include basic information of the sensors on the humanoid robot (specifically including type, identifier, manufacturer, model and version), physical properties (specifically including size, shape, installation position, Euler angle position, electrical interface, operating temperature, waterproof level and anti-electromagnetic interference capability), performance parameters (specifically including measurement range, resolution, accuracy, sampling rate, field of view and sensitivity), communication interface (specifically including communication protocol, data format, data synchronization mechanism, bus configuration), calibration parameters (specifically including internal parameter matrix, external field matrix, offset compensation and calibration timestamp), noise and error information (specifically including noise characteristics, system error, interference model), data output configuration (specifically including output mode, supported data compression algorithm, supported data filtering algorithm), driver dependency configuration (specifically including driver, dependencies, configuration file path) and other information (specifically including cost, life, certification and user manual link, etc.).

[0029] Exemplarily, kinematic parameters are used to indicate the joint motion relationships and motion restrictions of the humanoid robot, such as joint variable range, velocity acceleration restrictions, singular configuration definition, geometric and inertial parameters, homogeneous transformation matrix, self-motion parameters, and joint coupling relationships.

[0030] Exemplarily, the visual appearance includes color, texture, etc., and the collision model information includes mesh model, elastic coefficient, collision mask, level, trigger area, and component ownership.

[0031] By obtaining the model file of the humanoid robot, it is possible to ensure that the behavior of the humanoid robot complies with the laws of physics.

[0032] S102, parsing the model file to obtain the positional relationship between the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establishing a reference coordinate system according to the positional relationship between the two feet.

[0033] Optionally, the model file can be parsed by a parsing tool to extract the positional relationship of the two feet of the humanoid robot. Specifically, the parsing tool can be the urdf_parser or SDFormat library of ROS. The positional relationship of the two feet of the humanoid robot is used to indicate the connecting rod and joint information of the left and right feet of the humanoid robot.

[0034] Optionally, the kinematic parameters of each component, the geometric information in the physical structure, and the inertial information are read from the model file, and based on the kinematic parameters of each component, the geometric information in the physical structure, and the inertial information, the mass matrix, Coriolis force and centrifugal force matrix, and gravity term are calculated in combination with the dynamic model as component constraint information.

[0035] Specifically, the mass matrix can be obtained through the mass and inertia tensor of the connecting rod, the Coriolis force and centrifugal force matrices can be obtained through the mass matrix and the joint velocity, and the gravity term can be obtained through the connecting rod instruction, the center of mass position and the gravitational acceleration.

[0036] Exemplarily, the model file may be parsed to obtain the left foot position, the right foot position and the joint constraints, and the center point may be calculated to establish a reference coordinate system.

[0037] S103. Construct a parallel ankle joint mechanism model of the humanoid robot according to the reference coordinate system and component constraint information, and determine inverse kinematics information according to the parallel ankle joint mechanism model.

[0038] Optionally, after obtaining the reference coordinate system, the posture of the ankle joint end relative to the calf can be defined , that is, the posture of the foot, where and Represents the rotation angles around the X-axis and Y-axis respectively, and uses the rotation matrix RX( ) and RY( ) Calculate the position o′Pi of the parallel connecting rod hinge point Pi. Specifically, it can be calculated by referring to the following formula: o′Pi=RY(θ)RX( ) o′Pi′ Among them, o′Pi′ is the coordinate of the hinge point at its initial position.

[0039] For example, Figure 2 A schematic diagram of a parallel ankle joint mechanism model provided in an embodiment of the present application, referring to Figure 2 As shown, after calculating the position o′Pi of the parallel link hinge point Pi, the length of the active arm can be determined as L1, and the length of the passive rod can be determined as L2, thereby obtaining a parallel ankle joint mechanism model.

[0040] Wherein, Pi is the parallel connecting rod hinge point corresponding to the ith parallel joint in the parallel ankle joint mechanism model. Depending on the structure of the humanoid robot, the parallel ankle joint mechanism model of the humanoid robot can have different numbers of parallel joints. Ai is the ith parallel node in the Z-axis direction of the parallel ankle joint mechanism model, Bi is the ith parallel node in the YO'Z plane of the parallel ankle joint mechanism model, and Ci is the ith parallel node in the XYZ space of the parallel ankle joint mechanism model.

[0041] Optionally, after the parallel ankle joint mechanism model is obtained, inverse kinematics information may be calculated based on the geometric relationship in the parallel ankle joint mechanism model.

[0042] The inverse kinematics information includes the inverse position information, the inverse velocity information and the inverse acceleration information of the humanoid robot.

[0043] Exemplarily, the position inverse solution information includes the angle of the active arm of the humanoid robot, the velocity inverse solution information includes the angular velocity of the active arm of the humanoid robot, and the acceleration inverse solution information includes the angular acceleration of the active arm of the humanoid robot.

[0044] The active arm angle of the humanoid robot can be used to determine the arm posture and end effector position of the humanoid robot, thereby ensuring the accuracy of motion planning and inverse kinematics calculations. The active arm angular velocity of the humanoid robot can be used to determine the arm movement speed of the humanoid robot, thereby ensuring the accuracy of dynamic motion control and balance adjustment. The active arm angular acceleration of the humanoid robot can be used to determine the acceleration of the arm movement, thereby ensuring the accuracy of dynamic modeling and force control.

[0045] S104: Training the reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information.

[0046] Optionally, after obtaining the inverse kinematics information, the reinforcement learning strategy of the humanoid robot may be adjusted based on the inverse kinematics information and the component constraint information, and training may be performed based on the adjusted reinforcement learning strategy.

[0047] In this embodiment, by obtaining the model file of the humanoid robot and parsing the model file, the positional relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot are obtained, and according to the positional relationship of the two feet, a reference coordinate system is established, and the parallel ankle joint mechanism model of the humanoid robot can be constructed according to the reference coordinate system and the component constraint information, and the inverse kinematics information is determined according to the parallel ankle joint mechanism model, so that the reinforcement learning strategy of the humanoid robot can be trained according to the inverse kinematics information and the component constraint information, avoiding the series-parallel conversion by solving in the training process, improving the operation efficiency, and at the same time, reducing the requirements for the parallel sampling efficiency of the reinforcement learning, simplifying the training process, and accelerating the convergence speed of the reinforcement learning strategy. In addition, the training of the reinforcement learning strategy based on the parallel ankle joint model and the component constraint information can also improve the accuracy of motion planning, and the instability factors in the training can be reduced by optimizing the strategy training through inverse kinematics information and constraint information.

[0048] As a possible implementation, Figure 3 A schematic diagram of a process for determining inverse kinematics information in a humanoid robot reinforcement learning training method provided in an embodiment of the present application, referring to Figure 3 As shown, the inverse kinematics information is determined according to the parallel ankle joint mechanism model in S103, including: S301. Determine position inverse solution information according to the parallel ankle joint mechanism model.

[0049] Optionally, the parallel ankle joint mechanism model may be subjected to geometric analysis and solved to obtain the angle of the active arm of the humanoid robot as inverse position solution information.

[0050] Exemplarily, the solution may be obtained by using a vector ring method, a homogeneous coordinate transformation method or a screw theory method, or by directly establishing equations through geometric analysis.

[0051] S302: Determine the speed inverse solution information according to the position inverse solution information.

[0052] Optionally, after obtaining the position inverse solution information, the speed inverse solution information may be calculated based on the position inverse solution information.

[0053] For example, the velocity inverse solution information may be calculated by a multi-joint Jacobian matrix method or a numerical integration method, wherein the velocity inverse solution information may be the angular velocity of the active arm of the humanoid robot.

[0054] S303. Determine acceleration inverse solution information according to velocity inverse solution information.

[0055] Optionally, after the velocity inverse solution information is obtained, the acceleration inverse solution information may be calculated based on the velocity inverse solution information.

[0056] Exemplarily, the angular velocity of the active arm of the humanoid robot may be calculated by any one of a Jacobian matrix derivative method, a Hessian matrix method, and a Lagrange equation method to obtain the angular acceleration of the active arm of the humanoid robot.

[0057] Through the parallel ankle joint mechanism model, the position inverse solution information is first determined, and then the velocity inverse solution information is determined based on the position inverse solution information, and then the acceleration inverse solution information is determined based on the velocity inverse solution information. The inverse kinematics information of the humanoid robot can be gradually obtained in a layered manner, so that the terminal can be accurately positioned through the position inverse solution, the dynamic response can be optimized through the velocity inverse solution, and the impact can be eliminated through the acceleration inverse solution to achieve smooth movement of the humanoid robot. It can also avoid singular configurations, reduce computational complexity, reduce energy consumption, and ensure stability and real-time performance.

[0058] As a possible implementation, Figure 4 A schematic diagram of a process for determining position inverse solution information in the humanoid robot reinforcement learning training method provided in the embodiment of the present application, referring to Figure 4 As shown, according to the parallel ankle joint mechanism model, the position inverse solution information is determined, including: S401. Determine the posture of the foot, the length of the active arm, the length of the passive rod, the geometric offset, and the geometric constraints according to the parallel ankle joint mechanism model.

[0059] Optionally, the kinematics of the parallel ankle joint can be determined from the parallel ankle joint mechanism model to obtain the posture x of the foot, the length L1 of the active arm, the length L2 of the passive rod, the geometric offset r1, the hinge point Pi of the parallel link, and the geometric constraints. The geometric constraints may be the rod length constraints and the joint angle range of the humanoid robot.

[0060] S402, input the posture of the foot, the length of the active arm, the length of the passive rod, and the geometric offset into a pre-constructed vector loop equation, and solve to obtain at least one first active arm angle initial value and at least one second active arm angle initial value.

[0061] Optionally, the first active arm vector ring equation and the second active arm vector ring equation can be established in advance through geometric constraints, and the posture χ of the foot, the length L1 of the active arm, the length L2 of the passive rod, the geometric offset r1 and the hinge point Pi of the parallel connecting rod are respectively input into the first active arm vector ring equation and the second active arm vector ring equation, and the first active arm vector ring equation and the second active arm vector ring equation are solved to obtain at least one first active arm angle initial value and at least one second active arm angle initial value.

[0062] Exemplarily, the first active arm vector loop equation and the second active arm vector loop equation may be solved by numerical methods such as Newton's iteration method to obtain at least one first active arm angle initial value and at least one second active arm angle initial value.

[0063] S403, determining the first active arm angle target value and the second active arm angle target value according to each first active arm angle initial value, each second active arm angle initial value and geometric constraints, and using the first active arm angle target value and the second active arm angle target value as position inverse solution information.

[0064] Optionally, after obtaining at least one first active arm angle initial value and at least one second active arm angle initial value, the first active arm angle target value and the second active arm angle target value can be screened from each first active arm angle initial value and each second active arm angle initial value according to the joint angle range and the rod length constraint, and the first active arm angle target value and the second active arm angle target value are used as the inverse position solution information.

[0065] Through the parallel ankle joint mechanism model, the posture of the foot, the length of the active arm, the length of the passive rod, the geometric offset and the geometric constraints are determined, and the posture of the foot, the length of the active arm, the length of the passive rod, and the geometric offset are input into the pre-constructed vector ring equation to solve and obtain at least one first active arm angle initial value and at least one second active arm angle initial value, and according to each first active arm angle initial value, each second active arm angle initial value and the geometric constraints, the first active arm angle target value and the second active arm angle target value are determined, and the first active arm angle target value and the second active arm angle target value are used as position inverse solution information, which can directly model the geometric relationship based on the parallel ankle joint mechanism model, avoid kinematic approximation errors, ensure that the solution results strictly meet the actual physical constraints, and can also improve the movement efficiency of the humanoid robot. In addition, it helps to improve the humanoid robot's fault tolerance to physical deviations and enhance control stability.

[0066] As a possible implementation, Figure 5 A schematic diagram of a process for determining the inverse speed information in the humanoid robot reinforcement learning training method provided in the embodiment of the present application, referring to Figure 5 As shown, the above S302 determines the speed inverse solution information according to the position inverse solution information, including: S501 . Determine an active arm vector and a passive rod vector according to a first active arm angle target value and a second active arm angle target value.

[0067] Optionally, the active arm vector BiCi and the passive arm vector CiPi are determined according to the first active arm angle target value and the second active arm angle target value in the position inverse solution information.

[0068] S502: Calculate the rotation matrix of the end point according to the active arm vector and the passive rod vector.

[0069] Optionally, the rotation matrix of the end point Pi in the reference coordinate system is calculated based on the active arm vector BiCi and the passive rod vector CiPi.

[0070] S503, constructing a Jacobian matrix according to the active arm vector, the passive rod vector and the rotation matrix of the end point.

[0071] Optionally, vector operations may be performed based on the active arm vector BiCi, the passive rod vector CiPi, and the rotation matrix of the end point Pi to construct a Jacobian matrix.

[0072] Exemplarily, taking the i-th row of the Jacobian matrix as an example, the y-component of the vector cross product of the active arm vector BiCi and the passive rod vector CiPi can be calculated, and the transpose vector of the passive rod vector CiPi can be calculated, and the ratio of the transpose vector of the passive rod vector CiPi to the y-component can be calculated, and the product of the ratio of the transpose vector of the passive rod vector CiPi to the y-component and the rotation matrix of the end point Pi in the reference coordinate system can be calculated to obtain the i-th row of the Jacobian matrix, and so on, the Jacobian matrix can be obtained.

[0073] S504, obtaining the speed of the foot.

[0074] Optionally, the speed of the first foot plate and the speed of the second foot plate, that is, the speed of the end effector, may be acquired.

[0075] S505, calculating the product of the foot plate speed and the Jacobian matrix to obtain the first active arm angular velocity and the second active arm angular velocity, and using the Jacobian matrix, the first active arm angular velocity and the second active arm angular velocity as velocity inverse solution information.

[0076] Optionally, matrix multiplication can be used to calculate the product of the velocity of the first foot and the Jacobian matrix, and the product of the velocity of the second foot and the Jacobian matrix, respectively, to obtain the angular velocity of the first active arm and the angular velocity of the second active arm, and the Jacobian matrix, the angular velocity of the first active arm and the angular velocity of the second active arm are used as velocity inverse solution information.

[0077] The active arm vector and the passive rod vector are determined by the first active arm angle target value and the second active arm angle target value, and the rotation matrix of the end point is calculated based on the active arm vector and the passive rod vector, and the Jacobian matrix is ​​constructed based on the active arm vector, the passive rod vector and the rotation matrix of the end point, and the product of the foot speed and the Jacobian matrix is ​​calculated to obtain the first active arm angular velocity and the second active arm angular velocity. The Jacobian matrix, the first active arm angular velocity and the second active arm angular velocity are used as the speed inverse solution information, which can avoid complex iterative calculations, adapt to high-frequency control cycles, ensure real-time dynamic tracking, and avoid step-by-step approximation errors, which is especially suitable for the rapid change of direction requirements of high-rigidity ankle joints. In addition, the singular area is avoided in advance through speed redistribution or trajectory correction to prevent the mechanism from losing control due to sudden changes in joint speed. In addition, the optimal balance between accuracy, efficiency and controllability can be achieved.

[0078] As a possible implementation, Figure 6 A schematic diagram of a process for determining the inverse solution information of acceleration in the humanoid robot reinforcement learning training method provided in the embodiment of the present application, referring to Figure 6 As shown, the above S303 determines the acceleration inverse solution information according to the velocity inverse solution information, including: S601, determine the acceleration of the foot.

[0079] Optionally, the acceleration of the first foot plate and the acceleration of the second foot plate, that is, the acceleration of the end effector, may be acquired from the sensor side.

[0080] Optionally, the speed of the first foot and the speed of the second foot may be differentiated with respect to time to obtain the acceleration of the first foot and the acceleration of the second foot.

[0081] S602: Calculate the time derivative of the Jacobian matrix.

[0082] Optionally, the time derivative of the Jacobian matrix can be computed via the chain rule.

[0083] Exemplarily, the partial derivative of the Jacobian matrix with respect to the joint angle is calculated, and the partial derivative is multiplied by the joint velocity to obtain the time derivative of the Jacobian matrix.

[0084] S603, according to the Jacobian matrix, the speed of the foot, the acceleration of the foot and the time derivative of the Jacobian matrix, calculate the first active arm angular acceleration and the second active arm angular acceleration, and use the first active arm angular acceleration and the second active arm angular acceleration as acceleration inverse solution information.

[0085] Optionally, the first active arm angular acceleration is calculated according to the Jacobian matrix, the velocity of the first foot plate, the acceleration of the first foot plate and the time derivative of the Jacobian matrix.

[0086] Optionally, the angular acceleration of the second active arm is calculated according to the Jacobian matrix, the velocity of the second foot plate, the acceleration of the second foot plate and the time derivative of the Jacobian matrix.

[0087] Exemplarily, the product of the Jacobian matrix and the acceleration of the first foot is calculated as the first product, and the product of the time derivative of the Jacobian matrix and the velocity of the first foot is calculated as the second product, and the sum of the first product and the second product is calculated to obtain the first active arm angular acceleration.

[0088] Exemplarily, the product of the Jacobian matrix and the acceleration of the second foot is calculated as the third product, and the product of the time derivative of the Jacobian matrix and the velocity of the second foot is calculated as the fourth product, and the sum of the third product and the fourth product is calculated to obtain the angular acceleration of the second active arm.

[0089] By determining the acceleration of the foot, calculating the time derivative of the Jacobian matrix, and calculating the first active arm angular acceleration and the second active arm angular acceleration according to the Jacobian matrix, the speed of the foot, the acceleration of the foot and the time derivative of the Jacobian matrix, and using the first active arm angular acceleration and the second active arm angular acceleration as the acceleration inverse solution information, it is possible to eliminate dynamic errors in high-speed motion, achieve precise tracking of the terminal acceleration, avoid mechanical vibration caused by sudden changes in joint speed, and extend the fatigue life of the mechanism.

[0090] As a possible implementation, Figure 7 A schematic diagram of a flow chart of training a reinforcement learning strategy for a humanoid robot in a reinforcement learning training method for a humanoid robot provided in an embodiment of the present application, referring to Figure 7 As shown, the above S104 trains the reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information, including: S701, initialize the humanoid robot model and reinforcement learning environment.

[0091] The humanoid robot model is a humanoid robot with a serial model, and the type of the reinforcement learning environment is a serial model.

[0092] It can be understood that the serial model means that during the reinforcement learning strategy training process, the humanoid robot is modeled as an open-loop serial structure model rather than an actual closed-loop parallel structure (such as the parallel mechanism of the ankle). That is, the robot joints are connected in series, and the movement of each joint only depends on the state of its previous joint, thereby simplifying the dynamic calculations and making the training process more efficient.

[0093] S702. In the current training round, obtain the current state from the reinforcement learning environment.

[0094] The current state includes joint angles, velocities, and accelerations.

[0095] Optionally, it is determined whether the reinforcement learning strategy has converged. If not, the current state is obtained from the reinforcement learning environment in the current training round. The current state includes joint angles, velocities, and accelerations.

[0096] S703: Determine the target parallel action according to the current reinforcement learning strategy, the position inverse solution information and the current state.

[0097] Optionally, a target parallel action, i.e., a parallel mechanism action, can be generated based on the angle of the first active arm and the angle of the second active arm in the position inverse solution information through the current reinforcement learning strategy, wherein the target parallel action is used to indicate the target action when the humanoid robot model has a parallel ankle structure.

[0098] S704. Determine the expected series torque according to the target parallel action, the current state, the inverse kinematics information, and the component constraint information, apply the expected series torque to the reinforcement learning environment, and update the reinforcement learning strategy to train the reinforcement learning strategy of the humanoid robot. Optionally, the target parallel action can be converted into the joint state of the series model through inverse kinematics information and component constraint information, and the series expected torque can be calculated. The series expected torque is applied to the reinforcement learning environment, and the next state and reward are observed to update the reinforcement learning strategy, so as to realize the training of the reinforcement learning strategy for the humanoid robot. As a possible implementation, Figure 8 A schematic diagram of a process for determining the expected series torque in the humanoid robot reinforcement learning training method provided in the embodiment of the present application, referring to Figure 8 As shown, in the above S704, the desired series torque is determined according to the target parallel action, the current state, the inverse kinematics information and the component constraint information, including: S801. Calculate the basic torque according to the current angle in the current state and the angle of the target parallel action.

[0099] Optionally, taking the angle as an example, the basic torque can be calculated based on the current angle in the current state through the PD controller based on the current angle in the current state and the angle of the target parallel action. Optionally, the calculation can also be performed using angular velocity or angular acceleration as an example, which is not described in detail in this application.

[0100] Exemplarily, taking the first active arm as an example, the error between the current angle of the first active arm and the angle of the first active arm in the target parallel action can be calculated first, and the basic torque of the first active arm can be calculated based on the error and the PD control algorithm.

[0101] For example, taking the second active arm as an example, the error between the current angle of the second active arm and the angle of the second active arm in the target parallel action can be calculated first, and then the basic torque of the second active arm can be calculated based on the error and the PD control algorithm.

[0102] S802: Calculate and obtain compensation torque according to inverse kinematics information and component constraint information.

[0103] Optionally, taking the first active arm as an example, the compensation torque of the first active arm can be calculated according to the angle of the first active arm, the angular velocity of the first active arm and the angular acceleration of the first active arm in the inverse kinematics information and the mass matrix, Coriolis force matrix and gravity term in the component constraint information.

[0104] Exemplarily, the product of the mass matrix and the angular acceleration of the first active arm can be calculated as the fifth product, the product of the Coriolis force matrix and the angular velocity of the first active arm can be calculated as the sixth product, and the sum of the fifth product, the sixth product and the gravity term can be calculated as the first active arm compensation torque.

[0105] Optionally, taking the second active arm as an example, the compensation torque of the second active arm can be calculated based on the angle of the second active arm, the angular velocity of the second active arm and the angular acceleration of the second active arm in the inverse kinematics information and the mass matrix, Coriolis force matrix and gravity term in the component constraint information.

[0106] Exemplarily, the product of the mass matrix and the angular acceleration of the second active arm can be calculated as the seventh product, the product of the Coriolis force matrix and the angular velocity of the second active arm can be calculated as the eighth product, and the sum of the seventh product, the eighth product and the gravity term can be calculated as the second active arm compensation torque.

[0107] S803: Superimpose the basic torque and the compensation torque to obtain the parallel desired torque.

[0108] Optionally, taking the first active arm as an example, the basic torque of the first active arm and the compensation torque of the first active arm may be superimposed to obtain the parallel expected torque of the first active arm.

[0109] Optionally, taking the second active arm as an example, the basic torque of the second active arm and the compensation torque of the second active arm may be superimposed to obtain the parallel expected torque of the second active arm.

[0110] S804. Calculate and obtain a transposed Jacobian matrix according to the Jacobian matrix.

[0111] Optionally, a transposed matrix of the Jacobian matrix is ​​calculated to obtain a transposed Jacobian matrix.

[0112] S805 , calculating the product of the transposed Jacobian matrix and the parallel desired torque to obtain the series desired torque.

[0113] Optionally, the product of the transposed Jacobian matrix and the parallel desired torque is calculated as the series desired torque.

[0114] The basic torque is calculated through the current angle in the current state and the angle of the target parallel action, and the compensation torque is calculated based on the inverse kinematics information and component constraint information. The basic torque and the compensation torque are superimposed to obtain the parallel desired torque, and the transposed Jacobian matrix is ​​calculated based on the Jacobian matrix, and the product of the transposed Jacobian matrix and the parallel desired torque is calculated to obtain the series desired torque. This decouples the multi-physics field coupling problem into hierarchical sub-problems through matrix operations, and achieves global optimization among model accuracy, computational efficiency and control stability. It can avoid directly solving complex nonlinear equations, improve real-time computing efficiency, and reduce redundant energy consumption. At the same time, it can also improve the robustness of the system to external disturbances.

[0115] Based on the same inventive concept, the embodiments of the present application also provide a humanoid robot reinforcement learning training device corresponding to the humanoid robot reinforcement learning training method. Since the principle of solving the problem by the device in the embodiments of the present application is similar to the above-mentioned humanoid robot reinforcement learning training method in the embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0116] Reference Fig. 9 As shown, Fig. 9 A schematic diagram of a humanoid robot reinforcement learning training device provided in an embodiment of the present application, the device comprising: an acquisition module 901, a parsing module 902, a determination module 903 and a training module 904; An acquisition module 901 is used to acquire a model file of a humanoid robot, where the model file is used to indicate at least a physical structure, sensor information, and kinematic parameters of the humanoid robot; The parsing module 902 is used to parse the model file, obtain the position relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establish a reference coordinate system according to the position relationship of the two feet; A determination module 903 is used to construct a parallel ankle joint mechanism model of the humanoid robot according to the reference coordinate system and the component constraint information, and determine inverse kinematics information according to the parallel ankle joint mechanism model, wherein the inverse kinematics information includes position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot; The training module 904 is used to train the reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information.

[0117] Optionally, the determination module 903 is specifically configured to: According to the parallel ankle joint mechanism model, determine the inverse position solution information; Determine the velocity inverse solution information based on the position inverse solution information; According to the inverse solution information of velocity, the inverse solution information of acceleration is determined.

[0118] Optionally, the determination module 903 is specifically configured to: According to the parallel ankle joint mechanism model, the posture of the foot, the length of the active arm, the length of the passive rod, the geometric offset and the geometric constraints are determined; Inputting the posture of the foot plate, the length of the active arm, the length of the passive rod, and the geometric offset into a pre-constructed vector loop equation, and solving to obtain at least one initial value of the first active arm angle and at least one initial value of the second active arm angle; According to each first active arm angle initial value, each second active arm angle initial value and geometric constraints, the first active arm angle target value and the second active arm angle target value are determined, and the first active arm angle target value and the second active arm angle target value are used as position inverse solution information.

[0119] Optionally, the determination module 903 is specifically configured to: Determine an active arm vector and a passive rod vector according to a first active arm angle target value and a second active arm angle target value; According to the active arm vector and the passive rod vector, the rotation matrix of the end point is calculated; The Jacobian matrix is ​​constructed based on the active arm vector, the passive rod vector and the rotation matrix of the end point; Get the speed of the foot; The product of the foot plate velocity and the Jacobian matrix is ​​calculated to obtain the first active arm angular velocity and the second active arm angular velocity, and the Jacobian matrix, the first active arm angular velocity and the second active arm angular velocity are used as velocity inverse solution information.

[0120] Optionally, the determination module 903 is specifically configured to: Determine the acceleration of the foot; Compute the time derivative of the Jacobian matrix; According to the Jacobian matrix, the speed of the foot, the acceleration of the foot and the time derivative of the Jacobian matrix, the first active arm angular acceleration and the second active arm angular acceleration are calculated, and the first active arm angular acceleration and the second active arm angular acceleration are used as acceleration inverse solution information.

[0121] Optionally, the training module 904 is specifically configured to: Initialize a humanoid robot model and a reinforcement learning environment, wherein the humanoid robot model is a humanoid robot with a serial model, and the type of the reinforcement learning environment is a serial model; In the current training round, the current state is obtained from the reinforcement learning environment. The current state includes joint angles, velocities, and accelerations. Determine a target parallel action according to the current reinforcement learning strategy, the position inverse solution information and the current state, where the target parallel action is used to indicate a target action when the humanoid robot model has a parallel ankle structure; According to the target parallel action, the current state, inverse kinematics information and component constraint information, the expected series torque is determined, and the expected series torque is applied to the reinforcement learning environment, and the reinforcement learning strategy is updated to train the reinforcement learning strategy of the humanoid robot. Optionally, the training module 904 is specifically configured to: The basic torque is calculated based on the current angle in the current state and the angle of the target parallel action; The compensation torque is calculated based on the inverse kinematics information and the component constraint information; The basic torque and the compensation torque are superimposed to obtain the parallel desired torque; According to the Jacobian matrix, the transposed Jacobian matrix is ​​calculated; Calculate the product of the transposed Jacobian matrix and the parallel desired torque to get the series desired torque.

[0122] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference may be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.

[0123] The present application also provides an electronic device, such as Fig.10 As shown, Fig.10 The electronic device structure diagram provided in the embodiment of the present application includes: a processor 1001, a memory 1002, and optionally, a bus 1003. The memory 1002 stores machine-readable instructions executable by the processor 1001 (for example, Fig. 9 In the device, the acquisition module 901, the analysis module 902, the determination module 903 and the execution instructions corresponding to the training module 904 are obtained, etc.), when the electronic device is running, the processor 1001 communicates with the memory 1002 through the bus 1003, and when the machine-readable instructions are executed by the processor 1001, the steps of the above-mentioned humanoid robot reinforcement learning training method are executed.

[0124] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the above-mentioned humanoid robot reinforcement learning training method are executed.

[0125] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0126] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or part of the technical solution that contributes to the prior art or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), disk or optical disk and other media that can store program code.

[0127] The above are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be covered by the protection scope of the present application.

Claims

1. A humanoid robot reinforcement learning training method, characterized in that: include: Acquire a model file of a humanoid robot, wherein the model file is used to indicate at least a physical structure, sensor information, and kinematic parameters of the humanoid robot; Parsing the model file to obtain the positional relationship of the two feet of the humanoid robot and component constraint information of the humanoid robot, and establishing a reference coordinate system according to the positional relationship of the two feet; According to the reference coordinate system and the component constraint information, a parallel ankle joint mechanism model of the humanoid robot is constructed, and inverse kinematics information is determined according to the parallel ankle joint mechanism model, wherein the inverse kinematics information includes position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot; The reinforcement learning strategy of the humanoid robot is trained according to the inverse kinematics information and the component constraint information.

2. The humanoid robot reinforcement learning training method according to claim 1, characterized in that: Determining inverse kinematics information according to the parallel ankle joint mechanism model includes: Determining the position inverse solution information according to the parallel ankle joint mechanism model; Determining the velocity inverse solution information according to the position inverse solution information; The acceleration inverse solution information is determined according to the velocity inverse solution information.

3. The humanoid robot reinforcement learning training method according to claim 2, characterized in that: Determining the position inverse solution information according to the parallel ankle joint mechanism model includes: According to the parallel ankle joint mechanism model, determining the posture of the foot plate, the length of the active arm, the length of the passive rod, the geometric offset and the geometric constraint conditions; Inputting the posture of the foot plate, the length of the active arm, the length of the passive rod, and the geometric offset into a pre-constructed vector loop equation, and solving to obtain at least one first active arm angle initial value and at least one second active arm angle initial value; According to each of the first active arm angle initial values, each of the second active arm angle initial values ​​and the geometric constraints, the first active arm angle target value and the second active arm angle target value are determined, and the first active arm angle target value and the second active arm angle target value are used as the position inverse solution information.

4. The humanoid robot reinforcement learning training method according to claim 3, characterized in that: The determining the velocity inverse solution information according to the position inverse solution information comprises: Determine an active arm vector and a passive rod vector according to the first active arm angle target value and the second active arm angle target value; According to the active arm vector and the passive rod vector, a rotation matrix of the end point is calculated; Constructing a Jacobian matrix according to the active arm vector, the passive rod vector and the rotation matrix of the end point; Get the speed of the foot; The product of the foot plate speed and the Jacobian matrix is ​​calculated to obtain a first active arm angular velocity and a second active arm angular velocity, and the Jacobian matrix, the first active arm angular velocity and the second active arm angular velocity are used as the velocity inverse solution information.

5. The humanoid robot reinforcement learning training method according to claim 4, characterized in that: The determining the acceleration inverse solution information according to the velocity inverse solution information comprises: Determine the acceleration of the foot; calculating the time derivative of the Jacobian matrix; The first active arm angular acceleration and the second active arm angular acceleration are calculated according to the Jacobian matrix, the speed of the foot, the acceleration of the foot and the time derivative of the Jacobian matrix, and the first active arm angular acceleration and the second active arm angular acceleration are used as the acceleration inverse solution information.

6. The humanoid robot reinforcement learning training method according to claim 1, characterized in that: The step of training the reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information comprises: Initializing a humanoid robot model and a reinforcement learning environment, wherein the humanoid robot model is a humanoid robot with a serial model, and the type of the reinforcement learning environment is a serial model; In a current training round, obtaining a current state from the reinforcement learning environment, wherein the current state includes a joint angle, a velocity, and an acceleration; Determine a target parallel action according to a current reinforcement learning strategy, the position inverse solution information, and the current state, wherein the target parallel action is used to indicate a target action when the humanoid robot model has a parallel ankle structure; According to the target parallel action, the current state, the inverse kinematics information and the component constraint information, the expected series torque is determined, the expected series torque is applied to the reinforcement learning environment, and the reinforcement learning strategy is updated to train the reinforcement learning strategy of the humanoid robot.

7. The humanoid robot reinforcement learning training method according to claim 6, characterized in that: The determining of the expected series torque according to the target parallel action, the current state, the inverse kinematics information and the component constraint information includes: Calculating a basic torque according to a current angle in the current state and an angle of the target parallel action; Calculating a compensation torque according to the inverse kinematics information and the component constraint information; Superimposing the basic torque and the compensation torque to obtain a parallel desired torque; According to the Jacobian matrix, the transposed Jacobian matrix is ​​calculated; The product of the transposed Jacobian matrix and the parallel desired torque is calculated to obtain the series desired torque.

8. A humanoid robot reinforcement learning training device, characterized in that: include: An acquisition module, used for acquiring a model file of a humanoid robot, wherein the model file is used for indicating at least a physical structure, sensor information and kinematic parameters of the humanoid robot; A parsing module, used to parse the model file, obtain the position relationship of the two feet of the humanoid robot and the component constraint information of the humanoid robot, and establish a reference coordinate system according to the position relationship of the two feet; a determination module, configured to construct a parallel ankle joint mechanism model of the humanoid robot according to the reference coordinate system and the component constraint information, and determine inverse kinematics information according to the parallel ankle joint mechanism model, wherein the inverse kinematics information includes position inverse solution information, velocity inverse solution information, and acceleration inverse solution information of the humanoid robot; A training module is used to train the reinforcement learning strategy of the humanoid robot according to the inverse kinematics information and the component constraint information.

9. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the humanoid robot reinforcement learning training method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the humanoid robot reinforcement learning training method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Three-degree-of-freedom parallel ankle joint rehabilitation mechanism and control method thereof

    CN110559159A

  • Gait training method and device of quadruped robot based on deep reinforcement learning, electronic equipment and medium

    CN112596534A

  • Motion control method and system of table tennis robot and storage medium

    CN113650010A

  • Bionic biped robot balance control method and humanoid robot system

    CN118579173A

  • Plug-in type PTC heater assembly

    KR1020250028588A