A humanoid robot control method, system, storage medium, and program product

The robot control system is built through imitation learning methods, which solves the problem of inefficient control of humanoid robots, and realizes high-quality action generation and anthropomorphism, which is suitable for virtual human animation production and simulated robot control.

CN119238533BActive Publication Date: 2025-07-11HUA DATA TECH (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411651055.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-07-11
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

The existing humanoid robot control technology cannot efficiently generate high-quality actions and cannot anthropomorphize, resulting in inefficiency and inability to meet the needs of complex tasks.

Method used

The imitation learning method is adopted to preprocess expert action data, build the robot skeleton architecture, configure joint parameters, build state space, action space and reward functions, and use multi-frame control methods, combining reinforcement learning and policy network optimization to drive robot learning.

Benefits of technology

It improves the efficiency of robots to learn high-quality actions, can be anthropomorphized in various scenarios, has higher user acceptance and improved training speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119238533B_ABST
    Figure CN119238533B_ABST
Patent Text Reader

Abstract

The present invention provides a humanoid robot control method, system, storage medium and program product, belonging to the field of computer vision. The method includes: preprocessing expert action data and processing the expert action data into expert data equivalent to the skeletal structure of the target robot; building a robot with a humanoid structure in a simulation environment, configuring the joint parameters of the robot, and each joint degree of freedom is controlled by an independent physical control module; constructing a policy representation method for the robot, including a state space, an action space, a reward function, and a multi-frame control method; initializing the robot; minimizing the difference between the robot action and the expert action in each frame, maximizing the reward function, and driving the robot to learn. The present invention can assist the learning process of the humanoid robot, enabling the robot to be anthropomorphic while completing tasks and improving the training speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a humanoid robot control method, system, storage medium and program product. Background Art

[0002] With the continuous development of robot-related technologies, users have higher and higher requirements for the functionality and adaptability of robots. Most traditional industrial robots have a single application scenario and cannot be generalized to daily production and life. Thus, humanoid robot technology has emerged.

[0003] However, the current humanoid robot control technology cannot control the robot to complete various complex actions and tasks, and cannot be anthropomorphic. In the algorithm using reinforcement learning for robot control, it takes a long time and is inefficient for the robot to generate high-quality actions, which cannot meet the requirements of robot control.

[0004] In recent years, related technologies of imitation learning have developed rapidly, and it has been proved that it can effectively solve the problem of low efficiency caused by the too large exploration space in reinforcement learning. It is of great significance to study the humanoid robot control method based on imitation learning to make the robot generate diverse actions, complete complex tasks and be anthropomorphic at the same time. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the present invention provides a humanoid robot control method, system, storage medium and program product, which uses the imitation learning method for motion control, aiming to solve the problems that it takes a long time and is inefficient for the robot to generate high-quality actions and cannot meet the requirements of robot control, so as to make the robot generate diverse actions, complete complex tasks and be anthropomorphic at the same time.

[0006] In the first aspect, the present invention provides a humanoid robot control method, including the following steps:

[0007] Preprocess the expert action data and process the expert action data into expert data equivalent to the skeletal structure of the target robot;

[0008] Build a robot in a humanoid structure in a simulation environment, configure the joint parameters of the robot, and each joint degree of freedom is controlled by an independent physical control module;

[0009] Construct a policy representation method for the robot, including a state space, an action space, a reward function, and a multi-frame control method;

[0010] Initialize the robot;

[0011] Minimize the difference between the robot action and the expert action in each frame, maximize the reward function, and drive the robot to learn.

[0012] As a further improvement of the present invention, it further includes the steps of:

[0013] Put the states in the state space, the actions in the action space, and the rewards in the reward function into the experience buffer pool. After the experience buffer pool is filled, perform policy update;

[0014] Sample data from the experience buffer pool and perform the update of the policy network and the value network.

[0015] As a further improvement of the present invention, the preprocessed expert action data is processed into expert data equivalent to the skeletal structure of the target robot, including:

[0016] Construct an expert action data set;

[0017] Perform skeletal redirection on the expert action data to adapt to the skeletal parameters of the target robot;

[0018] Fine-tune by the inverse kinematics algorithm to form an imitation learning expert data set.

[0019] As a further improvement of the present invention, the building components of the robot include rigid bodies, collision bodies, and configurable joints;

[0020] The joint parameters of the robot include rigid body weight, collision range, joint degrees of freedom, and joint rotation range.

[0021] As a further improvement of the present invention, the state space includes the position of the pelvis relative to the ground origin, the centroid positions of each rigid body segment, rotation information, linear velocity information, angular velocity information, and time variable.

[0022] As a further improvement of the present invention, the position of the pelvis relative to the ground origin is the first three dimensions of the state space, calculated using the world coordinate system, and is used to provide the relative relationship between the robot and the ground;

[0023] The centroid positions of each rigid body segment, the linear velocity information, and the angular velocity information are all three-dimensional information, all calculated using the pelvis coordinate system;

[0024] The rotation information is a six-dimensional feature vector, calculated using the pelvis coordinate system;

[0025] The time variable is the last dimension of the state space, responsible for providing the time information when the robot executes actions, and is used to model the timing information.

[0026] As a further improvement of the present invention, the action space uses a control method of controlling the joint rotation on each rotational degree of freedom to generate motion. When the robot moves, the joint rotation is a continuous variable;

[0027] The representation methods of the action space include:

[0028] The Gaussian distribution is used to model the action space, and the rotation on each degree of freedom is modeled as an independent Gaussian distribution;

[0029] After completing the modeling of the action distribution, sampling is performed in the distribution, and the sampling result is the target rotation of the joint;

[0030] To convert the target rotation value into joint torque, the robot is driven to move through the joint torque.

[0031] As a further improvement of the present invention, in the reward function, the expression formula of the action imitation reward function of the robot is:

[0032]

[0033] w pos = 0.2, w rot = 0.5, w vel = 0.1, w end = 0.1, w cof = 0.1;

[0034] The form of the reward is the weighted average of the position reward, rotation reward, speed reward, end effector reward, and centroid reward;

[0035] where w is the weighting coefficient of each item, and r is the specific reward value of each item; the reward The items from left to right are: position reward, rotation reward, speed reward, end effector reward, and centroid reward.

[0036] As a further improvement of the present invention, the multi-frame control method includes:

[0037] The historical states of the previous n frames before the current state are concatenated together to form a new feature vector, which is used as the input of the policy network to capture the temporal information between actions;

[0038] The input representation of the policy network is:

[0039]

[0040] where, represents the state vector formed by concatenating the historical states of the previous n frames of the current state;

[0041] The policy network includes a fully connected network, several LSTM modules, and a linear layer network;

[0042] At each step of the simulation environment simulation update, the temporal information is fed into the policy network, and the policy network outputs the mean of the action distribution, and then a sample is generated from the action distribution to obtain an action a;

[0043] The action distribution is characterized as follows:

[0044]

[0045] where a is the action value, and its specific meaning is joint rotation; s is the robot state value returned by the simulation environment; π θ is the conditional probability distribution from state to action, which is modeled as a Gaussian distribution; the mean is output by the network and the variance is a fixed initial value ∑.

[0046] As a further improvement of the present invention, the initialization of the robot includes:

[0047] Using expert state initialization to sample states from the reference motion, so that the robot can sample any state on the motion trajectory in the initial stage.

[0048] As a further improvement of the present invention, minimizing the difference between the robot's action and the expert's action in each frame and maximizing the reward function to drive the robot to learn includes:

[0049] Reinforcement learning training, performing policy gradient ascent to maximize the reward function;

[0050] Monitoring the training process, and performing early stopping when the robot training collapses to prevent the policy from entering a valley and being unable to jump out;

[0051] Obtaining the trained model;

[0052] Putting it into the simulation environment, so that the robot can be anthropomorphic while completing the target action.

[0053] In a second aspect, the present invention provides a computer system, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above method.

[0054] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0055] In a fourth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0056] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0057] Using an imitation learning-based humanoid robot control method can assist the learning process of humanoid robots, enabling the robots to be anthropomorphic while completing tasks and improving the training speed. In various scenarios, the robot can learn human actions through the imitation learning method to perform complex actions. When new action requirements arise, the robot can learn new actions only through a piece of expert action data, with higher efficiency. Moreover, anthropomorphic robots are more acceptable to users and appear warmer. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a flowchart of a control method for a humanoid robot disclosed by the present invention;

[0059] Figure 2 It is a schematic diagram of the principle of a control method for a humanoid robot disclosed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the present invention will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. Among them, steps S1, S2,... in the embodiments described in the present invention do not limit the only execution steps of the present invention; various models, simulation environments, and software described in the present invention are not the only limiting ways of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0061] In the present invention, "module", "system", etc. refer to relevant entities applied to a computer, such as hardware, a combination of hardware and software, software, or software in execution. Specifically, for example, software includes but is not limited to processes running on a processor, processors, objects, executable software, execution threads, programs, and / or computers. Also, an application program or script program running on a server, and the server can both be software. One or more software can be in the process and / or thread of execution, and the software can be localized on one computer and / or distributed between two or more computers and can be run by various computer-readable media.

[0062] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0063] As Figure 1 and Figure 2 shown, in the embodiments of the present invention, on the first aspect, a control method for a humanoid robot is provided, and the specific process can be as follows:

[0064] S1: Preprocess the expert action data and process it into expert data that is equivalent to the skeletal structure of the target robot;

[0065] Preferably, preprocessing the expert action data and processing it into expert data that is equivalent to the skeletal structure of the target robot includes:

[0066] S11: Construct an expert action data set;

[0067] S12: Perform skeletal redirection on the expert action data to adapt to the skeletal parameters of the target robot;

[0068] S13: Fine-tune with the inverse kinematics algorithm to form an imitation learning expert data set.

[0069] S2: Build a robot using a humanoid structure in a simulation environment, configure the joint parameters of the robot, and each joint degree of freedom is controlled by an independent physical control module;

[0070] Preferably, the building components of the robot include rigid bodies, collision bodies, and configurable joints;

[0071] The joint parameters of the robot include rigid body weight, collision range, joint degree of freedom, and joint rotation range.

[0072] In an embodiment of the present invention, in the Unity simulation physics engine, use components such as rigid bodies, collision bodies, and configurable joints to build a humanoid simulation robot with a height of 1.75 meters, a weight of 65 kg, and multiple degrees of freedom and high-dimensional features, and configure joint parameters such as rigid body weight, collision range, joint degree of freedom, and joint rotation range to ensure that the humanoid simulation robot meets the dynamic requirements. Each joint degree of freedom is controlled by an independent physical control module to ensure the flexibility of the robot. The robot has a total of 18 rigid bodies, 17 joints, and 31 rotational degrees of freedom.

[0073] S3: Construct a policy representation method for the robot, including a state space, an action space, a reward function, and a multi-frame control method;

[0074] Preferably, the state space includes the position of the pelvis relative to the ground origin, the centroid position of each rigid body, rotation information, linear velocity information, angular velocity information, and time variable.

[0075] Preferably, the position of the pelvis relative to the ground origin is the first three dimensions of the state space, calculated using the world coordinate system, and is used to provide the relative relationship between the robot and the ground;

[0076] The centroid position of each rigid body, the linear velocity information, and the angular velocity information are all three-dimensional information, all calculated using the pelvis coordinate system;

[0077] The rotation information is a six - dimensional feature vector calculated using the pelvis coordinate system;

[0078] The time variable is the last dimension of the state space, responsible for providing the time information during the execution of the robot's actions and used for modeling the timing information.

[0079] In an embodiment of the present invention, the robot has a total of 18 rigid bodies, 17 joints, and 31 rotational degrees of freedom. For the state space, the first three dimensions are the position of the pelvis relative to the ground origin, which is calculated in the world coordinate system and used to provide the relative relationship between the robot and the ground. In addition, information about other parts of the robot's body is also considered, including the centroid position, rotation information, linear velocity information, and angular velocity information of each rigid body. Among them, the centroid position, linear velocity information, and angular velocity information of each rigid body are three - dimensional information, all calculated in the pelvis coordinate system; the rotation information is a six - dimensional feature vector calculated in the pelvis coordinate system. The last dimension is the time variable t, which is responsible for providing the time information during the execution of the robot's actions, used for modeling the timing information to facilitate capturing and simulating more complex behaviors. The value of the time variable t is related to the current system update frequency and the total duration of the currently executed action. Summarizing the above information, each joint - segment rigid body has three - dimensional position information, three - dimensional linear velocity information, three - dimensional angular velocity information, and six - dimensional state information. Therefore, the total dimension of the state space is 18*(3 + 6 + 3 + 3)+1 = 271, and the last dimension is the time - variable dimension. Each rigid body is represented by independent features, ensuring a more accurate description and control of the robot's actions.

[0080] Preferably, the action space generates motion by controlling the joint rotations on each rotational degree of freedom. When the robot moves, the joint rotations are continuous variables;

[0081] The representation method of the action space includes:

[0082] S31: Model the action space using a Gaussian distribution. The rotation on each degree of freedom is modeled as an independent Gaussian distribution;

[0083] S32: After completing the action - distribution modeling, sample in the distribution, and the sampling result is the target rotation of the joint;

[0084] S33: To convert the target rotation value into joint torque and drive the robot to move through the joint torque.

[0085] In an embodiment of the present invention, the robot has a total of 18 rigid bodies, 17 joints, and 31 rotational degrees of freedom. For the action space, motion is generated by controlling the joint rotations on each rotational degree of freedom. The degree of freedom of the robot is 31 - dimensional, so the action space is 31 - dimensional. When the robot moves, the joint rotations are continuous variables.

[0086] The representation method of the action space includes: to describe joint rotation more precisely, a Gaussian distribution is used to model the action space. The rotation on each degree of freedom is modeled as an independent Gaussian distribution, which provides a continuous and smooth way to describe actions. This construction method has sufficient flexibility to independently adjust the actions on each degree of freedom according to different needs and scenarios. After completing the action distribution modeling, sampling is performed in the distribution, and the sampling result is the target rotation of the joint. Finally, to convert the target rotation value into joint torque, PD-control is required for torque calculation, and the robot is driven to move through joint torque.

[0087] Preferably, in the reward function, the expression formula of the action imitation reward function of the robot at time t is:

[0088]

[0089] w pos = 0.2, w rot = 0.5, w vel = 0.1, w end = 0.1, w cof = 0.1;

[0090] The form of the reward is the weighted average of the position reward, rotation reward, speed reward, end effector reward, and centroid reward;

[0091] where w is the weighting coefficient for each item, and r is the specific reward value for each item; the rewards from left to right for each item are: position reward, rotation reward, speed reward, end effector reward, and centroid reward.

[0092] Preferably, the multi-frame control method includes:

[0093] The historical states of the previous n frames before the current state are concatenated together to form a new feature vector, which is used as the input of the policy network. And because there is one-dimensional time information in the state space design, the temporal information between actions can be captured;

[0094] The input representation of the policy network is:

[0095]

[0096] where represents the state vector formed by concatenating the historical states of the previous n frames of the current state;

[0097] The policy network includes a fully connected network, several LSTM modules, and a linear layer network. To better capture temporal information, several LSTM modules are stacked after the fully connected network, and the output layer of the fully connected network is changed to a feature output layer; the output of the LSTM layer passes through a linear layer network, and finally the output dimension size is equal to the size of the action space.

[0098] At each step of the simulation environment simulation update, the temporal information is fed into the policy network, and the policy network outputs the mean of the action distribution, and then samples an action a from the action distribution;

[0099] The action distribution is characterized as follows, where:

[0100]

[0101] where a is the action value, and its specific meaning is joint rotation; s is the robot state value returned by the simulation environment; π θ is the conditional probability distribution from state to action, which is modeled as a Gaussian distribution; the mean is output by the network and the variance is a fixed initial value ∑.

[0102] For the control method, different from the traditional single-frame control, the present invention adopts a multi-frame control method. Considering the correlation of continuous frame information, the method for representing the policy network that only uses the MLP network and adopts single-frame input is improved. The multi-frame control introduces more context information and enhances the model's understanding of temporal actions.

[0103] S4: Initialize the robot to ensure the optimal training effect of the robot;

[0104] Preferably, the initialization of the robot includes:

[0105] Adopt expert state initialization, sample states from the reference motion, and the robot can sample any state on the motion trajectory in the initial stage.

[0106] The traditional initial state distribution p(s0) is used to determine the initial state of the robot in each episode. Usually, the robot is initialized to a certain fixed state (Fixed Initialization, FI). The specific implementation of this initialization strategy is to initialize the character to the starting state of the reference motion in each episode, and then use the policy to drive the robot to match the actions in the reference motion frame by frame until the end of the entire action segment. This initialization strategy reduces the initial exploration space of the character, and high rewards are only obtained after accessing the states in the reference motion trajectory. Before accessing the high-reward state, the policy cannot judge whether the state is beneficial through the value function, which limits the exploration behavior of the character and is not conducive to training.

[0107] The present invention adopts Expert Initialization (EI). By sampling states from the reference motion, the robot can sample any state on the motion trajectory in the initial stage. Even before the robot explores these states, the EI strategy can discover that high rewards can be obtained in these states.

[0108] S5: Minimize the difference between the robot's actions and the expert's actions in each frame, maximize the reward function, and drive the robot to learn.

[0109] Preferably, the minimizing the difference between the robot's actions and the expert's actions in each frame, maximizing the reward function, and driving the robot to learn includes:

[0110] S51: Perform reinforcement learning training, execute policy gradient ascent to maximize the reward function;

[0111] S52: Monitor the training process, and perform early stopping when the robot's training crashes to prevent the policy from entering a trough and being unable to jump out;

[0112] S53: Obtain the trained model;

[0113] S54: Put it into the simulation environment to make the robot anthropomorphic while completing the target action.

[0114] In an embodiment of the present invention, the Proximal Policy Optimization (PPO) algorithm is used to perform reinforcement learning training on the robot's policy, execute policy gradient ascent to maximize the reward function. The PPO algorithm includes a policy network and a value network. The policy network is responsible for learning the robot's behavior policy, and the value network is responsible for evaluating the behavior of the policy network. At the same time, monitor the training process, and perform early stopping when the robot's training crashes to prevent the policy from entering a trough and being unable to jump out. Finally, obtain the trained model. Put it into the Unity simulation environment, and the robot can be made anthropomorphic while completing the target action.

[0115] In another embodiment of the present invention, on the basis of the steps S1 to S5, it may further include:

[0116] S6: Put the states in the state space, the actions in the action space, and the rewards in the reward function into the experience buffer pool. After the experience buffer pool is filled, perform policy update;

[0117] S7: Sample data from the experience buffer pool and perform the update of the policy network and the value network.

[0118] Such as Figure 2As shown in the figure, the principle of a humanoid robot control method based on imitation learning provided in the embodiments of the present invention is as follows: In the training stage, at the beginning of each episode, an initial state s0 is sampled from the expert dataset, and this state is used to initialize the character in the simulation environment. At each step of the simulation environment simulation update, the temporal information is fed into the policy network. The policy network outputs the mean of the action distribution, and then an action a is sampled from the action distribution; the state, action, and reward are put into the experience buffer. After the experience buffer is filled, the policy update is executed; a small batch of data is sampled from the experience buffer, and the policy network and the value network are updated.

[0119] In the embodiments of the present invention, in the second aspect, the present invention provides a computer system, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.

[0120] In the embodiments of the present invention, in the third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0121] In the embodiments of the present invention, in the fourth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0122] The present invention uses a humanoid robot control method based on imitation learning, which is applicable to fields such as virtual human animation production and simulation robot control. It can assist the learning process of humanoid robots, enabling the robots to be anthropomorphic while completing tasks, and improving the training speed. For example, in a scenario for food service, the robot can learn human actions through the imitation learning method to perform complex actions such as wiping the table, sweeping the floor, and delivering meals. When a new action requirement arises, the robot can learn a new action only through a piece of expert action data, with higher efficiency. Moreover, the anthropomorphic robot is more acceptable to users and appears warmer.

Claims

1. A humanoid robot control method, characterized in that, It includes the following steps: Preprocess the expert action data and process the expert action data into expert data equivalent to the target robot's skeletal structure; Build a robot using a humanoid structure in a simulation environment, configure the joint parameters of the robot, and each joint degree of freedom is controlled by an independent physical control module; Construct a policy representation method for the robot, including a state space, an action space, a reward function, and a multi-frame control method; The state space includes the position of the pelvis relative to the ground origin, the centroid position of each rigid body segment, rotation information, linear velocity information, angular velocity information, and a time variable; The multi-frame control method includes: Concatenate the historical states of the frames before the current state to form a new feature vector, which serves as the input to the policy network to capture the temporal information between actions; Concatenate the historical states of the frames before the current state to form a new feature vector, which serves as the input to the policy network to capture the temporal information between actions; The input representation of the policy network is: ; Among them, represents the state vector formed by concatenating the historical states of the previous frames of the current state; The policy network includes a fully connected network, several LSTM modules, and a linear layer network; At each step of simulating the update in the simulation environment, the timing information is fed into the policy network. The policy network outputs the mean of the action distribution, and then samples an action from the action distribution. ; The action distribution representation is: ; Among them, is the action value, and its specific meaning is joint rotation; s is the robot state value returned by the simulation environment; is the conditional probability distribution from state to action, which is modeled as a Gaussian distribution; the mean is output by the network, and the variance is a fixed initial value ; Initialize the robot; Minimize the difference between the robot's actions and the expert actions in each frame, maximize the reward function, and drive the robot to learn.

2. The humanoid robot control method according to claim 1, wherein It also includes the steps: Put the states in the state space, the actions in the action space, and the rewards in the reward function into an experience buffer pool. After the experience buffer pool is filled, perform policy update; Sample data from the experience buffer pool and perform the update of the policy network and the value network.

3. The humanoid robot control method according to claim 1, wherein The preprocessing of the expert action data, which processes the expert action data into expert data equivalent to the target robot's skeletal structure, includes: Construct an expert action data set; Perform skeletal redirection on the expert action data to adapt to the skeletal parameters of the target robot; Fine-tune by the inverse kinematics algorithm to form an imitation learning expert data set.

4. The humanoid robot control method according to claim 1, wherein, The building components of the robot include rigid bodies, collision bodies, and configurable joints; The joint parameters of the robot include rigid body weight, collision range, joint degree of freedom, and joint rotation range.

5. The humanoid robot control method according to claim 4, wherein, The position of the pelvis relative to the ground origin is the first three dimensions of the state space, calculated using the world coordinate system, and is used to provide the relative relationship between the robot and the ground; The centroid position of each rigid body segment, the linear velocity information, and the angular velocity information are all three-dimensional information, all calculated using the pelvis coordinate system; The rotation information is a six-dimensional feature vector, calculated using the pelvis coordinate system; The time variable is the last dimension of the state space, responsible for providing the time information when the robot's actions are executed, and is used to model the timing information.

6. The humanoid robot control method according to claim 1 or 2, characterized in that, The action space uses a control method of controlling the joint rotation on each rotational degree of freedom to generate motion. When the robot moves, the joint rotation is a continuous variable; The representation method of the action space includes: Use a Gaussian distribution to model the action space, and the rotation on each degree of freedom is modeled as an independent Gaussian distribution; After completing the action distribution modeling, sample in the distribution, and the sampling result is the target rotation of the joint; To convert the target rotation value into joint torque, drive the robot to move through the joint torque.

7. The humanoid robot control method according to claim 1 or 2, characterized in that In the reward function, the expression formula of the robot's action imitation reward function is: ; ; The form of the reward is the weighted average of position reward, rotation reward, velocity reward, end effector reward, and centroid reward; where w is the weighting coefficient of each term and r is the specific reward value of each term; the reward Each term from left to right is in turn: position reward, rotation reward, speed reward, end effector reward, center of mass reward.

8. The humanoid robot control method according to claim 1, wherein The initialization of the robot includes: Using expert state initialization to sample states from the reference motion, the robot can initially sample any state on the motion trajectory.

9. The humanoid robot control method according to claim 1, characterized in that, Minimizing the difference between the robot's actions and the expert's actions in each frame and maximizing the reward function to drive the robot to learn, including: Reinforcement learning training, performing policy gradient ascent to maximize the reward function; Monitoring the training process and performing early stopping when the robot's training crashes to prevent the policy from entering a valley and being unable to escape; Obtaining a trained model; Putting it into a simulation environment to make the robot anthropomorphic while completing the target action.

10. A computer system, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method described in any one of claims 1-9.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1-9.

12. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1-9.

Citation Information

Patent Citations

  • Deformation compensation method, system and equipment of exoskeleton robot and storage medium

    CN117067201A

  • Robot turning and squatting stable standing control method and device and related equipment

    CN117301061A