A control method, device, and storage medium for a humanoid robot to walk with straight knees
Through Wasserstein-GAIL imitation learning method, using human gait data to train robot decision-making strategies, the problem of generating straight knee gait in the existing technology is solved, and the robot can naturally walk straight knees in the real environment is realized, which improves the naturalness and aesthetics of humanoid robots.
Patent Information
- Application Number
- CN202510616094.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing humanoid robot control scheme is difficult to achieve natural straight knee gait. The model-driven method is limited by the assumption of the centroid height. The reinforcement learning method lacks targeted reward function guidance, which makes it difficult for the learning process to converge to a natural and upright straight knee gait.
The generative adversarial imitation learning method based on Wasserstein-GAIL is adopted, and the decision-making strategy of the robot is trained using human gait data, speed tracking reward function and style reward function are constructed, and combined with the PD controller, the imitation learning strategy is optimized to reduce the difference between simulation and reality.
The robot's ability to walk naturally straight on its knees in a real environment is realized, and the robot's naturalness and aesthetics in high-simulation walking tasks are improved, and the technical needs of high-end service and performance interaction are met.
Smart Images

Figure CN120116236B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of robot control, and particularly to a control method, device, and storage medium for a humanoid robot to walk with straight knees. Background Technique
[0002] With the continuous development of robot technology, people are constantly pursuing technological progress in humanoid robots, enabling them to be applied in scenarios with high requirements for "human image simulation", such as public services, exhibition displays, stage performances, and film and television production. And as the scenarios gradually become more complex, the motion capabilities of humanoid robots are also continuously improving. In particular, higher standards are put forward for the naturalness, aesthetics, and human similarity of the walking actions of humanoid robots. For example, in occasions such as T-stage model walking, welcome reception, and immersive performances, the "crouching gait" of traditional humanoid robots with bent knees not only appears clumsy and rigid, but also seriously affects the overall visual perception, reducing the interactive affinity and scene integration of humanoid robots.
[0003] Different from the "crouching gait" of traditional humanoid robots with bent knees, the "straight-knee gait" nowadays is closer to the biomechanical characteristics of humans in natural walking and has the following functional advantages: First, by making full use of the dynamic characteristics of the robot itself, efficient regulation of the centroid acceleration is achieved, thereby reducing energy consumption; second, under the same joint configuration conditions, a higher maximum walking speed can be achieved; third, the naturalness and coordination of the robot's actions are improved, which is beneficial to obtaining better acceptance in the human-robot interaction environment.
[0004] The current mainstream control schemes for humanoid robots are mostly model-driven methods and policy learning methods based on reinforcement learning. For the model-driven method (Model-based Control), a typical method is the gait generation controller based on the linear inverted pendulum model (LIPM). To ensure the simplicity of model calculation, the assumption of a constant centroid height is often required. However, this assumption will directly limit the robot from achieving a straight-knee gait, forcing the robot to maintain walking stability in a bent-knee form. If the centroid height limit is tried to be relaxed, the gait generation model will exhibit non-linear characteristics, resulting in a sharp increase in the complexity of system solution and even the occurrence of motion singularities (such as the complete extension of the knee joint leading to the loss of control degrees of freedom), affecting stability and control accuracy.
[0005] Secondly, there is the policy learning method based on reinforcement learning. Reinforcement learning can directly optimize control policies in high-dimensional control spaces and has good adaptability. However, in the learning of specific target actions such as straight-knee walking, due to the large number of redundant degrees of freedom of the humanoid robot itself, there are a large number of action policies that are physically feasible but have extremely different behavioral characteristics in the search space. Without the guidance of a targeted reward function, the learning process is prone to falling into local optima and difficult to converge to a natural and upright straight-knee gait. Moreover, the artificial design of a high-quality reward function is not only costly but also lacks generality.
[0006] In summary, the above two humanoid robot control schemes cannot train humanoid robots walking with straight knees well, reducing the performance of humanoid robots walking with straight knees in real environments. Summary of the Invention
[0007] This application discloses a control method, device, and storage medium for a humanoid robot to walk with straight knees, aiming to improve the performance of a humanoid robot walking with straight knees in different scenarios.
[0008] To meet the higher requirements of humanoid robots for naturalness, aesthetics, and human consistency in high-fidelity walking tasks and to break through the technical bottlenecks of existing model-driven control and traditional reinforcement learning methods in aspects such as policy expression ability, control stability, and training guidance, this application specifically proposes a generative adversarial imitation learning method based on optimal transport distance optimization (Wasserstein-GAIL) to achieve a walking control strategy for a humanoid robot with a natural straight-knee gait.
[0009] The first aspect of this application discloses a control method for a humanoid robot to walk with straight knees, including:
[0010] Collect human gait data, where the human gait data is the gait data of a human walking with straight knees;
[0011] Perform action redirection processing on the human gait data and convert it into reference motion sequence data of the target humanoid robot;
[0012] Model the walking task of the target humanoid robot as a velocity-conditioned Markov decision process, build an adversarial network under this modeling, and generate a maximum expected discounted return function according to the environmental variable distribution and decision-making strategy;
[0013] Construct a velocity tracking reward function based on the linear velocity and angular velocity of the centroid of the target humanoid robot;
[0014] Construct a soft-boundary Wasserstein loss function of the discriminator with gradient penalty terms sampled from the real data distribution, generated data distribution, specific distribution, and specific distribution, and construct a style reward function based on the output of the discriminator;
[0015] Construct a PD controller such that the PD controller converts the actions output by the decision-making strategy into torque signals for the target humanoid robot;
[0016] After performing simulation learning using the constructed adversarial network, select the physical parameters of the target humanoid robot during straight-knee walking training according to the relevance, and set a prior distribution for each physical parameter;
[0017] Conduct a real-scene test movement for the target humanoid robot and collect the operation data corresponding to the physical parameters;
[0018] Combine the collected operation data with the prior distribution corresponding to the physical parameters, and feedback the updated posterior distribution to the imitation learning stage to optimize the control strategy in the imitation learning stage.
[0019] Optionally, perform action redirection processing on the human gait data and convert it into reference motion sequence data for the target humanoid robot, including:
[0020] Perform skeleton topology unification and bone binding processing on the human skeleton and the humanoid robot skeleton to generate an original skeleton, and there are several key joints on the original skeleton;
[0021] Perform coordinate system unification processing and root normalization processing based on the human gait data and the original skeleton to generate the key joint data of the humanoid robot;
[0022] Perform multi-objective inverse kinematics optimization to solve the key joint data to generate a reference motion sequence. The multi-objective inverse kinematics optimization is used to map the Cartesian positions of the key joints and the poses of the end effectors to the corresponding joint angular directions.
[0023] Optionally, the steps of performing skeleton topology unification and bone binding processing on the human skeleton and the humanoid robot skeleton to generate the original skeleton include:
[0024] Construct the kinematic trees of the human skeleton and the humanoid robot skeleton;
[0025] Perform skeleton merging on the kinematic trees of the human skeleton and the humanoid robot skeleton to generate the original skeleton, and retain one bone between the key joints;
[0026] Select the key joints and record the length of each bone segment.
[0027] Optionally, the steps of performing multi-objective inverse kinematics optimization to solve the key joint data to generate a reference motion sequence include:
[0028] Construct a key joint position matching loss function based on the actual positions of the key joints of the humanoid robot and the expected target positions of the key joints in the key joint data;
[0029] Construct an end - effector pose matching loss function based on the actual pose of the end - effector of the humanoid robot and the expected target pose in the key joint data;
[0030] Construct a joint minimum displacement loss function based on the joint angle data between adjacent frames in the key joint data;
[0031] Generate an objective function based on the key joint position matching loss function, the end - effector pose matching loss function, and the joint minimum displacement loss function;
[0032] Perform multi - objective inverse kinematics optimization on the key joint data according to the objective function to generate a reference motion sequence.
[0033] Optionally, the steps of generating an objective function based on the key joint position matching loss function, the end - effector pose matching loss function, and the joint minimum displacement loss function include:
[0034] Determine the acquisition scene parameters of the human gait data;
[0035] Perform texture contrast analysis and slope contrast analysis on the acquisition scene parameters and the flat ground scene parameters to generate gait deformation data;
[0036] Determine the key joints that generate deformation on the original skeleton according to the gait deformation data;
[0037] Generate key joint position weights for the key joints that generate deformation;
[0038] Determine the pose change degree parameter according to the gait deformation data, and generate end - effector pose weights according to the pose change degree parameter;
[0039] Determine the minimum displacement loss weight;
[0040] Generate an objective function based on the key joint position matching loss function, the end - effector pose matching loss function, the joint minimum displacement loss function, the key joint position weights, the end - effector pose weights, and the minimum displacement loss weights.
[0041] Optionally, after the step of performing multi - objective inverse kinematics optimization on the key joint data to generate a reference motion sequence, and before the step of modeling the walking task of the target humanoid robot as a Markov decision process with a speed condition, constructing an adversarial network under this modeling, and generating a maximum expected discounted return function according to the environmental variable distribution and the decision strategy, the control method further includes:
[0042] Perform trajectory quality optimization processing on the reference motion sequence.
[0043] The second aspect of this application discloses a control device for a humanoid robot to walk with straight knees, including:
[0044] The first acquisition unit is used to acquire human gait data, and the human gait data is the gait data of a human walking with straight knees;
[0045] The redirection unit is used to perform action redirection processing on the human gait data and convert it into reference motion sequence data of the target humanoid robot;
[0046] The first construction unit is used to model the walking task of the target humanoid robot as a Markov decision process with a speed condition, construct an adversarial network under this modeling, and generate a maximum expected discounted return function according to the environmental variable distribution and decision-making strategy;
[0047] The second construction unit is used to construct a speed tracking reward function with the linear velocity and angular velocity of the centroid of the target humanoid robot;
[0048] The third construction unit is used to construct a soft boundary Wasserstein loss function of the discriminator with the gradient penalty terms of the real data distribution, generated data distribution, specific distribution, and specific distribution sampling, and construct a style reward function with the output of the discriminator;
[0049] The fourth construction unit is used to construct a PD controller so that the PD controller converts the action output by the decision-making strategy into a torque signal of the target humanoid robot;
[0050] The setting unit is used to select the physical parameters of the target humanoid robot in the straight-knee walking training according to the relevance after completing the simulation learning using the constructed adversarial network, and set a prior distribution for each physical parameter;
[0051] The second acquisition unit is used to perform a real-scene test movement on the target humanoid robot and acquire the operation data corresponding to the physical parameters;
[0052] The optimization unit is used to combine the acquired operation data with the prior distribution corresponding to the physical parameters, and feedback the updated posterior distribution to the imitation learning stage to optimize the control strategy in the imitation learning stage.
[0053] Optionally, the redirection unit includes:
[0054] The first generation module is used to perform skeleton topology unification and bone binding processing on the human skeleton and the humanoid robot skeleton to generate an original skeleton, and there are several key joints on the original skeleton;
[0055] The second generation module is used to perform coordinate system unification processing and root normalization processing according to the human gait data and the original skeleton to generate the key joint data of the humanoid robot;
[0056] A third generation module for performing multi-objective inverse kinematics optimization on the key joint data to generate a reference motion sequence. The multi-objective inverse kinematics optimization is used to map the Cartesian positions of the key joints and the poses of the end effectors to the corresponding joint angular directions.
[0057] Optionally, the first generation module includes:
[0058] Construct the kinematic trees of the human skeleton and the humanoid robot skeleton;
[0059] Merge the kinematic trees of the human skeleton and the humanoid robot skeleton to generate an original skeleton, and retain one bone between the key joints;
[0060] Select the key joints and record the lengths of each bone segment.
[0061] Optionally, the third generation module includes:
[0062] A first construction sub-module for constructing a key joint position matching loss function according to the actual positions of the key joints of the humanoid robot and the expected target positions of the key joints in the key joint data;
[0063] A second construction sub-module for constructing an end effector attitude matching loss function according to the actual pose of the end effector of the humanoid robot and the expected target pose in the key joint data;
[0064] A third construction sub-module for constructing a joint minimum displacement loss function according to the joint angle data between adjacent frames in the key joint data;
[0065] A first generation sub-module for generating an objective function according to the key joint position matching loss function, the end effector attitude matching loss function, and the joint minimum displacement loss function;
[0066] A second generation sub-module for performing multi-objective inverse kinematics optimization on the key joint data according to the objective function to generate a reference motion sequence.
[0067] Optionally, the first generation sub-module includes:
[0068] Determine the acquisition scene parameters of the human gait data;
[0069] Perform texture comparison analysis and slope comparison analysis on the acquisition scene parameters and the flat ground scene parameters to generate gait deformation data;
[0070] Determine the key joints that generate deformation on the original skeleton according to the gait deformation data;
[0071] Generate key joint position weights for the key joints that generate deformation;
[0072] Determine the pose change degree parameter according to the gait deformation data, and generate the end pose weight according to the pose change degree parameter.
[0073] Determine the minimum displacement loss weight.
[0074] Generate an objective function according to the key joint position matching loss function, the end effector attitude matching loss function, the joint minimum displacement loss function, the key joint position weight, the end pose weight, and the minimum displacement loss weight.
[0075] Optionally, after the third generation module and before the first construction unit, the control device further includes:
[0076] An optimization module for performing trajectory quality optimization processing on the reference motion sequence.
[0077] A third aspect of the present application provides a control device for a humanoid robot to walk with straight knees, including:
[0078] A processor, a memory, an input / output unit, and a bus;
[0079] The processor is connected to the memory, the input / output unit, and the bus;
[0080] The memory stores a program, and the processor calls the program to execute the control method as described in the first aspect and any optional control method of the first aspect.
[0081] A fourth aspect of the present application provides a computer-readable storage medium, on which a program is stored, and when the program is executed on a computer, it executes the control method as described in the first aspect and any optional control method of the first aspect.
[0082] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0083] In this application, first, human gait data is collected, and the human gait data is the gait data of a human walking with straight knees. The human gait data is processed by action redirection and converted into reference motion sequence data of the target humanoid robot. The walking task of the target humanoid robot is modeled as a Markov decision process with a speed condition. Under this modeling, an adversarial network is constructed to generate a maximum expected discounted reward function according to the environmental variable distribution and decision-making strategy. A speed tracking reward function is constructed based on the linear velocity and angular velocity of the centroid of the target humanoid robot. A soft-boundary Wasserstein loss function of the discriminator is constructed with gradient penalty terms sampled from the real data distribution, generated data distribution, specific distribution, and specific distribution, and a style reward function is constructed with the output of the discriminator. A PD controller is constructed so that the PD controller converts the actions output by the decision-making strategy into torque signals of the target humanoid robot. After performing simulation learning using the constructed adversarial network, physical parameters of the target humanoid robot during straight-knee walking training are selected according to the relevance, and a prior distribution is set for each physical parameter. A live test motion is performed on the target humanoid robot, and the operation data corresponding to the physical parameters is collected. The collected operation data is combined with the prior distribution corresponding to the physical parameters, and the updated posterior distribution is fed back to the imitation learning stage to optimize the control strategy in the imitation learning stage.
[0084] By introducing the Generative Adversarial Imitation Learning (GAIL) framework, using real human gait data as a demonstration, the decision-making strategy of the humanoid robot is trained to approximate human gait characteristics in the behavior space, so that the natural straight-knee walking ability can be directly obtained without the design of a reward function. On this basis, to further improve the stability and accuracy of the imitation effect, the Wasserstein distance based on the optimal transport theory is introduced to construct an improved GAIL discriminator. The Wasserstein distance has better gradient properties and stronger distribution fitting ability during the training process, which can improve the sensitivity of policy learning to the differences in detailed motions. This improvement enables the robot to not only capture the macroscopic morphological characteristics of human gait but also more accurately restore its dynamic coordination and subtle rhythm. On this basis, physical parameters are inferred according to prior knowledge and the operation data of the humanoid robot to reduce the influence of the difference between simulation and reality. Its cooperation with the existing improvements provides more accurate parameter support for straight-knee gait generation and further improves the performance of the humanoid robot walking with straight knees in different real environments. Brief Description of the Drawings
[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0086] Figure 1 Schematic diagram of an embodiment of the control method for a humanoid robot to walk with straight knees of the present applicant;
[0087] Figure 2 Schematic diagram of an embodiment of the method for action redirection processing of the present application;
[0088] Figure 3 Schematic diagram of an embodiment of the method for generating an original skeleton of the present application;
[0089] Figure 4 Schematic diagram of an embodiment of the method for generating a reference motion sequence of the present application;
[0090] Figure 5 Schematic diagram of an embodiment of the method for generating an objective function of the present application;
[0091] Figure 6 Schematic diagram of an embodiment of the method for processing a reference motion sequence of the present application;
[0092] Figure 7 Schematic diagram of an embodiment of the control device for a humanoid robot to walk with straight knees of the present applicant;
[0093] Figure 8 Schematic diagram of another embodiment of the control device for a humanoid robot to walk with straight knees of the present applicant;
[0094] Figure 9 Schematic diagram of the walking control strategy of a humanoid robot based on Wasserstein adversarial imitation learning of the present application. Detailed implementation manners
[0095] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0096] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0097] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0098] As used in the specification and claims of this application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.
[0099] In addition, in the description of the specification and claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0100] References to "one embodiment" or "some embodiments" or the like described in the specification of this application mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants mean "including but not limited to", unless otherwise specifically emphasized.
[0101] In the prior art, most current mainstream humanoid robot control schemes are model-driven methods and policy learning methods based on reinforcement learning. For the model-driven method (Model-based Control), typical methods such as the gait generation controller based on the linear inverted pendulum model (LIPM), in order to ensure the simplicity of model calculation, it is often necessary to set the assumption condition of a constant center of mass height. However, this assumption will directly limit the robot from achieving a straight-knee gait, causing the robot to have to maintain walking stability in a bent-knee form. If an attempt is made to relax the center of mass height limit, the gait generation model will exhibit non-linear characteristics, resulting in a sharp increase in the complexity of system solution and even the occurrence of motion singularities (such as the complete extension of the knee joint leading to the loss of control degrees of freedom), affecting stability and control accuracy.
[0102] Secondly, there is the policy learning method based on reinforcement learning. Reinforcement learning can directly optimize control policies in a high-dimensional control space and has good adaptability. However, in the learning of specific target actions such as straight-knee walking, due to the large number of redundant degrees of freedom of the humanoid robot itself, there are a large number of action policies that are physically feasible but have extremely different behavioral characteristics in the search space. Without the guidance of a targeted reward function, the learning process is prone to falling into local optima and difficult to converge to a natural and upright straight-knee gait. Moreover, the artificial design of a high-quality reward function is not only costly but also lacks generality.
[0103] To sum up, the above two humanoid robot control schemes cannot train humanoid robots walking with straight knees well, reducing the performance of humanoid robots walking with straight knees in the real environment.
[0104] To meet the higher requirements of humanoid robots for naturality, aesthetics, and human consistency in high-fidelity walking tasks, and to break through the technical bottlenecks of existing model-driven control and traditional reinforcement learning methods in aspects such as policy expression ability, control stability, and training guidance, this application proposes a generative adversarial imitation learning method based on optimal transport distance optimization (Wasserstein-GAIL) to achieve the walking control policy of humanoid robots with natural straight-knee gaits.
[0105] Regarding the means innovation for new business requirements, the technical solution of this application introduces the idea of imitation learning for emerging business scenarios with high requirements for robot action expressiveness, such as T-stage shows, immersive welcome guests, and stage performances. Taking real human gait data as a reference template, and through the learning and training of the humanoid robot policy, its action trajectory is highly similar to that of humans in terms of temporal characteristics, posture style, and dynamic rhythm, thus meeting the scenario requirements of "humanization" and "natural fluency".
[0106] This technical solution uses an improved GAIL framework as the core means and introduces the Wasserstein distance to replace the traditional classification discriminant loss function to improve the training stability and policy distribution fitting ability of imitation learning. This means not only effectively avoids the dependence on high-quality reward functions in traditional reinforcement learning but also significantly enhances the naturality and performance tension of the robot generating straight-knee gaits.
[0107] Regarding the means breakthrough for existing technical problems, to solve the problem that it is difficult to generate straight-knee gaits under the "center of mass height assumption" of existing model-driven methods, and the problems of "difficult reward function design and unclear learning objectives" faced by reinforcement learning methods, the Wasserstein-GAIL imitation learning framework proposed by the present invention, as a new control strategy generation method, has the following breakthrough points:
[0108] 1. Get rid of the constraints of model simplification: No longer rely on simplified dynamic models such as the linear inverted pendulum, fundamentally avoiding the unnaturalness of gait caused by model assumptions.
[0109] 2. Avoid the design of reward functions: By imitating human demonstration data for policy alignment, it replaces the complex and cumbersome manual reward item design process in traditional reinforcement learning.
[0110] 3. Improve the learning stability and generalization ability: After introducing the Wasserstein distance, the difference measurement between the human and robot policy distributions is more continuous and robust, effectively improving the convergence efficiency of the training process and the quality of policy generation.
[0111] In summary, this application constructs a Wasserstein-GAIL policy training mechanism, forming a straight-knee gait generation method that does not rely on modeling simplification and external reward guidance, has highly anthropomorphic characteristics and good adaptability, can meet the technical requirements of typical scenarios such as high-end services and performance interactions, and has significant engineering practical value and commercial application potential.
[0112] Based on this, this application discloses a control method, device and storage medium for a humanoid robot to walk with straight knees, which is used to improve the performance of a humanoid robot walking with straight knees in different scenarios.
[0113] Next, the technical solutions in this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of this application.
[0114] The method of this application can be applied to a server, device, terminal or other devices with logical processing capabilities. This application does not make any limitations in this regard. For the convenience of description, the following takes the execution entity as the terminal as an example for description.
[0115] Please refer to Figure 1 , this application provides an embodiment of a control method for a humanoid robot to walk with straight knees, including:
[0116] 101. Collect human gait data, where the human gait data is the gait data of a human walking with straight knees.
[0117] 102. Perform action redirection processing on the human gait data and convert it into reference motion sequence data of the target humanoid robot.
[0118] In the embodiment of the present application, it is necessary to collect gait data of a human walking with straight knees and then perform motion redirection processing. Specifically, the goal of this link is to obtain high-dimensional human gait data based on a motion capture device, and then through a series of processing processes, generate standard reference trajectory data that is adapted to the humanoid robot in terms of structure, size, and motion style for subsequent imitation learning modules to use. The specific steps of the motion redirection processing will be described in subsequent embodiments. The collection of human gait data can be sourced from the SMPL model, the AMASS database, or the MoCap system, usually a sequence of 25 to 52 3D joints. The structure is inconsistent with that of the robot, but there is a certain similarity.
[0119] 103. Model the walking task of the target humanoid robot as a velocity-conditioned Markov decision process. Under this modeling, construct an adversarial network to generate a maximum expected discounted return function according to the environmental variable distribution and decision-making strategy;
[0120] In this embodiment, a core point is to use the humanoid robot walking control strategy of Wasserstein adversarial imitation learning. Specifically, please refer to Figure 9 , Figure 9 which is a schematic diagram of the humanoid robot walking control strategy based on Wasserstein adversarial imitation learning. Among them, human demonstrations represents the natural motion reference data of the humanoid robot after motion redirection obtained from the previous step, learned robot motions represents the motion data of the humanoid robot generated during the training process. Then, the discriminator module scores the similarity of these two sets of data. If the motion data generated by the humanoid robot itself is closer to the robot reference data, it will receive a higher score reward, thereby encouraging the behavior generated by the robot during the learning process to be more similar to the natural motion parameter data, so as to achieve the effect of natural straight-knee anthropomorphic walking. The specific details will be introduced separately later. user commands is the user instruction input to the intelligent agent (Agent), and robot states is the state of the humanoid robot input to the intelligent agent. similarity reward is the similarity reward returned to the agent, generated by the corresponding reward function.
[0121] In this embodiment, the terminal models the walking task of the humanoid robot as a velocity-conditioned Markov decision process. The formal definition includes the state space, action space, target space, policy function, and reward function.
[0122] Among them, the state space includes the current pose state of the robot (such as joint angles, velocities, center-of-mass position, root pose, etc.). The action space is the target position of each joint output by the policy network. The target space is the linear velocity and angular velocity commands input by the user. The policy function represents the action distribution under the state and desired velocity. The reward function consists of two sub-items, one is the velocity tracking reward function, and the other is the style imitation reward function.
[0123] The training objective is to maximize the conditional expected cumulative reward (maximize the expected discounted return function) J(π):
[0124]
[0125] Among them, this function represents the expected total reward J(π) under the given policy π. v*~p(v) in the function J(π) represents sampling a specific environmental parameter from the distribution of environmental variables (p(v), such as task configuration), and τ~p(·|π,v*) represents generating a trajectory according to the policy π in this environment. The right sum ∑_tγ^tr(s_t,a_t,v*) is the discounted cumulative reward of this trajectory, where γ is the discount factor, r is the reward function, which depends on the current state s_t, action a_t, and environmental variable v*. Overall, J(π) measures the average performance of the policy in various possible environments.
[0126] 104. Construct a velocity tracking reward function based on the linear velocity and angular velocity of the center of mass of the target humanoid robot;
[0127] In this embodiment, it is necessary to construct a velocity tracking reward function r V , and this embodiment uses an exponential decay function to measure the deviation degree between the current actual velocity and the target velocity, specifically as follows:
[0128]
[0129] Among them, and are the current linear velocity and angular velocity of the center of mass of the robot, and are the target linear velocity and angular velocity of the center of mass of the remote controller, which come from external input. , are a set of hyperparameters used to control the importance of each tracking error, and are another set of hyperparameters used to adjust the tracking accuracy of the corresponding terms.
[0130] 105. Construct the soft boundary Wasserstein loss function of the discriminator with the gradient penalty terms sampled from the real data distribution, generated data distribution, specific distribution, and specific distribution, and construct the style reward function with the output of the discriminator;
[0131] In this embodiment, in order to make the robot behavior style close to the reference data (such as the real human gait), this embodiment introduces a Wasserstein critic network to score the style. This module is constructed based on the improved Wasserstein-1 distance, and a soft-boundary Wasserstein loss function is constructed. The following is a detailed description.
[0132] 1) The discriminator network input and feature extraction methods are as follows:
[0133] Reference data distribution: Extracted from the continuous N-frame state data in the action data;
[0134] Policy generation distribution: Also extract N frames of states;
[0135] Feature construction function: Extract local motion style-related features from the overall action data, such as joint speed distribution, step frequency, centroid offset, etc.
[0136] 2) Design of the discriminator loss function (soft-boundary Wasserstein loss function):
[0137] In this embodiment, the problem of infinite output range in the original WGAN-GP is improved, and output range compression is introduced to prevent extreme reward values and improve training stability. Therefore, the following soft-boundary Wasserstein loss function is proposed:
[0138]
[0139] In this embodiment, θ is set as the discriminator parameter, x is taken from the real data distribution P r , x~ is from the generated data distribution P g , x^ is sampled from a specific distribution P x^ . D θ (x) is the output of the discriminator for x, E is to find the expectation, tanh is the activation function, η and λ are hyperparameters, ∇ x^ is the gradient of x^.
[0140] It can be seen that this loss function is in the form of WGAN-GP. The connection with Wasserstein is that the discriminator Lipschitz continuity is ensured by weight clipping to approximately the Wasserstein distance; this function replaces weight clipping by adding a gradient penalty term, forcing the discriminator gradient norm to be close to 1 to meet the Lipschitz condition, thus stably approximating the Wasserstein distance and guiding the generator optimization. The style reward calculation method, and finally the imitation style reward is given by the following function:
[0141]
[0142] The function of this style function reward function is to ensure that if the generated action is rated as "close to the reference distribution" by the discriminator, a higher style reward will be obtained, motivating the robot to exhibit an anthropomorphic straight-knee natural walking strategy. The final target reward function is generated by weighted superposition of the speed tracking reward function and the imitation style reward.
[0143] 106. Construct a PD controller so that the PD controller converts the action output by the decision-making strategy into the torque signal of the target humanoid robot;
[0144] In this embodiment, the control execution mechanism of the terminal is the PD controller torque mapping. The action output by the decision-making strategy network is the expected position of each key joint. The following PD controller is used to convert it into the torque signal that can be used for the drive of the humanoid robot:
[0145]
[0146] Among them, τ generally represents the torque or control input quantity. θ^ usually represents the estimated value of the angle. θt is the target angle value. k p is the proportional gain coefficient, which is used to adjust the intensity of the control action related to the angle error. θt˙ is the derivative of the target angle θt with respect to time, that is, the change rate of the target angle. k d is the differential gain coefficient, which is used to adjust the intensity of the control action related to the change rate of the target angle. This formula as a whole may be used in control theory to calculate the control torque, and comprehensively determine the control quantity through the angle error and the change rate of the target angle.
[0147] 107. After using the constructed adversarial network for simulation learning, select the physical parameters of the target humanoid robot in the straight-knee walking training according to the relevance, and set the prior distribution for each physical parameter;
[0148] 108. Conduct a real-scene test movement for the target humanoid robot, and collect the operation data corresponding to the physical parameters;
[0149] 109. Combine the collected operation data with the prior distribution corresponding to the physical parameters, and feedback the updated posterior distribution to the imitation learning stage to optimize the control strategy in the imitation learning stage.
[0150] In the process of the humanoid robot moving from the simulation environment to the real environment application, the difference in its physical parameters is the key factor hindering its performance. In this embodiment, the Bayesian method is used to process the sim2real problem.
[0151] (1) Determine the physical parameters and the prior distribution
[0152] First, clarify the physical parameters of the humanoid robot that need to be optimized. In this embodiment, based on the project of straight-knee walking gait, the physical parameters of the humanoid robot cover parameters that have a greater impact on the straight-knee walking motion of the humanoid robot, such as joint friction coefficient, motor torque constant, and mass distribution parameters. Based on training experience, research data of similar robots, and the results of preliminary simulation experiments, a reasonable prior distribution is set for each of the above-determined physical parameters. For example, for the joint friction coefficient, assuming it follows a uniform distribution within a certain interval according to the material and design of mechanical components; for the motor torque constant, referring to the technical specification of the motor, a normal distribution is used to describe its possible value range.
[0153] (2) Real robot testing and data collection
[0154] Conduct a series of carefully designed test movements on a real humanoid robot. These test movements need to include single-joint movements, such as flexion and extension of the knee joint and rotation of the shoulder joint, to separately obtain information on the physical parameters related to each joint; they also need to include complex overall walking movements to comprehensively analyze the impact of multiple physical parameters on the overall movement of the robot. During the testing process, with the help of various sensors built into the robot, such as joint position sensors, force sensors, acceleration sensors, etc., rich data is collected in real time. The data types include the change of joint angles over time, the current consumption of the motor, the movement trajectory of the robot's center of mass, the forces and torques received by each joint, etc.
[0155] (3) Bayesian inference and parameter update
[0156] Use the Bayesian inference algorithm to combine the collected real data with the pre-set prior distribution to iteratively update the posterior distribution of the physical parameters. The Bayesian formula is as follows: P(θ∣D)=P(D)P(D∣θ)P(θ) where P(θ) is the prior distribution of the physical parameter θ, P(D∣θ) is the likelihood function of observing the data D under the parameter θ, P(D) is the marginal probability of the data D, and P(θ∣D) is the updated posterior distribution. In actual calculations, by calculating the likelihood function under different parameter values and combining the prior distribution, the posterior distribution is obtained. For example, using the joint angle and motor current data during the robot's walking, calculate the possibility of this data occurring under different assumptions of friction coefficient and torque constant, and then update the posterior distribution of these parameters.
[0157] (4) Parameter feedback and strategy optimization
[0158] Feed the updated physical parameters back to the imitation learning stage. In the adversarial imitation learning based on the Wasserstein distance, re-optimize the control policy using the new physical parameters. For example, when calculating the dynamic model of the robot, adopt the updated mass distribution parameters and joint friction coefficient, so that the straight-knee walking motion generated by the policy network is more in line with the physical characteristics of the real robot. At the same time, in the calculation of the speed tracking reward function and the style imitation reward function, consider the new physical parameters, adjust the weights and calculation methods of the rewards, and guide the robot to learn a straight-knee walking strategy that is more adaptable to the real environment. By continuously repeating the process of real machine testing, data collection, Bayesian inference, and policy optimization, gradually approach the optimal physical parameters and control strategies in the real environment, and ensure that the robot can achieve stable and natural straight-knee walking in the real scenario.
[0159] In this embodiment, first collect human gait data, where the human gait data is the gait data of a human walking with straight knees. Perform action redirection processing on the human gait data and convert it into the reference motion sequence data of the target humanoid robot. Model the walking task of the target humanoid robot as a velocity-conditioned Markov decision process, construct an adversarial network under this modeling, and generate the maximum expected discounted return function according to the environmental variable distribution and decision-making strategy. Construct a speed tracking reward function based on the linear velocity and angular velocity of the centroid of the target humanoid robot. Construct the soft-boundary Wasserstein loss function of the discriminator with the gradient penalty terms of the real data distribution, the generated data distribution, a specific distribution, and a specific distribution sampling, and construct a style reward function with the output of the discriminator. Construct a PD controller so that the PD controller converts the action output by the decision-making strategy into the torque signal of the target humanoid robot. After the simulation learning using the constructed adversarial network, select the physical parameters of the target humanoid robot in the straight-knee walking training according to the relevance, and set the prior distribution for each physical parameter. Conduct a real-scene test motion for the target humanoid robot and collect the operation data corresponding to the physical parameters. Combine the collected operation data with the prior distribution corresponding to the physical parameters, and feed the updated posterior distribution back to the imitation learning stage to optimize the control strategy in the imitation learning stage.
[0160] By introducing the Generative Adversarial Imitation Learning (GAIL) framework and using real human gait data as a demonstration, the decision-making strategy of the humanoid robot is trained to approximate human gait characteristics in the behavior space, enabling it to directly acquire the ability to walk with straight knees naturally without the need for reward function design. On this basis, to further improve the stability and accuracy of the imitation effect, the Wasserstein distance based on the optimal transport theory is introduced to construct an improved GAIL discriminator. The Wasserstein distance has better gradient properties and stronger distribution fitting ability during the training process, which can enhance the sensitivity of policy learning to subtle motion differences. This improvement enables the robot to not only capture the macroscopic morphological characteristics of human gait but also more accurately restore its dynamic coordination and delicate rhythm. On this basis, physical parameters are inferred according to prior knowledge and the operating data of the humanoid robot to reduce the impact of the difference between simulation and reality. Collaborating with existing improvements, it provides more accurate parameter support for straight-knee gait generation, further enhancing the performance of the humanoid robot walking with straight knees in different real environments.
[0161] Secondly, this embodiment also has the following beneficial effects:
[0162] 1. Comprehensiveness of the human-to-robot action redirection process: It integrates multiple modules such as bone topology alignment, coordinate transformation, inverse kinematics optimization, and temporal reconstruction to form a complete conversion process from human actions to robot actions.
[0163] 2. Design of the multi-objective inverse kinematics optimizer: Considering end-effector accuracy, joint orientation, physical limitations, and motion smoothness simultaneously to ensure that the output actions are both natural and executable.
[0164] 3. Introduction of the Wasserstein adversarial imitation framework: Replacing the JS divergence loss function in traditional GAIL, effectively alleviating the problems of unstable training and vanishing gradients.
[0165] 4. Design of the soft-boundary Wasserstein loss function: Controlling the reward output range to prevent training failure caused by zero or highly fluctuating rewards, and enhancing the stability and expressiveness of imitation learning.
[0166] 5. Joint training mechanism of style reward and speed reward: By fusing imitation rewards and speed tracking rewards, the robot can not only imitate the human action style but also accurately track the target speed.
[0167] Please refer to Figure 2 , this application provides an embodiment of a method for action redirection processing, including:
[0168] 201. Perform skeleton topology unification and bone binding processing on the human skeleton and the humanoid robot skeleton to generate an original skeleton, and there are several key joints on the original skeleton;
[0169] In this embodiment, after the terminal obtains the gait data of a human walking with straight knees, it is necessary to unify according to the skeletal structures of the human body and the humanoid robot so that subsequent data processing can be adapted to the humanoid robot. The terminal performs skeletal topology unification and bone binding processing on the human skeleton and the humanoid robot skeleton to generate an original skeleton, where there are several key joints on the original skeleton. The specific method for generating the original skeleton will be described in subsequent embodiments.
[0170] 202. Perform coordinate system unification processing and root normalization processing based on the human gait data and the original skeleton to generate the key joint data of the humanoid robot;
[0171] In this embodiment, the human gait data is generally represented in the world coordinate system, while the humanoid robot control requires a local coordinate. Therefore, the terminal needs to perform coordinate system unification processing and root normalization processing based on the human gait data and the original skeleton to generate the key joint data of the humanoid robot. First, use the SE(3) homogeneous transformation to convert each frame of human pose data into a local coordinate with the pelvis as the root. Next, the rotation data is expressed using quaternions or rotation vectors to ensure interpolation smoothness. Finally, all joint rotations are represented relative to the skeleton parent node so that it can be directly used as the target value for subsequent kinematic solutions.
[0172] 203. Perform multi-objective inverse kinematics optimization solution on the key joint data to generate a reference motion sequence. The multi-objective inverse kinematics optimization solution is used to map the Cartesian positions of the key joints and the poses of the end effectors to the corresponding joint angular directions.
[0173] In this embodiment, the terminal performs multi-objective inverse kinematics optimization solution on the key joint data to generate the corresponding reference motion sequence, where the multi-objective inverse kinematics optimization solution is used to map the Cartesian positions of the key joints and the poses of the end effectors to the corresponding joint angular directions. The specific method will be described in subsequent embodiments.
[0174] Please refer to Figure 3 , an embodiment of a method for generating an original skeleton provided by this application includes:
[0175] 301. Construct the kinematic trees of the human skeleton and the humanoid robot skeleton;
[0176] 302. Merge the kinematic trees of the human skeleton and the humanoid robot skeleton to generate an original skeleton, and retain one bone between the key joints;
[0177] 303. Select the key joints and record the length of each bone segment.
[0178] In this embodiment, the input human gait data is usually a sequence of 25 to 52 3D joint points. Although the structure is inconsistent with that of the robot, there is a certain similarity. The human skeleton and the humanoid robot skeleton can be abstracted as homeomorphic graphs. Utilizing this property, in this embodiment, we propose an intermediate skeleton structure called the original skeleton, which retains the geometric and hierarchical characteristics of the skeleton. On this basis, first, the kinematic trees of the human skeleton and the robot skeleton are constructed, and then the two are merged to generate a unified original skeleton, retaining only one bone between the key nodes. The user manually selects n key joints, such as "hip joint - knee joint - ankle joint" for the leg and "shoulder - elbow - wrist" for the arm, etc.; and records the length ratio of each bone segment.
[0179] Please refer to Figure 4 , an embodiment of a method for generating a reference motion sequence provided by this application includes:
[0180] 401. Construct a key joint position matching loss function according to the actual positions of the key joints of the humanoid robot and the desired target positions of the key joints in the key joint data;
[0181] 402. Construct an end - effector attitude matching loss function according to the actual pose of the end - effector of the humanoid robot and the desired target pose in the key joint data;
[0182] 403. Construct a joint minimum displacement loss function according to the joint angle data between adjacent frames in the key joint data;
[0183] 404. Generate an objective function according to the key joint position matching loss function, the end - effector attitude matching loss function, and the joint minimum displacement loss function;
[0184] 405. Perform multi - objective inverse kinematics optimization on the key joint data according to the objective function to generate a reference motion sequence.
[0185] In order to map the Cartesian positions of the key joints and the pose of the end - effector to the corresponding joint angles, in this embodiment, the full - body inverse kinematics is modeled as a gradient - based multi - objective optimization problem. This optimization problem includes the following three objectives:
[0186]
[0187] The specific meanings of the above three objectives C1, C2, and C3 are as follows:
[0188] C1 is the key joint position matching loss (the smaller the better), which is used to ensure that the positions of the key joints of the humanoid robot are as close as possible to the target positions in the human gait data. Among them r P kRepresents the expected target positions of the key joints (elbows, knees, shoulders, etc.) of the robot obtained from human body data, p k (θ) represents the actual positions of the key joints of the robot calculated based on the current robot joint angles, and k is the number of key joints.
[0189] C2 is the end effector pose matching loss (the smaller the better), which is used to control the position and pose of the end parts such as the hand, sole of the foot, and head to reproduce the original movement as much as possible. Among them r P e Represents the expected target position and pose of the robot end obtained from human body data, p e (θ) represents the actual position and pose of the end effector calculated based on the current robot joint angles.
[0190] C3 is the minimum joint displacement loss, whose function is to encourage the change of joint angle θ between adjacent frames to be as small as possible to maintain the continuity and smoothness of the movement. For a humanoid robot with a high degree of redundancy, C3 helps to select a more natural motion sequence in the case of multiple sets of solutions.
[0191] Finally, the generated objective function C is the weighted sum of the three objective functions. The weighting k needs to be adjusted according to specific situations (the specific description of the embodiments will be carried out later). Then, through continuous iteration and optimization of the joint angles, this (error) is made as small as possible. The form is as follows:
[0192]
[0193] where k i are the weights of each sub-objective corresponding to the key joint positions, end poses, and minimum displacement losses respectively. In the above optimization process, the pre-designed value range of the joint angles and the joint speed constraint conditions need to be satisfied to ensure that the generated data is more in line with the movement of the humanoid robot and also conforms to the human gait data.
[0194] Please refer to Figure 5 , this application provides an embodiment of a method for generating an objective function, including:
[0195] 501. Determine the acquisition scene parameters of the human gait data;
[0196] In this embodiment, it is necessary to first determine the scene for collecting human gait data and record some corresponding environmental parameters. The standard scene is a horizontal wooden floor and a windless environment. Environmental parameter calibration: Measure the ground flatness with a laser rangefinder, and control the error within ±0.5 mm.
[0197] In this embodiment, the height range of the subjects for human gait data is 160 - 190 cm, and the weight is 50 - 90 kg. They wear tight sports clothes to reduce clothing interference. Before collection, a 10 - minute dynamic warm - up is required, and 3 standardized gait tests (such as walking straight with straight knees, turning, etc.) are completed to establish baseline data.
[0198] 502. Conduct texture contrast analysis and slope contrast analysis on the collected scene parameters and flat - ground scene parameters to generate gait deformation data.
[0199] Next, the terminal conducts texture contrast analysis and slope contrast analysis on the collected scene parameters and flat - ground scene parameters to generate gait deformation data. Specifically, first, conduct texture contrast analysis. Taking wooden floor as the standard texture, establish a ground friction coefficient model (μ = F_friction / F_normal), and measure the friction characteristics of different materials (such as carpet μ = 0.3, ceramic tile μ = 0.6) through a pressure distribution tester.
[0200] Next, conduct slope contrast analysis. Taking the horizontal ground as the standard, construct a slope geometric model, and obtain the slope inclination angle (θ) through 3D laser scanning.
[0201] Through texture contrast analysis and slope contrast analysis, human gait deformation data is generated. Specifically, it can be directly based on empirical data, look up the deformation data in a table. Different textures and slopes cause human gait deformation, which can be obtained through historical data.
[0202] In this embodiment, to ensure accuracy, the collected human gait data is used for label generation. Specifically, the results of texture contrast analysis and slope contrast analysis are used as labels, and then the labeled human gait data is input into a pre - trained gait biomechanics model (such as OpenSim). Through inverse dynamics calculation of joint torque changes, different key joint deformation data, including rotational deformation and angular velocity deformation, are generated.
[0203] 503. Determine the key joints that generate deformation on the original skeleton according to the gait deformation data.
[0204] 504. Generate key joint position weights for the key joints that generate deformation.
[0205] Next, the terminal determines the key joints that generate deformation on the original skeleton according to the key joint deformation data (gait deformation data), and generates key joint position weights for the key joints that generate deformation. The greater the deformation, the greater the adjusted weight. This method analyzes the deformation of each key joint under scene transformation through a gait biomechanics model, and adjusts the weights of the joints accordingly, so as to increase the adaptability of the subsequent generated data to the real - collected scene and improve the subsequent training effect.
[0206] 505. Determine the pose change degree parameter according to the gait deformation data, and generate the end pose weight according to the pose change degree parameter.
[0207] In this embodiment, the terminal determines the pose change degree parameter according to the gait deformation data. Similarly, this pose change degree parameter needs to analyze the change of the human gait data after the end of this frame of action compared with the standard scenario under the same action, and generate the end pose weight according to the pose change degree parameter.
[0208] 506. Determine the minimum displacement loss weight.
[0209] Since the minimum displacement loss weight is not affected by the scenario, the same weight parameter can be used without change.
[0210] 507. Generate an objective function according to the key joint position matching loss function, the end effector pose matching loss function, the joint minimum displacement loss function, the key joint position weight, the end pose weight, and the minimum displacement loss weight.
[0211] Generate an objective function according to the key joint position matching loss function, the end effector pose matching loss function, the joint minimum displacement loss function, the key joint position weight, the end pose weight, and the minimum displacement loss weight. Its formula is similar to that in step 405 and will not be elaborated here. In this embodiment, by analyzing the change of the scenario, the deformation amount of the straight-knee gait data in different scenarios is determined, and then the deformation amount of the straight-knee gait data is decomposed into the deformation of the joint and the deformation of the pose. And finally, the corresponding weight parameters are adjusted so that after the subsequent multi-objective inverse kinematics optimization solution, it can be more suitable for this scenario and improve the subsequent training effect.
[0212] Please refer to Figure 6 , an embodiment of a method for processing a reference motion sequence provided by this application includes:
[0213] 601. Perform trajectory quality optimization processing on the reference motion sequence.
[0214] In this embodiment, after solving the entire joint angle sequence, in order to ensure the quality of the data, the terminal needs to further optimize the trajectory quality through the following steps:
[0215] First, calculate the linear velocity and angular velocity of the root and each joint using the difference between adjacent frames.
[0216] Then, perform linear interpolation on the position and spherical linear interpolation (Slerp) on the direction to generate a continuous trajectory between discrete frames.
[0217] Finally, the exponential moving average filter (EMA) is used to smooth the joint positions and velocities, eliminate jumps and abnormal spikes, and improve the naturalness of movements and the control robustness.
[0218] Please refer to Figure 7 , an embodiment of a control device for a humanoid robot to walk with straight knees provided by this application includes:
[0219] The first acquisition unit 701 is configured to acquire human gait data, and the human gait data is the gait data of a human walking with straight knees;
[0220] The redirection unit 702 is configured to perform action redirection processing on the human gait data and convert it into reference motion sequence data of the target humanoid robot;
[0221] Optionally, the redirection unit 702 includes:
[0222] The first generation module is configured to perform skeleton topology unification and bone binding processing on the human skeleton and the humanoid robot skeleton to generate an original skeleton, and there are several key joints on the original skeleton;
[0223] Optionally, the first generation module includes: constructing the kinematic trees of the human skeleton and the humanoid robot skeleton; performing skeleton merging on the kinematic trees of the human skeleton and the humanoid robot skeleton to generate an original skeleton, and retaining one bone between the key joints; selecting the key joints and recording the length of each section of the bone.
[0224] The second generation module is configured to perform coordinate system unification processing and root normalization processing based on the human gait data and the original skeleton to generate key joint data of the humanoid robot;
[0225] The third generation module is configured to perform multi-objective inverse kinematics optimization solution on the key joint data to generate a reference motion sequence, and the multi-objective inverse kinematics optimization solution is used to map the Cartesian positions of the key joints and the poses of the end effectors to the corresponding joint angular directions.
[0226] Optionally, the third generation module includes:
[0227] The first construction sub-module is configured to construct a key joint position matching loss function based on the actual position of the key joint of the humanoid robot and the expected target position of the key joint in the key joint data;
[0228] The second construction sub-module is configured to construct an end effector attitude matching loss function based on the actual pose of the end effector of the humanoid robot and the expected target pose in the key joint data;
[0229] The third construction sub-module is configured to construct a joint minimum displacement loss function based on the joint angle data between adjacent frames in the key joint data;
[0230] The first generation sub-module is used to generate an objective function according to the key joint position matching loss function, the end effector pose matching loss function, and the joint minimum displacement loss function;
[0231] Optionally, the first generation sub-module includes: determining the acquisition scene parameters of the human gait data; performing texture contrast analysis and slope contrast analysis on the acquisition scene parameters and the flat scene parameters to generate gait deformation data; determining the key joints that generate deformation on the original skeleton according to the gait deformation data; generating key joint position weights for the key joints that generate deformation; determining the pose change degree parameters according to the gait deformation data, and generating end pose weights according to the pose change degree parameters; determining the minimum displacement loss weight; generating an objective function according to the key joint position matching loss function, the end effector pose matching loss function, the joint minimum displacement loss function, the key joint position weights, the end pose weights, and the minimum displacement loss weight.
[0232] The second generation sub-module is used to perform multi-objective inverse kinematics optimization on the key joint data according to the objective function to generate a reference motion sequence.
[0233] The optimization module is used to perform trajectory quality optimization on the reference motion sequence.
[0234] The first construction unit 703 is used to model the walking task of the target humanoid robot as a Markov decision process with speed conditions, construct an adversarial network under this modeling, and generate a maximum expected discounted return function according to the environmental variable distribution and the decision-making strategy;
[0235] The second construction unit 704 is used to construct a speed tracking reward function with the linear velocity and angular velocity of the centroid of the target humanoid robot;
[0236] The third construction unit 705 is used to construct a soft boundary Wasserstein loss function of the discriminator with the gradient penalty terms of the real data distribution, the generated data distribution, the specific distribution, and the specific distribution sampling, and construct a style reward function with the output of the discriminator;
[0237] The fourth construction unit 706 is used to construct a PD controller, so that the PD controller converts the action output by the decision-making strategy into a torque signal of the target humanoid robot;
[0238] The setting unit 707 is used to select the physical parameters of the target humanoid robot in the straight-knee walking training according to the relevance after completing the simulation learning using the constructed adversarial network, and set a prior distribution for each physical parameter;
[0239] The second acquisition unit 708 is configured to perform a live test movement for the target humanoid robot and acquire the operation data corresponding to the physical parameters;
[0240] The optimization unit 709 is configured to combine the acquired operation data with the prior distribution corresponding to the physical parameters, and feedback the updated posterior distribution to the imitation learning stage to optimize the control strategy in the imitation learning stage.
[0241] Please refer to Figure 8 , this application provides a control device for a humanoid robot to walk with straight knees, including:
[0242] A processor 801, a memory 802, an input / output unit 803, and a bus 804.
[0243] The processor 801 is connected to the memory 802, the input / output unit 803, and the bus 804.
[0244] The memory 802 stores a program, and the processor 801 calls the program to execute the control methods as described in Figure 1 , Figure 2 and Figure 3 , Figure 4 , Figure 5 and Figure 6 .
[0245] This application provides a computer-readable storage medium, on which a program is stored, and when the program is executed on a computer, it executes the control methods as described in Figure 1 , Figure 2 and Figure 3 , Figure 4 , Figure 5 and Figure 6 .
[0246] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0247] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0248] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0249] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0250] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
Claims
1. A control method for a humanoid robot to walk with straight knees, characterized in that, Including: Collecting human gait data, where the human gait data is the gait data of a human walking with straight knees; Performing action redirection processing on the human gait data and converting it into reference motion sequence data of a target humanoid robot; Modeling the walking task of the target humanoid robot as a Markov decision process with a speed condition, constructing an adversarial network under this modeling, and generating a maximum expected discounted return function according to the environmental variable distribution and decision-making strategy; Constructing a speed tracking reward function based on the linear velocity and angular velocity of the centroid of the target humanoid robot; Constructing a soft-boundary Wasserstein loss function of the discriminator with gradient penalty terms of real data distribution, generated data distribution, specific distribution, and specific distribution sampling, and constructing a style reward function based on the output of the discriminator; Constructing a PD controller so that the PD controller converts the actions output by the decision-making strategy into torque signals of the target humanoid robot; After performing simulation learning using the constructed adversarial network, selecting the physical parameters of the target humanoid robot in the straight-knee walking training according to the relevance and setting a prior distribution for each physical parameter; Performing a real-scene test movement on the target humanoid robot and collecting the operation data corresponding to the physical parameters; Combining the collected operation data with the prior distribution corresponding to the physical parameters, and feeding back the updated posterior distribution to the imitation learning stage to optimize the control strategy in the imitation learning stage.
2. The control method according to claim 1, wherein Performing action redirection processing on the human gait data and converting it into reference motion sequence data of a target humanoid robot, including: Performing skeleton topology unification and bone binding processing on the human skeleton and the humanoid robot skeleton to generate an original skeleton, and there are several key joints on the original skeleton; Performing coordinate system unification processing and root normalization processing according to the human gait data and the original skeleton to generate key joint data of the humanoid robot; Performing multi-objective inverse kinematics optimization solution on the key joint data to generate a reference motion sequence, and the multi-objective inverse kinematics optimization solution is used to map the Cartesian position of the key joint and the pose of the end effector to the corresponding joint angular direction.
3. The control method according to claim 2, wherein The step of performing skeleton topology unification and bone binding processing on the human skeleton and the humanoid robot skeleton to generate an original skeleton includes: Constructing kinematic trees of the human skeleton and the humanoid robot skeleton; Performing skeleton merging on the kinematic trees of the human skeleton and the humanoid robot skeleton to generate an original skeleton, and retaining one bone between the key joints; Selecting key joints and recording the length of each bone segment.
4. The control method according to any one of claims 2 to 3, characterized in that The step of performing multi-objective inverse kinematics optimization solution on the key joint data to generate a reference motion sequence includes: Constructing a key joint position matching loss function according to the actual position of the key joint of the humanoid robot and the expected target position of the key joint in the key joint data; Constructing an end effector pose matching loss function according to the actual pose of the end effector of the humanoid robot and the expected target pose in the key joint data; Constructing a joint minimum displacement loss function according to the joint angle data between adjacent frames in the key joint data; Generate an objective function based on the key joint position matching loss function, the end effector pose matching loss function, and the joint minimum displacement loss function; Perform multi-objective inverse kinematics optimization on the key joint data according to the objective function to generate a reference motion sequence.
5. The control method according to claim 4, characterized in that, The step of generating an objective function based on the key joint position matching loss function, the end effector pose matching loss function, and the joint minimum displacement loss function includes: Determine the acquisition scene parameters of the human gait data; Perform texture comparison analysis and slope comparison analysis on the acquisition scene parameters and the flat ground scene parameters to generate gait deformation data; Determine the key joints that generate deformation on the original skeleton according to the gait deformation data; Generate key joint position weights for the key joints that generate deformation; Determine the pose change degree parameter according to the gait deformation data, and generate the end pose weight according to the pose change degree parameter; Determine the minimum displacement loss weight; Generate an objective function according to the key joint position matching loss function, the end effector pose matching loss function, the joint minimum displacement loss function, the key joint position weight, the end pose weight, and the minimum displacement loss weight.
6. The control method according to claim 4, wherein After the step of performing multi-objective inverse kinematics optimization on the key joint data to generate a reference motion sequence, and before the step of modeling the target humanoid robot walking task as a Markov decision process with a speed condition, constructing an adversarial network under this modeling, and generating a maximum expected discounted return function according to the environmental variable distribution and the decision-making strategy, the control method further includes: Perform trajectory quality optimization processing on the reference motion sequence.
7. A control device for a humanoid robot to walk with straight knees, characterized in that, Including: A first acquisition unit for acquiring human gait data, where the human gait data is the gait data of a human walking with straight knees; A redirection unit for performing action redirection processing on the human gait data and converting it into reference motion sequence data of the target humanoid robot; A first construction unit for modeling the target humanoid robot walking task as a Markov decision process with a speed condition, constructing an adversarial network under this modeling, and generating a maximum expected discounted return function according to the environmental variable distribution and the decision-making strategy; A second construction unit for constructing a speed tracking reward function with the linear velocity and angular velocity of the centroid of the target humanoid robot; A third construction unit for constructing a soft boundary Wasserstein loss function of the discriminator with the gradient penalty terms of the real data distribution, the generated data distribution, a specific distribution, and a specific distribution sampling, and constructing a style reward function with the output of the discriminator; A fourth construction unit for constructing a PD controller so that the PD controller converts the action output by the decision-making strategy into a torque signal of the target humanoid robot; A setting unit for selecting the physical parameters of the target humanoid robot in the straight-knee walking training according to the relevance after completing the simulation learning using the constructed adversarial network, and setting a prior distribution for each physical parameter; A second acquisition unit for performing a real-scene test motion on the target humanoid robot and acquiring the operation data corresponding to the physical parameters; Optimization unit, configured to combine the collected operation data with the prior distribution corresponding to the physical parameters, and feed back the updated posterior distribution to the imitation learning stage to optimize the control strategy in the imitation learning stage.
8. The control device according to claim 7, characterized in that Redirection unit, comprising: First generation module, configured to perform skeleton topology unification and bone binding processing on the human skeleton and the humanoid robot skeleton to generate an original skeleton, and there are several key joints on the original skeleton; Second generation module, configured to perform coordinate system unification processing and root normalization processing according to the human gait data and the original skeleton to generate key joint data of the humanoid robot; Third generation module, configured to perform multi-objective inverse kinematics optimization solution on the key joint data to generate a reference motion sequence, and the multi-objective inverse kinematics optimization solution is used to map the Cartesian position of the key joint and the pose of the end effector to the corresponding joint angular direction.
9. A control device for a humanoid robot to walk with straight knees, characterized in that, Comprising: A processor, a memory, an input / output unit, and a bus, the processor is connected to the memory, the input / output unit, and the bus, the memory stores a program, and the processor calls the program to execute the control method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, A program is stored on the computer-readable storage medium, and when the program is executed on a computer, it executes the control method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Humanoid robot torque control method based on position ring pre-training
CN118809606A
Multi-gait biped motion control method and device based on deep reinforcement learning
CN118938645A
Cited By
Humanoid robot reinforcement learning gait control method and system
CN121050241A