Quadruped robot basic behavior control method and device and medium
Through semi-supervised learning and generative adversarial network optimization, the four-legged robot can learn diverse and natural behavioral patterns, solve the problems of poor behavioral diversity and task adaptability, and achieve efficient behavior control and rapid response.
Patent Information
- Application Number
- CN202510457641.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
The existing four-legged robots have a large gap in behavioral diversity, task adaptability, and the simulator and the real world, and traditional control methods are difficult to achieve efficient adaptation.
By collecting behavioral data of real animals, using semi-supervised learning and generative adversarial networks, combining variable lower bounds to maximize mutual information and discriminant clustering technology, optimize the stability of the generative adversarial networks and learn diverse and natural behavioral patterns.
It significantly improves the behavioral diversity and data utilization efficiency of four-legged robots, realizes precise control and rapid response to behavior patterns, and reduces the need for large-scale labeling of data.
Smart Images

Figure CN120295111A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control technology, and particularly relates to a basic behavior control method, device and medium for a quadruped robot. Background Art
[0002] With the continuous development of robot technology, quadruped robots are increasingly widely used in complex environments. However, existing quadruped robots still have deficiencies in aspects such as behavior diversity, task adaptability, and the gap between the simulator and the real world. Traditional control methods often have difficulty achieving efficient adaptation to complex environments, and there are significant performance differences during the transfer between the simulator and the real world. Summary of the Invention
[0003] This application provides a basic behavior control method, device and medium for a quadruped robot, and its advantage is to learn the diverse and natural behavior patterns of real quadruped animals, thereby significantly improving the performance of quadruped robots in terms of behavior diversity, data utilization efficiency and behavior controllability.
[0004] The technical solution of this application is as follows:
[0005] On the one hand, this application provides a basic behavior control method for a quadruped robot, including the following steps:
[0006] (1) Motion capture data preparation: Collect the behavior data of real animals and map it onto the skeleton of the target robot;
[0007] (2) Semi-supervised learning of discrete behavior patterns: Learn discrete behavior patterns, including walking, pacing, trotting, jogging and jumping, by maximizing the mutual information of the variational lower bound;
[0008] (3) Unsupervised learning of continuous behavior styles: Capture the continuous style changes in behavior by maximizing the mutual information between continuous variables and behavior data;
[0009] (4) Imbalanced data optimization: Adopt discriminative clustering technology and combine it with the method of maximizing regularization information to extract intrinsic information from imbalanced data;
[0010] (5) Adversarial training and stability enhancement: Introduce a gradient penalty term to optimize the stability of the generative adversarial network.
[0011] Further, in the step (1) of performing motion capture data preparation, it includes the following steps:
[0012] (1.1) Collect the motion capture data of a real animal dog, including five behavior patterns: walking, pacing, trotting, jogging, and jumping; the motion capture data includes labeled data and unlabeled data. The labeled motion capture data clearly indicates the behavior pattern to which the motion capture data belongs; the unlabeled motion capture data has no annotation of the behavior pattern.
[0013] (1.2) Use motion retargeting technology to map the original skeletal animation in the motion capture data to a quadruped robot; define the center of mass and five key points of the quadruped on the original skeleton, and map them to the robot skeleton, and use inverse kinematics to calculate the joint rotation to obtain the angular positions of the joints of the quadruped robot.
[0014] Further, in step (2), the semi-supervised mutual information is maximized by the following method:
[0015] (2.1) Formulate the control problem of the quadruped robot as a discrete-time dynamics problem satisfying a partially observable Markov process: at each time step t, use the policy π to control the robot to execute an action, which causes the robot to transfer to the next state; on this basis, use the latent skill variable c to represent five common animal dog behavior patterns: walking, pacing, trotting, jogging, and jumping;
[0016] Maximize the mutual information between c and the state of the quadruped robot, denoted as I(c; o I ), where o I includes the pitch angle, roll angle, yaw angular velocity, joint angles, joint velocities, and foot contact states of the quadruped robot; maximize the mutual information between the true label and the state of the quadruped robot, denoted as I(y; o I ), where y is the true label of the motion capture data;
[0017] Subsequently, approximate the mutual information by calculating the variational lower bound and , and the specific form is:
[0018]
[0019] where Q(·|o I ) is an estimate of the true posterior probability P(·|o I ), and the variational lower bound is tight when Q = P; d EL is the state transition distribution of the labeled motion capture data; d π is the state transition distribution of the quadruped robot obtained by the policy π; p(y) and p(c) represent the prior probabilities of y and c respectively, which are independent of the policy π; and represent the entropies of y and c respectively;
[0020] Let \(Q1 = Q2\), denoted as \(Q\). c to obtain the final semi - supervised regularization term \(L\). SS :
[0021]
[0022] (2.2) In equation (3), is the supervised term, \(Q\). c is implemented using a neural network, and \(d\). EL is used to predict \(y\), and \(Q\) is optimized according to the cross - entropy loss function. c ; is the unsupervised term, \(Q\). c is fixed, and according to \(d\). π learns the latent semantic variables of \(y\), and reinforcement learning is used to optimize the policy \(\pi\). The output of \(Q\). c will be used as the semi - supervised imitation reward \(r\). ss Its form is:
[0023] r SS =\(\log Q\). c (c|o I ). (19)
[0024] Furthermore, in step (3), the unsupervised mutual information is maximized in the following way:
[0025] (3.1) The motion capture data contains data of five behavioral patterns: walking, pacing, trotting, jogging, and jumping. There are also style variations within the same behavioral pattern; a continuous variable \(\in\) is introduced to capture this internal continuous style variation. The mutual information \(I(\in; o\). I ) is maximized to achieve this; the variational lower bound \(L\). US (Q ∈ ) is used to approximate the mutual information, and the calculation method is:
[0026]
[0027] where \(\in\) is sampled from the uniform distribution \(U(-1,1)\), and \(Q\). ∈ (\(\in|o\). I ) is an estimate of the true posterior probability;
[0028] (3.2) In equation (5), \(L\). US is the unsupervised term, which is used as the objective function to optimize \(Q\). ∈ At the same time, the unsupervised imitation reward \(r\). ∈ is calculated according to the output of \(Q\). US and reinforcement learning is used to optimize the policy \(\pi\). The specific form is:
[0029] r US =\(\log Q\)∈ (∈|o I )。 (21)
[0030] Furthermore, in step (4), the unbalanced data is optimized in the following way:
[0031] (4.1) Adjust the sampling distribution of the latent variable c to align it with the state transition distribution of the unbalanced data; specifically, use Q c Predict the empirical label distribution based on the unlabeled data where N is the total number of unlabeled data, is the predicted label, and ψ is the parameter of Q c ; then use the empirical label distribution as the sampling distribution of the latent skill variable c;
[0032] (4.2) Use the regularized information maximization method to calculate the penalty term L RIM , and automatically identify the boundaries between various behavior patterns in the unlabeled motion capture data. The specific form is:
[0033]
[0034] where the first term is the clustering hypothesis, represents entropy; the second term is to avoid the degenerate solution by minimizing the KL divergence between and the distribution p(c) of the latent variable c; the third term R(ψ) is the parameter regularization performed on ψ to avoid complex solutions.
[0035] Furthermore, in step (5), adversarial training and stability enhancement are performed in the following way:
[0036] (5.1) The overall framework of the algorithm training is optimized based on the generative adversarial imitation learning method; minimize a discriminator objective L similar to the least squares generative adversarial network GAIL : where D is the discriminator;
[0037] (5.2) Introduce a gradient penalty term L GP to improve the training stability. The specific form is:
[0038]
[0039] where D is the discriminator and φ is the discriminator parameter;
[0040] (5.3) Combine all the optimization terms mentioned above to obtain the integrated objective function of the algorithm, which is used as the discriminator D and estimators Q c and Q ∈The optimization objective is updated by the gradient descent method:
[0041]
[0042] Among them, the first term is a discriminator objective similar to that of the least squares generative adversarial network, that is
[0043] (5.4) After updating the discriminator D and the estimator Q c and Q ∈ the policy π is updated by maximizing the total reward through the proximal policy optimization algorithm in reinforcement learning; the total reward includes the imitation reward r T and the task reward r T :
[0044] r = w I r I + w T r T (25)
[0045] The imitation reward includes the discriminator reward, the semi-supervised imitation reward and the unsupervised imitation reward:
[0046] r I = r D + r SS + r US (26)
[0047] Among them, the first term r D is the discriminator reward. For the samples generated by the policy of the quadruped robot, the discriminator objective is to predict its score as -1; while for the demonstration samples in the motion capture data, the objective is to predict its score as 1. The specific reward formula is r D = max[0, 1 - 0.25(D(o I ) - 1) 2 , the semi-supervised imitation reward r SS and the unsupervised imitation reward r US are given by equations (4) and (6) respectively;
[0048] The task reward includes the linear velocity tracking reward the angular velocity tracking reward the jump height reward and the stable height reward The specific form is as follows:
[0049]
[0050] Among them, and v xy represent the command and the actual linear velocity; and ω zRepresent the command and the actual angular velocity; h cmd and h represent the command and the actual center height, is an indicator function;
[0051] (5.5) Continuously repeat the process of adversarial training, and cyclically update the discriminator, estimator, and policy until convergence; finally, obtain the policy π for generating the actions of five behavioral patterns, realizing the imitation of the natural behavioral patterns of the real animal dog by the quadruped robot.
[0052] On the other hand, the present application provides a basic behavior control device for a quadruped robot, including a memory and a processor. When the computer program stored in the memory is called and executed by the processor, the method described above is implemented.
[0053] On the other hand, the present application provides a computer-readable medium. When the computer program stored in the computer-readable medium is called and executed by a computer, the method described above is implemented.
[0054] In summary, the beneficial effects of the present application are as follows: (1) Through the semi-supervised learning method, the method provided by the present application can learn diverse and natural behavioral patterns, including discrete behavioral categories and continuous behavioral style changes. (2) Compared with the fully supervised method, the present invention significantly reduces the need for large-scale labeled data and realizes efficient learning by using unlabeled data and a small amount of labeled data. (3) By introducing latent variables, the method of the present application can achieve precise control of behavioral patterns and can quickly respond to speed commands to adapt to different task requirements. Brief Description of the Drawings
[0055] Figure 1 It is a training schematic diagram of the integrated control system of the quadruped robot;
[0056] Figure 2 It is a method flow chart of BBC;
[0057] Figure 3 It is a schematic diagram of using the action remapping technology to process the motion capture data;
[0058] Figure 4 It is a schematic diagram of the five behavioral patterns that the quadruped robot needs to imitate the real animal dog;
[0059] Figure 5 It is a method flow chart of TSC;
[0060] Figure 6 It is a schematic diagram of the terrain elevation map, privileged information, and depth image;
[0061] Figure 7 It is a comparison schematic diagram of robust optimization of the depth image;
[0062] Figure 8 Flow chart optimized for the emulator Specific implementation mode
[0063] The following will describe in detail the specific implementation mode of the present application with reference to the accompanying drawings
[0064] As Figure 1 shown, an integrated quadruped robot leg controller includes the following steps
[0065] (1) Use the basic behavior controller (BBC) to learn diverse and natural behavior patterns
[0066] (2) Use the task-specific controller (TSC) to generate control instructions to achieve efficient control of the quadruped robot
[0067] (3) Use the evolutionary strategy simulator optimization method to optimize the simulator parameters and narrow the gap between the simulator and the real world
[0068] In the above step (1), as Figure 2 shown, a basic behavior controller (BBC) for a quadruped robot includes the following steps
[0069] Step 1: Mo-cap data preparation: Collect the behavior data of real animals and map it onto the target robot skeleton
[0070] (1.1) Collect the motion capture data of real animal dogs, including five behavior patterns: walking, pacing, trotting, jogging, and jumping. Among them, only a small part of the mo-cap data is labeled, that is, the behavior pattern to which this mo-cap data belongs is clearly marked; while the vast majority of mo-cap data is unlabeled and has no annotation of the belonging behavior pattern
[0071] (1.2) Use the motion retargeting technology to map the original skeletal animation in the mo-cap data onto the quadruped robot. As Figure 3 shown, define the center of mass and five key points of the quadruped on the original skeleton, map them onto the robot skeleton, and use inverse kinematics to calculate the joint rotations to obtain the angular positions of the joints of the quadruped robot
[0072] Step 2: Semi-supervised learning of discrete behavior patterns: Maximize the mutual information through the variational lower bound to learn the five discrete behavior patterns as Figure 4 shown, including walking, pacing, trotting, jogging, and jumping; specifically including the following steps
[0073] (2.1) Since the cost of labeling motion capture data is relatively high, it is not appropriate to use a supervised learning method to obtain a large amount of labeled data. Also, because there are multiple behavior patterns in motion capture data, it is difficult to directly use an unsupervised approach to extract behavior features from the original data. Therefore, a semi-supervised learning method is selected. Based on using a small amount of labeled data to guide the disentanglement of multi-modal behaviors, a large amount of unlabeled data is fully utilized for learning.
[0074] Specifically, the control problem of the quadruped robot is formulated as a discrete-time dynamics problem that satisfies a partially observable Markov process: at each time step t, the robot executes an action using the policy π, and this action causes the robot to transition to the next state. On this basis, the latent skill variable c is used to represent five common animal dog behavior patterns: walking, pacing, trotting, jogging, and jumping, as Figure 5 shown. By changing the latent skill variable c of the input policy π, the quadruped robot exhibits corresponding different behavior patterns.
[0075] To ensure that c can guide the policy π to generate different actions, it is necessary to maximize the mutual information between c and the state of the quadruped robot, denoted as I(c; o I ), where o I includes the pitch angle, roll angle, yaw angular velocity, joint angles, joint velocities, and foot contact states of the quadruped robot. On the other hand, when using a small amount of labeled motion capture data for supervised learning, it is necessary to maximize the mutual information between the true label and the state of the quadruped robot, denoted as I(y; o I ), where y is the true label of the motion capture data.
[0076] Subsequently, the variational lower bounds and are calculated to approximately handle the mutual information, and the specific form is:
[0077]
[0078] where Q(·|o I ) is an estimate of the true posterior probability P(·|o I ), and the variational lower bound is tight when Q = P; d EL is the state transition distribution of the labeled motion capture data; d π is the state transition distribution of the quadruped robot obtained from the policy π; p(y) and p(c) represent the prior probabilities of y and c respectively, which are independent of the policy π; and represent the entropies of y and c respectively.
[0079] In practice, let Q1 = Q2, denoted as Q c , and the final semi-supervised regularization term L SS is obtained:
[0080]
[0081] (2.2) In formula (3), is the supervised term, Q c is implemented using a neural network, and d is used EL to predict y, and Q is optimized according to the cross - entropy loss function c ; is the unsupervised term, Q c is fixed. According to d π learn the latent semantic variables of y, and use reinforcement learning to optimize the policy π. The output of Q c will be used as the semi - supervised imitation reward r ss , and its form is:
[0082] r SS = logQ c (c|o I )(34)
[0083] Step 3, Unsupervised learning of continuous behavior styles: Capture the continuous style changes in behavior by maximizing the mutual information between continuous variables and behavior data; specifically, it includes the following steps:
[0084] (3.1) The motion capture data contains data of five behavior patterns: walking, pacing, trotting, jogging, and jumping. However, there are also style changes within the same behavior pattern. Therefore, it is necessary to introduce a continuous variable ∈ to capture this internal continuous style change, and it is achieved by maximizing the mutual information I(∈; o I ) between ∈ and the quadruped robot state. Use the variational lower bound L US to approximate the mutual information, and the calculation method is:
[0085]
[0086] where ∈ is sampled from the uniform distribution U(-1, 1), and Q ∈ (∈|o I ) is an estimate of the true posterior probability and is implemented using a neural network.
[0087] (3.2) Since there are no style labels within the same behavior pattern in the motion capture data, so in formula (5), L US is the unsupervised term, and it is used as the objective function to optimize Q ∈ , and at the same time, calculate the unsupervised imitation reward r ∈ according to the output of Q, and use reinforcement learning to optimize the policy π, and its specific form is: US
[0088]
[0088] r US = logQ∈ (∈|o I )(36)
[0089] Step 4, Unbalanced data optimization: Use discriminative clustering techniques and combine with the regularized information maximization method to extract intrinsic information from unbalanced data; specifically including the following steps:
[0090] (4.1) Since in motion capture data, the proportions of data with different behavioral patterns are different. When performing semi-supervised learning of behavioral patterns, it is necessary to adjust the sampling distribution of the latent variable c to align it with the state transition distribution of the unbalanced data. The specific method is to use Q c Predict the empirical label distribution based on unlabeled data where N is the total number of unlabeled data, is the predicted label, and ψ are the parameters of Q c Then use the empirical label distribution as the sampling distribution of the latent skill variable c.
[0091] (4.2) To better utilize unlabeled motion capture data, choose to use the regular information maximization method to calculate the penalty term L RIM , and automatically identify the boundaries between various behavioral patterns in the unlabeled motion capture data. The specific form is:
[0092]
[0093] where the first term is the clustering hypothesis, represents entropy. The second term is to avoid degenerate solutions by minimizing the KL divergence between and the distribution p(c) of the latent variable c. The third term R(ψ) is the parameter regularization performed on ψ to avoid complex solutions.
[0094] Step 5, Adversarial training and stability enhancement: Introduce a gradient penalty term to optimize the stability of the generative adversarial network and make the training process more robust; specifically including the following steps:
[0095] (5.1) Optimize the overall framework of algorithm training based on the generative adversarial imitation learning method. To obtain more stable training and higher-quality results, it is necessary to minimize a discriminator objective L GAIL similar to that of the least squares generative adversarial network: where D is the discriminator.
[0096] (5.2) Due to the function approximation error in the discriminator, the adversarial imitation learning method often has instability. The discriminator will assign non-zero gradients to actual data samples, which may lead to overshooting and oscillation of the generator. Therefore, a gradient penalty term L GPTo improve the training stability, the specific form is as follows:
[0097]
[0098] Among them, D is the discriminator, and φ is the discriminator parameter;
[0099] (5.3) Combine all the optimization items mentioned in the above steps to obtain the integrated objective function of the algorithm, which is used as the optimization objective of the discriminator D and the estimator Q c and Q ∈ and update it through the gradient descent method:
[0100]
[0101] Among them, the first term is a discriminator objective similar to the least squares generative adversarial network, that is
[0102] (5.4) After updating the discriminator D and the estimator Q c and Q ∈ , update the policy π by maximizing the total reward through the proximal policy optimization algorithm in reinforcement learning. The total reward includes the imitation reward r T and the task reward r T :
[0103] r = w I r I + w T r T (40)
[0104] The imitation reward includes the discriminator reward, the semi-supervised imitation reward, and the unsupervised imitation reward:
[0105] r I = r D + r SS + r US (41)
[0106] Among them, the first term r D is the discriminator reward. For the samples generated by the policy of the quadruped robot, the discriminator target is to predict its score as -1; while for the demonstration samples in the motion capture data, the target is to predict its score as 1. The specific reward formula is r D = max[0, 1 - 0.25(D(o I ) - 1) 2 , the semi-supervised imitation reward r SS and the unsupervised imitation reward r US are given by equations (4) and (6) respectively.
[0107] The task reward includes the linear velocity tracking reward Angular velocity tracking reward Jump height reward And stable height reward The specific forms are as follows:
[0108]
[0109] Among them, And v xy Represent the command and the actual linear velocity; And ω z Represent the command and the actual angular velocity; h cmd And h represent the command and the actual center height, Is an indicator function.
[0110] (5.5) Continuously repeat the process of adversarial training, and cyclically update the discriminator, estimator, and policy until convergence. Finally, obtain the policy π that generates the actions of five behavioral patterns, realizing the imitation of the natural behavioral patterns of real animals, dogs, by the quadruped robot.
[0111] In the said step (two), as Figure 5 Shown, a task-specific controller (TSC) for a quadruped robot includes the following steps:
[0112] Step 1: Sample the terrain height around the quadruped robot. As Figure 6 Shown, each orange point beside the quadruped robot represents the location of that point as a sampling point for the surrounding terrain height, which will be updated in real time as the robot moves, helping the robot continuously obtain the surrounding terrain elevation map.
[0113] Step 2: The teacher policy is trained based on privileged information, including:
[0114] (2.1) A perfect basic behavior controller has been deployed on the quadruped robot as the underlying controller, which can control the robot to make different actions after receiving the corresponding command input. The required commands include behavior pattern commands, speed commands, root height commands, and action style commands. The behavior pattern commands are discrete and can enable the quadruped robot to select one of the five behavior patterns of walking, pacing, trotting, jogging, and jumping; the speed commands are used to control the forward linear velocity and yaw angular velocity of the quadruped robot; the root height commands are used to control the body height of the robot; the action style commands are used to control the stride amplitude and frequency of the robot.
[0115] On this basis, a task-specific controller is needed as the top-level policy to output the commands required by the underlying controller and control the quadruped robot to make corresponding actions to complete specific tasks. To solve the problem of sparse reward signals in complex scenarios of specific tasks, the present invention uses a privileged learning architecture and divides the algorithm into two stages: teacher policy training and student policy training.
[0116] During the teacher policy training process, it is necessary to collect data including privileged information, etc. in the simulator as the input of the teacher policy network π TSC which specifically includes the following:
[0117] (a) External perception information: According to the sampled surrounding terrain elevation map;
[0118] (b) Privileged information: The type of obstacle the quadruped robot is currently facing, the yaw angle error Δ yaw with the current target point, and the yaw angle error Δ′ yaw with the next target point; as Figure 6 shown, the blue dot represents the current target point that the quadruped robot needs to reach, and the red dot after the blue dot represents the next target point that needs to be reached. At the same time, the type of obstacle the robot currently belongs to will also be provided in real time;
[0119] (c) Proprioceptive information: Includes the pitch angle, roll angle, yaw angular velocity, joint angles, joint velocities, and foot-end contact states of the quadruped robot.
[0120] It is difficult to directly train in the simulation environment according to the real-world scene conditions due to problems such as sparse reward signals. Although the surrounding terrain elevation map and privileged information are difficult to obtain in reality, they can be obtained in real time in the simulation environment. By additionally using this information in the teacher policy training, the speed and effect of policy training will be improved.
[0121] (2.2) Use the hybrid proximal policy optimization algorithm in reinforcement learning to implement the hybrid action space policy training. The teacher policy hybrid action space a T represents that the teacher policy output action is a mixture of discrete and continuous variable commands: a T ={a D ,a C}}, where the teacher policy discrete command a D represents the behavior pattern command; the teacher policy continuous command a C represents the speed command, root height command, and action style command.
[0122] The hybrid proximal policy optimization algorithm needs to maximize the reward r TSC in the following form:
[0123] r TSC =w I r I +w TSC r TSC (46)
[0124] In Equation (1), w I and w TSCThey are the coefficients of the imitation reward and the task reward, respectively. r I is the imitation reward calculated by the discriminator network pre-trained in advance, and the calculation formula is r I = max[0, 1 - 0.25(D(o I )) - 1) 2 , where D is the discriminator network, and o I is the proprioceptive information of the quadruped robot. o I is used as the input of D. If the actions of the quadruped robot are closer to those of the natural animal dog, then the output D(o I ) of D will be close to 1, otherwise close to -1. The task reward r TSC includes the linear velocity tracking reward the yaw angle tracking reward the waypoint arrival reward and the termination reward The formulas are as follows:
[0125]
[0126] where v is the linear velocity, ν target is the target linear velocity, d wpt is the direction of the target waypoint, and θ z are the yaw angles of the target waypoint and the quadruped robot respectively, p wpt and p represent the position of the target waypoint and the quadruped robot respectively, is the indicator function.
[0127] The trained teacher policy is combined with the underlying controller with the assistance of privileged information, and can manipulate the quadruped robot to complete different task requirements in the simulation environment.
[0128] In step (3), the robustness of the depth image is enhanced through the following steps by the self-supervised robust optimization method:
[0129] (3.1) The depth image is the external perception information commonly used when the quadruped robot is deployed in real-scene tasks. The depth images in the real environment usually contain noises from various sources, making them different from the images in the simulation environment. To improve the robustness of the algorithm to the depth image, probability enhancement processing is performed on the depth images collected by the binocular depth camera on the quadruped robot, such as Figure 3 shown, including white noise, background noise, random cropping, edge noise, and Gaussian blur.
[0130] (3.2) Use the self-supervised learning method Bootstrap Your Own Latent (BYOL) to maximize the similarity between two enhanced views of the same depth image to learn task-related features, such asFigure 7 As shown, integrating this self-supervised loss into the depth image encoder network enhances the robustness of the top-level policy to complex real-world environments.
[0131] In step (4), the student policy is trained to imitate the teacher policy based on historical state privileged information, which includes the following steps:
[0132] (4.1) To obtain the top-level policy of the quadruped robot that can be deployed in real-world scenarios, when training the student policy network , use the robustified depth image to replace the surrounding terrain elevation map, and collect the historical sequence of proprioceptive information instead of the difficult-to-obtain privileged information as the input.
[0133] (4.2) Use a gated recurrent unit to process the input sequence and output the predicted privileged information and the latent vector of environmental information, and output the final mixed student action through the gated recurrent unit and a multi-layer perceptron where is the mixed action space of the student policy, is the discrete command of the student policy, is the continuous command of the student policy. Training the student policy uses supervised learning to imitate the teacher policy, and the form of the objective function to be minimized is as follows:
[0134]
[0135] where K is the number of discrete actions. The first term of the objective function corresponds to the cross-entropy loss between the teacher policy and the discrete command of the student policy, while the second term is used to capture the mean square error between the teacher policy and the continuous command of the student policy.
[0136] The trained student policy combined with the underlying controller enables the quadruped robot to successfully execute various complex tasks in the real-world environment.
[0137] In step (3), as Figure 8 shown, a simulator optimization method based on evolutionary strategy and generative adversarial imitation learning includes the following steps:
[0138] Step 1, parameter initialization: Initialize the parameter distribution of the simulator;
[0139] (1.1) Define the initial range of the simulator parameters;
[0140] (1.2) Randomly sample the initial parameter distribution.
[0141] Step 2, real data collection: Collect state-action trajectory data in the simulation environment and the real environment respectively; specifically, it includes the following steps:
[0142] (2.1) In the real environment, use a predefined controller to generate state-action trajectories;
[0143] (2.2) Use the collected data for adversarial training.
[0144] Step 3, Simulation data collection: Collect state-action trajectory data in the simulation environment and the real environment respectively; specifically including the following steps:
[0145] (3.1) In the simulation environment, use the same controller as in step (2.1) to generate state-action trajectories;
[0146] (3.2) Use the collected data for adversarial training.
[0147] Step 4, Adversarial training: Use a discriminator to distinguish the state-action transitions between the real world and the emulator, and at the same time optimize the parameter distribution through the Evolutionary Strategy (ES) to maximize the discriminator's score for real samples; specifically including the following steps:
[0148] (4.1) The optimization objective of the discriminator D is to distinguish the state-action transitions between the real world and the emulator, and the specific form is:
[0149]
[0150] where d B (s, a, s′) is simulation data, and d M (s, a, s′) is real data;
[0151] (4.2) The optimization objective of the Evolutionary Strategy (ES) is to maximize the discriminator's score for real samples
[0152] Step 5, Parameter update: Update the emulator according to the optimized parameter distribution; specifically including the following steps: Specifically including the following steps:
[0153] (5.1) Update the physical parameters of the emulator according to the optimized parameter distribution;
[0154] Step 6, Iterative optimization: Repeat data collection, adversarial training, and parameter update until the parameter distribution converges; specifically including the following steps:
[0155] (6.1) Repeat data collection and adversarial training until the discriminator can no longer distinguish between real and simulation data;
[0156] (6.2) Verify whether the optimized emulator meets the preset convergence conditions, such as the change range of the parameter distribution is less than the threshold.
[0157] Step 7, Policy optimization: Use the optimized emulator to retrain or fine-tune the control policy; specifically including the following steps:
[0158] (7.1) Use the optimized emulator physical parameters as the parameter range for domain randomization, retrain the control strategy, or fine-tune the same controller as in (2.1) to adapt the strategy to simulations closer to the real environment.
[0159] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the creative concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application.
Claims
1. A basic behavior control method for a quadruped robot, characterized in that, It includes the following steps: (1) Motion capture data preparation: Collect the behavior data of real animals and map it onto the target robot skeleton; (2) Semi-supervised learning of discrete behavior patterns: Learn discrete behavior patterns, including walking, pacing, trotting, jogging, and jumping, by maximizing the mutual information of the variational lower bound; (3) Unsupervised learning of continuous behavior styles: Capture the continuous style changes in behavior by maximizing the mutual information between continuous variables and behavior data; (4) Optimization of imbalanced data: Adopt discriminative clustering techniques and combine the method of maximizing regularized information to extract the intrinsic information from imbalanced data; (5) Adversarial training and stability enhancement: Introduce a gradient penalty term to optimize the stability of the generative adversarial network.
2. The basic behavior control method of the quadruped robot according to claim 1, characterized in that, In step (1), motion capture data preparation is carried out, including the following steps: (1.1) Collect the motion capture data of real animal dogs, including five behavior patterns, namely walking, pacing, trotting, jogging, and jumping; The motion capture data includes labeled data and unlabeled data. The labeled motion capture data clearly indicates the behavior pattern to which the motion capture data belongs; The unlabeled motion capture data has no annotation of the behavior pattern to which it belongs; (1.2) Use motion redirection technology to map the original skeletal animation in the motion capture data onto the quadruped robot; Define the centroid and five key points of the quadruped on the original skeleton, map them onto the robot skeleton, and use inverse kinematics to calculate the joint rotation to obtain the angular positions of the joints of the quadruped robot.
3. The basic behavior control method of the quadruped robot according to claim 1, characterized in that In step (2), the semi-supervised mutual information is maximized in the following way: (2.1) Formulate the control problem of the quadruped robot as a discrete-time dynamics problem that satisfies the partially observable Markov process: At each time step t, use the policy π to control the robot to execute an action, which makes the robot transfer to the next state; On this basis, use the latent skill variable c to represent five common animal dog behavior patterns: walking, pacing, trotting, jogging, and jumping; Maximize the mutual information between \(c\) and the state of the quadruped robot, denoted as \(I(c; o I )\), where \(o I includes the pitch angle, roll angle, yaw angular velocity, joint angles, joint velocities, and foot contact states of the quadruped robot; Maximize the mutual information between the true label and the state of the quadruped robot, denoted as \(I(y; o I )\), where \(y\) is the true label of the motion capture data; Subsequently, the variational lower bound is calculated and to approximate the mutual information, and the specific form is as follows: Among them, Q(·|o I ) is an estimate of the true posterior probability P(·|o I ), and the variational lower bound is tight when Q = P; d EL is the state transition distribution of labeled motion capture data; d π is the state transition distribution of the quadruped robot obtained by the policy π; p(y) and p(c) represent the prior probabilities of y and c respectively, which are independent of the policy π; and represent the entropies of y and c respectively; Let Q1 = Q2, denoted as Q c , and obtain the final semi-supervised regularization term L SS : (2.2) In formula (3), is the supervised item, Q c is implemented using a neural network. Using d EL to predict y, and optimizing Q according to the cross-entropy loss function c ; is the unsupervised item, Q c is fixed. According to d π to learn the latent semantic variables of y, using reinforcement learning to optimize the policy π, and the output of Q c will be used as the semi-supervised imitation reward r ss , and its form is: r SS = logQ c (c|o I ) (4).
4. The basic behavior control method of the quadruped robot according to claim 1, characterized in that In step (3), the unsupervised mutual information is maximized in the following way: (3.1) The motion capture data contains data of five behavior patterns: walking, pacing, trotting, jogging, and jumping. There are also style changes within the same behavior pattern; Introduce a continuous variable ∈ to capture this internal continuous style change by maximizing the mutual information I(∈; o I ) between ∈ and the quadruped robot state; use the variational lower bound L US (Q ∈ ) to approximate the mutual information, calculated as: where ∈ is sampled from the uniform distribution U(-1, 1), and Q ∈ (∈|o I ) is an estimate of the true posterior probability; In equation (5) of (3.2), L US is the unsupervised term, which optimizes Q as the objective function ∈ while calculating the unsupervised imitation reward r according to the output of Q ∈ and using reinforcement learning to optimize the policy π, whose specific form is US : r US = logQ ∈ (∈|o I ) (6).
5. The basic behavior control method of the quadruped robot according to claim 1, characterized in that In step (4), the imbalanced data is optimized in the following way: (4.1) Adjust the sampling distribution of the latent variable c to align it with the state transition distribution of the imbalanced data; specifically, use Q c Predict the empirical label distribution based on the unlabeled data where N is the total number of unlabeled data, is the predicted label, and ψ is the parameter of Q c ; then use the empirical label distribution as the sampling distribution of the latent skill variable c; (4.2) Calculate the penalty term L using the regular information maximization method RIM , and automatically identify the boundaries between various behavioral patterns in the unlabeled motion capture data, in the specific form of: Among them, the first term is the clustering assumption, representing the entropy of; the second term is to avoid degenerate solutions by minimizing the KL divergence between and the distribution p(c) of the latent variable c; the third term R(ψ) is the parameter regularization performed on ψ to avoid complex solutions.
6. The basic behavior control method of the quadruped robot according to claim 1, characterized in that In step (5), adversarial training and stability enhancement are carried out in the following way: (5.1) The overall framework of algorithm training is optimized based on the generative adversarial imitation learning method; minimize a discriminator objective L similar to that of the least squares generative adversarial network GAIL : where D is the discriminator; (5.2) Introduce the gradient penalty term $L$ GP to improve the training stability, and the specific form is: Among them, D is the discriminator, and φ is the discriminator parameter; (5.3) Combine all the optimization items mentioned in the above steps to obtain the integrated objective function of the algorithm, which serves as the optimization objectives for the discriminator D and the estimator Q c and Q ∈ and update them using the gradient descent method: Among them, the first item is a discriminator objective similar to the least squares generative adversarial network, that is (5.4) When updating the discriminator D and the estimator Q c and Q ∈ After that, the policy π is updated by maximizing the total reward through the proximal policy optimization algorithm in reinforcement learning; where the total reward includes the imitation reward r T and the task reward r T : r = w I r I + w T r T (10) The imitation rewards include discriminator rewards, semi-supervised imitation rewards, and unsupervised imitation rewards: r I =r D +r SS +r US (11) Among them, the first item r D is the discriminator reward. For the samples generated by the quadruped robot through the policy, the discriminator's goal is to predict its score as -1; while for the demonstration samples in the motion capture data, the goal is to predict its score as 1. The specific reward formula is r D = max[0, 1 - 0.25(D(o I ) - 1) 2 . The semi-supervised imitation reward r SS and the unsupervised imitation reward r US are given by equations (4) and (6) respectively; The task rewards include linear velocity tracking rewards angular velocity tracking rewards jump height rewards and stable height rewards The specific forms are as follows: wherein, and v xy represent the command and the actual linear velocity; and ω z represent the command and the actual angular velocity; h cmd and h represent the command and the actual center height, is an indicator function; (5.5) Continuously repeat the process of adversarial training, and cyclically update the discriminator, estimator, and policy until convergence; Finally, obtain the policy π that generates the actions of five behavior patterns, realizing the imitation of the natural behavior patterns of real animal dogs by the quadruped robot.
7. A basic behavior control device for a quadruped robot, characterized in that It includes a memory and a processor. The memory stores a computer program. When the computer program is called and executed by the processor, it implements the method described in any one of claims 1-6.
8. A computer-readable medium, characterized in that, The computer-readable medium stores a computer program. When the computer program is called and executed by a computer, it implements the method described in any one of claims 1-6.