Robot motion control method and equipment
By building a fusion framework based on policy network, adversarial evaluation network and feature extraction network, and combining a small number of expert data sets to train the policy network, the problem of high requirements for expert data sets in the existing technology is solved, and high-quality robot motion style imitation is achieved.
Patent Information
- Application Number
- CN202510119576.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-24
AI Technical Summary
In the prior art, when using adversarial networks to perform robotic human-imitation movements, expert data sets are needed to train evaluation networks, which leads to very high requirements for expert data sets and it is difficult to effectively imitate human movement styles.
By building a fusion framework based on policy networks, adversarial evaluation networks and feature extraction networks, and training the policy networks with a small number of expert data sets, reducing dependence on expert data sets, and pre-extracting terrain features and behavioral features through feature extraction networks, inputting the policy network to improve the flexibility of action behavior and the imitation quality of human behavior styles.
It is achieved to improve the quality of robot motion style imitation while reducing the dependence of expert data sets, making robot motion behavior closer to human behavior.
Smart Images

Figure CN120056097A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of humanoid robots, and in particular, to a robot motion control method and device. Background Art
[0002] Robot humanoid motion refers to the robot simulating the motion mode and behavior of humans through specific designs and technologies.
[0003] In related technologies, an adversarial network can be used to learn the motion style of humans.
[0004] However, in the process of implementing the present application, the inventors found that there are at least the following problems in the prior art: The method using an adversarial network requires separately training an evaluation network using an expert data set to improve the humanoid actions of the robot, and has very high requirements for the expert data set. Summary of the Invention
[0005] The embodiments of the present application provide a robot motion control method and device to improve the imitation quality of the motion style and reduce the dependence on the expert data set.
[0006] In a first aspect, the embodiments of the present application provide a robot motion control method, including:
[0007] Obtain first historical data, second historical data, and a current first action instruction; the first historical data includes observation data corresponding to a first preset number of time steps before the current time step, and the second historical data includes observation data corresponding to a second preset number of time steps before the current time step; the second preset number is less than the first preset number;
[0008] Input the first historical data into a trained feature extraction network to obtain first feature data; the first feature data includes terrain features and / or behavior features;
[0009] Input the current first action instruction, the first feature data, and the second historical data into a trained policy network to obtain a current first action; the trained policy network is obtained by training the policy network according to a reinforcement learning algorithm based on a fusion framework composed of a feature extraction network, a policy network, and an adversarial evaluation network;
[0010] Determine a corresponding current first torque according to the current first action, and drive the robot to execute the current first action according to the current first torque.
[0011] In a possible design, the trained feature extraction network is a convolutional neural network; the convolutional neural network includes a first hidden layer and a second hidden layer; the input layer of the first hidden layer has 42 channels, the output layer has 32 channels, the convolutional kernel size is 6, and the stride is 5; the input layer of the second hidden layer has 32 channels, the output layer has 16 channels, the convolutional kernel size is 4, and the stride is 2; the activation function of the convolutional neural network is the ReLU activation function.
[0012] In a possible design, a fusion framework is constructed based on the feature extraction network, the policy network, and the adversarial evaluation network; based on the fusion framework, the policy network is trained according to the Proximal Policy Optimization (PPO) algorithm to obtain the trained policy network.
[0013] In a possible design, the feature extraction network includes an encoder and a decoder; the encoder includes a convolutional neural network; the decoder includes a multi-layer perceptron.
[0014] In a possible design, the training of the policy network according to the Proximal Policy Optimization (PPO) algorithm based on the fusion framework to obtain the trained policy network includes:
[0015] For each time step, the third historical data is input into the feature extraction network to obtain the second feature data; the second feature data includes terrain features and / or behavior features; the third historical data includes the third preset number of actions before the previous action;
[0016] The current second action instruction, the previous reward, the second feature data, and the fourth historical data are input into the policy network to obtain the current second action; the fourth historical data includes the fourth preset number of actions between the previous actions; the fourth preset number is less than the third preset number;
[0017] The current first action is input into the controller of the robot to obtain the corresponding current second torque;
[0018] The current second torque is input into the environment model to obtain the current first state;
[0019] The current first state is input into the adversarial evaluation network to obtain the current style reward;
[0020] Based on a preset reward function, the current reward is determined according to the current task reward and the style reward;
[0021] The current second action, the current first state, the current style reward, and the current reward are stored in the experience replay buffer;
[0022] Sample sample data from the experience replay buffer; update the policy network according to the sample data based on the objective function of the PPO algorithm;
[0023] Update the adversarial evaluation network according to the sample data based on a preset objective function;
[0024] Update the feature extraction network according to the sample data based on a preset loss function.
[0025] In a possible design, the expression of the objective function of the PPO algorithm is:
[0026]
[0027] where, is the objective function of PPO, and the goal is to find the parameter θ to maximize the objective function; is the estimate of the advantage function, indicating the benefit of taking action a t in state w t relative to the average; w t is is the new policy p θ relative to the old policy in action a t the probability ratio, that is, the importance sampling ratio; is the truncated policy ratio, where ɛ is a constant, ɛ ≤ 0.2, and the clip function ensures that the importance sampling ratio is in the range of [1 - ɛ, 1 + ɛ].
[0028] In a possible design, the expression of the preset objective function is:
[0029]
[0030] where, represents the real state and the expert reference state; represents the evaluation value output by the adversarial evaluation network for the state pair output, and when it is closer to the expert state, the output value is closer to 1, otherwise it is closer to -1; is the real state and the expert reference state, the divergence difference between the two states; is the weight coefficient.
[0031] In a possible design, the expression of the preset loss function is:
[0032]
[0033] where, represents the historical observation data, Represents the next state of reconstruction, Represents the next state, where MSE (mean square error) is the mean square error; is the posterior distribution, that is, the conditional distribution of the latent variable z given the input x; p(zt) is the prior distribution, that is, the marginal distribution of the latent variable z; KL is the KL divergence, which measures the difference between the approximate posterior distribution and the prior distribution p(zt).
[0034] In a possible design, the expression of the preset reward function is:
[0035]
[0036] where, is the main reward, is the task reward, is the style reward, ω g is the weight of the task reward, ω s is the weight of the style reward.
[0037] In a second aspect, an embodiment of the present application provides a robot motion control device, including: at least one processor and a memory;
[0038] The memory stores computer-executable instructions;
[0039] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the method described in the first aspect above and various possible designs of the first aspect.
[0040] The robot motion control method and device provided in this embodiment. The method includes obtaining first historical data, second historical data, and a current first action instruction. The first historical data includes observation data corresponding to a first preset number of time steps before the current time step, and the second historical data includes observation data corresponding to a second preset number of time steps before the current time step. The second preset number is less than the first preset number. Input the first historical data into a trained feature extraction network to obtain first feature data, which includes terrain features and / or behavior features. Input the current first action instruction, the first feature data, and the second historical data into a trained policy network to obtain a current first action. The trained policy network is based on a fusion framework composed of a feature extraction network, a policy network, and an adversarial evaluation network, and is obtained by training the policy network according to the reinforcement learning algorithm. According to the current first action, determine the corresponding current first torque, and drive the robot to execute the current first action according to the current first torque. The robot motion control method provided in this embodiment constructs a fusion framework based on a policy network, an adversarial evaluation network, and a feature extraction network, and can easily train the policy network to achieve human-like motion by combining only a small amount of expert data sets. And by setting the feature extraction network, it can pre-extract feature data such as terrain features and behavior features and input them into the policy network, making the action behavior determined by the policy network more compliant and in line with the human behavior style, improving the imitation quality of the behavior style. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of the robot motion control method provided in the embodiment of the present application;
[0043] Figure 2 It is a schematic diagram of the application scenario of the robot motion control method provided in the embodiment of the present application;
[0044] Figure 3 It is a flowchart of the training process of the policy network provided in the embodiment of the present application;
[0045] Figure 4 It is a schematic diagram of the training principle of the fusion framework of the robot motion control method provided in the embodiment of the present application;
[0046] Figure 5 It is a schematic diagram of the structure of the robot motion control device provided in the embodiment of the present application;
[0047] Figure 6 This is a schematic diagram of the hardware structure of the robot motion control device provided by the embodiments of the present application. Detailed implementation manners
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0049] It should be noted that the robot motion control method and device provided by the present application can be used in the field of humanoid robot technology, and can also be used in any field other than the field of humanoid robot technology. The application fields of the robot motion control method and device provided by the present application are not limited.
[0050] Humanoid motion of a robot refers to the robot simulating the motion modes and behaviors of humans through specific designs and technologies.
[0051] In related technologies, the motion style of a human can be learned by directly learning based on human motion characteristics. For example, the motion trajectory of a human is subjected to feature learning based on a Fourier Latent Dynamics (FLD) encoder to obtain a Fourier parameter space and sample therefrom to form a reconstructed trajectory; the motion style of a human can also be learned through an adversarial network. However, it is difficult to learn the subtle changes in actions in the former method, and it is difficult to reproduce the lower limb motion while maintaining balance. In the latter method, an evaluation network needs to be separately trained using an expert data set to improve the humanoid actions of the robot, and the requirements for the expert data set are very high.
[0052] To solve the above technical problems, the inventors of the present application have found through research that a humanoid fusion framework based on a policy network, an adversarial evaluation network, and a feature extraction network can be used to easily train humanoid motion in combination with a small amount of expert data sets. By setting the feature extraction network, feature data such as terrain features and behavior features can be pre-extracted and input into the policy network, so that the action behaviors determined by the policy network are more compliant and conform to the human behavior style, improving the imitation quality of the behavior style. Based on this, the embodiments of the present application provide a robot motion control method.
[0053] The technical solutions of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0054] Figure 1 This is a schematic flowchart of the robot motion control method provided by the embodiments of the present application. As Figure 1 shown, the method includes:
[0055] 101. Obtain first historical data, second historical data, and a current first action instruction; the first historical data includes observation data corresponding to a first preset number of time steps before the current time step, and the second historical data includes observation data corresponding to a second preset number of time steps before the current time step; the second preset number is less than the first preset number.
[0056] The execution subject of this embodiment can be the controller of a humanoid robot or the humanoid robot.
[0057] Specifically, the first historical data contains observation data of more time steps than the second historical data. Therefore, the first historical data can also be called long historical data, and the second historical data can be called short historical data. As Figure 2 shown, the long historical data can include the action a corresponding to the moment t - 1 t-1 and the state O obtained after the humanoid robot executes the action a t-1 , to t the moment t-n corresponding to the action a t-n and the state O obtained after executing the action a t-n . n can be a positive integer greater than 50 and less than 150, such as 100. The short historical data can include the action a corresponding to the moment t - 1 t-n+1 and the state O obtained after the humanoid robot executes the action a t-1 , to the action a corresponding to the moment t - m t-1 and the state O obtained after executing the action a t . m can be a positive integer greater than or equal to 2 and less than or equal to 8, such as 5. Among them, O t-n includes the angular velocity ω t-n , the gravity vector g t-m+1 , the whole body joint position q t , and the whole body joint velocity q t '. t t t '.
[0058] Among them, the current task instruction can be an action instruction issued by the task instruction module, and can include parameters such as a destination, an average speed, and an average angular velocity.
[0059] In this embodiment, a storage space can be set to store the action a t-1 , state Qt (which may include the angular velocity ω t , the gravity vector g t , the whole-body joint positions q t and the whole-body joint velocities q t '), whether standing, whether walking, control instructions, etc. data, for each time step, the corresponding w can be extracted from this storage space t input policy network and the trained feature extraction network. w t includes long historical data, short historical data, sign of standing or walking (e.g., 0 for standing, 1 for walking).
[0060] 102. Input the first historical data into the trained feature extraction network to obtain first feature data; the first feature data includes terrain features and / or behavior features.
[0061] Specifically, by using the trained feature extraction network to extract features from the long historical data, terrain features or behavior features can be obtained. By using the trained feature extraction network to extract the representation work of the long historical data, the trained policy network can recognize and predict the terrain environment and behavior features through the historical latent representation, and the generated action behaviors are also more compliant; since the long historical data contains observation data at enough time steps, the terrain features or behavior features can be predicted more accurately, improving the prediction accuracy.
[0062] In some embodiments, the trained feature extraction network is a convolutional neural network; the convolutional neural network includes a first hidden layer and a second hidden layer; the number of input layer channels of the first hidden layer is 42, the number of output layer channels is 32, the convolutional kernel size is 6, and the stride is 5; the number of input layer channels of the second hidden layer is 32, the number of output layer channels is 16, the convolutional kernel size is 4, and the stride is 2; the activation function of the convolutional neural network is the ReLU activation function.
[0063] Specifically, the structure of the convolutional neural network (CNN) consists of two hidden layers. The configurations of the two hidden layers are defined in the way of [number of input layer channels, number of output layer channels, convolutional kernel size, stride size], which are [42, 32, 6, 5] and [32, 16, 4, 2] respectively, and the ReLU activation function is used.
[0064] 103. Input the current first action instruction, the first feature data, and the second historical data into the trained policy network to obtain the current first action; the trained policy network is a fusion framework composed of a feature extraction network, a policy network, and an adversarial evaluation network, and is obtained by training the policy network according to the reinforcement learning algorithm.
[0065] In this embodiment, the trained policy network consists of three layers, and each layer contains 256 neurons using ELU (Exponential Linear Unit) as the activation function. Among them, for the input layer: the number of neurons in the input layer depends on the number of input features of the specific problem. For example, if the input feature is an n-dimensional vector, then the input layer has n neurons. For the hidden layer: the first hidden layer: contains 256 neurons, and each neuron uses ELU as the activation function. The second hidden layer: also contains 256 neurons, and each neuron also uses ELU as the activation function. The main role of the hidden layer is to extract the features of the input data and map them into a higher-dimensional or lower-dimensional space to more easily solve classification or regression problems. For the output layer: the number of neurons in the output layer depends on the output requirements of the specific problem. For example, in a classification problem, the number of neurons in the output layer may be equal to the number of classes (using the softmax activation function); in a regression problem, the output layer may have only one neuron (without using an activation function or using a linear activation function). In the policy network of reinforcement learning, the output layer can correspond to the probability distribution of actions or the continuous values of actions. ELU (Exponential Linear Unit): ELU is a non-linear activation function.
[0066] 104. Determine the corresponding current first torque according to the current first action, and drive the robot to execute the current first action according to the current first torque.
[0067] Specifically, as Figure 2 shown, for the action a output by the trained policy network t-1 generate torque τ through a PD controller torque , and this torque serves as a control command, acting on the motors of the robot. The robot tracks the control command and interacts with the environment to generate a new state s t (the position q of the whole body joints t , the linear velocity v t , the angular velocity ω t , the height position z pos ). It will be compared with the expert set in the constructed adversarial evaluation network to form a style reward and act on the update of the policy network together with the task reward. Finally, the agent will improve its behavior style according to the style coefficient to make its actions more like those of humans.
[0068] The robot motion control method provided in this embodiment constructs a fusion framework based on a policy network, an adversarial evaluation network, and a feature extraction network. It can easily train the policy network to achieve human-like motion by combining only a small amount of expert datasets. Moreover, by setting up the feature extraction network, it can pre-extract feature data such as terrain features and behavior features and input them into the policy network, making the action behaviors determined by the policy network more compliant and in line with human behavior styles, thus improving the imitation quality of behavior styles.
[0069] Figure 3 It is a schematic flowchart of the training process of the policy network provided in the embodiment of this application. As Figure 3 shown, the method includes:
[0070] 301. Construct a fusion framework based on the feature extraction network, the policy network, and the adversarial evaluation network.
[0071] Specifically, first, an expert dataset required for the adversarial evaluation network can be constructed. Specifically, data of each joint and key point of the human body can be collected from a human motion capture device, and then the data mapped to the corresponding joints and end positions on the robot body can be used as the expert dataset. By this method, various full-body data including walking, jumping, squatting, etc. are collected as the expert dataset. Since the robot motion control method provided in the embodiment of this application uses a fusion framework constructed by a feature extraction network, a policy network, and an adversarial evaluation network for training, the number of samples in the expert dataset does not have to be too large. For example, it can include 200 samples.
[0072] In this embodiment, the adversarial evaluation network is composed of two layers of MLP (Multi-Layer Perceptron), and each layer contains 1024 and 512 neurons with'relu' (Rectified Linear Unit) as the activation function respectively.
[0073] In some embodiments, the feature extraction network includes an encoder and a decoder; the encoder includes a convolutional neural network; the decoder includes a multi-layer perceptron.
[0074] Specifically, the Variational Autoencoder (VAE) is composed of an encoder and a decoder. The encoder is based on the structure of a convolutional neural network (CNN) and consists of two hidden layers. The configurations of the two hidden layers are defined in the form of [number of input layer channels, number of output layer channels, convolutional kernel size, stride size], which are [42, 32, 6, 5] and [32, 16, 4, 2] respectively, and the ReLU activation function is used. The decoder is based on the structure of a multi-layer perceptron (MLP) and consists of three layers, each containing 256 neurons with'relu' as the activation function.
[0075] 302. Based on the fusion framework, train the policy network according to the Proximal Policy Optimization (PPO) algorithm to obtain the trained policy network.
[0076] Specifically, after constructing the fusion framework of the CNN network, the policy network, and the Adversarial Motion Primitives (AMP), the data of the expert set and the state data of the robot can be read to learn the action style of the expert set. Furthermore, based on the reinforcement learning PPO algorithm, the robot behavior policy can be trained in the Isaac Gym environment to learn the action execution policy. The obtained execution policy is transformed into a model. First, it is verified by sim2sim in physical environments such as Mujoco, and after passing the verification, it is converted into an ONNX model for real - machine deployment and debugging, and finally, a robot capable of performing human - like actions is obtained.
[0077] In some embodiments, the step of training the policy network according to the Proximal Policy Optimization (PPO) algorithm based on the fusion framework to obtain the trained policy network may include: for each time step, input the third historical data into the feature extraction network to obtain the second feature data; the second feature data includes terrain features and / or behavior features; the third historical data includes the third preset number of actions before the previous action; input the current second action instruction, the previous reward, the second feature data, and the fourth historical data into the policy network to obtain the current second action; the fourth historical data includes the fourth preset number of actions between the previous actions; the fourth preset number is less than the third preset number; input the current first action into the controller of the robot to obtain the corresponding current second torque; input the current second torque into the environment model to obtain the current first state; input the current first state into the Adversarial Motion Primitives (AMP) to obtain the current style reward; based on a preset reward function, determine the current reward according to the current task reward and the style reward; store the current second action, the current first state, the current style reward, and the current reward in the experience replay buffer; sample sample data from the experience replay buffer; update the policy network according to the sample data based on the objective function of the PPO algorithm; update the Adversarial Motion Primitives (AMP) according to the sample data based on a preset objective function; update the feature extraction network according to the sample data based on a preset loss function.
[0078] In this embodiment, during the update process of the feature extraction network, multiple samples can be extracted from the experience replay buffer. Each sample includes long historical data for inputting into the CNN encoder and the corresponding state Ot+1 (this state is the observed data obtained after the action corresponding to the long historical data is executed, and this state is used as the label value). Furthermore, for each sample, the long historical data in the sample can be input into the feature extraction network (i.e., the VAE composed of a CNN encoder and an MLP decoder) to obtain the output of the MLP, and the loss can be calculated based on this output value and the label value.
[0079] Exemplarily, as Figure 4 shown, first, the long historical data a t-1 , O t , a t-2 , O t-1 , …… are learned based on the CNN encoder in the VAE. Its output, together with the short historical data O t-4 ~O t , a t-4 ~a t-1 and the control instruction Ct are simultaneously input into the policy network. Through the extraction of the representation of the long historical data by the variational autoencoder (VAE), the agent can recognize and predict the terrain environment through the latent representation of the history, and the generated action behavior is also more compliant.
[0080] Secondly, for the output action a t-1 of the policy network, the torque τ torque is generated through the PD controller as the control instruction and acts on the motors of the robot. The robot tracks the control instruction and interacts with the environment to generate new body state data s t , which will be compared with the expert set in the constructed adversarial evaluation network to form a style reward and act on the update of the policy network together with the task reward. Finally, the policy network will improve its behavior style according to the style coefficient, making its actions more like those made by humans.
[0081] In some embodiments, the expression of the objective function of the PPO algorithm is:
[0082] (1)
[0083] where is the objective function of PPO, and the goal is to find the parameter θ to maximize the objective function; is the estimate of the advantage function, which represents the benefit of taking the action a t in the state w t relative to the average; w t is is the new policy p θ relative to the old policy at the action at The probability ratio on it, i.e., the importance sampling ratio; is the truncation policy ratio, where ɛ is a constant, ɛ ≤ 0.2, and the clip function ensures that the importance sampling ratio is within the range of [1 - ɛ, 1 + ɛ].
[0084] Specifically, the policy adopted by the policy network is based on the Markov Decision Processes (MDP) of reinforcement learning. A process in which an agent, i.e., a robot, in an environment selects actions according to the current state to maximize the long-term reward. State set (S): Represents the set of all possible states that the agent may be in. Action set (A): Represents the set of all actions that the agent can choose in each state. State transition probability (P): Describes the probability distribution of transitioning to other states after performing a certain action in a given state. Reward function (R): Defines the reward obtained by the agent for performing a certain action in a given state.
[0085] How to select the optimal action to obtain the maximum expected reward in a complex environment has become a difficult problem. In this embodiment, the Proximal Policy Optimization (PPO) reinforcement learning method is adopted, thus solving the difficulty of choosing the learning rate value when the policy gradient processes the continuous action space. If the learning rate value is too small, the convergence of deep reinforcement learning will be poor, resulting in a situation where the training cannot be completed. If the value is too large, the data will be inconsistent during the iteration of the new and old policies, causing large learning fluctuations or local oscillations. The PPO algorithm formula is shown in expression (1). Among them, is the objective function of PPO, and the goal is to find the parameter θ to maximize this objective function. is the estimation of the advantage function, indicating the benefit of taking action a t in state w t relative to the average value. w t includes short historical data (i.e., the fourth historical data), and the estimation of the advantage function is usually calculated by some methods, such as Generalized Advantage Estimation (GAE). is the probability ratio of the new policy pθ to the old policy pθk on action at. This ratio is also called the "importance sampling ratio". is the "truncation policy ratio" unique to PPO, where is a small constant, such as 0.1 or 0.2. This clip function ensures that the importance sampling ratio is within the range, which can avoid too large policy updates.
[0086] In some embodiments, the expression of the preset objective function is as follows:
[0087] (2)
[0088] Wherein, represents the true state and the expert reference state; represents the evaluation value output by the adversarial evaluation network for the state pair The closer the output value is to the expert state, the closer it is to 1, and vice versa, the closer it is to -1; is the divergence difference between the true state and the expert reference state, the two states; is the weight coefficient.
[0089] Specifically, the update strategy of the adversarial evaluation network is mainly based on the above expression (2). By taking the data read from the expert set and the data of the robot's own state as inputs, and updating by minimizing the distance between its output value and the label. Wherein, represents the true state and the expert reference state; represents the evaluation value output by the adversarial evaluation network for the state pair The closer the output value is to the expert state, the closer it is to 1, and vice versa, the closer it is to -1. is the divergence difference between the true state and the expert reference state, the two states. It is hoped that the difference will not be too large, and it is adjusted by the weight coefficient. The selection of the weight coefficient is also based on experience.
[0090] In some embodiments, the expression of the preset loss function is as follows:
[0091] (3)
[0092] Wherein, represents the historical observation data, that is, short historical data, represents the reconstructed next state, represents the next state, MSE (mean square error) is the mean square error; is the posterior distribution, that is, the conditional distribution of the latent variable z given the input x; p(zt) is the prior distribution, that is, the marginal distribution of the latent variable z; KL is the Kullback-Leibler divergence, which measures the difference between the approximate posterior distribution and the prior distribution p(zt). The KL divergence is also called relative entropy or information divergence, and is a method for measuring the difference between two probability distributions.
[0093] Specifically, in a variational autoencoder (VAE), the main function of the VAE encoder is to compress the input data into a low-dimensional representation, which is called a latent variable. This latent variable is the low-dimensional representation of the data in the latent space. It captures the main features of the data while removing redundant information. Through the latent variable, dimensionality reduction, reconstruction, and generation of data can be achieved, and future movements can be predicted. In a VAE, the encoder does not output a single latent variable value, but a distribution of latent variables. Specifically, the encoder outputs the mean (μ) and variance (σ²) of the latent variable or their logarithmic forms (logσ²). In this way, a specific value of the latent variable can be sampled from this distribution in the latent space. The VAE encoder is usually a neural network that accepts input data and outputs the mean and variance (or logarithmic variance) of the latent variable
[0094] The update process of the variational autoencoder (VAE) is as follows: First, input data: The historical observation data is fed into the encoder. Second, encoding: The encoder compresses the input data into a low-dimensional latent representation through neural network layers (CNN convolutional layers). In a VAE, this latent representation is not a single point, but the parameters of a distribution (mean and variance). Third, sampling: A point is sampled from the latent distribution. Through sampling, the VAE can generate diverse outputs, rather than just reconstructing the input. Then, decoding: The sampled latent representation is fed into the decoder, and the decoder attempts to reconstruct the original input data through neural network layers (MLP multi-layer perceptron). Furthermore, calculating the loss: The loss function of the VAE consists of two parts: the reconstruction error and the KL divergence (Kullback-Leibler Divergence). Here, the reconstruction error measures the difference between the reconstructed data and the predicted data, while the KL divergence measures the difference between the posterior distribution and the prior distribution of the latent representation. Finally, backpropagation and parameter update: According to the loss function, the gradient is calculated using the backpropagation algorithm, and the parameters of the encoder and decoder are updated through an optimization algorithm (Adam).
[0095] Among them, the preset loss function used for calculating the loss is as shown in the above expression (3).
[0096] In some embodiments, the expression of the preset reward function is:
[0097] (4)
[0098] Among them, is the main reward, is the task reward, is the style reward, ω g is the weight of the task reward, ωs The weight for style rewards.
[0099] Specifically, for the reward design part, as shown in Expression (4), the rewards are mainly divided into two parts: style rewards and task rewards.
[0100] For the task rewards, the method used in this patent mainly includes the x and y direction command tracking tasks and the smoothing rewards. The main reward methods are as follows:
[0101] (5)
[0102] Among them, rv is the linear velocity tracking reward, represents the reward coefficient, which can be adjusted according to experience, represents the task command, represents the actual velocity.
[0103] (6)
[0104] Among them, rw is the angular velocity tracking reward, represents the reward coefficient, which can be adjusted according to experience, represents the task command, represents the actual angular velocity.
[0105] The robot motion control method provided in this embodiment collects data through a fusion framework composed of a feature extraction network, a policy network, and an adversarial evaluation network, and then updates the policy network, the adversarial evaluation network, and the feature extraction network respectively based on the collected data, so as to obtain a well-trained policy network and a well-trained feature extraction network with more accurate style learning, making the behavior style of the humanoid robot closer to the behavior style of humans.
[0106] Figure 5 It is a schematic structural diagram of the robot motion control device provided in the embodiment of the present application. As Figure 5 shown, the robot motion control device 50 includes: an acquisition module 501, an input module 502, and an execution module 503.
[0107] The acquisition module 501 is used to acquire first historical data, second historical data, and a current first action instruction; the first historical data includes observation data corresponding to a first preset number of time steps before the current time step, and the second historical data includes observation data corresponding to a second preset number of time steps before the current time step; the second preset number is less than the first preset number.
[0108] An input module 502 is configured to input the first historical data into a trained feature extraction network to obtain first feature data; the first feature data includes terrain features and / or behavior features.
[0109] The input module 502 is further configured to input the current first action instruction, the first feature data, and the second historical data into a trained policy network to obtain a current first action; the trained policy network is a fusion framework composed of a feature extraction network, a policy network, and an adversarial evaluation network, and is obtained by training the policy network according to a reinforcement learning algorithm.
[0110] An execution module 503 is configured to determine a corresponding current first torque according to the current first action; and drive the robot to execute the current first action according to the current first torque.
[0111] The robot motion control device provided by the embodiments of the present application constructs a fusion framework based on a policy network, an adversarial evaluation network, and a feature extraction network, and can easily train a policy network to achieve humanoid motion by combining only a small amount of expert data sets. Moreover, by setting up a feature extraction network, feature data such as terrain features and behavior features can be extracted in advance and input into the policy network, making the action behaviors determined by the policy network more compliant and in line with human behavior styles, thereby improving the imitation quality of behavior styles.
[0112] In some embodiments, the trained feature extraction network is a convolutional neural network; the convolutional neural network includes a first hidden layer and a second hidden layer; the number of input layer channels of the first hidden layer is 42, the number of output layer channels is 32, the convolutional kernel size is 6, and the stride is 5; the number of input layer channels of the second hidden layer is 32, the number of output layer channels is 16, the convolutional kernel size is 4, and the stride is 2; the activation function of the convolutional neural network is a ReLU activation function.
[0113] In some embodiments, the device 50 further includes: a training module (not shown) configured to: construct a fusion framework based on a feature extraction network, a policy network, and an adversarial evaluation network; and train the policy network according to a proximal policy optimization (PPO) algorithm based on the fusion framework to obtain the trained policy network.
[0114] In some embodiments, the feature extraction network includes an encoder and a decoder; the encoder includes a convolutional neural network; the decoder includes a multi-layer perceptron.
[0115] In some embodiments, the training module is specifically configured to: for each time step, input the third historical data into the feature extraction network to obtain second feature data; the second feature data includes terrain features and / or behavior features; the third historical data includes a third preset number of actions before the previous action; input the current second action instruction, the previous reward, the second feature data, and the fourth historical data into the policy network to obtain the current second action; the fourth historical data includes a fourth preset number of actions between the previous actions; the fourth preset number is less than the third preset number; input the current first action into the controller of the robot to obtain the corresponding current second torque; input the current second torque into the environment model to obtain the current first state; input the current first state into the adversarial evaluation network to obtain the current style reward; based on a preset reward function, determine the current reward according to the current task reward and the style reward; store the current second action, the current first state, the current style reward, and the current reward into the experience replay buffer; sample sample data from the experience replay buffer; update the policy network according to the sample data based on the objective function of the PPO algorithm; update the adversarial evaluation network according to the sample data based on a preset objective function; update the feature extraction network according to the sample data based on a preset loss function.
[0116] In some embodiments, the expression of the objective function of the PPO algorithm is:
[0117]
[0118] Where, is the objective function of PPO, and the goal is to find the parameter θ to maximize the objective function; is the estimate of the advantage function, which represents the benefit of taking action a t in state w t relative to the average value; w t is is the new policy p θ relative to the old policy at action a t the probability ratio, that is, the importance sampling ratio; is the truncated policy ratio, where ɛ is a constant, ɛ ≤ 0.2, and the clip function ensures that the importance sampling ratio is within the range of [1 - ɛ, 1 + ɛ].
[0119] In some embodiments, the expression of the preset objective function is:
[0120]
[0121] Where, represents the real state and the expert reference state; represents the evaluation value output by the adversarial evaluation network for the state pair The closer the output value is to the expert state, the closer it is to 1, and vice versa, the closer it is to -1; is the divergence difference in the distribution between the real state and the expert reference state; is the weight coefficient.
[0122] In some embodiments, the expression of the preset loss function is:
[0123]
[0124] where represents historical observation data, represents the reconstructed next state, represents the next state, and MSE (mean square error) is the mean square error; is the posterior distribution, that is, the conditional distribution of the latent variable z given the input x; p(zt) is the prior distribution, that is, the marginal distribution of the latent variable z; KL is the KL divergence, which measures the difference between the approximate posterior distribution and the prior distribution p(zt).
[0125] In some embodiments, the expression of the preset reward function is:
[0126]
[0127] where is the main reward, is the task reward, is the style reward, ω g is the weight of the task reward, ω s is the weight of the style reward.
[0128] The robot motion control device provided by the embodiments of the present application can be used to execute the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0129] Figure 6 FIG. is a schematic hardware structure diagram of the robot motion control device provided by the embodiments of the present application, and the device can be a humanoid robot.
[0130] The device 60 may include one or more of the following components: a processing component 601, a memory 602, a power supply component 603, a multimedia component 604, an audio component 605, an input / output (I / O) interface 606, a sensor component 607, and a communication component 608.
[0131] The processing component 601 generally controls the overall operation of the device 60, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 601 may include one or more processors 609 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 601 may include one or more modules to facilitate the interaction between the processing component 601 and other components. For example, the processing component 601 may include a multimedia module to facilitate the interaction between the multimedia component 604 and the processing component 601.
[0132] The memory 602 is configured to store various types of data to support the operation of the device 60. Examples of such data include instructions for any application or method operating on the device 60, contact data, phone book data, messages, pictures, videos, etc. The memory 602 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.
[0133] The power component 603 provides power to various components of the device 60. The power component 603 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 60.
[0134] The multimedia component 604 includes a screen that provides an output interface between the device 60 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 604 includes a front camera and / or a rear camera. When the device 60 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0135] The audio component 605 is configured to output and / or input audio signals. For example, the audio component 605 includes a microphone (MIC), which is configured to receive external audio signals when the device 60 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 602 or transmitted via the communication component 608. In some embodiments, the audio component 605 further includes a speaker for outputting audio signals.
[0136] The I / O interface 606 provides an interface between the processing component 601 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0137] The sensor component 607 includes one or more sensors for providing a status assessment of various aspects of the device 60. For example, the sensor component 607 can detect the on / off state of the device 60, the relative positioning of components, such as the display and keypad of the device 60. The sensor component 607 can also detect a change in the position of the device 60 or a component of the device 60, the presence or absence of user contact with the device 60, the orientation or acceleration / deceleration of the device 60, and the temperature change of the device 60. The sensor component 607 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 607 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 607 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0138] The communication component 608 is configured to facilitate communication between the device 60 and other devices in a wired or wireless manner. The device 60 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 608 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 608 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0139] In an exemplary embodiment, the device 60 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0140] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 602 including instructions, and the above instructions can be executed by a processor 609 of the device 60 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0141] The above computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk. The readable storage medium may be any available medium accessible by a general-purpose or special-purpose computer.
[0142] An exemplary readable storage medium is coupled to the processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium may also be a component of the processor. The processor and the readable storage medium may be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium may also exist as discrete components in the device.
[0143] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program may be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical disks.
[0144] The embodiments of the present application also provide a computer program product including a computer program. When the computer program is executed by a processor, it implements the robot motion control method executed by the robot motion control device as described above.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A robot motion control method, characterized in that: include: Acquire first historical data, second historical data and current first action instruction; the first historical data includes observation data corresponding to a first preset number of time steps before the current time step, and the second historical data includes observation data corresponding to a second preset number of time steps before the current time step; the second preset number is less than the first preset number; Inputting the first historical data into a trained feature extraction network to obtain first feature data; the first feature data includes terrain features and / or behavior features; Inputting the current first action instruction, the first feature data and the second historical data into a trained policy network to obtain the current first action; the trained policy network is a fusion framework composed of a feature extraction network, a policy network and an adversarial evaluation network, and is obtained by training the policy network according to a reinforcement learning algorithm; According to the current first action, a corresponding current first torque is determined, and according to the current first torque, the robot is driven to perform the current first action.
2. The method according to claim 1, characterized in that: The trained feature extraction network is a convolutional neural network; the convolutional neural network includes a first hidden layer and a second hidden layer; the number of input layer channels of the first hidden layer is 42, the number of output layer channels is 32, the convolution kernel size is 6, and the step size is 5; the number of input layer channels of the second hidden layer is 32, the number of output layer channels is 16, the convolution kernel size is 4, and the step size is 2; the activation function of the convolutional neural network is the ReLU activation function.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: Construct a fusion framework based on feature extraction network, strategy network and adversarial evaluation network; Based on the fusion framework, the policy network is trained according to the proximal policy optimization (PPO) algorithm to obtain the trained policy network.
4. The method according to claim 3, characterized in that The feature extraction network includes an encoder and a decoder; the encoder includes a convolutional neural network; and the decoder includes a multi-layer perceptron.
5. The method according to claim 3, characterized in that: The method of training the policy network based on the fusion framework and according to the proximal policy optimization (PPO) algorithm to obtain the trained policy network includes: For each time step, inputting the third historical data into the feature extraction network to obtain second feature data; the second feature data includes terrain features and / or behavioral features; the third historical data includes a third preset number of actions before the last action; Inputting the current second action instruction, the previous reward, the second characteristic data and the fourth historical data into the strategy network to obtain the current second action; the fourth historical data includes a fourth preset number of actions between the previous action; the fourth preset number is less than the third preset number; Inputting the current first action into a controller of the robot to obtain a corresponding current second torque; Inputting the current second torque into the environment model to obtain a current first state; Inputting the current first state into the adversarial evaluation network to obtain a current style reward; Based on a preset reward function, determining a current reward according to a current task reward and the style reward; storing the current second action, the current first state, the current style reward, and the current reward in an experience replay buffer; Sampling sample data from the experience replay buffer; updating the policy network according to the sample data based on the objective function of the PPO algorithm; Based on a preset objective function, updating the adversarial evaluation network according to the sample data; Based on a preset loss function, the feature extraction network is updated according to the sample data.
6. The method according to claim 5, characterized in that The expression of the objective function of the PPO algorithm is: in, is the objective function of PPO, and the goal is to find the parameter θ to maximize the objective function; is an estimate of the advantage function, indicating that in state w t Take action a t Benefit relative to the average; w t for is the new strategy p θ Compared to the old strategy In action a t The probability ratio on , that is, the importance sampling ratio; is the cut-off strategy ratio, where ɛ is a constant, ɛ≤0.2, and the clip function ensures that the importance sampling ratio is in the range of [1-ɛ,1+ɛ].
7. The method according to claim 5, characterized in that The expression of the preset objective function is: in, Represents the real state and the expert reference state; Represents the adversarial evaluation network output for the state pair The output evaluation value is closer to 1 when it is closer to the expert state, and closer to -1 when it is less. is the distribution divergence difference between the true state and the expert reference state; is the weight coefficient.
8. The method according to claim 5, characterized in that The expression of the preset loss function is: in, is the observational data representing history, represents the next state of reconstruction, Represents the next state, MSE is the mean square error; is the posterior distribution, i.e., the conditional distribution of the latent variable z given the input x; p(zt) is the prior distribution, i.e., the marginal distribution of the latent variable z; KL is the KL divergence, which measures the approximate posterior distribution The difference from the prior distribution p(zt).
9. The method according to claim 5, characterized in that The expression of the preset reward function is: in, As the main reward, As a reward for the task, For style rewards, ω g is the weight of the task reward, ω s is the weight of the style reward.
10. A robot motion control device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the robot motion control method according to any one of claims 1 to 9.
Citation Information
Patent Citations
End-to-end on-orbit autonomous filling control system and method based on deep reinforcement learning
CN111844034A
Method for controlling a robot device and robot device controller
CN114063446A
Reinforcement learning model training and mobile robot control method and storage medium
CN116901091A
Mechanical arm control method and device based on pre-training model, equipment and medium
CN118418134A
Method for controlling a robotic device
US20220375210A1
Cited By
Whole-body control method and control device of robot, storage medium and robot
CN120395889A
Robot mechanical arm cooperative control method based on imitation learning
CN121132692A
Training method and device for squatting action control model of two-point-foot robot
CN121388606A
Robot motion control strategy network training method and device based on imitation learning, robot motion control method and device, equipment, robot and storage medium
CN121468592A
Motion control model training method and device and robot motion control method
CN122411128A