A diverse motion control method based on the latent space of action sequences

Through the method based on the hidden space of the action sequence, the hidden space of the controller is learned and the diversity controller is generated, which solves the problems of limited diversity and high resource demand in the prior art, and achieves the fine controllability and diversity improvement of the diversity of the controller.

CN116038711BActive Publication Date: 2025-07-25FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310058917.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2025-07-25
Estimated Expiration
2043-01-17

AI Technical Summary

Technical Problem

Existing diversity motion control methods usually only train a fixed number of diversified controllers, resulting in limited diversity, and it is difficult to select a suitable controller to adapt to changes when environmental changes are changed, and the training resource requirements are high.

Method used

Using a variety of motion control methods based on the hidden space of the action sequence, the basic action sequence is generated through the trajectory generator, and a priori action sequence distribution is constructed, and the hidden space of the controller is learned by the trajectory autoencoder, so as to achieve fine control of the core features of the controller, and generate infinite number of controllers with different core features.

Benefits of technology

The fine controllability of controller diversity is achieved, infinite multiple controllers can be generated to adapt to environmental changes, improve the degree of diversity, and avoid the limitation of resource demand with the increase in the number of controllers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116038711B_ABST
    Figure CN116038711B_ABST
Patent Text Reader

Abstract

The present invention provides a diverse motion control method based on the latent space of action sequences. This method selects the open-loop control method in motion control, abstracts the motion control process of the agent into a continuous action sequence, and regards this action sequence as the controller of the agent. The agent completes motion control by sequentially executing each action in the action sequence. Then, the unsupervised learning method of the variational autoencoder is used to learn the latent space of the controller, and the trained decoder is used to reconstruct the values in the latent space into an action sequence. Values in different latent spaces correspond to different controllers to control the agent to generate different motion patterns. Therefore, for a continuous latent space, theoretically, an infinite number of controllers can be obtained, thus greatly improving the degree of diversity. At the same time, the present invention realizes fine control of the diversity between controllers by controlling the values in the latent space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics, and particularly to a diverse motion control method based on the latent space of action sequences. Background Art

[0002] Motion control has always been an open and challenging task in the field of robotics. Whether it is classical motion control methods such as linear quadratic regulator (LQR) [1], model predictive control (MPC) [2], or current data-driven learning-based motion control methods such as proximal policy optimization (PPO) [3], deep deterministic policy gradient (DDPG) [4], etc., the goal is to obtain a single optimal or approximately optimal controller. However, a single controller becomes very vulnerable when facing a changing environment. Suppose our controller can only control the agent to move straight. Then when the environment changes, for example, when an obstacle appears on the straight path, the controller will collide with the obstacle, resulting in the breakdown of the agent's motion control. If multiple controllers can be learned at once, and each controller can control the agent to move in different directions, then a suitable controller can be selected to control the agent to bypass the obstacle. Therefore, it is necessary to learn multiple diverse controllers that are different from each other. When the environment where the agent is located changes, we can select a suitable controller from multiple controllers to control the agent to adapt to the environmental changes. This type of method that simultaneously learns multiple controllers that are different from each other is called diverse motion control.

[0003] Diversity-based motion control has always been a hot topic in the fields of evolutionary computation and reinforcement learning. Existing diversity-based motion control methods generally train a fixed number of controllers simultaneously. By maximizing a certain difference between the controllers, they can enable the agents to generate diverse motion behaviors. For example, in evolutionary computation, the heuristic Quality-Diversity (QD) algorithm is a representative method that can be used for diversity-based motion control. As a divergent search method, QD considers both the quality of the search solutions and the diversity between different solutions during training. QD will explicitly move away from previously visited places in the search space. In theory, the search of QD can also cover the entire search space. However, QD requires a manually defined behavior descriptor to measure the differences between search solutions, which undoubtedly limits the application of the QD method in complex tasks.

[0004] In addition, in reinforcement learning, researchers have also proposed a series of works to train multiple diverse control strategies. Most of these methods adopt a closed-loop control approach and train a fixed number of diverse control strategies through different diversity incentives [5-10]. For example, the Diversity via Determinants (DvD) algorithm [7] uses the determinant of a matrix to measure the diversity of controllers, and each value in the matrix represents the difference between controllers. By using the value of the determinant as an incentive reward and maximizing this incentive reward, DvD can search for diverse control strategies with different behavioral performances. However, such methods can often only train a fixed number of diverse control strategies. When the number of controllers gradually increases, the training resources required by the method will also increase accordingly. For example, for DvD, as the number of strategies increases, the calculation of the determinant will become intractable, which undoubtedly limits the degree of diversity.

[0005] In summary, most of the existing diversity-based motion control methods can only train a fixed number of diverse control strategies, which undoubtedly limits the degree of diversity. At the same time, the policy diversity is uncontrollable. It only encourages differences between policies, but lacks precise control over the differences between policies. It is very likely that there are no suitable candidates among these controllers that can be extrapolated to the changed environment. Finally, when the number of controllers increases, the training resources required by the method will also increase accordingly. Therefore, there are limitations in learning a fixed number of diverse control strategies. Summary of the Invention

[0006] To solve the above problems, a diversity-based motion control method is provided. The present invention adopts the following technical solutions:

[0007] The present invention provides a diverse motion control method based on the latent space of action sequences for controlling an agent to perform diverse motions. It is characterized by including the following steps: Step S1, training a trajectory generator to generate a basic action sequence and using the action sequence as the controller of the agent; Step S2, constructing a prior action sequence distribution based on the basic action sequence generated by the trajectory generator to represent the specific form of the controller space; Step S3, constructing a trajectory autoencoder that has an encoder and a decoder, continuously sampling action sequences based on the prior action sequence distribution and inputting them into the trajectory autoencoder for training to obtain the latent space of the sampled action sequences; Step S4, sampling the latent space and using the trained decoder to reconstruct the values in the latent space into action sequences for the agent to execute actions in the motion control task and exhibit diverse behavior control patterns based on the corresponding values in the latent space.

[0008] The diverse motion control method based on the latent space of action sequences provided by the present invention may also have the following technical feature. Among them, the trajectory generator is composed of a sine curve and a radial basis kernel function (RBF). The sine curve is two sine curve neurons, regarded as generators of control signals for generating basic motion patterns. The output of the neurons is shown in the following formula:

[0009]

[0010] In the formula, t is the current time step, f represents the frequency of curve oscillation, which has different settings for different control tasks, and T is the number of time steps corresponding to the period of curve oscillation;

[0011] The radial basis kernel function is used to shape the output of the sine curve. The activation function of the radial basis kernel function neuron is shown in the following formula:

[0012]

[0013] In the formula, H is the number of RBF neurons, T represents the motion period, u i,j represents the jth mean of the ith RBF neuron, and σ is the standard deviation of the RBF neuron.

[0014] The diverse motion control method based on the action sequence latent space provided by the present invention may further have the following technical features. The training process of the trajectory generator is as follows: Step S1-1, initialize the affine layer parameters of the trajectory generator, and optimize the affine layer parameters using the covariance matrix adaptation evolution strategy. Assign the obtained optimal parameters to the affine layer. Step S1-2, use the output of the sine curve as the torque control signal of the intelligent agent, and introduce a radial basis kernel function to shape the output of the sine curve. Step S1-3, after obtaining multiple periodic curves through the shaping in Step S1-2, perform a linear weighted combination of these periodic curves through one affine layer to obtain the final torque signal and output it, thereby controlling the intelligent agent to complete the motion control task.

[0015] The diverse motion control method based on the action sequence latent space provided by the present invention may further have the following technical features. Among them, the prior action sequence distribution is obtained by performing a distribution process on the basic action sequence. The specific construction process is as follows:

[0016] Model the prior distribution as a Gaussian distribution whose mean is the basic action sequence ξ output by the trajectory generator base , and the standard deviation is shown in the following formula:

[0017]

[0018] In the formula, σ is a constant used to control the values of the elements in Φ. Φ is a diagonal matrix, and the values on the diagonal are given by the function φ(t). φ(t) is a piecewise function. In the first half [0, T / 2] of the motion period, with η as the slope, the function value increases as the time step increases. In the second half [T / 2, T] of the motion control, with -η as the slope, the function value decreases as the time step increases. In this way, the expression form of the prior action sequence distribution is as follows:

[0019]

[0020] Thus, the prior distribution of the action sequence is constructed.

[0021] The diverse motion control method based on the action sequence latent space provided by the present invention may further have the following technical features. Among them, the encoder is used for encoding, that is, mapping the input action sequence to a low-dimensional continuous latent space. The redundant information in the action sequence is removed during the encoding process, aiming to represent the input action sequence through variables in a low-dimensional latent space. The decoder is used to reconstruct the latent space into an action sequence.

[0022] The diverse motion control method based on the latent space of action sequences provided by the present invention may also have the following technical features. The training method of the trajectory autoencoder is as follows: On the basis of maximizing the evidence lower bound, a weight factor is multiplied to learn the latent space of the action sequence. The expression for maximizing the evidence lower bound is as follows:

[0023]

[0024] In the formula, is the encoder, and its parameters are p θ (ξ|z) is the decoder, and its parameters are θ; p(z) is the prior distribution of the latent space z, which is a Gaussian distribution with a manually set mean of 0 and a variance of 1; D KL represents the KL divergence between two distributions, is the prior distribution of the action sequence ξ constructed in Equation 5; represents the reconstruction term, which aims to enable the trajectory autoencoder to reconstruct the input action ξ, represents the penalty term, which aims to constrain the encoder to be as close as possible to the prior distribution p(z) of the latent variable, thereby ensuring the simplicity of coding;

[0025] The value of the weight factor is proportional to the cumulative reward obtained by executing the action sequence. The weight factor makes the action sequences with high cumulative rewards have a greater proportion during update, so that the latent space focuses more on representing the core features of the action sequences with high cumulative rewards.

[0026] Functions and effects of the invention

[0027] According to a diverse motion control method based on the latent space of action sequences of the present invention, the open-loop control method in motion control is selected. The motion control process of the intelligent agent is abstracted into a continuous action sequence, and this action sequence is regarded as the controller of the intelligent agent. The intelligent agent completes motion control by sequentially executing each action in the action sequence. Then, the unsupervised learning method of the variational autoencoder is used to learn the latent space of the controller, and the trained decoder is used to reconstruct the value of the latent space into an action sequence. Different values in the latent space correspond to different controllers to control the intelligent agent to generate different motion patterns. Therefore, for a continuous latent space, theoretically, an infinite number of controllers can be obtained, thus greatly improving the degree of diversity. At the same time, by controlling the value of the latent space, fine control of the diversity between controllers is achieved.

[0028] Compared with the prior art, instead of being restricted to learning a fixed number of diversity controllers, this method examines the diversity-based motion control method from a novel perspective, focusing on learning the latent space that encodes the core features of the controllers. Controllers with different core features can control the robot to generate diverse behaviors, and theoretically, an infinite number of controllers with different core features can be realized. At the same time, this method can achieve fine-tunable diversity. When controllers with large differences from each other are needed, the differences between the latent space sampling values can be enlarged. Conversely, when multiple controllers with small differences from each other are needed, the differences between the latent space sampling values can be reduced. Description of the Drawings

[0029] Figure 1 is the framework diagram of the diverse motion control algorithm based on the action sequence latent space in the embodiment of the present invention;

[0030] Figure 2 is the flowchart of the diverse motion control method based on the action sequence latent space in the embodiment of the present invention;

[0031] Figure 3 is the schematic diagram of the signal output by the sine curve in the embodiment of the present invention;

[0032] Figure 4 is the schematic diagram of the output of the RBF neuron in the embodiment of the present invention;

[0033] Figure 5 is the schematic diagram of the torque curve generated by the trajectory generator in the embodiment of the present invention;

[0034] Figure 6 is the schematic diagram of the sample of the prior action sequence distribution in the embodiment of the present invention;

[0035] Figure 7 is the schematic diagram of the motion control task in the simulation environment in the experiment of the present invention;

[0036] Figure 8 is the training curve diagram of the trajectory generator in the motion control task in the experiment of the present invention;

[0037] Figure 9 is the schematic diagram of the result of the diversity experiment in the experiment of the present invention.

[0038] Figure 10 is the schematic diagram of the experimental result of the controllability of diversity in the experiment of the present invention. Detailed Embodiment

[0039] For the problem of diverse motion control, the present invention proposes a method for learning the latent space of a learning controller - the Diverse Motion Control Based on Latent Space of Action Sequence (LSAS-DMC) algorithm. The core idea of this method is: mapping the solution space where the controller is located into a low-dimensional continuous latent space, and there is a point-to-point corresponding relationship between the two spaces. Different controllers can be regarded as different points in the solution space, and at the same time, they also correspond to different points in the latent space. The latent space can be regarded as a low-dimensional space that encodes some core features of the controller (for example, the traveling direction and traveling speed), and the latent space is learned through an unsupervised learning method. Using the latent space, it is very convenient to generate different diverse controllers to complete the motion control task.

[0040] Figure 1 It is the framework diagram of the diverse motion control algorithm based on the latent space of action sequence in this embodiment.

[0041] The method adopted in this method is open-loop control. The simple and stable open-loop control method can well demonstrate the diversity of control. Under the setting of open-loop control, the motion control process of the agent is abstracted into an action sequence, that is, this action sequence is used as the controller of the agent, and the agent repeatedly executes the action sequence to complete the motion control.

[0042] As Figure 1 shown, this method consists of two stages. The first stage is to train a trajectory generator to generate a basic action sequence as the controller. The second stage is to train the latent space of the learning controller. First, the basic action sequence generated by the trajectory generator can be regarded as a point in the controller space and used as the initial solution. Subsequently, a prior distribution of the action sequence is constructed according to the output basic action sequence. Because the specific form of the controller (i.e., the action sequence) space is unknown, this means that it is impossible to obtain sufficient available controllers to learn the latent space representation of the controller. To solve this problem, this method is inspired by the work of trajectory optimization [11 - 12] of the robotic arm, and the prior distribution of the action sequence controller is constructed to characterize the controller space. Finally, action sequences are continuously sampled from the prior distribution of the controller, and a variational autoencoder (i.e., Figure 1The variational autoencoder (VAE) trains the latent space of the action sequence. As a generative model method in unsupervised learning, the variational autoencoder can well capture the core features of the action sequence, mapping the high-dimensional controller space to a low-dimensional latent variable space. Subsequently, samples are taken from the latent space of the controller, and the values in the latent space are reconstructed into an action sequence using the trained decoder. Finally, diverse behavior control patterns are executed and demonstrated in the motion control task.

[0043] In order to make the technical means, creative features, achieved objectives and effects realized by the present invention easy to understand, the following specifically describes a diverse motion control method based on the latent space of the action sequence of the present invention in combination with embodiments and the accompanying drawings.

[0044] <Embodiment>

[0045] Figure 2 is the flowchart of the diverse motion control method based on the latent space of the action sequence in the embodiment of the present invention.

[0046] As Figure 1 and Figure 2 shown, the specific process of the diverse motion control method based on the latent space of the action sequence is as follows:

[0047] Step S1, train a trajectory generator for generating a basic action sequence and use the action sequence as the controller of the intelligent agent.

[0048] The Trajectory Generator (TG) is a classic open-loop control method that uses the periodicity of robot motion to construct a motion controller. The structure of the trajectory generator used in this method is as Figure 1 shown on the left. It consists of a sine curve and a radial basis kernel function network.

[0049] The sine curve generates a rhythmic output that can simulate the rhythm of motion. The first part of the trajectory generator is two sine curve neurons, which can be regarded as generators of control signals, generating basic motion patterns. The output of the neurons is given by the following formula:

[0050]

[0051] where t is the current time step, and f represents the frequency of curve oscillation, which has different settings for different control tasks.

[0052] Figure 3 is a schematic diagram of the signal output by the sine curve in this embodiment.

[0053] The output of the sine curve is as Figure 3As shown, T is the number of time steps corresponding to the period of the curve oscillation, and in this embodiment, T is referred to as the motion period. Assuming that the time step length of the motion control task is 0.01 s and the frequency f of the sine curve is 10 Hz, the output of the sine curve has a period every 10 time steps, that is, T = 10.

[0054] Since the selection and design of the sine curve are highly subjective, directly using the output of the sine curve as the torque control signal often fails to well complete the motion control task. Therefore, the improved method used in this embodiment is to introduce a Radial Basis Function (RBF) after the output of the sine curve to shape the output of the sine curve. The input of the RBF is the output of the sine curve, and the activation function of the RBF neuron is given by the following formula:

[0055]

[0056] In the formula, H is the number of RBF neurons, T represents the motion period, u i,j represents the jth mean of the ith RBF neuron, and σ is the standard deviation of the RBF neuron. It can be seen from Equation 2 that the means of each RBF neuron are equally spaced and have the same standard deviation. The outputs of two sine functions are reshaped into H oscillating curves with the same amplitude, period, and different phases after passing through H RBF neurons. Figure 4 shows the outputs of 10 RBF neurons.

[0057] After obtaining multiple periodic curves through reshaping, a linear weighted combination of these periodic curves is performed through an affine layer to obtain the output of the final torque signal. Assuming that the action space of the motion control task is 3-dimensional (i.e., 3 torque signals need to be output), the parameter W of the affine layer is a [H, 3] matrix. By modifying the parameter matrix W, the trajectory generator can output torque signals in any pattern.

[0058] In this embodiment, the Covariance Matrix Adaptation Evolution Strategies (CMA-ES) is used to optimize the parameter matrix W. Since W has only a few hundred parameters, CMA-ES, as an advanced gradient-free optimization algorithm, can well search for appropriate parameters so that the trajectory generator can output torque signals to control the robot to complete the motion control task.

[0059] Figure 5 is a schematic diagram of the torque curve generated in the embodiment of the present invention.

[0060] Figure 5Shows the torque signal curves generated by the trajectory generator for the front thigh, front shin, and front foot of the HalfCheetah robot. Each time step t corresponds to a point on the curve, which constitutes the action a at the current time step. t , then the controller is a sequence of actions [a1, a2, …… a T within a motion cycle. This sequence of actions serves as the basic action sequence ξ base for use in the next step. Based on the above process description, in this embodiment, the pseudocode for training the trajectory generator is as follows:

[0061] Algorithm 1 Training of the Trajectory Generator

[0062] Input: Sine curve Motion cycle T, length LEN of Rollout, number of iterations N, number of RBF neurons H, and standard deviation σ, population size E of CMA - ES

[0063] Output: Basic action sequence ξ base = [a1, a2, …… a T

[0064]

[0065]

[0066] So far, through the trajectory generator, a basic action sequence can be obtained. The robot sequentially executes each action in the sequence to complete forward movement.

[0067] The reason for using the trajectory generator is that it is simple and easy to train. Moreover, the purpose of this method is to obtain only a basic controller for subsequent diverse control. Therefore, the trajectory generator is a very suitable choice.

[0068] Step S2: Construct a prior action sequence distribution based on the basic action sequence generated by the trajectory generator to represent the specific form of the controller space.

[0069] Compared with learning a fixed number of diverse controllers, this method attempts to learn the internal latent space of the controller. The basic idea is to map the controller to a low - dimensional continuous latent space by leveraging the latent space modeling ability of the generative model. The latent space contains some core features of the controller. Uniformly sample latent variables from the latent space, and different latent variables can be reconstructed into different controllers. Different controllers control the robot to complete motion control tasks and exhibit different diverse motion patterns.

[0070] ​The premise of latent space learning is to know the specific form of the controller space so that available controllers can be sampled from it to train the latent space representation of the controller. However, the space of the controller (i.e., the action sequence) is unknown. Therefore, a pre-defined prior distribution of the controller is introduced to replace the solution space of the controller. The basic action sequence ξ generated by the trajectory generator base can be regarded as a point in the controller space and used as the initial solution. Subsequently, based on ξ base a prior distribution of the action sequence is constructed. For the motion control of a robot, we expect that in each motion cycle T, the diversity change in the start and end phases is relatively small, while more diversity is shown in the middle phase of the motion cycle. This is because the starting state and ending state of the motion control task have a relatively small range of variability, and large-scale changes may lead to task failure (such as tipping over or rolling over), while diverse motion patterns can usually be shown in the middle phase of the motion.

[0071] Based on the above idea, this embodiment proposes a way to distribute the basic action sequence. Specifically:

[0072] The prior distribution is modeled as a Gaussian distribution N with the mean being the basic action sequence ξ output by the trajectory generator base , and the standard deviation is shown as follows:

[0073]

[0074] where σ is a constant used to control the values of the elements in Φ. Φ is a diagonal matrix, and the values on the diagonal are given by the function φ(t). φ(t) is a piecewise function. In the first half of the motion cycle [0, T / 2], the function value increases with the increase of the time step (with η as the slope), and in the second half of the motion control [T / 2, T], the function value decreases with the increase of the time step (with -η as the slope). The expression form of the prior trajectory (action sequence) distribution is as follows:

[0075]

[0076] Thus, the prior distribution of the action sequence controller is constructed.

[0077] Step S3: Construct a trajectory autoencoder, which has an encoder and a decoder. Based on the prior action sequence distribution, action sequences are continuously sampled and input into the trajectory autoencoder for training to obtain the latent space of the sampled action sequences.

[0078] Step S4: Sample the latent space, and use the trained decoder to reconstruct the values in the latent space into action sequences for the agent to execute actions in the motion control task and show diverse behavior control patterns based on the corresponding values in the latent space.

[0079] After constructing the prior distribution of the controller new action sequences ξ are continuously sampled from it. Subsequently, the action sequences are input into the trajectory autoencoder, and the action sequences are reconstructed through the "encoding - decoding" process.

[0080] The essence of the trajectory autoencoder is a variational autoencoder, which is a mainstream type of generative model. Specifically, the trajectory autoencoder consists of an encoder and a decoder p θ (ξ|z) in two parts. The purpose of the encoder is to map the input action sequence ξ into a low - dimensional continuous latent space z, and this process is called "encoding". The encoding process eliminates redundant information in the action sequence and aims to represent the input action sequence through variables in a low - dimensional latent space. The purpose of the decoder p θ (ξ|z) is to reconstruct the latent space z into the action sequence ξ. Through the two processes of encoding and decoding, the latent space z becomes the "bottleneck" between the encoder and the decoder. This means that the latent space z must represent the core features of the action sequence ξ input by the encoder so that the decoder can well reconstruct the latent space z back into the original action sequence ξ.

[0081] The latent space is a low - dimensional continuous space that represents the core features of the action sequence controller. For motion control, the core features include but are not limited to motion speed, motion direction, motion posture, etc. The diversity of motion control is precisely reflected in the differences of these core features. Therefore, if controllers with different core features can be designed, the motion of the intelligent agent can be made diverse. However, these core features are very difficult to directly represent through the controller, and the latent space can well model these core features.

[0082] In this embodiment, the set latent space is a 1 - dimensional real - valued continuous space, which means that it is very convenient to sample from the latent space. At the same time, the decoder p θ (ξ|z) decodes the values in the latent space into the action sequence controller. Different values in the latent space correspond to different core features, and action sequences with different core features can control the intelligent agent to generate different motion patterns. With the help of the latent space, diverse action sequences can be generated to show the diversity of motion control. Since the latent space is a continuous space, theoretically, an infinite number of different controllers can be generated. In addition, the latent space makes the differences in the core features between controllers controllable, so that the differences between the motion patterns of the intelligent agent are controllable. For example, controllers with similar values in the latent space have similar core features, and vice versa. Different values in the latent space result in different core features of the decoded action sequences, so the learning of the latent space makes the motion patterns of the intelligent agent finely controllable.

[0083] In this embodiment, the training method of the trajectory autoencoder is similar to that of the variational autoencoder. The latent space of the action sequence is learned by maximizing the evidence lower bound (ELBO). ELBO is given by Equation 6:

[0084]

[0085] where is the encoder with parameters p θ (ξ|z) is the decoder with parameters θ; p(z) is the prior distribution of the latent space z, which is a Gaussian distribution with a manually set mean of 0 and variance of 1; D KL represents the KL divergence between two distributions, is the prior distribution of the action sequence ξ constructed in Equation 5; represents the reconstruction term, which aims to enable the trajectory autoencoder to reconstruct the input action ξ, represents the penalty term, which aims to constrain the encoder to be as close as possible to the prior distribution p(z) of the latent variable, thereby ensuring the simplicity of encoding.

[0086] However, there are problems with directly using Equation 6 to train the trajectory autoencoder. Because the prior distribution is a Gaussian distribution with ξ base as the mean, the action sequence ξ sampled from the prior distribution may be very different from ξ base (although the probability is very small), which often leads to the failure of motion control. Figure 6 shows the action sequence sampled from . The thick line represents the basic action sequence output by the trajectory generator, and the thin line is the action sequence sampled from the prior trajectory distribution. The diagonal matrix Φ makes the variance of the action sequence smaller at the beginning and end of the motion period T, and larger in the middle stage, with more room for variation.

[0087] To avoid learning the latent space representation of such extreme action sequences, this embodiment multiplies Equation 6 by a weight factor λ, and the value of the weight factor is proportional to the cumulative reward R(ξ) obtained by executing the action sequence ξ, that is, λ ∝ exp(R(ξ)). The weight factor makes the action sequences with high cumulative rewards have a greater proportion during the update, so that the latent space focuses more on representing the core features of the action sequences with high cumulative rewards. According to the above training process and the updated formula, the pseudo-code process corresponding to the algorithm is shown in Algorithm 2 below.

[0088] Algorithm 2 Learning of the Controller Latent Space

[0089] Input: Prior distribution of the action sequence controller Standard normal distribution p(z), dimension m of the latent space, number of iterations N, batch size B, maximum length LEN of Rollout

[0090] Output: Encoder And decoder p θ (ξ|z)

[0091]

[0092]

[0093] The above Algorithm 1 and Algorithm 2 together constitute the process of the LSAS-DMC algorithm of the present invention.

[0094] Figure 7 It is a schematic diagram of the motion control task in the simulation environment in the embodiment of the present invention.

[0095] As Figure 7 shown, in this embodiment, 4 motion control tasks in the simulation environment are selected to evaluate the performance of the LSAS-DMC algorithm. The selected tasks cover the motion control of various robots, and the purpose of each control task is to control the robot to move forward. It should be noted that Figure 7 (c) and (d) in use the same robot, but they are two different motion control tasks. (c) aims to learn the Trot gait, while (d) aims to learn the Gallop gait. Since the HalfCheetah-v3 task is a 2D motion control task, the main evaluation is whether LSAS-DMC can discover new motion postures (such as jumping, somersaulting, etc.). For the remaining motion control tasks, the main evaluation is whether LSAS-DMC can achieve motion control of different core features (such as motion direction and speed).

[0096] Parameter settings: The length of all motion control task Rollouts is 200. The number of iterations for training the trajectory generator is 500, the number of RBF neurons is 100, and the standard deviation is -0.5π. The encoder and decoder of the trajectory autoencoder are both MLP neural networks, the hidden layer is [256, 512, 256], the dimension of the latent variable is 1, the update batch size is 32, the learning rate is 0.001, and the number of iterations is 200.

[0097] Train the trajectory generator using the above parameters:

[0098] LSAS-DMC needs to train a trajectory generator to generate a basic action sequence controller. For the trajectory generator, the parameter to be optimized is the parameter W of the affine layer. We expect to modify the parameter W so that the action sequence corresponding to the torque signal generated by the trajectory generator can control the robot to complete the motion control task. The dimension of W generally does not exceed 1000 dimensions. For example, for the Ant-v3 task, its action dimension is 8, that is, the trajectory generator needs to output 8 torque signals. The number of RBF neurons selected in this experiment is 100, and the number of parameters of W is 800. For lower-dimensional parameters, the gradient-free optimization algorithm CMA-ES can optimize well. The training process of the trajectory generator is shown in Algorithm 1 in detail.

[0099] Figure 8 The training curves of the trajectory generator in 4 motion control tasks are shown. The horizontal axis is the number of iterations of CMA-ES, and the vertical axis is the cumulative reward of Rollout. The optimization goal of CMA-ES is to maximize the cumulative reward.

[0100] Figure 8 The dark color in it represents the cumulative reward corresponding to the optimal parameters searched during the entire training process, the light color represents the average result of training with 5 random seeds, and the shaded part is the standard deviation. After 500 iterations of update, the reward curve gradually converges. For all tasks, CMA-ES can search for excellent parameter W to complete the motion control task. This also reflects the superiority of the trajectory generator, which is simple and easy to train and can conveniently provide an action sequence controller. After the training is completed, the action sequence output under the optimal parameters is used as the basic action sequence ξ base 。

[0101] Diversity comparison and diverse motion control demonstration:

[0102] Use the basic action sequence ξ generated by the trajectory generator base Construct the prior distribution of the controller According to the process of Algorithm 2, train the trajectory autoencoder for each motion control task. The number of iterations for training the trajectory autoencoder is 200, and the set dimension of the latent space z is 1 dimension. Through experiments, it is found that for 4 motion control tasks, the values of the latent space are distributed between -2 and 2. Therefore, during diverse motion control, this experiment uniformly samples the required number of values of the latent space z in the interval from -2 to 2 and inputs them into the decoder p θIn ($\xi|z$), the decoding generates the action sequence controller $\xi$. The agent repeatedly executes $\xi$ and performs Rollout until the motion control task ends. As mentioned before, the values in the latent space encode the core features of the action sequence controller. Different values in the latent space correspond to action sequences with different core features, thus controlling the robot to exhibit different motion patterns. The subsequent diversity control experiments are all carried out based on the above method.

[0103] In this experiment, the sota methods DvD and LSAS-DMC in the diversity motion control method are also selected for diversity comparison.

[0104] For DvD, 10 diverse neural network controllers are trained simultaneously for each motion control task. When testing the diversity, each controller controls the robot to perform a Rollout of length 200 in the simulation environment, and the trajectories generated by the 10 controllers are collected. For LSAS-DMC, 10 values are uniformly sampled from the interval [-2, 2) of the latent space. For each motion control task, the values in the latent space are decoded to generate 10 action sequence controllers, and each controller controls the robot to perform a Rollout of length 200 in the simulation environment and collect the trajectories. The trajectories include a total of 200 actions executed by the robot from the start to the end of the motion.

[0105] In the experiment, the method introduced in the DvD paper is used to evaluate the diversity of the trajectories. Specifically, for any two trajectories $\tau$ i and $\tau$ j , the Gaussian kernel function $k$ (i,j) = exp(-||$\tau$ i -$\tau$ i || 2 / 2) is used to evaluate the difference between them. $k$ (i,j) = 1 indicates that the two trajectories are the same, and $k$ (i,j) = 0 indicates that the two trajectories are orthogonal. For each trajectory, the difference $k$ with the remaining 9 trajectories is calculated, and finally a difference matrix $K(\tau$ i ,$\tau$ j ), i, j = 1,..., 10 can be formed. By calculating the determinant of the difference matrix, the overall difference of the sampled trajectories is evaluated, which also reflects the diversity of the trajectories, i.e., Div = det(K). Table 1 below shows the diversity evaluated by the above method.

[0106] Table 1 Comparison of diversity scores for motion control tasks

[0107]

[0108] As can be seen from Table 1 above, for the 4 motion control tasks, LSAS-DMC outperforms DvD in 3 tasks, demonstrating the superiority of LSAS-DMC in diverse motion control. Compared with the method of learning a fixed number of diverse controllers, the advantage of LSAS-DMC lies in solving the diverse control problem from the perspective of generating tasks. By directly learning the latent space of the controller to generate diverse controllers, it is not limited by the number of controllers, improves the degree of diversity, and avoids the limitations of learning a fixed number of controllers. At the same time, the diversity between controllers is controllable, and this advantage is demonstrated based on the following experiments.

[0109] Demonstration of diverse motion control:

[0110] For motion control tasks, we expect the robot to exhibit diversity in motion direction and speed. Different forward directions can well avoid obstacles in the environment, and different forward speeds can cope with uneven roads. Therefore, both of these can be regarded as the core features that diverse motion control wants to capture, and controllers with different core features can control the robot to generate diverse motion patterns. This experiment verifies whether the latent space trained by LSAS-DMC encodes the core features through a diversity experiment, and whether the generated action sequence controllers can control the robot to generate diverse behaviors. Specifically:

[0111] Uniformly sample 30 values from the interval [-2, 2) for the latent space and generate the action sequence controllers corresponding to the motion control tasks, that is, each motion control task has 30 decoded action sequence controllers. In each motion control task, the action sequence controls the robot to perform a Rollout of 200 time steps, and records the position and speed of the robot during the motion.

[0112] Figure 9 It is a schematic diagram of the results of the diversity experiment in the embodiment of the present invention.

[0113] Figure 9 Shows the diversity of motion direction and speed in three tasks (different grayscales represent different values of the latent space). The first row shows the changes in the x-axis and y-axis positions of the robot during the motion, and the second row represents the changes in the x-axis and y-axis speeds during the motion. Curves of different colors represent the controllers decoded from different values of the latent space. From Figure 8 It can be seen that in the 3 control tasks, the motion positions of the robot are all radial. According to different values of the latent variable, the traveling directions of the robot are different, and the traveling speeds are also different, which well demonstrates the diversity of motion control. This shows that the latent space of the controller learned by LSAS-DMC encodes the core features of the controller well. Different values of the latent space represent controllers with different core features, and control the robot to exhibit diverse behaviors.

[0114] In addition to controlling the robot to exhibit diverse behaviors, another advantage of LSAS-DMC is that it makes diversity finely controllable. The values in the latent space can be regarded as a mapping of the core features of the controller. Therefore, the degree of diversity of the controller can be influenced by the values in the latent space, enabling the robot to better cope with environmental changes. For example, when it is necessary to bypass a large obstacle, values in the latent space with a large interval can be taken, making the differences between the robot's movement directions larger, so that the obstacle can be bypassed; or when the robot needs to pass through a small pass, values in the latent space with a small interval can be taken, making the differences between the robot's movement directions smaller, so as to find the most suitable controller to control the robot to pass through the pass.

[0115] Figure 10 Intuitively demonstrates the controllability of LSAS-DMC for diversity. The differences between the controllers corresponding to the latent space values z = 2.0 and z = -2.0 are large, while the differences between the controllers corresponding to z = 0.2 and z = -0.2 are small. LSAS-DMC realizes the controllability of diversity through the latent space, which improves the practicality of the method. It not only pursues diversity, but more importantly, pursues controllable diversity to meet actual needs, which undoubtedly reflects the superiority of this method.

[0116] Figure 10 Only three motion control tasks were involved in the diversity experiment shown, and HalfCheetah was missing. This is because this task is a 2D motion control and cannot reflect diversity in the movement direction, so it is not applicable to the above diversity experiment. However, HalfCheetah can exhibit different motion postures during forward movement. Therefore, for the HalfCheetah task, we focused on evaluating whether LSAS-DMC can discover diverse motion postures. Figure 10 Shows the discovery of diverse motion postures by LSAS-DMC. It can be seen that as the latent space values vary, the robot moves forward with different motion postures. We selected several obvious motion postures. For example, when z = 1.6, the robot moves forward in a somersault posture; when z = 0.4, the robot moves forward in a leaning-down posture; when z = -0.6, the robot jumps forward; when z = -2.0, the robot moves forward in a pouncing posture. Different latent space values enable the controller to control the robot to produce motion postures similar to those of animals in the real world, demonstrating the diverse motion control ability of this method.

[0117] Functions and effects of the embodiment

[0118] According to the diverse motion control method based on the latent space of action sequences in this embodiment, the open-loop control method in motion control is selected, and the motion control process of the agent is abstracted into a continuous action sequence. This action sequence is regarded as the controller of the agent, and the agent sequentially executes each action in the action sequence to complete motion control. This method also samples from the latent space of the controller, and uses the trained decoder to reconstruct the values in the latent space into an action sequence. The values in different latent spaces correspond to different controllers to control the agent to generate different motion patterns. For the continuous latent space, theoretically, an infinite number of controllers can be obtained, thus greatly enhancing the degree of diversity. At the same time, by controlling the values in the latent space, fine control of the diversity between controllers is achieved.

[0119] Compared with the prior art, this method does not stick to learning a fixed number of diverse controllers, but examines the diversity-based motion control method from a novel perspective, focusing on learning the latent space that encodes the core features of the controller. Controllers with different core features can control the robot to generate diverse behaviors, and theoretically, an infinite number of controllers with different core features can be realized. At the same time, this method can achieve finely controllable diversity. When controllers with large differences from each other are needed, the differences between the sampled values in the latent space can be enlarged. On the contrary, when multiple controllers with small differences from each other are needed, the differences between the sampled values in the latent space can be reduced.

[0120] In addition, this method also demonstrates the effectiveness of the LSAS-DMC method in a series of motion control tasks in continuous action spaces, and it can control the agent to generate diverse behaviors.

[0121] The above embodiments are only used to illustrate the specific implementation manners of the present invention, and the present invention is not limited to the description scope of the above embodiments.

[0122] The above references are as follows:

[0123] [1]Li W,Todorov E.Iterative linear quadratic regulator design fornonlinear biological movement systems[C] / / ICINCO(1).2004:222-229.[2]Camacho EF,Alba C B.Model predictive control[M].Springer science&business media,2013.

[0124] [3]Schulman J,Wolski F,Dhariwal P,et al.Proximal policy optimizationalgorithms[J].arXiv preprint arXiv:1707.06347,2017.

[0125] [4]Lillicrap T P,Hunt J J,PritzelA,et al.Continuous control with deepreinforcement learning[J].arXiv preprint arXiv:1509.02971,2015.

[0126] [5]Eysenbach B,Gupta A,Ibarz J,et al.Diversity is all you need:Learning skills without a reward function[J].arXiv preprint arXiv:1802.06070,2018.

[0127] [6]Kumar S,KumarA,Levine S,et al.One solution is not all you need:Few-shot extrapolation via structured maxent rl[J].Advances in NeuralInformation Processing Systems,2020,33:8198-8210.

[0128] [7]Parker-Holder J,Pacchiano A,Choromanski K M,et al.Effectivediversity in population based reinforcement learning[J].Advances in NeuralInformation Processing Systems,2020,33:18050-18062.

[0129] [8]Sharma A,Gu S,Levine S,et al.Dynamics-aware unsupervised discoveryofskills[J].arXiv preprint arXiv:1907.01657,2019.

[0130] [9]Pierrot T,MacéV,Chalumeau F,et al.Diversity Policy Gradient forSample Efficient Quality-Diversity Optimization[C] / / ICLR Workshop on AgentLearning in Open-Endedness.2022.

[0131]

[10] Zahavy T,Schroecker Y,Behbahani F,et al.Discovering Policies withDOMiNO:Diversity Optimization Maintaining Near Optimality[J].arXiv preprintarXiv:2205.13521,2022.

[0132]

[11] Kalakrishnan M,Chitta S,Theodorou E,et al.STOMP:Stochastictrajectory optimization for motion planning[C] / / 2011 IEEE internationalconference on robotics and automation.IEEE,2011:4569-4574.

[0133]

[12] Osa,Takayuki."Motion planning by learning the solution manifoldin trajectory optimization."The International Journal of Robotics Research41.3(2022):281-311.

Claims

1. A diverse motion control method based on the latent space of action sequences, used to control an agent to perform diverse motions, characterized in that, It includes the following steps: Step S1, training a trajectory generator to generate a basic action sequence and using the action sequence as the controller of the agent; Step S2, constructing a prior action sequence distribution based on the basic action sequence generated by the trajectory generator to represent the specific form of the controller space; Step S3, constructing a trajectory autoencoder which has an encoder and a decoder. Based on the prior action sequence distribution, continuously sample action sequences and input them into the trajectory autoencoder for training to obtain the latent space of the sampled action sequences; Step S4, sampling the latent space and using the trained decoder to reconstruct the values of the latent space into action sequences for the agent to execute actions in the motion control task and display diverse behavior control modes based on the corresponding values of the latent space. The trajectory generator consists of a sine curve and a radial basis kernel function RBF; The sine curve is composed of two sine curve neurons, regarded as the generator of control signals for generating basic motion patterns. The output of the neuron is as follows: In the formula, t is the current time step, and f represents the frequency of curve oscillation, which has different settings for different control tasks; The radial basis kernel function is used to shape the output of the sine curve. The activation function of the radial basis kernel function neuron is as follows: where H is the number of RBF neurons, T represents the motion period, and u i,j represents the j-th mean value of the i-th RBF neuron, and σ is the standard deviation of the RBF neuron, The training process of the trajectory generator is as follows: Step S1-1, initializing the affine layer parameters of the trajectory generator and optimizing the affine layer parameters using the covariance matrix adaptation evolution strategy, and assigning the obtained optimal parameters to the affine layer; Step S1-2, using the output of the sine curve as the torque control signal of the agent and introducing the radial basis kernel function to shape the output of the sine curve; Step S1-3, after obtaining multiple periodic curves through the shaping in Step S1-2, performing a linear weighted combination of these periodic curves through an affine layer to obtain the final torque signal and output it, thereby controlling the agent to complete the motion control task.

2. A diverse motion control method based on the latent space of action sequences according to claim 1, wherein: Among them, The prior action sequence distribution is obtained by performing a distribution process on the basic action sequence. The specific construction process is as follows: Model the prior distribution as a Gaussian distribution whose mean is the base action sequence ξ output by the trajectory generator base , and the standard deviation is shown as follows: In the formula, σ is a constant used to control the values of the elements in φ. φ is a diagonal matrix, and the values on the diagonal are given by the function φ(t). φ(t) is a piecewise function. In the first half [0, T / 2] of the motion cycle, with η as the slope, the function value increases as t increases. In the second half [T / 2, T] of the motion cycle, with -η as the slope, the function value decreases as t increases. Thus, the expression form of the prior action sequence distribution is as follows: Thus, the prior action sequence distribution is constructed.

3. A diverse motion control method based on the latent space of action sequences according to claim 1, wherein: Among them, The encoder of the trajectory autoencoder is used for encoding, that is, mapping the input action sequence to a low-dimensional continuous latent space. The encoding process eliminates the redundant information in the action sequence, aiming to represent the input action sequence through variables in a low-dimensional latent space. The decoder is used to reconstruct the latent space into an action sequence.

4. A diverse motion control method based on the latent space of an action sequence according to claim 1, characterized in that: Among them, The training method of the trajectory autoencoder is: multiplying by a weight factor on the basis of maximizing the evidence lower bound to learn the latent space of the action sequence, The expression for maximizing the evidence lower bound is as follows: wherein, is an encoder, and its parameter is p θ (ξ|z) is a decoder, and its parameter is θ; p(z) is the prior distribution of the latent space z, which is a Gaussian distribution with a manually set mean of 0 and a variance of 1; D KL represents the KL divergence between the prior distribution of the latent space z and the standard Gaussian distribution, is the prior distribution of the action sequence ξ; represents the reconstruction term, which aims to enable the trajectory autoencoder to reconstruct the input action sequence ξ, represents the penalty term; The value of the weight factor is proportional to the cumulative reward obtained by executing the action sequence.

Citation Information

Patent Citations

  • Motion sequence generation method based on conditional generative adversarial networks (GANs)

    CN108596149A

  • Method for reinforcement learning exploration and utilization of trajectory space determinant point process

    CN113239629A