Method and apparatus for generating action of simulation character and training model thereof

By jointly training the environment model and the joint torque generation model, the problems of slow training speed and sluggish movements in the motion generation process of simulated characters are solved, and the flexibility of end-to-end training and motion generation is improved.

CN115908652BActive Publication Date: 2026-04-28PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEKING UNIV
Filing Date
2022-09-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, the training speed of the motion generation process for simulated characters is slow, and the generated simulated characters have stiff movements that cannot be effectively applied to downstream tasks.

Method used

An environment model is used instead of a simulator and is jointly trained with the joint torque generation model of the simulated character. The environment model generates predicted simulated posture data based on the torque of each joint of the simulated character, and uses it for autoregressive training of the joint torque generation model and the environment model to achieve end-to-end training.

Benefits of technology

The training speed of the motion generation process for simulated characters has been improved, and the generated simulated characters have more flexible movements, making them better applicable to downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908652B_ABST
    Figure CN115908652B_ABST
Patent Text Reader

Abstract

The application provides a simulation role action generation method and device and a simulation role action generation model training method and device. The simulation role action generation method comprises the following steps: obtaining simulation posture data corresponding to a current time of a simulation role and a current control signal used for controlling the simulation role; generating target simulation posture data based on the simulation posture data corresponding to the current time of the simulation role, the current control signal, a pre-trained joint torque generation model of the simulation role, and a simulation environment; the joint torque generation model is used for determining torque of each joint of the simulation role; the joint torque generation model is obtained through joint training with an environment model; the environment model is used for generating predicted simulation posture data based on the torque of each joint of the simulation role, and the predicted simulation posture data is used for autoregressive training of the joint torque generation model and the environment model of the simulation role. Based on this, the problem of slow training speed and generated simulation action stiffness in the simulation role action generation process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and apparatus for generating the motion of a simulated character and training its model. Background Technology

[0002] In the field of computer graphics, numerous studies have focused on data-driven motion generation for simulated characters. Data-driven motion generation for simulated characters is a widely studied and important problem in graphics, aiming to generate new character movements from existing motion data, enabling them to perform different downstream tasks in various environments. To make the movements of simulated characters more realistic, motion capture technology is generally used to acquire human motion data, such as running and jumping. Simulated characters learn and imitate these movements in a simulated environment and apply them to various unseen environments and tasks. In computer graphics, simulated characters apply the learned human movements to tasks including walking on different terrains, interactive control, and the transfer of control from simulated characters to controlling real robots.

[0003] Similar generative tasks have been extensively studied in other fields such as image processing, resulting in numerous generative models. However, in the task of generating motion for simulated characters, the end-to-end gradient descent training required for the generative model is no longer suitable due to the black-box nature of the simulator. Existing motion generation techniques for simulated characters either abandon end-to-end training in favor of multi-stage training or use sampling for optimization instead of direct gradient descent. This leads to generative models performing significantly worse on this task than tasks like image generation, which can utilize direct gradient descent. Specifically, this manifests as longer training times, more static motion in the generated simulated characters, and the need for extensive repetitive training for new downstream tasks.

[0004] Therefore, how to solve the technical problem of slow training speed and stiff animation of simulated characters in the existing technology is an important research direction. Summary of the Invention

[0005] This invention provides a method and apparatus for generating and training the motion of a simulated character, thereby solving the technical problems of slow training speed and stiff motion of the generated simulated characters in the prior art.

[0006] This invention provides a method for generating motion of a simulated character, comprising: acquiring simulated posture data of the simulated character at the current moment and a current control signal for controlling the simulated character; generating target simulated posture data based on the simulated posture data of the simulated character at the current moment, the current control signal, a pre-trained joint torque generation model of the simulated character, and a simulation environment; the target simulated posture data being the simulated posture data of the simulated character at the next moment corresponding to the current control signal; wherein, the joint torque generation model of the simulated character is used to determine the torque of each joint of the simulated character; the joint torque generation model of the simulated character is obtained through joint training with an environment model; the environment model is used to generate predicted simulated posture data based on the torque of each joint of the simulated character, and the predicted simulated posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model.

[0007] Based on any of the above embodiments, in this embodiment, the joint torque generation model of the simulated character includes a strategy model. The strategy model is determined based on the historical simulation posture data and historical control signals corresponding to the simulated character. The strategy model is used to determine the offset in the encoding space, and the historical control signal g includes the orientation d input by the user. i and speed v i Accordingly, the training process of the strategy model includes: acquiring historical simulation posture data s corresponding to the simulation character. t and historical control signal g; based on the historical simulation posture data and historical control signal corresponding to the simulation character, the strategy model of the simulation character is determined according to the first preset loss function and the second preset loss function; wherein, the first preset loss function and the second preset loss function are used together to control the direction and speed of the simulation character's movement; wherein, the first preset loss function is: The second preset loss function is: l r =|M d (s t ,v t ,d t )|;wherein, d' i v i 'Simulation pose data s of the simulated character at the next moment, generated from the environment model' t+1 The root node orientation and root node velocity, M d For the policy model, h is a constant representing the input orientation d. i and speed v i The total number, v t ,d t respectively with v i ,d i Correspondingly, these represent the user's input orientation and speed at a certain moment.

[0008] Based on any of the above embodiments, in this embodiment, the historical control signal g further includes the posture category j of the simulated character. Correspondingly, the training process of the strategy model includes: determining a target loss function based on a first preset loss function, a second preset loss function, and a third preset loss function; wherein, the third preset loss function is used to distinguish the posture category of the historical simulated posture data, and the third preset loss function is: Part One Used for classifying motion capture target pose data, Part 2 (1+D(s) t ,s t+1 ,j)) 2 Used for classifying simulated attitude data; where m j The preset reference action is used; a strategy model for the simulation role is determined based on the objective loss function and a third preset loss function, using adversarial learning; where the objective loss function is: Where Cross is the cross-entropy function, e j Let x be the probability distribution where only x = j is 1.

[0009] Based on any of the above embodiments, in this embodiment, the joint torque generation model of the simulated character further includes a motion encoding model and a motion decoding model. Both the motion encoding model and the motion decoding model are based on the historical simulation posture data and motion capture target posture data s corresponding to the simulated character. t+1 Confirmed, the motion capture target pose data s t+1 The simulation pose data is used to represent the pose data of the simulated character at the next moment corresponding to each historical moment; the historical simulation pose data corresponding to the simulated character includes: the simulation pose data of the simulated character at the previous historical moment; correspondingly, the training process of the motion encoding model and the motion decoding model includes: acquiring the simulation pose data of the simulated character at the previous historical moment and its corresponding motion capture target pose data s. t+1 Based on the simulation posture data of the simulation character at the previous historical moment and the corresponding number of motion capture target postures, the parameters of the pre-set motion encoding model and the pre-set motion decoding model are updated using the fourth preset loss function; wherein, the fourth preset loss function is: Wherein q(z|s t ,s t+1 p(z) is a variational distribution or posterior distribution. t |s t ) is the prior distribution, D KL The KL divergence between distributions; logp(s)t+1 |s t ,z t )∝-||s t+1 -s' t+1 ||, where s' t+1 The simulation posture data of the simulated character at the next moment is output by the environment model, representing the predicted simulation posture data at the next moment.

[0010] Based on any of the above embodiments, in this embodiment, the training process of the environment model includes: acquiring the torque of each joint of the simulated character from the motion decoding model; updating the parameters of the environment model based on the simulation posture data of the simulated character at the previous historical moment, the torque of each joint of the simulated character, and a preset loss function of the environment model; wherein, the loss function of the environment model is: Where b is a constant, representing the unfolded length of the autoregressive training, and s i With s t Corresponding to, s i For the simulated posture data of the simulated character at a certain historical moment, s' i The simulated posture data of the simulated character at a certain moment is output by the environment model, representing the predicted simulated posture data.

[0011] Based on any of the above embodiments, in this embodiment, the historical simulation posture data corresponding to the simulation character further includes: simulation posture data for the next historical moment corresponding to the simulation character. Correspondingly, after updating the parameters of the environment model, the method further includes: acquiring the simulation posture data for the next historical moment corresponding to the simulation character and its corresponding predicted simulation posture data; wherein, the simulation posture data for the next historical moment corresponding to the simulation character is determined by the torque of each joint of the simulation character determined by the motion decoding model after parameter update, and the simulation posture of the simulation character at the previous historical moment; updating the parameters of the motion encoding model and the motion decoding model based on the predicted simulation posture data for the next historical moment corresponding to the simulation character and its corresponding motion capture target posture data, and the fourth preset loss function; and regenerating the torque of each joint of the simulation character based on the updated motion encoding model and the motion decoding model; updating the parameters of the environment model based on the simulation posture data for the next historical moment corresponding to the simulation character and its corresponding predicted simulation posture data, the regenerated torque of each joint of the simulation character, and the preset loss function of the environment model.

[0012] Based on any of the above embodiments, in this embodiment, after updating the parameters of the motion encoding model and the motion decoding model, the method further includes: determining and saving the simulation posture data of the simulation character at the next historical moment based on the simulation posture data of the simulation character at the previous historical moment and the torque of each joint of the simulation character determined by the motion decoding model after updating the parameters.

[0013] This invention also provides a training method for a motion generation model of a simulated character. The motion generation model includes: a joint torque generation model of the simulated character and an environment model; wherein the joint torque generation model of the simulated character is obtained through joint training with the environment model; the environment model is used to generate predicted simulated posture data based on the torque of each joint of the simulated character, and the predicted simulated posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model; the method includes: acquiring historical simulated posture data corresponding to the simulated character and its corresponding motion capture target posture data, as well as historical control signals; the motion capture target posture data is used to represent The simulation character's posture data for the next moment corresponding to each historical moment is predicted. Based on the historical simulation posture data corresponding to the simulation character and its corresponding motion capture target posture data, the predicted simulation posture data, the fourth preset loss function, and the preset environment model's loss function, the motion encoding model, motion decoding model, and environment model in the simulation character's joint torque generation model are iteratively updated. Based on the historical simulation posture data corresponding to the simulation character, the predicted simulation posture data, the historical control signals, the updated motion encoding model, motion decoding model, and environment model, the strategy model in the simulation character's joint torque generation model is determined according to the target loss function.

[0014] This invention also provides a motion generation device for a simulated character, the device comprising: a first acquisition module, configured to acquire simulated posture data of the simulated character at the current moment and a current control signal for controlling the simulated character; and a generation module, configured to generate target simulated posture data based on the simulated posture data of the simulated character at the current moment, the current control signal, a pre-trained joint torque generation model of the simulated character, and a simulation environment; the target simulated posture data being the simulated posture data of the simulated character at the next moment corresponding to the current control signal; wherein, the joint torque generation model of the simulated character is used to determine the torque of each joint of the simulated character; the joint torque generation model of the simulated character is obtained through joint training with an environment model; the environment model is used to generate predicted simulated posture data based on the torque of each joint of the simulated character, and the predicted simulated posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model.

[0015] This invention also provides a training device for a motion generation model of a simulated character. The motion generation model includes: a joint torque generation model of the simulated character and an environment model; wherein the joint torque generation model of the simulated character is obtained through joint training with the environment model; the environment model is used to generate predicted simulated posture data based on the torque of each joint of the simulated character, and the predicted simulated posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model; the device includes: a second acquisition module, used to acquire historical simulated posture data corresponding to the simulated character and its corresponding motion capture target posture data, as well as historical control signals; the motion capture target posture data is used to represent the expected... The simulation character's posture data for the next moment corresponding to each historical moment; the update module, used to iteratively update the motion encoding model, motion decoding model, and environment model in the joint torque generation model of the simulation character based on the historical simulation posture data corresponding to the simulation character, its corresponding motion capture target posture data, predicted simulation posture data, a fourth preset loss function, and the loss function of the preset environment model; the determination module, used to determine the strategy model in the joint torque generation model of the simulation character based on the historical simulation posture data corresponding to the simulation character, the predicted simulation posture data, the historical control signals, the updated motion encoding model and motion decoding model, and the environment model, according to the target loss function.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method for generating the action of the simulated character and training its model as described above.

[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for generating the action of the simulated character and training its model as described above.

[0018] The present invention provides a method and apparatus for generating motion of a simulated character and training its model. By using an environment model instead of the original simulator, and jointly training it with the joint torque generation model of the simulated character, the environment model can generate predicted simulated posture data based on the torque of each joint of the simulated character. The generated predicted simulated posture data is then used in the autoregressive training of the joint torque generation model and the environment model. This allows the joint torque generation model of the simulated character to directly learn the relationship between the posture and motion of the simulated character and the torque of each joint. This enables end-to-end training of the joint torque generation model of the simulated character, thereby improving the training speed of the motion generation process and the quality of the generated simulated character's motion, making the generated simulated character's motion more flexible. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 This is one of the flowcharts illustrating the motion generation method for simulated characters provided by this invention;

[0021] Figure 2 This is a flowchart illustrating the training method for the strategy model provided by the present invention;

[0022] Figure 3 This is a flowchart illustrating the training method for the action encoding model and action decoding model provided by the present invention;

[0023] Figure 4 This is one of the flowcharts illustrating the training method for the environment model provided by this invention;

[0024] Figure 5 This is the second flowchart illustrating the motion generation method for simulated characters provided by this invention;

[0025] Figure 6 This is one of the flowcharts illustrating the training method for the motion generation model of the simulated character provided by the present invention;

[0026] Figure 7 This is the second flowchart illustrating the environmental model training method provided by the present invention;

[0027] Figure 8 This is the third flowchart illustrating the environmental model training method provided by the present invention;

[0028] Figure 9 This is a schematic block diagram of the action decoding model provided by the present invention;

[0029] Figure 10 This is the third flowchart illustrating the motion generation method for simulated characters provided by this invention;

[0030] Figure 11 This is a schematic diagram of the motion generation device for simulated characters provided by the present invention;

[0031] Figure 12 This is a schematic diagram of the structure of the training device for the motion generation model of the simulated character provided by the present invention;

[0032] Figure 13This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0034] To facilitate understanding, we will first explain the relevant content of the existing technology.

[0035] There are two main methods for generating motion in existing technologies: multi-stage methods and sampling estimation methods.

[0036] Multi-stage methods typically use reinforcement learning and other techniques based on motion capture data to calculate the torques that enable the simulated character to perform the same actions as the motion capture. Then, the generative model learns the strategy to generate torques with the same distribution as in the first step. Therefore, multi-stage methods are used to enable the simulated character to generate torques that match the motion capture based on the generative model.

[0037] Sampling estimation methods use generative models for action generation, but they rely on a large number of samples to estimate the direction of policy improvement. Typically, the process involves three steps: First, the simulator is controlled by the current policy to track motion capture data and perform multiple steps. Second, the simulation data collected in the first step is used to estimate the direction of policy improvement and refine the policy. Third, steps one and two are repeated.

[0038] Based on the above description, it can be understood that both approaches encode action experience into a latent space. The policy of the downstream task (e.g., handle control) will use vectors in this latent space as its output. Existing technologies learn the policy of downstream tasks by reusing the sampling estimation method described above. Therefore, existing technologies have the following drawbacks: multi-stage methods must spend a lot of time pre-compiling torques. Furthermore, since the learned torque is generated, it has no independent understanding of the action; the generated action must be highly similar to the given dataset, otherwise it will deviate from human behavior patterns, and the action quality will drop significantly. Sampling estimation methods estimate the improvement direction by collecting a large amount of simulation data, which takes a lot of time, and the estimated improvement direction is not accurate, resulting in slow training speed and slow convergence on large-scale datasets. Downstream tasks also suffer from slow training because they need to repeat the sampling estimation.

[0039] Therefore, in order to solve the problems of slow training speed and stiff generated simulated actions in the above-mentioned simulation character action generation process, the present invention proposes a method for generating simulated character actions and training its model.

[0040] It's understandable that, compared to other generative tasks, the characteristic of simulated character motion generation is that the character needs to move within a simulated environment. A corresponding example is kinematics-based character motion generation, which doesn't utilize a simulated environment. Kinematics-based character motion generation can be simply described as being given the character's current posture (or posture over a past period), and using neural network technology to directly generate the character's posture at the next moment, thus extending into character motion. In this process, since the character's posture at the next moment is directly generated by the neural network, a loss function is directly defined for the posture, and the gradient after differentiation can be backpropagated through the neural network. In contrast, simulated characters are subject to physical constraints, requiring the neural network to generate the torque applied to each joint, which is then passed to the simulator, and the simulation result becomes the posture at the next moment. However, the non-differentiability of the simulator leads to the interruption of training gradient propagation. This characteristic results in commonly used generative models for simulated character motion generation tasks, such as variational autoencoders and generative adversarial networks, being directly applicable to kinematic character motion generation but not to simulated character motion generation. Therefore, how to effectively generate simulated character actions and efficiently apply them to downstream tasks is the technical problem that this invention aims to solve.

[0041] The following is combined Figures 1-13 The present invention describes the method and apparatus for generating the motion of a simulated character and training its model.

[0042] Figure 1 This is one of the flowcharts illustrating the motion generation method for the simulated character provided by the present invention. For example... Figure 1 As shown, the motion generation method 100 for the simulated character may include steps 110 to 120, which are described in detail below. It should be understood that the method 100 can be executed by a motion generation device for the simulated character, and the method 100 includes:

[0043] Step 110: Obtain the simulation posture data of the simulation character at the current moment and the current control signal used to control the simulation character.

[0044] The current control signal used to control the simulated character is the control signal used to control the simulated character at the current moment, and is used to control the simulated character to perform certain actions according to the user's needs. For example, it can be a control signal from a control handle, used to control the simulated character to perform actions such as running, jumping, or hopping on one leg.

[0045] Step 120: Generate target simulation posture data based on the simulation posture data of the simulation character at the current moment, the current control signal, the pre-trained joint torque generation model of the simulation character, and the simulation environment.

[0046] The target simulation attitude data refers to the simulation attitude data of the simulated character at the next moment corresponding to the current control signal. This target simulation attitude data is generated in the simulation environment.

[0047] The joint torque generation model of the simulated character is used to determine the torque of each joint of the simulated character; the joint torque generation model of the simulated character is obtained through joint training with the environment model; the environment model is used to generate predicted simulation posture data based on the torque of each joint of the simulated character, and the predicted simulation posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model.

[0048] The historical simulation posture data corresponding to the simulation character refers to the simulation posture data corresponding to the simulation character over a period of time in the past, and the historical simulation posture data corresponding to the simulation character includes the simulation posture data at the initial moment corresponding to the simulation character.

[0049] It is understandable that, apart from the initial pose data, the historical pose data of the simulated character can be determined based on the initial pose data of the simulated character, the pose data of the motion capture target, the motion encoding and decoding models in the joint torque generation model of the simulated character, and the environment model. For details, please refer to... Figure 7 Related description. The pose data of the motion capture target can be obtained through motion capture technology.

[0050] Understandably, in practice, the torque of each joint of a simulated character can be calculated based on a neural network. The simulator is then trained using a large amount of collected simulation data to obtain the simulated character's posture for the next moment. However, due to the non-differentiable nature of the simulator, gradient descent training is not feasible, resulting in slow training speed and stiff, unnatural movements in the generated simulated characters. Therefore, this invention replaces the original simulator with an environment model, which is jointly trained with the simulated character's joint torque generation model. This allows the environment model to generate predicted simulated posture data (the simulated character's posture for the next moment) based on the torque of each joint. This predicted posture data is then used in the autoregressive training of both the simulated character's joint torque generation model and the environment model. This allows the simulated character's joint torque generation model to directly learn the relationship between the simulated character's posture and the torque of each joint. This end-to-end training of the simulated character's joint torque generation model improves the training speed of the motion generation process and enhances the quality of the generated simulated character's movements, resulting in more flexible movements.

[0051] Based on the above embodiments, in this embodiment, the joint torque generation model of the simulated character includes a strategy model. The strategy model is determined based on the historical simulation posture data and historical control signals corresponding to the simulated character. The strategy model is used to determine the offset in the encoding space, and the historical control signal g includes the orientation d input by the user. i and speed v i Accordingly, such as Figure 2 As shown, the training process of the policy model includes the following steps:

[0052] Step 210: Obtain the historical simulation posture data s corresponding to the simulation character. t And the historical control signal g.

[0053] Step 220: Based on the historical simulation posture data and historical control signals corresponding to the simulation character, determine the strategy model of the simulation character according to the first preset loss function and the second preset loss function; wherein, the first preset loss function and the second preset loss function are used together to control the direction and speed of the simulation character's movement.

[0054] The first preset loss function is: The second preset loss function is: l r =|M d (s t ,v t ,d t )|;wherein, d' i v i 'Simulation pose data s of the simulated character at the next moment, generated from the environment model't+1 The root node orientation and root node velocity, M d For the policy model, h is a constant representing the input orientation d. i and speed v i The total number, v t ,d t respectively with v i ,d i Correspondingly, these represent the user's input orientation and speed at a certain moment.

[0055] It is understandable that, in order to facilitate the control of the simulated character to make actions according to a certain orientation and speed, the above-mentioned first preset loss function and second preset loss function are defined.

[0056] Based on the above embodiments, in this embodiment, the historical control signal g further includes the posture category j of the simulated character. The posture category j is used to define the action posture type corresponding to the simulated character, that is, g = (d, v, j). Accordingly, the training process of the strategy model includes:

[0057] The target loss function is determined based on the first preset loss function, the second preset loss function, and the third preset loss function; wherein, the third preset loss function is used to distinguish the attitude categories of historical simulation attitude data, and the third preset loss function is: Part One Used for classifying motion capture target pose data, Part 2 (1+D(s) t ,s t+1 ,j)) 2 Used for classifying simulated attitude data; where m j This is a preset reference action; among which, Both D and are discriminators, such as multilayer perceptrons. A strategy model for determining the simulation role is based on the target loss function and a third preset loss function, using adversarial learning.

[0058] The objective loss function is: Where Cross is the cross-entropy function, e j Let x be the probability distribution where only x = j is 1. Where c(s) t ,s t+1 ) is a classifier whose input is action s. t ,s t+1 The output is the probability distribution of each action, and c can be a multilayer perceptron.

[0059] Based on the above embodiments, in this embodiment, the joint torque generation model of the simulated character further includes a motion encoding model and a motion decoding model. Both the motion encoding model and the motion decoding model are based on the historical simulation posture data and motion capture target posture data s corresponding to the simulated character. t+1 Confirmed, the motion capture target pose data s t+1 This is used to represent the pose data of the simulated character at the next predicted historical moment; the historical pose data corresponding to the simulated character includes: the pose data of the simulated character at the previous historical moment; correspondingly, such as Figure 3 As shown, the training process of the action encoding model and the action decoding model includes the following steps:

[0060] Step 310: Obtain the simulation posture data of the simulation character at the previous historical moment and its corresponding motion capture target posture data.

[0061] It is understandable that during the first training session, the simulation posture data of the simulated character at the previous historical moment can correspond to the simulation posture data of the simulated character at the initial moment, combined with... Figure 4 As can be seen from the steps, since the action encoding model, the action decoding model, and the environment model are obtained through multiple iterations of training, as the number of training iterations increases, the simulation posture data of the previous historical moment corresponding to the simulation character can be the simulation posture data of other moments, and is relative to the simulation posture data of the next historical moment corresponding to the simulation character.

[0062] Step 320: Based on the simulation posture data of the simulation character at the previous historical moment and the number of motion capture target postures corresponding to it, and the fourth preset loss function, update the parameters of the preset motion encoding model and the preset motion decoding model.

[0063] The fourth preset loss function is: Wherein q(z|s t ,s t+1 p(z) is a variational distribution or posterior distribution. t |s t ) is the prior distribution, D KL The KL divergence between distributions; log p(s) t+1 |s t ,z t )∝-||s t+1 -s' t+1 ||, where s' t+1 The simulation posture data of the simulated character at the next moment is output by the environment model, representing the predicted simulation posture data at the next moment.

[0064] The pre-defined motion encoding model can be a multilayer perceptron (MLP), specifically, for example, a 3-layer MLP with 256 hidden units per layer. The pre-defined motion decoding model can consist of multiple MLPs, such as a Mixture of Experts (MoE) structure, with each MLP having, for example, 4 layers with 512 hidden units per layer. Its input consists of the latent space vector z output by the motion encoding model and the simulated pose data of the simulated character at historical moments. For details, please refer to [reference needed]. Figure 9 Related descriptions.

[0065] It is understandable that the encoding process corresponding to the action coding model can further include: prior encoding, posterior encoding, and encoding resampling. For details, please refer to... Figure 5 This will not be discussed in detail here.

[0066] It's also understandable that for the aforementioned fourth preset loss function, it's typically necessary to define the prior distribution and variational distribution of the encoding. The prior distribution is usually defined as a standard normal distribution, and during testing, simply sampling on the standard normal distribution is sufficient to generate new actions. However, for simulated animation, such a prior distribution, lacking closed-loop information (e.g., simulated pose data from historical moments), can lead to the character being prone to falling and the generated actions being discontinuous. Therefore, a prior distribution dependent on the simulated character's historical pose data is proposed, i.e., p(z|s t The posterior distribution is q(z|s). t ,s t+1 ). Among them, p(z) t |s t ) and p(s t+1 |s t ,z t Both are implemented using a multilayer perceptron. Specifically, two 3-layer perceptrons can be used, denoted by M1 and M2 respectively, and assuming that the distributions are both normally distributed, then: p(z|s t )~N(M1(s t ),σ 2 ), q(z|s t ,s t+1 )~N(M1(s t )+M2(s t ,s t+1 ),σ 2 ), where σ can be 0.3.

[0067] As mentioned earlier, the motion encoding model, motion decoding model, and environment model are obtained through autoregressive training. Therefore, after the initial update of the environment model parameters, since the simulation posture data of the next historical moment corresponding to the simulated character is unknown, the next step of autoregressive training cannot be performed. Therefore, after the initial update of the parameters of the motion encoding model and motion decoding model, the torque of each joint of the simulated character obtained based on the motion encoding model and motion decoding model, as well as the simulation posture data of the simulated character corresponding to the previous historical moment, can be input into the simulator. The simulator will then generate the actual simulation posture data of the next historical moment corresponding to the simulated character, thereby facilitating the autoregressive training to obtain the environment model based on the actual simulation posture data of the next historical moment corresponding to the simulated character.

[0068] As mentioned earlier, the motion encoding model, motion decoding model, and environment model are obtained through multiple autoregressive trainings. Therefore, after the initial update of the environment model parameters, since the simulation posture data of the next historical moment corresponding to the simulated character is unknown, the next autoregressive training cannot be performed. Therefore, after the initial update of the parameters of the motion encoding model and motion decoding model, the torque of each joint of the simulated character obtained based on the motion encoding model and motion decoding model, the simulation posture data of the simulated character corresponding to the previous historical moment, and the initially updated environment model can be used to determine the simulation posture data of the next historical moment corresponding to the simulated character. This facilitates the autoregressive training of the motion encoding model, motion decoding model, and environment model based on the simulation posture data of the next historical moment corresponding to the simulated character.

[0069] In addition, it should be noted that the simulation pose data s' of the simulated character at the next moment output by the environment model t+1 It is the simulation posture data (predicted simulation posture data) of the simulation character at the next moment output by the environment model, which is different from the simulation posture data of the simulation character at the next historical moment generated by the simulator.

[0070] Therefore, in one embodiment, in order to obtain more simulation posture data of the simulation character, step 330 may be included after step 320.

[0071] Step 330: Based on the simulation posture data of the simulation character at the previous historical moment and the torque of each joint of the simulation character determined by the motion decoding model after parameter update, determine and save the simulation posture data of the simulation character at the next historical moment.

[0072] It is understandable that this step can be performed by the simulator to generate simulation posture data for the next historical moment corresponding to the simulation character, so as to facilitate the subsequent training of the environment model. This enables the environment model to generate differentiable simulation posture data for the next historical moment, and then iterative training can be performed multiple times based on the differentiable simulation posture data for the next historical moment, together with the aforementioned motion encoding model and motion decoding model, to achieve end-to-end training and accelerate the training speed.

[0073] Based on the above embodiments, in this embodiment, the training process of the environment model is as follows: Figure 4 As shown, it includes:

[0074] Step 410: Obtain the torque of each joint of the simulated character from the motion decoding model.

[0075] Step 420: Update the parameters of the environment model based on the simulation posture data of the simulation character at the previous historical moment, the torque of each joint of the simulation character, and the loss function of the preset environment model.

[0076] The loss function for the environment model is: Where b is a constant, representing the unfolded length of the autoregressive training, and s i With s t Corresponding to, s i For the simulated posture data of the simulated character at a certain historical moment, s' i The simulated posture data of the simulated character at a certain moment is output by the environment model, representing the predicted simulated posture data.

[0077] Among them, s i Specifically, it can be defined as the quaternions of position, rotation, velocity, and angular velocity of each rigid body part of the simulated character. Before being input into the network, all s i All values ​​are processed relative to the root node of the simulated character, i.e., position minus root node position, rotation minus root node rotation, velocity minus root node velocity minus rotation, and angular velocity minus root node rotation. Additional information is added to handle collisions with the ground. The absolute height of each rigid body part and the upward orientation of the root node are also added.

[0078] It is understandable that after the initial update of the environment model parameters, the torques of each joint of the simulated character obtained based on the motion encoding model and motion decoding model, the simulated posture data of the simulated character at the previous historical moment, and the initially updated environment model can be used to determine the predicted simulated posture data (predicted simulated posture data) of the simulated character at the next historical moment. This facilitates the autoregressive training of the motion encoding model, motion decoding model, and environment model based on the predicted simulated posture data of the simulated character at the next historical moment. Furthermore, since s'i The simulated pose data of the simulated character at a certain moment is output by the environment model, therefore s' i Differentiability ensures gradient propagation and improves the training framework to guarantee training stability.

[0079] As mentioned earlier, the motion encoding model, motion decoding model, and environment model are obtained through autoregressive iterative training. Therefore, the motion encoding model, motion decoding model, and environment model can be further trained based on the simulation posture data of the simulation character at the next historical moment. It can be understood that continuing to train the motion encoding model, motion decoding model, and environment model based on the simulation posture data of the simulation character at the next historical moment allows the environment model to learn more historical simulation posture data and information about the torques of each joint of the simulation character, which is beneficial for generating more flexible simulation character movements.

[0080] In other words, after updating the parameters of the environment model in step 420, the action encoding model, action decoding model and environment model can continue to be trained. Therefore, the method also includes steps 430 to 450, which are described in detail below.

[0081] Step 430: Obtain the simulation posture data of the simulation role at the next historical moment and its corresponding predicted simulation posture data.

[0082] The simulation posture data of the simulation character at the next historical moment is determined by the torque of each joint of the simulation character determined by the motion decoding model after parameter update, and the simulation posture of the simulation character at the previous historical moment.

[0083] It is understood that the process of determining the simulation posture data of the simulation character at the next historical moment in this step can be referred to the relevant description in step 330, and will not be repeated here.

[0084] It can also be understood that the simulated attitude data for the next historical moment is generated by the simulator, while the predicted simulated attitude data for the next historical moment is generated by the environment model. Specifically, after updating the parameters of the environment model in step 420 above, the predicted simulated attitude data for the next historical moment can be generated based on the environment model with updated parameters.

[0085] Step 440: Based on the predicted simulation posture data corresponding to the next historical moment of the simulation character and its corresponding motion capture target posture data, as well as the fourth preset loss function, update the parameters of the motion encoding model and the motion decoding model; and regenerate the torque of each joint of the simulation character based on the updated motion encoding model and the motion decoding model.

[0086] Step 450: Based on the simulation posture data of the simulation character at the next historical moment and its corresponding predicted simulation posture data, the torque of each joint of the regenerated simulation character, and the loss function of the preset environment model, update the parameters of the environment model.

[0087] It is understood that steps 410 to 450 above can be repeated multiple times with steps 310 to 330 above until the preset number of iterations is reached.

[0088] Figure 5 Another schematic flowchart of the simulation character motion generation method provided by the present invention.

[0089] Typically, given a dataset of motion capture target poses, the goal is to transfer it to a simulation environment and encode motion techniques into a latent space so that it can be reused for specific simulated character motion generation tasks. A motion capture target pose dataset generally consists of a series of human poses s, which can be represented by τ = (s0, s1, ... s2). i The expression denoted by ) represents the i-th frame of data. The interval between frames is fixed and can be 1 / 120 of a second. A variational autoencoder can be used to model the pose transition, assuming the current pose s is known. i We can assume possible future postures s i+1 Determined by an unknown latent variable z, we can be concerned with its conditional distribution p(s). i+1 |z,s i ).

[0090] In order to determine the pose s based on existing target pose data i+1 This could be a new action. Typically, it's assumed that the prior distribution of z is an easily sampled distribution, and the goal is to generate a reasonable pose s by sampling on this distribution p(z). i+1 According to the theory of variational autoencoders, it is possible to optimize the right-hand side of the formula below.

[0091] Wherein q(z|s t ,s t+1 If the distribution is a variational distribution or a posterior distribution, a neural network is generally used for fitting. KL Let be the KL divergence between distributions. Typically, if the simulation character generates state s'... t+1 Then logp(s) can be estimated. t+1 |s t ,z t )∝-||s t+1 -s' t+1 ||.

[0092] However, due to the s' in simulated character animation t+1 Generated by the simulator, it cannot propagate gradients, therefore s' t+1 The gradient of the loss function defined above will be interrupted at this point. This invention uses a neural network to represent the environment model to fit the simulator to ensure gradient propagation and improves the training framework to ensure training stability. Specifically, as shown in Figure 5, the simulation posture data s of the simulated character at a certain historical moment... i Prior encoding is performed on the motion capture target pose data s corresponding to the simulated pose data at a certain historical moment. i+1 By performing posterior encoding and then resampling the encoded information obtained from both prior and posterior encoding, the variable z in the latent space can be obtained. Then, the variable z is correlated with the simulation posture data s of the simulation role at a specific historical moment. i Input the motion decoding model, output the torque 'a' of each joint of the simulated character, and combine the torque 'a' of each joint with the simulated posture data 's' of the simulated character at a specific historical moment. i The input is given to the environment model, which then outputs the simulation attitude data s' for the next moment corresponding to a given historical moment. i+1 It's understandable that the s' i+1 It can be used as input for the next autoregressive training of the environment model to eliminate accumulated errors.

[0093] This invention also provides a training method for a motion generation model of a simulated character, which can be executed by a training device for the motion generation model of a simulated character. The motion generation model includes: a joint torque generation model of the simulated character and an environment model. The joint torque generation model of the simulated character is used to determine the torque of each joint of the simulated character; the joint torque generation model of the simulated character is obtained through joint training with the environment model; the environment model is used to generate predicted simulated posture data based on the torque of each joint of the simulated character, and the predicted simulated posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model. Specifically, as... Figure 6 As shown, the training method for the motion generation model of the simulated character includes the following steps:

[0094] Step 610: Obtain the historical simulation posture data corresponding to the simulation character and its corresponding motion capture target posture data, as well as the historical control signals.

[0095] Among them, the motion capture target pose data is used to represent the pose data of the simulated character at the next moment corresponding to each historical moment.

[0096] Step 620: Based on the historical simulation posture data corresponding to the simulation character and its corresponding motion capture target posture data, predicted simulation posture data, fourth preset loss function, and preset environment model loss function, iteratively update the motion encoding model, motion decoding model, and environment model in the joint torque generation model of the simulation character.

[0097] As mentioned earlier, the historical simulation posture data for the simulated character includes the simulation posture data from the previous historical moment and the simulation posture from the next historical moment. Therefore, the corresponding motion encoding model, motion decoding model, and environment model can be obtained through multiple iterative training iterations. Furthermore, the predicted simulation posture data generated by the environment model will be used during the iterative training process. See the previous text for details. Figure 3 and Figure 4 The relevant descriptions will not be repeated here.

[0098] Step 630: Based on the historical simulation posture data corresponding to the simulation character, the predicted simulation posture data, the historical control signals, the updated motion encoding model and motion decoding model, and the environment model, determine the strategy model in the joint torque generation model of the simulation character according to the target loss function.

[0099] Combination Figure 3 and Figure 4 Based on the relevant descriptions, pre-trained motion encoding models, motion decoding models, and environment models can be obtained. Therefore, the generation of simulated character actions can be completed based on the pre-trained motion encoding models, motion decoding models, and environment models. However, in order to adapt to specific action generation tasks, the corresponding policy models can be further trained based on the historical simulation posture data, predicted simulation posture data, and historical control signals corresponding to the simulated character, so that the simulated character can generate corresponding actions according to the corresponding policies.

[0100] It is understandable that since the prior model in the action encoding model, as well as the action decoding model and the environment model, are all pre-trained, they can be directly used to assist in training the policy model in the simulation action generation task. Compared with existing technologies, there is no need to perform a large amount of sampling to update the policy, which greatly accelerates the training process.

[0101] Specifically, the process of step 630 can be referred to the relevant description above, and will not be repeated here.

[0102] Figure 7 This is another schematic diagram illustrating the environment model training method provided by the present invention. To obtain an environment model that better matches the task of generating simulated character actions, an alternating training approach can be used, where the environment model and the generation model are trained concurrently. To train the environment model, it is first necessary to obtain real data from the simulator. Simulator data can be generated before each training session using the following method:

[0103] First, such as Figure 7 As shown, it will be as follows Figure 5 The environment model shown was replaced with a simulator simulation, and the resulting (s) i ,a i ,s i+1 Stored in the buffer.

[0104] Secondly, such as Figure 8 As shown, to eliminate accumulated error, similar to training a variational autoencoder, the data (s) stored in the buffer is used. i ,a i ,s i+1 An autoregressive training environment model is used. Specifically, a continuous sequence (s0, a0, s1, a1…) is taken from the buffer, and s0 and a… are used as the training environment model. i Generate predictions s'1, s'2, ... from the environmental model.

[0105] Figure 9 This is a schematic diagram of the action decoding model provided by the present invention. Figure 9 As shown, a network is first used to map the input to the weights w of each MLP. i Then, the input is injected into each MLP, and the results are weighted and averaged using weights w to serve as the input for the next layer. To emphasize the role of z, each layer can also additionally concatenate z with the output of the previous layer.

[0106] In one embodiment, during the specific training process, the environment model can be optimized 8 times, followed by the motion encoding model and motion decoding model (i.e., variational autoencoder) optimized 8 times. The buffer size is 50,000 pairs, which means storing 50,000 pairs of historical simulation posture data corresponding to the simulated character and their corresponding motion capture target posture data. 2048 pairs are collected before each training session and 512 pairs are sampled from them for batch training. Furthermore, the RAdam trainer is used for training, with a learning rate of 2e-3 for the environment model and a learning rate of 1e-5 for the variational autoencoder (motion encoding model and motion decoding model).

[0107] Figure 10 This is a schematic diagram illustrating the simulation character action generation process provided by the present invention. It can be understood that after training the variational autoencoder, an encoding space is obtained. Sampling in this space can generate new simulation character actions, which can then be used as the action space for the downstream task (the simulation character action generation task) to control the simulation character to complete the downstream task. Typically, the utilization of the encoding space often employs reinforcement learning techniques, involving extensive sampling. However, such sampling efficiency is low. Therefore, this invention uses a pre-trained environment model to directly train the downstream task, greatly accelerating the training process.

[0108] like Figure 10 As shown in the figure, the prior encoder M1, the action decoding model, and the environment model are pre-trained. What needs to be introduced and trained additionally is the "downstream policy" (corresponding to the policy model mentioned earlier), which can be denoted as M. d Its input is the current state s. t The input to the downstream task (corresponding to the user input control signal g mentioned earlier) is the input vector, and the output is the offset z in the encoding space. The role of the downstream policy is to select a suitable vector z in the encoding space so that it can complete the downstream task of the input. Here, z is the vector obtained after encoding resampling (the same as the posterior distribution of the first stage described above). The training process of other parts is the same as described above. Figure 2 The content is the same, so it will not be repeated here. That is to say, all parameters of the variational autoencoder and the environment model in the downstream task policy remain unchanged, and the downstream policy is trained with RAdam. The downstream policy can be a 3-layer MLP with 256 hidden units.

[0109] The following describes the motion generation device for the simulated character and the training device for its model provided by the present invention. The motion generation device for the simulated character described below can be referred to in correspondence with the motion generation method for the simulated character described above. The training device for the motion generation model of the simulated character described below can be referred to in correspondence with the motion generation model for the simulated character described above.

[0110] Figure 11 This is a schematic diagram of the motion generation device for simulated characters provided by the present invention, such as... Figure 11 As shown, the motion generation device for simulated characters provided in this embodiment of the invention includes:

[0111] The first acquisition module 1110 is used to acquire the simulation posture data of the simulation character at the current moment and the current control signal used to control the simulation character.

[0112] The generation module 1120 is used to generate target simulation posture data based on the simulation posture data of the simulation character at the current moment, the current control signal, the pre-trained joint torque generation model of the simulation character, and the simulation environment; the target simulation posture data is the simulation posture data of the simulation character at the next moment corresponding to the current control signal.

[0113] The joint torque generation model of the simulated character is used to determine the torque of each joint of the simulated character; the joint torque generation model of the simulated character is obtained through joint training with the environment model; the environment model is used to generate predicted simulation posture data based on the torque of each joint of the simulated character, and the predicted simulation posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model.

[0114] The motion generation device for simulated characters provided by this invention uses an environment model instead of the original simulator and performs joint training with the joint torque generation model of the simulated character. This enables the environment model to generate predicted simulated posture data based on the torque of each joint of the simulated character. The generated predicted simulated posture data is then used in the autoregressive training of the joint torque generation model and the environment model. This allows the joint torque generation model of the simulated character to directly learn the relationship between the posture and motion of the simulated character and the torque of each joint. This enables end-to-end training of the joint torque generation model of the simulated character, thereby improving the training speed of the motion generation process and the quality of the generated simulated character's motion, making the generated simulated character's motion more flexible.

[0115] Based on any of the above embodiments, in this embodiment, the joint torque generation model of the simulated character includes a strategy model. The strategy model is determined based on the historical simulation posture data and historical control signals corresponding to the simulated character. The strategy model is used to determine the offset in the encoding space, and the historical control signal g includes the orientation d input by the user. i and speed v i Accordingly, the generation module 1120 includes:

[0116] The first acquisition unit is used to acquire the historical simulation posture data s corresponding to the simulation character. t And the historical control signal g;

[0117] The first determining unit is used to determine the strategy model of the simulation character based on the historical simulation posture data and historical control signals corresponding to the simulation character, according to a first preset loss function and a second preset loss function; wherein, the first preset loss function and the second preset loss function are used together to control the direction and speed of the simulation character's movement; wherein, the first preset loss function is: The second preset loss function is: l r =|M d (s t ,v t ,d t )|;wherein, d' i v i 'Simulation pose data s of the simulated character at the next moment, generated from the environment model't+1 The root node orientation and root node velocity, M d For the policy model, h is a constant representing the input orientation d. i and speed v i The total number, v t ,d t respectively with v i ,d i Correspondingly, these represent the user's input orientation and speed at a certain moment.

[0118] Based on any of the above embodiments, in this embodiment, the historical control signal g further includes the posture category j of the simulated character, and correspondingly, the generation module 1120 includes:

[0119] The second determining unit is used to determine a target loss function based on a first preset loss function, a second preset loss function, and a third preset loss function; wherein, the third preset loss function is used to distinguish the attitude categories of historical simulation attitude data, and the third preset loss function is: Part One Used for classifying motion capture target pose data, Part 2 (1+D(s) t ,s t+1 ,j)) 2 Used for classifying simulated attitude data; where m j This is a preset reference action;

[0120] The third determining unit is used to determine the strategy model of the simulation role based on the objective loss function and the third preset loss function, using adversarial learning; wherein, the objective loss function is: Where Cross is the cross-entropy function, e j Let x be the probability distribution where only x = j is 1.

[0121] Based on any of the above embodiments, in this embodiment, the joint torque generation model of the simulated character further includes a motion encoding model and a motion decoding model. Both the motion encoding model and the motion decoding model are based on the historical simulation posture data and motion capture target posture data s corresponding to the simulated character. t+1 Confirmed, the motion capture target pose data s t+1 The generation module 1120 is used to represent the posture data of the simulated character at the next moment corresponding to each historical moment; the historical simulation posture data corresponding to the simulated character includes: the simulation posture data of the simulated character at the previous historical moment; correspondingly, the generation module 1120 includes:

[0122] The second acquisition unit is used to acquire the simulation posture data of the simulation character at the previous historical moment and its corresponding motion capture target posture data s.t+1 ;

[0123] The first update unit is used to update the parameters of the pre-set motion encoding model and the pre-set motion decoding model based on the simulation posture data of the simulation character at the previous historical moment and the number of motion capture target postures corresponding to it, as well as the fourth preset loss function; wherein, the fourth preset loss function is: Wherein q(z|s t ,s t+1 p(z) is a variational distribution or posterior distribution. t |s t ) is the prior distribution, D KL The KL divergence between distributions; logp(s) t+1 |s t ,z t )∝-||s t+1 -s' t+1 ||, where s' t+1 The simulation posture data of the simulated character at the next moment is output by the environment model, representing the predicted simulation posture data at the next moment.

[0124] Based on any of the above embodiments, in this embodiment, the generation module 1120 includes:

[0125] The third acquisition unit is used to acquire the torque of each joint of the simulated character from the motion decoding model;

[0126] The third update unit is used to update the parameters of the environment model based on the simulation posture data of the simulation character at the previous historical moment, the torque of each joint of the simulation character, and the loss function of the preset environment model; wherein, the loss function of the environment model is: Where b is a constant, representing the unfolded length of the autoregressive training, and s i With s t Corresponding to, s i For the simulated posture data of the simulated character at a certain historical moment, s' i The simulated posture data of the simulated character at a certain moment is output by the environment model, representing the predicted simulated posture data.

[0127] Based on any of the above embodiments, in this embodiment, the historical simulation posture data corresponding to the simulation character further includes: simulation posture data of the simulation character at the next historical moment; correspondingly, the device further includes:

[0128] The fourth acquisition unit is used to acquire the simulation posture data of the next historical moment corresponding to the simulation character and the corresponding predicted simulation posture data; wherein, the simulation posture data of the next historical moment corresponding to the simulation character is determined by the torque of each joint of the simulation character determined by the motion decoding model after updating parameters, and the simulation posture of the simulation character at the previous historical moment.

[0129] The fourth update unit is used to update the parameters of the motion encoding model and the motion decoding model based on the predicted simulation posture data corresponding to the next historical moment of the simulation character and its corresponding motion capture target posture data, as well as the fourth preset loss function; and to regenerate the torque of each joint of the simulation character based on the updated motion encoding model and the motion decoding model.

[0130] The fifth update unit is used to update the parameters of the environment model based on the simulation posture data of the simulation character at the next historical moment and its corresponding predicted simulation posture data, the torque of each joint of the regenerated simulation character, and the loss function of the preset environment model.

[0131] Based on any of the above embodiments, in this embodiment, the device further includes:

[0132] The saving module is used to determine and save the simulation posture data of the simulation character at the next historical moment based on the simulation posture data of the simulation character at the previous historical moment and the torque of each joint of the simulation character determined by the motion decoding model after updating the parameters.

[0133] The present invention also provides a training device for a motion generation model of a simulated character.

[0134] Figure 12 This is a schematic diagram of a training device for a motion generation model of a simulated character provided by the present invention. The motion generation model includes: a joint torque generation model of the simulated character and an environment model; wherein, the joint torque generation model of the simulated character is obtained through joint training with the environment model; the environment model is used to generate predicted simulated posture data based on the torque of each joint of the simulated character, and the predicted simulated posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model. Figure 12 As shown, the training device for the motion generation model of the simulated character provided in this embodiment of the invention includes:

[0135] The second acquisition module 1210 is used to acquire historical simulation posture data corresponding to the simulation character and its corresponding motion capture target posture data, as well as historical control signals; the motion capture target posture data is used to represent the posture data of the simulation character at the next moment corresponding to each historical moment.

[0136] The update module 1220 is used to iteratively update the motion encoding model, motion decoding model and environment model in the joint torque generation model of the simulated character based on the historical simulation posture data corresponding to the simulated character and its corresponding motion capture target posture data, predicted simulation posture data, fourth preset loss function and preset environment model loss function.

[0137] The determination module 1230 is used to determine the strategy model in the joint torque generation model of the simulation character based on the historical simulation posture data corresponding to the simulation character, the predicted simulation posture data, the historical control signals, the updated motion encoding model and motion decoding model, and the environment model, according to the target loss function.

[0138] The training device for the motion generation model of the simulated character provided by the present invention can directly train the policy model through the trained motion encoding model, motion decoding model and environment model. Compared with the prior art, it does not require a large amount of sampling to update the policy, which greatly accelerates the training process.

[0139] Figure 13 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 13As shown, the electronic device may include a processor 1310, a communication interface 1320, a memory 1330, and a communication bus 1340. The processor 1310, communication interface 1320, and memory 1330 communicate with each other via the communication bus 1340. The processor 1310 can call logic instructions from the memory 1330 to execute methods for generating the simulated character's actions and training its model. The method for generating the motion of a simulated character includes: acquiring the simulated posture data of the simulated character at the current moment and the current control signal used to control the simulated character; generating target simulated posture data based on the simulated posture data of the simulated character at the current moment, the current control signal, a pre-trained joint torque generation model of the simulated character, and a simulation environment; the target simulated posture data is the simulated posture data of the simulated character at the next moment corresponding to the current control signal; wherein, the joint torque generation model of the simulated character is used to determine the torque of each joint of the simulated character; the joint torque generation model of the simulated character is obtained through joint training with an environment model; the environment model is used to generate predicted simulated posture data based on the torque of each joint of the simulated character, and the predicted simulated posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model. The training method for the motion generation model of the simulated character includes: acquiring historical simulation posture data and corresponding motion capture target posture data of the simulated character, as well as historical control signals; the motion capture target posture data is used to represent the posture data of the simulated character at the next moment corresponding to each historical moment; based on the historical simulation posture data and corresponding motion capture target posture data of the simulated character, the predicted simulation posture data, the fourth preset loss function, and the loss function of the preset environment model, iteratively updating the motion encoding model, motion decoding model, and environment model in the joint torque generation model of the simulated character; based on the historical simulation posture data, the predicted simulation posture data, the historical control signals, the updated motion encoding model, motion decoding model, and environment model, determining the strategy model in the joint torque generation model of the simulated character according to the target loss function.

[0140] Furthermore, the logical instructions in the aforementioned memory 1330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0141] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the simulation character action generation and model training method provided by the present invention. The simulation character action generation method includes: acquiring simulation posture data of the simulation character at the current moment and a current control signal for controlling the simulation character; generating target simulation posture data based on the simulation posture data of the simulation character at the current moment, the current control signal, a pre-trained joint torque generation model of the simulation character, and a simulation environment; the target simulation posture data is the simulation posture data of the simulation character at the next moment corresponding to the current control signal; wherein the joint torque generation model of the simulation character is used to determine the torque of each joint of the simulation character; the joint torque generation model of the simulation character is obtained through joint training with an environment model; the environment model is used to generate predicted simulation posture data based on the torque of each joint of the simulation character, and the predicted simulation posture data is used for autoregressive training of the joint torque generation model of the simulation character and the environment model. The training method for the motion generation model of the simulated character includes: acquiring historical simulation posture data and corresponding motion capture target posture data of the simulated character, as well as historical control signals; the motion capture target posture data is used to represent the posture data of the simulated character at the next moment corresponding to each historical moment; based on the historical simulation posture data and corresponding motion capture target posture data of the simulated character, the predicted simulation posture data, the fourth preset loss function, and the loss function of the preset environment model, iteratively updating the motion encoding model, motion decoding model, and environment model in the joint torque generation model of the simulated character; based on the historical simulation posture data, the predicted simulation posture data, the historical control signals, the updated motion encoding model, motion decoding model, and environment model, determining the strategy model in the joint torque generation model of the simulated character according to the target loss function.

[0142] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for generating motion of a simulated character and training its model provided by the present invention. The method for generating motion of a simulated character includes: acquiring simulation posture data of the simulated character at the current moment and a current control signal for controlling the simulated character; generating target simulation posture data based on the simulation posture data of the simulated character at the current moment, the current control signal, a pre-trained joint torque generation model of the simulated character, and a simulation environment; the target simulation posture data is the simulation posture data of the simulated character at the next moment corresponding to the current control signal; wherein the joint torque generation model of the simulated character is used to determine the torque of each joint of the simulated character; the joint torque generation model of the simulated character is obtained through joint training with an environment model; the environment model is used to generate predicted simulation posture data based on the torque of each joint of the simulated character, and the predicted simulation posture data is used for autoregressive training of the joint torque generation model of the simulated character and the environment model. The training method for the motion generation model of the simulated character includes: acquiring historical simulation posture data and corresponding motion capture target posture data of the simulated character, as well as historical control signals; the motion capture target posture data is used to represent the posture data of the simulated character at the next moment corresponding to each historical moment; based on the historical simulation posture data and corresponding motion capture target posture data of the simulated character, the predicted simulation posture data, the fourth preset loss function, and the loss function of the preset environment model, iteratively updating the motion encoding model, motion decoding model, and environment model in the joint torque generation model of the simulated character; based on the historical simulation posture data, the predicted simulation posture data, the historical control signals, the updated motion encoding model, motion decoding model, and environment model, determining the strategy model in the joint torque generation model of the simulated character according to the target loss function.

[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0144] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0145] It is understood that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A motion generation method of an imitation character, characterized by, The method comprises: obtaining simulation posture data corresponding to a current time of the simulation character and a current control signal for controlling the simulation character; generating target simulation posture data based on the simulation posture data corresponding to the current time of the simulation character, the current control signal, a joint torque generation model of the pre-trained simulation character, and a simulation environment; the target simulation posture data is simulation posture data of the simulation character at a next time corresponding to the current control signal; wherein the joint torque generation model of the simulation character is used to determine the torque of each joint of the simulation character; the joint torque generation model of the simulation character is obtained by joint training with an environment model; the environment model is used to generate predicted simulation posture data based on the torque of each joint of the simulation character, and the predicted simulation posture data is used for autoregressive training of the joint torque generation model of the simulation character and the environment model; The joint torque generation model of the simulated character comprises a policy model, the policy model being determined based on historical simulation pose data corresponding to the simulated character and historical control signals, the policy model being used to determine an offset in an encoding space, the historical control signals comprising a user input orientation and a velocity , and accordingly, a training process of the policy model comprises: Acquiring historical simulation pose data corresponding to the simulation character and historical control signals ; determine a policy model of the simulation character according to a first preset loss function and a second preset loss function based on the historical simulation pose data and the historical control signal corresponding to the simulation character; wherein the first preset loss function and the second preset loss function are used together to control the direction and speed of the simulation character in advance; wherein the first preset loss function is: , and the second preset loss function is: ; wherein are respectively a root node orientation and a root node speed of simulation pose data of the simulation character generated by the environment model at a next moment , and is a policy model, h is a constant, and represents the total number of input orientations and speeds , and and correspond respectively, and represent the user input orientation and speed at a certain moment respectively.

2. The motion generating method of an imitation character according to Claim 1, characterized by, The historical control signal The posture category of the simulation role is also included Correspondingly, the training process of the strategy model comprises: The target loss function is determined according to the first preset loss function, the second preset loss function and the third preset loss function; wherein the third preset loss function is used to distinguish the pose category of the historical simulation pose data, and the third preset loss function is: wherein the first part is used for classifying the motion capture target pose data, and the second part is used for classifying the simulation pose data; wherein is a preset reference motion. The policy model of the simulation role is determined based on the target loss function and the third preset loss function based on the adversarial learning, wherein the target loss function is: where Cross is a cross-entropy function, is a probability distribution x with only 1.

3. The motion generating method of an imitation character according to Claim 1, characterized by, The joint torque generation model of the simulation character further comprises a motion encoding model and a motion decoding model, and the motion encoding model and the motion decoding model are both based on the historical simulation pose data and the motion capture target pose data corresponding to the simulation character It is determined that the motion capture target pose data for representing the pose data of the simulation character at the next moment corresponding to each historical moment the historical simulation posture data corresponding to the simulation character includes simulation posture data corresponding to a previous historical time of the simulation character; correspondingly, the training process of the action encoding model and the action decoding model comprises: Obtain the simulation pose data of the previous historical moment corresponding to the simulation character and the corresponding motion capture target pose data ; The parameters of the pre-set action encoding model and the pre-set action decoding model are updated based on simulation pose data of a previous historical moment corresponding to the simulation character and the corresponding motion capture target pose data, and a fourth preset loss function; wherein the fourth preset loss function is: ; wherein, is a variational distribution or a posterior distribution, is a prior distribution, is the KL divergence between the distributions; , wherein, is the simulation pose data of the simulation character at the next moment output by the environment model, and represents the simulation pose data predicted at the next moment.

4. The action generating method of an imitation character according to Claim 3, characterized by, the training process of the environment model comprises: obtaining the torque of each joint of the simulation character from the action decoding model; Based on the simulation posture data of the previous historical moment corresponding to the simulation character, the torque of each joint of the simulation character and the loss function of the preset environment model, the parameters of the environment model are updated; wherein the loss function of the environment model is: wherein b is a constant, representing the expansion length of the autoregressive training, wherein corresponding to , is the simulation posture data of a certain historical moment corresponding to the simulation character, is the simulation posture data of a certain moment of the simulation character output by the environment model, representing the predicted simulation posture data.

5. The motion generating method of an imitation character according to Claim 4, characterized by, the historical simulation posture data corresponding to the simulation character also includes simulation posture data corresponding to a next historical time of the simulation character, and correspondingly, after updating the parameters of the environment model, the method further comprises: obtaining simulation posture data corresponding to a next historical time of the simulation character and its corresponding predicted simulation posture data; wherein the simulation posture data corresponding to the next historical time of the simulation character is determined by the torque of each joint of the simulation character determined by the updated action decoding model and the simulation posture corresponding to the previous historical time of the simulation character; updating the parameters of the action encoding model and the action decoding model based on the predicted simulation posture data corresponding to the next historical time of the simulation character and its corresponding motion capture target posture data, and a fourth preset loss function, and regenerating the torque of each joint of the simulation character based on the updated action encoding model and the action decoding model; updating the parameters of the environment model based on the simulation posture data corresponding to the next historical time of the simulation character and its corresponding predicted simulation posture data, the regenerated torque of each joint of the simulation character, and a preset loss function of the environment model.

6. The motion generating method of an imitation character according to Claim 3, characterized by, After updating the parameters of the action encoding model and the action decoding model, the method further comprises: determining and saving the simulation posture data corresponding to the next historical time of the simulation character based on the simulation posture data corresponding to the previous historical time of the simulation character and the torque of each joint of the simulation character determined by the updated action decoding model.

7. A training method of a motion generation model of an imitation character, characterized by, The action generation model comprises a joint torque generation model of a simulation character and an environment model; the joint torque generation model of the simulation character is used to determine the torque of each joint of the simulation character; the joint torque generation model of the simulation character is obtained through joint training with the environment model; the environment model is used to generate predicted simulation pose data based on the torque of each joint of the simulation character, and the predicted simulation pose data is used for autoregressive training of the joint torque generation model of the simulation character and the environment model; the method comprises the following steps: obtaining historical simulation pose data corresponding to a simulation character and action capture target pose data corresponding to the historical simulation pose data, and historical control signals; the action capture target pose data is used to represent the pose data of the simulation character at a next time point corresponding to each historical time point; iteratively updating an action encoding model and an action decoding model in the joint torque generation model of the simulation character and the environment model based on the historical simulation pose data corresponding to the simulation character, the action capture target pose data corresponding to the historical simulation pose data, a fourth preset loss function, and a preset loss function of the environment model; determining a policy model in the joint torque generation model of the simulation character based on the historical simulation pose data corresponding to the simulation character, the historical control signals, the updated action encoding model and the action decoding model, and the environment model according to a target loss function; The joint torque generation model of the simulated character includes a strategy model, which is determined based on the historical simulated posture data and historical control signals of the simulated character. The strategy model is used to determine the offset in the encoding space. The historical control signals... This includes the orientation input by the user. and speed Accordingly, the training process of the strategy model includes: acquiring historical simulation posture data corresponding to the simulation character. and historical control signals Based on the historical simulation posture data and historical control signals corresponding to the simulated character, a strategy model for the simulated character is determined according to a first preset loss function and a second preset loss function; wherein, the first preset loss function and the second preset loss function are used together to control the direction and speed of the simulated character's movement; wherein, the first preset loss function is: The second preset loss function is: ;in, The simulation pose data of the simulated character at the next moment, generated from the environment model. The orientation and velocity of the root node. For the policy model, h is a constant representing the orientation of the input. and speed Total number respectively with Correspondingly, these represent the user's input orientation and speed at a certain moment.

8. An action generating apparatus for simulating a character, characterized by comprising: The simulation character action generation method of any one of claims 1-6 is implemented, and the device comprises: a first obtaining module configured to obtain simulation pose data of a current time corresponding to a simulation character and a current control signal used to control the simulation character; a generating module configured to generate target simulation pose data based on the simulation pose data of the current time corresponding to the simulation character, the current control signal, a pre-trained joint torque generation model of the simulation character, and a simulation environment; the target simulation pose data is simulation pose data of the simulation character at a next time corresponding to the current control signal; The joint torque generation model of the simulation character is used to determine the torque of each joint of the character; the joint torque generation model of the simulation character is obtained through joint training with the environment; the environment model is used to generate predicted simulation pose data based on the torque of each joint; the predicted simulation pose data is used for autoregressive training of the joint torque generation model of a simulation character and the environment model.

9. A training device of a motion generation model of an imitation character, characterized by, The simulation character action generation model training method of claim 7 is implemented, and the action generation model comprises a joint torque generation model of a simulation character and an environment model; the simulation character joint torque generation model is obtained through joint training with the environment model; the environment model is used to generate predicted simulation posture data based on the torque of each joint of the simulation character; the predicted simulation pose data is used for autoregressive training of the joint torque generation model and the environment model of the simulation character; the device comprises: The second acquisition module is configured to acquire historical simulation pose data corresponding to the simulation character and motion capture target pose data corresponding to the historical simulation pose data, and historical control signals; the motion capture target pose data is used to represent a pose data of the simulation character at a next time corresponding to each historical time; The updating module is configured to iteratively update, based on the historical simulation pose data corresponding to the simulation character, the motion capture target pose data corresponding to the historical simulation pose data, a fourth preset loss function, and a loss function of a preset environment model, to obtain an action encoding model and an action decoding model in a joint torque generation model of the simulation character and the environment model. The determining module is configured to determine, based on the historical simulation pose data corresponding to the simulation character, the historical control signals, the updated action encoding model and the action decoding model, and the environment model, a policy model in the joint torque generation model of the simulation character according to a target loss function.

10. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the simulation character action generation method according to any one of claims 1 to 6 or the simulation character action generation model training method of claim 7 when executing the program. 11.A non-transitory computer-readable storage medium having stored thereon a computer program. The computer program implements the steps of the simulation character action generation method according to any one of claims 1 to 6 or the simulation character action generation model training method of claim 7 when executed by the processor.