Training Method, Action Generation Method and Device for Virtual Character Action Generation Model

By decomposing the movements of virtual images into basic postures and amplitude rhythms, and generating and adjusting the movements separately, the problems of poor realism and low matching with speech in the prior art are solved, and a higher sense of realism and matching of action are achieved.

CN114972590BActive Publication Date: 2025-06-17BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210676086.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-06-17
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

When generating body movements of virtual images, the prior art fails to effectively consider the characteristics of body postures of virtual images when performing movements, resulting in poor realism of generated movements, serious deformation of body movements, and low degree of matching with voice.

Method used

The speaker's movements are broken down into two aspects: basic posture and amplitude rhythm, and the basic posture and posture adjustment sequence are generated respectively. Through virtual image action generation model training, the realism of the movement and matching with speech are improved.

Benefits of technology

By decomposing movements as the basic posture and posture adjustment sequence, the generated virtual image movements are real, and the matching between movements and voices is enhanced, solving the problems of poor realism and low matching in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972590B_ABST
    Figure CN114972590B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method, an action generation method, and an apparatus for a virtual avatar action generation model, and relates to the field of artificial intelligence. The training method for the virtual avatar action generation model includes: for each first image sequence of the virtual avatar, obtaining the ground truth of the first image sequence, the ground truth of the second image sequence corresponding to the first image sequence, and the speech corresponding to the second image sequence; using the virtual avatar action generation model to process the ground truth of the first image sequence, the ground truth of the second image sequence, and the speech, and obtaining a processing result, including: generating a predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence; generating a predicted value of the pose adjustment sequence of the second image sequence according to the speech; generating a predicted value of the second image sequence according to the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence; and training the virtual avatar action generation model according to the processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and particularly to a training method for a virtual character action generation model, a virtual character action generation method and apparatus, an electronic device, and a computer-readable storage medium. Background Art

[0002] When people talk, they often accompany natural body movements. These body movements can assist in the organization and presentation of language content, bringing a better listening experience to the audience. The virtual human technology refers to generating a corresponding speaker video given a piece of speech. It includes a key step: lip synchronization generation and body movement generation. Lip synchronization generation refers to generating mouth movements corresponding to the speech, and body movement generation refers to generating body movements corresponding to the speech.

[0003] Compared with lip synchronization generation, body movement generation is a more difficult problem. Some methods use a model similar to lip synchronization generation to model body movement generation, using a convolutional neural network or a sequence model, taking speech features as input, directly regressing the positions of body key points, and then using mean square error, etc. as a supervision signal for training. Summary of the Invention

[0004] The inventors believe that: the prior art directly regresses the positions of body key points according to speech features, without considering the characteristics of the body posture when the virtual character performs actions, resulting in very poor realism of the generated actions, serious deformation of body movements, and low matching degree between the actions and the speech.

[0005] In view of the above technical problems, the present disclosure proposes a solution, which decomposes the actions of the speaker into two aspects: basic posture and amplitude rhythm, and generates the basic posture and the posture adjustment sequence respectively, improving the realism of the generated actions and the matching with the speech.

[0006] According to a first aspect of the present disclosure, a method for training a virtual character action generation model includes: for each first image sequence of a virtual character, obtaining the ground truth of the first image sequence, the ground truth of a second image sequence corresponding to the first image sequence, and the speech corresponding to the second image sequence, wherein in all images of the first image sequence, the virtual character corresponds to the same basic pose, and in all images of the second image sequence, the virtual character corresponds to the same basic pose; using the virtual character action generation model to process the ground truth of the first image sequence, the ground truth of the second image sequence, and the speech, and obtaining a processing result, including: generating a predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence; generating a predicted value of a pose adjustment sequence of the second image sequence according to the speech, wherein the pose adjustment sequence is a parameter sequence of at least one of the action amplitude and action rhythm of the virtual character in the second image sequence; generating a predicted value of the second image sequence according to the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence; training the virtual character action generation model according to the processing result.

[0007] In some embodiments, the generating a predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence includes: generating a posterior distribution of a latent variable according to the ground truth of the first image sequence and the ground truth of the second image sequence, wherein the latent variable is a random variable, and the posterior distribution of the latent variable is related to the basic pose of the second image sequence and is independent of both the action amplitude and action rhythm of the virtual character; generating a predicted value of the basic pose of the second image sequence according to the posterior distribution of the latent variable.

[0008] In some embodiments, the generating a posterior distribution of a latent variable according to the ground truth of the first image sequence and the ground truth of the second image sequence includes: calculating a first intermediate variable according to the ground truth of the first image sequence and the ground truth of the second image sequence; generating a posterior distribution of the latent variable with the first intermediate variable as a sample as the posterior distribution of the latent variable.

[0009] In some embodiments, the calculating a first intermediate variable according to the ground truth of the first image sequence and the ground truth of the second image sequence includes: using a first encoder to calculate the encodings of the ground truth of the first image sequence and the ground truth of the second image sequence; calculating a first intermediate variable according to the encodings of the ground truth of the first image sequence and the ground truth of the second image sequence, wherein the first intermediate variable includes information about the change from the basic pose of the first image sequence to the basic pose of the second image sequence.

[0010] In some embodiments, the generating a posterior distribution of the latent variable with the first intermediate variable as a sample as the posterior distribution of the latent variable includes: using a second encoder to calculate the latent space encoding of the first intermediate variable as the posterior distribution of the latent variable.

[0011] In some embodiments, generating a predicted value of the base pose of the second image sequence according to the posterior distribution of the latent variable includes: determining the value of the latent variable according to the posterior distribution of the latent variable; and generating a predicted value of the base pose of the second image sequence according to the value of the latent variable.

[0012] In some embodiments, determining the value of the latent variable according to the posterior distribution of the latent variable with the first intermediate variable as a sample includes: determining the value of the latent variable by sampling the posterior distribution of the latent variable.

[0013] In some embodiments, generating a predicted value of the base pose of the second image sequence according to the value of the latent variable includes: using a second decoder corresponding to the second encoder to calculate a decoding result of the value of the latent variable; and using a first decoder corresponding to the first encoder to generate a predicted value of the base pose of the second image sequence according to the decoding result of the value of the latent variable.

[0014] In some embodiments, the posterior distribution of the latent variable is a first distribution or a second distribution; when there is no change from the base pose of the first image sequence to the base pose of the second image sequence, the posterior distribution of the latent variable is the first distribution; when there is a change from the base pose of the first image sequence to the base pose of the second image sequence, the posterior distribution of the latent variable is the second distribution, and the second distribution is different from the first distribution.

[0015] In some embodiments, training the virtual avatar action generation model according to the processing result includes: calculating a first loss function according to the posterior distribution of the latent variable, the preset prior distribution of the latent variable, and the speech; and training the virtual avatar action generation model according to the first loss function.

[0016] In some embodiments, calculating the first loss function according to the posterior distribution of the latent variable, the preset prior distribution of the latent variable, and the speech includes: predicting whether there is a change from the base pose of the first image sequence to the base pose of the second image sequence according to the speech to obtain a prediction result; when the prediction result is that there is no change from the base pose of the first image sequence to the base pose of the second image sequence, calculating the first loss function according to the expectation and variance of the first distribution; and when the prediction result is that there is a change from the base pose of the first image sequence to the base pose of the second image sequence, calculating the first loss function according to the information entropy of the prior distribution of the latent variable and the second distribution.

[0017] In some embodiments, predicting whether there is a change from the base pose of the first image sequence to the base pose of the second image sequence to obtain a prediction result includes: converting the speech into text; and predicting whether there is a change from the base pose of the first image sequence to the base pose of the second image sequence according to the keywords in the text to obtain a prediction result.

[0018] In some embodiments, generating a predicted value of a pose adjustment sequence of a second image sequence according to speech includes: extracting Mel cepstral coefficients of the speech; generating a predicted value of a pose adjustment sequence of the second image sequence according to the Mel cepstral coefficients of the speech.

[0019] In some embodiments, training a virtual avatar motion generation model according to a processing result includes: calculating a second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence; training the virtual avatar motion generation model according to the second loss function.

[0020] In some embodiments, calculating the second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence includes: calculating the second loss function according to the ground truth of the second image sequence, the mean value of the ground truth of the second image sequence, and the predicted value of the pose adjustment sequence of the second image sequence.

[0021] In some embodiments, calculating the second loss function according to the mean value of the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence includes: calculating a second intermediate variable according to the difference between the ground truth of the second image sequence and the mean value of the ground truth of the second image sequence, where the second intermediate variable is related to at least one of the motion amplitude and motion rhythm of the virtual avatar and is independent of the basic pose of the virtual avatar; calculating the second loss function according to the difference between the predicted value of the pose adjustment sequence of the second image sequence and the second intermediate variable.

[0022] In some embodiments, training a virtual avatar motion generation model according to a processing result includes: calculating a third loss function according to the ground truth of the second image sequence and the predicted value of the second image sequence; training the virtual avatar motion generation model according to the third loss function.

[0023] In some embodiments, training a virtual avatar motion generation model according to a processing result includes: calculating a first loss function according to the posterior distribution of a latent variable, a preset prior distribution of the latent variable, and speech; calculating a second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence; calculating a third loss function according to the ground truth of the second image sequence and the predicted value of the second image sequence; training the virtual avatar motion generation model according to the weighted sum of the first loss function, the second loss function, and the third loss function.

[0024] In some embodiments, processing the ground truth of the first image sequence, the ground truth of the second image sequence, and the speech by using the virtual avatar motion generation model to obtain a processing result includes: calculating the decoding result of the encoding of the ground truth of the first image sequence and the decoding result of the encoding of the ground truth of the second image sequence by using a first decoder corresponding to a first encoder; training the virtual avatar motion generation model according to the processing result includes: calculating a fourth loss function according to the decoding result of the encoding of the ground truth of the first image sequence, the decoding result of the encoding of the ground truth of the second image sequence, the ground truth of the first image sequence, and the ground truth of the second image sequence; training the virtual avatar motion generation model according to the fourth loss function.

[0025] In some embodiments, calculating the fourth loss function according to the decoding result of the encoding of the ground truth of the first image sequence, the decoding result of the encoding of the ground truth of the second image sequence, the ground truth of the first image sequence, and the ground truth of the second image sequence includes: calculating the fourth loss function according to the norm of the difference between the decoding result of the encoding of the ground truth of the first image sequence and the ground truth of the first image sequence, and the norm of the difference between the decoding result of the encoding of the ground truth of the second image sequence and the ground truth of the second image sequence.

[0026] In some embodiments, training the virtual avatar motion generation model according to the fourth loss function includes: calculating a first loss function according to the posterior distribution of the latent variable, the preset prior distribution of the latent variable, and the speech; calculating a second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence; calculating a third loss function according to the ground truth of the second image sequence and the predicted value of the second image sequence; training the virtual avatar motion generation model according to the weighted sum of the first loss function, the second loss function, the third loss function, and the fourth loss function.

[0027] In some embodiments, the second image sequence is the next image sequence of the first image sequence, the speech is the speech uttered by the virtual avatar when executing the second image sequence, and the predicted value of the basic pose of the second image sequence includes the predicted value of the basic pose of at least one of the head, upper limbs, lower limbs, and torso of the virtual avatar in the second image sequence.

[0028] In some embodiments, generating the predicted value of the second image sequence according to the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence includes: based on the predicted value of the basic pose of the second image sequence, using the predicted value of the pose adjustment sequence of the second image sequence to adjust at least one of the motion amplitude or motion rhythm of the virtual avatar, and taking the adjustment result as the predicted value of the second image sequence.

[0029] In some embodiments, generating the predicted value of the second image sequence based on the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence includes: generating a sequence of predicted values of the basic pose of the second image sequence according to the predicted value of the basic pose of the second image sequence, where the basic poses of the virtual avatars in the sequence of predicted values of the basic pose of the second image sequence are all the predicted values of the basic pose; obtaining the predicted value of the second image sequence by adding the sequence of predicted values of the basic pose of the second image sequence to the predicted value of the pose adjustment sequence of the second image sequence.

[0030] According to a second aspect of the present disclosure, there is provided a method for generating actions of a virtual avatar, including: obtaining an initial basic pose and speech of the virtual avatar; using a virtual avatar action generation model to generate a basic pose of an image sequence corresponding to the speech according to the initial basic pose and the speech, where one image sequence corresponds to one basic pose, and the virtual avatar action generation model is trained according to the virtual avatar action generation model training method described in any embodiment of the present disclosure; using the virtual avatar action generation model to generate a pose adjustment sequence of the image sequence according to the speech, where the pose adjustment sequence is a parameter sequence of at least one of the action amplitude and action rhythm of the virtual avatar in the image sequence; generating an image sequence according to the basic pose and the pose adjustment sequence as the action of the virtual avatar.

[0031] In some embodiments, generating an image sequence according to the basic pose and the pose adjustment sequence as the action of the virtual avatar includes: based on the basic pose of the image sequence, using the pose adjustment sequence of the image sequence to adjust at least one of the action amplitude or action rhythm of the virtual avatar, and taking the adjustment result as the image sequence.

[0032] In some embodiments, generating an image sequence according to the basic pose and the pose adjustment sequence as the action of the virtual avatar includes: generating a sequence of basic poses of the image sequence according to the basic pose of the image sequence, where the virtual avatars in the sequence of basic poses of the image sequence are all the same basic pose; generating the image sequence by adding the sequence of basic poses of the image sequence to the pose adjustment sequence of the image sequence.

[0033] In some embodiments, obtaining the initial basic pose and the speech corresponding to the image sequence includes: dividing the speech into multiple speech segments, where the multiple speech segments are arranged in time sequence; taking the obtained initial basic pose as the initial basic pose of the first speech segment; for other speech segments except the first speech segment, taking the basic pose of the previous speech segment of the current speech segment generated by using the virtual avatar action generation model as the initial basic pose of the current speech segment.

[0034] According to a third aspect of the present disclosure, there is provided a training device for a virtual avatar action generation model, including: an acquisition unit configured to, for each first image sequence of the virtual avatar, acquire the ground truth of the first image sequence, the ground truth of the second image sequence corresponding to the first image sequence, and the voice corresponding to the second image sequence, wherein in all images of the first image sequence, the virtual avatar corresponds to the same basic pose, and in all images of the second image sequence, the virtual avatar corresponds to the same basic pose; a processing unit configured to use the virtual avatar action generation model to process the ground truth of the first image sequence, the ground truth of the second image sequence, and the voice to obtain a processing result, including: generating a predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence; generating a predicted value of the pose adjustment sequence of the second image sequence according to the voice, wherein the pose adjustment sequence is a parameter sequence of at least one of the action amplitude and action rhythm of the virtual avatar in the second image sequence; generating a predicted value of the second image sequence according to the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence; a training unit configured to train the virtual avatar action generation model according to the processing result.

[0035] According to a fourth aspect of the present disclosure, there is provided an action generation device for a virtual avatar, including: an acquisition unit configured to acquire the initial basic pose and voice of the virtual avatar; an initial basic pose generation unit configured to use the virtual avatar action generation model to generate the basic pose of the image sequence corresponding to the voice according to the initial basic pose and the voice, wherein one image sequence corresponds to one basic pose, and wherein the virtual avatar action generation model is trained by the virtual avatar action generation model training device according to any embodiment of the present disclosure; a pose adjustment sequence generation unit configured to use the virtual avatar action generation model to generate the pose adjustment sequence of the image sequence according to the voice, wherein the pose adjustment sequence is a parameter sequence of at least one of the action amplitude and action rhythm of the virtual avatar in the image sequence; an image sequence generation unit configured to generate an image sequence according to the basic pose and the pose adjustment sequence as the action of the virtual avatar.

[0036] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: a memory; and a processor coupled to the memory, the processor being configured to execute the virtual avatar action generation model training method according to any embodiment of the present disclosure or execute the virtual avatar action generation method according to any embodiment of the present disclosure based on instructions stored in the memory. Description of the Drawings

[0037] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0038] With reference to the accompanying drawings, the present disclosure can be more clearly understood from the following detailed description, wherein:

[0039] Figure 1 A flowchart showing a method for training an avatar motion generation model according to some embodiments of the present disclosure;

[0040] Figure 2 A flowchart showing the processing of training samples according to some embodiments of the present disclosure;

[0041] Figure 3 A schematic diagram showing a method for generating a basic pose during model training according to some embodiments of the present disclosure;

[0042] Figure 4 A schematic diagram showing a method for generating a pose adjustment sequence during model training according to some embodiments of the present disclosure;

[0043] Figure 5 A schematic diagram showing the prediction value of generating a second image sequence according to some embodiments of the present disclosure;

[0044] Figure 6 A flowchart showing a method for generating an avatar's motion according to some embodiments of the present disclosure;

[0045] Figure 7 A block diagram showing a training device for an avatar motion generation model according to some embodiments of the present disclosure;

[0046] Figure 8 A block diagram showing an action generation device for an avatar according to some embodiments of the present disclosure;

[0047] Figure 9 A block diagram showing an electronic device according to some other embodiments of the present disclosure;

[0048] Figure 10 A block diagram showing a computer system for implementing some embodiments of the present disclosure. Detailed Description of Specific Embodiments

[0049] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0050] At the same time, it should be understood that, for the sake of convenience of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0051] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present disclosure or its application or use.

[0052] Techniques, methods, and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered as part of the specification.

[0053] In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0054] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.

[0055] In the prior art, a deterministic model similar to mouth shape generation is used to model limb movement generation, directly generating the actions of a virtual character from the input speech, without considering the characteristics of the body postures in the actions of the virtual image, resulting in actions lacking realism, severely deformed limbs, and poor matching with the speech.

[0056] In addition, the relationship from speech to limb movement is many-to-many. A single piece of speech may correspond to multiple different action sequences, with a great deal of uncertainty. For example, when a person is giving a speech, to relieve tension and appear more natural, they often use some habitual body postures. These body postures are not necessarily strongly related to the speech and are just subconscious actions.

[0057] However, in the prior art, a one-to-one mapping model is used to fit the many-to-many relationship, ignoring this inherent uncertainty, resulting in the actions generated using these models being one-to-one corresponding to the input speech, lacking diversity, and making the actions of the virtual image overly rigid.

[0058] To solve the above problems, the present disclosure proposes a method for training a virtual image action generation model.

[0059] The present disclosure proposes the concept of a pose pattern (i.e., a basic pose). The body posture of a speaker is regarded as a random vector that follows a multi-modal distribution in a high-dimensional space. A mode of this distribution is a maximum point in the probability density function of the distribution, reflecting the habitual body postures of the virtual image, including the postures of at least one of the head, upper limbs, lower limbs, and torso. For example, the common gestures and standing postures of the virtual image, etc. This maximum point can be defined as the pose mode of the virtual image and used as the basic pose of the virtual image.

[0060] The basic pose is only related to the action type of the virtual character. For example, clapping hands and raising hands are different basic poses. The basic pose is independent of the amplitude and size of the virtual character's actions. For example, whether the hand is raised higher or lower when raising hands, it only corresponds to one basic pose: raising hands; regardless of the speed of the clapping rhythm, it only corresponds to one basic pose: clapping hands. The basic pose can also be understood as a standard pose.

[0061] Figure 1 The flowchart showing the training method of the virtual character action generation model according to some embodiments of the present disclosure.

[0062] As Figure 1 shown, the training method of the virtual character action generation model includes steps S1 - S3.

[0063] In step S1, for each first image sequence of the virtual character, obtain the ground truth of the first image sequence, the ground truth of the second image sequence corresponding to the first image sequence, and the voice corresponding to the second image sequence, where in all images of the first image sequence, the virtual character corresponds to the same basic pose, and in all images of the second image sequence, the virtual character corresponds to the same basic pose.

[0064] For example, according to the basic pose, the video of the virtual character can be divided into multiple sub - videos. Each sub - video constitutes an action, which is an image sequence, and an image sequence can include multiple image frames. In each image sequence, the virtual character has only one basic pose for its action, but the amplitude and rhythm of the action can vary.

[0065] In some embodiments, the second image sequence is the next image sequence of the first image sequence, the voice is the voice emitted by the virtual character when performing the second image sequence, and the predicted value of the basic pose of the second image sequence includes the predicted value of the basic pose of at least one of the head, upper limbs, lower limbs, and torso of the virtual character in the second image sequence.

[0066] For the training sample (M (i-1) , M i , S i ), the ground truth M i-1 of the first image sequence and the ground truth M i of the second image sequence are the ground truths of two actions of the virtual character, and M i-1 and M i are temporally correlated, and the voice is the voice corresponding to M i . For example, M i is the next action of M i-1 , and the virtual character performs M i while emitting the voice S i .

[0067] In step S2, the ground truth of the first image sequence, the ground truth of the second image sequence, and the speech are processed using the virtual character motion generation model to obtain a processing result.

[0068] The actions of the speaker can be decomposed into two aspects: the basic pose (such as clapping) and the amplitude rhythm (such as the position of the hands and the clapping speed when clapping). M i-1 and M i These two image sequences are equivalent to two actions of the virtual character, and each action contains both the basic pose information and the information of at least one of the action amplitude and rhythm.

[0069] Figure 2 The flowchart showing the processing of training samples according to some embodiments of the present disclosure.

[0070] As Figure 2 shown, using the virtual character motion generation model, the ground truth of the first image sequence, the ground truth of the second image sequence, and the speech are processed to obtain a processing result, including step S20 - step S22.

[0071] The virtual character motion generation model is a limb motion generation framework including two modules, including a basic pose module and a pose adjustment module. When training the model, in step S20, the action and amplitude information in M i-1 and M i are removed using the pose module to generate the predicted value of the basic pose of the second image sequence. In step S21, the pose adjustment module generates a pose adjustment sequence according to the speech prosody to adjust at least one of the action amplitude and rhythm of the virtual character. Steps S20 and S21 have no sequence and can be executed in parallel.

[0072] In step S20, according to the ground truth of the first image sequence and the ground truth of the second image sequence, the predicted value of the basic pose of the second image sequence is generated.

[0073] In some embodiments, generating the predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence includes: generating the posterior distribution of the latent variable according to the ground truth of the first image sequence and the ground truth of the second image sequence, where the latent variable is a random variable, and the posterior distribution of the latent variable is related to the basic pose and is irrelevant to both the action amplitude and the action rhythm; generating the predicted value of the basic pose of the second image sequence according to the posterior distribution of the latent variable.

[0074] For example, here the latent variable z is introduced, and z is an unobservable random variable and cannot be directly based on the training samples M i-1 and M iIt can be obtained. The posterior probability distribution of the latent variable with respect to the basic posture can be obtained by observing the training samples. Therefore, the training samples M i-1 and M i are input into the basic posture module to calculate the posterior probability distribution of the latent variable. According to the posterior distribution of the latent variable, the predicted value of the basic posture of the second image sequence can be restored

[0075] By introducing the latent variable z, the non-determinism from speech to limb movement is modeled, and the process of generating the predicted value of the basic posture is transformed into a random process related to the basic posture, so that the relationship between the input data and the generated result of the basic posture module is many-to-many. That is, for the same training sample, different predicted values of the basic posture can be generated Improve the realism and diversity of the generated actions.

[0076] In some embodiments, according to the ground truth of the first image sequence and the ground truth of the second image sequence, the posterior distribution of the latent variable is generated, including: calculating a first intermediate variable according to the ground truth of the first image sequence and the ground truth of the second image sequence; generating the posterior distribution of the latent variable with the first intermediate variable as a sample as the posterior distribution of the latent variable.

[0077] For example, the samples M i-1 and M i are input into the basic posture module to calculate the first intermediate variable τ. Taking τ as an observable variable, the posterior distribution P(z|τ) of z with respect to τ is obtained as the posterior distribution of z.

[0078] In some embodiments, calculating the first intermediate variable according to the ground truth of the first image sequence and the ground truth of the second image sequence includes: using a first encoder to calculate the encodings of the ground truth of the first image sequence and the ground truth of the second image sequence; calculating the first intermediate variable according to the encodings of the ground truth of the first image sequence and the ground truth of the second image sequence, where the first intermediate variable includes information about the change from the basic posture of the first image sequence to the basic posture of the second image sequence.

[0079] Figure 3 A schematic diagram showing a method for generating a basic posture during model training according to some embodiments of the present disclosure.

[0080] As Figure 3 shown, M i-1 and M i are input into the basic posture module, and the first encoder is used to calculate the encodings e i-1 of M i-1 and the encodings e i of M i , and according to e i-1 and e iDetermine the first intermediate variable τ. For example, e i-1 and e i The difference between them can be used as τ. It is not necessary to directly observe the basic postures of M i-1 and M i from the training samples M i-1 and M i . By the method of taking the difference, τ can include the information of the transformation from the basic posture of M i-1 to the basic posture of M i .

[0081] In some embodiments, generating the posterior distribution of the latent variable with the first intermediate variable as the sample as the posterior distribution of the latent variable includes: using a second encoder to calculate the expectation and variance of the posterior distribution of the latent variable with the first intermediate variable as the sample; generating the posterior distribution of the latent variable with the first intermediate variable as the sample according to the expectation and variance as the posterior distribution of the latent variable.

[0082] For example, through the second encoder, calculate the mathematical expectation μ θ (τ) and variance ∑ θ (τ) of P(z|τ). It can be assumed that P(z|τ) is a normal distribution, then P(z|τ) can be expressed as

[0083] The basic posture module does not directly obtain the basic postures of M i-1 and M i , but in the process of encoding τ into the latent space using the second encoder, the information of the posture adjustment parameters in the input samples M i-1 and M i is removed, and only the information of the basic posture transformation is retained, so that the generated posterior distribution P(z|τ) is related to the basic posture and independent of the posture adjustment parameters.

[0084] Encoding τ into the latent space through the second encoder can enable the model to learn P(z|τ), thereby transforming the generation process of the basic posture into a random process.

[0085] In some embodiments, the second encoder employs a CVAE (Conditional Variational AutoEncoder).

[0086] In some embodiments, determine the value of the latent variable according to the posterior distribution of the latent variable; generate the predicted value of the basic posture of the second image sequence according to the value of the latent variable.

[0087] For example, after determining p(z|τ), the value of z can be determined. Based on the value of z, the predicted value of the basic pose of the second image sequence can be restored. Since P(z|τ) is related to the basic pose and independent of the pose adjustment parameters, the z generated according to p(z|τ) is also related to the basic pose and independent of the pose adjustment parameters.

[0088] In some embodiments, the posterior distribution of the latent variable is the first distribution or the second distribution; when there is no change from the basic pose of the first image sequence to the basic pose of the second image sequence, the posterior distribution of the latent variable is the first distribution; when there is a change from the basic pose of the first image sequence to the basic pose of the second image sequence, the posterior distribution of the latent variable is the second distribution, and the second distribution is different from the first distribution.

[0089] For example, when there is no change from the basic pose of the first image sequence to the basic pose of the second image sequence, the posterior distribution of the latent variable is P0, and when there is a change from the basic pose of the first image sequence to the basic pose of the second image sequence, the posterior distribution of the latent variable is P1.

[0090] In some embodiments, determining the value of the latent variable according to the posterior distribution of the latent variable with the first intermediate variable as a sample includes: determining the value of the latent variable by sampling the posterior distribution of the latent variable.

[0091] For example, when the posterior distribution of the latent variable is P0, the results obtained by sampling the same distribution multiple times are the same. That is to say, when there is no change from the basic pose of the first image sequence to the basic pose of the second image sequence, the value of the generated latent variable is determined, so that the basic pose of the second image sequence generated at this time is also determined, that is, the same as the basic pose of the first image sequence.

[0092] When the posterior distribution of the latent variable is P1, the results obtained by sampling the same distribution multiple times may be different. Due to the uncertainty of the sampling results, the diversity of the generated basic poses can be realized.

[0093] In some embodiments, generating the predicted value of the basic pose of the second image sequence according to the value of the latent variable includes: using the second decoder corresponding to the second encoder to calculate the decoding result of the value of the latent variable; using the first decoder corresponding to the first encoder to generate the predicted value of the basic pose of the second image sequence according to the decoding result of the value of the latent variable.

[0094] For example, first input z into the second decoder to obtain τ * , and then, and τ *They are respectively input into the first decoder, and the decoding results are added to obtain is the reconstruction result of e i , The process of i can be understood as: e is first encoded into the latent space by the second encoder and then restored by the second decoder to obtain z is related to the basic pose and has nothing to do with the pose adjustment parameters. Therefore,

[0095] In step S21, according to the speech, a predicted value of the pose adjustment sequence of the second image sequence is generated, where the pose adjustment sequence is a parameter sequence of at least one of the action amplitude and action rhythm of the virtual character in the second image sequence.

[0096] Figure 4 FIG. shows a schematic diagram of a method for generating a pose adjustment sequence during model training according to some embodiments of the present disclosure.

[0097] As Figure 4 shown, in the rhythm adjustment module, a fully convolutional network is used as the generator. First, the prosodic features of the speech are extracted, and then, according to the prosodic features of the speech, a predicted value of the pose adjustment sequence is generated through the fully convolutional network By generating a predicted value of the pose adjustment sequence according to the prosodic features of the speech the action rhythm of the generated image sequence can be controlled by the prosodic information in the speech without affecting the predicted value of the basic pose.

[0098] In some embodiments, generating a predicted value of the pose adjustment sequence of the second image sequence according to the speech includes: extracting the mel cepstral coefficients of the speech; generating a predicted value of the pose adjustment sequence of the second image sequence according to the mel cepstral coefficients of the speech.

[0099] In step S22, according to the predicted value of the basic pose and the predicted value of the pose adjustment sequence of the second image sequence, a predicted value of the second image sequence is generated.

[0100] Figure 5 FIG. shows a schematic diagram of generating a predicted value of the second image sequence according to some embodiments of the present disclosure.

[0101] As Figure 5As shown, for the convenience of calculation, the predicted value of a basic pose is extended into a sequence of predicted values of a basic pose. All the image frames in the sequence of predicted values of the basic pose are the same as the predicted value of the basic pose. The pose adjustment module is a sequence of parameters generated by the pose adjustment module, and each parameter is a multi-dimensional vector, which is an adjustment coefficient for the movement amplitude or rhythm. Based on the sequence of predicted values of the basic pose in the second image sequence, the parameters in the predicted value of the pose adjustment sequence are used to adjust the movement amplitude of the virtual character in each image frame. The image frames in the sequence of predicted values of the basic pose may only include key frames, and the parameters in the predicted value of the pose adjustment sequence can be used to adjust the duration from one key frame to another key frame, so as to adjust the movement change rhythm of the virtual character.

[0102] In step S3, according to the processing result, train the virtual character action generation model.

[0103] In some embodiments, training the virtual character action generation model according to the processing result includes: calculating a first loss function according to the posterior distribution of the latent variable, the preset prior distribution of the latent variable, and the speech; training the virtual character action generation model according to the first loss function.

[0104] For example, the preset prior distribution P(z) of the latent variable follows a normal distribution, that is Then, use the first loss function to realize the fitting of the posterior distribution to the prior distribution, and use the speech as a condition to indicate that the pose pattern changes or remains unchanged.

[0105] In some embodiments, calculating the first loss function according to the posterior distribution of the latent variable, the preset prior distribution of the latent variable, and the speech includes: predicting whether the basic pose of the first image sequence changes to the basic pose of the second image sequence according to the speech, and obtaining a prediction result; when the prediction result is that the basic pose of the first image sequence does not change to the basic pose of the second image sequence, calculate the first loss function according to the expectation and variance of the first distribution; when the prediction result is that the basic pose of the first image sequence changes to the basic pose of the second image sequence, calculate the first loss function according to the prior distribution of the latent variable and the information entropy of the second distribution.

[0106] For example, the first loss function can be calculated by the following formula:

[0107] The first loss function Is expressed as:

[0108]

[0109] Among them, c i Is the prediction result, I0(c i ) and I1(c i) is a function of (c i ) :

[0110] When c i = 0, I0(c i ) = 1, when c i = 1, I0(c i ) = 0;

[0111] When c i = 0, I1(c i ) = 0, when c i = 1, I1(c i ) = 1.

[0112] Where τ is an intermediate variable, μ θ (τ) is the mathematical expectation of the distribution , ∑ θ (τ) is the variance of the distribution . ||μ θ (τ)|| represents the norm of μ θ (τ). ||∑ θ (τ)|| represents the norm of ∑ θ (τ).

[0113] Where represents calculating the divergence between the distribution and the standard normal distribution .

[0114] e i contains information about the movement amplitude and rhythm, but the reconstructed does not contain information about the movement amplitude and rhythm. During training, for samples where c i = 0 (M i , M i-1 ), they have the same basic pose, but may have different movement amplitudes and rhythms. We use the restriction to make the posterior distributions (μ θ (τ), ∑ θ (τ)) obtained by the second encoder for different samples the same, so that the different movement amplitude and rhythm information contained in (e i , e i-1 ) is ignored.

[0115] In some embodiments, predicting whether the basic pose of the first image sequence changes to the basic pose of the second image sequence to obtain a prediction result includes: converting speech into text; predicting whether the basic pose of the first image sequence changes to the basic pose of the second image sequence according to the keywords in the text to obtain a prediction result.

[0116] For example, first convert the speech S i into text, and then obtain semantic features through a keyword matching method. According to the semantic features, predict whether the basic pose of the first image sequence changes to the basic pose of the second image sequence, and obtain c i .

[0117] In the first loss function, guided by the semantics of S i , different situations are encoded into corresponding latent spaces, so that the two situations of whether the basic pose of the first image sequence changes to the basic pose of the second image sequence correspond to different distributions P0 and P1 respectively. That is to say, use c i to indicate whether the speech corresponds to the basic pose changing or remaining unchanged. The value of c i is only related to S i and has nothing to do with M i-1 . S i only determines whether the pose pattern changes or remains unchanged, that is, the value of S i can only determine whether the next basic pose is different from the basic pose in M i-1 . However, S i does not affect which specific pose the next basic pose is, thus realizing the diversity of actions.

[0118] In some embodiments, according to the processing result, train the virtual character action generation model, including: calculating a second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence; training the virtual character action generation model according to the second loss function.

[0119] By comparing the generated by the pose adjustment module with the ground truth, the second loss function can ensure the accuracy of the generated by the pose adjustment module according to the prosody of the speech.

[0120] In some embodiments, calculating the second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence includes: calculating the second loss function according to the ground truth of the second image sequence, the mean value of the ground truth of the second image sequence, and the predicted value of the pose adjustment sequence of the second image sequence.

[0121] For example, the mean value of the ground truth of the second image sequence can be the arithmetic mean of the ground truth of the second image sequence in time series, which can be expressed as

[0122] In some embodiments, a second loss function is calculated based on the mean value of the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence, including: calculating a second intermediate variable based on the difference between the ground truth of the second image sequence and the mean value of the ground truth of the second image sequence, wherein the second intermediate variable is related to at least one of the action amplitude and action rhythm of the virtual character and is independent of the basic pose of the virtual character; calculating the second loss function based on the difference between the predicted value of the pose adjustment sequence of the second image sequence and the second intermediate variable.

[0123] For example, the second intermediate variable is The second loss function can be calculated by the following formula

[0124]

[0125] where the symbol |||| represents the norm of the vector within ||||.

[0126] By finding the difference between M i and , information related to the basic pose can be removed, and the offset of each image in the sequence of the training sample M i relative to the mean value can be calculated. The purpose of training the pose adjustment module is to make the difference between the generated and the above offset smaller. Therefore, the second loss function can ensure the accuracy of the pose adjustment module in generating according to the prosody information in the speech without affecting the predicted value of the basic pose.

[0127] According to some embodiments of the present disclosure, a virtual character action generation model is trained based on the processing result, including: calculating a third loss function based on the ground truth of the second image sequence and the predicted value of the second image sequence; training the virtual character action generation model according to the third loss function.

[0128] For example, the third loss function can be calculated by the following formula

[0129]

[0130] is the loss function corresponding to the predicted value of the second image sequence, and by it can be ensured that the final result generated by the virtual character action generation model is correct.

[0131] According to some embodiments of the present disclosure, training a virtual avatar motion generation model according to a processing result includes: calculating a first loss function according to the posterior distribution of a latent variable, a preset prior distribution of the latent variable, and speech; calculating a second loss function according to the ground truth of a second image sequence and the predicted value of a pose adjustment sequence of the second image sequence; calculating a third loss function according to the ground truth of the second image sequence and the predicted value of the second image sequence; and training the virtual avatar motion generation model according to a weighted sum of the first loss function, the second loss function, and the third loss function.

[0132] For example, the first loss function is obtained through the following formula The second loss function and the third loss function for a weighted sum:

[0133]

[0134] According to some embodiments of the present disclosure, processing the ground truth of a first image sequence, the ground truth of a second image sequence, and speech by using a virtual avatar motion generation model to obtain a processing result includes: calculating a decoding result of an encoding of the ground truth of the first image sequence and a decoding result of an encoding of the ground truth of the second image sequence by using a first decoder corresponding to a first encoder; training the virtual avatar motion generation model according to the processing result, including: calculating a fourth loss function according to the decoding result of the encoding of the ground truth of the first image sequence, the decoding result of the encoding of the ground truth of the second image sequence, the ground truth of the first image sequence, and the ground truth of the second image sequence; and training the virtual avatar motion generation model according to the fourth loss function.

[0135] For example, the encoding of the ground truth of the first image sequence can be decoded to obtain a decoding result f dec (e i-1 ), and the encoding of the ground truth of the second image sequence can be decoded to obtain a decoding result f dec (e i ). The decoding results f dec (e i-1 ) and f dec (e i ) are sequences related only to the base pose. According to the decoding results f dec (e i-1 ), f dec (e i ), the ground truth of the first image sequence, and the ground truth of the second image sequence, a regularization term, i.e., the fourth loss function, is calculated.

[0136] According to some embodiments of the present disclosure, a fourth loss function is calculated based on the encoded decoding result of the first image sequence true value, the encoded decoding result of the second image sequence true value, the first image sequence true value and the second image sequence true value, including: calculating the fourth loss function based on the norm of the difference between the encoded decoding result of the first image sequence true value and the first image sequence true value, and the norm of the difference between the encoded decoding result of the second image sequence true value and the second image sequence true value.

[0137] For example, the fourth loss function can be calculated according to the following formula:

[0138] L reg =||M i-1 -f dec (e i-1 )||+||M i -f dec (e i )||

[0139] The role of the fourth loss function is to ensure that the randomly sampled z is not ignored by the model, that is, to ensure that according to e i-1 (i.e., the code corresponding to the action at the previous moment) generates the code corresponding to the action at the next moment This process is a random process related to hidden variables, not a completely deterministic process, not completely based on e i-1 Sure That is to say, Rather than

[0140] According to some embodiments of the present disclosure, a virtual image action generation model is trained according to a fourth loss function, including: calculating a first loss function according to the posterior distribution of latent variables, a preset prior distribution of latent variables, and speech; calculating a second loss function according to the true value of the second image sequence and the predicted value of the posture adjustment sequence of the second image sequence; calculating a third loss function according to the true value of the second image sequence and the predicted value of the second image sequence; and training the virtual image action generation model according to a weighted sum of the first loss function, the second loss function, the third loss function, and the fourth loss function.

[0141] For example, the first loss function can be calculated according to the following formula The second loss function And the third loss function And the fourth loss function The weighted sum of:

[0142]

[0143] According to some embodiments of the present disclosure, a fourth loss function is calculated based on the decoding result of the encoding of the first image sequence ground truth, the decoding result of the encoding of the second image sequence ground truth, the first image sequence ground truth, and the second image sequence ground truth, including: calculating the fourth loss function based on the norm of the difference between the decoding result of the encoding of the first image sequence ground truth and the first image sequence ground truth, and the norm of the difference between the decoding result of the encoding of the second image sequence ground truth and the second image sequence ground truth.

[0144] For example, the fourth loss function can be calculated by the following formula:

[0145]

[0146] When training the virtual character motion generation model according to the present disclosure, the speaker's motion is decomposed into two aspects: the basic pose (such as clapping) and the amplitude rhythm (such as the position of the hand and the clapping speed when clapping). The basic pose and the pose adjustment sequence are generated respectively, and then the pose of the virtual character is adjusted with the pose adjustment sequence, making full use of the pose information in the image and the audio features in the speech, which can improve the realism of the motion generated by the virtual character motion generation model and the matching with the speech.

[0147] To verify the effect of the training method of the virtual character motion generation model according to the present disclosure, the inventors used Figure 1 the training method of the motion generation model shown in the figure to train the virtual character motion generation model, and tested the trained virtual character motion generation model on the "Speech2Gesture" dataset and the "TedGesture" dataset.

[0148] The test results show that the motions of the virtual characters generated by the virtual character motion generation model are superior to the prior art in three aspects: realism (the degree of similarity to the body motions of people in reality when speaking), motion diversity, and matching with the speech (matching with the rhythm and amplitude of the speech prosody).

[0149] Figure 6 Shows a method for generating motions of a virtual character according to some embodiments of the present disclosure.

[0150] As Figure 6 shown, the method for generating motions of a virtual character includes steps S4 - S7.

[0151] In step S4, the initial basic pose and speech of the virtual character are obtained.

[0152] In step S5, using the virtual character motion generation model, according to the initial basic pose and speech, generate the basic pose of the image sequence corresponding to the speech, where one image sequence corresponds to one basic pose, and the virtual character motion generation model is trained according to the virtual character motion generation model training method of any embodiment of the present disclosure.

[0153] In some embodiments, generating the basic pose of the image sequence corresponding to the speech according to the initial basic pose and speech includes the following steps:

[0154] First, infer c i from the speech S i (using the method of keyword matching).

[0155] Then, use the basic pose module in the virtual character motion generation model to sample the value of z from P0(z) or P1(z) that has been learned during model training according to the value of c i . When c i = 0, sample from the distribution z ~ P0(z), and when c i = 1, sample from the distribution z ~ P1(z).

[0156] Then, use the second decoder to decode z to obtain

[0157] Finally, use the first decoder to decode to obtain the basic pose of the image sequence to be generated corresponding to the speech

[0158] In step S6, using the virtual character motion generation model, generate the pose adjustment sequence of the image sequence according to the speech, where the pose adjustment sequence is a parameter sequence of at least one of the action amplitude and action rhythm of the virtual character in the image sequence.

[0159] For example, using the rhythm adjustment module, use a fully convolutional network as the generator. First, extract the prosodic features of the speech S i , and then, according to the prosodic features of the speech, generate the predicted value of the pose adjustment sequence through the fully convolutional network By generating the predicted value of the pose adjustment sequence according to the prosodic features of the speech it is possible to control the action rhythm of the generated image sequence with the prosodic information in the speech without affecting the predicted value of the basic pose.

[0160] In step S7, generate the image sequence as the action of the virtual character according to the basic pose and the pose adjustment sequence.

[0161] For example, the predicted value of a basic pose can be extended into a sequence of predicted values of a basic pose, and all the image frames in the sequence of predicted values of the basic pose are the same as the predicted value of the basic pose. The pose adjustment module is a sequence of parameters generated by the pose adjustment module, and each parameter is a multi-dimensional vector, which is an adjustment coefficient for the movement amplitude or rhythm. Based on the sequence of predicted values of the basic pose in the second image sequence, the parameters in the predicted values of the pose adjustment sequence are used to adjust the movement amplitude of the virtual character in each image frame. The image frames in the sequence of predicted values of the basic pose may only include key frames, and the parameters in the predicted values of the pose adjustment sequence can be used to adjust the duration from one key frame to another, so as to adjust the movement change rhythm of the virtual character.

[0162] In some embodiments, generating an image sequence as the action of a virtual character according to a basic pose and a pose adjustment sequence includes: based on the basic pose of the image sequence, using the pose adjustment sequence of the image sequence to adjust at least one of the movement amplitude or movement rhythm of the virtual character, and taking the adjustment result as the image sequence.

[0163] In some embodiments, generating an image sequence as the action of a virtual character according to a basic pose and a pose adjustment sequence includes: generating a sequence of basic poses of the image sequence according to the basic pose of the image sequence, where the virtual characters in the sequence of basic poses of the image sequence are all the same basic pose; generating the image sequence by adding the sequence of basic poses of the image sequence and the pose adjustment sequence of the image sequence.

[0164] The present disclosure decomposes the actions of the speaker into two aspects: a basic pose (such as clapping) and an amplitude rhythm (such as the position of the hand and the clapping speed when clapping), generates the basic pose and the pose adjustment sequence respectively, and then uses the pose adjustment sequence to adjust the pose of the virtual character, making full use of the pose information in the image and the audio features in the speech, which can improve the realism of the generated actions and the matching with the speech.

[0165] In some embodiments, obtaining an initial basic pose and speech corresponding to an image sequence includes: dividing the speech into multiple speech segments, where the multiple speech segments are arranged in time sequence; taking the obtained initial basic pose as the initial basic pose of the first speech segment; for other speech segments except the first speech segment, taking the basic pose of the previous speech segment of the current speech segment (generated by using the virtual character action generation model) as the initial basic pose of the current speech segment.

[0166] For example, for a long speech, it is divided into multiple segments of speech and then gradually generated in sequence. When generating the actions corresponding to each segment of speech, the initial basic pose input into the virtual character generation model is the basic pose generated for the previous segment of speech.

[0167] Figure 7 A block diagram showing a training device for an avatar motion generation model according to some embodiments of the present disclosure.

[0168] As Figure 7 shown, the training device 7 of the avatar motion generation model includes an acquisition unit 71, a processing unit 72, and a training unit 73.

[0169] The acquisition unit 71 is configured to, for each first image sequence of the avatar, acquire the ground truth of the first image sequence, the ground truth of the second image sequence corresponding to the first image sequence, and the speech corresponding to the second image sequence, where in all images of the first image sequence, the avatar corresponds to the same basic pose, and in all images of the second image sequence, the avatar corresponds to the same basic pose. For example, perform the steps as Figure 1 shown in step S1.

[0170] The processing unit 72 is configured to use the avatar motion generation model to process the ground truth of the first image sequence, the ground truth of the second image sequence, and the speech to obtain a processing result, including: generating a predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence; generating a predicted value of the pose adjustment sequence of the second image sequence according to the speech, where the pose adjustment sequence is a parameter sequence of at least one of the motion amplitude and motion rhythm of the avatar in the second image sequence; generating a predicted value of the second image sequence according to the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence. For example, perform the steps as Figure 1 shown in step S2.

[0171] The training unit 73 is configured to train the avatar motion generation model according to the processing result. For example, perform the steps as Figure 1 shown in step S3.

[0172] The training device of the avatar motion generation model of the present disclosure makes full use of the pose information in the image and the audio features in the speech to train the model, which can improve the realism of the actions generated by the avatar motion generation model and the matching with the speech.

[0173] Figure 8 A block diagram showing an action generation device for an avatar according to some embodiments of the present disclosure.

[0174] As Figure 8 shown, the training device 8 of the avatar motion generation model includes an acquisition unit 81, a processing unit 82, and a training unit 83.

[0175] The acquisition unit 81 is configured to acquire the initial basic pose and speech of the avatar. For example, perform the steps as Figure 6Step S4 shown above.

[0176] An initial basic pose generation unit 82 is configured to use an avatar motion generation model to generate a basic pose of an image sequence corresponding to speech according to an initial basic pose and speech, where one image sequence corresponds to one basic pose, and the avatar motion generation model is trained by an avatar motion generation model training device according to any embodiment of the present disclosure, for example, performing steps such as Figure 6 Step S5 shown above.

[0177] A pose adjustment sequence generation unit 83 is configured to use an avatar motion generation model to generate a pose adjustment sequence of an image sequence according to speech, where the pose adjustment sequence is a parameter sequence of at least one of the motion amplitude and motion rhythm of the avatar in the image sequence, for example, performing steps such as Figure 6 Step S6 shown above.

[0178] An image sequence generation unit 84 is configured to generate an image sequence as the motion of the avatar according to the basic pose and the pose adjustment sequence, for example, performing steps such as Figure 6 Step S7 shown above.

[0179] The avatar motion generation device of the present disclosure decomposes the speaker's motion into two aspects: a basic pose (such as clapping) and an amplitude rhythm (such as the position of the hand and the clapping speed when clapping), generates the basic pose and the pose adjustment sequence respectively, and then adjusts the pose of the avatar with the pose adjustment sequence, making full use of the pose information in the image and the audio features in the speech, which can improve the realism of the motion generated by the avatar motion generation model and the matching with the speech.

[0180] Figure 9 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.

[0181] As Figure 9 shown, the electronic device 9 includes a memory 91; and a processor 92 coupled to the memory 91. The memory 91 is used to store instructions corresponding to an embodiment of the avatar motion generation model training method or an embodiment of the avatar motion generation method. The processor 92 is configured to execute the avatar motion generation model training method according to any embodiment of the present disclosure based on the instructions stored in the memory 91, or execute the avatar motion generation method according to any embodiment of the present disclosure.

[0182] Figure 10 A block diagram of a computer system for implementing some embodiments of the present disclosure is shown.

[0183] As Figure 10As shown, the computer system 100 may be embodied in the form of a general-purpose computing device. The computer system 100 includes a memory 1010, a processor 1020, and a bus 1000 that connects different system components.

[0184] The memory 1010 may include, for example, a system memory, a non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs. The system memory may include a volatile storage medium, such as a random access memory (RAM) and / or a cache memory. The non-volatile storage medium stores, for example, instructions for corresponding embodiments that execute at least one of the virtual avatar action generation model training method or the virtual avatar action generation method. The non-volatile storage medium includes, but is not limited to, a disk memory, an optical memory, a flash memory, etc.

[0185] The processor 1020 may be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, or discrete hardware components such as transistors. Accordingly, each unit such as a judgment unit and a determination unit may be implemented by a central processing unit (CPU) running instructions for executing corresponding steps in the memory, or may be implemented by a dedicated circuit for executing the corresponding steps.

[0186] The bus 1000 may use any bus structure among a variety of bus structures. For example, the bus structure includes, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus.

[0187] The computer system 100 may further include an input / output interface 1030, a network interface 1040, a storage interface 1050, etc. These interfaces 1030, 1040, 1050, and the memory 1010 and the processor 1020 may be connected via the bus 1000. The input / output interface 1030 provides a connection interface for input / output devices such as a display, a mouse, and a keyboard. The network interface 1040 provides a connection interface for various networking devices. The storage interface 1050 provides a connection interface for external storage devices such as a floppy disk, a USB flash drive, and an SD card.

[0188] Here, various aspects of the present disclosure have been described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, and computer program products according to the embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of the blocks, can be implemented by computer-readable program instructions.

[0189] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable device to produce a machine such that the instructions executed by the processor create means for implementing the functions specified in one or more boxes in the flowchart and / or block diagram.

[0190] These computer-readable program instructions may also be stored in a computer-readable memory, which instructions cause the computer to operate in a particular manner, thereby producing a manufacture including instructions for implementing the functions specified in one or more boxes in the flowchart and / or block diagram.

[0191] The present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.

[0192] Through the virtual avatar action generation model training method, virtual avatar action generation method, device, electronic device, and computer-readable storage medium in the above embodiments, the pose information in the image and the audio features in the speech are fully utilized, which can improve the realism of the actions of the generated virtual avatar and the matching with the speech.

[0193] So far, the training method of the virtual avatar action generation model, the virtual avatar action generation method, device, electronic device, and computer-readable storage medium according to the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.

Claims

1. A training method for a virtual character action generation model, comprising: For each first image sequence of the virtual avatar, obtain the ground truth of the first image sequence, the ground truth of the second image sequence corresponding to the first image sequence, and the speech corresponding to the second image sequence, where in all the images of the first image sequence, the virtual avatar corresponds to the same basic pose, and in all the images of the second image sequence, the virtual avatar corresponds to the same basic pose, and the second image sequence is the next image sequence of the first image sequence; Use the virtual avatar action generation model to process the ground truth of the first image sequence, the ground truth of the second image sequence, and the speech to obtain a processing result, including generating a predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence; generating a predicted value of the pose adjustment sequence of the second image sequence according to the speech, where the pose adjustment sequence is a parameter sequence of at least one of the action amplitude and action rhythm of the virtual avatar in the second image sequence; generating a predicted value of the second image sequence according to the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence, where the processing result includes the predicted value of the second image sequence; training the virtual avatar action generation model according to the processing result, where generating the predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence includes generating a posterior distribution of the latent variable according to the ground truth of the first image sequence and the ground truth of the second image sequence, where the latent variable is a random variable, and the posterior distribution of the latent variable is related to the basic pose of the second image sequence and is independent of both the action amplitude and action rhythm of the virtual avatar; generating the predicted value of the basic pose of the second image sequence according to the posterior distribution of the latent variable.

2. The training method for a virtual character action generation model according to claim 1, wherein, Generating the posterior distribution of the latent variable according to the ground truth of the first image sequence and the ground truth of the second image sequence includes: Calculating a first intermediate variable according to the ground truth of the first image sequence and the ground truth of the second image sequence; Generating a posterior distribution of the latent variable with the first intermediate variable as a sample as the posterior distribution of the latent variable.

3. The training method for a virtual character action generation model according to claim 2, wherein, Calculating the first intermediate variable according to the ground truth of the first image sequence and the ground truth of the second image sequence includes: Using a first encoder to calculate the encoding of the ground truth of the first image sequence and the encoding of the ground truth of the second image sequence; Calculating a first intermediate variable according to the encoding of the ground truth of the first image sequence and the encoding of the ground truth of the second image sequence, where the first intermediate variable includes information about the change from the basic pose of the first image sequence to the basic pose of the second image sequence.

4. The training method for a virtual character action generation model according to claim 3, wherein, Generating a posterior distribution of the latent variable with the first intermediate variable as a sample as the posterior distribution of the latent variable includes: Using a second encoder to calculate the latent space encoding of the first intermediate variable as the posterior distribution of the latent variable.

5. The training method for a virtual character action generation model according to claim 4, wherein, Generating the predicted value of the basic pose of the second image sequence according to the posterior distribution of the latent variable includes: Determining the value of the latent variable according to the posterior distribution of the latent variable; Generating the predicted value of the basic pose of the second image sequence according to the value of the latent variable.

6. The training method for a virtual character action generation model according to claim 5, wherein, Determining the value of the latent variable according to the posterior distribution of the latent variable with the first intermediate variable as a sample includes: Determine the value of the latent variable by sampling the posterior distribution of the latent variable.

7. The training method for a virtual character action generation model according to claim 5, wherein, Generating a predicted value of the base pose of the second image sequence according to the value of the latent variable includes: Using a second decoder corresponding to the second encoder to calculate the decoding result of the value of the latent variable; Using a first decoder corresponding to the first encoder, and generating a predicted value of the base pose of the second image sequence according to the decoding result of the value of the latent variable.

8. The training method for a virtual character action generation model according to claim 1, wherein, The posterior distribution of the latent variable is the first distribution or the second distribution; When there is no change from the base pose of the first image sequence to the base pose of the second image sequence, the posterior distribution of the latent variable is the first distribution; When there is a change from the base pose of the first image sequence to the base pose of the second image sequence, the posterior distribution of the latent variable is the second distribution, and the second distribution is different from the first distribution.

9. The training method for a virtual character action generation model according to claim 8, wherein, Training the virtual character action generation model according to the processing result includes: Calculating a first loss function according to the posterior distribution of the latent variable, the preset prior distribution of the latent variable, and the speech; Training the virtual character action generation model according to the first loss function.

10. The training method for a virtual character action generation model according to claim 9, wherein, Calculating the first loss function according to the posterior distribution of the latent variable, the preset prior distribution of the latent variable, and the speech includes: Predicting whether there is a change from the base pose of the first image sequence to the base pose of the second image sequence according to the speech, and obtaining a prediction result; When the prediction result is that there is no change from the base pose of the first image sequence to the base pose of the second image sequence, calculating the first loss function according to the expectation and variance of the first distribution; When the prediction result is that there is a change from the base pose of the first image sequence to the base pose of the second image sequence, calculating the first loss function according to the information entropy of the prior distribution of the latent variable and the second distribution.

11. The training method of the virtual character motion generation model according to claim 10, wherein, Predicting whether there is a change from the base pose of the first image sequence to the base pose of the second image sequence, and obtaining a prediction result includes: Converting the speech into text; Predicting whether there is a change from the base pose of the first image sequence to the base pose of the second image sequence according to the keywords in the text, and obtaining a prediction result.

12. The training method of the virtual character motion generation model according to claim 1, wherein, Generating a predicted value of the pose adjustment sequence of the second image sequence according to the speech includes: Extracting the mel cepstral coefficients of the speech; Generating a predicted value of the pose adjustment sequence of the second image sequence according to the mel cepstral coefficients of the speech.

13. The training method of the virtual character motion generation model according to claim 1, wherein, Training the virtual character action generation model according to the processing result includes: Calculating a second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence; Training the virtual character action generation model according to the second loss function.

14. The training method of the virtual character motion generation model according to claim 13, wherein, Calculating the second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence includes: Calculating the second loss function according to the ground truth of the second image sequence, the mean value of the ground truth of the second image sequence, and the predicted value of the pose adjustment sequence of the second image sequence.

15. The training method of the virtual character motion generation model according to claim 14, wherein, Calculating the second loss function according to the mean value of the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence includes: Calculate a second intermediate variable according to the difference between the ground truth of the second image sequence and the mean of the ground truth of the second image sequence, where the second intermediate variable is related to at least one of the movement amplitude and movement rhythm of the avatar, and is independent of the base pose of the avatar; Calculate a second loss function according to the difference between the predicted value of the pose adjustment sequence of the second image sequence and the second intermediate variable.

16. The training method of the virtual character motion generation model according to claim 1, wherein, The training of the avatar motion generation model according to the processing result includes: Calculate a third loss function according to the ground truth of the second image sequence and the predicted value of the second image sequence; Train the avatar motion generation model according to the third loss function.

17. The training method of the virtual character motion generation model according to claim 1, wherein, The training of the avatar motion generation model according to the processing result includes: Calculate a first loss function according to the posterior distribution of the latent variable, the preset prior distribution of the latent variable, and the speech; Calculate a second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence; Calculate a third loss function according to the ground truth of the second image sequence and the predicted value of the second image sequence; Train the avatar motion generation model according to the weighted sum of the first loss function, the second loss function, and the third loss function.

18. The training method of the virtual character motion generation model according to claim 3, wherein, The processing of the ground truth of the first image sequence, the ground truth of the second image sequence, and the speech by using the avatar motion generation model to obtain a processing result includes: Use the first decoder corresponding to the first encoder to calculate the decoding result of the encoding of the ground truth of the first image sequence and the decoding result of the encoding of the ground truth of the second image sequence; The training of the avatar motion generation model according to the processing result includes: Calculate a fourth loss function according to the decoding result of the encoding of the ground truth of the first image sequence, the decoding result of the encoding of the ground truth of the second image sequence, the ground truth of the first image sequence, and the ground truth of the second image sequence; Train the avatar motion generation model according to the fourth loss function.

19. The training method of the virtual image motion generation model according to claim 18, wherein, The calculation of the fourth loss function according to the decoding result of the encoding of the ground truth of the first image sequence, the decoding result of the encoding of the ground truth of the second image sequence, the ground truth of the first image sequence, and the ground truth of the second image sequence includes: Calculate the fourth loss function according to the norm of the difference between the decoding result of the encoding of the ground truth of the first image sequence and the ground truth of the first image sequence, and the norm of the difference between the decoding result of the encoding of the ground truth of the second image sequence and the ground truth of the second image sequence.

20. The training method of the virtual image motion generation model according to claim 19, wherein, The training of the avatar motion generation model according to the fourth loss function includes: Calculate a first loss function according to the posterior distribution of the latent variable, the preset prior distribution of the latent variable, and the speech; Calculate a second loss function according to the ground truth of the second image sequence and the predicted value of the pose adjustment sequence of the second image sequence; Calculate a third loss function according to the ground truth of the second image sequence and the predicted value of the second image sequence; Train the avatar motion generation model according to the weighted sum of the first loss function, the second loss function, the third loss function, and the fourth loss function.

21. The training method of the virtual image motion generation model according to any one of claims 1-20, wherein, The speech is the speech uttered by the avatar when executing the second image sequence, and the predicted value of the base pose of the second image sequence includes the predicted value of the base pose of at least one of the head, upper limbs, lower limbs, and torso of the avatar in the second image sequence.

22. The training method of the virtual image motion generation model according to any one of claims 1-20, wherein, Generating a predicted value of the second image sequence based on the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence includes: Based on the predicted value of the basic pose of the second image sequence, using the predicted value of the pose adjustment sequence of the second image sequence to adjust at least one of the action amplitude or action rhythm of the virtual image, and taking the adjustment result as the predicted value of the second image sequence.

23. The training method of the virtual image motion generation model according to any one of claims 1-20, wherein, Generating a predicted value of the second image sequence based on the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence includes: Generating a sequence of predicted values of the basic pose of the second image sequence according to the predicted value of the basic pose of the second image sequence, wherein the basic poses of the virtual image in the sequence of predicted values of the basic pose of the second image sequence are all the predicted values of the basic pose; Obtaining the predicted value of the second image sequence by adding the sequence of predicted values of the basic pose of the second image sequence to the predicted value of the pose adjustment sequence of the second image sequence.

24. A virtual image motion generation method, comprising: Obtaining the initial basic pose and voice of the virtual image; Using the virtual image action generation model, generating the basic pose of the image sequence corresponding to the voice according to the initial basic pose and the voice, wherein one image sequence corresponds to one basic pose, and the virtual image action generation model is trained according to the virtual image action generation model training method described in any one of claims 1-23; Using the virtual image action generation model, generating a pose adjustment sequence of the image sequence according to the voice, wherein the pose adjustment sequence is a parameter sequence of at least one of the action amplitude and action rhythm of the virtual image in the image sequence; Generating an image sequence according to the basic pose and the pose adjustment sequence as the action of the virtual image.

25. The virtual image motion generation method according to claim 24, wherein, The generating an image sequence according to the basic pose and the pose adjustment sequence as the action of the virtual image includes: Based on the basic pose of the image sequence, using the pose adjustment sequence of the image sequence to adjust at least one of the action amplitude or action rhythm of the virtual image, and taking the adjustment result as the image sequence.

26. The virtual image motion generation method according to claim 24, wherein, The generating an image sequence according to the basic pose and the pose adjustment sequence as the action of the virtual image includes: Generating a sequence of basic poses of the image sequence according to the basic pose of the image sequence, wherein the virtual images in the sequence of basic poses of the image sequence are all the same basic pose; Generating an image sequence by adding the sequence of basic poses of the image sequence to the pose adjustment sequence of the image sequence.

27. The virtual image motion generation method according to claim 24, wherein, The obtaining the initial basic pose and the voice corresponding to the image sequence includes: Dividing the voice into multiple voice segments, wherein the multiple voice segments are arranged in time sequence; Taking the obtained initial basic pose as the initial basic pose of the first voice segment; For other voice segments except the first voice segment, taking the basic pose of the previous voice segment of the current voice segment generated by using the virtual image action generation model as the initial basic pose of the current voice segment.

28. A training device for a virtual image action generation model, comprising: An acquisition unit, configured to acquire, for each first image sequence of a virtual avatar, the ground truth of the first image sequence, the ground truth of the second image sequence corresponding to the first image sequence, and the voice corresponding to the second image sequence, where, in all images of the first image sequence, the virtual avatar corresponds to the same basic pose, in all images of the second image sequence, the virtual avatar corresponds to the same basic pose, and the second image sequence is the next image sequence of the first image sequence; A processing unit, configured to process the ground truth of the first image sequence, the ground truth of the second image sequence, and the voice by using a virtual avatar motion generation model to obtain a processing result, including generating a predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence; generating a predicted value of the pose adjustment sequence of the second image sequence according to the voice, where the pose adjustment sequence is a parameter sequence of at least one of the motion amplitude and motion rhythm of the virtual avatar in the second image sequence; generating a predicted value of the second image sequence according to the predicted value of the basic pose of the second image sequence and the predicted value of the pose adjustment sequence, where the processing result includes the predicted value of the second image sequence; A training unit, configured to train the virtual avatar motion generation model according to the processing result, where, the generating a predicted value of the basic pose of the second image sequence according to the ground truth of the first image sequence and the ground truth of the second image sequence includes generating a posterior distribution of a latent variable according to the ground truth of the first image sequence and the ground truth of the second image sequence, where the latent variable is a random variable, and the posterior distribution of the latent variable is related to the basic pose of the second image sequence and is irrelevant to both the motion amplitude and motion rhythm of the virtual avatar; generating a predicted value of the basic pose of the second image sequence according to the posterior distribution of the latent variable.

29. A virtual image action generation device, comprising: An acquisition unit, configured to acquire the initial basic pose and voice of the virtual avatar; An initial basic pose generation unit, configured to generate the basic pose of the image sequence corresponding to the voice according to the initial basic pose and the voice by using a virtual avatar motion generation model, where one image sequence corresponds to one basic pose, and the virtual avatar motion generation model is trained by the virtual avatar motion generation model training device according to claim 28; A pose adjustment sequence generation unit, configured to generate the pose adjustment sequence of the image sequence according to the voice by using a virtual avatar motion generation model, where the pose adjustment sequence is a parameter sequence of at least one of the motion amplitude and motion rhythm of the virtual avatar in the image sequence; An image sequence generation unit, configured to generate an image sequence according to the basic pose and the pose adjustment sequence as the motion of the virtual avatar.

30. An electronic device, comprising: A memory; and A processor coupled to the memory, the processor being configured to execute the virtual avatar motion generation model training method according to any one of claims 1 to 23 or execute the virtual avatar motion generation method according to any one of claims 24 to 27 based on instructions stored in the memory.

31. A computer-readable storage medium, on which computer program instructions are stored, and when the instructions are executed by a processor, the virtual image action generation model training method according to any one of claims 1 to 23, or the virtual image action generation method according to any one of claims 24 to 27 is implemented.

Citation Information

Patent Citations

  • Human motion capture and virtual animation generation method based on deep learning

    CN110033505A

  • Method and system for driving human face animation through real-time voice

    CN110751708A