Method, device and equipment for driving 3D (three-dimensional) digital human eye expression and head action by voice

By extracting features from speech signals and using the head and eyeball generation model and time-sequence Transformer model to generate head and eyeball movements synchronized with the speech signal, the problem of unnatural eye animation in the prior art is solved, and diversified 3D digital human eye movements and head movements are achieved.

CN120495478APending Publication Date: 2025-08-15XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510349608.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art in voice-driven 3D digital human animation ignores the animation of eye gaze, resulting in the synthetic 3D speaking face appearing unnatural and stiff, and requires additional multi-modal prompts that it is impossible to generate diverse eye and head movements from the voice signal only.

Method used

By extracting embedded features from the speech signal, a pre-trained head and eyeball generation model is used to generate potential spatial representations, and a time-series Transformer model is used for cross-modal alignment, generating a sequence of head posture and eyeball rotation parameters synchronized with the speech signal, and finally generating a 3D digital human animation containing head movements, eye movements and facial expressions.

Benefits of technology

It realizes automatic and diversified generation of 3D digital eye movements, head movements and facial expressions from only the voice signals, improving the naturalness and reality of the animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495478A_ABST
    Figure CN120495478A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method, device and equipment for driving 3D digital human eye expression and head action through voice. The method comprises the following steps: extracting a voice embedding feature from an input voice signal; respectively inputting the voice embedded features into a pre-trained head motion generation model and an eyeball motion generation model to obtain a potential space representation of head motion and a potential space representation of eyeball motion; performing cross-modal alignment on the potential space representations of the head movement and the eyeball movement by using a time sequence Transform model, and generating a head posture parameter sequence and an eyeball rotation parameter sequence which are synchronous with the voice signal; and according to the head posture parameter sequence, the eyeball rotation parameter sequence and a facial expression parameter sequence generated based on voice driving, generating a 3D digital human animation including head actions, eyeball movements and facial expressions. According to the technical scheme provided by the embodiment of the invention, 3D digital human eye movement, head movement and facial expression can be automatically and diversely generated only from the voice signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device and apparatus for driving the eye and head movements of a 3D digital human through voice. Background Art

[0002] Despite significant progress in speech-driven 3D digital human animation, existing methods primarily focus on mouth movements, head movements, and emotional facial expressions, neglecting eye gaze animation. This results in synthesized 3D speaking faces often gazing straight ahead with static, unnatural, and stiff eyes. While some methods have achieved eye gaze animation, they often require additional multimodal cues as input, such as scene images, environmental information, or director's scripts, and are unable to generate diverse eye gaze and head movements from speech signals alone. Summary of the Invention

[0003] The embodiments of the present application provide a method, apparatus, and device for voice-driven 3D digital human eye and head movements, thereby automatically and diversely generating 3D digital human eye movements, head movements, and facial expressions from voice signals alone, at least to a certain extent.

[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.

[0005] According to one aspect of an embodiment of the present application, a method for voice-driven 3D digital human eye and head movements is provided, comprising:

[0006] Extract speech embedding features from the input speech signal;

[0007] Inputting the speech embedding features into a pre-trained head movement generative model and an eye movement generative model, respectively, wherein the head movement generative model is constructed based on a variational autoencoder to generate a latent space representation of head movement, and the eye movement generative model is constructed based on a vector quantized variational autoencoder to generate a latent space representation of eye movement;

[0008] cross-modally aligning the latent spatial representations of the head movement and the eye movement using a temporal Transformer model to generate a head posture parameter sequence and an eye rotation parameter sequence synchronized with the speech signal;

[0009] A 3D digital human animation including head movements, eye movements and facial expressions is generated according to the head posture parameter sequence, the eye rotation parameter sequence and the facial expression parameter sequence generated based on voice drive.

[0010] According to one aspect of an embodiment of the present application, a voice-driven 3D digital human eye and head movement device is provided, comprising:

[0011] An extraction module, used to extract speech embedding features from the input speech signal;

[0012] a first generation module, configured to input the speech embedding features into a pre-trained head movement generation model and an eye movement generation model, respectively, wherein the head movement generation model is constructed based on a variational autoencoder to generate a latent space representation of head movement, and the eye movement generation model is constructed based on a vector quantized variational autoencoder to generate a latent space representation of eye movement;

[0013] A second generation module is configured to perform cross-modal alignment on the latent space representations of the head movement and the eye movement using a temporal Transformer model to generate a head posture parameter sequence and an eye rotation parameter sequence synchronized with the speech signal;

[0014] The processing module is used to generate a 3D digital human animation including head movements, eye movements and facial expressions according to the head posture parameter sequence, the eye rotation parameter sequence and the facial expression parameter sequence generated based on voice drive.

[0015] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for driving the eye and head movements of a 3D digital human via voice is implemented as described in the above embodiment.

[0016] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the voice-driven 3D digital human eye and head movement method as described in the above embodiments.

[0017] According to one aspect of an embodiment of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice-driven 3D digital human eye and head movement method provided in the above-described embodiments.

[0018] In the technical solutions provided in some embodiments of the present application, speech embedding features are extracted from the input speech signal and the speech embedding features are respectively input into a pre-trained head movement generation model and an eye movement generation model, wherein the head movement generation model is constructed based on a variational autoencoder to generate a latent space representation of head movement, and the eye movement generation model is constructed based on a vector quantized variational autoencoder to generate a latent space representation of eye movement. Then, a temporal Transformer model is used to perform cross-modal alignment on the latent space representations of head movement and eye movement to generate a head posture parameter sequence and an eye rotation parameter sequence synchronized with the speech signal. Then, based on the head posture parameter sequence, the eye rotation parameter sequence and the facial expression parameter sequence generated based on speech drive, a 3D digital human animation including head movements, eye movements and facial expressions is generated. In this way, 3D digital human eye movements, head movements and facial expressions can be automatically and diversely generated from only the speech signal.

[0019] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:

[0021] Figure 1 A schematic diagram of a process for a method of driving 3D digital human eye and head movements with voice according to an embodiment of the present application is shown;

[0022] Figure 2 A schematic diagram of a process for constructing a training data set according to an embodiment of the present application is shown;

[0023] Figure 3 A schematic diagram of a model pre-training process according to an embodiment of the present application is shown;

[0024] Figure 4 A block diagram of a voice-driven 3D digital human eye and head movement device according to an embodiment of the present application is shown;

[0025] Figure 5 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0026] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.

[0027] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0028] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0029] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0030] Figure 1 A flow chart of a method for voice-driven 3D digital human eye and head movements according to an embodiment of the present application is shown.

[0031] The method can be applied to a terminal device or a server, wherein the terminal device may include but is not limited to one or more of a smart phone, a tablet computer, a portable computer, and a desktop computer; the server may be a physical server or a cloud server.

[0032] like Figure 1 As shown, the voice-driven 3D digital human eye and head movement method includes at least steps S110 to S140, which are described in detail as follows (the following description takes the application of the method to a terminal device as an example, hereinafter referred to as "terminal"):

[0033] In step S110 , speech embedding features are extracted from the input speech signal.

[0034] In this embodiment, the terminal can receive the user's input voice signal through its own configured audio acquisition device (such as a microphone, etc.). Then, the terminal can use a speech encoder (such as wav2vec2.0, etc.) to extract time series features from the voice signal, thereby obtaining corresponding voice embedding features to represent the semantic and prosodic information of the voice.

[0035] In step S120, the speech embedding features are respectively input into the pre-trained head movement generation model and the eye movement generation model, wherein the head movement generation model is constructed based on a variational autoencoder to generate a latent space representation of head movement, and the eye movement generation model is constructed based on a vector quantized variational autoencoder to generate a latent space representation of eye movement.

[0036] In this embodiment, after acquiring the speech embedding features, the terminal can input them into the pre-trained head movement generation model and eye movement generation model respectively. Specifically, the embodiment of the present application provides a cascaded VAE-VQVAE model, which is built based on variational auto-encoders (VAE). The head movement generation model can generate a latent space representation of the head movement of a 3D digital human based on the input speech embedding features.

[0037] The eye movement generation model is built based on the Vector Quantized-Variational Autoencoder (VQVAE). It introduces a discrete codebook based on VAE and improves diversity by quantizing the latent space. This eye movement generation model can generate a latent space representation of the eye movement of a 3D digital human based on the input speech embedding features.

[0038] After constructing the head movement generation model and the eye movement generation model, they need to be pre-trained so that they can complete the corresponding prediction function. In some embodiments of the present application, pre-training the head movement generation model and the eye movement generation model includes:

[0039] Constructing a training dataset, wherein the training dataset includes audio-mesh sequences involving eye gaze, head and facial movements;

[0040] The head movement generation model and the eye movement generation model are pre-trained based on the training data set.

[0041] In this embodiment, the terminal can pre-build a training data set, which can include an audio-mesh sequence designed for eye gaze, head and facial movements. The audio-mesh sequence is a time-series data set containing speech audio, 3D head posture, eye rotation and facial expression parameters. After obtaining the training data set, the terminal can use it to train the head movement generation model and the eye movement generation model so that the two can complete the corresponding prediction functions.

[0042] In one embodiment, constructing a training dataset includes:

[0043] Filtering a video to be processed containing a clear human eye area from the selected 2D audio-visual data;

[0044] Performing 3D face reconstruction on the video to be processed frame by frame to obtain corresponding FLAME parameters, which include facial motion parameters and head motion parameters;

[0045] Project the eyeball center of the 3D face model onto a 2D plane and fit it to the human eyes in the video frame to obtain the eye movement parameters;

[0046] Temporal convolution smoothing processing is performed according to the head movement parameters, eye movement parameters and facial movement parameters to obtain a continuous audio-grid sequence.

[0047] In this embodiment, Figure 2 As shown, considering that 2D audio-visual datasets are easier to access and have a wider coverage, a training dataset can be constructed by reconstructing 3D face meshes and 3D sight lines from 2D videos. The terminal can obtain the selected 2D audio-visual data from its own storage space or the Internet. After obtaining the 2D audio-visual data to be selected, the terminal can filter it to obtain a video to be processed containing a clear human eye area. It can be understood that the visibility of the subject's eyes in the video is crucial for the subsequent 3D eye gaze fitting. Therefore, if the subject's eyes are blocked, the terminal can exclude the video. For videos where the eyes are invisible due to head rotation, the terminal can preliminarily exclude them by detecting the significant difference in the distance between the left and right eye corners and the side of the face, and then exclude videos with obstructions such as sunglasses (which essentially cover the eyes) through manual screening.

[0048] Then, in order to generate 3D facial motion and 3D head motion frame by frame, the terminal can perform 3D face reconstruction on the processed video frame by frame. In one example, in order to ensure the highest quality, the 3D face reconstruction network adopted in the embodiment of the present application integrates four methods: EMOCA, SPECTRE, DECA and MICA. The above four methods all use FLAME as the basis for 3D face representation. At the output end, DECA provides overall head rotation information to form 3D head motion; MICA outputs facial shape vectors, EMOCA is responsible for expression vectors, and DECA also provides mandibular posture information, which together constitute 3D facial motion. SPECTRE is mainly used in the network training stage to enhance the performance of lip pronunciation through lip reading loss. In this way, through the above-mentioned 3D face reconstruction network, FLAME parameters including facial motion parameters and head motion parameters can be obtained.

[0049] In order to obtain eye movement data, the terminal can project the eye center of the reconstructed 3D face model onto a 2D plane and fit it to the human eyes in the video frame to obtain eye movement parameters. After obtaining the above data, it needs to be optimized. Specifically, when processing video frames containing blinking movements, the eye posture can be kept unchanged based on the last known eye-opening state. Because during blinking, the iris position is difficult to accurately detect. In this way, the continuity and naturalness of facial animation can be guaranteed, and the problem of sudden eye movement that may occur during blinking can be effectively avoided. Subsequently, temporal convolution smoothing is applied to the above three types of data (i.e., head movement parameters, facial movement parameters, and eye movement parameters) to reduce unnecessary jitter. They are then integrated into a FLAME model to obtain a continuous audio-grid sequence, achieving a smoother and more realistic animation effect.

[0050] Please continue to refer to Figure 1 In step S130, a temporal Transformer model is used to perform cross-modal alignment on the latent space representations of the head movement and the eye movement to generate a head posture parameter sequence and an eye rotation parameter sequence synchronized with the speech signal.

[0051] In this embodiment, the temporal Transformer model is a model based on the Transformer architecture, which is specifically used to process sequence data and can capture long-term dependencies between elements in the sequence. In the present application, the temporal Transformer model is used to perform cross-modal alignment of the latent space representations of head movement and eye movement. It should be understood that cross-modal alignment refers to aligning data of different modalities (such as speech signals and action signals) in the feature space so that the data of different modalities can correspond to and fuse with each other. In the present application, cross-modal alignment is used to align the speech signal with the latent space representations of head movement and eye movement to generate a sequence of action parameters synchronized with the speech signal. In one example, the temporal Transformer model includes an encoder and a decoder. The encoder is used to process the speech embedding feature sequence, and the decoder is used to generate a head posture parameter sequence and an eye rotation parameter sequence. The speech embedding feature sequence is encoded by the encoder to extract high-level semantic features. At the same time, the latent space representations of head movement and eye movement are fused with the speech features, and the self-attention mechanism and cross-modal attention mechanism in the Transformer model are used to capture the complex relationship between the speech signal and the action signal to achieve cross-modal alignment.

[0052] In this way, the temporal Transformer model can effectively capture the complex relationship between speech and motion signals, achieving high-precision cross-modal alignment and ensuring that the generated motion parameter sequence is highly synchronized with the speech signal. Furthermore, through cross-modal alignment, the model can generate a rich variety of head poses and eye rotations, enhancing the naturalness and realism of 3D digital human animation.

[0053] Please continue to refer to Figure 1 In step S140, a 3D digital human animation including head movements, eye movements and facial expressions is generated according to the head posture parameter sequence, the eye rotation parameter sequence and the facial expression parameter sequence generated based on voice drive.

[0054] In this embodiment, the terminal can adopt a voice-driven facial expression parameter generation method to generate a corresponding facial expression parameter sequence based on voice embedded features. It should be noted that the voice-driven facial expression generation method can adopt an existing model, which will not be repeated in this application.

[0055] Then, the terminal can combine the head posture parameter sequence, the eye rotation parameter sequence and the facial expression parameter sequence to generate a complete 3D digital human animation including head movements, eye movements and facial expressions.

[0056] So, based on Figure 1In the embodiment shown, speech embedding features are extracted from the input speech signal and input into a pre-trained head movement generation model and an eye movement generation model, respectively. The head movement generation model is constructed based on a variational autoencoder to generate a latent space representation of head movement, and the eye movement generation model is constructed based on a vector quantized variational autoencoder to generate a latent space representation of eye movement. Then, a temporal Transformer model is used to perform cross-modal alignment on the latent space representations of head movement and eye movement to generate a head posture parameter sequence and an eye rotation parameter sequence synchronized with the speech signal. Then, based on the head posture parameter sequence, the eye rotation parameter sequence, and the facial expression parameter sequence generated based on speech drive, a 3D digital human animation including head movements, eye movements, and facial expressions is generated. In this way, 3D digital human eye movements, head movements, and facial expressions can be automatically and diversely generated from only the speech signal.

[0057] In some embodiments of the present application, the head movement generation model and the eye movement generation model are pre-trained based on the training dataset, including pre-training of a vector quantized variational autoencoder and joint training of a variational autoencoder and a vector quantized variational autoencoder;

[0058] Among them, the pre-training of the vector quantized variational autoencoder includes:

[0059] Input the true value of the eye movement parameter into the encoder of the vector quantized variational autoencoder to generate a feature vector;

[0060] Mapping the feature vector to the nearest entry in the codebook using an element-by-element quantization function to obtain a quantized feature vector;

[0061] Reconstructing the quantized feature vector into an eye movement parameter sequence using a decoder of a vector quantized variational autoencoder, and optimizing the parameters of the vector quantized variational autoencoder by minimizing the reconstruction loss and the codebook alignment loss;

[0062] Joint training of variational autoencoders and vector quantized variational autoencoders, including:

[0063] The true value of the head motion parameters and the speech embedding features are input into the encoder of the variational autoencoder to generate the corresponding Gaussian distributed latent code;

[0064] Based on the latent code and the speech embedding features, generating a corresponding head movement parameter sequence through a decoder of a variational autoencoder;

[0065] Taking the head movement parameter sequence as a condition, generating an eye movement parameter sequence through autoregression of a temporal Transformer model;

[0066] The reconstruction loss, speed loss, KL divergence loss and feature regularization loss of the variational autoencoder are jointly optimized.

[0067] In this embodiment, the ultimate goal of the method provided in this application is to collectively synthesize eye movements, blinks, head movements, and facial movements from speech within the unified FLAME framework for 3D facial representation. For facial movement generation, existing speech-driven 3D face methods can be directly used as generators. For eye and head movements, a cascaded VAE-VQVAE model and a temporal Transformer model are designed to predict and generate them.

[0068] Generative models such as VAEs and VQVAEs are commonly used to model the weak correlations between speech signals and non-verbal cues in conversations, such as head pose, hand gestures, and body posture. These models are able to learn compact representations of non-verbal motion data. By mapping speech into a compact latent space, capturing the weak correlations between speech and motion is significantly reduced, thereby improving the quality of motion synthesis.

[0069] However, there is a key difference between VAE and VQVAE in terms of motion diversity. VAE tends to output the mean of a Gaussian distribution, which limits the diversity of the output. VQVAE uses a discrete codebook to capture a diverse set of potential representations, allowing a wider range of possible outputs. Since the rotation range of the head usually exceeds the rotation range of the eyeballs, the change in head movement is usually higher than the change in eye gaze movement. VAE is able to produce head movements with sufficient diversity, while VQVAE may sometimes produce overly diverse and exaggerated head movements. In contrast, it is difficult for VAE to produce sufficiently diverse eye gaze movements, while VQVAE effectively achieves this. Therefore, the embodiments of the present application use VAE to learn the head movement latent space and use VQVAE to learn the eye gaze movement latent space.

[0070] Therefore, in the method provided in the embodiment of the present application, the head movement generation model and the eye movement generation model are pre-trained based on the training data set, including pre-training of the vector quantized variational autoencoder and joint training of the variational autoencoder and the vector quantized variational autoencoder.

[0071] like Figure 3 As shown, the speech signal is converted to 3D eye movement in one framework. (left and right eye) and 3D head movement The model training consists of two stages: first, a discrete latent space of eye movements is learned by pre-training a VQVAE, and then a VAE-based speech-to-head mapping and a VQVAE-based speech-to-eye movement mapping are jointly trained in an autoregressive manner.

[0072] In the first stage, we pre-train VQVAE to model the latent space of eye movements as a discrete codebook The codebook is obtained by In the second stage, conditional VAE is used to learn the Reconstructing head movements First, a speech encoder is used to extract speech embedding A from the input speech 1: , and then predict the head movement and the previously predicted eye movements As a condition, the speech is embedded in A 1: is mapped to the target motion code. This mapping is achieved through multimodal alignment and codebook The search is completed. Finally, the eye movement decoder is used to further decode the eye movement of the current frame t+1. And in the next round of autoregression, it is used as a frame of past motion. When using the model for inference, the true value is not used. and the head motion encoder in VAE, and the head latent code c in VAE is randomly sampled from a standard Gaussian.

[0073] In order to learn the discrete latent space of eye movement, the present embodiment constructs a discrete codebook To form a discrete latent space, any frame in the eye gaze sequence can be represented by the codebook term z k The VQVAE model is based on Transformer and contains three main components: motion encoder E VQ , Motion Decoder D VQ and codebook These components are in true value Pre-training is performed during the self-reconstruction process. Figure 3 As shown, First, it is encoded into a feature vector Then, we use the element-wise quantization function Q(·) to convert Quantized into feature vector Z q .

[0074] Specifically, this function will Each vector in is mapped to the codebook Recent entries in:

[0075]

[0076] Finally, Z q Decoded and reconstructed into

[0077]

[0078] The training loss of VQVAE includes reconstruction loss and two intermediate code-level losses:

[0079]

[0080] Among them, the first one is the motion reconstruction loss, and the last two are used to reduce the codebook and embedded features The codebook is updated based on the distance between them. sg(·) indicates the stop gradient operation, and β is a weighting factor used to balance the weights between different loss terms.

[0081] For the mapping of speech to head and speech to eye movement, a conditional variational autoencoder (VAE) model is used to model the mapping from speech to head movement. VAE Acceptance and Speech EmbeddingA 1: Spliced real movement As input. Encoder E VAE Output continuous latent code c, which follows a Gaussian distribution With a learned mean μ h and the learned variance The latent code c is further compared with A 1: Spliced and then input to the head motion decoder D VAE , predicting the head motion as All T-frame motion They are all generated through one reconstruction, and the formula is as follows:

[0082]

[0083] After the VAE model, an autoregressive model based on Vector Quantized Variational Autoencoder (VQVAE) is cascaded for eye gaze movement generation. Predicted head movement is used as a condition to guide the mapping from speech to eye gaze movements. Specifically, the Transformer cross-modal decoder D cross-modal To simulate three different modalities, namely speech audio, head movement and eye gaze movement. cross-modal Equipped with causal self-attention to track eye gaze movements in the past T frames In the context of learning the dependencies between each frame. In addition, D cross-modalIt is also equipped with a cross-modal attention mechanism that incorporates past eye gaze movements Head movement and speech embedding A 1: Alignment. Set the cross-attention key (K) and value (V) to the concatenation of the predicted head motion embedding and speech embedding, and set the query (Q) to the past eye gaze embedding.

[0084] Cross-modal decoder D cross-modal Output features as follows:

[0085]

[0086] in and are linear projection layers used to extract eye gaze motion embedding and head motion embedding, respectively.

[0087] It is further quantified by the following formula: and decoded by the pre-trained VQVAE decoder as

[0088] New predicted movement Used to update past movements in preparation for the next prediction.

[0089] Training the head motion encoder E VAE , head motion decoder D VAE , cross-modal decoder D cross-modal , projection layer and and part of the speech encoder, while maintaining the codebook and eye gaze decoder D VQ Unchanged. The training loss consists of four items: reconstruction loss Speed loss KL divergence loss and feature regularization loss

[0090] Reconstruction losses Measures the difference between predicted and true values for head and eye gaze:

[0091]

[0092] Speed loss The difference between the first derivative of the predicted value and the first derivative of the true value used to measure head and eye gaze:

[0093]

[0094] KL divergence loss Force a Gaussian distribution to predict head motion Close to standard Gaussian distribution Its simplified form is:

[0095]

[0096] Feature regularization loss Measuring predicted eye gaze movement characteristics and the quantized features in the codebook Deviation between:

[0097]

[0098] The final training loss function is as follows:

[0099]

[0100] where λ1 is set to 1e-4.

[0101] In some embodiments of the present application, the method further comprises:

[0102] Counting the blink frequency distribution of the training data set and generating random blink intervals based on Gaussian distribution;

[0103] The eye expression parameters of the 3D digital human are adjusted within the frame window of the random blink interval to generate a continuous blink animation.

[0104] In this embodiment, in order to generate more vivid animations, the method provided in the embodiment of the present application also realizes the generation of blinks. It should be understood that blinking is mainly affected by personal habits, environmental factors and other conditions. The terminal can perform statistical analysis on the training data set to obtain a suitable blink frequency that matches the human gaze animation. Specifically, the number of blinks in each video in the data set can be calculated, and then these counts can be converted into blinks per minute based on the duration of the video.

[0105] Blink detection can be performed by calculating the eye aspect ratio (EAR), defined as the ratio of the eye's height to its width. When the eyes are closed, the EAR approaches zero. Statistical analysis shows that the number of blinks per minute follows a Gaussian distribution with a mean of 40.10 and a variance of 15.42. When generating animations, the device first extracts the number of blinks per minute from this distribution and converts it into an interval between blinks. Subsequently, the expression parameters are adjusted within a five-frame window of each blink to create the blinking action.

[0106] In this way, the model trained by the above method can generate vivid and diverse eye movements and head movements that match the audio rhythm after audio is input into the model. Combined with the voice-driven 3D face model method, it can finally generate 3D digital human animation that includes eye movements, blinking, head movements and facial expressions.

[0107] The following describes an embodiment of the device of this application, which can be used to implement the voice-driven 3D digital human eye and head movement method described in the above-mentioned embodiments of this application. For details not disclosed in the device embodiment of this application, please refer to the embodiment of the voice-driven 3D digital human eye and head movement method described in the above-mentioned embodiment of this application.

[0108] Figure 4 A block diagram of a voice-driven 3D digital human eye and head movement device according to an embodiment of the present application is shown.

[0109] Reference Figure 4 As shown, according to one embodiment of the present application, a voice-driven 3D digital human eye and head movement device includes:

[0110] An extraction module, used to extract speech embedding features from the input speech signal;

[0111] a first generation module, configured to input the speech embedding features into a pre-trained head movement generation model and an eye movement generation model, respectively, wherein the head movement generation model is constructed based on a variational autoencoder to generate a latent space representation of head movement, and the eye movement generation model is constructed based on a vector quantized variational autoencoder to generate a latent space representation of eye movement;

[0112] A second generation module is configured to perform cross-modal alignment on the latent space representations of the head movement and the eye movement using a temporal Transformer model to generate a head posture parameter sequence and an eye rotation parameter sequence synchronized with the speech signal;

[0113] The processing module is used to generate a 3D digital human animation including head movements, eye movements and facial expressions according to the head posture parameter sequence, the eye rotation parameter sequence and the facial expression parameter sequence generated based on voice drive.

[0114] In one embodiment, the processing module is further configured to pre-train the head movement generation model and the eye movement generation model, including:

[0115] Constructing a training dataset, wherein the training dataset includes audio-mesh sequences involving eye gaze, head and facial movements;

[0116] The head movement generation model and the eye movement generation model are pre-trained based on the training data set.

[0117] In one embodiment, the processing module is further configured to construct a training data set, including:

[0118] Filtering a video to be processed containing a clear human eye area from the selected 2D audio-visual data;

[0119] Performing 3D face reconstruction on the video to be processed frame by frame to obtain corresponding FLAME parameters, which include facial motion parameters and head motion parameters;

[0120] Project the eyeball center of the 3D face model onto a 2D plane and fit it to the human eyes in the video frame to obtain the eye movement parameters;

[0121] Temporal convolution smoothing processing is performed according to the head movement parameters, eye movement parameters and facial movement parameters to obtain a continuous audio-grid sequence.

[0122] In one embodiment, the head movement generation model and the eye movement generation model are pre-trained based on the training dataset, including pre-training of a vector quantized variational autoencoder and joint training of the variational autoencoder and the vector quantized variational autoencoder;

[0123] Among them, the pre-training of the vector quantized variational autoencoder includes:

[0124] Input the true value of the eye movement parameter into the encoder of the vector quantized variational autoencoder to generate a feature vector;

[0125] Mapping the feature vector to the nearest entry in the codebook using an element-by-element quantization function to obtain a quantized feature vector;

[0126] Reconstructing the quantized feature vector into an eye movement parameter sequence using a decoder of a vector quantized variational autoencoder, and optimizing the parameters of the vector quantized variational autoencoder by minimizing the reconstruction loss and the codebook alignment loss;

[0127] Joint training of variational autoencoders and vector quantized variational autoencoders, including:

[0128] The true value of the head motion parameters and the speech embedding features are input into the encoder of the variational autoencoder to generate the corresponding Gaussian distributed latent code;

[0129] Based on the latent code and the speech embedding features, generating a corresponding head movement parameter sequence through a decoder of a variational autoencoder;

[0130] Taking the head movement parameter sequence as a condition, generating an eye movement parameter sequence through autoregression of a temporal Transformer model;

[0131] The reconstruction loss, speed loss, KL divergence loss and feature regularization loss of the variational autoencoder are jointly optimized.

[0132] In one embodiment, the processing module is further configured to:

[0133] Counting the blink frequency distribution of the training data set and generating random blink intervals based on Gaussian distribution;

[0134] The eye expression parameters of the 3D digital human are adjusted within the frame window of the random blink interval to generate a continuous blink animation.

[0135] Figure 5 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.

[0136] It should be noted that Figure 5 The computer system of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0137] like Figure 5 As shown, the computer system includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage part 508 to the random access memory (RAM) 503, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 503. The CPU 501, ROM 502 and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0138] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, and the like; an output section 507 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 508 including a hard disk; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. Removable media 511, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 510 as needed, so that computer programs read therefrom can be installed into the storage section 508 as needed.

[0139] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from a removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the various functions defined in the system of the present application are executed.

[0140] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0142] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.

[0143] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the method described in the above embodiments.

[0144] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0145] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0146] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.

[0147] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for driving 3D digital human eye and head movements with voice, characterized in that: include: Extract speech embedding features from the input speech signal; Inputting the speech embedding features into a pre-trained head movement generative model and an eye movement generative model, respectively, wherein the head movement generative model is constructed based on a variational autoencoder to generate a latent space representation of head movement, and the eye movement generative model is constructed based on a vector quantized variational autoencoder to generate a latent space representation of eye movement; cross-modally aligning the latent spatial representations of the head movement and the eye movement using a temporal Transformer model to generate a head posture parameter sequence and an eye rotation parameter sequence synchronized with the speech signal; A 3D digital human animation including head movements, eye movements and facial expressions is generated according to the head posture parameter sequence, the eye rotation parameter sequence and the facial expression parameter sequence generated based on voice drive.

2. The method according to claim 1, characterized in that Pre-training the head movement generation model and the eye movement generation model includes: Constructing a training dataset, wherein the training dataset includes audio-mesh sequences involving eye gaze, head and facial movements; The head movement generation model and the eye movement generation model are pre-trained based on the training data set.

3. The method according to claim 2, characterized in that: Construct a training dataset, including: Filtering a video to be processed containing a clear human eye area from the selected 2D audio-visual data; Performing 3D face reconstruction on the video to be processed frame by frame to obtain corresponding FLAME parameters, which include facial motion parameters and head motion parameters; Project the eyeball center of the 3D face model onto a 2D plane and fit it to the human eyes in the video frame to obtain the eye movement parameters; Temporal convolution smoothing processing is performed according to the head movement parameters, eye movement parameters and facial movement parameters to obtain a continuous audio-grid sequence.

4. The method according to claim 2, characterized in that Pre-training the head movement generation model and the eye movement generation model based on the training data set, including pre-training of a vector quantized variational autoencoder and joint training of the variational autoencoder and the vector quantized variational autoencoder; Among them, the pre-training of the vector quantized variational autoencoder includes: Input the true value of the eye movement parameter into the encoder of the vector quantized variational autoencoder to generate a feature vector; Mapping the feature vector to the nearest entry in the codebook using an element-by-element quantization function to obtain a quantized feature vector; Reconstructing the quantized feature vector into an eye movement parameter sequence using a decoder of a vector quantized variational autoencoder, and optimizing the parameters of the vector quantized variational autoencoder by minimizing the reconstruction loss and the codebook alignment loss; Joint training of variational autoencoders and vector quantized variational autoencoders, including: The true value of the head motion parameters and the speech embedding features are input into the encoder of the variational autoencoder to generate the corresponding Gaussian distributed latent code; Based on the latent code and the speech embedding features, generating a corresponding head movement parameter sequence through a decoder of a variational autoencoder; Taking the head movement parameter sequence as a condition, generating an eye movement parameter sequence through autoregression of a temporal Transformer model; The reconstruction loss, speed loss, KL divergence loss and feature regularization loss of the variational autoencoder are jointly optimized.

5. The method according to claim 2 or 3, characterized in that The method further comprises: Counting the blink frequency distribution of the training data set and generating random blink intervals based on Gaussian distribution; The eye expression parameters of the 3D digital human are adjusted within the frame window of the random blink interval to generate a continuous blink animation.

6. A voice-driven 3D digital human eye and head movement device, characterized by: include: An extraction module, used to extract speech embedding features from the input speech signal; a first generation module, configured to input the speech embedding features into a pre-trained head movement generation model and an eye movement generation model, respectively, wherein the head movement generation model is constructed based on a variational autoencoder to generate a latent space representation of head movement, and the eye movement generation model is constructed based on a vector quantized variational autoencoder to generate a latent space representation of eye movement; A second generation module is configured to perform cross-modal alignment on the latent space representations of the head movement and the eye movement using a temporal Transformer model to generate a head posture parameter sequence and an eye rotation parameter sequence synchronized with the speech signal; The processing module is used to generate a 3D digital human animation including head movements, eye movements and facial expressions according to the head posture parameter sequence, the eye rotation parameter sequence and the facial expression parameter sequence generated based on voice drive.

7. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voice-driven 3D digital human eye and head movement method according to any one of claims 1 to 5 is implemented.

8. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the voice-driven 3D digital human eye and head movement method as described in any one of claims 1 to 5.