A voice-driven 3D digital human generation method based on deep learning
By combining the Meta Former model with the UE5 engine, the applicability and rendering efficiency of voice-driven 3D digital human generation have been improved, solving the problems of applicability and high computing resource usage in existing technologies. It is suitable for different digital human roles and reduces the consumption of computing resources.
Patent Information
- Application Number
- CN202510927244.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-07
AI Technical Summary
Existing voice-driven technology has limited applicability and rendering effects in the field of digital humans, and it consumes a lot of computing resources, making it difficult to apply in different scenarios.
The Meta Former model is used to align and predict the features of audio data and facial data. Through the combination of linear layer, feature alignment layer, periodic position encoding layer, target mask layer, memory mask layer and action decoder, renderable 3D digital human mouth shape data is generated and rendered in real time using the UE5 engine.
It improves the applicability and rendering efficiency of voice-driven 3D digital human generation, makes it applicable to different digital human characters, and reduces the use of computing resources.
Smart Images

Figure CN120431222B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning audio processing technology, and more particularly to a method for generating a 3D digital human through voice-driven deep learning. Background Art
[0002] With the continuous advancement of artificial intelligence technology, deep learning-based audio-driven technology is gaining increasing attention. Audio-driven technology aims to use deep learning algorithms to build audio feature extraction models and, combined with computer vision or rendering technology, convert input audio into a video talking head or digital human.
[0003] Current research in audio feature extraction models is primarily based on classic deep learning network models, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs). For example, VOCA (Voice Operated Character Animation) mines audio features by building multi-layer convolutional neural networks and fully connected layers. It then uses a decoder in the fully connected layers to map these audio features to facial movements, and finally uses a mesh system to render realistic 3D facial animations. Wav2Lip, a speech-driven lip sync generation model, is centered on a deep learning-based generative adversarial network (GAN). By introducing a lip synchronization discriminator, the model can solve the problem of being unable to accurately adjust the lip movements of any identity when processing dynamic speaking videos, resulting in the video being out of sync with the new audio; Voice2Face is a model that can generate facial and tongue animations, using a conditional variational autoencoder to generate mesh animations from speech; FaceFormer is a model based on the Transformer architecture that uses a self-attention mechanism to process time series data. By training on a large number of labeled audio-animation pair datasets, it can generate realistic 3D facial animations from audio input.
[0004] However, existing voice-driven technologies use predicted facial data to generate videos or render mesh systems. Their effectiveness depends heavily on the quality of the video, and mesh systems are difficult to apply in real-world scenarios. Existing technologies also lack the ability to transfer between different digital human characters, limiting their applicability in different scenarios. Furthermore, rendering digital humans and exporting videos consume significant computing resources.
[0005] Therefore, how to improve the applicability and rendering effect of voice-driven technology in the field of digital humans is an urgent problem that technicians in this field need to solve. Summary of the Invention
[0006] In view of this, the present invention provides a voice-driven 3D digital human generation method based on deep learning, which realizes the prediction of voice-driven 3D digital human mouth shape data, improves the versatility of the predicted data, and improves the efficiency of digital human rendering.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A method for generating a 3D digital human through voice-driven deep learning includes the following steps:
[0009] Step 1: Collect audio data and corresponding facial data and perform preprocessing;
[0010] Step 2: Use the preprocessed audio data and facial data to train the Meta Former model to obtain a facial prediction model;
[0011] Step 3: Collect the audio to be converted and input it into the facial prediction model to obtain predicted facial data;
[0012] Step 4: Transmit the predicted facial data to the UE5 engine through the client to generate a digital human.
[0013] Preferably, the Wav2Vec2.0 model is used to preprocess the audio data, extract preliminary features, and obtain audio vectors.
[0014] Preferably, the Meta Former model includes a linear layer, a feature alignment layer, an action encoder, a periodic position encoding layer, a target mask layer, a memory mask layer, and an action decoder;
[0015] Linear layer, sets the character vector and converts it into character features;
[0016] Feature alignment layer, which converts the audio vector into a frame-based vector and corresponds one-to-one with each frame image in the facial data;
[0017] A periodic position encoder is configured to encode the predicted frame image using an improved sinusoidal position encoding method, obtain a periodic position code, add the periodic position code to the predicted frame image, and encode the periodic position code into a renderable format;
[0018] Target mask layer, generates target mask matrix according to the predicted frame image;
[0019] Memory mask layer, generates a memory mask matrix based on the predicted frame image;
[0020] The action decoder performs inference prediction based on the aligned audio vector and facial data, character features, target mask matrix and memory mask matrix, generates predicted frame images and encodes them. All predicted frame images arranged according to periodic position coding constitute the predicted facial data.
[0021] Preferably, the feature alignment layer includes a feature extraction layer, a linear insertion layer, and a feature projection layer connected in sequence, which can be expressed as:
[0022] ;
[0023] Among them, audio represents the input audio vector; FE represents the feature extraction operation, which converts the audio vector into a vector in frames; LI represents the linear interpolation operation, which adjusts the number of feature points in the audio vector to be consistent with the frame rate of the facial data; FP represents the feature projection operation, which converts the adjusted audio vector into a format that matches the number of facial frames.
[0024] Preferably, facial data and audio vectors are input into the MetaFormer model in batches for training. The training process of the MetaFormer model includes:
[0025] Step 21: Initialize the character vector and embed it into character features through a linear layer; the character features have the same size as the facial data;
[0026] Step 22: Obtain the total length of the facial data as the total number of frames;
[0027] Step 23: Input the audio vector and facial data into the feature alignment layer for feature alignment;
[0028] Step 24: Start batch cyclic prediction based on the total number of frames, set the number of cycles k=1, and project the character features into the same dimension as the facial data as the initial prediction frame image;
[0029] Step 25: Add periodic position codes to the predicted frame image through the periodic position encoder, generate a target mask matrix based on the predicted frame image in the target mask layer, generate a memory mask matrix based on the predicted frame image in the memory mask layer, and encode the periodic position codes corresponding to the predicted frame image into a renderable format;
[0030] Step 26: Input the aligned audio vector and facial data, character features, target mask matrix, and memory mask matrix into the action decoder for inference to obtain a new predicted frame image;
[0031] Step 27: Project the new predicted frame image to the same dimension as the facial data, and increase the number of loops k by 1;
[0032] Step 28: If the number of loops k is less than or equal to the total number of frames, return to step 25; otherwise, all predicted frame images arranged according to the periodic position coding constitute predicted facial data, and the loss is calculated based on the predicted facial data and the collected facial data;
[0033] Step 29: Perform backpropagation based on the loss to update the parameters of the Meta Former model and return to step 21 until the model iteration stopping condition is reached. The model parameters include the parameters of the linear layer, the linear insertion layer of the feature alignment layer, the feature projection layer, and the action decoder.
[0034] Preferably, the improved sinusoidal position encoding method is expressed as:
[0035] ;
[0036] in, Represents the scaling parameter, the default value is 100; t is the current frame time; d is the model dimension; is the dimension index; p is the time period, that is, the processing frequency, which defaults to 30fps; the periodic position encoder PPE cyclically injects position information in each time period of an audio vector, that is, taking the time period of the audio feature as a node, injects position information into the predicted frame image inferred from the corresponding node; Indicates the periodic position coding corresponding to the even-numbered predicted frame image, using sinusoidal periodic coding; Indicates the periodic position coding corresponding to the odd-numbered predicted frame image, using cosine periodic coding; mod represents the remainder after calculating the division; adding periodic position coding to the predicted frame image is expressed as:
[0037] ;
[0038] ;
[0039] Wherein, Sn represents the current predicted frame image; is the weight, is the deviation, is the vector value of the previous predicted frame image; x represents the predicted frame image with the number of loops k greater than 1 and less than or equal to the total number of frames T; Represents all predicted frame images inferred; Represents the predicted frame image at time t; Represents the predicted frame image at time t after adding periodic position coding.
[0040] Preferably, the target mask layer generates a target mask matrix according to the current predicted frame image, which is used in the action decoder to mask the ungenerated part of the audio vector that has not yet participated in the inference prediction, so as to avoid affecting the current predicted part that participates in the current inference prediction, and at the same time represents the association weight between the current predicted part of the audio vector and the generated part that has participated in the inference prediction; the target mask matrix Expressed as:
[0041] ;
[0042] Where p represents the time period; -∞ represents the masked area; i represents the column of the matrix, and j represents the row of the matrix. Represents the weight of the target code in the j-th row and the i-th column; the number represents the weight, and the smaller the value, the lower the weight when predicting the next frame.
[0043] Preferably, the memory mask layer generates a memory mask matrix based on the current predicted frame image, which is used to align the latest series of frame positions in the action decoder; the memory mask matrix Expressed as:
[0044] ;
[0045] Among them, k represents the number of current continuous frames; i represents the column of the matrix, and j represents the row of the matrix. Represents the memory weight of the j-th row and the i-th column.
[0046] Preferably, the Huber loss function is used to calculate the loss , expressed as:
[0047] ;
[0048] ;
[0049] in, is the threshold; represents the difference between the inferred facial data and the true value; Y represents the true facial data aligned with the audio vector and fed into the action decoder; and f(x) represents the predicted facial data generated by model inference. The loss function calculates the difference between the predicted value f(x) and the true value Y and continuously optimizes the model parameters using backpropagation to minimize the loss. This process enables the model to gradually learn the mapping between speech features and facial actions.
[0050] Preferably, the client includes a digital human module of a receiving module and a display module; the receiving module receives the predicted facial data and the audio to be converted; the display module is developed based on Linux / Windows system components and plays the audio to be converted; the digital human module is developed based on the JSONLiveLink plug-in and sends the predicted facial data to the UE5 engine in the time sequence of the predicted frame images in the predicted facial data. The UE5 engine renders and generates a digital human based on the predicted facial data, and feeds the digital human back to the display module for synchronous playback with the audio to be converted.
[0051] Preferably, facial data is detailed motion features of different facial organs, recording only relative positions, such as the direction and amplitude of movement of the mouth, left and right eyes, eyebrows, and chin. 52 key features are used, including: eyeBlinkLeft (blinking the left eye), eyeSquintLeft (squinting the left eye), eyeBlinkRight (blinking the right eye), eyeWideRight (widening the right eye), jawForward (chin forward), jawLeft (chin to the left), mouthClose (mouth closed), mouthLeft (mouth to the left), mouthFrownLeft (frowning the left corner of the mouth), mouthRollUpper (rolling the upper lip), browDownLeft (brow down), and cheekPuff (puffing the cheek).
[0052] It can be seen from the above technical solution that compared with the existing technology, the present invention discloses a voice-driven 3D digital human generation method based on deep learning. Through the processing of the Meta Former model and real-time transmission of the client to the UE5 engine for rendering, the facial data predicted based on the audio can be adapted to different digital human characters, getting rid of the problem that traditional digital humans rely on rendered video output and greatly improving rendering efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0054] Figure 1 A schematic diagram of the process flow of the voice-driven 3D digital human generation method based on deep learning provided by the present invention;
[0055] Figure 2 A schematic diagram of the Meta Former model structure provided by the present invention;
[0056] Figure 3 A schematic diagram of the Meta Former model training process provided by the present invention;
[0057] Figure 4 Schematic diagram of the Wav2Vec2.0 model structure provided by the present invention;
[0058] Figure 5 A schematic diagram of the target mask matrix structure provided by the present invention;
[0059] Figure 6 A schematic diagram of the memory mask matrix structure provided by the present invention;
[0060] Figure 7 A schematic diagram of the client transmission structure provided by the present invention;
[0061] Figure 8 A schematic diagram of the facial data fitting curve changes in the embodiment provided by the present invention;
[0062] Figure 9 A schematic diagram of changes in a facial data vector graph in an embodiment of the present invention;
[0063] Figure 10 Schematic diagram of loss changes during the model training process provided by the present invention. DETAILED DESCRIPTION
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0065] The embodiment of the present invention discloses a method for generating a 3D digital human driven by voice based on deep learning, such as Figure 1 As shown, the following steps are included:
[0066] S1: Collect audio data and corresponding facial data and perform preprocessing;
[0067] S2: Use the preprocessed audio data and facial data to train the Meta Former model to obtain a facial prediction model;
[0068] S3: Collect the audio to be converted and input it into the facial prediction model to obtain predicted facial data;
[0069] S4: The predicted facial data is transmitted to the UE5 engine through the client to generate a digital human.
[0070] The Wav2Vec2.0 model is then used to preprocess the audio data, extract preliminary features, and obtain audio vectors. The Wav2Vec2.0 model uses a series of convolutional and fully connected layers to perform preliminary feature extraction and dimensionality conversion, processing t seconds of audio data from one end into an audio vector of length t × sr, where sr is the model's sampling rate.
[0071] Furthermore, the Wav2Vec2.0 model structure is as follows Figure 4As shown in the figure, the system comprises a feature encoder, a quantization module, and a context modeling module. The feature encoder uses a multi-layer CNN network to process raw audio data and output a latent speech representation Z. The quantization module converts the continuous latent representation into a discrete quantized representation Q through vector quantization or product quantization. The context modeling module uses a masked mechanism to randomly mask the latent speech representation through a multi-layer Transformer network to generate a context representation C and obtain a speech vector. The model uses a contrastive loss L to compare the context representation with the quantized representation to optimize the model parameters.
[0072] Furthermore, the Meta Former model includes a linear layer, a feature alignment layer, an action encoder, a periodic position encoding layer, a target mask layer, a memory mask layer, and an action decoder;
[0073] The linear layer sets the character vector and converts it into character features. This conversion is a custom optional identifier before model training, such as differentiated training for male and female characters, or hosting style and daily style.
[0074] Feature alignment layer, which converts the audio vector into a frame-based vector and corresponds one-to-one with each frame image in the facial data;
[0075] A periodic position encoder is configured to encode the predicted frame image using an improved sinusoidal position encoding method, obtain a periodic position code, add the periodic position code to the predicted frame image, and encode the periodic position code into a renderable format;
[0076] Target mask layer, generates target mask matrix according to the predicted frame image;
[0077] Memory mask layer, generates a memory mask matrix based on the predicted frame image;
[0078] The action decoder performs inference prediction based on the aligned audio vector and facial data, character features, target mask matrix, and memory mask matrix, generating and encoding predicted frames. All predicted frames arranged according to periodic positional encoding form the predicted facial data. The length of the preprocessed audio vector is t × sr, while the length of the facial data is t × 30. Due to the different data lengths of the two modalities, a feature alignment layer is required before training.
[0079] Furthermore, the feature alignment layer includes a feature extraction layer, a linear insertion layer, and a feature projection layer connected in sequence, which can be expressed as:
[0080] ;
[0081] Where audio represents the input audio vector; FE represents the feature extraction operation; LI represents the linear interpolation operation; and FP represents the feature projection operation. Key information is extracted from the audio using the Wav2Vec network through feature extraction. Based on the frame rate of the facial data, the number of feature points that need to be added or subtracted from the audio vector corresponding to the extracted key information is determined, and linear interpolation is used to achieve this. The adjusted audio vector is then converted into a format that matches the number of facial frames using feature projection.
[0082] The feature extraction layer consists of a series of one-dimensional convolution modules. By adjusting the convolution kernel size and stride, the audio signal is converted into a frame-level vector representation. Since the feature extraction layer processes at a high precision of 50 frames per second, it is necessary to reconvert the audio vector to a 30-frame-per-second representation through a linear interpolation layer. The feature projection layer adjusts the vector dimension to 512 for subsequent unified processing. After feature alignment, the dimension of the audio feature is , and the dimension of facial data is The sampling rate refers to the number of audio samples collected per second, and the frame rate refers to the number of facial data frames collected per second. The length of the audio vector corresponding to each frame is the ratio of the sampling rate to the frame rate, that is, sr / fps. The frame length len is: .
[0083] Furthermore, facial data and audio vectors are fed into the MetaFormer model in batches for training. The training process of the MetaFormer model includes:
[0084] S21: Initialize the character vector and embed it into character features through a linear layer; the character features have the same size as the facial data;
[0085] S22: Obtain the total length of the facial data as the total number of frames;
[0086] S23: Input the audio vector and facial data into the feature alignment layer for feature alignment;
[0087] S24: Start batch cyclic prediction based on the total number of frames, set the number of cycles k=1, project the character features into the same dimension as the facial data, and use it as the initial prediction frame image;
[0088] S25: adding a periodic position code to the predicted frame image through a periodic position encoder, generating a target mask matrix based on the predicted frame image at a target mask layer, generating a memory mask matrix based on the predicted frame image at a memory mask layer, and encoding the periodic position code corresponding to the predicted frame image into a renderable format;
[0089] S26: Input the aligned audio vector and facial data, character features, target mask matrix, and memory mask matrix into the action decoder for inference to obtain a new predicted frame image;
[0090] S27: Project the new predicted frame image to the same dimension as the facial data, and increase the number of loops k by 1;
[0091] S28: If the number of loops k is less than or equal to the total number of frames, return to S25; otherwise, all predicted frame images arranged according to the periodic position coding and the corresponding periodic position coding in the renderable format constitute predicted facial data, and the loss is calculated based on all predicted frame images and facial data in the predicted facial data;
[0092] S29: Backpropagate the loss to update the parameters of the Meta Former model, and return to S21 until the model iteration stopping condition is reached. The model parameters include the linear layer, the linear insertion layer of the feature alignment layer, the feature projection layer, and the parameters of the action decoder. Backpropagation is used to calculate the gradient of the loss with respect to each parameter. After calculating the gradient, an optimization algorithm (such as gradient descent and its variants, such as Adam and RMSprop) is used to update the model parameters. The update formula is generally parameter = parameter - learning rate × gradient. The learning rate is a hyperparameter.
[0093] Furthermore, the improved sinusoidal position encoding method is expressed as:
[0094] ;
[0095] in, Represents the scaling parameter, the default value is 100; t is the current frame time; d is the model dimension; is the dimension index; p is the time period, that is, the processing frequency, which defaults to 30fps; the periodic position encoder PPE cyclically injects position information in each time period of an audio vector, that is, taking the time period of the audio feature as a node, injects position information into the predicted frame image inferred from the corresponding node; Indicates the periodic position coding corresponding to the even-numbered predicted frame image, using sinusoidal periodic coding; Indicates the periodic position coding corresponding to the odd-numbered predicted frame image, using cosine periodic coding; mod represents the remainder after calculating the division; adding periodic position coding to the predicted frame image is expressed as:
[0096] ;
[0097] ;
[0098] Wherein, Sn represents the current predicted frame image; is the weight, is the deviation, is the vector value of the previous predicted frame image; x represents the predicted frame image with the number of loops k greater than 1 and less than or equal to the total number of frames T; y represents the predicted frame image with the number of loops k equal to 1; Represents all predicted frame images inferred; Represents the predicted frame image at time t; Represents the predicted frame image at time t after adding periodic position coding.
[0099] Furthermore, the target mask layer generates a target mask matrix based on the predicted frame image, which is used in the action decoder to mask the ungenerated part of the audio vector that has not yet participated in the inference prediction, so as to avoid affecting the current predicted part that participates in the current inference prediction, and at the same time represents the association weight between the current predicted part of the audio vector and the generated part that has participated in the inference prediction; the target mask matrix Expressed as:
[0100] ;
[0101] Among them, p represents the processing frequency (time period); -∞ represents the masked area; i represents the column of the matrix, and j represents the row of the matrix. Indicates the weight of the target code in row j and column i; the number indicates the weight, the smaller the value, the lower the weight when predicting the next frame. Figure 5 The figure shows a 15×15 target mask matrix, indicating that 15 frames of data have been generated. The target mask matrix for the 16th frame will be generated according to the above formula. The visualization result is similar to this matrix expanded vertically and horizontally by one row (column) to the lower right corner. Different numbers in the target mask matrix represent different data, such as: infinitesimal represents the ungenerated part, 0 represents the current predicted part, and less than 0 represents the generated part.
[0102] Furthermore, the memory mask layer generates a memory mask matrix based on the current predicted frame image, which is used to align the latest series of frame positions in the action decoder; the memory mask matrix Expressed as:
[0103] ;
[0104] Among them, k represents the number of current continuous frames; i represents the column of the matrix, and j represents the row of the matrix. Represents the memory weight of the j-th row and the i-th column.
[0105] Furthermore, the Huber loss function is used to calculate the loss , expressed as:
[0106] ;
[0107] ;
[0108] in, is the threshold; Represents the gap between the inferred facial data and the true value; Y represents the true facial action data, that is, the facial data aligned with the audio vector input to the action decoder; f(x) represents the predicted facial data generated by model inference. The loss function calculates the difference between the predicted value f(x) and the true value Y, and uses back propagation to continuously optimize the model parameters to minimize the loss value. This process enables the model to gradually learn the mapping relationship from speech features to facial actions. Since facial data only records relative positions, cumulative errors or noise interference may occur during transmission or processing. The use of the Huber loss function can improve the system's robustness to outliers, while still fully capturing changes in details when the error is small. Using Huber loss can better balance the accuracy and robustness of the model. When the error is small (that is, below the threshold ), Huber loss is similar to mean square error, which squares the error, similar to mean square error (MSE), emphasizing the impact of smaller errors. When the error is large, the loss function is converted to a linear relationship, similar to mean absolute error (MAE), which is smoother for outliers and reduces the impact of outliers on model training.
[0109] Furthermore, the client includes a digital human module consisting of a receiving module and a display module; the receiving module receives the predicted facial data and the audio to be converted; the display module is developed based on Linux / Windows system components and plays the audio to be converted; the digital human module is developed based on the JSONLiveLink plug-in and sends the predicted facial data to the UE5 engine in the time sequence of the predicted frame images. The UE5 engine renders and generates a digital human based on the predicted facial data, and feeds the digital human back to the display module for synchronous playback with the audio to be converted.
[0110] The client is developed in Python, using TCP connections, independent of operating systems (Windows / Linux). Utilizing multi-threading, the threads sending predicted frames and playing the converted audio execute simultaneously without interfering with each other. The sending speed of predicted facial data, which corresponds to the speed of facial animation in the UE5 engine, can be manually controlled. For example, if the predicted facial data is 30fps (30 frames per second), the frame delay can be set to 0.033 seconds to ensure accurate frame transmission and matching of the audio.
[0111] Furthermore, facial data contains detailed motion features of various facial organs, recording only the relative positions, such as the direction and amplitude of movement of the mouth, left and right eyes, eyebrows, and chin. 52 key features are used, including: eyeBlinkLeft (blinking the left eye), eyeSquintLeft (squinting the left eye), eyeBlinkRight (blinking the right eye), eyeWideRight (widening the right eye), jawForward (chin forward), jawLeft (chin to the left), mouthClose (mouth closed), mouthLeft (mouth to the left), mouthFrownLeft (frowning the left corner of the mouth), mouthRollUpper (rolling the upper lip), browDownLeft (brow down), and cheekPuff (puffing the cheek).
[0112] On the other hand, in a specific embodiment, the facial data includes a single mouth feature. During the training process, the left lip (mouthLeft) fitting curve of the facial data corresponding to a training audio segment changes as follows: Figure 8 As shown, orange represents the original facial data, blue represents the fitted facial data, the horizontal axis represents the frame length, and the vertical axis represents the facial feature value. Figure 8 (a)-(d) show the curve changes obtained after 5, 35, 70, and 100 iterations of training, respectively. By observing the changes in the fitted curves during training, we can see that the initialized facial features show a uniform distribution over the frame length, which is mainly due to the introduction of periodic position encoding. As training progresses, the fitted curves gradually approach the original curves, indicating that the facial features predicted by the model are becoming increasingly similar to the original facial features.
[0113] During the training process, the facial data of a training audio and the vector graph of the predicted facial data change during the training process, such as Figure 9 As shown, the horizontal axis represents all facial feature vectors, the vertical axis represents the frame length, and the brightness and darkness of the color represent the size of the value. Figure 9 (a) represents the facial feature vector of the original facial data, Figure 9 (b)-(e) show the facial feature vectors of the predicted facial data obtained after 5, 35, 70, and 100 iterations of training, respectively. The vectors and their variations show that the initialized facial feature maps differ significantly from the original ones. However, as training progresses, the facial feature maps gradually approach the original ones in both facial feature and frame dimensions. This demonstrates that the model has adequately learned holistic facial reasoning.
[0114] like Figure 10The figure shows how the training loss changes with the number of training iterations (epochs), with the horizontal axis representing the number of iterations and the vertical axis representing the loss value. As you can see, the loss curve is relatively stable, indicating that the model is able to effectively learn and stably master knowledge during training.
[0115] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0116] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating 3D digital humans based on voice-driven deep learning, characterized in that: The following steps are involved: Step 1: Collect audio data and corresponding facial data and perform preprocessing; Step 2: Use the preprocessed audio data and facial data to train the MetaFormer model to obtain a facial prediction model. Use the Wav2Vec model to preprocess the audio data, extract preliminary features, and obtain audio vectors. The MetaFormer model includes a linear layer, a feature alignment layer, an action encoder, a periodic position encoder, a target mask layer, a memory mask layer, and an action decoder. Linear layer, sets the character vector and converts it into character features; Feature alignment layer, which converts the audio vector into a frame-based vector and corresponds one-to-one with each frame image in the facial data; A periodic position encoder is configured to encode the predicted frame image using an improved sinusoidal position encoding method, obtain a periodic position code, add the periodic position code to the predicted frame image, and encode the periodic position code into a renderable format; Target mask layer, generates target mask matrix according to the predicted frame image; Memory mask layer, generates a memory mask matrix based on the predicted frame image; The action decoder performs inference prediction based on the aligned audio vector and facial data, character features, target mask matrix, and memory mask matrix, generates predicted frame images, and encodes them. All predicted frame images constitute the predicted facial data. The target mask layer generates a target mask matrix based on the current predicted frame image. The target mask matrix Expressed as: Where p represents the time period; i represents the column of the matrix, j represents the row of the matrix, and B F (i, j) represents the target code weight of the j-th row and i-th column; The memory mask layer generates a memory mask matrix based on the current predicted frame image. The memory mask matrix B A Expressed as: Among them, k represents the number of the current continuous frames; i represents the column of the matrix, j represents the row of the matrix, B A (i, j) represents the memory weight of the j-th row and i-th column; The improved sinusoidal position encoding method is expressed as: PPE (t,2α) =sin((t mod p) / 10000 2α / d ) / t PPE (t,2α+1) =cos((t mod p) / 10000 2α / d ) / τ Among them, τ represents the scaling parameter; t is the current frame time; d is the model dimension; α is the dimension index; p is the time period; PPE (t,2α) Indicates the periodic position encoding corresponding to the even-numbered predicted frame image; PPE (t,2α+1) Indicates the periodic position code corresponding to the odd-numbered predicted frame image; mod indicates the remainder after calculating the division; adding the periodic position code to the predicted frame image is expressed as: Wherein, Sn represents the current predicted frame image; W f is the weight, b f is the deviation, is the vector value of the previous predicted frame image; x represents the predicted frame image with the number of cycles k greater than 1 and less than or equal to the total number of frames T; f(x) represents all the predicted frame images inferred; f t Represents the predicted frame image at time t; Represents the predicted frame image at time t after adding periodic position coding; Step 3: Collect the audio to be converted and input it into the facial prediction model to obtain predicted facial data; Step 4: Transmit the predicted facial data to the UE5 engine through the client to generate a digital human.
2. The method for generating 3D digital humans based on voice-driven deep learning according to claim 1, characterized in that: The feature alignment layer consists of a feature extraction layer, a linear insertion layer, and a feature projection layer connected in sequence, which can be expressed as: FP(LI(FE(audio))) Among them, audio represents the input audio vector; FE represents the feature extraction operation; LI represents the linear interpolation operation, which corresponds the audio vector to each frame image in the facial data one by one; FP represents the feature projection operation, which maps the audio vector to a format that matches the facial data.
3. The method for generating 3D digital humans based on voice-driven deep learning according to claim 1, characterized in that: The training process of the Meta Former model includes: Step 21: Initialize the character vector and embed it into character features; the character features have the same size as the facial data; Step 22: Obtain the total length of the facial data as the total number of frames; Step 23: Align the audio vector and facial data; Step 24: Start batch cyclic prediction based on the total number of frames, set the number of cycles k = 1, and project the character features into the same dimension as the facial data as the initial prediction frame image; Step 25: Add periodic position codes to the predicted frame image, generate a target mask matrix and a memory mask matrix based on the predicted frame image, and encode the periodic position codes corresponding to the predicted frame image into a renderable format; Step 26: Perform inference based on the aligned audio vector and facial data, character features, target mask matrix, and memory mask matrix to obtain a new predicted frame image; Step 27: Project the new predicted frame image to the same dimension as the facial data, and increase the number of loops k by 1; Step 28: If the number of loops k is less than or equal to the total number of frames, return to step 25; otherwise, all predicted frame images constitute predicted facial data, and the loss is calculated based on the predicted facial data and the collected facial data; Step 29: Perform backpropagation based on the loss, update the parameters of the MetaFormer model, and return to step 21 until the model iteration stopping condition is reached.
4. The method for generating a 3D digital human based on voice-driven deep learning according to claim 3, characterized in that: The loss L is calculated using the Huber loss function δ (a), expressed as: a=Yf(x) Where δ is the threshold; a represents the gap between the predicted facial data and the facial data; Y represents the facial data; and f(x) represents the predicted facial data predicted by inference.
5. The method for generating 3D digital humans based on voice-driven deep learning according to claim 1, characterized in that: The client includes a digital human module, a receiving module, and a display module; the receiving module receives the predicted facial data and the audio to be converted; the display module is developed based on Linux / Windows system components and plays the audio to be converted; the digital human module is developed based on the JSONLiveLink plug-in and sends the predicted facial data to the UE5 engine in the time sequence of the predicted frame images in the predicted facial data. The UE5 engine renders and generates a digital human based on the predicted facial data, and feeds the digital human back to the display module for synchronous playback with the audio to be converted.
Citation Information
Patent Citations
Voice-driven 3D character facial expression method based on deep learning
CN113763519A
Facial animation generation method and device, electronic equipment and storage medium
CN116912375A