A method and system for generating digital human speaking videos via text-driven technology

By employing a text-driven digital character speaking video generation method, and utilizing techniques such as the Face-Alignment model and multimodal attention mechanism, the problems of insufficient action and facial expression and poor synchronization between speech and action are solved, thereby improving generation efficiency and quality and achieving high-quality video generation.

CN120812364BActive Publication Date: 2025-11-14HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511303017.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-11-14
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing digital character speaking video generation technologies suffer from a lack of movement and facial expressions, poor synchronization between voice and movement, and bottlenecks in generation efficiency and quality. As a result, the generated videos lack realism and high quality, making it difficult to meet real-time requirements and user needs.

Method used

This method generates digital human speaking videos using a text-driven approach. It utilizes the Face-Alignment model for face detection and cropping, combines the CLIP text encoder and BERT model to extract semantic and emotional features, employs a multimodal attention mechanism and a U-shaped network for image feature reshaping, uses WaveNet to generate audio, and generates high-quality videos through a time alignment module.

Benefits of technology

It improves the generation effect and efficiency of digital character speaking videos, realizes natural and smooth changes in movements and expressions, enhances the synchronization between voice and movement, and generates videos with high realism and image quality, meeting users' demand for high-quality video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812364B_ABST
    Figure CN120812364B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for generating digital character speaking videos using text-driven methods, belonging to the field of artificial intelligence video generation technology. The method first processes the videos involved in the digital character video dataset; then extracts text and image features; next, it reconstructs image features and uses WaveNet to generate audio from the text features; then, it repairs the generated multi-frame images; finally, it concatenates the repaired images and the generated audio in chronological order to generate a digital character video and evaluates the result. This method possesses powerful control capabilities and diverse control types, requires no retraining of the base model, effectively improves the generation effect of digital character speaking videos, and ensures a high degree of consistency between the character's actions, expressions, and speech content in the video, significantly enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence video generation technology, and more specifically, relates to a method and system for generating digital character speaking videos driven by text. Background Technology

[0002] In today's era of rapid digital information dissemination, digital character speaking videos, as a highly expressive content format, have gained widespread application and attention in various fields such as virtual live streaming, smart education, advertising, and film and television production. Using digital virtual characters as a medium, they convey information through speech, attracting user attention and enhancing the effectiveness of information delivery. However, existing digital character speaking video generation technologies have significant shortcomings, severely restricting their further development and application in various fields.

[0003] Limited Actions and Expressions: Currently, most digital character speaking video generation technologies rely primarily on pre-set action and expression libraries. In practical applications, the actions and expressions of digital characters are extremely limited, making it difficult to adapt naturally and smoothly to different speech content and emotional expression needs. Whether it's the emotional fluctuations when telling a story or the facial expressions during interactive communication, they appear stiff and rigid, failing to showcase rich and diverse emotional layers and a vivid and natural communication style. This makes digital characters lack realism and approachability, making it difficult to establish an effective emotional connection with the audience.

[0004] Poor synchronization between voice and action: Achieving accurate synchronization between voice and action is one of the key technical challenges in generating digital character speaking videos. Existing technologies often have significant errors when dealing with the synchronization of voice and action. Inconsistencies between lip movements and speech occur frequently; the lip movements of a digital character cannot accurately correspond to the pronunciation of the speech. Discrepancies between facial expressions and semantics are also common; for example, a digital character may display sadness while expressing happiness. These problems severely affect the realism of the video and the viewing experience, reducing the accuracy and credibility of information transmission.

[0005] Generation Efficiency and Quality Bottlenecks: From a technical implementation perspective, existing methods for generating digital human speaking videos lack efficient fusion techniques when processing multimodal data (such as text, speech, and images). This results in extremely high computational complexity during the generation process, requiring significant computing resources and time, making it difficult to meet real-time requirements. For example, in live streaming scenarios, it is impossible to generate high-quality digital human speaking videos in a timely manner. Furthermore, the generated videos also suffer from significant deficiencies in image quality and detail; the characters are not realistic enough, and the background is not clear enough, failing to meet users' demands for high-quality video content.

[0006] These issues not only limit the application of digital character speaking videos in existing fields, but also hinder their expansion into emerging fields such as immersive virtual reality experiences and high-end film and television special effects production. Summary of the Invention

[0007] To address the aforementioned problems, this invention provides a method and system for generating digital character speaking videos driven by text, thereby improving the generation effect, efficiency, and quality, which is an important problem that urgently needs to be solved in this field.

[0008] The above objectives are achieved through the following technical solutions:

[0009] This invention first provides a method for generating digital character speaking videos driven by text, the method comprising the following steps:

[0010] (1) For the videos involved in the digital person video dataset, video processing is performed: First, the original video is preprocessed by frame rate normalization; then, the preprocessed video is converted into images frame by frame; then, the Face-Alignment model is used to detect faces in the obtained images and the facial regions are cropped to obtain the processed images.

[0011] (2) Text and image feature extraction: The specific feature extraction method for the additional text information used to describe the facial expressions and dialogue content of the characters in the digital character video is as follows: The pre-trained CLIP text encoder is used to process the text information to generate semantic features; at the same time, a sentiment classification auxiliary branch based on the BERT pre-trained model is introduced to extract the sentiment features of the text simultaneously; finally, the semantic features and sentiment features are fused to form a two-way text feature that has both semantic content and sentiment attributes.

[0012] For the processed image obtained in step (1), the image features are extracted using a vector quantization variational autoencoder.

[0013] (3) Image feature reshaping: The text features and image features obtained in step (2) are first added to the image features through a diffusion process, and then the text features and the image features with added noise are fused through a multimodal attention mechanism and fed into a U-shaped network for iterative denoising. After multiple iterations, the denoised image features are generated. This process is repeated to generate multiple reshaped image features.

[0014] (4) Audio generation: Audio is generated using WaveNet based on the text features obtained in step (2);

[0015] (5) Image generation: Multiple reconstructed image features are obtained from step (3). Before being sent to the decoder, they need to be processed by the time alignment module. Specifically, some image features are first specified from the overall image features as motion features. Then, noise is added to the motion features and they are spliced ​​with the current batch of image features along the time axis. After that, the image features processed by the time alignment module are sent to the decoder to reconstruct multiple frames of images. Then, the generated multiple frames of images are repaired by the face restoration module.

[0016] (6) Video generation: The image repaired in step (5) and the audio generated in step (4) are spliced ​​together in chronological order to generate a digital character video;

[0017] (7) Evaluation of generated results.

[0018] Furthermore, the preprocessing operation of performing frame rate standardization on the original video in step (1) is aimed at the problem of inconsistent video frame rates in the HDTF dataset. All video frame rates are standardized to 25 frames per second. For video frames with a frame rate lower than 25 frames per second, frames are padded, and for video frames with a frame rate greater than 25 frames per second, redundancy is removed.

[0019] In step (1), the Face-Alignment model is used to perform face detection on the obtained image and crop out the facial region. Specifically, the 68 feature points of the face are located by the facial key point detection technology in the Face-Alignment model. The minimum bounding rectangle of the coordinate set of these 68 feature points is calculated as the initial facial region. The boundaries of the initial facial region are expanded outward to form an extended region. The extended region is extracted from the original image as the face cropping result. Finally, the bilinear interpolation algorithm is used to scale the cropping region to a standard size of 512×512 pixels.

[0020] Furthermore, step (2) specifically includes the following sub-steps:

[0021] (2-1) The additional text information used to describe the facial expressions and dialogue content of the digital characters in the dataset is processed using a pre-trained CLIP text encoder to generate semantic features:

[0022] The input text is After processing by the CLIP text encoder, the semantic features of the text are obtained. The calculation method is as follows:

[0023] ,

[0024] in For CLIP text encoder, Represents the set of real numbers. Semantic feature dimension;

[0025] (2-2) Extracting sentiment features from text using the sentiment classification auxiliary branch of the BERT pre-trained model. The calculation formula is as follows:

[0026] ,

[0027] in For the BERT model, This represents a multilayer perceptron sentiment classifier. For the emotional feature dimension;

[0028] (2-3) Merge the semantic features obtained in step (2-1) and the sentiment features obtained in step (2-2) to obtain bidirectional text features containing semantic content and sentiment attributes. The calculation formula is as follows:

[0029] ,

[0030] in This is the weight vector;

[0031] (2-4) Based on the processed image obtained in step (1) Image features are extracted using a vector quantization variational autoencoder. The calculation formula is as follows:

[0032] ,

[0033] It is a vector quantization variational autoencoder that converts the image The continuous features obtained after compression It is an encoder function that converts an image into continuous features;

[0034] Encoder function This can be further represented as the multi-layer operations of a convolutional neural network:

[0035] ,

[0036] in, It is a convolutional layer. It is the Sigmoid activation function. The number of convolutional layers.

[0037] Furthermore, step (3) specifically includes the following sub-steps:

[0038] (3-1) The image features obtained in step (2) By gradually adding noise, it is eventually transformed into pure Gaussian noise. This process is governed by a predefined noise schedule. Control, among which Indicates the first The noise intensity of the first step, the second step The calculation formula for the noise addition process in step 1 is as follows:

[0039] ,

[0040] ,

[0041] in For the first Features after adding noise For the first Features after adding noise It is random Gaussian noise. Let be the noise weights at step t;

[0042] After recursive expansion, directly from the original image features Generate the first Features after adding noise:

[0043] ,

[0044] in The total noise weight is expressed as: , It is the initial random Gaussian noise;

[0045] (3-2) will the first Features after adding noise The bidirectional text features obtained in step (2) By using a multimodal attention mechanism for fusion, the image generation process can understand text semantics. The calculation formula is as follows:

[0046] ,

[0047] in, This represents the function of the multimodal attention mechanism. Let Q be the normalization function, and let Q be the query matrix. Q comes from image features; K is the key matrix and V is the value matrix, both of which come from bidirectional text features. ; It represents the dimension of the feature, and the superscript T indicates transpose. A fully connected layer representing bidirectional textual features;

[0048] (3-3) U-shaped network, that is, the network with the first... Features after adding noise Two-way text features and time step Given the input, predict the added noise. Gradually restore the original image features.

[0049] No. The denoising update step is calculated using the following formula:

[0050] ,

[0051] in The standard deviation of noise. For additional noise; The calculation formula is as follows;

[0052] ,

[0053] in, Noise added to the prediction; No. The features after adding noise are initially random noise. ; It is a two-way text feature; For time steps; This is the variance scheduling parameter, which controls the noise addition rate. It is set using a cosine function for linear scheduling, and the calculation formula is as follows:

[0054] ,

[0055] in, and These are the minimum and maximum values ​​of the variance, respectively.

[0056] (3-4) Through multiple independent diffusion denoising loops, several different but consistent denoised image features are generated, each loop starting from a different random noise point. Beginning, the first The denoised image features generated by the round are calculated using the following formula:

[0057] ,

[0058] ,

[0059] in For the first Denoising image features generated by the round, For the diffusion process described above, For the first The initial random noise of the wheel, This involves multiple iterations, where N is the number of iterations for image generation.

[0060] Finally, the images are concatenated to form a reconstructed image feature set after denoising. The calculation formula is as follows:

[0061] .

[0062] Furthermore, step (4) specifically includes the following sub-steps:

[0063] (4-1) Two-way text features obtained from step (2) Including semantic and sentiment information, it needs to be transformed into a conditional embedding vector suitable for WaveNet through a mapping layer. The calculation formula is as follows:

[0064] ,

[0065] in For audio features, It is a linear transformation matrix. For conditional embedding dimensions;

[0066] (4-2) WaveNet captures long-range dependencies in audio sequences through multi-layer dilated causal convolutions, while also incorporating audio features Integrating into the calculation of each layer, the first The layer dilation convolution operation is calculated using the following formula:

[0067] ,

[0068] in For the first Layer expansion characteristics, For the first Layer expansion characteristics, For the first Layer dilation convolution operation, The kernel size is [size]. For the first Layer feature dimension It is the expansion factor;

[0069] Next, conditions are fused, and the calculation formula is as follows:

[0070] ,

[0071] in For the first Layer conditional fusion features It is the Sigmoid activation function. This indicates element-wise multiplication. It is the hyperbolic tangent function. These are the scaling parameters for the first batch. This is the offset parameter for the first batch. The scaling parameters for the second batch. This is the offset parameter for the second batch. From audio features Generated through linear layers:

[0072] ,

[0073] in For linear layers, Audio features;

[0074] (4-3) The output of WaveNet is the probability distribution of audio samples, which is then converted into a continuous audio signal through an autoregressive generation process. The final output probability distribution is calculated as follows:

[0075] ,

[0076] in For probability distribution, For a moment Audio samples, This represents the total number of layers in WaveNet. For the final conditional fusion feature;

[0077] The audio signal is converted to a continuous audio signal through an autoregressive generation process, and the calculation formula is as follows:

[0078] ,

[0079] in For a moment Audio samples, The sample value with the highest conditional probability. Audio features;

[0080] Finally, the audio is spliced ​​together to form a continuous audio stream. The calculation formula is as follows:

[0081] ,

[0082] in For continuous audio, For a moment Audio samples.

[0083] Furthermore, step (5) specifically includes the following steps:

[0084] (5-1) Step (3) yields the reconstructed image feature set as follows: Each feature For the semantic representation of an image, let the motion features be... ,in, Indicates the first Motion characteristics of the image;

[0085] The following formula is used to add noise to motion features:

[0086] ,

[0087] in, For the first Motion characteristics of an image after adding noise Indicates the first Motion characteristics of the image For noise;

[0088] The noise-added motion features are concatenated with the corresponding image features to obtain the input features for the decoder:

[0089] ,

[0090] in, For the first Input features of the image This indicates a join operation along the feature dimension. For the first Zhang's denoised image features For the first Motion characteristics of an image after adding noise;

[0091] (5-2) The decoder upsamples the input features through multiple upsampling operations. Remodeling into feature maps The calculation formula is as follows;

[0092] ,

[0093] in In the The first image Second upsampling, For upsampling, The kernel size is [size]. For the first The upsampling factor of the layer, This represents the total number of upsampling layers;

[0094] Then, the RGB pixel values ​​are output through an activation function, calculated using the following formula:

[0095] ,

[0096] in For the generated first Zhang Image For the first The upsampling results of the images, Use the Sigmoid activation function;

[0097] (5-3) Finally, the face restoration module is used to restore the generated multi-frame video images. The calculation formula is as follows:

[0098] ,

[0099] For the first The high-resolution image after super-resolution processing. For super-resolution networks, These are the parameters of the pre-trained super-resolution model;

[0100] Then, a detector is used to obtain the face bounding boxes for each image. And the coordinates of the key points, to obtain the affine transformation matrix. :

[0101] ,

[0102] For the first The affine transformation matrix of the image; A function for calculating the transformation parameters. for Reference key points for the face in the image, such as the coordinates of the eyes, nose tip, and mouth corners;

[0103] The facial features are extracted from the image through affine transformation, aligned to a standard pose, and scaled to a fixed size.

[0104] ,

[0105] For the first A face image after pose alignment It is a quadratic mapping. For the first Affine transformation of the image;

[0106] Align the face input generator G to generate the repaired result:

[0107] ,

[0108] For the first Zhang's restored facial image For generator networks; These are the parameters of the pre-trained generator;

[0109] To ensure temporal consistency between adjacent frames, the optical flow field between adjacent images needs to be calculated first:

[0110] ,

[0111] in For the first Image number 1 The optical flow field of +1 facial image, Let be the optical flow field function. For the first Zhang's restored facial image For the first +1 restored face image;

[0112] The final repair result was generated through weighted fusion:

[0113] ,

[0114] For the first The final result of the face restoration. These are the weighting coefficients. For optical flow-based pixel remapping operations, For the first Image number 1 The light flow field of a facial image;

[0115] Finally, the repaired and unrepaired regions are merged: first, by using inverse affine transformation. The final repair result The face pose is transformed back to the original image from the standardized pose, and then compared with the original image after super-resolution processing. Based on face bounding box Merge regions:

[0116] ,

[0117] For the first Zhang merged the restored face image with the original image background to create the complete image. A function for fusing images. For the function of changing posture, For the first The final result of the face restoration. For the first Inverse affine transformation of the image For the first The high-resolution image after super-resolution processing. The bounding box for the face.

[0118] This invention also provides a text-driven digital person video generation system, including a video processing module, a text and image feature extraction module, an image feature generation module, an audio generation module, and a video generation module. The video processing module first performs frame rate normalization preprocessing on the original video, detects faces using a Face-Alignment model and crops out facial regions, then adjusts it to a 512×512 pixel image. The text and image feature extraction module extracts corresponding features based on text and image information. The image feature generation module performs noise addition and removal based on the extracted image features and fuses them with text features to obtain reconstructed image features. The audio generation module generates audio using WaveNet based on text features. The video generation module uses a time alignment module to ensure the generated image sequence has continuity over time based on the reconstructed image features, finally stitches together multiple images and audio to construct a digital person video model, generates a fused digital person video, and evaluates the generation result. Attached Figure Description

[0119] Figure 1 This illustration shows a flowchart of the steps in the text-driven digital character speaking video generation method according to an embodiment of the present invention.

[0120] Figure 2 The specific component modules of image processing in an embodiment of the present invention are shown;

[0121] Figure 3 The specific components of image generation in an embodiment of the present invention are shown. Detailed Implementation

[0122] The technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the relevant principles and requirements. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0123] This embodiment presents a text-driven digital persona video generation system. This system achieves high-quality digital persona video generation through the collaborative work of multiple modules, and the application of a multimodal attention mechanism and a time alignment module. Specifically, the system includes a video processing module, a text and image feature extraction module, an image feature generation module, an audio generation module, and a video generation module. The video processing module first performs frame rate normalization preprocessing on the original video, detects faces using a Face-Alignment model and crops out the facial region, then adjusts it to a 512×512 pixel image. The text and image feature extraction module extracts corresponding features based on text and image information. The image feature generation module performs noise addition and denoising on the extracted image features and fuses them with text features to obtain reconstructed image features. The audio generation module generates audio using WaveNet based on text features. The video generation module, based on the reconstructed image features, uses a time alignment module to ensure the continuity of the generated image sequence over time, finally stitches together multiple images and audio to construct a digital persona video model, generates a fused digital persona video, and evaluates the generation result.

[0124] Based on the above system, a method for generating text-driven digital character speaking videos according to the present invention includes the following steps:

[0125] (1) Perform video processing on the videos involved in the digital person video dataset:

[0126] To fully validate the effectiveness of the text-driven digital character video generation method, this embodiment constructs a multi-source training dataset with a total duration of approximately 134 hours. This dataset covers diverse scenarios, aiming to comprehensively ensure the model's diversity and generalization ability, enabling it to achieve excellent performance in various complex real-world application scenarios.

[0127] Data source composition:

[0128] The HDTF dataset plays a crucial role as a benchmark dataset for professional scenarios. It contains 6 hours of high-quality raw video clips, all at a uniform resolution of 1920×1080, providing a clear and stable foundation for subsequent processing and analysis. The dataset includes over 200 speaker samples from different genders, ages (ranging from 10 to 70 years old), and ethnicities, constructing a balanced demographic distribution that fully reflects the characteristics of different population groups. Recorded in a professional recording studio environment with a three-camera fixed shooting system, it ensured the stability of video recording and multi-angle coverage. The collected content includes standard scenarios such as news broadcasts and interviews, which feature standardized language expressions and behavioral patterns. Simultaneously, facial landmarks are annotated using the 68-point ASM standard, accurate to the 0.1 pixel level. This high-precision annotation method accurately captures subtle facial changes. Furthermore, speech-text alignment labels are generated simultaneously, providing important reference data for subsequent audio and video processing and analysis.

[0129] Publicly available data from video websites greatly enriches the dataset, complementing the diversity of real-world social scenarios. It contains 72 hours of publicly available video, covering over 50 languages ​​and resolutions ranging from 480p to 4K, encompassing video resources of various qualities and styles. The data includes real-world social scenarios such as vlog self-narration, live interactions, and short video announcements, simulating various non-ideal shooting conditions in real life, such as complex lighting conditions, dynamic backgrounds, and non-standard speaking postures. By incorporating this data, the model can better adapt to various complex shooting environments and character behaviors in the real world.

[0130] Large-Scale Film Dataset: This dataset is primarily used for artistic expression and emotional depth mining. It contains 56 hours of film video, selected from over 100 films and television works of different eras and genres, covering various formats and showcasing rich artistic styles and expressive techniques. The dataset includes 12 basic facial expressions and 36 combined facial expressions from professional actors, recording facial muscle dynamics through micro-expression capture technology, enabling the nuanced portrayal of characters' emotional changes. Furthermore, the dataset covers multi-camera language, film-level makeup and styling, and cross-cultural expression scenes, providing rich material for models to learn diverse emotional expressions and artistic representations.

[0131] To address the issue of inconsistent video frame rates in the aforementioned dataset, this embodiment first performs preprocessing to unify the frame rate. For videos with a frame rate lower than 25 frames per second, frame interpolation is employed. By analyzing the image information between adjacent frames, an algorithm generates appropriate intermediate frames to achieve the standard frame rate of 25 frames per second. For videos with a frame rate higher than 25 frames per second, redundancy removal is performed, eliminating duplicate or similar frames and adjusting the frame rate to 25 frames per second without affecting the integrity of the video content. After frame rate unification, the video is converted into images, providing the foundational data for subsequent processing.

[0132] Next, the Face-Alignment model is used for further image processing. This model first detects faces, using advanced computer vision algorithms to quickly and accurately identify facial regions in the image. Then, facial landmark detection technology is applied to locate 68 feature points of the face, including key positions of the eyes, nose, mouth, and eyebrows. The initial facial region is obtained by calculating the minimum bounding rectangle of the set of these feature point coordinates. To ensure complete inclusion of all key facial information, the boundaries of the initial region are expanded outwards to form extended regions. Finally, these extended regions are extracted from the original image as the face cropping result. To standardize image size, bilinear interpolation is used to scale the cropped region to a standard size of 512×512 pixels for subsequent feature extraction and processing.

[0133] (2) Text and image feature extraction:

[0134] For text, additional information describing facial expressions and dialogue content is added to the dataset. This information is processed using a pre-trained CLIP text encoder to generate 1024-dimensional semantic features, achieving a multi-level abstract representation of text semantics. Secondly, a sentiment classification auxiliary branch is introduced. By constructing a multi-task learning framework, sentiment features such as sentiment polarity (positive / negative / neutral) and tone intensity (1-5 level quantization) are extracted simultaneously, enhancing the model's ability to capture the semantic sentiment dimension and forming bidirectional text features containing both semantic content and sentiment attributes. Specifically, the process includes the following sub-steps:

[0135] (2-1) Input text is After processing by the CLIP text encoder, semantic features are obtained. The calculation method is as follows:

[0136] ,

[0137] in For CLIP text encoder, Represents the set of real numbers. This is the semantic feature dimension. Through this process, the semantic information in the text is transformed into numerical features that can be processed by a computer.

[0138] (2-2) To enhance the model's ability to capture the semantic sentiment dimension, the sentiment classification auxiliary branch of the BERT pre-trained model is used to extract the sentiment features of the text. This branch is based on the BERT pre-trained model. By constructing a multi-task learning framework, it simultaneously extracts sentiment features from the text, such as sentiment polarity (divided into positive, negative, and neutral) and tone intensity (quantized in 1-5 levels). It also extracts the sentiment features of the text simultaneously through cross-entropy loss. The calculation formula is as follows:

[0139] ,

[0140] in For the BERT model, This represents a multilayer perceptron sentiment classifier. For the emotional feature dimension;

[0141] (2-3) Combine the semantic features obtained in step (2-1) and the sentiment features obtained in step (2-2) to obtain bidirectional text features that contain both semantic content and sentiment attributes. The calculation formula is as follows:

[0142] ,

[0143] in The weight vector is used to balance the proportions of semantic and emotional features in the final text features by adjusting the weights. This semantic information, as content input, will be fused with visual and audio features in the multimodal attention module, providing rich semantic and emotional guidance for video generation.

[0144] (2-4) For the image obtained in step (1), the processed image Image features are extracted using a vector quantization variational autoencoder. The calculation formula is as follows:

[0145] ,

[0146] It is a vector quantization variational autoencoder that converts the image The continuous features obtained after compression It is an encoder function that converts an image into continuous features;

[0147] Encoder function This can be further represented as the multi-layer operations of a convolutional neural network:

[0148] ,

[0149] in, It is a convolutional layer. It is the Sigmoid activation function. The number of convolutional layers.

[0150] Its encoder structure converts the input image into low-dimensional latent features, providing input for the U-shaped network. The encoder employs a 5-layer convolutional network structure, with each layer equipped with a 3×3 convolutional kernel (stride 2). This structural design effectively extracts local features from the image. A BatchNorm layer and a LeakyReLU activation function are then cascaded. The BatchNorm layer accelerates the model's training process and improves training stability, while the LeakyReLU activation function introduces non-linear factors, enhancing the model's expressive power. Ultimately, the encoder outputs 64-dimensional latent features. In this process, by introducing KL divergence loss, a balance is achieved between the expressive power of the latent space and regularization, effectively suppressing overfitting and ensuring that the extracted image features have good generalization ability. Through this operation, the features are discretized, preserving not only key information such as the image's contours, facial expressions, and background structure, but also significantly reducing the data volume by removing redundant data, providing a more efficient and compact input for subsequent processing.

[0151] (3) Image feature reconstruction:

[0152] The text features and image features obtained in step (2) are first subjected to multiple iterations of noise through a diffusion process. Then, the text features and the noisy image features are fused together through a multimodal attention mechanism and fed into a U-shaped network for iterative denoising. After multiple iterations, the denoised image features are generated. This process is repeated multiple times to generate multiple reconstructed image features. Specifically, the process includes the following sub-steps:

[0153] (3-1) The image features obtained in step (2) By gradually adding noise, it is eventually transformed into pure Gaussian noise. This process is governed by a predefined noise schedule. Control, among which Indicates the first The noise intensity of the first step, the second step The calculation formula for the noise addition process in step 1 is as follows:

[0154] ,

[0155] ,

[0156] in For the first Features after adding noise For the first Features after adding noise It is random Gaussian noise. For the first Noise weighting of the step;

[0157] After recursive expansion, directly from the original image features Generate the first Features after adding noise:

[0158] ,

[0159] in The total noise weight can be expressed as: , It is the initial random Gaussian noise. Through this process, the original image features are gradually transformed into pure Gaussian noise, laying the foundation for subsequent denoising and feature reshaping.

[0160] (3-2) The first Features after adding noise The bidirectional text features obtained in step (2) By using a multimodal attention mechanism for fusion, the image generation process can understand text semantics. The calculation formula is as follows:

[0161] ,

[0162] in, This represents the function of the multimodal attention mechanism. For normalization function, For querying the matrix, Q comes from image features; For the key matrix and For value matrices, and Both originate from bidirectional textual features: ; It represents the dimension of the feature, and the superscript T indicates transpose. The fully connected layer representing bidirectional text features, through a multimodal fusion module, enables the image generation process to understand text semantics, organically combining text information with image features, and providing guidance for generating multi-frame video images of digital human facial expressions that conform to the text description.

[0163] (3-3) U-shaped network, that is, the network with the first... Features after adding noise Two-way text features and time step Given the input, predict the added noise. Gradually restore the original image features.

[0164] No. The denoising update step is calculated using the following formula:

[0165] ,

[0166] in The standard deviation of noise. For additional noise, The calculation formula is as follows;

[0167] ,

[0168] in, The noise added to the prediction, where, No. The features after adding noise are initially random noise. , It is a two-way text feature. For time steps;

[0169] This is the variance scheduling parameter, which controls the noise addition rate. It is set using a cosine function for linear scheduling, and the calculation formula is as follows:

[0170] ,

[0171] in, and These are the minimum and maximum values ​​of the variance, respectively. For time steps.

[0172] (3-4) Through multiple independent diffusion denoising loops, several different denoised image features that all conform to the text description are generated. Each loop starts from a different random noise starting point. Beginning, the first The denoised image features generated by the round are calculated using the following formula:

[0173] ,

[0174] ,

[0175] in For the first Denoising image features generated by the round, For the diffusion process described above, For the first The initial random noise of the wheel, This involves multiple iterations, where N is the number of iterations for image generation.

[0176] Finally, the images are concatenated to form a set of reconstructed image features after denoising. The calculation formula is as follows:

[0177] .

[0178] The training process of the image feature generation module is as follows:

[0179] D1: Preparing the Training Dataset: Collecting a large amount of text-image pair data, where the text describes the content and emotion of the image in detail, and the image is a corresponding digital human facial expression image. This data will serve as the basis for model training, enabling the model to learn the correspondence between text and images.

[0180] D2: Initialization Parameters: The parameters of the U-shaped network are randomly initialized, and a control network is used to lock the parameters of a large pre-trained model and copy its encoding layers, fully preserving the performance and capabilities of the large model, which serves as a powerful backbone network for learning various conditional controls. The trainable copy is connected to the original locked model through zero-weight convolutional layers. During training, the weights gradually increase, thereby improving the training effect and ensuring that the generated digital character image matches the emotions and behavioral tendencies implied by the text in the initial stage.

[0181] D3: Training: Input the text and images from the training dataset into the U-shaped network. Continuously adjust the model's parameters by minimizing the loss function between the generated and real images until the model converges. The loss function used is the mean squared error loss, which measures the pixel difference between the generated and real images. By optimizing this loss function, the model's generated images become closer to the real images.

[0182] (4) Audio generation: Based on the text features obtained in step (2), audio is generated using WaveNet, specifically through the following sub-steps:

[0183] (4-1) Two-way text features obtained from step (2) Including semantic and sentiment information, it needs to be transformed into a conditional embedding vector suitable for WaveNet through a mapping layer. The calculation formula is as follows:

[0184] ,

[0185] in For audio features, It is a linear transformation matrix. For conditional embedding dimensions.

[0186] (4-2) WaveNet captures long-range dependencies in audio sequences through multi-layer dilated causal convolutions, while also incorporating audio features Integrating into the calculation of each layer, the first The layer dilation convolution operation is calculated using the following formula:

[0187] ,

[0188] in For the first Layer expansion characteristics, For the first Layer expansion characteristics, For the first Layer dilation convolution operation, The kernel size is [size]. For the first Layer feature dimension The dilation factor is used. Through this dilated convolution operation, WaveNet can effectively capture long-range dependencies in audio signals, generating more natural and smooth audio.

[0189] Next, conditions are fused, and the calculation formula is as follows:

[0190] ,

[0191] in For the first Layer conditional fusion features It is the Sigmoid activation function. This indicates element-wise multiplication. It is the hyperbolic tangent function. These are the scaling parameters for the first batch. This is the offset parameter for the first batch. The scaling parameters for the second batch. This is the offset parameter for the second batch. From audio features Generated through linear layers:

[0192] ,

[0193] in For linear layers, For audio features, conditional fusion is used to organically combine the audio features with the features obtained after dilated convolution;

[0194] (4-3) The output of WaveNet is the probability distribution of audio samples, which is then converted into a continuous audio signal through an autoregressive generation process. The final output probability distribution is calculated as follows:

[0195] ,

[0196] in For probability distribution, For a moment Audio samples, This represents the total number of layers in WaveNet. For the final conditional fusion feature;

[0197] The audio signal is converted to a continuous audio signal through an autoregressive generation process, and the calculation formula is as follows:

[0198] ,

[0199] in For a moment Audio samples, The sample value with the highest conditional probability. Audio features;

[0200] Finally, the audio is spliced ​​together to form a continuous audio stream. The calculation formula is as follows:

[0201] ,

[0202] in For continuous audio, For a moment Audio samples.

[0203] (5) Image Generation: Multiple reconstructed image features obtained from step (3) need to be processed by the time alignment module before being sent to the decoder. Specifically, some image features are first specified from the overall image features as motion features, then noise is added to the motion features, and they are then spliced ​​with the current batch of image features along the time axis. After the above processing is completed, these image features are sent to the decoder to reconstruct multiple images. Then, the multi-frame images generated are repaired by the face restoration model GPEN, and finally, these images and audio are spliced ​​together in time order to generate a video. Specifically, the following sub-steps are included:

[0204] (5-1) Step (3) yields the reconstructed image feature set as follows: Each feature For the semantic representation of an image, let the motion features be... ,in, Indicates the first Motion characteristics of the image;

[0205] Adding noise to motion features, the formula is as follows:

[0206] ,

[0207] in, For the first Motion characteristics of an image after adding noise Indicates the first Motion characteristics of the image For noise;

[0208] The noise-added motion features are concatenated with the corresponding image features to obtain the input features for the decoder.

[0209] in, For the first Input features of the image This represents a connection operation along the feature dimension. For the first Zhang's denoised image features For the first Motion characteristics of an image after adding noise.

[0210] (5-2) The decoder upsamples the input features through multiple upsampling operations. Remodeling into feature maps The calculation formula is as follows;

[0211] ,

[0212] in In the The first image Layer sampling, For upsampling, The kernel size is [size]. For the first The upsampling factor of the layer, This represents the total number of upsampling layers;

[0213] Then, the RGB pixel values ​​are output through an activation function, calculated using the following formula:

[0214] ,

[0215] in For the generated first Zhang Image For the first The upsampling results of the images, This is the Sigmoid activation function.

[0216] (5-3) Finally, the face restoration module is used to restore the generated multi-frame video images. The face restoration model can enhance the clarity and details of the facial area, eliminate pixel noise, optimize the facial features of the digital character, and make the character image more realistic and natural. The calculation formula is as follows:

[0217] ,

[0218] For the first The high-resolution image after super-resolution processing. For super-resolution networks, , These are the parameters of the pre-trained super-resolution model;

[0219] Then, a detector is used to obtain the face bounding boxes for each image. And the coordinates of the key points, to obtain the affine transformation matrix. :

[0220] ,

[0221] For the first The affine transformation matrix of the image; A function for calculating the transformation parameters. For the coordinates of the eyes, nose tip, and mouth corners of the face in n images;

[0222] The facial features are extracted from the image through affine transformation, aligned to a standard pose, and scaled to a fixed size.

[0223] ,

[0224] For the first A face image after pose alignment. It is a quadratic mapping. For the first Affine transformation of the image;

[0225] Align the face input generator G to generate the repaired result:

[0226] ,

[0227] For the first Zhang's restored facial image For generator networks, These are the parameters of the pre-trained generator;

[0228] To ensure temporal consistency between adjacent frames, the optical flow field between adjacent images needs to be calculated first:

[0229] ,

[0230] in For the first Image number 1 The optical flow field of +1 facial image, Let be the optical flow field function. For the first Zhang's restored facial image For the first +1 restored face image.

[0231] The final repair result was generated through weighted fusion:

[0232] ,

[0233] For the first The final result of the face restoration. , These are the weighting coefficients. For optical flow-based pixel remapping operations, For the first Image number 1 The light flow field of a facial image;

[0234] Finally, the repaired and unrepaired regions are merged: first, by using inverse affine transformation. The final repair result The face pose is transformed back to the original image from the standardized pose, and then compared with the original image after super-resolution processing. Based on face bounding box Merge regions:

[0235] ,

[0236] For the first Zhang merged the restored face image with the original image background to create the complete image. A function for fusing images. For the function of changing posture, For the first The final result of the face restoration. For the first Inverse affine transformation of the image For the first The high-resolution image after super-resolution processing. The bounding box for the face.

[0237] By optimizing the above steps, the model can repair blurred facial features and remove blemishes. For example, for facial blurring caused by low resolution or the generation process, the model enhances the face by learning high-frequency detail features; for noise points and artifacts, it smooths them using contextual information. In particular, the model performs fine-tuning of eye areas such as eyelashes and iris details, mouth areas such as lip texture and tooth clarity, and skin texture, making facial expressions more vivid and natural.

[0238] (6) Video generation: stitching together multiple images and audio to generate a digital character video.

[0239] (7) System Result Evaluation. To scientifically evaluate the generated image and video quality, this embodiment uses Frachet distance and 16-frame video Frachet distance as evaluation metrics. Cosine similarity of facial identity, expression, and head movement FID and pose FID are measured respectively. Finally, the lip synchronization error distance is measured to achieve accurate evaluation of the audiovisual alignment effect. The evaluation results are shown in Table 1:

[0240] Table 1:

[0241]

[0242] This system can solve various problems existing in traditional video generation technologies. Through innovative modular architecture and algorithm design, it can achieve high-quality, realistic and smooth generation of digital character videos, meeting users' higher demands for digital character videos.

Claims

1. A method for generating digital character speaking videos via text-driven methods, characterized in that, The method includes the following steps: (1) For the videos involved in the digital person video dataset, video processing is performed: First, the original video is preprocessed by frame rate normalization; then, the preprocessed video is converted into images frame by frame; then, the Face-Alignment model is used to detect faces in the obtained images and the facial regions are cropped to obtain the processed images. (2) Text and image feature extraction: The specific feature extraction method for the additional text information used to describe the facial expressions and dialogue content of the characters in the digital character video is as follows: The pre-trained CLIP text encoder is used to process the text information to generate semantic features; at the same time, a sentiment classification auxiliary branch based on the BERT pre-trained model is introduced to extract the sentiment features of the text simultaneously; finally, the semantic features and sentiment features are fused to form a two-way text feature that has both semantic content and sentiment attributes. For the processed image obtained in step (1), the image features are extracted using a vector quantization variational autoencoder. (3) Image feature reshaping: The text features and image features obtained in step (2) are first added to the image features through a diffusion process, and then the text features and the image features with added noise are fused through a multimodal attention mechanism and fed into a U-shaped network for iterative denoising. After multiple iterations, the denoised image features are generated. This process is repeated to generate multiple reshaped image features. (4) Audio generation: Audio is generated using WaveNet based on the text features obtained in step (2); (5) Image generation: Multiple reconstructed image features are obtained from step (3). Before being sent to the decoder, they need to be processed by the time alignment module. Specifically, some image features are first specified from the overall image features as motion features. Then, noise is added to the motion features and they are spliced ​​with the current batch of image features along the time axis. After that, the image features processed by the time alignment module are sent to the decoder to reconstruct multiple frames of images. Then, the generated multiple frames of images are repaired by the face restoration module. (6) Video generation: The image repaired in step (5) and the audio generated in step (4) are spliced ​​together in chronological order to generate a digital character video; (7) Evaluation of generated results.

2. The method for generating digital character speaking videos via text-driven methods according to claim 1, characterized in that, The preprocessing operation of performing frame rate standardization on the original video in step (1) is aimed at the problem of inconsistent video frame rates in the HDTF dataset. All video frame rates are standardized to 25 frames per second. For video frames with a frame rate lower than 25 frames per second, frames are padded, and for video frames with a frame rate greater than 25 frames per second, redundancy is removed. In step (1), the Face-Alignment model is used to perform face detection on the obtained image and crop out the facial region. Specifically, the 68 feature points of the face are located by the facial key point detection technology in the Face-Alignment model. The minimum bounding rectangle of the coordinate set of these 68 feature points is calculated as the initial facial region. The boundaries of the initial facial region are expanded outward to form an extended region. The extended region is extracted from the original image as the face cropping result. Finally, the bilinear interpolation algorithm is used to scale the cropping region to a standard size of 512×512 pixels.

3. The method for generating digital character speaking videos via text-driven methods according to claim 1, characterized in that, Step (2) specifically includes the following sub-steps: (2-1) The additional text information used to describe the facial expressions and dialogue content of the digital characters in the dataset is processed using a pre-trained CLIP text encoder to generate semantic features: The input text is After processing by the CLIP text encoder, the semantic features of the text are obtained. The calculation method is as follows: , in For CLIP text encoder, Represents the set of real numbers. Semantic feature dimension; (2-2) Extracting sentiment features from text using the sentiment classification auxiliary branch of the BERT pre-trained model. The calculation formula is as follows: , in For the BERT model, This represents a multilayer perceptron sentiment classifier. For the emotional feature dimension; (2-3) Merge the semantic features obtained in step (2-1) and the sentiment features obtained in step (2-2) to obtain bidirectional text features containing semantic content and sentiment attributes. The calculation formula is as follows: , in This is the weight vector; (2-4) Based on the processed image obtained in step (1) Image features are extracted using a vector quantization variational autoencoder. The calculation formula is as follows: , It is a vector quantization variational autoencoder that converts the image The continuous features obtained after compression It is an encoder function that converts an image into continuous features; Encoder function This can be further represented as the multi-layer operations of a convolutional neural network: , in, It is a convolutional layer. It is the Sigmoid activation function. The number of convolutional layers.

4. The method for generating digital character speaking videos via text-driven methods according to claim 1, characterized in that, Step (3) specifically includes the following sub-steps: (3-1) The image features obtained in step (2) By gradually adding noise, it is eventually transformed into pure Gaussian noise. This process is governed by a predefined noise schedule. Control, among which Indicates the first The noise intensity of the first step, the second step The calculation formula for the noise addition process in step 1 is as follows: , , in For the first Features after adding noise For the first Features after adding noise It is random Gaussian noise. Let be the noise weights at step t; After recursive expansion, directly from the original image features Generate the first Features after adding noise: , in The total noise weight is expressed as: , It is the initial random Gaussian noise; (3-2) will the first Features after adding noise The bidirectional text features obtained in step (2) By using a multimodal attention mechanism for fusion, the image generation process can understand text semantics. The calculation formula is as follows: , in, This represents the function of the multimodal attention mechanism. Let Q be the normalization function, and let Q be the query matrix. Q comes from image features; K is the key matrix and V is the value matrix, both of which come from bidirectional text features. ; It represents the dimension of the feature, and the superscript T indicates transpose. A fully connected layer representing bidirectional textual features; (3-3) U-shaped network, that is, the network with the first... Features after adding noise Two-way text features and time step Given the input, predict the added noise. Gradually restore the original image features. No. The denoising update step is calculated using the following formula: , in The standard deviation of noise. For additional noise; The calculation formula is as follows; , in, Noise added to the prediction; No. The features after adding noise are initially random noise. ; It is a two-way text feature; For time steps; This is the variance scheduling parameter, which controls the noise addition rate. It is set using a cosine function for linear scheduling, and the calculation formula is as follows: , in, and These are the minimum and maximum values ​​of the variance, respectively. (3-4) Through multiple independent diffusion denoising loops, several different but consistent denoised image features are generated, each loop starting from a different random noise point. Beginning, the first The denoised image features generated by the round are calculated using the following formula: , , in For the first Denoising image features generated by the round, For the diffusion process described above, For the first The initial random noise of the wheel, This involves multiple iterations, where N is the number of iterations for image generation. Finally, the images are concatenated to form a reconstructed image feature set after denoising. The calculation formula is as follows: 。 5. The method for generating digital character speaking videos via text-driven methods according to claim 1, characterized in that, Step (4) specifically includes the following sub-steps: (4-1) Two-way text features obtained from step (2) Including semantic and sentiment information, it needs to be transformed into a conditional embedding vector suitable for WaveNet through a mapping layer. The calculation formula is as follows: , in For audio features, It is a linear transformation matrix. For conditional embedding dimensions; (4-2) WaveNet captures long-range dependencies in audio sequences through multi-layer dilated causal convolutions, while also incorporating audio features Integrating into the calculation of each layer, the first The calculation formula for layer dilation convolution is as follows: , in For the first Layer expansion characteristics, For the first Layer expansion characteristics, For the first Layer dilation convolution operation, The kernel size is [size]. For the first Layer feature dimension It is the expansion factor; Next, conditions are fused, and the calculation formula is as follows: , in For the first Layer conditional fusion features It is the Sigmoid activation function. This indicates element-wise multiplication. It is the hyperbolic tangent function. These are the scaling parameters for the first batch. This is the offset parameter for the first batch. The scaling parameters for the second batch. This is the offset parameter for the second batch. From audio features Generated through linear layers: , in For linear layers, Audio features; (4-3) The output of WaveNet is the probability distribution of audio samples, which is then converted into a continuous audio signal through an autoregressive generation process. The final output probability distribution is calculated as follows: , in For probability distribution, For a moment Audio samples, This represents the total number of layers in WaveNet. For the final conditional fusion feature; The audio signal is converted to a continuous audio signal through an autoregressive generation process, and the calculation formula is as follows: , in For a moment Audio samples, The sample value with the highest conditional probability. Audio features; Finally, the audio is spliced ​​together to form a continuous audio stream. The calculation formula is as follows: , in For continuous audio, For a moment Audio samples.

6. The method for generating digital character speaking video via text-driven method according to claim 1, characterized in that, Step (5) specifically includes the following steps: (5-1) Step (3) yields the reconstructed image feature set as follows: Each feature For the semantic representation of an image, let the motion features be... ,in, Indicates the first Motion characteristics of the image; The following formula is used to add noise to motion features: , in, For the first Motion characteristics of an image after adding noise Indicates the first Motion characteristics of the image For noise; The noise-added motion features are concatenated with the corresponding image features to obtain the input features for the decoder: , in, For the first Input features of the image This indicates a join operation along the feature dimension. For the first Zhang's denoised image features For the first Motion characteristics of an image after adding noise; (5-2) The decoder upsamples the input features through multiple upsampling operations. Remodeling into feature maps The calculation formula is as follows; , in In the The first image Second upsampling, For upsampling, The kernel size is [size]. For the first The upsampling factor of the layer, This represents the total number of upsampling layers; Then, the RGB pixel values ​​are output through an activation function, calculated using the following formula: , in For the generated first Zhang Image For the first The upsampling results of the images, Use the Sigmoid activation function; (5-3) Finally, the face restoration module is used to restore the generated multi-frame video images. The calculation formula is as follows: , For the first The high-resolution image after super-resolution processing. For super-resolution networks, These are the parameters of the pre-trained super-resolution model; Then, a detector is used to obtain the face bounding boxes for each image. And the coordinates of the key points, to obtain the affine transformation matrix. : , For the first The affine transformation matrix of the image; A function for calculating the transformation parameters. for Reference key points for the face in the image, such as the coordinates of the eyes, nose tip, and mouth corners; The facial features are extracted from the image through affine transformation, aligned to a standard pose, and scaled to a fixed size. , For the first A face image after pose alignment It is a quadratic mapping. For the first Affine transformation of the image; Align the face input generator G to generate the repaired result: , For the first Zhang's restored facial image For generator networks; These are the parameters of the pre-trained generator; To ensure temporal consistency between adjacent frames, the optical flow field between adjacent images needs to be calculated first: , in For the first Image number 1 The optical flow field of +1 facial image, Let be the optical flow field function. For the first Zhang's restored facial image For the first +1 restored face image; The final repair result was generated through weighted fusion: , For the first The final result of the face restoration. These are the weighting coefficients. For optical flow-based pixel remapping operations, For the first Image number 1 The light flow field of a facial image; Finally, the repaired and unrepaired regions are merged: first, by using inverse affine transformation. The final repair result The face pose is transformed back to the original image from the standardized pose, and then compared with the original image after super-resolution processing. Based on face bounding box Merge regions: , For the first Zhang merged the restored face image with the original image background to create the complete image. A function for fusing images. For the function of changing posture, For the first The final result of the face restoration. For the first Inverse affine transformation of the image For the first The high-resolution image after super-resolution processing. The bounding box for the face.

7. A text-driven digital character video generation system, characterized in that, This system is used to run the method described in any one of claims 1-6. The system includes a video processing module, a text and image feature extraction module, an image feature generation module, an audio generation module, and a video generation module. The video processing module first performs frame rate normalization preprocessing on the original video, detects faces using a Face-Alignment model and crops out facial regions, then adjusts the image to a 512×512 pixel image. The text and image feature extraction module extracts corresponding features based on text and image information. The image feature generation module adds and removes noise based on the extracted image features and fuses them with text features to obtain reconstructed image features. The audio generation module generates audio using WaveNet based on the text features. The video generation module, based on the reconstructed image features, uses a time alignment module to ensure that the generated image sequence has continuity over time. Finally, it stitches together multiple images and audio to construct a digital character video model, generates a fused digital character video, and evaluates the generation result.

Citation Information

Patent Citations

  • Character-driven AIGC video generation method and device

    CN119255064A

  • Digital human video generation method based on multi-modal large model

    CN120472059A