Voice-driven facial animation generation method

By introducing cross-modal fusion mechanism and depth residual enhancement mechanism in the speech-driven facial animation generation technology, the network structure and data set processing are optimized, and the problems of inaccurate lip synchronization and insufficient generation quality in the prior art are solved, and more efficient and high-quality facial animation generation is achieved.

CN120088373APending Publication Date: 2025-06-03XIAMEN UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510534569.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing speech-driven facial animation generation technology has problems such as inaccurate lip synchronization, poor image quality, and insufficient data set diversity and representation, especially when dealing with Chinese pronunciation.

Method used

A speech-driven facial animation generation method is adopted, and by introducing a cross-modal fusion mechanism and a depth residual enhancement mechanism, the network structure is optimized to improve the model's ability to extract and reconstruct complex features, and the training data set is optimized to adapt to multi-lingual environments.

Benefits of technology

It significantly improves the accuracy of lip synchronization and the naturalness and quality of generating animations, enhances the generalization ability and robustness of the model, and provides more efficient and high-quality technical solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088373A_ABST
    Figure CN120088373A_ABST
Patent Text Reader

Abstract

The invention provides a voice-driven facial animation generation method, and belongs to the technical field of facial animation generation. According to the method, on the basis of a traditional model, a cross attention mechanism and a deep residual enhancement mechanism are introduced, the fusion process of audio and visual features is optimized, the synchronism of lip shape driving is enhanced, and the user experience is improved. And the definition and the stability of the generated image are improved. Meanwhile, a pre-trained lip shape synchronization discriminator is improved, and a training data set is optimized, so that the lip shape synchronization discriminator is more in line with Chinese pronunciation characteristics, and the adaptability and robustness of the model in a multi-language environment are further improved. The method aims at remarkably improving the generation quality of the facial animation and the lip shape synchronization effect through an innovative network architecture and a data processing mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of facial animation generation, and particularly relates to a method for generating voice-driven facial animation. Background Art

[0002] With the rapid development of artificial intelligence and computer vision technologies, voice-driven facial animation generation technology has gradually become a research hotspot. This technology drives the generation of facial animations by analyzing voice signals, enabling virtual characters or digital humans to make corresponding lip movements and facial expressions according to the voice content, thereby achieving a more natural and vivid human-computer interaction experience. This technology has broad application prospects in fields such as virtual reality, intelligent education, the entertainment industry, and news broadcasting, and can significantly enhance the user's sense of participation and immersion.

[0003] However, there are still some problems to be solved urgently in the current voice-driven facial animation generation technology. First, the mapping relationship between voice signals and facial animations is complex and difficult to accurately model. Voice signals contain rich semantic information and emotional features, while facial animations need to accurately reflect these information, and the conversion between the two requires highly complex algorithm support. Second, the existing generation models have deficiencies in lip synchronization and generation quality. For example, the generated lip animations may have problems such as jitter, unnatural mouth shape changes, or mismatches with the voice content, affecting the user experience. In addition, the existing datasets have limitations in terms of diversity and representativeness. Especially when dealing with the pronunciation of specific languages (such as Chinese), the generation effect is often not ideal.

[0004] To solve these problems, researchers have been exploring new methods and technologies. For example, some studies have tried to introduce attention mechanisms to enhance the model's ability to capture voice signals and facial features, but these methods still have deficiencies in cross-modal feature fusion. Other studies focus on improving datasets to enhance the performance of the model by increasing the diversity and quality of data, but these methods still have a large room for improvement in enhancing lip synchronization and generation quality.

[0005] In view of this, the present application is proposed. Summary of the Invention

[0006] The present invention provides a method that can at least partially improve the above problems.

[0007] To achieve the above object, the present invention adopts the following technical solutions: A method for generating voice-driven facial animation, which includes: Obtaining an audio file and a facial file to be processed, and preprocessing the audio file to obtain a Mel spectrogram; Preprocess the facial file and mel spectrogram through cross-attention cross-modal fusion using the trained generator to generate the target facial image video; Use the pre-trained lip-sync discriminator to discriminate the temporal consistency and speech matching degree between the target facial image and the audio file, generate the discrimination result, and output the final facial animation according to the discrimination result.

[0008] In summary, a study was conducted on the technology of speech-driven facial animation generation, and an innovative generation method was proposed, namely a speech-driven facial animation generation method. The speech-driven facial animation generation method effectively solves the problems of inaccurate lip-sync and poor generated image quality existing in the prior art by introducing advanced feature fusion mechanisms and network optimization strategies. In terms of feature fusion, the method adopts a novel cross-modal interaction mechanism that can dynamically adjust the weights of audio and visual features, thereby achieving a more accurate lip-sync effect. At the same time, by optimizing the network structure, the model's ability to extract and reconstruct complex features is enhanced, significantly improving the clarity and stability of the generated facial animation. In addition, the present invention also optimizes the training dataset to better adapt to the multilingual environment, further improving the generalization ability and robustness of the model. After systematic experimental verification.

[0009] The speech-driven facial animation generation method shows significant advantages in multiple performance indicators, providing a more efficient and high-quality technical solution for fields such as virtual digital human interaction and intelligent education, and having important application value and broad development prospects. Description of the Drawings

[0010] Figure 1 is a schematic flowchart of the speech-driven facial animation generation method provided by an embodiment of the present invention; Figure 2 is a framework diagram of the speech-driven facial animation generation method provided by an embodiment of the present invention; Figure 3 is a structural diagram of the cross-attention mechanism provided by an embodiment of the present invention; Figure 4 is a structural diagram of the multi-head attention mechanism provided by an embodiment of the present invention; Figure 5 is a schematic diagram of the overall structure of the deep residual enhancement mechanism provided by an embodiment of the present invention; Figure 6 is a structural diagram of the SE network provided by an embodiment of the present invention; Figure 7 is a structural diagram of the generator encoder network provided by an embodiment of the present invention; Figure 8 is a structural diagram of the generator network decoder provided by an embodiment of the present invention; Figure 9 It is the discriminator network architecture diagram provided by an embodiment of the present invention; Figure 10 It is the schematic diagram of data preprocessing before model training provided by an embodiment of the present invention. Detailed implementation manners

[0011] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0012] Refer to Figure 1 、 Figure 2 As shown, a first embodiment of the present invention discloses a method for generating speech-driven facial animations, which can be executed by a speech-driven facial animation generation device (hereinafter referred to as the generation device), and particularly, executed by one or more processors in the generation device to implement the following method: S1. Obtain an audio file and a facial file to be processed, and preprocess the audio file to obtain a Mel spectrogram; Specifically, step S1 includes: obtaining an audio file to be processed, and extracting and processing the audio signal in the audio file by using an acoustic feature extraction module; Performing conversion processing on the extracted audio signal to obtain a Mel spectrogram.

[0013] Preferably, the facial file includes a random facial reference frame and a real facial semi-mask synchronization frame.

[0014] In this embodiment, the speech-driven facial animation generation method is designed based on a baseline model, aiming to eliminate technical bottlenecks such as lip jitter and mouth shape mismatch that occur during facial animation generation. Its overall architecture consists of two core modules, a generator and a lip synchronization discriminator. The main task of the generator is to generate a matching lip movement sequence based on complex speech features; the lip synchronization discriminator is mainly responsible for evaluating the generation effect, comparing the generation result with real data to ensure high accuracy and synthesis effect of the generated lip sequence; finally, the lip movement is returned to the facial area of the target image.

[0015] Specifically, in this embodiment, first obtain an audio file to be processed, and use an acoustic feature extraction module to extract and process the audio signal in the audio file. The audio file can be any audio file containing speech information, such as recorded human voices or synthetic speech. This module can accurately separate speech-related features from complex audio signals and provide accurate data input for subsequent conversion processing. This process can not only remove irrelevant noise interference, but also enhance the key features of the speech signal, thereby improving the robustness and accuracy of the entire system.

[0016] Subsequently, the extracted audio signal is processed through conversion to obtain a Mel spectrogram, aiming to complete the structural characterization process of the audio signal. The Mel spectrogram is a representation method that can effectively represent the spectral characteristics of the audio signal and is widely used in the field of speech processing. By converting the audio signal into a Mel spectrogram, it is possible to better capture the semantic information and emotional characteristics of the speech, providing a rich semantic basis for subsequent facial animation generation. This conversion process involves a series of complex signal processing techniques, including the application of the short-time Fourier transform and the Mel filter bank, ensuring that the spectrogram can accurately reflect the frequency distribution and temporal variations of the audio signal.

[0017] In addition, the acquisition of facial files is equally crucial. Facial files include random facial reference frames and real facial semi-mask synchronization frames. The random facial reference frames can be any facial images, providing the basic structure and appearance information of the face. The real facial semi-mask synchronization frames are a special type of facial image that, through masking technology, highlights the lower part of the face, especially the lip region, enabling the model to focus more on lip shape changes during animation generation. This combination of facial files not only provides rich facial details but also guides the model to focus on key areas through the masking mechanism, thereby improving the accuracy and naturalness of the generated animation.

[0018] Based on this, it is possible to effectively acquire and preprocess audio and facial data, laying a solid foundation for subsequent facial animation generation. This process not only ensures the quality and usability of the data but also improves the efficiency and performance of the entire system by optimizing the data processing flow.

[0019] S2. Use the trained generator to perform cross-attention cross-modal fusion preprocessing on the facial file and the Mel spectrogram to generate a target facial image video; Specifically, step S2 includes: using an audio encoder to perform multi-layer neural network extraction processing on the Mel spectrogram to obtain audio features; Using a facial encoder to perform multi-layer neural network extraction processing on the facial file to obtain facial visual features; Based on the cross-attention mechanism, perform cross-modal fusion processing on the audio features and the facial visual features. At the same time, enhance the feature expression ability according to the multi-head attention mechanism to obtain fused features; Use a facial decoder to perform multi-layer non-linear transformation processing on the fused features to generate a target facial image video.

[0020] In this embodiment, it first involves the extraction of audio features. An audio encoder is used to perform multi-layer neural network extraction processing on the Mel spectrogram. The audio encoder gradually extracts the deep semantic information in the Mel spectrogram through a series of convolutional layers and activation functions, and finally obtains the audio features. This process can not only capture the key frequency and time features of the speech signal, but also enhance the expressive ability of the features through the non-linear transformation of the multi-layer network. The accurate extraction of audio features is crucial for subsequent cross-modal fusion, providing a solid foundation for generating natural and smooth lip animations.

[0021] At the same time, the extraction of facial visual features is also carried out synchronously; a facial encoder is used to perform multi-layer neural network extraction processing on the obtained facial files. The facial encoder also extracts deep visual features from the facial images through multi-layer convolutional operations. These features include not only the overall structural information of the face, but also pay special attention to the details of key areas such as the lips. Through techniques such as spectral normalization and the Leaky ReLU activation function, the facial encoder can effectively improve the stability and distinctiveness of the features. The precise extraction of facial visual features enables the generator to better understand the structure and details of the input face, thus maintaining the naturalness and consistency of the face during the generation process.

[0022] After the extraction of audio features and facial visual features is completed, cross-modal fusion processing is performed on the two based on the cross-attention mechanism. The cross-attention mechanism precisely regulates the contribution degrees of audio and visual features by dynamically allocating attention weights. This mechanism allows the generator to "intelligently" select the most relevant key information in the audio features and facial information, and assigns different attention weights to the features of different regions of the face. In this way, the generator can ensure the consistency between the lip animation and the speech content during the generation process, significantly improving the accuracy of lip synchronization. In addition, the introduction of the multi-head attention mechanism further enhances the feature expression ability, enabling the output features to dynamically focus on the associations between different features. The multi-head attention mechanism divides the feature space into multiple independent subspaces, and the model can parallelly capture the interaction relationships between audio and visual features in different subspaces. This multi-head attention structure not only improves the feature expression ability, but also enhances the model's ability to capture complex features, making the generated facial animation more delicate and natural.

[0023] Finally, a face decoder is used to perform multi-layer non-linear transformation processing on the fused features to generate a target facial image video, and skip connections are used to retain early features. The face decoder gradually restores the spatial dimensions of the feature map through a series of upsampling operations and convolutional layers, and finally generates a facial image video synchronized with the input audio. During the decoding process, a deep residual enhancement mechanism is introduced. Through standard convolutional operations, squeeze-and-excitation networks, and residual connection modules, the spatial detail performance of the generated images is effectively improved, and the information loss during feature transmission is reduced, thereby enhancing the ability to retain image details. This mechanism not only reduces the information loss during feature transmission but also enhances the ability to retain image details, making the generated facial animations clearer and more stable in visual quality.

[0024] Through the operations in step S2, the efficient fusion of audio and visual features and the generation of high-quality facial animations are achieved. The introduction of the cross-attention mechanism and the multi-head attention mechanism significantly improves the accuracy of lip synchronization and the naturalness of the generated animations. The application of the deep residual enhancement mechanism further optimizes the detail performance and stability of the generated images. The organic combination of these technical features makes this method significantly innovative and practical in the field of speech-driven facial animation generation, and can provide high-quality technical solutions for fields such as virtual digital human interaction and intelligent education.

[0025] Please refer to Figure 7 , preferably, an audio encoder is used to perform multi-layer neural network extraction processing on the Mel spectrogram to obtain audio features, specifically: Taking the Mel spectrogram as the input of the audio encoder, and using multi-layer convolutional operations to gradually extract the audio features of the Mel spectrogram to obtain audio features; Among them, progressive multiple convolutional layers are used to gradually increase the channel dimension during the extraction process. At the same time, batch normalization and ReLU activation functions are used to ensure the stability of extraction and non-linear expression.

[0026] In this embodiment, the generation network structure consists of an encoder and a decoder, which can effectively complete animation generation or image reconstruction tasks, and encode and decode features to achieve the reconstruction and transformation of complex information. Among them, the encoder structure is responsible for compressing and abstractly representing the input data, using convolutional operations to gradually reduce the spatial dimension layer by layer, and retaining the extracted audio or visual features in the deep network.

[0027] Specifically, in this embodiment, the Mel spectrogram obtained through preprocessing is used as the input of the audio encoder. The Mel spectrogram is a representation method that can effectively represent the spectral characteristics of audio signals. It is widely used in the field of speech processing and can capture the key frequency and time characteristics of speech signals. The audio encoder extracts audio features from the Mel spectrogram layer by layer through multi-layer convolutional operations. This process adopts a progressive design of multiple convolutional layers. Each layer of convolutional operation can not only extract deeper features but also gradually increase the channel dimension, thus enriching the expressive ability of features. This design of gradually increasing the channel dimension enables the model to gradually extract high-level semantic information from low-level frequency features, providing a rich semantic basis for subsequent feature fusion.

[0028] During the extraction process, in order to ensure the stability and non-linear expression of the extraction, batch normalization and ReLU activation function are particularly adopted. Batch normalization is an effective regularization technique that can reduce internal covariate shift, accelerate the training process, and improve the stability and generalization ability of the model at the same time. The ReLU activation function introduces non-linearity into the model, enabling the model to learn more complex feature mapping relationships. By combining batch normalization and ReLU activation function, the audio encoder can not only stably extract audio features but also effectively express the non-linear features in audio signals, thus improving the model's understanding ability of speech signals.

[0029] This optimized audio feature extraction method has significant beneficial effects. First, by gradually increasing the channel dimension layer by layer, the model can capture richer audio features, thereby improving the accuracy of lip synchronization and the naturalness of generated animations. Second, the use of batch normalization and ReLU activation function not only improves the stability of feature extraction but also enhances the non-linear expression ability of the model, making the generated facial animations more realistic and natural in visual effects. This design can also significantly improve the training efficiency of the model, reduce the training time, and make the entire system more efficient and practical.

[0030] In addition, after convolutional operations, the residual structure is used multiple times to effectively alleviate the problem of gradient disappearance in deep networks, making the model more robust in the process of learning audio feature representation. After multiple non-linear transformations, audio features are gradually encoded into high-dimensional feature vectors K and V provide audio information for the subsequent cross-modal cross-attention mechanism.

[0031] Preferably, a face encoder is used to perform multi-layer neural network extraction processing on the face file to obtain face visual features, specifically: Use the facial file as the input of the facial encoder, perform layer-by-layer convolutional processing on the facial file, and transition the low-level image texture features of the facial file to high-level semantic information; Jointly use spectral normalization and the Leaky ReLU activation function to process the facial file after convolutional processing to obtain facial visual features.

[0032] Specifically, in this embodiment, the preprocessed facial file is used as the input of the facial encoder. The facial file can be a random facial reference frame or a real facial semi-mask synchronization frame, and these files provide the basic structure and appearance information of the face. The operation process of the facial encoder follows multi-scale extraction, but an optimization method is introduced for image feature processing on the network. Through layer-by-layer convolutional processing, it gradually transitions from low-level image texture features to high-level semantic information, realizing a deep understanding of the complex space of the facial image. Each layer of convolutional operation can not only extract deeper features but also gradually increase the channel dimension, thereby enriching the expressive ability of the features. This layer-by-layer extraction design enables the model to gradually extract high-level semantic information from low-level texture features, providing a rich semantic basis for subsequent feature fusion.

[0033] During the convolutional processing, in order to ensure the stability of extraction and non-linear expression, spectral normalization and the Leaky ReLU activation function are jointly used; multi-layer residual connections are used to enhance the information transmission ability, enabling the facial encoder to capture more detailed and discriminative facial image features. Spectral normalization is an effective regularization technique that can stabilize the training process, reduce the problems of gradient explosion and gradient disappearance, and at the same time improve the generalization ability of the model. The Leaky ReLU activation function introduces non-linear factors into the model, enabling the model to learn more complex feature mapping relationships. By combining spectral normalization and the Leaky ReLU activation function, the facial encoder can not only stably extract facial visual features but also effectively express the non-linear features in the facial image, thereby improving the model's understanding ability of the facial image. The deep feature vector output Q Provides visual speech information for subsequent cross-modal fusion. The cross-attention mechanism is used as a mediator to connect audio and facial features, realizing the dynamic integration of multi-modal information.

[0034] This optimized method for facial visual feature extraction has significant beneficial effects. First, through layer-by-layer convolutional processing, the model can capture richer facial visual features, thereby improving the naturalness and detail performance of the generated animation. Second, the use of spectral normalization and the Leaky ReLU activation function not only improves the stability of feature extraction but also enhances the model's non-linear expression ability, making the generated facial animation more realistic and natural in visual effects. In addition, this design can significantly improve the training efficiency of the model, reduce the training time, and make the entire system more efficient and practical.

[0035] Please refer to Figures 3 to 4 Preferably, cross-modal fusion processing is performed on audio features and facial visual features based on the cross-attention mechanism. Meanwhile, the feature expression ability is enhanced according to the multi-head attention mechanism to obtain fused features, specifically as follows: Map the facial visual features I ∈ R C*H*W and audio features A ∈ R C*T into a feature space containing queries, keys, and values. The mathematical expression is: , , , where W q , W k , W v are all learnable linear mapping matrices. Among them, R is the set of real numbers (all eigenvalues are real numbers), C is the number of channels (the visual channel number refers to the input color information. If an RGB image is input, the number of channels is 3. The audio channel number refers to the number of channels of the input audio. If stereo audio is input, the number of channels is 2), H is the height of the feature map, W is the width of the feature map, T is the length of the time series; Calculate the attention distribution according to Q and K , and generate a fused feature representation V by weighting according to . Among them, d k is the dimension of K ; Based on the multi-head attention mechanism, the feature space is divided into h independent subspaces, and the interaction relationship between audio and visual features in different subspaces is captured in parallel. The mathematical expression is: . Among them, W o is a linear mapping matrix; Extract and aggregate features in multiple subspaces to obtain fused features , where a are learnable weight parameters.

[0036] Speech features and visual facial image features essentially belong to heterogeneous information spaces. How to effectively connect and deeply fuse these two different types of modal features has been widely discussed in the fields of computer vision and multimodal learning. Traditional feature fusion methods directly fuse features of different modalities using simple feature concatenation or concatenation methods, which cannot effectively capture the complex semantic associations and dynamic dependencies between cross-modal features. This results in the generated lip shape sequence being unable to accurately express the fine-grained changes in speech. To solve this problem, the present invention proposes the speech-driven facial animation generation method.

[0037] Specifically, in this embodiment, in the actual facial animation generation task, the generator is responsible for mapping and converting audio information and facial images into a synchronous and accurate lip movement sequence. To achieve high-quality speech-driven facial animation generation, the cross-modal fusion process of audio features and facial visual features is particularly optimized. This process is one of the core links of the entire system. Through a carefully designed cross-attention mechanism and multi-head attention mechanism, it is ensured that audio features and facial visual features can be efficiently fused, thereby providing rich semantic information and accurate feature expressions for subsequent facial animation generation. Among them, the advantage of the cross-attention mechanism for cross-modal interaction is to use dynamic and adaptive allocation of attention weights to accurately regulate the contribution degrees of visual and audio modal features. The cross-attention mechanism enables the generator to "intelligently" select the most relevant key information in audio features and the current facial information, and at the same time assigns different attention weights to features in different regions of the face, which can ensure the consistency between the lip animation and the speech content during the generation process.

[0038] Specifically, the cross-modal fusion processing process is as follows: First, to effectively implement cross-modal feature mapping, this method adopts a linear transformation strategy; map facial visual features and audio features into a feature space containing Query (abbreviated as Q), Key (abbreviated as K), and Value (abbreviated as V). This mapping process is achieved through a learnable linear mapping matrix. Through this mapping, facial visual features and audio features can interact and fuse in a unified feature space, align in terms of dimensions, and provide a basis for subsequent attention mechanism calculations.

[0039] Immediately afterwards, according to the mapped Q and K calculate the attention distribution, and according to VWeighted generation of fused feature representations. This process is achieved through the scaled dot - product attention mechanism. In this way, the model can dynamically allocate attention weights and precisely regulate the contribution degrees of audio and visual features. This dynamic allocation mechanism enables the model to "intelligently" select the most relevant key information from audio features and facial information, thus ensuring the consistency between lip animations and speech content during the generation process. This operation allows the model to dynamically adjust the attention weights of different features, thereby more precisely capturing the subtle features of audio - video synchronization.

[0040] After calculating the attention distribution and weighted fusion calculation to complete the process of the scaled dot - product attention mechanism, to further improve the feature representation ability, the present invention introduces a multi - head attention structure in the cross - attention mechanism module. Among them, the multi - head attention mechanism divides the feature space into h independent sub - spaces and simultaneously captures the interaction relationships between audio and visual features in different sub - spaces. This process is achieved through a linear mapping matrix. Through the multi - head attention mechanism, the model can extract and aggregate features in multiple different representation sub - spaces, thereby enhancing the expressive ability of the generator.

[0041] After extracting and aggregating features in multiple sub - spaces, fused features are obtained. To balance the retention of original visual features and cross - modal information fusion, the cross - attention module adopts a weighted fusion strategy method. The main role of this weighted fusion strategy is to control the influence degree of the cross - attention mechanism on the original features; it not only retains the key features of the original facial image but also can introduce adaptive audio semantic information, thus achieving a more natural and smooth facial animation effect during the generation process.

[0042] Please refer to Figure 5 、 Figure 6 、 Figure 8 , preferably, use a facial decoder to perform multi - layer non - linear transformation processing on the fused features to generate the target facial image video. Specifically: Perform de - convolution processing on the fused features to improve the spatial resolution and reconstruct each layer of features; Based on the up - sampling method, expand the spatial information to generate the target facial image video. Among them, a depth residual enhancement mechanism DREM is introduced during the up - sampling process to achieve multi - level feature fusion. Specifically: Based on the convolutional layer, use standard convolutional operations to extract the incoming local spatial information and capture the global information of the facial image; According to the squeeze - excitation SE network, perform self - adaptive calibration on the feature channels, and use a residual connection to add the original image features and the processed feature maps element - by - element. Among them, the squeeze operation F sq realizes the compressed description of global context information, and the mathematical expression is The excitation operationF ex Using a two-layer fully connected network to learn the dependencies between channels, the mathematical expression is , U The input feature map X The transposed feature map, Z c is the global description of the cth channel, s is the channel weight, δ is the ReLU activation function, σ is the Sigmoid activation function, W 1 is the weight of the first fully connected layer, W 2 is the weight of the second fully connected layer; According to the channel weight s For weighted processing, the mathematical expression is: Ũ c =s c *U c , the original features are recalibrated at the channel level.

[0043] In this embodiment, the decoder structure is responsible for restoring the feature details, and uses upsampling operations such as transposed convolution to restore the spatial size of the feature map layer by layer; the encoder and decoder structures can effectively capture the input data structure and retain key information during the generation process. The audio sequence and dynamic semantic information extracted by the audio encoder and the visual space image information extracted by the facial encoder are fused through a cross-attention mechanism to provide richer features for subsequent decoders. As the back part of the generation network, the decoder completes the task of restoring the multimodal encoded features into facial images layer by layer. By adding deconvolution operations and deep residual enhancement mechanisms, the image restoration capability is effectively improved.

[0044] The feature reconstruction of the decoder is a process of progressive multi-scale information recovery. Specifically, the processing process of the facial decoder is as follows: First, the feature representation output by the encoder is received, the fused features are deconvolved to improve the spatial resolution of the feature map, and each layer of features is reconstructed. The deconvolution operation is an effective upsampling method that can gradually restore the spatial size of the feature map and maintain the continuity of high-dimensional information, thereby generating facial image videos synchronized with the input audio. To enhance feature expression, in each layer of feature reconstruction operation, the network uses an upsampling method to effectively expand spatial information while maintaining the continuity of high-dimensional information; it uses complex mechanisms to deeply explore the nonlinear expression capabilities of features, alleviate the gradient vanishing in the training process, and improve the naturalness of the generated facial images. This progressive multi-scale information recovery process makes the generated facial images more natural and clear in visual quality.

[0045] Furthermore, to improve the training stability of the generator and enhance the feature expression ability, a Deep Residual Enhanced Mechanism (DREM for short) is designed. DREM uses a multimodal feature fusion method to comprehensively integrate the squeeze-and-excitation network, efficient convolution operations, and cross-scale feature interaction modules to enhance the fine-grained processing of visual features during facial animation generation. To improve the output ability of the decoder for complex feature generation targets, a squeeze-and-excitation (SE) network, standard convolution operations, residual connection modules, and pointwise convolutions are fused in DREM to design a multi-level feature enhancement framework.

[0046] In this embodiment, the advantage of the deep residual enhancement mechanism is to use continuous convolution operations and feature weighting modules to extract and dynamically adjust the input features. In the decoding stage of the decoder, using DREM can effectively improve the spatial detail performance of the generated images. Specifically, during the upsampling process, a deep residual enhancement mechanism is introduced to achieve multi-level feature fusion. DREM extracts the incoming local spatial information through standard convolution operations to capture the global information of the facial image. This convolution operation can not only expand the receptive field but also enhance the feature expression ability. Then, according to the squeeze-and-excitation network, the feature channels are adaptively calibrated, and channel-level feature recalibration is used to dynamically weight each feature channel to enhance the key features.

[0047] The SE network realizes the compressed description of the global context information through the compression operation, and the excitation operation uses a two-layer fully connected network to learn the dependencies between channels. The cooperation of compression and excitation is used to adaptively enhance the key features; in this way, the SE network can dynamically adjust the channel weights, highlight the key channels, and suppress redundant information, improving the feature expression ability.

[0048] Simply put, this operation is used to compress each number of channels into a scalar to achieve a global channel description; the global average pooling operation is used to perform a global low-dimensional embedding operation on the feature map. To enable the two fully connected layers to learn the correlations between channels, the network uses activation functions to adjust the channel weights. Among them, the excitation operation uses a two-layer fully connected network to model the channel dependencies. The first fully connected network compresses the number of channels to reduce the computational complexity, while the second fully connected network restores the number of channels and dynamically adjusts the weights using the ReLU and Sigmoid activation functions.

[0049] Weighted processing is performed according to channel weights to re-calibrate the original features at the channel level, effectively highlighting key channels and suppressing redundant information; this weighted processing can not only enhance key features but also retain the key information of the original image. Finally, residual connection is used to add the original image features and the processed feature map element by element, thus achieving effective information transmission. Residual connection not only alleviates the problem of gradient disappearance in the training of deep networks, enables feature fusion to be achieved at multiple levels, improves the overall quality of the generated image, but also enhances the stability of training. The decoder finally realizes the restoration of the image from low dimension to high dimension, making the generated facial animation not only more natural in visual quality but also more advantageous in detail representation ability.

[0050] In addition, DREM also introduces multiple optimization strategies to improve the overall efficiency. Using pointwise convolution to perform linear transformation in the channel dimension can reconstruct the feature channel distribution, reduce the model complexity, and promote multi-channel information interaction; using residual connection to combine multi-module characteristics and retain the original input information, after skip connection, the input features are added to the output features after convolution and SE network to form residual information. This operation alleviates the problem of gradient disappearance in the training of deep networks and also enhances the stability of training. The design of DREM improves the ability of the facial animation generation task. The SE network realizes the dynamic and accurate selection of key features, and the standard convolution and pointwise convolution operations enhance the spatial-channel expression ability. The residual connection provides structural support for the deep optimization of the network, and the multi-dimensional feature enhancement mechanism makes the generator process cross-modal features more intelligently and efficiently.

[0051] Generally speaking, this optimized facial decoder design has significant beneficial effects: First, through deconvolution operation and upsampling method, the model can gradually restore the spatial size of the feature map and generate high-quality facial image videos. Second, the introduction of the Deep Residual Enhancement Mechanism (DREM), through standard convolution operation and SE network, significantly improves the expression ability and detail performance of features. In addition, the use of residual connection not only improves the stability of training but also enhances the model's ability to process complex features. The organic combination of these technical features makes the generated facial image videos more realistic and natural in visual effects.

[0052] Please refer to Figure 9 S3, use the pre-trained lip-sync discriminator to discriminate the temporal consistency and speech matching degree between the target facial image and the audio file, generate a discrimination result, and output the final facial animation according to the discrimination result.

[0053] Specifically, step S3 includes: using a facial encoder to extract facial features from the target facial image video to obtain facial image embedding features , an audio encoder is used to extract audio semantic features from the target facial image video to obtain audio embedding features , wherein, the facial encoder only processes and extracts the lower half region containing lip movement information, and the audio encoder extracts a spectrogram through STFT and then extracts audio semantic features through a multi-layer CNN; The cosine similarity between the audio embedding and the facial image embedding is calculated by minimizing the loss function to obtain the synchronization probability between the audio embedding and the facial image embedding , to quantify the synchronization between the audio and the facial animation frames, and its calculation formula is: , wherein, is a constant to prevent the denominator from being zero; The cosine similarity is converted to the interval [0, 1], wherein the cosine similarity The closer it is to 1, the higher the synchronization between the audio and the facial animation frames, and vice versa, the lower the synchronization between the audio and the facial animation frames; The binary cross-entropy loss function is used to optimize the parameters of the lip synchronization discriminator to obtain , wherein, y i is the true label, y i = 0 indicates out-of-sync, y i = 1 indicates in-sync, N is the batch data volume; Based on the optimized lip synchronization discriminator, the synchronization loss L sync and the reconstruction loss L 1 are used as the loss function of the generator to optimize the generator, wherein, , , is the output image of the generator, is the true image of the generator.

[0054] In this embodiment, a lip synchronization discriminator is used to evaluate the synchronization relationship between the generated lip shape and the input audio, and the generation process is continuously optimized by means of pre-training to ensure that the final generated result is natural and smooth. To achieve the goal, the lip synchronization discriminator needs to deeply extract both audio and facial visual features to realize the dynamic correspondence between the lip movement and the speech content in the facial animation.

[0055] Use a pre-trained lip-sync discriminator to discriminate the temporal consistency and speech matching degree between the target facial image video and the audio file. This process involves extracting facial features and audio semantic features from the target facial image video. Specifically, the lip-sync discriminator continues the CNN-based architecture of the Wav2Lip model and consists of two parts: a facial encoder and an audio encoder. The discriminator has a similar structure to the generator. The facial encoder is used to extract facial features from the target facial image video to obtain facial image embedding features. Considering the particularity of the lip-sync task and to improve the training efficiency, the facial encoder only processes and extracts the lower half region containing mouth shape information during this process. This design not only reduces the computational complexity but also improves the pertinence of feature extraction. At the same time, the audio encoder is used to extract audio semantic features from the target facial image video to obtain audio embedding features. The audio encoder extracts the spectrogram through the short-time Fourier transform (STFT), and then extracts the audio semantic features through a multi-layer convolutional neural network (CNN). Among them, residual connections are added after convolution for both encoders to ensure the stability of transmission. After feature extraction, high-dimensional audio embeddings and facial embeddings are generated respectively, representing facial dynamic information and audio semantic content. These embeddings serve as the basis for similarity calculation and are used to evaluate the synchronization of lip shape and audio.

[0056] To improve the robustness of the lip-sync discriminator, pre-train it on the improved and optimized dataset so that the discriminator can learn cross-modal features and speech-lip mapping relationships. The advantages of pre-training are as follows: the feature embeddings learned during the pre-training process can provide a stable foundation for the generator; the diverse data distribution can enhance the model's adaptability to different scenarios; it alleviates overfitting and accelerates the convergence speed of the model during generator training.

[0057] Furthermore, as a pre-trained model, the lip-sync discriminator evaluates the synchronization relationship between the audio and the facial animation frames and guides the generator training. To quantify the synchronization between the audio and the facial animation frames, the cosine similarity between the audio embedding and the facial image embedding is calculated by minimizing the loss function. To map the similarity to the probability space, the synchronization probability of the audio embedding and the facial image embedding is obtained to represent the synchronization correlation. The cosine similarity is converted to the interval [0,1]. The closer the cosine similarity is to 1, the higher the synchronization between the audio and the facial animation frames; conversely, the lower the synchronization.

[0058] To optimize the parameters of the lip-sync discriminator, the binary cross-entropy loss function is used to optimize the parameters of the lip-sync discriminator. By minimizing the loss function, the synchronization probability between the audio embedding and the image embedding is calculated, and the model gradually learns the temporal alignment relationship between the audio and the facial animation. After the lip-sync discriminator is trained, the parameters are fixed and the weights are saved to guide the training of the generator. That is, based on the optimized lip-sync discriminator, the synchronization loss and the reconstruction loss are used as the loss functions of the generator to optimize the generator. The synchronization loss uses the pre-trained lip-sync discriminator to evaluate the quality of the generated results. The trained lip-sync discriminator outputs the synchronization probability as a "supervision tool" to provide an optimization direction for the generator. By minimizing the synchronization loss and maximizing the generator objective, it is ensured that the generated facial animation frames and the input audio are in good synchronization. The reconstruction loss is used to measure the difference between the generated image and the real image. The reconstruction loss calculates the mean absolute error at the pixel level to make the generated effect of the generator closer to the visual output of the real image.

[0059] It can be seen from this that two requirements are mainly considered when using the multi-objective loss function in facial animation generation. The synchronization loss ensures the temporal consistency between the audio and the vision, and the reconstruction loss ensures the visual quality of the generated image. The dual constraints improve the efficiency.

[0060] This optimized design of the lip-sync discriminator has significant beneficial effects: 1. Through the pre-trained lip-sync discriminator, the temporal consistency and speech matching degree between the generated facial animation and the audio can be accurately evaluated, thus providing effective feedback for the generator. 2. Using the cosine similarity and the binary cross-entropy loss function, the synchronization between the audio and the facial animation frames can be quantified, and the parameters of the discriminator can be optimized, enabling the model to better learn the mapping relationship between the audio and the facial animation during the training process. 3. The combination of the synchronization loss and the reconstruction loss not only ensures the temporal consistency between the audio and the vision, but also ensures the visual quality of the generated image, making the generated facial animation more realistic and natural in terms of visual effects.

[0061] Please refer to Figure 10 , preferably, before performing the cross-attention cross-modal fusion preprocessing on the facial file and the mel spectrogram using the trained generator, it also includes obtaining a training data set and preprocessing the training data set. Specifically: Obtain multiple Chinese speech data sample videos, and perform screening processing on the Chinese speech data sample videos to screen out the videos with unqualified facial orientation angles and unqualified facial clarity; Integrate the remaining Chinese speech data sample videos, and on this basis, add other sample videos to combine and obtain a data set; Perform temporal segmentation on the dataset, dividing it into short segments with a duration of 2 seconds, setting the duration to randomly fluctuate within 0.2 seconds and ensuring that there are accurate words in the video. The sampling frame rate of the video data is 25fps, and the sampling frequency for audio data feature extraction is 16KHz; Divide the video segments into a training set, a validation set, and a test set according to a ratio of 7:2:1 for training and testing the generator and the lip-sync discriminator; Resample the audio files in the dataset and use the short-time Fourier transform to map the time-domain signal of the audio to the frequency-domain space. Use Mel filters to map the linear spectrum to the Mel scale to generate a high-dimensional Mel spectrogram representation, obtaining an audio training dataset.

[0062] In deep learning tasks, a high-quality and representative training dataset is the key to the performance of model applications. Currently, mainstream datasets in the audio-visual field include LRW-1000, LSR2, etc. Although they provide rich video data, there are the following main limitations in application: the facial features of some videos are blurred, and the side video angles are single; the main construction scenarios of the datasets are in English, without considering the Chinese language pronunciation structure, and the synthesized Chinese animation effect is poor; datasets such as CMLR contain Chinese content, and the samples are from news broadcast scenarios, and the single scenario cannot meet the needs of animation generation. These limitations result in poor facial animation generation effects under Chinese audio drive.

[0063] Specifically, in this embodiment, to overcome application limitations, before the trained generator is used to perform cross-attention cross-modal fusion preprocessing on the facial file and the Mel spectrogram, it is first necessary to obtain and preprocess the training dataset. Based on the dataset LSR2, systematic data optimization and expansion work has been carried out: First, obtain multiple Chinese speech data sample videos. These video samples are the basis for system training and cover Chinese speech content of different speakers, different speech rates, and intonations. To ensure data quality, these Chinese speech data sample videos are screened, and videos with unqualified facial orientation angles and unclear facial clarity are screened out. This screening process is completed through computer vision technology combined with manual verification to ensure that the remaining video samples meet the requirements in terms of facial features and speech quality. This strict screening mechanism not only improves the overall quality of the dataset but also provides more accurate inputs for subsequent model training, helping to improve the naturalness and synchronization of the generated animations. In this step, high-quality Chinese speech data videos are specifically introduced, including this year's news broadcasts and content recorded by professional laboratory equipment, to ensure that the videos cover the characteristics of different subjects, speech rates, and intonations, and to supplement English sample videos such as YouTube White House speeches to enhance the generalization ability of the model and make the data closer to diverse real-world scenarios. In addition, low-quality video screening and removal are also carried out. During the screening process, factors such as facial orientation angles and facial clarity are considered, and computer vision screening methods are combined with manual verification to remove samples with unclear faces and inappropriate perspectives.

[0064] On the basis of screening, the remaining Chinese speech data sample videos (accounting for 20%) are integrated, and other sample videos, such as English sample videos, are added on this basis. The total duration is 30 hours to enhance the diversity and generalization ability of the dataset. The SyncNet model method is used to filter out unqualified samples using a set threshold to achieve audio-visual alignment of the face, and all videos are standardized in a unified format to ensure the stability of data quality. By integrating video samples of different languages and scenarios, a comprehensive dataset is constructed. The construction of this diverse dataset enables the model to be trained in multiple languages and scenarios, thereby improving the adaptability and robustness of the model and enabling it to better handle the task of speech-driven facial animation generation in different environments. After optimizing the dataset, improvements have been made in terms of quantity scale, sample quality, and diverse scenarios, providing data support for the model training of facial animation generation and enhancing the driving effect and adaptability in Chinese scenarios.

[0065] Next, in the facial animation generation task, preprocessing and alignment operations on audio-visual data are essential steps to achieve stable lip synchronization. The data processing operations will be systematically carried out using a similar processing method as the baseline model to precisely correspond audio features and visual features in the time dimension, providing an effective data basis for model training. Specifically, the integrated dataset is segmented in time series. The video is divided into short segments with a duration of 2 seconds, with the duration randomly fluctuating within 0.2 seconds, and ensuring that each segment contains accurate words. This time series segmentation method not only ensures data consistency but also provides a finer time resolution for model training, helping the model better learn the temporal relationship between speech and facial movements. The sampling frame rate of video data is set to 25fps, while the sampling frequency for audio data feature extraction is set to 16KHz. A face detection and alignment algorithm is used to extract and align all samples. The selection of these parameters is based on current best practices in video and audio processing, ensuring a high-quality representation of data in both time and frequency domains, providing a solid foundation for subsequent feature extraction and model training. This ensures stable input data during training, guaranteeing efficient training and consistent video quality.

[0066] Finally, the segmented video segments are divided into a training set, a validation set, and a test set according to a ratio of 7:2:1, which contains more than 40,000 sample segments, with a cumulative number of frames reaching 1.75 million. This division ratio has been verified through multiple experiments, aiming to ensure that the model can fully learn data features during the training process, while accurately evaluating and optimizing the model's performance through the validation set and the test set. The training set is used for the model training process, the validation set is used to adjust the model's hyperparameters and prevent overfitting, and the test set is used to finally evaluate the model's performance. This scientific data division method not only improves the model training efficiency but also ensures the model's generalization ability on unseen data, providing a reliable performance guarantee for the practical application of the system.

[0067] Data preprocessing operations can reduce the alignment complexity of the model during training. Using training samples in both Chinese and English can enhance the completeness of speech feature representation, and can improve the distribution balance, richness of feature expression, and Chinese phoneme pronunciation of the constructed data samples.

[0068] Furthermore, preprocessing is also required before training the audio information. The audio file of the source video is loaded at a sampling rate of 16KHz for resampling to ensure the restoration of the speech signal while meeting the requirements of the lip-sync time resolution. In terms of frequency-domain feature extraction, the short-time Fourier transform is used to map the time-domain signal to the frequency-domain space, and then the Mel filter is used to map the linear spectrum to the Mel scale to generate a high-dimensional Mel spectrogram representation, which not only retains the key spectral features of the speech signal but also achieves precise alignment with the video frames with a fixed time resolution. The stable temporal mapping relationship between the audio feature sequence and the video feature sequence frames is achieved by calculating the frame number sequence of the input video in combination with the audio window size and the frame shift parameter.

[0069] The preprocessing operation before training the video information is completed by differentiating different model components. The generator first extracts the video sequence of consecutive frames from the training dataset, performs unified cropping and scaling operations to ensure the spatial consistency of the input facial images, and uses the lower face region masking mechanism to set the pixels in the lower part of the face to zero, forcing the model to focus its attention on the generation task of the lip region. Performing such operations not only ensures the uniformity of the input data but also effectively eliminates the interference of visual features in non-target regions on model training. The lip-sync discriminator model uses the lower face region cropping method to precisely retain the region below the center of the vertical image pixels, enabling the model to pay more attention to the synchronization of the lip-sync actions and reducing the influence of non-critical regions such as the eyes and forehead on the discrimination process. In terms of temporal alignment, precise synchronization between the two is achieved using the frame number and frame rate mapping mechanism, and an index relationship is established between the video sequence and the audio features to ensure the alignment of the audio segment and the corresponding video frame in the time dimension. Using the temporal alignment mechanism can provide a real and reliable evaluation benchmark for the lip-sync discriminator to accurately evaluate the synchronization degree between the lip actions and the speech content in the face; it can also provide a stable discrimination signal for the optimization of the generator. Using the audio window design can ensure the extraction quality while maximizing the temporal alignment accuracy.

[0070] In this embodiment, a multi-dimensional evaluation framework is constructed with reference to the baseline model ReSyncED evaluation system. The purpose is to comprehensively evaluate the generation efficiency of the model using diverse test data. The video tests of the evaluation framework are extensive and representative, capable of simulating real application scenarios. First, test cases are generated. The test videos of the multi-dimensional evaluation framework are generated in a hierarchical cross manner. The specific process is as follows: At the visual input level, 8 source contents with different facial features (4 videos and 4 images) are selected to help evaluate the generalization ability of the model for different visual features; at the audio input level, 10 test audios with a duration of more than 10 seconds are selected, including real-scene human voice audios and audios synthesized by text-to-speech (TTS) technology, to help evaluate the stability of the model in long-sequence generation tasks and its adaptability to different audio features; in terms of language distribution, a mixed evaluation method with a Chinese-English ratio of 2:3 is used to help verify the performance of the model under different speech features and lip shape transformations; finally, the 8 source contents and 10 test audios are fully combined to form an 8×10 complete test matrix, generating a total of 80 different driving facial animation test cases to ensure the reliability and comprehensiveness of the experimental results.

[0071] During the experimental evaluation, the same 80 test cases are used to evaluate the performance of each comparative model. A complete evaluation system that combines comprehensive objective and subjective evaluation indicators is used: The objective evaluation indicators focus on lip synchronization evaluation and facial generation quality evaluation; the subjective evaluation, based on other research findings, designs four-dimensional evaluation indicators. Comprehensive evaluation can analyze the performance of the model in different application scenarios, enabling quantitative evaluation of the performance of each model in different dimensions and providing guidance for model improvement.

[0072] To further obtain the objective evaluation indicator for lip synchronization, in terms of the synchronization of synthetic facial animations, two of the most widely used lip synchronization evaluation indicators are used for performance evaluation: Lip Sync Error Distance (LSE-D) and Lip Sync Error Confidence (LSE-C). These two indicators quantitatively evaluate the matching degree of audio and video from different dimensions. LSE-D calculates the Euclidean distance between the facial image frame features and the corresponding audio features in the depth space to measure synchronization. The smaller LSE-D is, the higher the consistency between the frame where the image sequence is located and the corresponding audio features in the time domain and frequency domain. LSE-C reflects the credibility of the generation result. The implementation is to first extract the average distance mdist of the facial sequence frames: determine which minimum distance minval and its corresponding frame index, and use Calculate the difference between the median and the minimum distance. The higher the LSE-C, the more audio feature information exists in a specific frame, indicating a higher lip-sync accuracy of the generated result at that time point. By analyzing the evaluation metrics LSE-C and LSE-D, we can understand the performance of the model in aligning audio and visual features, and thus improve the algorithm and training method more specifically.

[0073] Obtain the objective evaluation metrics for the quality of animation generation. The evaluation of the quality of facial animation generation measures the performance of the generated images from multiple perspectives. The metrics used include key features such as image sharpness, texture details, contrast, and information entropy, which reflect the generation effect from multiple perspectives. The main metrics are as follows: The sharpness (Brenner) metric is used to evaluate the sharpness of the edges of an image. This metric quantifies the sharpness characteristics of the graph by calculating the gray difference between adjacent pixels. The calculation formula is . In the calculation formula, I(i, j) represents the gray value of the image at (i, j) coordinates. The higher the Brenner value, the better the sharpness of the image edges and details, and the more able to express the detailed information of the face. The energy (Energy) evaluation metric is used to evaluate the uniformity and richness of image texture. The higher the Energy value, the richer and more realistic the image texture. The calculation formula is .

[0074] The variance value (Variance) of the contrast evaluation metric is used to measure the dispersion degree of image pixel values. The calculation formula is , μ represents the average gray value of the image, MN represents the total number of pixels, Variance The higher it is, the better the contrast of the image, and the more able to highlight the levels of facial key points. The squared mean of the gray differences (SMD2) metric is used to evaluate the sharpness of local regions of an image. This metric calculates the product of the differences between pixels and their neighboring pixels to evaluate the sharpness of partial regions. The calculation formula is . The information entropy (Variance) metric is used to evaluate the information richness of an image, which can reflect the complexity and detailed performance of the image. The calculation formula is , where P i represents the probability that the gray value i appears. Using the lip-sync evaluation metric and the generation quality evaluation metric can comprehensively evaluate the performance of the generated animation. The above metrics can dynamically reflect the ability to maintain details, providing a quantitative basis for subsequent improvement of model performance.

[0075] Finally, subjective evaluation indicators are obtained. Objective evaluation indicators can quantitatively reflect the performance ability of the model, but relying solely on objective data to measure the effect cannot comprehensively capture the subtle differences in the human perception system. Therefore, a subjective evaluation method is designed to supplement the objective evaluation system. In subjective evaluation, 68 evaluators with different backgrounds were recruited to subjectively measure the generation status. The evaluation team included postgraduate students in the field of computer vision, professional video producers, and ordinary users. Based on a systematic review of the most widely used evaluation methods in the existing technology, subjective evaluation indicators in four evaluation dimensions were proposed: lip-sync accuracy, image clarity, facial naturalness, and overall visual quality.

[0076] The specific design of the four evaluation dimensions is as follows: Lip-sync accuracy evaluates the visual consistency between the audio content and the lip movement, focusing on the matching degree between the lip movement and the speech features; Image clarity evaluates the fusion status between the generated area and the original image and the detail fidelity; Facial naturalness evaluates the continuity and stability of facial movements, focusing on whether there are artifacts such as deformation or temporal incoherence during generation; Overall visual quality comprehensively evaluates the realism and immersion of the animation generation. To ensure the consistency of subjective evaluation, the members of the evaluation team were given a simple standardized training to explain the key points and detailed criteria for scoring each indicator. In the actual evaluation, the evaluation members needed to score 5 test samples generated by different models on a 5-point scale, with a higher score indicating a better effect. The test samples were presented in a random order and blind testing was performed during the evaluation to reduce the influence of the order effect and ensure the accuracy and objectivity of the evaluation.

[0077] In this embodiment, the performance of this method is verified and analyzed. The computer environment for facial animation generation is as follows: The hardware configuration uses an AMD Ryzen7 7900X processor and an ASUS GeForce RTX 4090 GPU. The software environment uses Python3.8 for programming and development, the deep learning framework uses PyTorch 1.11.0, and the CUDA acceleration library version is 11.3 to ensure the efficient execution of model training and inference. The pre-training strategy is adopted in the training process of the lip-sync discriminator, and the number of iterations is 12×10 5 . The cosine similarity is used as the loss function metric standard during the training process, and the loss curve gradually converges. At 6×10 5 steps, the loss steadily drops below 0.3 and gradually stabilizes below 0.2 later, indicating that the model's recognition ability for lip-sync features reaches the expected level. The training cycle of the generator model is 600, the batch size is 64, the Adam optimizer is selected as the optimization algorithm, and the L1 loss and Lsync loss are used as the loss functions to ensure visual quality and synchronization accuracy.

[0078] The evaluation of facial animation generation performance is carried out from four dimensions: visual effect display, visual quality assessment, evaluation of the impact of dataset optimization, and quantification of lip synchronization. The experimental data consists of 80 video samples generated by a data evaluation system and is processed and analyzed through a systematic data evaluation process. The trained model is used as the inference model for visual display. It can be seen from the images that the generated effect is excellent in terms of pronunciation accuracy. Under the dual-context drive of Chinese and English, the opening and closing of the character's mouth shape is relatively natural and smooth, meeting the human pronunciation standards. The following will evaluate the generation effect in combination with different indicators.

[0079] To evaluate the quality of animation generation, multi-dimensional evaluation indicators are used to evaluate the quality of the generated face. Before the evaluation, the S3FD algorithm is used to accurately detect and extract the facial area of the generated animation to ensure that all experimental samples are exactly the same in terms of image facial range, spatial resolution, and time series length, so as to obtain more reliable evaluation results. In the quantification evaluation stage, Brenner, Energy, Variance, SMD2, and Entropy are used to evaluate the quality of the generated facial images. The table shows the comparison results of the model and the baseline model on these indicators.

[0080] Table 1 Comparison results of image generation quality and baseline

[0081] It can be seen from the data in the table that the improved model shows performance advantages in all generation quality evaluation indicators. The improvements in the Brenner and Energy indicators are 2.25% and 3.08% respectively, and the improvements in the Variance, SMD2, and Entropy indicators are 1.74%, 2.4%, and 1.68% respectively. The data shows that the generated face is more plump compared to the baseline model, and the complexity of face generation has also increased. When generating facial details, the improved model has also enhanced the image restoration degree in the lip area and other aspects of the generation effect, effectively proving the role of using cross-modal feature interaction, enhanced features, and optimized datasets.

[0082] The two most commonly used indicators of lip synchronization, LSE-C and LSE-D, are used to evaluate the impact of dataset optimization and lip synchronization quality, and to measure the consistency between the lip movement in the generated facial animation and the input audio time sequence. The higher the value of the LSE-C indicator, the better the lip synchronization effect; the lower the value of the LSE-D indicator, the smaller the lip synchronization deviation and the more the lip movement fits the audio. Training is carried out on the LSR2 dataset and the improved dataset respectively. The results obtained from different datasets are shown in Table 2.

[0083] Table 2 Comparison results of lip synchronization for different datasets

[0084] After comparison on different datasets, the proposed model has been improved compared with the baseline models. Good training results have been obtained by training on the improved dataset, which has improved by 4.48% compared with the baseline model LSE-C and 5.18% compared with LSE-D. The improved dataset has laid a foundation for lip synchronization in generating facial animations.

[0085] To comprehensively verify the effectiveness of the method proposed in the present invention, comparative experiments have been carried out by comparing it with multiple representative open-source models and baseline methods under audio driving in different scenarios. The comparison results of multiple models are shown in Table 3 and Table 4.

[0086] Table 3 Lip synchronization effect under natural scene audio driving

[0087] Table 4 Comparison results of lip synchronization effect under TTS audio driving

[0088] Table 3 shows the experimental results under natural scene audio driving. It can be observed from the data that the index of the proposed model on LSE-C reaches 9.38, an increase of 4.8% compared with the baseline model of 8.95; LSE-D decreases to 6.48, a decrease of 4.84% compared with the baseline, and it has more advantages than other open-source model results. The data shows that the proposed cross-attention mechanism and additional residual structure improve the level of the model to capture cross-modal features and achieve more accurate lip synchronization. Table 4 is the comparison result in the synthetic speech driving scenario, and the method used in the present invention still shows obvious advantages. The TTS driving of all models is lower than the natural scene driving effect. The reason may be that synthetic speech usually lacks some subtle features in natural speech, but the improved model still maintains the optimal performance in both scenarios, proving its robustness.

[0089] An ablation experiment will be designed to evaluate the effects of various components in the facial animation generation model and analyze the influence of the cross-modal cross-attention mechanism and the deep residual enhancement mechanism on the baseline model. The experiment will keep the video frames as the input source, adopt the same evaluation method as the comparative experiment, and introduce PSNR as an additional evaluation index to evaluate the fidelity of the generated animation. Incorporating PSNR as the main evaluation in the ablation experiment is mainly because this index can more accurately measure the coherence of the generated results in the time dimension and is suitable for the evaluation scenario of continuous video sequences rather than static images. Table 5 shows the quantitative analysis results of the ablation experiment, comparing the performance of the baseline model, the model with the deep residual enhancement mechanism removed, the model with the cross-attention mechanism removed, and the complete model.

[0090] Table 5 Results of ablation experiment

[0091] After ablation experiments, the complete model achieved improvements in all metrics compared to the baseline model. When the cross-attention mechanism was removed, the model performance decreased, indicating that this mechanism plays an important role in the complex association between the audio and visual modalities; after removing the deep residual enhancement mechanism, LSE-C and LSE-D also decreased, proving the effectiveness of this mechanism in feature extraction ability and maintaining visual details; when neither or only a single component was removed, the effect was better than the baseline model, demonstrating the role of each component in achieving the lip synchronization task.

[0092] In summary, the speech-driven facial animation generation method significantly improves the quality of facial animation generation and lip synchronization effect through an optimized network architecture and data processing flow. This method not only achieves innovation in technology but also shows significant beneficial effects in practical applications, providing a high-quality technical solution for fields such as virtual digital human interaction and intelligent education.

[0093] In terms of technical implementation, the speech-driven facial animation generation method first optimizes the extraction process of audio and facial features. Through progressive convolutional layers and batch normalization techniques, the audio encoder can extract deep semantic information in the Mel spectrogram layer by layer, while the facial encoder extracts rich visual features from facial files through spectral normalization and the Leaky ReLU activation function. This optimized feature extraction method not only improves the expressive ability of features but also enhances the model's ability to capture complex features, providing a solid foundation for subsequent cross-modal fusion. Further, the cross-attention mechanism and multi-head attention mechanism are introduced to achieve efficient fusion of audio features and facial visual features. By dynamically allocating attention weights, the model can "intelligently" select the facial features most relevant to the speech content, thus ensuring a high degree of synchronization between the lip animation and the speech during the generation process. The introduction of the multi-head attention mechanism further enhances the feature expression ability, making the generated facial animation more delicate and natural in visual effects.

[0094] In terms of data processing, this method specifically optimizes the construction and preprocessing process of the training dataset. Through a strict screening mechanism, video samples with unqualified facial orientation angles and unclear facial clarity are removed, ensuring the high quality of the dataset. At the same time, by integrating video samples of multiple languages and scenarios, a diverse dataset is constructed, significantly improving the generalization ability and adaptability of the model. In addition, through temporal segmentation and scientific data partitioning, the present invention provides high-quality inputs for model training, further enhancing the overall performance of the system.

[0095] In terms of the optimization of the generator and the lip-sync discriminator, a deep residual enhancement mechanism and a binary cross-entropy loss function are adopted. The deep residual enhancement mechanism effectively improves the spatial detail performance of the generated images through standard convolution operations, squeeze-and-excitation networks, and residual connection modules, making the generated facial animations clearer and more stable in visual quality. The lip-sync discriminator can accurately evaluate the temporal consistency and speech matching degree between the generated facial animations and the audio through an optimized loss function and pre-training strategy, providing effective feedback for the generator and further improving the synchronization of the generated animations.

[0096] Generally speaking, the speech-driven facial animation generation method realizes high-quality speech-driven facial animation generation through a series of technological innovations and optimizations. The optimized feature extraction method, cross-modal fusion mechanism, data processing flow, and the designs of the generator and discriminator not only improve the quality and synchronization of the generated animations but also enhance the generalization ability and adaptability of the system. The organic combination of these technical features makes the present invention have important application values and broad development prospects in the fields of virtual digital human interaction, intelligent education, etc.

[0097] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A speech-driven facial animation generation method, characterized in that: include: Acquire an audio file and a facial file to be processed, and preprocess the audio file to obtain a mel-spectrogram; The trained generator is used to perform cross-attention cross-modal fusion preprocessing on the facial file and the Mel-spectrogram to generate the target facial image video; The pre-trained lip sync discriminator is used to discriminate the timing consistency and voice matching between the target facial image and the audio file, generate a discrimination result, and output the final facial animation based on the discrimination result.

2. The method for generating facial animation driven by speech according to claim 1, characterized in that: Get the audio file to be processed, pre-process the audio file, and obtain a Mel-spectrogram, specifically: Acquire an audio file to be processed, and use an acoustic feature extraction module to extract and process an audio signal in the audio file; The extracted audio signal is converted to obtain a Mel-spectrogram.

3. The voice-driven facial animation generation method according to claim 1, characterized in that: The face file includes a random face reference frame and a real face half-mask synchronization frame.

4. The method for generating facial animation driven by speech according to claim 1, characterized in that: The pre-trained generator is used to perform cross-attention cross-modal fusion preprocessing on the facial file and the Mel-spectrogram to generate the target facial image video, specifically: Using an audio encoder to perform a multi-layer neural network extraction process on the Mel-spectrogram to obtain audio features; Using a facial encoder to perform multi-layer neural network extraction processing on the facial file to obtain facial visual features; Based on the cross-attention mechanism, the audio features and facial visual features are cross-modally fused. At the same time, the feature expression ability is enhanced according to the multi-head attention mechanism to obtain the fused features. The facial decoder is used to perform multi-layer nonlinear transformation on the fused features to generate the target facial image video.

5. The method for generating facial animation driven by speech according to claim 4, characterized in that: The audio encoder is used to perform multi-layer neural network extraction processing on the Mel-spectrogram to obtain audio features, specifically: Using the Mel-spectrogram as the input of an audio encoder, and extracting audio features of the Mel-spectrogram layer by layer using a multi-layer convolution operation to obtain audio features; Among them, multiple progressive convolutional layers are adopted to increase the channel dimension layer by layer during the extraction process. At the same time, batch normalization and ReLU activation function are used to ensure the stability and nonlinear expression of the extraction.

6. The method for generating facial animation driven by speech according to claim 4, characterized in that: The facial encoder is used to perform multi-layer neural network extraction processing on the facial file to obtain facial visual features, specifically: Using the facial file as input of a facial encoder, performing layer-by-layer convolution processing on the facial file to transition low-level image texture features of the facial file to high-level semantic information; The convolutional facial files are processed by combining spectral normalization and Leaky ReLU activation function to obtain facial visual features.

7. The method for generating facial animation driven by speech according to claim 4, characterized in that: Based on the cross-attention mechanism, the audio features and facial visual features are cross-modally fused. At the same time, the feature expression ability is enhanced according to the multi-head attention mechanism to obtain the fusion features, which are as follows: Facial visual features I∈R C*H*W and audio features A∈R C*T Mapped to the feature space containing query, key and value, the mathematical expression is: , , ,in, W q , W k , W v are all learnable linear mapping matrices, where R is the set of real numbers, C is the number of channels, H is the height of the feature map, W is the width of the feature map, T is the length of the time series; according to Q and K Calculate the attention distribution and use V Weighted generation of fusion feature representation ,in, d k for K Dimensions; Based on the multi-head attention mechanism, the feature space is divided into h independent subspaces, and the interactive relationship between audio and visual features in different subspaces is captured in parallel. The mathematical expression is: ,in, W o is the linear mapping matrix; Extract and aggregate features in multiple subspaces to obtain fused features ,in, a are learnable weight parameters.

8. The method for generating facial animation driven by speech according to claim 7, characterized in that: Use the facial decoder to perform multi-layer nonlinear transformation processing on the fused features to generate the target facial image video, specifically: Deconvolution is performed on the fused features to improve the spatial resolution and reconstruct the features of each layer; The spatial information is expanded based on the upsampling method to generate the target facial image video. In the upsampling process, the deep residual enhancement mechanism DREM is introduced to achieve multi-level feature fusion, specifically: The convolutional layer uses standard convolution operations to extract the incoming local spatial information and capture the global information of the facial image; The feature channels are adaptively calibrated according to the compression-excitation SE network, and the original image features are added element by element to the processed feature map using residual connections. F sq To achieve the compressed description of global context information, the mathematical expression is: , incentive operation F ex Using a two-layer fully connected network to learn the dependencies between channels, the mathematical expression is , U The input feature map X The transposed feature map, Z c is the global description of the cth channel, s is the channel weight, δ is the ReLU activation function, σ is the Sigmoid activation function, W 1 is the weight of the first fully connected layer, W 2 is the weight of the second fully connected layer; According to the channel weight s For weighted processing, the mathematical expression is: Ũ c =s c *U c , the original features are recalibrated at the channel level.

9. The method for generating facial animation driven by speech according to claim 1, characterized in that: Use the pre-trained lip sync discriminator to discriminate the timing consistency and voice matching between the target facial image and the audio file, generate the discrimination result, and output the final facial animation according to the discrimination result, specifically: Use a facial encoder to extract facial features from the target facial image video to obtain facial image embedding features , use an audio encoder to extract audio semantic features from the target facial image video to obtain audio embedding features , where the facial encoder only processes and extracts the lower half of the area containing lip shape information, and the audio encoder extracts the spectrogram through STFT, and then extracts the audio semantic features through multi-layer CNN; The cosine similarity between the audio embedding and the facial image embedding is calculated by minimizing the loss function to obtain the synchronization probability of the audio embedding and the facial image embedding. , to quantify the synchronization between audio and facial animation frames, which is calculated as: ,in, To prevent the denominator from being zero; The cosine similarity Convert to the interval [0,1], where cosine similarity The closer it is to 1, the higher the synchronization between the audio and the facial animation frame. Conversely, the lower the synchronization between the audio and the facial animation frame. The binary cross entropy loss function is used to optimize the parameters of the lip sync discriminator and obtain ,in, y i is the true label, y i =0 means no synchronization, y i =1 means synchronization, N is the batch data size; Based on the optimized lip sync discriminator, the synchronization loss is used L sync and reconstruction loss L 1 is used as the loss function of the generator to optimize the generator, where , , Output image for the generator, is the real image of the generator.

10. The voice-driven facial animation generation method according to claim 1, characterized in that: Before using the trained generator to perform cross-attention cross-modal fusion preprocessing on the facial file and the Mel-spectrogram, it also includes obtaining a training dataset and preprocessing the training dataset, specifically: Acquire multiple Chinese voice data sample videos, and screen and process the Chinese voice data sample videos to remove videos with substandard facial orientation angles and substandard facial clarity; Integrate the remaining Chinese speech data sample videos, and add other sample videos on this basis to obtain the data set; The dataset is segmented into short clips of 2 seconds in length. The duration is randomly fluctuated at 0.2 seconds to ensure that there are accurate words in the video. The sampling frame rate of the video data is 25fps, and the sampling frequency of the audio data feature extraction is 16KHz. The video clips are divided into training set, validation set and test set according to 7:2:1 for training and testing the generator and lip sync discriminator; The audio files in the dataset are resampled, and the short-time Fourier transform is used to map the time domain signal of the audio to the frequency domain space. The Mel filter is used to map the linear spectrum to the Mel scale to generate a high-dimensional Mel spectrum graph representation to obtain the audio training dataset.

Citation Information

Patent Citations

  • Cross-modal multi-feature fusion audio and video speech recognition method and system

    CN112053690A

  • Speaking face video generation method and device based on convolutional neural network

    CN113378697A

  • Data fusion classification method based on residual extrusion excitation

    CN117853796A

  • Visual speech recognition method and system based on cross attention fusion

    CN118675526A

  • KR20250035420A