Audio-driven facial animation generation method, system, equipment and medium
By combining the methods of keyframe generation and interpolation, the diffusion model and pre-trained audio encoder are used to solve the problem of identity offset and quality reduction in audio-driven facial animation generation, and high-quality, long-term and coherent facial animation generation are achieved, improving the realism and time consistency of the video.
Patent Information
- Application Number
- CN202510696915.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing audio-driven facial animation generation methods have problems of identity offset and quality degradation in the long video generation process, making it difficult to maintain high-quality time consistency and long-term time dependence.
Using a method of combining keyframe generation and interpolation, a diffusion model and a pre-trained audio encoder are used to extract audio features, and a keyframe sequence is generated through audio attention and time step embedding, and a loss function optimization model of RGB space and potential feature space is combined to achieve smooth transition and temporal consistency.
Generate high-quality, long-term coherent and natural audio-driven facial animations, effectively capturing long-term time dependencies and improving the video's realistic and time consistency.
Smart Images

Figure CN120259503A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video generation, and in particular, to an audio-driven facial animation generation method, system, device, and medium. Background Art
[0002] With the development of generative models such as Generative Adversarial Networks (GANs) and Diffusion Models (DMs), significant progress has been made in the field of audio-driven facial animation. These methods have greatly enhanced the realism and expressiveness of facial animations, enabling promising applications in virtual assistants, education, virtual reality, and assisting people with communication disorders. As a result, there has been a sharp increase in high-resolution, natural, long-term audio-driven facial animations.
[0003] Although early methods of audio-driven facial animation were limited in terms of head rotation or only focused on generating the mouth region, current methods have evolved to produce results that are almost indistinguishable from real videos. Despite this progress, most methods struggle to handle long audio inputs, experiencing identity drift and a decline in overall quality after the initial few seconds. To extend the generation length, some methods incorporate additional spatial information, such as the target head position or landmarks, as model inputs. While this can improve temporal consistency, it confines the animation to predefined facial movements, limiting expressiveness. Other methods use motion skeletons to provide context for previous movements; however, many autoregressive methods suffer from error accumulation over time, reducing the overall quality. Summary of the Invention
[0004] The present invention aims to solve the problems of identity drift and quality decline during the generation of long videos, and provides an audio-driven facial animation generation method, system, device, and medium that combines key frame generation with interpolation to generate videos that maintain high quality over time and capture long-term temporal dependencies.
[0005] To this end, the present invention adopts the following technical solutions.
[0006] In a first aspect, the present invention provides an audio-driven facial animation generation method, which includes: Step 1), performing audio feature extraction using a combined embedding from two pre-trained audio encoders; Step 2), feeding the extracted audio features into the audio attention and time step embedding of a diffusion model; Step 3), conditioning on the audio input and identity frames, the diffusion model generates a sequence of key frames at a low frame rate; Step 4), conditioning on two consecutive frames in the key frame sequence, the diffusion model interpolates between the key frames; Step 5), optimize the diffusion model by combining the loss functions in the RGB space and the latent feature space.
[0007] In the first stage, conditioned on the identity frame and audio input, generate a sequence of key frames at a low frame rate, spanning multiple seconds and eliminating the need for motion frames. In the second stage, use interpolation to fill in the intermediate frames, ensuring smooth transitions and temporal consistency. The present invention divides the generation into two parts, implicitly separating motion and identity control, thereby producing more natural motion and improving identity preservation over time. For longer sequences, this process can be repeated, with interpolation generating seamless transitions between segments.
[0008] Furthermore, in the said step 1), the two pre-trained audio encoders are WavLM and BEATs. The WavLM captures the language content from speech, and the BEATs is trained to extract features from acoustic signals including non-speech sounds.
[0009] Even further, in the said step 2), the mechanism of feeding the audio features into the diffusion model includes: 2.1) Audio attention: Combine the embeddings As keys and values in the cross-attention layer within the U-Net architecture, enabling the diffusion model to focus on relevant audio features; Wherein, represents the WavLM pre-trained audio encoder, represents the BEATs pre-trained audio encoder; 2.2) Time step embedding: Add to the time step embedding So , represents the fused feature of the time step embedding and the audio features, and MLP represents a fully connected neural network. Further promote the diffusion model to align the image and audio frames and improve lip synchronization.
[0010] Furthermore, the specific process of generating the key frame sequence in the said step 3) is as follows: 3.1) Given a noisy input sequence , the goal is to generate a sequence where a person speaks in synchronization with the given audio; 3.2) To provide identity and background information, connect the identity frame with the noise input through the VAE encoder, effectively utilizing the skip connections of the U-Net architecture to retain input details; 3.3) Conditioned on the audio input and the identity frame, generate a key frame sequence of length at a low frame rate through the diffusion model, with an interval of Frames, these key frames effectively capture remote temporal dependencies and serve as anchors for subsequent interpolation stages.
[0011] Furthermore, the specific process of step 4) is as follows: 4.1) Obtain two consecutive frames from the key frame sequence and as conditional frames; 4.2) Create a sequence for matching the input shapes of the start and end frames, where represents the learned embedding of the missing frames; this sequence is concatenated with the noise input by channel and interpolated through the diffusion model to obtain coherent intermediate frames.
[0012] Even further, the specific process of optimizing the diffusion model in step 5) is as follows: 5.1) Decode the latent feature space back to the RGB space to obtain the decoded frame Apply the L2 loss between the decoded frame and the ground truth frame and add it to the L2 loss between the latent feature space and ; represents the ground truth feature space; 5.2) Optimize the parameters of the diffusion model during the backpropagation process.
[0013] In a second aspect, the present invention provides an audio-driven facial animation generation system for implementing the above audio-driven facial animation generation method, which includes: Audio feature extraction unit: Extract audio features using the combined embeddings from two pre-trained audio encoders; Audio feature feeding unit: Feed the extracted audio features into the audio attention and time step embeddings of the diffusion model; Key frame sequence generation unit: Conditional on the audio input and identity frames, the diffusion model generates a key frame sequence at a low frame rate; Interpolation unit: Conditional on two consecutive frames in the key frame sequence, the diffusion model interpolates between the key frames; Diffusion model optimization unit: Optimize the diffusion model by combining the loss functions in the RGB space and the latent feature space.
[0014] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above audio-driven facial animation generation method are implemented.
[0015] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above audio-driven facial animation generation method are implemented.
[0016] The beneficial effects of the present invention are as follows: By combining keyframe generation and interpolation, and utilizing the extended temporal context, the present invention can generate videos that maintain high quality over time and capture long-term temporal dependencies, effectively maintaining the temporal consistency and realism of long sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of a method for generating audio-driven facial animation according to the present invention; Figure 2 is a schematic diagram of the process of a method for generating audio-driven facial animation according to the present invention; Figure 3 is a composition diagram of a system for generating audio-driven facial animation according to the present invention; Figure 4 is a schematic diagram of a logical structure of an electronic device provided in the specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The present invention will be further described and explained below with reference to the accompanying drawings of the specification and the specific embodiments.
[0019] Embodiment 1 As Figure 1 and Figure 2 shown, this embodiment is a method for generating audio-driven facial animation, including the following steps: Step 1), extracting audio features using the combined embeddings from two pre-trained audio encoders.
[0020] The two pre-trained audio encoders are WavLM and BEATs. WavLM captures the linguistic content from speech, and BEATs is trained to extract features from a wider range of acoustic signals (including non-speech sounds). Let represent the WavLM pre-trained audio encoder, represent the BEATs pre-trained audio encoder.
[0021] Step 2), feeding the extracted audio features into the audio attention and time step embeddings of the diffusion model.
[0022] On the one hand, the combined embedding serves as the keys and values in the cross-attention layer within the U-Net architecture, enabling the diffusion model to focus on relevant audio features; on the other hand, is added to the diffusion time step embedding thereby , represents the fused feature of the time step embedding and the audio feature. MLP represents a fully connected neural network. This further promotes the diffusion model to align the image and audio frames and improves lip synchronization.
[0023] Step 3): Conditional on the audio input and the identity frame, the diffusion model generates a sequence of key frames at a low frame rate.
[0024] First, given a noisy input sequence ; the goal is to generate a sequence where a person speaks in sync with the given audio; Second, to provide identity and background information, the identity frame is connected to the noisy input through the VAE encoder, effectively using the skip connections of the U-Net architecture to retain input details; Finally, conditional on the audio input and the identity frame, a sequence of key frames of length T with an interval of S frames is generated through the diffusion model. These key frames effectively capture long-range temporal dependencies and serve as anchors for the subsequent interpolation stage.
[0025] Step 4): Conditional on two consecutive frames in the key frame sequence, the diffusion model interpolates between the key frames.
[0026] First, two consecutive frames and are obtained from the key frame sequence as conditional frames; Then, to match the input shapes of the start and end frames, a sequence is created, where represents the learned embedding of the missing frames. This sequence is concatenated with the noisy input by channel and interpolated through the diffusion model to obtain coherent intermediate frames.
[0027] Step 5): Optimize the diffusion model by combining the loss functions in the RGB space and the latent feature space.
[0028] First, decode the latent feature space back to the RGB space to obtain the decoded frame . Apply the L2 loss between the decoded frame and the ground truth frame and add it to the L2 loss between the latent feature space and ; represents the ground truth feature space; Then, optimize the parameters of the diffusion model during backpropagation.
[0029] Compared with the existing methods for generating audio-driven facial animations, the present invention solves the problems of identity drift and quality degradation during the generation of long videos, and realizes the generation of long-time, coherent and natural audio-driven facial animations.
[0030] Application Example To verify the effectiveness of the method of the present invention (i.e., Example 1), the training process of the method of the present invention uses HDTF (High Definition Talking Face) and CelebV-Text as datasets. Among them, the HDTF dataset focuses on flow-oriented high-resolution audio-visual face generation, involving speaker videos in the past two years, including 362 different people, and the video stream is cropped into 512*512 faces; CelebV-Text is a dataset containing 70,000 different facial video clips, and each clip has 20 text descriptions.
[0031] The present invention uses an objective evaluation method to evaluate the performance of the model trained by the method. For HDTF and CelebV-Text, the aesthetic quality index (AQ), Fréchet inception distance (FID), and learned perceptual image patch similarity (LPIPS) are used as evaluation indicators for image quality; and the Fréchet video distance (FVD) and smoothness index are used to evaluate general video quality.
[0032] Table 1 shows the quantitative results of the method of the present invention on the HDTF and CelebV-Text datasets. Bold indicates the best results.
[0033] Table 1: Quantitative Results of Audio-Driven Facial Animation Generation on HDTF and CelebV-Text Datasets
[0034] The results show that the method of the present invention achieves the lowest FID and FVD, indicating higher realism and temporal coherence. At the same time, the method achieves the highest AQ and LPIPS, confirming the visual attractiveness of the generated animations. Although SadTalker achieves the highest smoothness, the method of the present invention ranks closely behind and is better than SadTalker in other metrics, with better overall performance.
[0035] Example 2 This embodiment provides an audio-driven facial animation generation system for implementing the audio-driven facial animation generation method described in Example 1, as Figure 3 shown, which consists of an audio feature extraction unit, an audio feature feeding unit, a key frame sequence generation unit, an interpolation unit, and a diffusion model optimization unit.
[0036] The aforementioned audio feature extraction unit: It uses the combined embeddings from two pre-trained audio encoders for audio feature extraction; this unit is used to implement the function of step 1) in Embodiment 1, which will not be elaborated here.
[0037] The aforementioned audio feature feeding unit: It feeds the extracted audio features into the audio attention and time step embeddings of the diffusion model; this unit is used to implement the function of step 2) in Embodiment 1, which will not be elaborated here.
[0038] The aforementioned key frame sequence generation unit: Conditional on the audio input and identity frames, the diffusion model generates a key frame sequence at a low frame rate; this unit is used to implement the function of step 3) in Embodiment 1, which will not be elaborated here.
[0039] The aforementioned interpolation unit: Conditional on two consecutive frames in the key frame sequence, the diffusion model interpolates between the key frames; this unit is used to implement the function of step 4) in Embodiment 1, which will not be elaborated here.
[0040] The aforementioned diffusion model optimization unit: It optimizes the diffusion model by combining the loss functions in the RGB space and the latent feature space; this unit is used to implement the function of step 5) in Embodiment 1, which will not be elaborated here.
[0041] It should be noted that each unit in the above audio-driven facial animation generation system can be implemented in whole or in part by software, hardware, and their combinations. The above units can be embedded in the processor of the electronic device in hardware form or be independent of it, or be stored in the memory of the electronic device in software form, so that the processor can call and execute the operations corresponding to each of the above units. For the specific limitations of an audio-driven facial animation generation system, refer to the limitations of an audio-driven facial animation generation method (i.e., Embodiment 1) in the above text. The two have the same functions and effects, which will not be elaborated here.
[0042] Embodiment 3 This embodiment provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to Embodiment 1 of the present invention.
[0043] Embodiment 4 This embodiment provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to Embodiment 1 of the present invention.
[0044] Reference Figure 4, the structural block diagram of the electronic device 400 that can be used as the server or client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0045] As Figure 4 shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 402 or the computer program loaded from the storage unit 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0046] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device that can input information into the electronic device 400. The input unit 406 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 407 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include but is not limited to a magnetic disk, an optical disk. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0047] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above. For example, in some embodiments, the aforementioned audio-driven facial animation generation system can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. In some embodiments, the computing unit 401 can be configured to execute the aforementioned audio-driven facial animation generation system in any other suitable manner (e.g., by means of firmware).
[0048] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0049] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0050] As used in this invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0051] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0052] The systems and techniques described here can be implemented in a computing system including a back-end component (e.g., as a data server), or a computing system including a middleware component (e.g., an application server), or a computing system including a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described here), or in a computing system including any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0053] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0054] Those skilled in the art can clearly and easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art based on the disclosure of the present invention should fall within the protection scope of the present invention.
Claims
1. An audio-driven facial animation generation method, characterized in that, Including: Step 1), performing audio feature extraction using the combined embeddings from two pre-trained audio encoders; Step 2), feeding the extracted audio features into the audio attention and time step embeddings of the diffusion model; Step 3), conditioning on the audio input and identity frames, the diffusion model generates a sequence of key frames at a low frame rate; Step 4), conditioning on two consecutive frames in the key frame sequence, the diffusion model interpolates between the key frames; Step 5), optimizing the diffusion model by combining loss functions in the RGB space and the latent feature space.
2. The audio-driven facial animation generation method according to claim 1, wherein In the said Step 1), the two pre-trained audio encoders are WavLM and BEATs. WavLM captures the language content from speech, and BEATs is trained to extract features from acoustic signals including non-speech sounds.
3. The audio-driven facial animation generation method according to claim 2, wherein In the said Step 2), the mechanism of feeding the audio features into the diffusion model includes: 2.1) Audio attention: combined embedding As keys and values in the cross-attention layer within the U-Net architecture; Among them, represents the WavLM pre-trained audio encoder, represents the BEATs pre-trained audio encoder; 2.2) Time step embedding: Add to the time step embedding , so , represents the fused feature of the time step embedding and the audio feature, and MLP represents a fully connected neural network.
4. The audio-driven facial animation generation method according to claim 1, wherein, The specific process of generating the key frame sequence in the said Step 3) is: 3.1) Given a noisy input sequence , the goal is to generate a sequence of a person speaking in synchronization with the given audio; 3.2) To provide identity and background information, the identity frame is connected to the noise input through the VAE encoder, and the skip connections of the U-Net architecture are effectively utilized to retain the input details; 3.3) Conditional on the audio input and identity frames, generate a sequence of keyframes of length at a low frame rate through a diffusion model, with an interval of frames. These keyframes effectively capture long-range temporal dependencies and serve as anchors for the subsequent interpolation stage.
5. The audio-driven facial animation generation method according to claim 1, wherein The specific process of the said Step 4) is: 4.1) Obtain two consecutive frames from the key frame sequence and as conditional frames; 4.2) Create a sequence to match the input shapes of the start and end frames , where represents the learned embedding of the missing frames; this sequence is concatenated with the noise input by channel and interpolated through the diffusion model to obtain coherent intermediate frames.
6. The audio-driven facial animation generation method according to claim 5, characterized in that, The specific process of optimizing the diffusion model in the said Step 5) is: 5.1) Decode the latent feature space back to the RGB space to obtain the decoded frame . Apply the L2 loss between the decoded frame and the ground truth frame , and add it to the L2 loss between the latent feature space and , where represents the ground truth feature space. 5.2) Optimizing the parameters of the diffusion model during the backpropagation process.
7. An audio-driven facial animation generation system for implementing the audio-driven facial animation generation method according to any one of claims 1-6, characterized in that, Including: Audio feature extraction unit: performing audio feature extraction using the combined embeddings from two pre-trained audio encoders; Audio feature feeding unit: feeding the extracted audio features into the audio attention and time step embeddings of the diffusion model; Key frame sequence generation unit: conditioning on the audio input and identity frames, the diffusion model generates a sequence of key frames at a low frame rate; Interpolation unit: conditioning on two consecutive frames in the key frame sequence, the diffusion model interpolates between the key frames; Diffusion model optimization unit: optimizing the diffusion model by combining loss functions in the RGB space and the latent feature space.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice-driven facial animation method and device, equipment and medium
CN117115312A
Speaking face video generation method and device based on multi-modal information control
CN117456587A
Method, system and equipment for generating movie video clip by text
CN117478978A
Generating video using potential diffusion model
CN118053090A
Night parking space detection method, device and equipment based on parking space detection model
CN118247997A