Speaking face generation method and device, equipment and medium
By encoding, masking, and fusing audio features in the speech face generation method, the problems of inconsistent consecutive frames and insufficient generalization ability in speech face generation are solved, and high-quality, cross-language speech face video generation is achieved.
Patent Information
- Application Number
- CN202510881700.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
AI Technical Summary
Existing methods for generating speaking faces lack consistency between consecutive frames, leading to phenomena such as lip flickering, identity drift, and head shaking. Furthermore, when training data is insufficient, the model struggles to generalize to new speakers, new languages, or audio with strong accents, resulting in poor generalization ability.
By acquiring reference face images and encoding the original video frame sequence, a latent representation is generated. The original video frame sequence is then processed with lip masking to generate lip features. Combined with random sampling noise and audio features, lip alignment features are generated. Finally, lip synchronization frames are decoded to generate high-fidelity, strong temporal coherence, and cross-language generalization speaking face video.
While maintaining high-resolution visual quality, it achieves deep alignment of audio content, facial expressions, and temporal dynamics, improving the quality and generalization ability of speaker face generation, and achieving the goals of high fidelity, strong temporal coherence, and cross-language generalization.
Smart Images

Figure CN120807728A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a talking face generation method and device, equipment and a medium. BACKGROUND
[0002] With the rapid development of deep learning, computer vision and generative model technology, cross-modal generation tasks have gradually become a research hotspot in the field of artificial intelligence. Among them, talking face generation (TFG), which generates continuous and natural talking videos according to input audio or text-driven face images, is attracting widespread attention. This task involves the conversion and alignment of information from multiple modalities such as speech, text, images, and actions, and is a highly integrated generation problem.
[0003] In the medical field, talking face generation technology has innovative application potential in the medical field and can assist in diagnosis, treatment and patient communication through dynamic facial synthesis technology. For example, by generating dynamic images of a patient's face, doctors can more intuitively observe details such as facial muscle movements and expression changes to assist in diagnosing diseases such as Bell's palsy, Parkinson's disease, and facial nerve palsy. For example, by analyzing the subtle changes in a patient's facial expressions, features such as movement slowness or stiffness can be identified. Using talking face generation technology, dynamic images of a patient's face can be recorded regularly to compare changes at different time points, helping doctors assess disease progression or treatment effectiveness.
[0004] In the financial field, talking face generation technology is becoming an important tool for digital transformation, intelligent services, and cost reduction and efficiency improvement in traditional industries such as finance, securities, and insurance. Specific applications include: intelligent virtual customer service / digital teller, which replaces traditional text or voice customer service with virtual talking faces to simulate human tellers for services such as account opening consultation, financial advice, and process explanation, enhancing customer trust and loyalty compared to traditional all-day service, unified speech, and personalized expression; virtual analysts / consultant assistants use virtual faces to analyze the stock market, interpret hot topics, and analyze announcements for users, achieving low-cost, scalable expert explanations; combined with large models, content can be delivered to thousands of people.
[0005] Current talking face generation methods can be divided into the following categories: geometric-driven two-stage methods: first predict intermediate representations (such as key points, 3D mesh, pose parameters, etc.) from audio, then drive a rendering network to generate video. End-to-end image-to-video generation models: directly generate continuous face image sequences from audio or text input. Diffusion model and Transformer-based methods introduce diffusion structures and attention mechanisms to improve alignment and semantic expression capabilities while maintaining visual quality. Multi-modal methods driven by speech and text: adding language semantic information (such as emotion, accent) to improve personalization and controllability.
[0006] In the existing speaker face generation method, there is a lack of consistency between consecutive frames when generating a speaker face, resulting in phenomena such as lip flickering, identity drift, head shaking, etc. When the training data is insufficient, the model is difficult to generalize to new speakers, new languages or strong accent audio, and the generalization ability is poor. SUMMARY
[0007] In view of the deficiencies of the prior art described above, the present application provides a speaker face generation method, device, equipment and medium, aiming to solve the problems in the prior art that the speaker face generation method lacks consistency between consecutive frames when generating a speaker face, resulting in phenomena such as lip flickering, identity drift, head shaking, etc. When the training data is insufficient, the model is difficult to generalize to new speakers, new languages or strong accent audio, and the generalization ability is poor.
[0008] The technical scheme of the present application is as follows:
[0009] The first embodiment of the present application provides a speaker face generation method, the method comprising:
[0010] obtaining a reference face image and an original video frame sequence, respectively encoding the reference face image and the original video frame sequence to generate latent representations, the latent representations including reference face image features and original video frame sequence features;
[0011] performing lip mask processing on each image in the original video frame sequence to obtain a local mouth image, and encoding the local mouth image to generate a lip feature;
[0012] obtaining random sampling noise, and generating joint latent features based on the latent representations, the lip feature and the random sampling noise;
[0013] obtaining audio features, and generating audio-lip alignment features based on the reference face image features, the joint latent features and the audio features;
[0014] decoding the audio-lip alignment features to generate lip synchronization frames, and assembling the lip synchronization frames to generate a speaker face video.
[0015] Another embodiment of the present application provides a speaker face generation device, the device comprising:
[0016] a first encoding module for obtaining a reference face image and an original video frame sequence, respectively encoding the reference face image and the original video frame sequence to generate latent representations, the latent representations including reference face image features and original video frame sequence features;
[0017] a second encoding module, configured to perform a lip mask processing on each frame image in the original video frame sequence to obtain a local mouth image, encode the local mouth image, and generate a lip feature;
[0018] a joint feature generation module, configured to obtain random sampling noise, and generate a joint latent feature based on the latent representation, the lip feature, and the random sampling noise;
[0019] a lip-sound alignment module, configured to obtain an audio feature, and generate a lip-sound alignment feature based on the reference face image feature, the joint latent feature, and the audio feature;
[0020] a speaker face video generation module, configured to generate a lip-synchronized frame after decoding the lip-sound alignment feature, and generate a speaker face video by assembling the lip-synchronized frame.
[0021] Another embodiment of the present application provides a computer device, comprising at least one processor; and
[0022] a memory connected in communication with the at least one processor; wherein
[0023] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the speaker face generation method.
[0024] Another embodiment of the present application also provides a computer readable storage medium storing computer executable instructions, and the computer executable instructions are executed by one or more processors to enable the one or more processors to perform the steps of the speaker face generation method.
[0025] Beneficial effects: The speaking face generation method, apparatus, device and medium of the embodiments of the present invention can be used through the speaking face generation method, apparatus, device and medium, including: obtaining a reference face image and an original video frame sequence, encoding the reference face image and the original video frame sequence respectively, and generating a latent representation, wherein the latent representation includes reference face image features and original video frame sequence features; performing lip masking processing on each frame image in the original video frame sequence to obtain a local mouth image, encoding the local mouth image, and generating lip features; obtaining random sampling noise, and generating a joint latent feature based on the latent representation, lip features and random sampling noise; obtaining audio features, and generating a lip alignment feature based on the reference face image features, the joint latent feature and the audio features; generating a lip synchronization frame after decoding the lip synchronization feature, and assembling the lip synchronization frames to generate a speaking face video. While maintaining high-resolution visual quality, this invention achieves deep alignment of audio content, facial expressions and temporal dynamics, achieving the triple goals of "high fidelity, strong temporal coherence, and cross-language generalization". It lays a solid foundation for the application of multimodal diffusion models in speech-driven scenarios and improves the quality and generalization ability of speaking face generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0027] Figure 1 A schematic diagram of an application environment of an embodiment of a method for generating a speaking face according to the present invention;
[0028] Figure 2 This is a flow chart of a preferred embodiment of a method for generating a speaking face according to the present invention;
[0029] Figure 3 This is a functional module diagram of a preferred embodiment of a speaking face generation device of the present invention;
[0030] Figure 4 A schematic structural diagram of a preferred embodiment of a computer device of the present invention;
[0031] Figure 5 This is another structural diagram of a preferred embodiment of a computer device of the present invention. DETAILED DESCRIPTION
[0032] In order to make the objects, technical solutions and effects of the present application clearer, more explicit and more comprehensible, the present application will be further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0033] The embodiments of the present application will be described below in conjunction with the accompanying drawings. Those skilled in the art can know that, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0034] The terms "first", "second", and the like in the specification of the present application and claims and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. Here, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or equipment containing a series of units do not have to be limited to those units, but can include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0035] The method provided by the present application can be applied in the field of artificial intelligence (artificial intelligence, AI), which is a theory, method, technology and application system for simulating, extending and expanding human intelligence by using a digital computer or a machine controlled by a digital computer, perceiving an environment, acquiring knowledge and using the knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is the design principle and implementation method of various intelligent machines, so that the machine has the functions of perception, reasoning and decision-making. The research in the field of artificial intelligence includes robots, natural speech processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, AI basic theory, etc.
[0036] The speaker face generation method provided by the embodiments of the present application can be applied in, for example Figure 1In an application environment of the method, a client communicates with a server through a network. The client accesses a network or a service platform of the server. The server can obtain a reference face image and an original video frame sequence, encode the reference face image and the original video frame sequence respectively, generate latent representation, and the latent representation includes reference face image features and original video frame sequence features. Each image in the original video frame sequence is subjected to lip mask processing to obtain a local mouth image, and the local mouth image is encoded to generate a lip feature. Random sampling noise is obtained, and based on the latent representation, the lip feature and the random sampling noise, a joint latent feature is generated. An audio feature is obtained, and based on the reference face image feature, the joint latent feature and the audio feature, an audio-lip alignment feature is generated. After decoding the audio-lip alignment feature, a lip synchronization frame is generated, and the lip synchronization frame is assembled to generate a speaker face video. In the application, while maintaining high-resolution visual quality, depth alignment of audio content, facial expression and time sequence is achieved, achieving the three goals of "high fidelity, strong time sequence coherence and cross-language generalization", laying a solid foundation for the application of multi-modal diffusion models in voice-driven scenarios, and improving the quality and generalization ability of speaker face generation. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The application will be described in detail below through specific embodiments.
[0037] To solve the above problems, the embodiment of the application provides a speaker face generation method, please refer to Figure 2 , Figure 2 The flowchart of the preferred embodiment of the speaker face generation method of the application is shown in Figure 2 , which includes:
[0038] Step S100, obtaining a reference face image and an original video frame sequence, respectively encoding the reference face image and the original video frame sequence to generate latent representation, and the latent representation includes reference face image features and original video frame sequence features.
[0039] The embodiment of the application is mainly used to solve the key problems of unstable lip movement timing, weak semantic expression, poor cross-speaker generalization ability and poor lip-face fusion in the current speaker face generation process, to realize the depth alignment of voice-vision and the collaborative modeling of structure-texture-expression, and to generate a speaker face video with stronger realism, emotional expression and cross-domain adaptability. The method in the embodiment of the application is mainly applied to the speaker face generation of high-fidelity digital humans.
[0040] In the medical field, high-fidelity digital humans rely on computer graphics, artificial intelligence, speech synthesis and recognition, motion capture, and other technologies to create virtual avatars with highly realistic appearances, natural language interaction capabilities, and smooth body movements. Through high-precision three-dimensional reconstruction technology, the geometric information of the real human body is converted into a digital model, enabling arbitrary rotation, combination, and scaling of human structures, providing intuitive support for medical teaching and surgical planning. Combined with a medical knowledge base, digital humans can understand patient symptom descriptions, provide intelligent triage, preliminary diagnosis, and other services, significantly improving medical efficiency. By connecting with medical devices, digital humans can obtain real-time patient vital signs, surgical instrument positions, and other critical information, providing data support for surgical decision-making.
[0041] For example, in clinical diagnosis and treatment assistance, after patients describe their symptoms, digital humans can accurately recommend departments for treatment, reducing patient waiting time. Collecting patient basic information and generating electronic medical records provides doctors with diagnostic references. Analyzing medical image data, automatically identifying lesion areas, and assisting doctors in developing treatment plans.
[0042] In the financial field, high-fidelity digital humans rely on computer vision, speech synthesis and recognition, natural language processing, machine learning, and other core technologies to achieve realistic simulation of appearance, movement, emotion, and language. For example, by capturing three-dimensional face data with a depth camera and optimizing reconstruction, combined with speech synthesis technology to generate natural and smooth speech, digital humans can speak and understand user intentions like real people. This technology integration not only improves the naturalness of interaction but also gives digital humans emotional expression capabilities, allowing them to adjust their tone and expression according to the conversation context, enhancing user experience immersion.
[0043] For example, as an intelligent customer service, high-fidelity digital humans can provide 24-hour online answers to user inquiries about account inquiries, money transfers, financial product consultations, and other issues. For example, the digital employee "Xiaozhao" of China Merchants Bank can generate personalized financial planning based on user asset size and risk tolerance, significantly improving service efficiency. Digital financial consultants analyze global market data and quantitative models to provide multi-market investment strategy recommendations to users. For example, the "Zhangle Global Access" digital investment consultant of Huatai Securities can quickly analyze the impact of the Federal Reserve's interest rate cut on different industry sectors and push investment logic and target analysis reports to users.
[0044] The reference face image and the original video frame sequence are obtained, and the reference face image is encoded respectively to generate a reference face image feature. The original video frame sequence is encoded to generate an original video frame sequence feature. During encoding, the input data (such as images, audio) is mapped to a continuous latent space to generate a latent vector. A convolutional neural network (CNN) or a fully connected network is usually used, and the specific structure depends on the type of input data. For example, for image data, the encoder may consist of multiple convolutional layers and down-sampling layers, outputting a low-dimensional continuous latent vector. The reference face image is encoded by using a face key point detection model to detect the key points of the reference face image, obtaining the face key point coordinates, determining the face mask area according to the face key point coordinates, generating a face mask image, and encoding the face mask image based on the encoder to generate a face image feature, which is used for generating and positioning the speaker face. The encoding of the original video frame sequence is to extract features from each video frame in the video frame sequence to generate an original video frame sequence feature.
[0045] In step S100, the reference face image and the original video frame sequence are obtained, and the reference face image and the original video frame sequence are encoded respectively to generate a latent representation, including:
[0046] In step S101, a vector quantization variational autoencoder sharing weights is constructed.
[0047] In step S102, the reference face image and the original video frame sequence are obtained.
[0048] In step S103, the reference face image is encoded based on the vector quantization variational autoencoder to generate a reference face image feature.
[0049] In step S104, the original video frame sequence is encoded based on the vector quantization variational autoencoder to generate an original video frame sequence feature.
[0050] In step S105, the reference face image feature and the original video frame sequence feature are obtained to generate a latent representation.
[0051] The embodiment of the present application constructs a vector quantization variational autoencoder sharing weights, undertakes a latent space mapping task, and effectively compresses high-resolution images / videos through a shared encoder-decoder weight at the visual end. The VQ-VAE (Vector Quantized Variational Autoencoder) is a generative model combining VAE (Variational Autoencoder) and vector quantization technology, aiming to generate data through discrete latent variables to improve the quality and diversity of generated samples. It mainly maps the input data to a continuous latent space through an encoder, maps it to a vector in a discrete codebook through vector quantization, and finally reconstructs the data through a decoder.
[0052] Given a reference face image I ref and an original video frame sequence First, a VQ-VAE encoder sharing weights is constructed. Based on the VQ-VAE encoder, the reference face image and the original video frame sequence are encoded respectively to generate corresponding reference face image features and original video frame sequence features. The reference face image features and the original video frame sequence features are obtained corresponding latent representations
[0053] Step S200, performing lip masking processing on each image in the original video frame sequence to obtain a local mouth image, and encoding the local mouth image to generate a lip feature.
[0054] Performing lip masking processing on each image in the original video frame sequence, the lip masking processing is a key technology in computer vision and image processing, mainly used to separate the lip area from the face image or video. The main methods of lip masking processing include: based on traditional image processing method, for example, segmentation can be performed by color threshold segmentation method, and the color feature (such as red tone) of the lip area is used for segmentation. Convert the image from the RGB color space to the HSV or YCbCr color space. Set the threshold according to the hue range of the lips (such as Hue in HSV is in the range of 0-10 or 160-180). Apply threshold segmentation to generate a binary mask. Or use edge detection and morphological operation method, use the edge feature (such as contour) of the lips for segmentation. Use edge detection algorithm (such as Canny) to extract the lip edge. Apply morphological operations (such as dilation, erosion) to fill the area within the edge.
[0055] The lip mask processing method can also use a deep learning-based method. For example, a semantic segmentation model is used to perform pixel-level classification on the image using a convolutional neural network (CNN) or a fully convolutional network (FCN) to generate a lip mask. Common models: U-Net: a classic semantic segmentation model suitable for small data sets. DeepLab: uses Atrous Convolution and Pyramid Pooling (ASPP), suitable for high-resolution images. BiSeNet: combines spatial and contextual paths to balance speed and accuracy. High accuracy, strong robustness, and adaptability to complex scenes.
[0056] A generative adversarial network (GAN) can also be used to generate high-quality lip masks, especially for low-resolution or blurred images.
[0057] After the image is processed with the lip mask, a local mouth image M t is obtained, which is also encoded as a latent feature m t .
[0058] In step S200, each frame image in the original video frame sequence is processed with a lip mask to obtain a local mouth image, and the local mouth image is encoded to generate a lip feature, including:
[0059] Step S201: obtaining each frame image in the original video frame sequence;
[0060] Step S202: detecting the lip position in each frame image, and processing each frame in the original video frame sequence with a lip mask according to the lip position to obtain a local mouth image;
[0061] Step S203: encoding the local mouth image based on the vector quantization variational autoencoder to generate a lip feature.
[0062] In the embodiment of the application, after being obtained, the face key point detection + convex hull generation method can also be used for lip covering. The face key point detection model (such as Dlib, MediaPipe) is used to extract the lip key points. The convex hull is generated according to the key points, which is used as a lip mask, and a local mouth image is obtained. The local mouth image is encoded based on the vector quantization variational autoencoder to generate a lip feature.
[0063] Step S300: obtaining random sampling noise, and generating a joint latent feature based on the latent representation, the lip feature, and the random sampling noise.
[0064] The network structure of the embodiment of the application further comprises a feature extraction module. The feature extraction module comprises an identity preservation module and a cross-motion facial learning module. Static identity embedding information and dynamic mouth-muscle coupling features are captured respectively. This part adopts a parallel double-branch cross-attention transformer: the query end takes the reference latent feature z ref as the key / value, and the target end takes each frame of latent feature z t as the query, to realize quasi-real-time identity migration and pose alignment. In order to enhance local expression details, a residual adaptive instance normalization layer is built in the module to fine-tune skin color and illumination in the form of channel standard deviation matching. In the mouth-muscle coupling modeling of the cross-motion facial learning module, an attention mechanism is adopted. The local branch pays attention to the lip m t fine-grained deformation, and the global branch pays attention to the facial muscle texture z t . Through global-to-local cross-attention gating fusion in the middle, coupling features between facial action units are explicitly captured.
[0065] In step S300, random sampling noise is obtained, and a joint latent feature is generated based on the latent representation, the lip feature and the random sampling noise, comprising:
[0066] In step S301, the reference face image feature and the original video frame sequence feature are aligned to generate an identity feature.
[0067] In step S302, the lip feature and the identity feature are coupled based on a cross-attention algorithm to generate a coupled embedding feature.
[0068] In step S303, random sampling noise is obtained, and a joint latent feature is generated based on feature fusion of the identity feature, the coupled embedding feature and the random sampling noise.
[0069] The identity preservation module uses a cross-attention function A(·). Cross-attention is a key mechanism in the Transformer architecture, which is used to establish dynamic associations between two different sequences or feature representations. Unlike self-attention, cross-attention achieves cross-modal or cross-sequence information fusion by comparing the query of one sequence with the key and value of another sequence. zref With z t t alignment, output identity features as follows:
[0070]
[0071] The cross-motion face learning module takes (m t ,z t ) as input, learns mouth-muscle coupling embedding through a Self-Cross Transformer unit
[0072] Randomly sampled noise ξ t ~ N(0, I) is generated into joint latent features through a feature fusion module
[0073]
[0074] Step S400, acquiring audio features, generating audio-lip alignment features based on the reference face image features, joint latent features and audio features.
[0075] After acquiring the audio input by the user, feature extraction is performed on the audio to generate audio features. Align the audio features with the reference face image features and the joint latent features to generate audio-lip alignment features. Audio feature extraction is the basic step for audio signal processing, speech recognition, music information retrieval and other tasks. Its goal is to extract representative features from the original audio signal for subsequent model training or analysis. The main methods of audio feature extraction include time domain, frequency domain, time-frequency domain and deep learning features: time domain features are directly extracted from the waveform signal of the audio, reflecting the change of the signal with time. Frequency domain features convert the audio signal from time domain to frequency domain through Fourier transform, reflecting the frequency components of the signal. Time-frequency domain features combine time domain and frequency domain information to reflect the joint distribution of the signal in time and frequency. Deep learning models can learn high-level features directly from raw audio signals or low-level features.
[0076] Wherein, step S400, that is, acquiring audio features, generating audio-lip alignment features based on the reference face image features, joint latent features and audio features, includes:
[0077] Step S401, performing cross-scale identity alignment on the reference face image and the joint latent features to generate identity alignment features;
[0078] Step S402, acquiring input audio, performing feature extraction on the input audio to generate corresponding audio features;
[0079] Step S403, inputting the audio features into the lip generation model to generate audio-lip alignment features.
[0080] With initial noise, the conditional diffusion process D with time step τ ∈ [1, K] is used to gradually denoise; at each U-Net resolution layer, the identity feature maintainer injects the multi-scale features of z ref to achieve cross-scale identity alignment.
[0081] The phoneme-level embedding φ(a) extracted from the input audio a on the speech side is fed into the audio-guided lip generator, which adjusts the feature distribution of the mouth-related channels in the U-Net through multi-head cross-attention to ensure audio-lip alignment.
[0082] The Feature Fusion Module uses convolutional layers and channel attention blocks to complete information-noise weighting in the channel dimension. This module includes a 3x3 convolution and a squeeze-and-excitation block. The convolution is used to aggregate spatial context, and the SE-Block automatically adjusts the proportion of identity, motion, and noise through learned scaling coefficients γ to achieve an interpretable "information-noise weight" scheduling.
[0083] The core Stable Diffusion Backbone+Identity Maintainer performs temporal denoising and maintains multi-scale identity consistency; this subnetwork introduces a learnable residual mapping R 2 at U-Net scales 64 2 ,128 2 ,256 2 ,512 l to counteract the "identity degradation" phenomenon in the later stages of adversarial training.
[0084] On the speech side, the Audio-Guided Lip Generator injects audio conditions through cross-attention; to address the diversity of accents, the 64-dimensional representation of the audio is first extracted and then mapped to the phoneme-level embedding. The use of cross-attention within the generator enables effective fusion of audio-lip, audio-muscle, and audio-identity from multiple angles, significantly improving lip synchronization in complex environmental scenarios.
[0085] The Cross-modal Temporal Modeling Module explicitly models the spatio-temporal dependencies between N-frame latent sequences, outputting smoother lip and head movements. This hierarchical architecture balances the needs of identity-action decoupling and multi-modal alignment, laying the foundation for high-fidelity and strongly consistent video generation. The module input is the latent sequence The bilinear similarity cos with offset is adopted β The inter-frame relationship is measured, and an inter-frame weight matrix is generated through time bias multi-head attention (TB-MHSA), so as to realize explicit "neighbor time domain smoothing".
[0086] The cross-modal time sequence modeling module performs
[0087]
[0088] The inter-frame residual is explicitly minimized.
[0089] In step S403, the audio features are input into the lip generation model to generate audio-lip alignment features, including:
[0090] In step S431, a lip generation model based on a deep learning network is constructed.
[0091] In step S432, a loss function of the lip generation model is set.
[0092] In step S433, the audio features are input into the lip generation model, and the lip generation model is trained based on the loss function. After the training is completed, audio-lip alignment features are generated.
[0093] A lip generation model based on a deep learning network is constructed, and a loss function of the lip generation model is set. The loss function includes:
[0094] Noise prediction loss. The diffusion model core module is trained to learn to accurately predict the noise term ε in the process from the clean latent to the noisy latent.
[0095] Perceptual loss. It ensures the semantic consistency of the lip contour, mouth shape structure and character details; maintains the realism of the image at the structure and texture level, especially the clarity and stability of the mouth area.
[0096] Flow-guided frame-to-frame smoothness loss. It ensures that the continuously generated frames change smoothly in vision, preventing phenomena such as mouth shape jumping and mouth shaking.
[0097] Timestep-aware Spatial Loss. Enhance the temporal structure consistency in multi-frame input training, reduce the impact of strong noise stage on training stability. Add LPIPS perceptual loss to maintain structural authenticity; Collaborate with Temporal Attention mechanism to model temporal relationships.
[0098] Total Loss Function: Combine the four losses with weights for training, balance denoising, structure clarity, inter-frame transition, and temporal coherence.
[0099] Step S500, after decoding the audio-lip alignment feature, generate the lip-sync frame, assemble the lip-sync frame to generate the speaker face video.
[0100] Final audio-lip alignment feature Map back to the pixel domain through the VQ-VAE decoder to get high-definition lip-sync frames Assemble into output video according to time index After getting the lip-sync frame, ensure that the resolution, color space (such as RGB), and encoding format of all frames are consistent.
[0101] The representation of time index is as follows: Frame number: integer sequence (such as 0, 1, 2,...). Timestamp: in seconds or milliseconds (such as 0.0, 0.033, 0.066,... for 30FPS video). Frame rate (FPS): the frame rate of the video determines the interval of the time index (such as 30FPS means the interval between each frame is 1 / 30 seconds).
[0102] Map frames to time index. Sequential matching: ensure that the order of frames is consistent with the time index. Missing frame processing: if some time index is missing frames, interpolation or skipping may be needed.
[0103] Video encoding and synthesis process as follows: encoding format: select video encoding format (such as H.264, H.265, MPEG-4). Choose video container (such as MP4, AVI, MOV). Use libraries such as FFmpeg, OpenCV, MoviePy for encoding and synthesis.
[0104] Among them, step S500, i.e. step after decoding the audio-lip alignment feature, generate the lip-sync frame, assemble the lip-sync frame to generate the speaker face video, including:
[0105] Step S501, based on the vector quantization variational autoencoder, the audio-lip alignment feature is decoded to generate the lip-sync frame;
[0106] Step S502, obtain the time index corresponding to the lip-sync frame;
[0107] Step S503, based on the time index, the lip synchronization frame is assembled, and a speaker face video is generated.
[0108] The lip synchronization frame is decoded based on the vector quantization variational self-decoder to generate a lip synchronization frame. Wherein the time index is a temporal indexing, which is an important concept when processing sequential data such as time series, video, speech, etc., especially in the self-cross transformer for modeling temporal dependencies and cross-modal alignment.
[0109] In time series data (such as stock prices, sensor data), time index is used to determine the position of each time step, helping the model capture long-term dependencies in time. For example, when predicting future stock prices, the model needs to know the relationship between the current time step and the historical time step.
[0110] Cross-modal alignment. In multi-modal tasks (such as video caption generation, speech recognition), time index is used to align the time dimension of different modalities. For example, in video caption generation, video frames (visual modality) and text descriptions (language modality) need to be aligned through time index to ensure that the generated text is consistent with the video content in time.
[0111] Dynamic attention allocation. Time index can be combined with attention mechanism to make the model dynamically focus on information at different time steps. For example, in speech recognition, the model can dynamically focus on key frames in the speech signal through time index, ignoring noise or redundant information.
[0112] The implementation of time index includes explicit time encoding, relative time encoding and timestamp embedding. Timestamp embedding is used when processing non-uniform time steps (such as event sequences), which directly embeds the timestamp (such as the numerical value of the timestamp or the time interval) into the model. For example, in medical event sequence prediction, the timestamp of event occurrence can be embedded into the model to help the model capture the time interval relationship between events. After combining the lip synchronization frame based on the time index, a speaker face video is generated.
[0113] The embodiment of the application realizes a stable diffusion architecture based framework, designs a structure-phenomenon collaborative feature extractor (identity preservation module and cross-motion face learning module) to realize the natural fusion of lip and facial texture, illumination and motion, designs an identity feature preserver to ensure the authenticity of the generated video, introduces a speech guided lip generator network to improve the emotion restoration ability, and introduces a cross-modal time sequence modeling module to enhance the dynamic consistency between frames. The framework not only supports fine control and continuous generation of target human face driven by any speech input, but also has good personal transfer ability and multi-language expansion potential, and is suitable for virtual digital people, intelligent customer service, digital marketing and other practical scenarios.
[0114] It should be noted that the above steps do not necessarily have a certain order, and those skilled in the art can understand from the description of the embodiments of the present application that the above steps can have different execution orders in different embodiments, that is, they can be executed in parallel, or they can be executed in exchange, etc.
[0115] Another embodiment of the present application provides a speaker face generation device, which corresponds to the speaker face generation method described above. As shown in the figure, the device 1 comprises: Figure 3
[0116] The first encoding module 100 is configured to obtain a reference face image and an original video frame sequence, encode the reference face image and the original video frame sequence respectively, and generate latent representations, wherein the latent representations comprise reference face image features and original video frame sequence features.
[0117] The second encoding module 200 is configured to perform lip mask processing on each frame image in the original video frame sequence to obtain a local mouth image, encode the local mouth image, and generate a lip feature.
[0118] The joint feature generation module 300 is configured to obtain random sampling noise, generate joint latent features based on the latent representations, the lip feature, and the random sampling noise.
[0119] The audio-lip alignment module 400 is configured to obtain audio features, generate audio-lip alignment features based on the reference face image features, the joint latent features, and the audio features.
[0120] The speaker face video generation module 500 is configured to decode the audio-lip alignment features to generate lip synchronization frames, assemble the lip synchronization frames, and generate a speaker face video.
[0121] The specific implementation is described in the method embodiment, which will not be repeated here.
[0122] In one embodiment, the first encoding module 100 is specifically configured to:
[0123] construct a vector quantization variational autoencoder sharing weights;
[0124] obtain a reference face image and an original video frame sequence;
[0125] encode the reference face image based on the vector quantization variational autoencoder to generate reference face image features;
[0126] encode the original video frame sequence based on the vector quantization variational autoencoder to generate original video frame sequence features;
[0127] Based on the reference face image feature and the original video frame sequence feature, a latent representation is obtained.
[0128] The specific implementation is described in the method embodiment, which will not be repeated here.
[0129] In one embodiment, the second encoding module 200 is specifically configured to:
[0130] Obtain each frame image in the original video frame sequence;
[0131] Detect the lip position in each frame image, perform lip mask processing on each frame in the original video frame sequence according to the lip position, and obtain a local mouth image;
[0132] Encode the local mouth image based on the vector quantization variational autoencoder to generate a lip feature.
[0133] The specific implementation is described in the method embodiment, which will not be repeated here.
[0134] In one embodiment, the joint feature generation module 300 is specifically configured to:
[0135] Align the reference face image feature and the original video frame sequence feature to generate an identity feature;
[0136] Couple the lip feature and the identity feature based on a cross-attention algorithm to generate a coupled embedding feature;
[0137] Obtain a random sampling noise, and perform feature fusion based on the identity feature, the coupled embedding feature, and the random sampling noise to generate a joint latent feature.
[0138] The specific implementation is described in the method embodiment, which will not be repeated here.
[0139] In one embodiment, the audio-lip alignment module 400 is specifically configured to:
[0140] Perform cross-scale identity alignment on the reference face image and the joint latent feature to generate an identity alignment feature;
[0141] Obtain an input audio, perform feature extraction on the input audio, and generate a corresponding audio feature;
[0142] Input the audio feature into a lip generation model to generate an audio-lip alignment feature.
[0143] The specific implementation is described in the method embodiment, which will not be repeated here.
[0144] In one embodiment, the audio-lip alignment module 400 is further configured to:
[0145] Build a lip generation model based on deep learning network;
[0146] Setting a loss function for the lip generation model;
[0147] The audio features are input into a lip generation model, and the lip generation model is trained based on a loss function. After the training is completed, a voice-lip alignment feature is generated.
[0148] The specific implementation method is shown in the method embodiment, which will not be repeated here.
[0149] In one embodiment, the speaking face video generation module 500 is specifically configured to:
[0150] Decoding the lip alignment features based on a vector quantized variational autodecoder to generate a lip synchronization frame;
[0151] Get the time index corresponding to the lip sync frame;
[0152] The lip synchronization frames are assembled based on the time index to generate a speaking face video.
[0153] The specific implementation method is shown in the method embodiment, which will not be repeated here.
[0154] This invention provides a device for generating speaking faces, achieving end-to-end collaborative optimization at the algorithm design and loss function levels. While maintaining high-resolution visual quality, it achieves deep alignment of audio content, facial expressions, and temporal dynamics, achieving the triple goals of "high fidelity, strong temporal coherence, and cross-language generalization," and achieving a systematic breakthrough compared to the existing optimal baseline. This provides a new paradigm for building commercially viable virtual digital humans and lays a solid foundation for the application of multimodal diffusion models in speech-driven scenarios.
[0155] Another embodiment of the present invention provides a computer device, which may be a server, and its internal structure diagram may be as shown in FIG. Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a speaking face generation method.
[0156] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a speaking face generation method.
[0157] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0158] Obtain a reference facial image and an original video frame sequence, encode the reference facial image and the original video frame sequence respectively, and generate a latent representation, wherein the latent representation includes features of the reference facial image and features of the original video frame sequence;
[0159] Performing lip masking processing on each frame image in the original video frame sequence to obtain a local mouth image, encoding the local mouth image to generate lip features;
[0160] Obtaining random sampling noise, and generating a joint latent feature based on the latent representation, the lip feature, and the random sampling noise;
[0161] Acquire audio features, and generate lip alignment features based on the reference facial image features, the joint latent features, and the audio features;
[0162] After decoding the lip alignment features, lip synchronization frames are generated, and the lip synchronization frames are assembled to generate a speaking face video.
[0163] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0164] Obtain a reference facial image and an original video frame sequence, encode the reference facial image and the original video frame sequence respectively, and generate a latent representation, wherein the latent representation includes features of the reference facial image and features of the original video frame sequence;
[0165] performing lip mask processing on each frame image in the original video frame sequence to obtain a local mouth image, encoding the local mouth image to generate a lip feature;
[0166] obtaining random sampling noise, generating a joint latent feature based on the latent representation, the lip feature and the random sampling noise;
[0167] obtaining an audio feature, generating a lip-audio alignment feature based on the reference face image feature, the joint latent feature and the audio feature;
[0168] generating a lip synchronization frame after decoding the lip-audio alignment feature, and assembling the lip synchronization frame to generate a speaker face video.
[0169] It should be noted that the functions or steps described above with respect to the computer readable storage medium or the computer device can correspond to the related descriptions of the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0170] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM).
[0171] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the objectives of the present embodiments as needed.
[0172] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the relevant technology can be embodied in the form of a software product. This computer software product can be present in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or certain parts of the embodiment.
[0173] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.
[0174] Conditional language such as "can," "may," or "might," among others, unless specifically stated otherwise or otherwise understood within the context as used, is generally intended to convey that particular embodiments can include, while other embodiments do not, particular features, elements, and / or operations. Thus, such conditional language is also generally intended to imply that features, elements, and / or operations are anyway required for one or more embodiments or that one or more embodiments must include logic for determining, with or without input or prompting, whether such features, elements, and / or operations are included or to be performed in any particular embodiment.
[0175] What has been described herein in this specification and the accompanying drawings includes examples that can provide methods and apparatus for generating speaking faces. Of course, it is not possible to describe every conceivable combination of elements and / or methods for the purpose of describing the various features of the present disclosure, but it will be appreciated that many additional combinations and permutations of the disclosed features are possible. It will therefore be apparent that various modifications can be made to the present disclosure without departing from the scope or spirit of the present disclosure. In addition, or in the alternative, other embodiments of the present disclosure may be apparent from consideration of the present specification and the accompanying drawings and from practice of the present disclosure as presented herein. It is intended that the examples set forth in this specification and the accompanying drawings be considered in all respects to be illustrative and not restrictive. Although specific terms are employed herein, they are used in a general and descriptive sense and not for purposes of limitation.
Claims
1. A method for generating a speaking face, characterized in that , the method comprises: Obtain a reference facial image and an original video frame sequence, encode the reference facial image and the original video frame sequence respectively, and generate a latent representation, wherein the latent representation includes features of the reference facial image and features of the original video frame sequence; Performing lip masking processing on each frame image in the original video frame sequence to obtain a local mouth image, encoding the local mouth image to generate lip features; Obtaining random sampling noise, and generating a joint latent feature based on the latent representation, the lip feature, and the random sampling noise; Acquire audio features, and generate lip alignment features based on the reference facial image features, the joint latent features, and the audio features; After decoding the lip alignment features, lip synchronization frames are generated, and the lip synchronization frames are assembled to generate a speaking face video.
2. The method for generating a speaking face according to claim 1, wherein: The obtaining of a reference face image and an original video frame sequence, encoding the reference face image and the original video frame sequence respectively, and generating a potential representation includes: Construct a vector quantized variational autoencoder with shared weights; Obtain reference face images and original video frame sequences; Encoding the reference facial image based on the vector quantization variational autoencoder to generate reference facial image features; Encoding the original video frame sequence based on the vector quantization variational autoencoder to generate original video frame sequence features; A potential representation is obtained based on the reference facial image features and the original video frame sequence features.
3. The method for generating a speaking face according to claim 1, wherein: The step of performing lip masking on each frame image in the original video frame sequence to obtain a partial mouth image, encoding the partial mouth image, and generating lip features includes: Get each frame image in the original video frame sequence; Detecting the lip position in each frame of the image, and performing lip masking processing on each frame in the original video frame sequence according to the lip position to obtain a partial mouth image; The local mouth image is encoded based on the vector quantization variational autoencoder to generate lip features.
4. The method for generating a speaking face according to claim 1, wherein: The obtaining of random sampling noise and generating a joint latent feature based on the latent representation, the lip feature and the random sampling noise include: Aligning the reference facial image features with the original video frame sequence features to generate identity features; The lip feature and the identity feature are coupled based on a cross attention algorithm to generate a coupled embedded feature; Obtain random sampling noise, perform feature fusion based on the identity feature, the coupled embedded feature, and the random sampling noise, and generate a joint latent feature.
5. The method for generating a speaking face according to claim 1, wherein: The acquiring of audio features and generating lip alignment features based on the reference facial image features, the combined latent features, and the audio features includes: Performing cross-scale identity alignment on the reference face image and the joint latent feature to generate an identity-aligned feature; Acquire input audio, perform feature extraction on the input audio, and generate corresponding audio features; The audio features are input into a lip generation model to generate a voice-lip alignment feature.
6. The method for generating a speaking face according to claim 5, wherein: Inputting the audio features into the lip generation model to generate the voice-lip alignment features includes: Build a lip generation model based on deep learning network; Setting a loss function for the lip generation model; The audio features are input into a lip generation model, and the lip generation model is trained based on a loss function. After the training is completed, a voice-lip alignment feature is generated.
7. The method for generating a speaking face according to claim 1, wherein: After decoding the lip alignment features, generating lip synchronization frames, assembling the lip synchronization frames to generate a speaking face video, including: Decoding the lip alignment features based on a vector quantized variational autodecoder to generate a lip synchronization frame; Get the time index corresponding to the lip sync frame; The lip synchronization frames are assembled based on the time index to generate a speaking face video.
8. A device for generating a speaking face, characterized in that: The device comprises: A first encoding module is configured to obtain a reference facial image and an original video frame sequence, encode the reference facial image and the original video frame sequence respectively, and generate a latent representation, wherein the latent representation includes features of the reference facial image and features of the original video frame sequence; a second encoding module, configured to perform lip masking on each frame image in the original video frame sequence to obtain a partial mouth image, and encode the partial mouth image to generate lip features; a joint feature generation module, configured to obtain random sampling noise and generate a joint latent feature based on the latent representation, the lip feature and the random sampling noise; a lip alignment module, configured to obtain audio features and generate lip alignment features based on the reference facial image features, the joint latent features, and the audio features; The speaking face video generation module is used to generate lip synchronization frames after decoding the voice-lip alignment features, and assemble the lip synchronization frames to generate a speaking face video.
9. A computer device, characterized in that: The computer device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the speaking face generation method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the one or more processors can execute the steps of the speaking face generation method according to any one of claims 1 to 7.