Digital human voice lip synchronization method based on edge perception and multiple modes
Through the method of enhancing network based on adversarial textures, the lip movement trajectory and texture details are optimized, and the lip blur and instability problems in the existing lip synchronization technology are solved, and the high-quality lip synchronization effect is achieved, which improves the realism and user experience of digital people.
Patent Information
- Application Number
- CN202510457780.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
The existing lip synchronization technology has problems such as blurred lips texture, unstable movement, unclear boundaries, and inaccurate matching of lips and voice, which affects the realism and user experience of digital people.
Using an adversarial texture enhancement network method, the lip movement trajectory is optimized through the graph neural network, combined with the generation adversarial network to improve texture details, introduce edge-aware loss to optimize boundary clarity, and improve the authenticity and nature of lip synchronization through multimodal feature alignment and identity consistency loss.
Generate high-resolution, natural and smooth sound-lip synchronous videos, which improves the interactive experience of digital people, making their applications more realistic and smooth in the fields of virtual anchors, film and television post-production, and game animation.
Smart Images

Figure CN120374810A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, especially the cross - application of computer vision and deep learning, and specifically relates to a digital human audio - lip synchronization method based on an adversarial texture enhancement network. Background Art
[0002] With the development of artificial intelligence and computer vision technologies, digital human technology has been widely applied in fields such as virtual anchors, intelligent customer service, post - production of films and games animation. Among them, audio - lip synchronization, as a core link of digital human technology, determines the authenticity and naturalness of the synthesized video. An ideal audio - lip synchronization system needs to ensure that the lip movement of the virtual character can accurately match the audio content while maintaining visual smoothness and naturalness. However, the current mainstream audio - lip synchronization technologies still face many challenges, which affect the realism of digital humans and the user experience.
[0003] Currently, audio - lip synchronization technologies are mainly divided into rule - based, statistical learning - based, and deep - learning - based methods. Early methods drove lip animation through phoneme - lip shape mapping rules, but due to the lack of adaptability to actual speech changes, the generated lip movements were rigid and unnatural. Statistical learning - based methods such as Hidden Markov Models (HMMs) and Gaussian Mixture Models (GMMs) can, to a certain extent, establish a probability mapping between speech and lip movements, but are prone to the problem of insufficient generalization ability in the face of complex speech environments. In recent years, with the development of deep learning, neural - network - based methods have become mainstream. For example, Wav2Lip, as a relatively advanced current audio - lip synchronization model, can use an autoregressive convolutional neural network (CNN) and LSTM for lip prediction, and fuse audio and video information through an attention mechanism to achieve relatively high - quality lip synchronization. However, Wav2Lip still has problems such as blurred lip texture, unstable lip movement, unclear boundaries, and inaccurate lip - speech matching. Summary of the Invention
[0004] Object of the Invention: The present invention aims to provide a digital human audio - lip synchronization method based on an adversarial texture enhancement network, which improves the realism and naturalness of digital human lip synchronization by optimizing the lip movement trajectory, enhancing the texture details of the lip region, and strengthening the multi - modal feature alignment ability.
[0005] Technical Solution: The present invention discloses a digital human audio - lip synchronization method based on an adversarial texture enhancement network, including the following steps:
[0006] Step 1: Obtain audio and video data, collect video segments of the target person, and at the same time obtain the corresponding audio data to ensure the time synchronization of the audio and video data;
[0007] Step 2: Use the Wav2Lip model to extract audio and video features and perform preliminary audio-visual alignment to generate a preliminary lip image, extract the coordinate sequence of lip key points as the dynamic lip features;
[0008] Step 3: Based on the dynamic lip features extracted in Step 2, construct a lip motion prediction model of the graph neural network GCN to propagate and aggregate the motion information in the lip region;
[0009] Step 4: Use a generative adversarial network to enhance the texture of the lip image generated in Step 2. Through perceptual loss and adversarial loss, improve the texture details, edge sharpness, and light and shadow layering in the lip region. Integrate the multi-scale feature enhancement strategy to extract the embedded feature vectors in consecutive frame lip images;
[0010] Step 5: Construct an audio-visual feature consistency evaluation module. Combine the identity consistency loss and the edge perception loss to perform multi-dimensional discrimination and optimization on the audio-visual features processed in Step 3 and Step 4;
[0011] Step 6: Synthesize the optimized lip region into the original video and output the final audio-lip synchronous video.
[0012] Furthermore, the specific method of Step 1 includes:
[0013] Step 1.1: Collect video clips of the target person and obtain the corresponding audio data;
[0014] Step 1.2: Perform normalization processing on the audio data;
[0015] Step 1.3: Perform face detection, locate the lip region, and crop the lip region image;
[0016] Step 1.4: Label the lip region.
[0017] Furthermore, the specific method of Step 2 includes:
[0018] Step 2.1: Use the pre-trained Wav2Lip model to perform audio-visual alignment processing on the audio and video data in Step 1. Generate a synchronous lip region image in an audio-driven manner. Use the pre-trained lip synchronization discriminator SyncNet to calculate the lip synchronization score, and use the UNet structure to generate the lip shape to obtain the lip image;
[0019] Step 2.2: Extract the coordinate sequence of lip key points in each frame from the lip image obtained in Step 2.1 for modeling the lip motion trajectory. By calculating the Euclidean distance between key points in different frames, characterize the dynamic features of the lip shape. The Euclidean distance calculation function formula is as follows:
[0020]
[0021] Among them, d is the Euclidean distance, and (x1, y1) and (x2, y2) respectively represent the spatial coordinates of the lip key points in different frame images.
[0022] Furthermore, the specific method of step 3 includes:
[0023] Step 3.1: Using the lip dynamic features extracted in step 2 as input features, taking the lip key points of each frame as the nodes of the graph, constructing a temporal graph structure, and establishing the connection relationship between nodes using the adjacency matrix;
[0024] Step 3.2: Defining the graph structure, each lip key point v i as a node, the connection relationship between nodes is determined by the adjacency matrix A, and the feature h of each node is calculated i , where the initial feature comes from the visual feature of the lip region, the feature matrix X is used for node feature representation, and each node contains information of multiple channels;
[0025] Step 3.3: Using the graph neural network GCN and the lip key point coordinate sequence in each frame to perform multi-layer graph convolution operations to propagate and aggregate the motion information of the lip region, and the function formula is as follows:
[0026]
[0027] where, H 1 is the node feature of the l-th layer of the input layer, H l represents the node feature of the l-th layer, A is the adjacency matrix, D is the degree matrix, W l is the trainable weight of this layer, σ is the activation function, and σ is used to introduce non-linearity so that the lip motion information can be propagated and learned more complexly, and at the same time avoid gradient disappearance or explosion.
[0028] Furthermore, the specific method of step 4 includes:
[0029] Step 4.1: Using the lip image generated in step 2 as input, constructing a generative adversarial network, generating a high-definition lip image by the generator, and the discriminator judges the difference in visual texture between the generated image and the real image;
[0030] Step 4.2: Introducing an edge-aware loss during the optimization process to optimize the boundary of the lip region, and the edge-aware loss function combines gradient enhancement regularization and contrast loss;
[0031] First, calculate the gradient map of the lip region:
[0032]
[0033] Among them, and respectively represent the gradient information of the image in the x and y directions;
[0034] Then construct the boundary gradient loss:
[0035] L edge = ∑||G(I gen ) - G(I gt )||2
[0036] Among them, I gen is the generated image, and I gt is the real image; introduce the contrast loss:
[0037] L contrast = ∑||F lip (I gen ) - F nom-lip (I gen )||2
[0038] Among them, F lip represents the features of the lip region, and F nom-lip represents the features of the non-lip region, ensuring that the boundary of the lip region is natural and clear;
[0039] The overall form of the final edge-aware loss is as follows:
[0040] L total = λ1L edge + λ2L contrast
[0041] Among them, λ1 and λ2 are hyperparameters used to balance the influence of the gradient loss and the contrast loss.
[0042] Step 4.3: Incorporate the multi-scale feature enhancement strategy, introduce feature maps of multiple different spatial scales in the generator, and perform joint modeling, fusion, and reconstruction;
[0043] Step 4.4: Extract the shape, texture, and edge information in the lip image from the intermediate layer of the generator as the embedded feature vector to provide the subsequent consistency discrimination input for Step 5.
[0044] Furthermore, the specific method of the said Step 5 includes:
[0045] Step 5.1: Calculate the cosine similarity of the embedded feature vectors of the lip region between consecutive frames to measure the smoothness and coherence of lip movement. The calculation formula is as follows:
[0046]
[0047] Among them, f t , f t+1They respectively represent the embedded feature vectors of the lip regions in the t-th frame and the (t + 1)-th frame. The closer to 1, the more coherent the movement is;
[0048] Step 5.2: Use the optical flow method to estimate the motion direction and speed of each pixel point through the pixel changes between consecutive frames, model and correct the motion of the lip region between frames, so as to optimize its temporal coherence;
[0049] Step 5.3: Apply the identity consistency loss to ensure the consistency of the identity features of the lip region throughout the video. The calculation formula is as follows:
[0050]
[0051] where, I t is the facial image of the t-th frame in the video, and I ref is the reference frame;
[0052] Step 5.4: Use the perceptual loss to optimize the lip texture and enhance the sense of hierarchy of the lip region. The calculation function formula of the perceptual loss is:
[0053]
[0054] where, φ l (·) represents the feature map extracted in the l-th layer of the neural network, I gen is the generated image, and I gt is the real reference image;
[0055] Step 5.5: If the above multiple evaluation indicators do not reach the preset threshold, feedback the error signal to the training process of Step 3 or Step 4, then adjust the model parameters and perform iterative optimization. If the evaluation result has reached the expected effect, output the final lip region image of each frame.
[0056] Furthermore, the specific method of Step 6 includes:
[0057] Step 6.1: According to the lip region image of each frame output by Step 5, perform position alignment and region replacement with the original video frame according to the corresponding timestamp;
[0058] Step 6.2: Re-encode the processed image frame sequence at the original audio frame rate to generate a lip-sync video with the same duration and resolution as the input video, and then output the final video.
[0059] Beneficial effects:
[0060] The present invention optimizes the lip trajectory through a graph neural network (GCN) to make the lip movement smoother and more natural. At the same time, it combines Wav2Lip for step-by-step convolutional fusion to gradually extract audio and lip movement features, improving the matching accuracy of audio and lips. In addition, a generative adversarial network (GAN) is used to optimize the texture details of the lip region, and the perceptual loss and style matching loss are used to enhance the light and shadow levels and authenticity of the lips. To further improve the video quality, the method also introduces an edge-aware loss to optimize the lip boundary, making the lip region clearer and reducing visual artifacts. In addition, by combining the cross-modal attention mechanism and the identity consistency loss, the lips are kept stable between different frames, reducing the problem of frame skipping or drifting. Finally, the method can generate high-resolution, natural and smooth audio-lip synchronous videos, enhancing the interaction experience of digital humans and making them suitable for real-time application scenarios such as virtual anchors, post-production of films and television, and game animations, generating more realistic and smooth audio-lip synchronous videos. Description of the Drawings
[0061] Figure 1 It is a flowchart of a digital human audio-lip synchronization method and device based on edge perception and multi-modal;
[0062] Figure 2 It is a flowchart of training data;
[0063] Figure 3 It is a structure diagram of a generative adversarial network (GAN);
[0064] Figure 4 It is a model diagram of a graph neural network (GCN);
[0065] Figure 5 It is an overall model diagram of Wav2lip. Detailed Embodiments
[0066] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art fall within the scope defined by the appended claims of this application.
[0067] The present invention discloses a digital human audio-lip synchronization method based on an adversarial texture enhancement network, including the following steps:
[0068] Step 1: Obtain video segments of the target person for audio and video data collection, and at the same time obtain the corresponding audio data. Ensure the time synchronization of the audio and video data to ensure that each frame corresponds to the correct voice information.
[0069] Step 1.1: Collect video clips of the target person and obtain the corresponding audio data.
[0070] Step 1.2: Normalize the audio data.
[0071] Step 1.3: Perform face detection, locate the lip region, and crop the lip region image.
[0072] Step 1.4: Annotate the lip region.
[0073] Step 2: Use the Wav2Lip model to extract audio and video features and perform preliminary audio-visual alignment, generate a preliminary lip region image, and extract lip dynamic features for subsequent modeling.
[0074] Step 2.1: Use the pre-trained Wav2Lip model to perform audio-visual alignment on the audio and video data in Step 1, generate a synchronized lip region image in an audio-driven manner, calculate the lip synchronization score using the pre-trained lip synchronization discriminator SyncNet, and generate the lip shape using the UNet structure.
[0075] Step 2.2: Extract the sequence of lip key point coordinates in each frame from the lip image obtained in Step 2.1 for modeling the lip movement trajectory. Characterize the dynamic change features of the lip shape by calculating the Euclidean distance between key points in different frames. The Euclidean distance calculation function formula is as follows:
[0076]
[0077] where d is the Euclidean distance, and (x1, y1) and (x2, y2) represent the spatial coordinates of the lip key points in different frame images respectively.
[0078] Step 3: Build a lip movement prediction model based on the features extracted in Step 2 using a graph neural network.
[0079] Step 3.1: Use the lip dynamic features extracted in Step 2 as input features, take the lip key points of each frame as nodes of the graph, construct a temporal graph structure, and establish the connection relationship between nodes using the adjacency matrix.
[0080] Step 3.2: Define the graph structure. Each lip key point v i is used as a node, and the connection relationship between nodes is determined by the adjacency matrix A. Calculate the feature h i of each node, where the initial feature comes from the visual features of the lip region, and the node features are represented using the feature matrix X. Each node contains information in multiple channels.
[0081] Step 3.3: Use GCN to perform multi-layer graph convolution operations on the lip key point sequence obtained in Step 2.2 to propagate and aggregate the motion information in the lip region. The function formula is as follows:
[0082]
[0083] Among them, H 1 is the node feature of the l-th layer of the input layer, H l represents the node feature of the l-th layer, A is the adjacency matrix, D is the degree matrix, W l is the trainable weight of this layer, σ is the activation function, and σ is used to introduce non-linearity so that the lip motion information can be propagated and learned more complexly, while avoiding gradient vanishing or explosion.
[0084] Step 4: Use a generative adversarial network to enhance the texture of the lip region image generated in Step 2, and improve the texture details, edge sharpness, and light and shadow layering of the lip region through perceptual loss and adversarial loss.
[0085] Step 4.1: Use the lip region image generated in Step 2 as the input to construct a generative adversarial network. The generator generates a high-definition lip image, and the discriminator is used to judge the difference in visual texture between the generated image and the real image.
[0086] Step 4.2: Introduce edge-aware loss during the optimization process to optimize the boundary of the lip region. The edge-aware loss function combines gradient enhancement regularization and contrast loss. First, calculate the gradient map of the lip region:
[0087]
[0088] Among them, and represent the gradient information of the image in the x and y directions respectively; then, construct the boundary gradient loss:
[0089] L edge = ∑||G(I gen ) - G(I gt )||2
[0090] Among them, is I gen the generated image, and I gt is the real image; introduce the contrast loss:
[0091] L contrast = ∑||F lip (I gen ) - F nom-lip (I gen )||2
[0092] Among them, F lip represents the feature of the lip region, Fnon-lip Represents the features of the non-lip region, ensuring that the boundary of the lip region is natural and clear; finally, the overall form of the edge-aware loss is as follows:
[0093] L total = λ1L edge + λ2L contrast
[0094] where λ1 and λ2 are hyperparameters used to balance the effects of gradient loss and contrast loss.
[0095] Step 4.3: Incorporate a multi-scale feature enhancement strategy by introducing feature maps that utilize multiple different spatial scales simultaneously (such as from low-resolution to high-resolution levels) in the generator for joint modeling, fusion, and reconstruction.
[0096] Step 4.4: Extract information such as shape, texture, and edges in the lip image from the intermediate layer of the generator as an embedded feature vector to provide subsequent consistency discrimination input for Step 5.
[0097] Step 5: Construct an audio-visual feature consistency evaluation module to perform multi-dimensional discrimination and optimization on the results of Steps 3 and 4 in combination with the edge-aware loss.
[0098] Step 5.1: Extract the embedded feature vectors of the lip region in consecutive frames and calculate the cosine similarity between frames to measure the smoothness and coherence of lip movement. The calculation formula is as follows:
[0099]
[0100] where f t and f t+1 represent the embedded feature vectors of the lip region in the t-th frame and the (t + 1)-th frame respectively. The closer to 1, the more coherent the movement;
[0101] Step 5.2: Adopt the optical flow method to estimate the movement direction and speed of each pixel point through the pixel changes between consecutive frames, model and correct the movement of the lip region between frames, thereby optimizing its temporal coherence;
[0102] Step 5.3: Apply the identity consistency loss to ensure the consistency of the identity features of the lip region throughout the video. The calculation formula is as follows:
[0103]
[0104] where I t is the facial image of the t-th frame in the video, and I ref is the reference frame;
[0105] Step 5.4: Optimize the lip texture using perceptual loss to enhance the layering of the lip region. According to Step 4.2, similarly, we can also obtain the formula for the perceptual loss calculation function here:
[0106]
[0107] where φ l (·) represents the feature map extracted in the l-th layer of the neural network, I gen is the generated image, and I gt is the real reference image;
[0108] Step 5.5: If the above multiple evaluation metrics do not reach the preset threshold, the error signal can be fed back to the training process of Step 3 or Step 4. Subsequently, adjust the model parameters and perform iterative optimization. If the evaluation result has reached the expected effect, output the final lip region image for each frame.
[0109] Step 6: Synthesize the optimized lip region into the original video and output the final audio-lip synchronous video.
[0110] Step 6.1: According to the lip region image for each frame output in Step 5, align the position and replace the region with the original video frame according to the corresponding timestamp.
[0111] Step 6.2: Re-encode the processed image frame sequence at the original audio frame rate to generate an audio-lip synchronous video with the same duration and resolution as the input video, and then output the final video.
[0112] Experimental example:
[0113] Obtain the LRW dataset. LRW is a large-scale audio-visual dataset collected by the Visual Geometry Group of the University of Oxford in cooperation with the BBC. It contains 500 word categories, with approximately 1000 samples in each category. The video clip length is 29 frames (about 1.16 seconds), and the word appears in the middle of the video. Convert the video to frames, extract the lip region, and then extract the audio from the video and convert it to Mel spectrogram or other audio features for the model to use. Train the model using the preprocessed data. The evaluation metrics are as follows:
[0114] LSE-D (Lip Sync Error-Distance): Measure the lip synchronization error between the generated video and the real video. The lower the value, the better the synchronization effect.
[0115] LSE-C (Lip Sync Error-Confidence): Evaluate the confidence of synchronization. The higher the value, the higher the synchronization quality.
[0116] SSIM (Structural Similarity Index): Measures the similarity between the generated image and the real image. The higher the value, the better the image quality.
[0117] The experimental results are shown in Table 1:
[0118] Method LSE-D LSE-C SSIM The method of the present invention 6.03 6.41 0.84 Wav2Lip 7.21 5.32 0.78
[0119] From the above experimental evaluation metrics, it can be seen that the present invention can generate high-resolution, natural and smooth audio-lip synchronous videos, improving the interaction experience of digital humans.
[0120] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly. It cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A digital human audio-lip synchronization method based on an adversarial texture enhancement network, characterized in that It includes the following steps: Step 1: Obtain audio and video data, collect video clips of the target person, and simultaneously obtain the corresponding audio data to ensure the time synchronization of the audio and video data; Step 2: Use the Wav2Lip model to extract audio and video features and perform preliminary audio-visual alignment to generate a preliminary lip image, and extract the lip key point coordinate sequence as the lip dynamic feature; Step 3: Based on the lip dynamic feature extracted in Step 2, construct a lip movement prediction model of the graph neural network GCN to propagate and aggregate the movement information of the lip area; Step 4: Use a generative adversarial network to enhance the texture of the lip image generated in Step 2, improve the texture details, edge sharpness, and light and shadow layering of the lip area through perceptual loss and adversarial loss, fuse the multi-scale feature enhancement strategy, and extract the embedded feature vector in the consecutive frame lip images; Step 5: Construct an audio-visual feature consistency evaluation module, combine the identity consistency loss and the edge perception loss, and perform multi-dimensional discrimination and optimization on the audio-visual features processed in Step 3 and Step 4; Step 6: Synthesize the optimized lip area into the original video and output the final audio-lip synchronous video.
2. The digital human audio-lip synchronization method based on the adversarial texture enhancement network according to claim 1, wherein The specific method of Step 1 includes: Step 1.1: Collect video clips of the target person and obtain the corresponding audio data; Step 1.2: Perform normalization processing on the audio data; Step 1.3: Perform face detection, locate the lip area, and crop the lip area image; Step 1.4: Annotate the lip area.
3. The digital human audio-lip synchronization method based on the adversarial texture enhancement network according to claim 1, wherein The specific method of Step 2 includes: Step 2.1: Use the pre-trained Wav2Lip model to perform audio-visual alignment processing on the audio and video data in Step 1, generate a synchronous lip area image in an audio-driven manner, use the pre-trained lip synchronization discriminator SyncNet to calculate the lip synchronization score, and use the UNet structure to generate the lip shape to obtain the lip image; Step 2.2: Extract the lip key point coordinate sequence in each frame according to the lip image obtained in Step 2.1 for modeling the lip movement trajectory. By calculating the Euclidean distance between the key points in different frames, the dynamic feature of the lip shape is characterized. The formula of the Euclidean distance calculation function is as follows: where d is the Euclidean distance, and (x1, y1) and (x2, y2) respectively represent the spatial coordinates of the lip key points in different frame images.
4. The digital human audio-visual lip synchronization method based on the adversarial texture enhancement network according to claim 1, wherein The specific method of Step 3 includes: Step 3.1: Use the lip dynamic feature extracted in Step 2 as the input feature, take the lip key points of each frame as the nodes of the graph, construct a temporal graph structure, and use the adjacency matrix to establish the connection relationship between the nodes; Step 3.2: Define the graph structure. Each lip key point v i is used as a node, and the connection relationship between nodes is determined by the adjacency matrix A. Calculate the feature h i of each node, where the initial feature comes from the visual feature of the lip region. The node features are represented by the feature matrix X, and each node contains information of multiple channels; Step 3.3: Use the graph neural network GCN and the lip key point coordinate sequence in each frame to perform multi-layer graph convolution operations to propagate and aggregate the movement information of the lip area. The function formula is as follows: Among them, H 1 is the node feature of the l-th layer of the input layer, and H l represents the node feature of the l-th layer. A is the adjacency matrix, D is the degree matrix, and W l is the trainable weight of this layer. σ is the activation function. σ is used to introduce non-linearity so that the lip movement information can be propagated and learned more complexly, while avoiding gradient vanishing or explosion.
5. The digital human audio-visual lip synchronization method based on the adversarial texture enhancement network according to claim 1, wherein The specific method of Step 4 includes: Step 4.1: Take the lip image generated in Step 2 as the input, construct a generative adversarial network, the generator generates a high-definition lip image, and the discriminator judges the difference between the generated image and the real image in visual texture; Step 4.2: Introduce an edge-aware loss during the optimization process to optimize the boundary of the lip region. The edge-aware loss function combines gradient-enhanced regularization and contrast loss; First, calculate the gradient map of the lip region: Among them, and respectively represent the gradient information of the image in the x and y directions; Then, construct the boundary gradient loss: Among them, it is I gen Generate an image, I gt is the real image; introduce a contrastive loss: L contrast = ∑ ||F lip (I gen ) - F non-lip (I gen ) ||² Among them, F lip represents the feature of the lip region, and F non-lip represents the feature of the non-lip region, ensuring that the boundary of the lip region is natural and clear; Finally, the overall form of the edge-aware loss is as follows: L total = λ1L edge + λ2L contrast where λ1 and λ2 are hyperparameters used to balance the effects of gradient loss and contrast loss. Step 4.3: Incorporate a multi-scale feature enhancement strategy by introducing feature maps of multiple different spatial scales in the generator for joint modeling, fusion, and reconstruction; Step 4.4: Extract shape, texture, and edge information in the lip image from the intermediate layer of the generator as an embedded feature vector to provide subsequent consistency discrimination input for Step 5.
6. The digital human voice-lip synchronization method based on the adversarial texture enhancement network according to claim 1, wherein The specific method of the said Step 5 includes: Step 5.1: Calculate the cosine similarity of the embedded feature vectors in the lip region between consecutive frames to measure the smoothness and coherence of lip movement. The calculation formula is as follows: where, f t and f t+1 respectively represent the embedded feature vectors of the lip region in the t-th frame and the (t + 1)-th frame. The closer to 1, the more coherent the movement is; Step 5.2: Adopt the optical flow method to estimate the motion direction and speed of each pixel point through the pixel changes between consecutive frames, model and correct the motion of the lip region between frames, thereby optimizing its temporal coherence; Step 5.3: Apply the identity consistency loss to ensure the consistency of identity features in the lip region throughout the video. The calculation formula is as follows: Among them, I t is the facial image of the t-th frame in the video, and I ref is the reference frame; Step 5.4: Use perceptual loss to optimize the lip texture and enhance the sense of hierarchy in the lip region. The calculation function formula of the perceptual loss: L prec = ∑ l ||φ l (I gen ) - φ l (I gt )||2 Among them, φ l (·) represents the feature map extracted in the l-th layer neural network, I gen is the generated image, and I gt is the real reference image; Step 5.5: If the above multiple evaluation indicators do not reach the preset threshold, feedback the error signal to the training process of Step 3 or Step 4, then adjust the model parameters and perform iterative optimization. If the evaluation result has achieved the expected effect, output the final lip region image of each frame.
7. The digital human voice-lip synchronization method based on the adversarial texture enhancement network according to the claims, characterized in that, The specific method of the said Step 6 includes: Step 6.1: According to the lip region image of each frame output by Step 5, perform position alignment and region replacement with the original video frame according to the corresponding timestamp; Step 6.2: Re-encode the processed image frame sequence at the original audio frame rate to generate an audio-lip synchronized video with the same duration and resolution as the input video, and then output the final video.
Citation Information
Cited By
Multi-mode collaborative video quality evaluation network based on graph structure
CN121788524A
Method for generating multi-level pronunciation visualization based on phoneme feature amplification, electronic device and storage medium
CN122474076A