A method and device for speech-driven lip shape generation
Feature extraction and fusion are performed through the deep audio feature extractor and the audio-video sequence feature fusion, combined with the lip action generator and affine transformation module, high-resolution synthetic video data is generated, solving the problem of unnatural lip matching in high-resolution videos in the prior art, and realizing the preservation of facial texture details in high-resolution videos and efficient synchronization of audio and videos.
Patent Information
- Application Number
- CN202411775994.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-05
AI Technical Summary
The prior art is difficult to achieve natural and precise lip matching in high-resolution videos, especially with a small amount of reference data, and facial texture details cannot be fully retained when generating high-resolution videos.
By obtaining the original video data and audio data containing the complete face, the deep audio feature extractor and the audio-video sequence feature fusion device are used to extract and fusion, and the lip action generator and affine transformation module are combined to generate high-resolution synthetic video data.
Natural and precise lip matching in high-resolution videos ensures full retention of facial texture details and improves audio and video synchronization and robustness.
Smart Images

Figure CN119252275B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method and device for speech-driven lip shape generation. Background Art
[0002] Voice-driven lip-syncing technology refers to the technology that controls and generates lip movements in the virtual or real world through audio signals (such as human voice). This technology is often used in facial animation of virtual characters, real-time communication and speech recognition equipment by analyzing the input voice signal and converting it into corresponding lip movements.
[0003] With the widespread application of virtual characters in media production, film and television industries, etc., generating realistic "talking heads" has become an important research topic. However, with a small amount of reference data, traditional facial synchronization technology has difficulty in achieving natural and accurate lip matching in high-resolution videos. Several existing methods generate pixels in the mouth area directly from latent vectors through convolutional neural networks. Although they have achieved certain results in low-resolution scenes, they still have serious blurring problems when generating high-resolution videos and cannot fully preserve facial texture details. In addition, the synchronization of lip movement with speech signals and the maintenance of facial expressions and head postures also bring huge challenges to existing technologies.
[0004] In the prior art, there is a lack of a lip-sync generation method for speech-driven videos with high resolution and sufficient retention of facial texture details. Summary of the invention
[0005] In order to solve the technical problem that the existing technology cannot fully preserve the facial texture details when generating high-resolution videos, the embodiment of the present invention provides a method and device for speech-driven lip shape generation. The technical solution is as follows:
[0006] In one aspect, a method for speech-driven lip-sync generation is provided, the method being implemented by a lip-sync generation device, the method comprising:
[0007] Acquire original video data containing a complete face; and obtain original audio data based on the original video data;
[0008] Based on the ffmpeg tool, image processing is performed on the original video data to obtain spliced frame image data and facial feature points; two-dimensional convolution processing is performed on the spliced frame image data to obtain spliced image features;
[0009] According to the original audio data, feature extraction is performed by a deep audio feature extractor to obtain audio features;
[0010] According to the spliced image features and the audio features, feature fusion is performed by an audio-video sequence feature fuser to obtain fusion features;
[0011] According to the facial feature points and the fusion features, a lip motion generator is used to generate a video to obtain synthetic video data;
[0012] Perform calculation according to the original video data and the synthesized video data to obtain a loss function;
[0013] According to the loss function, reversely optimize the lip movement generator to obtain an optimized lip movement generator;
[0014] Acquire target audio data; based on a preset reference sequence video, generate video according to the target audio data through the deep audio feature extractor, the audio-video sequence feature fuser and the optimized lip movement generator to obtain target synthetic video data.
[0015] On the other hand, a device for speech-driven lip-shape generation is provided, which is applied to a method for speech-driven lip-shape generation, and the device comprises:
[0016] The original data acquisition module is used to acquire the original video data containing the complete face; and obtain the original audio data according to the original video data;
[0017] A spliced image feature acquisition module is used to perform image processing based on the original video data based on the ffmpeg tool to obtain spliced frame image data and facial feature points; perform two-dimensional convolution processing on the spliced frame image data to obtain spliced image features;
[0018] An audio feature acquisition module, used to extract features through a deep audio feature extractor according to the original audio data to obtain audio features;
[0019] A feature fusion module, used to perform feature fusion through an audio-video sequence feature fuser according to the spliced image features and the audio features to obtain fusion features;
[0020] A video synthesis module, used for generating a video through a lip motion generator according to the facial feature points and the fusion features to obtain synthesized video data;
[0021] A loss function calculation module, used to perform calculations based on the original video data and the synthesized video data to obtain a loss function;
[0022] A generator optimization module, used for reversely optimizing the lip movement generator according to the loss function to obtain an optimized lip movement generator;
[0023] The target video synthesis module is used to obtain target audio data; based on a preset reference sequence video, according to the target audio data, video generation is performed through the deep audio feature extractor, the audio-video sequence feature fuser and the optimized lip movement generator to obtain target synthesized video data.
[0024] On the other hand, a lip shape generation device is provided, comprising: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned methods for speech-driven lip shape generation is implemented.
[0025] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned methods for speech-driven lip shape generation.
[0026] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0027] The present invention proposes a method for speech-driven lip generation, which can not only extract local features through a deep audio feature extractor, but also enhance the global information in the audio signal, improve the alignment accuracy of the audio-video sequence, and ensure high robustness in complex speech recognition tasks. Through the self-attention mechanism, the long-distance dependency in the audio signal is effectively captured, so that the model can extract important features from the entire audio sequence, and improve the accuracy and naturalness of lip generation. The deep temporal convolution module can fully learn the time dependency and enhance the temporal modeling capability when audio and video are synchronized, thereby effectively supporting high-precision lip generation.
[0028] Based on the bidirectional cross-attention mechanism, audio and video features can influence each other, thereby improving the model's ability to model the dependency between audio and video. This design can more accurately synchronize the lip movements in audio and video, thereby improving the accuracy and naturalness of lip generation. By integrating the bidirectional interaction of audio and video features, the model's ability to capture the association between facial expressions and speech is enhanced, making the matching of generated facial expressions and pronunciation time more accurate. Using the fusion feature deep convolution structure to further fuse features improves the fidelity of details when generating facial images, ensuring that the mouth area is synchronized with the audio while the overall facial expressions and details are realistically reproduced.
[0029] The affine transformation module can effectively adapt to different angles, scales and positions, ensuring that mouth features can be accurately extracted and generated under various viewing angles, thereby enhancing the accuracy and robustness of speech-driven lip generation. The affine transformation is simple and efficient to calculate, suitable for real-time applications, and can reduce the computational burden in the generation process and improve the system response speed. The upsampling module ensures that the generated facial image remains highly consistent at the detail level by gradually restoring the spatial resolution of the image, and can accurately present the details of the mouth shape changing with the audio, thereby achieving high-quality synchronous generation of lip shape and facial expression. The present invention is a lip generation method for speech-driven video with high resolution and sufficient retention of facial texture details. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0031] Figure 1 is a flow chart of a method for speech-driven lip shape generation provided by an embodiment of the present invention;
[0032] Figure 2 is a block diagram of a speech-driven lip shape generation device provided by an embodiment of the present invention;
[0033] Figure 3 It is a structural schematic diagram of a lip shape generating device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0035] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0036] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0037] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0038] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0039] The embodiment of the present invention provides a method for speech-driven lip-forming generation, which can be implemented by a lip-forming generation device, which can be a terminal or a server. Figure 1 The flowchart of the method for speech-driven lip generation shown in the figure may include the following steps:
[0040] S1. Obtain original video data containing a complete face; and obtain original audio data based on the original video data.
[0041] In a feasible implementation, the original video data containing complete human faces used in the present invention mainly comes from multiple video platforms, including iQiyi, Xiaohongshu, WeChat Video Account and Bilibili. These platforms provide rich video content, which provides diversified data support for the lip shape generation of the present invention, especially high-resolution video and precise audio alignment, ensuring the quality and diversity of training data.
[0042] We selected high-definition videos with clear facial expressions, significant lip movements, and many dialogue scenes from iQiyi as training data. The provided Chinese content can effectively improve the model's support for Chinese speech, especially the optimization of lip generation for Chinese pronunciation.
[0043] By crawling high-definition videos from Xiaohongshu, we obtained more natural voice and facial expression data, which enhanced the adaptability of the model in short video scenarios.
[0044] By collecting high-resolution short videos from WeChat Video Account, especially dialogue and statement videos, the model is ensured to have a stronger ability to adapt to lip movements in life and natural contexts.
[0045] By screening videos with clear voice content from Bilibili, especially some speeches, game commentary and dubbing videos, we obtained rich facial expression and audio synchronization data, especially Chinese content, which enhanced the training effect of Chinese voice. The specific number of high-definition videos is shown in Table 1 (crawled data information table).
[0046] Table 1
[0047]
[0048] S2. Based on the ffmpeg tool, image processing is performed according to the original video data to obtain spliced frame image data and facial feature points; two-dimensional convolution processing is performed on the spliced frame image data to obtain spliced image features.
[0049] Optionally, based on the ffmpeg tool, image processing is performed according to the original video data to obtain spliced frame image data and facial feature points, including:
[0050] Use ffmpeg tool to process the original video data to obtain the original frame image data and facial feature points;
[0051] According to the facial feature points, the mouth area of the original frame image data is blocked to obtain the blocked frame image data;
[0052] The blocked frame image data and the original frame image data are spliced to obtain spliced frame image data.
[0053] In a feasible implementation mode, the present invention decomposes the video file into frames by FFmpeg and converts it into a continuous image sequence. Each frame represents a time point in the video and is usually extracted at a fixed frame rate (25fps). Facial feature points are extracted from each frame of the video. By using a facial key point detection tool, the key point positions of the face in each frame can be detected and marked. Special attention is paid to the feature points in the mouth area, which are crucial for generating lip shape and facial expressions. The extracted facial feature points will be used for mouth movement analysis in subsequent processing.
[0054] The masking technology is used to generate an occlusion area based on the mouth feature point area, and the mouth area in the video frame is occluded. This process ensures that the generated facial expression is consistent with the movement of the mouth, while reducing the interference of irrelevant background. The spliced frame image data contains the processed mouth area and the complete video frame, ready for the subsequent feature fusion and generation steps.
[0055] S3. Based on the original audio data, feature extraction is performed through a deep audio feature extractor to obtain audio features.
[0056] Among them, the deep audio feature extractor includes an audio data processing module, a self-attention feature extraction module and a deep temporal convolution module;
[0057] The audio data processing module includes a plurality of Mel filters;
[0058] The self-attention feature extraction module includes multiple one-dimensional convolutional neural networks and multiple attention heads; the convolution kernels configured in the multiple one-dimensional convolutional neural networks are different;
[0059] The deep temporal convolution module includes a long short-term memory network layer, a temporal convolution network layer, and a speech deep convolution network layer.
[0060] In a feasible implementation, the present invention uses a deep audio feature extractor to introduce a self-attention mechanism and a deep temporal convolution structure in the audio feature extraction process, which can capture the global dependency information in the audio signal. The local features of the audio are extracted through a multi-layer convolutional neural network, and the self-attention mechanism is applied to the feature map, thereby achieving weighted fusion of audio features and enhancing the robustness of speech recognition.
[0061] Optionally, based on the original audio data, feature extraction is performed by a deep audio feature extractor to obtain audio features, including:
[0062] Perform short-time Fourier transform on the original audio data to obtain a spectrum diagram;
[0063] Based on multiple Mel filters, feature extraction is performed on the spectrum graph to obtain multiple Mel spectrum features;
[0064] Based on the multi-head self-attention mechanism, self-attention features are extracted according to multiple Mel spectrum features to obtain multiple self-attention features;
[0065] Calculate based on multiple Mel spectrum features and multiple self-attention features to obtain enhanced features;
[0066] Perform temporal deep feature extraction on the enhanced features to obtain audio features.
[0067] In a feasible implementation, in this step, the original audio signal (the original audio signal is obtained from different sources to ensure diversity and richness, the sampling rate is set to 16kHz, and the format is monophonic) is input into the data preprocessing module, and a short-time Fourier transform is performed to generate a spectrogram in the time-frequency domain, and a window function (such as a Hamming window) is used to reduce spectrum leakage. The spectrogram is further Mel-filtered, the spectrogram is mapped to the Mel frequency domain, and Mel spectrum features are obtained. Mel filters of different numbers and shapes are applied to capture different frequency features. Feature enhancement and normalization are further performed, such as logarithmic compression to enhance the dynamic range of the feature, and the mean and variance of the Mel spectrum are normalized to reduce the volume difference in different environments.
[0068] Construct a multi-layer one-dimensional convolutional neural network with convolution kernels of different sizes to extract local features of the audio. Each convolution layer is followed by a nonlinear activation function and a maximum pooling layer to reduce the feature dimension. Then, the output feature map is used to calculate the key, value, and query to generate weighted features. The key and value are generated from the feature map through linear transformations respectively. The query is generated through another fully connected layer. The values are weighted by attention weights to obtain a weighted feature matrix. Multiple attention heads are introduced to generate multiple feature representations through different linear transformations and attention mechanisms. The outputs of each head are concatenated to form the final self-attention output feature. The output feature is added to the self-attention output feature to obtain the enhanced feature to retain the local feature information. The enhanced features are layer-normalized to improve the stability of the model and avoid gradient vanishing or exploding phenomena.
[0069] The enhanced features are input into the bidirectional long short-term memory network layer to capture the previous and next time sequence information, and the output is the context feature. The bidirectional long short-term memory network allows information to propagate from both directions of the sequence, enhancing the ability of temporal modeling. After the bidirectional long short-term memory network, a temporal convolutional network is applied to further extract temporal features to ensure effective modeling of temporal dependencies. The causal convolutional structure is used to ensure that the model does not use future information, and deep convolution is added for further deep feature extraction (using convolution kernel), ultimately ensuring that the deep audio information corresponding to the lip sequence can be extracted.
[0070] S4. According to the spliced image features and the audio features, feature fusion is performed through an audio-video sequence feature fuser to obtain fused features.
[0071] Among them, the audio-video sequence feature fusion module includes a feature temporal integration module and a bidirectional cross-attention mechanism feature fusion module;
[0072] The feature timing integration module uses a time axis alignment algorithm to synchronize the timing of the spliced image features with the timing of the audio features;
[0073] The bidirectional cross-attention mechanism feature fusion module includes a bidirectional interactive attention calculation layer, a weighted fusion layer, and a fusion feature deep convolution layer.
[0074] In a feasible implementation, different from the previous practice of connecting audio features and video sequence features only through channels, the present invention designs an audio-sequence feature fuser based on a bidirectional cross-attention mechanism, which can better capture the correspondence between the audio sequence and the lip movement sequence, and at the same time, in the process of generating the overall mouth area, it can also better capture the lip area most affected by the audio features.
[0075] The input features of the audio-video sequence feature fuser are audio features and video features. Audio features usually refer to the feature representation obtained after the audio signal is extracted, usually Mel spectrum features or other time-frequency domain features, with time and frequency information. Here is the audio feature tensor processed by the deep audio feature extractor. Video features refer to the feature representation of each frame obtained after the video frame sequence is extracted by the convolutional neural network, which contains rich visual information of facial features.
[0076] Before feature fusion, it is necessary to ensure that the audio features and video features are aligned in the time dimension. Use a time axis alignment algorithm (such as dynamic time warping) to ensure time synchronization between audio and video frames to avoid information misalignment.
[0077] The bidirectional cross-attention mechanism can significantly improve the ability to understand multimodal information by effectively fusing video feature sequences and audio feature sequences. In this mechanism, audio features are used as queries, and the keys of video features are mutually calculated to obtain the impact of audio on video content. At the same time, video features are also used as queries, and the keys of audio features are mutually calculated to capture the impact of video on audio content. This two-way interaction enables the model to fully understand the relationship between audio and video, and further enhances the overall information representation of audio and video through weighted fused features. The two weighted fused feature vectors are fed into the deep convolutional structure ( Convolution kernels) are deeply fused and sent to the next stage for operation.
[0078] S5. According to the facial feature points and the fusion features, a lip motion generator is used to generate a video to obtain synthetic video data.
[0079] Optionally, according to the facial feature points and the fusion features, a lip motion generator is used to generate a video to obtain synthetic video data, including:
[0080] Perform convolution processing on the fused features to obtain the processed fused features;
[0081] Based on the facial feature points, an affine transformation is performed on the processed fusion features to obtain the optimized fusion features;
[0082] Convolution upsampling is performed based on the optimized fusion features to obtain synthetic video data.
[0083] In a feasible implementation manner, different from the traditional method of generating the final synthetic frame only through the sequence after the fusion of the audio sequence and the video sequence, the present invention adopts a new generation idea and uses an affine transformation module for generation.
[0084] The benefit of using affine transformation to generate the mouth region is that it can maintain the relative shape and size of the mouth region while achieving adaptability to different angles, scales, and positions. Through affine transformation, the mouth region can be rotated, translated, scaled, and so on, thereby ensuring that the mouth features can still be accurately extracted from multiple perspectives, enhancing the robustness and accuracy of speech-driven lip generation. In addition, affine transformation is simple to calculate and highly efficient, making it suitable for real-time applications.
[0085] In the input module of the lip movement generator, the main task is to receive and process the fused features of audio and video. At this stage, the features of the deep fused tensor fused by the bidirectional cross attention mechanism are further extracted through the stacking of several convolutional layers.
[0086] The affine transformation module performs spatial transformation on the processed audio-video features so that the generated facial image can correctly reflect the movement of the mouth driven by the audio. Affine transformation mainly includes translation, rotation, scaling and shearing operations. By performing affine transformation on the features, it is ensured that the facial image generated at each moment has the appropriate angle, size and position, especially in the mouth area, so that the facial expression is accurately matched with the pronunciation time in the audio. Affine transformation ensures the spatial transformation of the mouth area at each time step while maintaining the coherence of other facial features.
[0087] The convolutional structure upsampling module is responsible for upsampling the low-resolution feature map after affine transformation to the final high-resolution facial image. This module usually uses the transposed convolution (or deconvolution) layer in the convolutional neural network, which gradually increases the spatial resolution of the feature map through a series of convolutional layers while retaining the detail information in the image. Among them, upsampling refers to using the transposed convolution structure to upsample the low-resolution feature map, expand the spatial size of the feature map through interpolation, and add more spatial details. In each upsampling stage, the convolution layer is used to further extract features to ensure that the details and structures in the facial image remain consistent. As the upsampling proceeds, the convolution layer gradually restores the spatial information of the image, and optimizes it according to the audio-driven mouth shape features so that the dynamic changes of the mouth can be presented naturally.
[0088] The upsampled feature map is passed through several convolutional layers to generate the final facial image. This process produces an image with high resolution and fine facial expression details, showing the movement of the mouth as the audio signal changes. The generated facial image can accurately synchronize the mouth shape and expression in the audio signal, ensuring that the generated facial image is natural, realistic and expressive.
[0089] S6. Calculate according to the original video data and the synthesized video data to obtain a loss function.
[0090] In one feasible implementation, the loss function includes cross entropy loss, perceptual loss and L1 loss. Cross entropy loss is a loss function used for classification tasks, which measures the error by comparing the category probability distribution output by the model with the distribution of the true category. The closer the probability output by the model is to the probability of the true category, the smaller the loss; conversely, the larger the loss. Perceptual loss is mainly used for image generation tasks, which calculates the loss by comparing the difference between the generated image and the real image in high-level features (using the VGG16 convolutional network). This loss focuses on the perceptual quality of the image, especially capturing the structure and texture details of the image. L1 loss, also known as mean absolute error, is a loss function that measures the difference between the generated result and the true value. It calculates the absolute difference between the predicted value and the actual value and averages it over all samples. L1 loss tends to produce smooth results and helps remove noise.
[0091] S7. According to the loss function, reversely optimize the lip movement generator to obtain an optimized lip movement generator.
[0092] In one feasible implementation, the cross entropy loss, perceptual loss and L1 loss are used to calculate the loss between the synthetic image and the real image, and back propagation is performed to optimize the model parameters. The training is continued until the model loss no longer decreases. The training process is repeated three times, each time according to , , The resolution image is fed into the model to achieve high-definition optimization.
[0093] S8. Obtain target audio data; based on a preset reference sequence video, according to the target audio data, perform video generation through a deep audio feature extractor, an audio-video sequence feature fuser, and an optimized lip movement generator to obtain target synthetic video data.
[0094] In a feasible implementation, a reference sequence video is selected, which should contain a complete human face, and is processed according to the processing method in the training phase. Mel-spectrogram features are also extracted from the audio to be synthesized.
[0095] By feeding the Mel-spectrogram features of the audio and the spliced images of the reference frames into the model for inference in sequence, a high-definition video driven by new audio can be obtained.
[0096] The present invention proposes a method for speech-driven lip generation, which can not only extract local features through a deep audio feature extractor, but also enhance the global information in the audio signal, improve the alignment accuracy of the audio-video sequence, and ensure high robustness in complex speech recognition tasks. Through the self-attention mechanism, the long-distance dependency in the audio signal is effectively captured, so that the model can extract important features from the entire audio sequence, and improve the accuracy and naturalness of lip generation. The deep temporal convolution module can fully learn the time dependency and enhance the temporal modeling capability when audio and video are synchronized, thereby effectively supporting high-precision lip generation.
[0097] Based on the bidirectional cross-attention mechanism, audio and video features can influence each other, thereby improving the model's ability to model the dependency between audio and video. This design can more accurately synchronize the lip movements in audio and video, thereby improving the accuracy and naturalness of lip generation. By integrating the bidirectional interaction of audio and video features, the model's ability to capture the association between facial expressions and speech is enhanced, making the matching of generated facial expressions and pronunciation time more accurate. Using the fusion feature deep convolution structure to further fuse features improves the fidelity of details when generating facial images, ensuring that the mouth area is synchronized with the audio while the overall facial expressions and details are realistically reproduced.
[0098] The affine transformation module can effectively adapt to different angles, scales and positions, ensuring that mouth features can be accurately extracted and generated under various viewing angles, thereby enhancing the accuracy and robustness of speech-driven lip generation. The affine transformation is simple and efficient to calculate, suitable for real-time applications, and can reduce the computational burden in the generation process and improve the system response speed. The upsampling module ensures that the generated facial image remains highly consistent at the detail level by gradually restoring the spatial resolution of the image, and can accurately present the details of the mouth shape changing with the audio, thereby achieving high-quality synchronous generation of lip shape and facial expression. The present invention is a lip generation method for speech-driven video with high resolution and sufficient retention of facial texture details.
[0099] Figure 2 1 is a block diagram of a speech-driven lip-shape generation device according to an exemplary embodiment, wherein the device is used in a speech-driven lip-shape generation method. Figure 2 The device includes an original data acquisition module 210, a spliced image feature acquisition module 220, an audio feature acquisition module 230, a feature fusion module 240, a video synthesis module 250, a loss function calculation module 260, a generator optimization module 270 and a target video synthesis module 280. Among them:
[0100] The original data acquisition module 210 is used to acquire the original video data containing the complete face; and obtain the original audio data according to the original video data;
[0101] The spliced image feature acquisition module 220 is used to perform image processing based on the original video data based on the ffmpeg tool to obtain spliced frame image data and facial feature points; perform two-dimensional convolution processing on the spliced frame image data to obtain spliced image features;
[0102] The audio feature acquisition module 230 is used to extract features through a deep audio feature extractor according to the original audio data to obtain audio features;
[0103] A feature fusion module 240 is used to perform feature fusion according to the spliced image features and the audio features through an audio-video sequence feature fuser to obtain fusion features;
[0104] The video synthesis module 250 is used to generate a video through a lip motion generator according to facial feature points and fusion features to obtain synthesized video data;
[0105] A loss function calculation module 260, used to calculate according to the original video data and the synthesized video data to obtain a loss function;
[0106] A generator optimization module 270, configured to perform reverse optimization on the lip movement generator according to the loss function to obtain an optimized lip movement generator;
[0107] The target video synthesis module 280 is used to obtain target audio data; based on a preset reference sequence video, according to the target audio data, video generation is performed through a deep audio feature extractor, an audio-video sequence feature fuser and an optimized lip movement generator to obtain target synthesized video data.
[0108] Optionally, the spliced image feature acquisition module 220 is further used to:
[0109] Use ffmpeg tool to process the original video data to obtain the original frame image data and facial feature points;
[0110] According to the facial feature points, the mouth area of the original frame image data is blocked to obtain the blocked frame image data;
[0111] The blocked frame image data and the original frame image data are spliced to obtain spliced frame image data.
[0112] Among them, the deep audio feature extractor includes an audio data processing module, a self-attention feature extraction module and a deep temporal convolution module;
[0113] The audio data processing module includes a plurality of Mel filters;
[0114] The self-attention feature extraction module includes multiple one-dimensional convolutional neural networks and multiple attention heads; the convolution kernels configured in the multiple one-dimensional convolutional neural networks are different;
[0115] The deep temporal convolution module includes a long short-term memory network layer, a temporal convolution network layer, and a speech deep convolution network layer.
[0116] Optionally, the audio feature acquisition module 230 is further used to:
[0117] Perform short-time Fourier transform on the original audio data to obtain a spectrum diagram;
[0118] Based on multiple Mel filters, feature extraction is performed on the spectrum graph to obtain multiple Mel spectrum features;
[0119] Based on the multi-head self-attention mechanism, self-attention features are extracted according to multiple Mel spectrum features to obtain multiple self-attention features;
[0120] Calculate based on multiple Mel spectrum features and multiple self-attention features to obtain enhanced features;
[0121] Perform temporal deep feature extraction on the enhanced features to obtain audio features.
[0122] Among them, the audio-video sequence feature fusion module includes a feature temporal integration module and a bidirectional cross-attention mechanism feature fusion module;
[0123] The feature timing integration module uses a time axis alignment algorithm to synchronize the timing of the spliced image features with the timing of the audio features;
[0124] The bidirectional cross-attention mechanism feature fusion module includes a bidirectional interactive attention calculation layer, a weighted fusion layer, and a fusion feature deep convolution layer.
[0125] Optionally, the video synthesis module 250 is further configured to:
[0126] Perform convolution processing on the fused features to obtain the processed fused features;
[0127] Based on the facial feature points, an affine transformation is performed on the processed fusion features to obtain the optimized fusion features;
[0128] Convolution upsampling is performed based on the optimized fusion features to obtain synthetic video data.
[0129] The present invention proposes a method for speech-driven lip generation, which can not only extract local features through a deep audio feature extractor, but also enhance the global information in the audio signal, improve the alignment accuracy of the audio-video sequence, and ensure high robustness in complex speech recognition tasks. Through the self-attention mechanism, the long-distance dependency in the audio signal is effectively captured, so that the model can extract important features from the entire audio sequence, and improve the accuracy and naturalness of lip generation. The deep temporal convolution module can fully learn the time dependency and enhance the temporal modeling capability when audio and video are synchronized, thereby effectively supporting high-precision lip generation.
[0130] Based on the bidirectional cross-attention mechanism, audio and video features can influence each other, thereby improving the model's ability to model the dependency between audio and video. This design can more accurately synchronize the lip movements in audio and video, thereby improving the accuracy and naturalness of lip generation. By integrating the bidirectional interaction of audio and video features, the model's ability to capture the association between facial expressions and speech is enhanced, making the matching of generated facial expressions and pronunciation time more accurate. Using the fusion feature deep convolution structure to further fuse features improves the fidelity of details when generating facial images, ensuring that the mouth area is synchronized with the audio while the overall facial expressions and details are realistically reproduced.
[0131] The affine transformation module can effectively adapt to different angles, scales and positions, ensuring that mouth features can be accurately extracted and generated under various viewing angles, thereby enhancing the accuracy and robustness of speech-driven lip generation. The affine transformation is simple and efficient to calculate, suitable for real-time applications, and can reduce the computational burden in the generation process and improve the system response speed. The upsampling module ensures that the generated facial image remains highly consistent at the detail level by gradually restoring the spatial resolution of the image, and can accurately present the details of the mouth shape changing with the audio, thereby achieving high-quality synchronous generation of lip shape and facial expression. The present invention is a lip generation method for speech-driven video with high resolution and sufficient retention of facial texture details.
[0132] Figure 3 is a structural schematic diagram of a lip shape generating device provided by an embodiment of the present invention, such as Figure 3 As shown, the mouth shape generating device may include the above Figure 2 Optionally, the lip-forming device 310 may include a first processor 2001 .
[0133] Optionally, the lip shape generating device 310 may further include a memory 2002 and a transceiver 2003 .
[0134] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.
[0135] Combine the following Figure 3 The components of the lip-sync generating device 310 are described in detail:
[0136] The first processor 2001 is the control center of the lip shape generating device 310, and can be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiment of the present invention, such as one or more microprocessors (digital signal processor, DSP), or one or more field programmable gate arrays (field programmable gate array, FPGA).
[0137] Optionally, the first processor 2001 may execute various functions of the lip generation device 310 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002 .
[0138] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 3 CPU0 and CPU1 are shown in FIG.
[0139] In a specific implementation, as an embodiment, the lip shape generation device 310 may also include multiple processors, such as Figure 3 The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0140] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.
[0141] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently and access the first processor 2001 through the interface circuit ( Figure 3 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0142] The transceiver 2003 is used to communicate with a network device or a terminal device.
[0143] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0144] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently and communicate with the first processor 2001 through the interface circuit ( Figure 3 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0145] It should be noted that Figure 3 The structure of the lip shape generation device 310 shown in the figure does not constitute a limitation on the router, and the actual knowledge structure recognition device may include more or less components than those shown in the figure, or combine certain components, or arrange the components differently.
[0146] In addition, the technical effects of the lip-forming generation device 310 can refer to the technical effects of the voice-driven lip-forming generation method described in the above method embodiment, and will not be repeated here.
[0147] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0148] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0149] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0150] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0151] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0152] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0153] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0154] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0155] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0156] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0157] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0158] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0159] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for speech-driven lip generation, characterized in that: The method comprises: Acquire original video data containing a complete face; and obtain original audio data based on the original video data; Based on the ffmpeg tool, image processing is performed on the original video data to obtain spliced frame image data and facial feature points; two-dimensional convolution processing is performed on the spliced frame image data to obtain spliced image features; According to the original audio data, feature extraction is performed by a deep audio feature extractor to obtain audio features; According to the spliced image features and the audio features, feature fusion is performed by an audio-video sequence feature fuser to obtain fused features; According to the facial feature points and the fusion features, a lip motion generator is used to generate a video to obtain synthetic video data; Perform calculation according to the original video data and the synthesized video data to obtain a loss function; According to the loss function, reversely optimize the lip movement generator to obtain an optimized lip movement generator; Acquire target audio data; based on a preset reference sequence video, generate video according to the target audio data through the deep audio feature extractor, the audio-video sequence feature fuser and the optimized lip movement generator to obtain target synthetic video data.
2. The method for speech-driven lip-sync generation according to claim 1, characterized in that: The method of performing image processing based on the original video data based on the ffmpeg tool to obtain spliced frame image data and facial feature points includes: Using ffmpeg tool, the original video data is processed to obtain original frame image data and facial feature points; According to the facial feature points, masking the mouth area of the original frame image data to obtain masked frame image data; The blocked frame image data and the original frame image data are spliced to obtain spliced frame image data.
3. The method for speech-driven lip-sync generation according to claim 1, characterized in that: The deep audio feature extractor includes an audio data processing module, a self-attention feature extraction module and a deep temporal convolution module; The audio data processing module includes a plurality of Mel filters; The self-attention feature extraction module includes multiple one-dimensional convolutional neural networks and multiple attention heads; the multiple one-dimensional convolutional neural networks are configured with different convolution kernels; The deep temporal convolution module includes a long short-term memory network layer, a temporal convolution network layer and a speech deep convolution network layer.
4. The method for speech-driven lip-sync generation according to claim 1, characterized in that: The step of extracting features by a deep audio feature extractor according to the original audio data to obtain audio features includes: Performing short-time Fourier transform on the original audio data to obtain a spectrogram; Based on multiple Mel filters, feature extraction is performed on the spectrum graph to obtain multiple Mel spectrum features; Based on the multi-head self-attention mechanism, self-attention features are extracted according to the multiple Mel spectrum features to obtain multiple self-attention features; Calculate according to the multiple Mel spectrum features and the multiple self-attention features to obtain enhanced features; Performing time series deep feature extraction on the enhanced features to obtain audio features.
5. The method for speech-driven lip-sync generation according to claim 1, characterized in that: The audio-video sequence feature fuser includes a feature temporal integration module and a bidirectional cross-attention mechanism feature fusion module; The feature timing integration module uses a time axis alignment algorithm to synchronize the timing of the spliced image features with the timing of the audio features; The bidirectional cross-attention mechanism feature fusion module includes a bidirectional interactive attention calculation layer, a weighted fusion layer and a fusion feature deep convolution layer.
6. The method for speech-driven lip-sync generation according to claim 1, characterized in that: The step of generating a video by a lip motion generator according to the facial feature points and the fusion feature to obtain synthetic video data includes: Performing convolution processing on the fused features to obtain processed fused features; Based on the facial feature points, performing affine transformation on the processed fusion features to obtain optimized fusion features; Convolution upsampling is performed according to the optimized fusion features to obtain synthetic video data.
7. A speech-driven lip-shape generation device, the speech-driven lip-shape generation device is used to implement the speech-driven lip-shape generation method according to any one of claims 1 to 6, characterized in that: The device comprises: The original data acquisition module is used to acquire the original video data containing the complete face; and obtain the original audio data according to the original video data; A spliced image feature acquisition module is used to perform image processing based on the original video data based on the ffmpeg tool to obtain spliced frame image data and facial feature points; perform two-dimensional convolution processing on the spliced frame image data to obtain spliced image features; An audio feature acquisition module, used to extract features through a deep audio feature extractor according to the original audio data to obtain audio features; A feature fusion module, used to perform feature fusion through an audio-video sequence feature fuser according to the spliced image features and the audio features to obtain fusion features; A video synthesis module, used to generate a video through a lip motion generator according to the facial feature points and the fusion features to obtain synthesized video data; A loss function calculation module, used to perform calculations based on the original video data and the synthesized video data to obtain a loss function; A generator optimization module, used for reversely optimizing the lip movement generator according to the loss function to obtain an optimized lip movement generator; The target video synthesis module is used to obtain target audio data; based on a preset reference sequence video, according to the target audio data, video generation is performed through the deep audio feature extractor, the audio-video sequence feature fuser and the optimized lip movement generator to obtain target synthesized video data.
8. The speech-driven lip-forming device according to claim 7, characterized in that: The audio feature acquisition module is further used to: Performing short-time Fourier transform on the original audio data to obtain a spectrogram; Based on multiple Mel filters, feature extraction is performed on the spectrum graph to obtain multiple Mel spectrum features; Based on the multi-head self-attention mechanism, self-attention features are extracted according to the multiple Mel spectrum features to obtain multiple self-attention features; Calculate according to the multiple Mel spectrum features and the multiple self-attention features to obtain enhanced features; Performing time series deep feature extraction on the enhanced features to obtain audio features.
9. A mouth shape generating device, characterized in that: The lip shape generating device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video synthesis method and device, equipment and storage medium
CN112866586A
Visual language recognition method and device, electronic equipment and storage medium
CN114581812A