Video generation method and system based on AI voice cloning and mouth shape synchronization

By generating natural speech and synchronizing lip movements using an autoregressive Transformer and a lip-sync model, the problem of mismatch between speech content and lip movements in video production is solved, reducing costs and improving production efficiency.

CN120980180APending Publication Date: 2025-11-18成都安易迅科技有限公司

Patent Information

Application Number
CN202510869614.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In existing video production technologies, the mismatch between voice content and lip movements leads to high production costs and long production cycles. Furthermore, TTS technology cannot achieve voice cloning for specific individuals and lacks personalization and dynamic adaptation capabilities.

Method used

By using an autoregressive Transformer-based speech synthesis model and lip movement model, combined with voiceprint features and lip key points, natural speech is generated and lip movements are synchronized. The inter-frame consistency loss function and attention mechanism are used for data matching to achieve synchronization between video lip movements and speech content.

Benefits of technology

It achieves matching of video lip movements with audio content, reducing production costs, shortening the production cycle, and supporting dynamic responses to user-modified input text in real time, thus improving response efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980180A_ABST
    Figure CN120980180A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method and system based on AI voice cloning and mouth shape synchronization, and the method comprises the steps: carrying out the fusion of voiceprint features of an input video and an input text through a voice synthesis model after the input video and the input text are obtained, so as to generate a natural voice; analyzing lip key points of the input video by using a lip shape displacement model, and matching lip shape change data according to the lip key points and natural voice; and generating an output video according to the input video and the lip shape change data. According to the method, the phoneme duration prediction of the speech synthesis model and the lip displacement model can be coupled through time sequence convolution, so that the mouth shape of the output video is matched with the speech content, lightweight video restoration is realized by redrawing the lip region, the dynamic response to the input text modified by the user in real time is supported, and the response efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a video generation method and system based on AI voice cloning and lip synchronization. BACKGROUND

[0002] In the production process of a video containing voice content, a professional team needs to complete actor performance video recording and dubbing work, resulting in high production cost and long production cycle of the video containing voice content. Moreover, when modifying the voice content after the video recording is completed, video recording and dubbing work need to be performed again, further increasing the production cost and production cycle.

[0003] In order to reduce the production cost and shorten the production cycle, a voice audio can be generated through a text-to-speech (TTS) technology, that is, after a TTS model performs semantic and grammatical analysis on a text, appropriate segments are selected from a pre-recorded voice segment library for splicing according to the analysis result, and voice parameters are generated through an acoustic model, so as to be converted into a voice audio signal through a vocoder. Then, the generated voice audio is used to replace the dubbing audio, so that the voice content in the video meets the requirements.

[0004] However, after the voice audio generated through the TTS technology is used to replace the voice content in the video, the lip shape in the video playback picture does not match the voice content, and it is difficult to obtain the video effect when recording. Moreover, the TTS technology cannot realize voice cloning of a specific character, lacks personalization, and the replacement process relies on a pre-recorded video library, has weak dynamic adaptation capability, and is difficult to match any new text. SUMMARY

[0005] Therefore, the embodiments of the present application provide a video generation method and system based on AI voice cloning and lip synchronization, to solve the problem of mismatch between video lip shape and voice content.

[0006] According to one aspect of the present application, a video generation method based on AI voice cloning and lip synchronization is provided, and the method comprises:

[0007] obtaining an input video and an input text;

[0008] fusing a voiceprint feature of the input video and the input text using a speech synthesis model to generate natural speech; the speech synthesis model is a generative language model based on an autoregressive Transformer basis; the voiceprint feature includes a fundamental frequency, a formant, and a prosody pattern;

[0009] The lip key points of the input video are analyzed using a lip displacement model, and according to the lip key points, lip shape change data is matched according to the natural speech; the lip displacement model is a latent diffusion model based on audio conditions; the lip displacement model uses an inter-frame consistency loss function and performs data matching based on an attention mechanism;

[0010] An output video is generated according to the input video and the lip shape change data.

[0011] In some embodiments, the voiceprint features of the input video are fused with the input text using a speech synthesis model to generate natural speech, including:

[0012] The speech synthesis model is called, and the speech synthesis model includes a feature extraction module and a prosody synthesis module;

[0013] The voiceprint features are extracted from the input video using the feature extraction module;

[0014] The prosody features of the input text are predicted by the prosody synthesis module, and the prosody synthesis module uses an autoregressive manner to perform prosody feature prediction on the input text with the voiceprint features as a reference;

[0015] The natural speech is generated according to the prosody features.

[0016] In some embodiments, the prosody features of the input text are predicted by the prosody synthesis module, including:

[0017] The input text is subjected to word segmentation processing to obtain a keyword set;

[0018] Boundary parameters are marked in the keyword set, and the boundary parameters include prosodic word boundaries, prosodic phrase boundaries, and intonation phrase boundaries;

[0019] A prosodic embedding vector is extracted from the voiceprint features using a reference encoder;

[0020] The prosody features of the input text are predicted in an autoregressive manner based on the boundary parameters and the prosodic embedding vector.

[0021] In some embodiments, the natural speech is generated according to the prosody features, including:

[0022] A fundamental frequency curve is established according to the prosody features, and the fundamental frequency curve includes a first curve and a second curve, the first curve is a quadratic function fundamental frequency curve constructed for normal syllables; and the second curve is a linear function fundamental frequency curve established for incomplete syllables;

[0023] minimizing a difference between fundamental frequency curves of adjacent syllables using a stitching cost function to generate an overall sentence fundamental frequency curve;

[0024] generating a mel-spectrogram from the overall sentence fundamental frequency curve and the input text, and converting the mel-spectrogram into an audio signal by a vocoder.

[0025] In some embodiments, the method further comprises:

[0026] extracting audio track data from the input video;

[0027] performing speech recognition on the audio track data to obtain speech recognition text;

[0028] generating a difference text set by comparing the speech recognition text with the input text, the difference text set comprising difference keywords that are different from the input text in the speech recognition text;

[0029] determining a replacement interval based on the difference text set;

[0030] fusing a voiceprint feature of the input video corresponding to the replacement interval with the input text using a speech synthesis model to generate natural speech.

[0031] In some embodiments, the lip key points of the input video are resolved using a lip displacement model, comprising:

[0032] extracting a face image frame by frame from the input video;

[0033] detecting lip key points in the face image to obtain a key point set, the number of lip key points contained in the key point set being greater than or equal to a preset landmark point threshold;

[0034] extracting coordinates and coordinate displacements of the lip key points in adjacent two frames of the face image.

[0035] In some embodiments, the lip shape change data is matched according to the natural speech in accordance with the lip key points, comprising:

[0036] setting a mask frame of the target image according to the lip key points;

[0037] selecting a reference frame from the target image;

[0038] channel-level splicing the mask frame, the reference frame, and a noise latent variable to serve as an input of an image segmentation network;

[0039] integrating the voiceprint feature into the image segmentation network through a cross-attention layer to generate lip shape change data according to the voiceprint feature and the target image.

[0040] In some embodiments, generating an output video according to the input video and the lip shape change data comprises:

[0041] replacing a mouth area in the input video with the lip shape change data to generate a replacement frame;

[0042] obtaining an evaluation parameter of the replacement frame and a preset parameter threshold;

[0043] if the evaluation parameter is greater than or equal to the preset parameter threshold, generating the output video according to the replacement frame;

[0044] if the evaluation parameter is less than the parameter threshold, inputting the input video and the lip shape change data into a video replacement model to generate the output video by the video replacement module; the video replacement model is configured to perform replacement on a mouth area of the input video based on a generative adversarial network.

[0045] In some embodiments, inputting the input video and the lip shape change data into a video replacement model to generate the output video by the video replacement module comprises:

[0046] extracting a face image from the input video;

[0047] inputting the face image, the lip shape change data, and a mel-spectrogram of the natural speech into a pre-trained video replacement model;

[0048] generating a mouth image according to the mel-spectrogram of the natural speech by the video replacement model;

[0049] replacing a mouth area of a video frame of the input video with the mouth image to generate the output video.

[0050] According to another aspect of the present application, a video generation system based on AI speech cloning and mouth shape synchronization is provided, the system comprising:

[0051] an acquisition module configured to acquire an input video and an input text;

[0052] a speech synthesis module configured to fuse a voiceprint feature of the input video and the input text using a speech synthesis model to generate a natural speech; the speech synthesis model is a generative language model based on an autoregressive Transformer basis; the voiceprint feature comprises a fundamental frequency, a formant, and a prosody pattern;

[0053] a lip matching module configured to analyze lip key points of the input video using a lip displacement model, and match lip change data according to the natural speech according to the lip key points;

[0054] a video output module configured to generate an output video according to the input video and the lip change data.

[0055] According to yet another aspect of the present application, a computer device is provided, which includes a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, and the processor implements the above-mentioned video generation method based on AI speech cloning and lip synchronization when executing the program.

[0056] According to still another aspect of the present application, a storage medium is provided, which stores a computer program, and the program is executable on a processor to implement the above-mentioned video generation method based on AI speech cloning and lip synchronization.

[0057] By means of the above technical solutions, the present application provides a video generation method and system based on AI speech cloning and lip synchronization, which can first fuse the voiceprint features of the input video and the input text using a speech synthesis model to generate natural speech after obtaining the input video and the input text, then analyze the lip key points of the input video using a lip displacement model, and match lip change data according to the natural speech according to the lip key points, and finally generate an output video according to the input video and the lip change data. The method can couple the phoneme duration prediction of the speech synthesis model and the lip displacement model through a time convolution, so that the lip shape of the output video matches the speech content, and the method also supports dynamic response to real-time modification of the input text by the user, thereby improving the response efficiency.

[0058] The above description is only a summary of the technical solutions of the present application, in order to more clearly understand the technical means of the present application, the specific embodiments of the present application can be implemented according to the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0059] The accompanying drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:

[0060] Figure 1 A flowchart of a video generation method based on AI speech cloning and lip synchronization provided by the embodiments of the present application is shown.

[0061] Figure 2 A video generation process schematic diagram provided for an embodiment of the present application;

[0062] Figure 3 A natural speech generation process schematic diagram provided for an embodiment of the present application;

[0063] Figure 4 An output video generation process schematic diagram provided for an embodiment of the present application;

[0064] Figure 5 An evaluation parameter processing process schematic diagram provided for an embodiment of the present application;

[0065] Figure 6 A video generation process schematic diagram based on a replacement interval provided for an embodiment of the present application;

[0066] Figure 7 A video generation system structure schematic diagram based on AI voice cloning and lip synchronization provided for an embodiment of the present application. DETAILED DESCRIPTION

[0067] Hereinafter, the present application will be described in detail with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0068] In an embodiment of the present application, video generation refers to a data processing process of obtaining an output video by replacing part of the speech content in an input video with speech content corresponding to input text. The input video is a video containing speech content. Since the video containing speech content requires a professional team to complete the actor performance video recording and dubbing work in the production process, the production cost of the video containing speech content is high, and the production cycle is long. Moreover, when modifying the speech content after video recording is completed, video recording and dubbing work need to be performed again, further increasing the production cost and the production cycle.

[0069] In order to reduce the production cost and shorten the production cycle, in some embodiments, the speech audio can be generated according to the input text by using a text-to-speech technology, and the generated speech audio can be used to replace the dubbing audio, so that the speech content in the video meets the speech content requirements of the input text. The TTS technology is a technology for converting text content into speech output, which can convert text information into audible speech audio signals.

[0070] For example, when generating speech audio through TTS technology, preprocessing can be performed on the input text first, that is, analyzing the input text, including word segmentation, syntax analysis, semantic understanding, etc., so as to correctly process punctuation, numbers, abbreviations, etc. in the text. Then, text-to-phoneme conversion is realized through a TTS model, so as to convert the text into a phoneme sequence. Among them, a phoneme is the smallest unit of speech, such as the initial and final consonants in pinyin. Then, through speech synthesis, pre-recorded speech segments are spliced (splicing synthesis) or speech signals are generated according to the phoneme sequence using a vocoder (parametric synthesis). Therefore, after semantic and syntactic analysis of the text through the TTS model, appropriate segments are selected from the pre-recorded speech segment library for splicing according to the analysis results, and speech parameters are generated through an acoustic model, so as to be converted into a speech audio signal through a vocoder.

[0071] However, after the speech content in the video is replaced by the speech audio generated through the TTS technology, the mouth shape in the video playback picture does not match the speech content, and it is difficult to obtain the video effect when recording. Moreover, the TTS technology cannot realize voice cloning of a specific person, lacks personalization, and the replacement process relies on a pre-recorded video library, has weak dynamic adaptation capability, and is difficult to match any new text.

[0072] To solve the problem of mismatch between video mouth shape and speech content, an AI voice cloning and mouth shape synchronization-based video generation method is provided in some embodiments of the present application. The method can be applied to electronic devices with data processing capability. The electronic devices include but are not limited to computers, servers, mobile terminals, smart wearable devices, industrial control machines, etc. For ease of description, in the embodiments of the present application, the electronic device is taken as the execution subject of the method. It should be understood that the method can also be applied to other types of execution subjects, which are not shown one by one in the embodiments of the present application. As shown in Figure 1 , Figure 2 The method includes the following steps.

[0073] S101, obtaining an input video and an input text;

[0074] To generate an output video, the electronic device can first obtain an input video and an input text. The input video is the original video to be replaced with speech content, which can be a video recorded during the performance of an actor. The input video can be a real-time recorded video or a pre-recorded video. Taking an advertisement video as an example, when an actor A performs according to the script requirements and makes speaking actions, a 5-second video segment of a mute version of the actor A's endorsement can be recorded by a video recording device such as a camera, so as to generate an original advertisement video.

[0075] When the input video is acquired, the user can upload the input video to the electronic device through an interactive action performed on the electronic device. For example, a video generation application for performing the AI voice cloning and lip synchronization based video generation method can be pre-configured in the electronic device. When the video generation application is run by the electronic device, an input interaction interface can be presented, which can include a video upload control for uploading a video and a text input control for inputting text. When the user clicks the upload control and specifies the storage path of the input video, the video generation application can acquire the input video based on the storage path.

[0076] The electronic device can also acquire the input video by sending a video acquisition instruction to a video source device. For example, the generated original advertisement video can be stored in the storage medium of a server. When it is necessary to modify the voice content in the advertisement video, the electronic device can send a video acquisition instruction to the server, and the server can feed back a video download link to the electronic device in response to the video acquisition instruction. The electronic device can then acquire the input video, i.e., “input_video”, by accessing the video download link.

[0077] The input text, as constraint data for replacing the voice content in the input video, can be text input by the user in real time. For example, for a 5-second video clip of a silent version of a driving video endorsed by actor A, the user can input text data “new intelligent driving, safely reach every journey” through the text input control in the input interaction interface. After the user clicks the button control of the confirmation function, the electronic device can acquire the input text through the text input control.

[0078] The electronic device can also be text recognized from a file uploaded by the user. The uploaded file can be a text file, a multimedia file, an image file, etc. For example, in addition to the video upload control and the text input control, the input interaction interface can also include a file upload control. The user can specify the storage path of the uploaded file through the file upload control. After the user clicks the button control of the confirmation function, the electronic device can acquire the uploaded file based on the storage path.

[0079] After the uploaded file is acquired, the electronic device can also determine the type of the uploaded file. If the type of the uploaded file is a text file, the input text can be directly read from the file. If the type of the uploaded file is a non-text file, a character recognition process can be called, and the uploaded file can be subjected to character recognition based on the character recognition process to acquire the input text from the uploaded file.

[0080] For example, when the user uploads a picture with the text "new intelligent driving, safely reach every journey" through a file upload control in the input interaction interface, the electronic device can call an optical character recognition (OCR) tool through the file format of the uploaded file. And perform text recognition on the picture through the OCR tool to obtain the input text, i.e. "ad_text".

[0081] S102, using a speech synthesis model to fuse the voiceprint features of the input video and the input text to generate natural speech.

[0082] After obtaining the input video, the electronic device can perform voiceprint feature extraction on the input video based on a speech synthesis model, and fuse the voiceprint features of the input video with the input text to generate natural speech. Wherein, the speech synthesis model is a pre-trained autoregressive Transformer-based generative language model. The speech synthesis model can realize text encoding and semantic understanding, that is, the speech synthesis model can use a pre-trained text base large model as a text encoder, so that the speech synthesis model can perform deep semantic understanding on the input text and convert it into high-level text tokens, providing a more accurate semantic basis for speech synthesis.

[0083] The speech synthesis model can also realize speech token generation, that is, generate speech tokens for input text through an autoregressive Transformer-based language model. The generated speech tokens are intermediate representations of speech synthesis, which can combine the semantic information of text with the acoustic features of speech. The speech synthesis model can train a larger codebook through a tokenizer, achieve higher activation rate, and significantly improve the accuracy of pronunciation and the consistency of timbre.

[0084] The speech synthesis model can also convert speech tokens into mel-spectrogram through the built-in flow matching model, introducing timbre and other acoustic details through speaker embedding and reference speech. The speech synthesis model can achieve smooth transition of acoustic parameters based on ODE-based technology, thereby generating more natural speech. And in the streaming mode, the speech synthesis model can receive text streams and generate speech in real time, with short first packet response time. Through the Chunk-Aware Causal Flow Matching (CA-CFM) technology, redundant calculations of word-by-word processing are avoided.

[0085] The voice synthesis model can use a pre-trained vocoder to convert the mel spectrum into an original audio signal while restoring the phase information of the audio, thereby generating the final voice output. In addition, the voice synthesis model can also support emotion, speaking style and fine-grained control instructions. Users can adjust the emotional expression, speech rate, intonation, etc. of the voice through simple instructions, so as to realize more personalized and lively voice synthesis.

[0086] In the process of generating natural speech using the voice synthesis model, the electronic device can call the model pre-trained by the electronic device, or run a third-party model. For example, the voice synthesis model is a CosyVoice model. The CosyVoice model adopts a modeling framework combining a large language model (LLM) with a factorization machine (FM), uses a pre-trained text base large model such as Qwen2.5-0.5B, replaces the text encoding and random Transformer (Text Encoder, randomTransformer) structure, and performs text semantic modeling.

[0087] In generating the output video, the electronic device can input the input video into the voice synthesis model to perform voiceprint feature recognition on the audio part in the input video through the voice synthesis model to obtain the voiceprint feature of the input video. The voiceprint feature includes fundamental frequency, formant, and prosody pattern. The fundamental frequency (F0) is the basic frequency of vocal cord vibration and can determine the pitch perception of the voice. The formant is the spectral energy concentration area produced by the vocal tract resonance, which can reflect the timbre characteristics. The prosody pattern (Prosody) is a suprasegmental feature, which can include rhythm, stress and intonation, etc. In continuous speech, the voiceprint features can be dynamically coupled, such as in a happy emotional voice state, the happy time has a wider fundamental frequency range, an increased formant bandwidth, and a faster prosodic rhythm, etc. Through accurate recognition and modeling of the voiceprint features, a voice content similar to the voice style of the input video can be synthesized.

[0088] The voice synthesis model can also fuse the generated voiceprint features with the input text to obtain natural speech that conforms to the voiceprint features in the input video and the content of the input text. Therefore, as shown in some embodiments, when the electronic device performs fusion of the voiceprint features of the input video and the input text using the voice synthesis model to generate natural speech, the voice synthesis model can be called first, and the voice synthesis model includes a feature extraction module and a prosody synthesis module. Figure 3

[0089] ​The feature extraction module is reused to extract the voiceprint features from the input video, and the prosody synthesis module is used to predict the prosody features of the input text, so as to generate the natural speech according to the prosody features. The prosody synthesis module uses an autoregressive manner to perform prosody feature prediction on the input text with the voiceprint features as a reference.

[0090] The prosody synthesis module is a key component in the speech synthesis model for modeling and generating natural prosody. The prosody synthesis module can make the synthesized speech more natural and expressive by capturing features such as pitch, speech rate, volume, and rhythm. The prosody synthesis module can implement prosody modeling, i.e., modeling the prosody features of speech, including pitch, speech rate, and volume. The prosody synthesis module can allow multiple reference audios to be used to extract prosody features by introducing a multi-reference timbre encoder, thereby improving the diversity and naturalness of the synthesized speech.

[0091] The prosody synthesis module can use an autoregressive manner to predict the prosody features of the target text, and can generate prosody labels corresponding to the target text according to the input text and the reference audio. The prosody features can also be controlled by learning a discrete representation of the prosody features.

[0092] For example, the electronic device can use the CosyVoice model to extract voiceprint features from the input video, i.e., based on the audio part of the input video, extract fundamental frequency, formant, prosody pattern, and other voiceprint features. Then, through the Prosody-LM sub-module, the input text and the voiceprint features are fused to generate natural speech in wav format.

[0093] In some embodiments, to generate natural speech, the electronic device can perform word segmentation on the input text through the prosody synthesis module to obtain a set of keywords and label boundary parameters in the set of keywords when predicting the prosody features of the input text through the prosody synthesis module. The boundary parameters include prosodic word boundaries, prosodic phrase boundaries, and intonational phrase boundaries. Then, the prosody embedding vector is extracted from the voiceprint features using a reference encoder, and the prosody features of the input text are predicted in an autoregressive manner based on the boundary parameters and the prosody embedding vector.

[0094] After predicting the prosodic features of the input text, the electronic device can also generate the natural speech based on the prosodic features through a prosodic synthesis module. Specifically, the prosodic synthesis module establishes a fundamental frequency curve based on the prosodic features. This fundamental frequency curve includes a first curve and a second curve. The first curve is a quadratic function fundamental frequency curve constructed for normal syllables; the second curve is a linear function fundamental frequency curve constructed for incomplete syllables. Then, a concatenation cost function is used to minimize the difference in fundamental frequency curves between adjacent syllables to generate the fundamental frequency curve of the entire sentence. A Mel spectrogram is generated based on the fundamental frequency curve of the entire sentence and the input text, and the Mel spectrogram is converted into an audio signal using a vocoder.

[0095] For example, when using prosodic synthesis modules such as Prosody-LM, electronic devices can first perform text preprocessing and analysis. This involves cleaning and segmenting the input text to ensure proper formatting. Then, prosodic structure annotation is performed, using a prosodic prediction model to annotate the text and determine the boundaries of prosodic words (PW), prosodic phrases (PPH), and intonation phrases (IPH). For instance, for the sentence "The weather is really nice today," the annotation results might be: Prosodic words (PW): ['today', 'weather', 'nice']; Prosodic phrases (PPH): ['today', 'nice']; Intonation phrases (IPH): ['The weather is really nice today'].

[0096] After text annotation, the electronic device can extract prosodic features based on the prosodic synthesis module. By providing reference audio, a reference encoder extracts prosodic embedding vectors from the reference audio. These vectors can contain prosodic features of the reference audio, such as pitch, speech rate, and intensity. Global prosodic tokens (GSTs) are then used to represent different prosodic styles. Finally, the reference encoder predicts the GST weight combination for a given reference audio.

[0097] After extracting prosodic features, electronic devices can perform predictions based on these features, i.e., autoregressive prediction. Prosodic synthesis modules such as Prosody-LM can progressively predict the prosodic features of the target text through autoregression and generate prosodic tokens based on contextual information such as prosodic words and phrase boundaries, as well as the prosodic embedding of reference audio. Then, through prosodic modeling, the prediction of prosodic features is implicitly embedded in the language model generation process. That is, in CosyVoice, prosodic information is determined by Text-Speech LM when predicting semantic tokens based on text context and reference audio cues.

[0098] According to the prediction result of the prosodic feature, the electronic device can generate the prosodic feature, i.e., modeling the fundamental frequency curve, using a quadratic function at 2 +bt+c represents the fundamental frequency curve for a normal syllable; and using a linear function mt+n represents the fundamental frequency curve for an unfinished syllable. By predicting boundary prosodic parameters such as the initial value, the terminal value, the initial slope, and the terminal slope of the fundamental frequency, and minimizing the difference between the fundamental frequency curves of adjacent syllables by using a stitching cost function, the fundamental frequency curve of the entire sentence is generated.

[0099] The prosodic feature is further embedded by combining the prosodic feature with an acoustic model. The predicted prosodic feature is embedded into an acoustic model such as FastSpeech 2, so that the prosodic feature can directly modify the predicted prosodic parameter value. The acoustic model can generate a mel-spectrogram according to the embedded prosodic feature, and then convert it into an audio signal through a vocoder such as HiFi-GAN.

[0100] As can be seen, based on the natural speech generation method described in the above embodiments, the electronic device can realize voiceprint cloning and speech generation based on the input video, by using the CosyVoice model to extract voiceprint features such as the fundamental frequency, formant, and prosodic pattern in the input video. And by using the Prosody-LM sub-module to fuse the input text and the cloned voiceprint, natural speech is generated, i.e., cloned_voice = cosyvoice.clone(video = input_video, text = ad_text).

[0101] S103, using a lip displacement model to analyze the lip key points of the input video, and according to the lip key points, matching the lip shape change data according to the natural speech.

[0102] After generating the natural speech, the electronic device can identify the lip features and regions of the input video through the lip displacement model to analyze the lip key points of the input video. The lip displacement model is a latent diffusion model based on audio conditions; the lip displacement model uses an inter-frame consistency loss function and performs data matching based on an attention mechanism.

[0103] As an audio-conditioned latent diffusion model, the lip displacement model can use a latent diffusion model such as Stable Diffusion to directly model complex audio-visual associations in the latent space. The lip displacement model can use Whisper to convert the audio spectrogram into an audio embedding, and integrate the audio embedding into a U-Net model through a cross-attention layer.

[0104] The lip displacement model can also extract temporal representations through a large-scale self-supervised video model, enhancing the temporal consistency of generated frames and real frames. During training, the difference between the temporal representations of generated frames and real frames is calculated as an additional loss, reducing video flicker. The lip displacement model can use a pre-trained synchronization network (SyncNet) to supervise the generated video, ensuring precise synchronization of lip movements and audio.

[0105] To parse the lip key points of the input video, in some embodiments, the electronic device can first extract face images frame by frame from the input video when performing lip key point parsing of the input video using the lip displacement model, and then detect the lip key points in the face images to obtain a key point set. Then, by extracting the coordinates and coordinate displacements of the lip key points in adjacent two frames of the face images, the lip key points of the input video are obtained.

[0106] The number of lip key points contained in the key point set is greater than or equal to a preset marker point threshold. The preset marker point threshold can be set according to the video frame resolution of the input video or the required resolution of the output video. For example, the lip displacement model is a latent diffusion model based on the LatentSync engine. When performing dynamic lip synchronization, the electronic device can first call the LatentSync engine and parse the lip key points of the input video through the LatentSync engine. In order to generate an output video with a resolution of 1080P and a frame rate of ≥30fps, the preset marker point threshold can be set to 68, that is, the LatentSync engine can parse 68 lip marker points in each face image.

[0107] After parsing the lip key points, the electronic device can match the lip movement change data according to the natural speech according to the lip key points. The lip movement change data is a data used to represent the lip movement change, which can include the association between phonemes and lip movement changes. Phonemes are the basic units of speech, and lip movement changes refer to the movement and shape changes of the lips when speaking. Since different phonemes will cause different lip movement changes when pronouncing, and lip movement changes can be represented by the lip shape in multiple frames of images, the lip movement change data can be multiple frames of lip images and phonetic labels set for the lip images.

[0108] To match the lip shape variation data, in some embodiments, the electronic device, when performing matching the lip shape variation data according to the natural speech according to the lip key points, can first set a mask frame of the target image according to the lip key points, and select a reference frame from the target image. Then, the mask frame, the reference frame and the noise latent variable are channel-level spliced as inputs of an image segmentation network. Thus, the voiceprint features are integrated into the image segmentation network through a cross-attention layer, so as to generate the lip shape variation data according to the voiceprint features and the target image.

[0109] For example, when the electronic device uses the LatentSync engine to match the phonemes and the lip shape variation, it can first perform audio feature extraction, that is, use the pre-trained audio feature extractor such as Whisper in the LatentSync engine to convert the audio into a mel-spectrogram and extract the audio embedding. In order to provide more rich temporal information, the model will bundle audio from multiple surrounding frames together as input.

[0110] Then, through video preprocessing, each frame of image is extracted from the input video, and face detection and alignment are performed. Affine transformation can be used to realize face frontalization, so that the model can more effectively learn facial features. Then, a mask frame is generated for each frame of image respectively, which is used to cover the mouth area in the face image. Then, a reference frame is selected from the input video, and the mask frame, the reference frame and the noise latent variable are channel-level spliced as inputs of the U-Net.

[0111] The LatentSync engine can be based on a latent diffusion model, and directly model the complex audio-video relationship in the latent space using the latent diffusion model conditioned on the audio embedding. The audio embedding is integrated into the U-Net through a cross-attention layer, so as to generate the lip shape variation synchronized with the audio according to the audio features and the input image. A large-scale self-supervised video model such as VideoMAE-v2 can also be used to extract temporal representations, and the difference between the generated frame and the real frame is calculated as an additional loss to enhance temporal consistency.

[0112] The electronic device can also use a pre-trained SyncNet to supervise the generated video, to ensure accurate synchronization of the lip shape and the audio. An inter-frame consistency loss function is added in the pixel space as the SyncNet loss, to optimize the accuracy of the lip shape synchronization. The generated latent representation can also be decoded into an image in the pixel space, and a learned perceptual image patch similarity (LPIPS) loss function is used to optimize the visual quality of the generated image.

[0113] As can be seen, through the lip movement data matching method described in the above embodiments, the electronic device can use the inter-frame consistency loss function to ensure natural lip transitions and avoid jitter. Furthermore, by aligning the speech spectrum with the multimodal lip movements and matching phonemes with lip movements through an attention mechanism, lip movement data is obtained. That is, lip_movement = latentsync.predict(cloned_voice.phonemes).

[0114] S104. Generate an output video based on the input video and the lip shape change data.

[0115] like Figure 4 As shown, after generating lip movement data, the electronic device can perform lip region replacement on the input video, that is, generate the output video based on the input video and lip movement data. Specifically: output_video = render(input_video, lip_movement, mask_region = MOUTH_AREA).

[0116] like Figure 5 As shown, in order to generate an output video, in some embodiments, when the electronic device generates an output video based on the input video and the lip movement data, it can first replace the lip shape region in the input video with the lip movement data to generate a replacement frame, and obtain the evaluation parameters and preset parameter thresholds of the replacement frame. The preset parameter thresholds can be set according to the resolution of the input video or the required resolution of the output video. The evaluation parameters of the replacement frame can be obtained through manual scoring or based on an image quality scoring model.

[0117] For example, an image quality scoring model can compare the lip shape image in the replacement frame with a preset standard lip shape, calculate the similarity between the two, and then calculate evaluation parameters based on the similarity. Another example is that an image quality scoring model can perform edge detection on the lip region in the replacement frame, quantize the detected edge size, edge color difference, and other features, and then perform a weighted sum of the quantization results to calculate evaluation parameters.

[0118] After obtaining the evaluation parameters and preset parameter thresholds, the electronic device can compare the evaluation parameters with the preset parameter thresholds. If the evaluation parameters are greater than or equal to the preset parameter thresholds, it means that the lip displacement model matching the lip change data can meet the output video requirements. Therefore, the output video can be generated according to the replacement frame, that is, the video frames in the corresponding time interval of the input video are replaced by the replacement frame.

[0119] If the evaluation parameter is less than the parameter threshold, it indicates that the lip movement model cannot match the lip shape change data to meet the output video requirements. Therefore, the input video and the lip shape change data can be input into a video replacement model to generate the output video through the video replacement module. The video replacement model is configured to perform lip shape replacement on the input video based on a generative adversarial network (GAN). For example, the video replacement model can use GAN inpainting techniques to replace the original video's lip shape region while preserving the original background and facial expression.

[0120] To generate the output video, in some embodiments, when the electronic device executes the process of inputting the input video and the lip movement data into a video replacement model to generate the output video through the video replacement module, it can first extract a face image from the input video, and then pass the face image, the lip movement data, and the Mel-spectrum of the natural speech as input to a pre-trained video replacement model. The video replacement model then generates a lip shape image based on the Mel-spectrum of the natural speech, and uses the lip shape image to replace the lip shape regions of the video frames in the input video to generate the output video.

[0121] For example, an electronic device can first invoke a video replacement model and use video processing libraries such as OpenCV's VideoCapture within the model to parse the input video file, reading the video frame by frame. Simultaneously, the corresponding audio file for natural speech is converted to a unified format, and the Mel-spectrum of the audio is extracted. Then, based on face detection tools such as Dlib or MTCNN, face localization is performed on each frame to obtain the coordinates of the face regions. Furthermore, the detected faces are aligned.

[0122] After face detection and localization, lip movements can be generated based on a GAN model. The face image and the Mel-spectrum of the audio are taken as input and fed into a pre-trained GAN model, allowing the model to generate lip-sync images based on audio features. The generated lip-sync images are then fused with frames from the input video to replace the corresponding lip-sync regions in the original video. Image fusion techniques, such as the Laplacian pyramid blending algorithm, are used to ensure a natural transition between the replaced lip-sync regions and the surrounding image. Furthermore, electronic devices can reconstruct the video, reassembling each processed frame and preserving the original audio. Finally, instigation networks such as GFP-GAN are used to optimize the image quality of the generated video, improving its clarity.

[0123] As can be seen, through the video synthesis steps described in the above embodiments, the video generation method based on AI voice cloning and lip-sync can achieve cross-modal alignment by temporally coupling the phoneme duration prediction of CosyVoice with the lip movement model of LatentSync through convolution, thus solving the problem of model fragmentation. It also employs local GAN ​​training, redrawing only the lip region to achieve lightweight video restoration and reduce computational requirements. Furthermore, it supports real-time text modification by users and generates a new version of the output video in a short time, achieving dynamic response and real-time interactive expansion. This method provides an end-to-end video generation solution, enabling the input of any person's video and input text, and outputting a video in which that person reads the text with a cloned voice and perfectly synchronized lip movements, thereby reducing the production cost of customized videos.

[0124] In some embodiments, as a refinement and extension of the specific implementation of the above embodiments, and in order to fully illustrate the specific implementation process of this embodiment, some embodiments of this application also provide a video generation method based on AI voice cloning and lip-syncing, such as... Figure 6 As shown, before fusing the speaker characteristics of the input video with the input text using a speech synthesis model to generate natural speech, the method further includes:

[0125] S201. Extract audio track data from the input video;

[0126] S202. Perform speech recognition on the audio track data to obtain speech-recognized text;

[0127] S203. By comparing the speech recognition text with the input text, a set of difference texts is generated. The set of difference texts includes difference keywords, which are keywords in the speech recognition text that are different from those in the input text.

[0128] S204. Determine the replacement interval based on the set of differing texts;

[0129] S205. Use a speech synthesis model to fuse the voiceprint features of the input video corresponding to the replacement interval with the input text to generate natural speech.

[0130] After acquiring input video and input text, electronic devices can separate audio track data from the input video and perform speech recognition on the audio track data using a speech-to-text tool to obtain speech-recognized text. For example, when the input video contains audio with the content "New autonomous driving, safe arrival on every journey," a speech-to-text tool based on speech recognition, natural language processing (NLP), and machine learning (ML) technologies can extract the recognized text information from the audio, thus obtaining the speech-recognized text "New autonomous driving, safe arrival on every journey."

[0131] After obtaining the speech-recognized text, the electronic device can compare the speech-recognized text with the input text to identify keywords in the speech-recognized text that differ from the input text. For example, when the speech-recognized text is "New autonomous driving, safe arrival on every journey," while the input text is "New intelligent driving, safe arrival on every journey," the difference keyword can be identified as "autonomous" through comparison.

[0132] Then, based on the set of differing texts, the replacement interval is determined, that is, the playback position of the keywords in the audio track data is determined, as well as the context information (adjacent keywords) of the differing keywords. For example, the playback position of the differing keyword "automatic" is 0:00:02:035; and its adjacent keywords are "new" and "driving". Since the interval between "new autonomous driving" and "driving" is very short according to the pronunciation characteristics of actor A, the replacement interval can be determined as [0:00:02:000, 0:00:03:137], that is, the playback time period corresponding to the voice of "new autonomous driving" in the input video.

[0133] Once the replacement interval is determined, the electronic device can use a speech synthesis model to fuse the speaker characteristics of the input video corresponding to the replacement interval with the input text to generate natural speech. This means that in the subsequent process of replacing speech content and lip movements, processing is only performed on the replacement interval. Therefore, by setting replacement intervals, not only can the amount of data processing be reduced, but the contextual information of the different keywords corresponding to the replacement interval can also be taken into account for intelligent replacement, thus improving the quality of video generation.

[0134] In some embodiments, as a specific implementation of the video generation method based on AI voice cloning and lip-syncing described in the above embodiments, some embodiments of this application also provide a video generation system based on AI voice cloning and lip-syncing, such as... Figure 7 As shown, the device includes:

[0135] According to another aspect of this application, a video generation system based on AI voice cloning and lip-syncing is provided, the system comprising:

[0136] The acquisition module is used to acquire input video and input text;

[0137] The speech synthesis module is used to fuse the voiceprint features of the input video with the input text using a speech synthesis model to generate natural speech; the speech synthesis model is a generative language model based on autoregressive Transformer; the voiceprint features include fundamental frequency, formants and prosodic patterns;

[0138] The lip-shape matching module is used to parse the lip key points of the input video using a lip-shape displacement model, and to match lip shape change data based on the natural speech according to the lip key points; the lip-shape displacement model is a latent diffusion model based on audio conditions; the lip-shape displacement model adopts an inter-frame consistency loss function and performs data matching based on an attention mechanism;

[0139] The video output module is used to generate an output video based on the input video and the lip shape change data.

[0140] By applying the technical solutions of the above embodiments, the video generation system based on AI voice cloning and lip-sync provided in the above embodiments can, after the acquisition module acquires the input video and input text, the speech synthesis module uses a speech synthesis model to fuse the voiceprint features of the input video with the input text to generate natural speech; then, the lip-sync matching module uses a lip-sync displacement model to analyze the lip key points of the input video, and matches lip-sync change data according to the lip key points and natural speech. Finally, the video output module generates the output video based on the input video and lip-sync change data. The system can couple the phoneme duration prediction of the speech synthesis model with the lip-sync displacement model through temporal convolution to match the lip shape of the output video with the speech content, and uses lightweight video repair by redrawing the lip region. It also supports dynamic response to user-modified input text in real time, improving response efficiency.

[0141] It should be noted that other corresponding descriptions of the functional units involved in the video generation system based on AI voice cloning and lip-syncing provided in the embodiments of this application can be found in the corresponding descriptions in the video generation method based on AI voice cloning and lip-syncing provided in the above embodiments, and will not be repeated here.

[0142] This application also provides a computer device, specifically a personal computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interfaces and a display device. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores location information. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0143] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.

[0144] In one embodiment, a computer-readable storage medium is also provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0145] In one embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0146] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0147] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.

[0148] Any references to memory, database, or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc.

[0149] Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can take many forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0150] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchain. The processors involved in the embodiments provided in this application may be, but are not limited to, general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc.

[0151] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0152] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A video generation method based on AI voice cloning and lip-syncing, characterized in that, The method includes: Get the input video and input text; The input video's voiceprint features are fused with the input text using a speech synthesis model to generate natural speech; the speech synthesis model is a generative language model based on the autoregressive Transformer; the voiceprint features include fundamental frequency, formants, and prosodic patterns. The lip shift model is used to parse the lip key points of the input video, and lip shape change data is matched according to the natural speech based on the lip key points; the lip shift model is a latent diffusion model based on audio conditions; the lip shift model adopts the inter-frame consistency loss function and performs data matching based on the attention mechanism; An output video is generated based on the input video and the lip shape change data.

2. The method according to claim 1, characterized in that, The input video's speaker characteristics are fused with the input text using a speech synthesis model to generate natural speech, including: The speech synthesis model is invoked, which includes a feature extraction module and a prosody synthesis module; The feature extraction module is used to extract the voiceprint features from the input video; The prosodic synthesis module predicts the prosodic features of the input text, and the prosodic synthesis module uses the voiceprint features as a reference to perform prosodic feature prediction on the input text in an autoregressive manner. The natural speech is generated based on the prosodic features.

3. The method according to claim 2, characterized in that, Predicting the prosodic features of the input text using the prosodic synthesis module includes: Perform word segmentation on the input text to obtain a keyword set; Boundary parameters are labeled in the keyword set, including prosodic word boundaries, prosodic phrase boundaries, and intonation phrase boundaries; Use a reference encoder to extract prosodic embedding vectors from voiceprint features; Based on the boundary parameters and the prosodic embedding vector, the prosodic features of the input text are predicted by autoregression.

4. The method according to claim 3, characterized in that, Generating the natural speech based on the prosodic features includes: A fundamental frequency curve is established based on the prosodic features. The fundamental frequency curve includes a first curve and a second curve. The first curve is a quadratic function fundamental frequency curve constructed for normal syllables. The second curve is a linear function fundamental frequency curve constructed for incomplete syllables. The fundamental frequency curve difference between adjacent syllables is minimized using a splicing cost function to generate the fundamental frequency curve of the entire sentence; A Mel spectrogram is generated based on the fundamental frequency curve of the entire sentence and the input text, and the Mel spectrogram is converted into an audio signal using a vocoder.

5. The method according to claim 1, characterized in that, The method further includes: Extract audio track data from the input video; Perform speech recognition on the audio track data to obtain the speech-recognized text; By comparing the speech recognition text with the input text, a set of difference texts is generated. The set of difference texts includes difference keywords, which are keywords in the speech recognition text that are different from those in the input text. The replacement interval is determined based on the set of differing texts; A speech synthesis model is used to fuse the voiceprint features of the input video corresponding to the replacement interval with the input text to generate natural speech.

6. The method according to claim 1, characterized in that, The lip keypoints of the input video are analyzed using a lip displacement model, including: Extract face images frame by frame from the input video; Detect lip key points in the face image to obtain a key point set, wherein the number of lip key points in the key point set is greater than or equal to a preset marker point threshold. Extract the coordinates and displacements of the key points of the lips in two adjacent frames of the face image.

7. The method according to claim 6, characterized in that, Based on the aforementioned lip key points, and according to the natural speech matching lip shape change data, including: Set the mask frame of the target image based on the key points of the lips; Select a reference frame from the target image; The mask frame, the reference frame, and the noise latent variable are concatenated at the channel level to serve as the input to the image segmentation network. The voiceprint features are integrated into the image segmentation network through a cross-attention layer to generate lip shape change data based on the voiceprint features and the target image.

8. The method according to claim 1, characterized in that, Generating an output video based on the input video and the lip shape change data includes: The lip shape change data is used to replace the lip shape region in the input video to generate a replacement frame; Obtain the evaluation parameters and preset parameter thresholds of the replacement frame; If the evaluation parameter is greater than or equal to a preset parameter threshold, the output video is generated based on the replacement frame; If the evaluation parameter is less than the parameter threshold, the input video and the lip shape change data are input into the video replacement model to generate the output video through the video replacement module; the video replacement model is configured to perform lip shape replacement on the input video based on a generative adversarial network.

9. The method according to claim 8, characterized in that, The input video and the lip shape change data are input into a video replacement model to generate the output video through the video replacement module, including: Extract facial images from the input video; The face image, the lip shape change data, and the Mel spectrogram of the natural speech are used as inputs and passed to the pre-trained video replacement model. The video replacement model generates lip-sync images based on the Mel spectrogram of the natural speech. The lip-sync image is used to replace the lip-sync region of the video frame of the input video to generate the output video.

10. A video generation system based on AI voice cloning and lip-syncing, characterized in that, The system includes: The acquisition module is used to acquire input video and input text; The speech synthesis module is used to fuse the voiceprint features of the input video with the input text using a speech synthesis model to generate natural speech; the speech synthesis model is a generative language model based on autoregressive Transformer; the voiceprint features include fundamental frequency, formants and prosodic patterns; The lip-shape matching module is used to parse the lip key points of the input video using a lip-shape displacement model, and to match lip shape change data based on the natural speech according to the lip key points; the lip-shape displacement model is a latent diffusion model based on audio conditions; the lip-shape displacement model adopts an inter-frame consistency loss function and performs data matching based on an attention mechanism; The video output module is used to generate an output video based on the input video and the lip shape change data.

Citation Information

Patent Citations

  • Personalized and dynamic text-to-speech sound cloning using incompletely trained text-to-speech model

    CN117597728A

  • Three-dimensional digital human head animation generation method based on voice rhythm decomposition

    CN118015162A

  • Lip sound synchronous processing and model training method, electronic equipment and storage medium

    CN118658100A

  • Text-driven digital human high-precision sound and lip synchronization system and text-driven digital human high-precision sound and lip synchronization method

    CN119228957A

Cited By

  • Digital human speech synthesis and mouth shape synchronization method and system

    CN122067510A

  • Digital human voice synthesis and lip synchronization method and system

    CN122067510B