Method for training video speech generation model, video synthesis method and related equipment
By constructing and pre-training a model with the same structure as the audio decoder, the problems of low speech quality and insufficient adaptability in video speech synthesis are solved, and high-quality video speech synthesis that can adapt to the needs of different scenarios is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2024-07-19
- Publication Date
- 2026-05-29
AI Technical Summary
Existing video speech synthesis technologies suffer from low-quality speech generation and difficulty in adapting to the needs of different target scenarios, resulting in poor synthesis effects.
A first audio-to-audio model and a second video-to-audio model are constructed, wherein the first audio decoder and the second audio decoder have the same structure. The first model is pre-trained by collecting a large amount of mono data, the parameters of the first audio decoder are saved and used to initialize the second audio decoder, and then the second model is trained in the target scene until the preset convergence condition is met.
It improves the quality of generated speech in video speech synthesis, enabling the model to adapt to the needs of different scenarios and enhancing the synthesis effect.
Smart Images

Figure CN119028359B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of financial technology, and in particular to training methods for video speech generation models, video synthesis methods, and related equipment. Background Technology
[0002] Currently, there are some applications related to video-to-speech synthesis in the financial customer service field. Video-to-speech synthesis combines video pixels with natural language text prompts to generate audio synchronized with the video content, thus creating highly flexible audio tracks for the video, including dialogue, sound effects, music, and more. However, existing video-to-speech synthesis methods still suffer from low-quality generated speech and difficulty adapting to different target scenarios, thereby reducing the effectiveness of video-to-speech synthesis. Summary of the Invention
[0003] In view of the shortcomings of the prior art, the purpose of this invention is to provide a training method, a video synthesis method and related equipment for video speech generation models that can be applied to financial technology or other related fields. Its main purpose is to improve the quality of generated speech in video speech synthesis, thereby improving the synthesis effect.
[0004] The technical solution of the present invention is as follows:
[0005] The first aspect of this invention provides a training method for a video speech generation model, comprising:
[0006] A first audio-to-audio model and a second video-to-audio model are constructed. The first model includes an audio encoder and a first audio decoder. The second model includes a video frame encoder, an identity encoder, and a second audio decoder. The first audio decoder and the second audio decoder have the same structure.
[0007] Collect a large amount of mono data to pre-train the first model, and save the parameters of the first audio decoder when the first model completes pre-training;
[0008] The second audio decoder is initialized according to the parameters of the first audio decoder;
[0009] The video dataset collected in the target scene is input into the initialized second model for training. The training is completed when the preset convergence condition is met, and the video speech generation model is obtained.
[0010] In one embodiment, the step of collecting a large amount of mono data to pre-train the first model and saving the parameters of the first audio decoder when the first model completes pre-training includes:
[0011] Collect a large amount of mono data.
[0012] The mono data is input into the first model, the speech features are extracted by the audio encoder, and the speech features are reconstructed by the first audio decoder to generate a reconstructed audio signal.
[0013] The weights of the first model are updated by backpropagation based on the generated reconstructed audio signal until pre-training is complete;
[0014] Obtain and save the parameters of the first audio decoder in the first model that has completed pre-training.
[0015] In one embodiment, initializing the second audio decoder according to the parameters of the first audio decoder specifically includes:
[0016] The parameters of the first audio decoder are used as the initial parameters of the second audio decoder to complete the initialization of the second audio decoder.
[0017] In one embodiment, a video dataset collected in the target scene is input into an initialized second model for training until a preset convergence condition is met, thus completing the training and obtaining a video-speech generation model, including:
[0018] Acquire a video dataset collected in the target scene, and extract video frames and original audio from the video dataset;
[0019] The video frames are input into the second model, and visual features are extracted by the video encoder and facial features of the speaker are extracted by the identity encoder.
[0020] The second audio decoder predicts and generates the corresponding target audio based on the visual features and the speaker's facial features.
[0021] The weights of the second model are updated under the supervision of the loss function based on the target audio and the original audio until a preset convergence condition is met, thus obtaining the video speech generation model.
[0022] In one embodiment, the signal form of the original audio includes the original waveform and Mel spectrum.
[0023] A second aspect of the present invention provides a video synthesis method, comprising:
[0024] Obtain the silent video to be synthesized;
[0025] The silent video to be synthesized is input into a pre-trained video-speech generation model to generate the corresponding audio;
[0026] A video with sound is synthesized from the silent video and the generated audio.
[0027] The pre-trained video speech generation model is obtained using the training method described above.
[0028] A third aspect of the present invention provides a training apparatus for a video speech generation model, comprising:
[0029] A building module is used to build a first audio-to-audio model and a second video-to-audio model. The first model includes an audio encoder and a first audio decoder. The second model includes a video frame encoder, an identity encoder, and a second audio decoder. The first audio decoder and the second audio decoder have the same structure.
[0030] The first training module is used to collect a large amount of mono data to pre-train the first model and save the parameters of the first audio decoder when the first model completes the pre-training.
[0031] An initialization module is used to initialize the second audio decoder according to the parameters of the first audio decoder;
[0032] The second training module is used to input the video dataset collected in the target scene into the initialized second model for training until the preset convergence condition is met, at which point the training is completed and a video speech generation model is obtained.
[0033] A fourth aspect of the present invention provides a video synthesis apparatus, comprising:
[0034] The acquisition module is used to acquire the silent video to be synthesized;
[0035] The audio generation module is used to input the silent video to be synthesized into a pre-trained video-speech generation model to generate the corresponding audio.
[0036] A synthesis module is used to synthesize a video with sound from the silent video and the generated audio.
[0037] The pre-trained video speech generation model is obtained using the training method described above.
[0038] A fifth aspect of the present invention provides an electronic device comprising at least one processor; and,
[0039] A memory communicatively connected to the at least one processor; wherein,
[0040] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the training method of the video speech generation model or the video synthesis method described above.
[0041] A sixth aspect of the present invention provides a non-volatile computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to perform the above-described training method for the video-speech generation model or to perform the above-described video synthesis method.
[0042] Beneficial Effects: This invention discloses a training method for a video speech generation model, a video synthesis method, and related equipment. Compared to existing technologies, this invention constructs a first audio-to-audio model and a second video-to-audio model. The first audio decoder in the first model and the second audio decoder in the second model have the same structure. A large amount of mono data is collected to pre-train the first model, and the parameters of the first audio decoder are saved when the pre-training is complete. The second audio decoder is initialized based on the parameters of the first audio decoder. The video dataset collected in the target scene is input into the initialized second model for training until a preset convergence condition is met, thus completing the training and obtaining the video speech generation model. Initializing the model using a pre-trained audio decoder allows the model to retain pre-trained speech features while adapting to the characteristics of the target scene dataset, improving the quality of generated speech in video speech synthesis and thus enhancing the synthesis effect. Attached Figure Description
[0043] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0044] Figure 1 A flowchart illustrating a training method for a video speech generation model provided in an embodiment of the present invention;
[0045] Figure 2 A flowchart of step S200 in the training method of the video speech generation model provided in the embodiment of the present invention;
[0046] Figure 3 A flowchart of step S400 in the training method of the video speech generation model provided in the embodiment of the present invention;
[0047] Figure 4 A flowchart of a video synthesis method provided in an embodiment of the present invention;
[0048] Figure 5 A schematic diagram of the functional modules of the training device for the video speech generation model provided in an embodiment of the present invention;
[0049] Figure 6 This is a schematic diagram of the functional modules of the video synthesis device provided in an embodiment of the present invention;
[0050] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The embodiments of the invention are described below in conjunction with the accompanying drawings.
[0052] Currently, there are some applications related to video-to-speech synthesis in the financial customer service field. Video-to-speech synthesis combines video pixels with natural language text prompts to generate audio synchronized with the video content, thus creating highly flexible audio tracks for the video, including dialogue, sound effects, music, and more. However, existing video-to-speech synthesis methods still suffer from low-quality generated speech and difficulty adapting to different target scenarios, thereby reducing the effectiveness of video-to-speech synthesis.
[0053] To address the aforementioned problems, this invention proposes a training method for a video-speech generation model and a video synthesis method. The system architecture of the video-speech generation model training method or video synthesis method provided in this embodiment may include a first terminal device, a second terminal device, a third terminal device, a network, and a server. The network serves as the medium for providing communication links between the first terminal device, the second terminal device, the third terminal device, and the server. The network may include various connection types, such as wired and / or wireless communication links, etc.
[0054] Users can use a first terminal device, a second terminal device, and a third terminal device to interact with the server via the network to receive or send messages, etc. Various communication client applications can be installed on the first terminal device, the second terminal device, and the third terminal device, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).
[0055] The first terminal device, the second terminal device, and the third terminal device can be various electronic devices with a display screen and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0056] A server can be a server that provides various services, such as a backend management server that supports content viewed by users using a first terminal device, a second terminal device, or a third terminal device (this is just an example). The backend management server can analyze and process data such as received user requests and feed back the processing results (such as web pages, information, or data obtained or generated based on user requests) to the terminal devices. A server can also be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") in terms of management difficulty and weak business scalability. A server can also be a server for a distributed system or a server integrated with blockchain technology.
[0057] It should be noted that the training method for the video-speech generation model or the video synthesis method provided in the embodiments of the present invention can generally be executed by a first terminal device, a second terminal device, or a third terminal device. Accordingly, the training device for the video-speech generation model or the video synthesis device provided in the embodiments of the present invention can also be disposed in the first terminal device, the second terminal device, or the third terminal device.
[0058] Alternatively, the training method for the video-to-speech generation model or the video synthesis method provided in the embodiments of the present invention can generally also be executed by a server. Correspondingly, the training device for the video-to-speech generation model or the video synthesis device provided in the embodiments of the present invention can generally be located in a server. The training method for the video-to-speech generation model or the video synthesis method provided in the embodiments of the present invention can also be executed by a server or server cluster that is different from the server and capable of communicating with the first terminal device, the second terminal device, the third terminal device, and / or the server. Correspondingly, the training device for the video-to-speech generation model or the video synthesis device provided in the embodiments of the present invention can also be located in a server or server cluster that is different from the server and capable of communicating with the first terminal device, the second terminal device, the third terminal device, and / or the server.
[0059] It should be understood that the number of terminal devices, networks, and servers listed above is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be used.
[0060] Please see Figure 1 , Figure 1 This is a flowchart of one embodiment of the training method for the video speech generation model provided by the present invention. Figure 1 As shown, the method specifically includes the following steps:
[0061] S100. Construct an audio-to-audio first model and a video-to-audio second model. The first model includes an audio encoder and a first audio decoder. The second model includes a video frame encoder, an identity encoder, and a second audio decoder. The first audio decoder and the second audio decoder have the same structure.
[0062] In this embodiment, a first audio-to-audio model (A2A model) and a second video-to-audio model (V2A model) are constructed respectively. The A2A model usually refers to an audio-based artificial intelligence model that can be used for various tasks, such as speech recognition, speech synthesis, and audio classification. The V2A model is an artificial intelligence model that generates synchronous audio corresponding to the input video based on the visual features of the input video, thereby achieving the effect of adding audio signals to silent videos.
[0063] The first model in this embodiment includes an audio encoder and a first audio decoder. The structure of the audio encoder is the reverse of that of the first audio decoder, i.e., using transposed convolution instead of convolution, or using convolution instead of transposed convolution, etc. Specifically, the audio encoder structure can employ a convolutional neural network (CNN), a recurrent neural network (RNN), or a Transformer network based on a self-attention mechanism, etc. The second model includes a video frame encoder, an identity encoder, and a second audio decoder, and the structures of the first and second audio decoders are identical. In other words, the decoder structure of the second model is the same as that of the first model, while the encoder structure varies depending on the input modality. By using decoders with the same structure for audio decoding, the training and learning of the audio data by the first model can be retained for the second model, thereby improving the quality and efficiency of its generated speech.
[0064] S200: Collect a large amount of mono data to pre-train the first model, and save the parameters of the first audio decoder when the first model completes pre-training.
[0065] In this embodiment, a large amount of mono data is first collected to pre-train the first model, thereby obtaining an audio-to-audio generation model. The purpose of pre-training is to enable the first audio decoder to learn general speech features without relying on a specific video dataset. When the first model completes pre-training, the parameters of the first audio decoder are saved for parameter sharing with the second model, so that the second model can retain the learning results of the first model on speech features.
[0066] In one embodiment, such as Figure 2 As shown, the step of collecting a large amount of mono data to pre-train the first model and saving the parameters of the first audio decoder when the first model completes pre-training includes:
[0067] S201: Acquires a large amount of mono data.
[0068] S202. Input the mono data into the first model, extract speech features through the audio encoder, and reconstruct the speech features through the first audio decoder to generate a reconstructed audio signal.
[0069] S203. Update the weights of the first model by backpropagation based on the generated reconstructed audio signal until pre-training is complete;
[0070] S204. Obtain and save the parameters of the first audio decoder in the first model that has completed pre-training.
[0071] In this embodiment, a large amount of mono data is collected for training the first model. Mono data can come from various sources, such as speech, music, and ambient sound. Mono audio is the most basic audio format, containing only one channel; therefore, all audio information is mixed within this single channel. Because mono data is relatively small, it is easier to process, which also reduces model complexity and improves training efficiency to some extent. Before training, the mono data can be preprocessed, such as cleaning the data to remove noise and irrelevant parts, or standardizing the audio signal to give it a uniform format and volume, etc., to improve training effectiveness.
[0072] The mono data is then input into the first model, where an audio encoder encodes the input audio signal into a low-dimensional feature vector to extract speech features, such as Mel-frequency cepstral coefficients (MFCC) and Mel-frequency energy features (MEL). The first audio decoder then reconstructs the audio signal based on the extracted features, generating a reconstructed audio signal. Based on the reconstructed audio signal output from the first model, backpropagation is performed to update the weights of the first model under the supervision of loss functions such as mean squared error (MSE) and cross-entropy loss, until the model converges, completing pre-training (e.g., reaching a preset number of training iterations or the loss function value is less than a preset value). After pre-training, the weights and architecture of the first model are saved, and the parameters of the first decoder are extracted for subsequent training of the second model.
[0073] S300. Initialize the second audio decoder according to the parameters of the first audio decoder.
[0074] In this embodiment, based on the pre-training results of the first model, the first audio decoder has already learned general speech features without relying on a specific video dataset. The second audio decoder in the second video-to-audio model is initialized according to the parameters of the first audio decoder. Specifically, since the two audio decoders have the same structure, the parameters of the first audio decoder are directly used as the initial parameters of the second audio decoder, thus completing the initialization of the second audio decoder. This allows the second model to retain the pre-trained speech features during training, improving the video-to-speech generation effect.
[0075] S400. Input the video dataset collected in the target scene into the initialized second model for training until the preset convergence condition is met, then the training is completed and the video speech generation model is obtained.
[0076] In this embodiment, the initialized second model is fine-tuned and trained using a video dataset collected in the target scenario. The target scenario can be flexibly set according to needs, such as financial service scenarios, feedback and complaint scenarios, education and training scenarios, or interactive scenarios involving dialects, etc. Since different target scenarios may have different speech characteristics, the second model is fine-tuned and trained using a video dataset collected in the target scenario. The purpose of fine-tuning is to adapt the video-to-speech second model to the specific video dataset while retaining the pre-trained speech features, so that the video-to-speech generation can adapt to the needs of different scenarios and improve the speech synthesis effect.
[0077] In one embodiment, such as Figure 3 As shown, the step of inputting the video dataset collected in the target scene into the initialized second model for training until a preset convergence condition is met, thus completing the training and obtaining the video-speech generation model, includes:
[0078] S301. Obtain the video dataset collected in the target scene, and extract the video frames and original audio from the video dataset;
[0079] S302. Input the video frame into the second model, extract visual features through the video encoder, and extract the facial features of the speaker through the identity encoder;
[0080] S303. The second audio decoder predicts and generates the corresponding target audio based on the visual features and the speaker's facial features.
[0081] S304. Update the weights of the second model under the supervision of the loss function based on the target audio and the original audio until the preset convergence condition is met, and obtain the video speech generation model.
[0082] In this embodiment, during the fine-tuning training phase, a video dataset collected in the target scene is acquired. The video dataset contains the sound types that the second model needs to learn in the target scene. The video dataset is preprocessed, including extracting video frames and audio signals. The extracted video frames and original audio signals can be obtained. The signal form of the original audio can include the original waveform and Mel spectrum. Similarly, the audio signal form of the first model pre-training can also include the original waveform and Mel spectrum, so as to realize the generation of video speech in different signal forms to meet different needs and scenarios.
[0083] The video frames are then input into the second model, where a video encoder extracts visual features (e.g., 3D convolution and ResNet-18 can be used to extract visual features from the video frames) and an identity encoder extracts speaker facial features. This identity encoder can use a pre-trained face or speaker encoder, such as a deep neural network-based identity encoder or a 3DMM-based identity encoder. The extracted visual and speaker facial features are then fed into a second audio decoder for predictive decoding to generate the target audio corresponding to the video frames. Specifically, for the two types of audio signals, the target audio can be generated in different ways. For example, for the original waveform, a discriminator based on a generative adversarial network can be used to improve the realism of the generated waveform; for the megohmmeter spectrum, a Conformer-based method can be used, employing a pre-trained neural vocoder to generate waveforms from the megohmmeter spectrum. This allows the model to work on different target outputs, generating high-quality speech from both the original waveform and the megohmmeter spectrum, meeting different needs and scenarios while avoiding the need for an additional vocoder to convert the output format.
[0084] During training, the target audio generated by the second model through forward propagation is backpropagated under the supervision of loss functions such as mean squared error (MSE) and cross-entropy loss to update the weights of the second model until the model converges, thus completing the training, for example, by reaching a preset number of training iterations or by the loss function value being less than a preset value. After training, the second model can extract low-dimensional features from the input of a silent video and generate the corresponding speech output. The model is initialized by a pre-trained audio decoder, which allows the model to retain the pre-trained speech features while adapting to the characteristics of the target scene dataset, thereby improving the performance of the video-to-speech model, improving the quality of the generated speech in video-to-speech synthesis, and thus improving the synthesis effect.
[0085] It should be noted that there is no necessary order between the above steps. Those skilled in the art will understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0086] Another embodiment of the present invention provides a video synthesis method, such as... Figure 4 As shown, the method includes the following steps:
[0087] S500: Obtain the silent video to be synthesized;
[0088] S600. Input the silent video to be synthesized into a pre-trained video-speech generation model to generate the corresponding audio;
[0089] S700: Based on the silent video and the generated audio, a video with sound is synthesized.
[0090] The pre-trained video speech generation model is obtained using the training method described above.
[0091] In this embodiment, after the video speech generation model is trained using the training method described above, end-to-end video synthesis can be achieved in various application scenarios, such as financial service scenarios, feedback and complaint scenarios, education and training scenarios, or interactive scenarios involving dialects. The video speech generation model can generate corresponding audio based on the visual content of the video, enabling the efficient addition of detailed audio tracks, including dialogue, sound effects, and music, to silent videos, achieving the effect of audiovisual synchronization.
[0092] Specifically, a silent video to be synthesized is acquired and input into a pre-trained video-speech generation model. The model can extract visual features from the video frames, such as lip-sync features, identity features, background features, etc. Based on the extracted visual features, it decodes the audio to generate the corresponding audio. Then, based on the silent video and the generated audio, a video with sound is synthesized, achieving high-quality video-speech synthesis that adapts to different scenario requirements.
[0093] Another embodiment of the present invention provides a training device for a video speech generation model, such as... Figure 5 As shown, the training device 1 includes:
[0094] Module 11 is used to construct a first audio-to-audio model and a second video-to-audio model. The first model includes an audio encoder and a first audio decoder. The second model includes a video frame encoder, an identity encoder, and a second audio decoder. The first audio decoder and the second audio decoder have the same structure.
[0095] The first training module 12 is used to collect a large amount of mono data to pre-train the first model and save the parameters of the first audio decoder when the first model completes the pre-training.
[0096] Initialization module 13 is used to initialize the second audio decoder according to the parameters of the first audio decoder;
[0097] The second training module 14 is used to input the video dataset collected in the target scene into the initialized second model for training until the preset convergence condition is met, and then the training is completed to obtain the video speech generation model.
[0098] The module referred to in this invention is a series of computer program instruction segments that can perform specific functions. It is more suitable than a program for describing the training and execution process of a video speech generation model. For specific implementation methods of each module, please refer to the corresponding method embodiments described above, which will not be repeated here.
[0099] In one embodiment, the first training module 12 includes:
[0100] The acquisition unit is used to acquire large amounts of mono data.
[0101] The first input unit is used to input the mono data into the first model, extract speech features through the audio encoder, and reconstruct the speech features through the first audio decoder to generate a reconstructed audio signal.
[0102] A pre-training unit is used to update the weights of the first model based on the generated reconstructed audio signal through backpropagation until pre-training is complete;
[0103] The parameter storage unit is used to acquire and store the parameters of the first audio decoder in the first model that has completed pre-training.
[0104] In one embodiment, the initialization module 13 is specifically used for:
[0105] The parameters of the first audio decoder are used as the initial parameters of the second audio decoder to complete the initialization of the second audio decoder.
[0106] In one embodiment, the second training module 14 includes:
[0107] The acquisition unit is used to acquire a video dataset collected in the target scene and extract video frames and original audio from the video dataset.
[0108] The second input unit is used to input the video frame into the second model, extract visual features through the video encoder, and extract the facial features of the speaker through the identity encoder;
[0109] The prediction generation unit is used to predict and generate the corresponding target audio based on the visual features and the speaker's facial features using the second audio decoder;
[0110] The training update unit is used to update the weights of the second model under the supervision of the loss function based on the target audio and the original audio until a preset convergence condition is met, thus obtaining the video speech generation model.
[0111] In one embodiment, the signal form of the original audio includes the original waveform and Mel spectrum.
[0112] Another embodiment of the present invention provides a video synthesis apparatus, such as... Figure 6 As shown, device 2 includes:
[0113] Acquisition module 21 is used to acquire the silent video to be synthesized;
[0114] The audio generation module 22 is used to input the silent video to be synthesized into a pre-trained video speech generation model to generate corresponding audio.
[0115] Synthesis module 23 is used to synthesize a video with sound based on the silent video and the generated audio.
[0116] The pre-trained video speech generation model is obtained using the training method described above.
[0117] The module referred to in this invention is a series of computer program instruction segments that can perform specific functions. It is more suitable than a program for describing the video synthesis execution process. For specific implementation methods of each module, please refer to the corresponding method embodiments above, which will not be repeated here.
[0118] Another embodiment of the present invention provides an electronic device, such as... Figure 7 As shown, the electronic device 10 includes:
[0119] One or more processors 110 and memory 120, Figure 7 The following description uses a processor 110 as an example. The processor 110 and the memory 120 can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.
[0120] Processor 110 is used to perform various control logics of electronic device 10. It can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Furthermore, processor 110 can also be any conventional processor, microprocessor, or state machine. Processor 110 can also be implemented as a combination of computing devices, such as a combination of DSP and microprocessor, multiple microprocessors, one or more microprocessors combined with DSP and / or any other such configuration.
[0121] The memory 120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the training method of the video-speech generation model in the embodiments of the present invention. The processor 110 executes various functional applications and data processing of the electronic device 10 by running the non-volatile software programs, instructions, and units stored in the memory 120, that is, implementing the training method of the video-speech generation model or the video synthesis method in the above method embodiments.
[0122] The memory 120 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device 10. Furthermore, the memory 120 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 120 may optionally include memory remotely located relative to the processor 110, and these remote memories may be connected to the electronic device 10 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0123] One or more units are stored in memory 120. When executed by one or more processors 110, they perform the training method of the video speech generation model in any of the above method embodiments, for example, performing the following steps:
[0124] A first audio-to-audio model and a second video-to-audio model are constructed. The first model includes an audio encoder and a first audio decoder. The second model includes a video frame encoder, an identity encoder, and a second audio decoder. The first audio decoder and the second audio decoder have the same structure.
[0125] Collect a large amount of mono data to pre-train the first model, and save the parameters of the first audio decoder when the first model completes pre-training;
[0126] The second audio decoder is initialized according to the parameters of the first audio decoder;
[0127] The video dataset collected in the target scene is input into the initialized second model for training. The training is completed when the preset convergence condition is met, and the video speech generation model is obtained.
[0128] In one embodiment, the step of collecting a large amount of mono data to pre-train the first model and saving the parameters of the first audio decoder when the first model completes pre-training includes:
[0129] Collect a large amount of mono data.
[0130] The mono data is input into the first model, the speech features are extracted by the audio encoder, and the speech features are reconstructed by the first audio decoder to generate a reconstructed audio signal.
[0131] The weights of the first model are updated by backpropagation based on the generated reconstructed audio signal until pre-training is complete;
[0132] Obtain and save the parameters of the first audio decoder in the first model that has completed pre-training.
[0133] In one embodiment, initializing the second audio decoder according to the parameters of the first audio decoder specifically includes:
[0134] The parameters of the first audio decoder are used as the initial parameters of the second audio decoder to complete the initialization of the second audio decoder.
[0135] In one embodiment, a video dataset collected in the target scene is input into an initialized second model for training until a preset convergence condition is met, thus completing the training and obtaining a video-speech generation model, including:
[0136] Acquire a video dataset collected in the target scene, and extract video frames and original audio from the video dataset;
[0137] The video frames are input into the second model, and visual features are extracted by the video encoder and facial features of the speaker are extracted by the identity encoder.
[0138] The second audio decoder predicts and generates the corresponding target audio based on the visual features and the speaker's facial features.
[0139] The weights of the second model are updated under the supervision of the loss function based on the target audio and the original audio until a preset convergence condition is met, thus obtaining the video speech generation model.
[0140] In one embodiment, the signal form of the original audio includes the original waveform and Mel spectrum.
[0141] Alternatively, the video synthesis method in any of the above method embodiments can be performed, for example, by executing the following steps:
[0142] Obtain the silent video to be synthesized;
[0143] The silent video to be synthesized is input into a pre-trained video-speech generation model to generate the corresponding audio;
[0144] A video with sound is synthesized from the silent video and the generated audio.
[0145] The pre-trained video speech generation model is obtained using the training method described above.
[0146] This invention also provides a non-volatile computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by one or more processors, they perform the training method for the video-speech generation model in any of the above method embodiments, for example, performing the following steps:
[0147] A first audio-to-audio model and a second video-to-audio model are constructed. The first model includes an audio encoder and a first audio decoder. The second model includes a video frame encoder, an identity encoder, and a second audio decoder. The first audio decoder and the second audio decoder have the same structure.
[0148] Collect a large amount of mono data to pre-train the first model, and save the parameters of the first audio decoder when the first model completes pre-training;
[0149] The second audio decoder is initialized according to the parameters of the first audio decoder;
[0150] The video dataset collected in the target scene is input into the initialized second model for training. The training is completed when the preset convergence condition is met, and the video speech generation model is obtained.
[0151] In one embodiment, the step of collecting a large amount of mono data to pre-train the first model and saving the parameters of the first audio decoder when the first model completes pre-training includes:
[0152] Collect a large amount of mono data.
[0153] The mono data is input into the first model, the speech features are extracted by the audio encoder, and the speech features are reconstructed by the first audio decoder to generate a reconstructed audio signal.
[0154] The weights of the first model are updated by backpropagation based on the generated reconstructed audio signal until pre-training is complete;
[0155] Obtain and save the parameters of the first audio decoder in the first model that has completed pre-training.
[0156] In one embodiment, initializing the second audio decoder according to the parameters of the first audio decoder specifically includes:
[0157] The parameters of the first audio decoder are used as the initial parameters of the second audio decoder to complete the initialization of the second audio decoder.
[0158] In one embodiment, a video dataset collected in the target scene is input into an initialized second model for training until a preset convergence condition is met, thus completing the training and obtaining a video-speech generation model, including:
[0159] Acquire a video dataset collected in the target scene, and extract video frames and original audio from the video dataset;
[0160] The video frames are input into the second model, and visual features are extracted by the video encoder and facial features of the speaker are extracted by the identity encoder.
[0161] The second audio decoder predicts and generates the corresponding target audio based on the visual features and the speaker's facial features.
[0162] The weights of the second model are updated under the supervision of the loss function based on the target audio and the original audio until a preset convergence condition is met, thus obtaining the video speech generation model.
[0163] In one embodiment, the signal form of the original audio includes the original waveform and Mel spectrum.
[0164] Alternatively, the video synthesis method in any of the above method embodiments can be performed, for example, by executing the following steps:
[0165] Obtain the silent video to be synthesized;
[0166] The silent video to be synthesized is input into a pre-trained video-speech generation model to generate the corresponding audio;
[0167] A video with sound is synthesized from the silent video and the generated audio.
[0168] The pre-trained video speech generation model is obtained using the training method described above.
[0169] As examples, non-volatile storage media can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) as external cache memory. By way of illustration and not limitation, RAM can be obtained in many forms such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The memory components or memories disclosed in the operating environment described herein are intended to include one or more of these and / or any other suitable types of memory.
[0170] In summary, the video speech generation model training method, video synthesis method, and related equipment disclosed in this invention include the following training method: constructing a first audio-to-audio model and a second video-to-audio model, wherein the first audio decoder in the first model and the second audio decoder in the second model have the same structure; collecting a large amount of mono data to pre-train the first model, and saving the parameters of the first audio decoder when the first model completes pre-training; initializing the second audio decoder according to the parameters of the first audio decoder; inputting the video dataset collected in the target scene into the initialized second model for training until a preset convergence condition is met, thus completing the training and obtaining the video speech generation model. By initializing the model through pre-trained audio decoders, the model can retain the pre-trained speech features while adapting to the characteristics of the target scene dataset, improving the quality of generated speech in video speech synthesis, thereby improving the synthesis effect.
[0171] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The computer program can be stored in a non-volatile, computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a memory, magnetic disk, floppy disk, flash memory, optical storage, etc.
[0172] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A training method for a video speech generation model, characterized in that, include: A first audio-to-audio model and a second video-to-audio model are constructed. The first model includes an audio encoder and a first audio decoder. The second model includes a video frame encoder, an identity encoder, and a second audio decoder. The first audio decoder and the second audio decoder have the same structure. Collect a large amount of mono data to pre-train the first model, and save the parameters of the first audio decoder when the first model completes pre-training; The second audio decoder is initialized according to the parameters of the first audio decoder; The video dataset collected in the target scene is input into the initialized second model for training. The training is completed when the preset convergence condition is met, and the video speech generation model is obtained.
2. The training method for the video speech generation model according to claim 1, characterized in that, The process of collecting a large amount of mono data to pre-train the first model and saving the parameters of the first audio decoder when the first model completes pre-training includes: Collect a large amount of mono data. The mono data is input into the first model, the speech features are extracted by the audio encoder, and the speech features are reconstructed by the first audio decoder to generate a reconstructed audio signal. The weights of the first model are updated by backpropagation based on the generated reconstructed audio signal until pre-training is complete; Obtain and save the parameters of the first audio decoder in the first model that has completed pre-training.
3. The training method for the video speech generation model according to claim 1, characterized in that, The initialization of the second audio decoder based on the parameters of the first audio decoder specifically includes: The parameters of the first audio decoder are used as the initial parameters of the second audio decoder to complete the initialization of the second audio decoder.
4. The training method for the video speech generation model according to claim 1, characterized in that, The process of inputting the video dataset collected in the target scene into the initialized second model for training, until a preset convergence condition is met, completes the training and yields the video-speech generation model, including: Acquire a video dataset collected in the target scene, and extract video frames and original audio from the video dataset; The video frame is input into the second model, and visual features are extracted by the video frame encoder and facial features of the speaker are extracted by the identity encoder. The second audio decoder predicts and generates the corresponding target audio based on the visual features and the speaker's facial features. The weights of the second model are updated under the supervision of the loss function based on the target audio and the original audio until a preset convergence condition is met, thus obtaining the video speech generation model.
5. The training method for the video speech generation model according to claim 4, characterized in that, The original audio signal format includes the original waveform and Mel spectrum.
6. A video synthesis method, characterized in that, include: Obtain the silent video to be synthesized; The silent video to be synthesized is input into a pre-trained video-speech generation model to generate the corresponding audio; A video with sound is synthesized from the silent video and the generated audio. The pre-trained video speech generation model is obtained using the training method described in any one of claims 1-5.
7. A training device for a video speech generation model, characterized in that, include: A building module is used to build a first audio-to-audio model and a second video-to-audio model. The first model includes an audio encoder and a first audio decoder. The second model includes a video frame encoder, an identity encoder, and a second audio decoder. The first audio decoder and the second audio decoder have the same structure. The first training module is used to collect a large amount of mono data to pre-train the first model and save the parameters of the first audio decoder when the first model completes the pre-training. An initialization module is used to initialize the second audio decoder according to the parameters of the first audio decoder; The second training module is used to input the video dataset collected in the target scene into the initialized second model for training until the preset convergence condition is met, at which point the training is completed and a video speech generation model is obtained.
8. A video synthesis apparatus, characterized in that, include: The acquisition module is used to acquire the silent video to be synthesized; The audio generation module is used to input the silent video to be synthesized into a pre-trained video-speech generation model to generate the corresponding audio. A synthesis module is used to synthesize a video with sound from the silent video and the generated audio. The pre-trained video speech generation model is obtained using the training method described in any one of claims 1-5.
9. An electronic device, characterized in that, Includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the training method of the video speech generation model according to any one of claims 1-5, or to perform the video synthesis method according to claim 6.
10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the training method of the video speech generation model according to any one of claims 1-5, or to perform the video synthesis method according to claim 6.