Voice-driven video synthesis method and device, and storage medium
By decoupling face standardization and lip-sync synthesis models, and combining them with detail processing models, the problem of mismatch between lip-sync and audio in video synthesis was solved, achieving high synchronization and high quality target speaking video generation.
Patent Information
- Application Number
- CN202511495294.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, the lip movements of the characters do not match the new audio during the video synthesis process, which can easily lead to lip-syncing errors.
A face standardization model is used to decouple identity features and lip movements to generate a closed-mouth face region image. Combined with a lip synthesis model, a speaking face region image is generated based on audio data. The image resolution and details are improved through a detail processing model. Finally, the images are fused to generate the target speaking video.
It significantly improves the synchronization and quality of lip movements and audio data in target speaking videos, reduces the sense of disjointed facial expressions, and enhances the realism and resolution of oral cavity details.
Smart Images

Figure CN121037652A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision and artificial intelligence, in particular to a speech-driven video synthesis method, device and storage medium. BACKGROUND
[0002] With the development of artificial intelligence technology, the synthesis of video has also been widely applied. For a speaking video, the person in the speaking video speaks the original audio, and the person's lip shape is also the lip shape corresponding to the original audio. Based on the speaking video, a person speaking new audio can be synthesized using artificial intelligence technology, and the lip shape of the person in the synthesized video becomes the lip shape corresponding to the new audio.
[0003] However, in the related art, in the synthesized video, the problem of mismatch between the person's lip shape and the new audio, i.e., the problem of lip shape error, easily occurs. SUMMARY
[0004] The present application aims to solve the above technical problems in the related art by providing a speech-driven video synthesis method, device and storage medium.
[0005] To achieve the above-mentioned purposes, the technical solutions adopted by the embodiments of the present application are as follows: In a first aspect, the embodiments of the present application provide a speech-driven video synthesis method, comprising: According to the multiple frames of original images in the original speaking video, multiple frames of face region images and corresponding face identity feature images are determined respectively; Using a face standardization model, the identity features and lip shape actions of the corresponding face region images are decoupled according to each frame of face identity feature image, and the corresponding closed-mouth face region image is output; Using a lip synthesis model, multiple frames of speaker face region images are generated according to the multiple frames of closed-mouth face region images and audio data; wherein the multiple frames of speaker face region images have lip shape actions and corresponding face actions matched with the audio data; The multiple frames of speaker face region images and the multiple frames of original images are fused to obtain a target speaking video matched with the audio data.
[0006] Optionally, the fusion processing of the multiple frames of speaker face region images and the multiple frames of original images to obtain the target speaking video matched with the audio data comprises: Using a detail processing model, the internal details of the mouth in each frame of speaker face region image of the multiple frames of speaker face region images are repaired to obtain multiple frames of high-definition speaker face region images; According to the multi-frame high-definition speaker face region image and the multi-frame original image, a fusion process is performed to obtain the target speaker video.
[0007] Optionally, the fusion process according to the multi-frame high-definition speaker face region image and the multi-frame original image to obtain the target speaker video comprises: A mask image with the same size as each frame of the high-definition speaker face region image is generated, wherein the mask image is a mask image with a Gaussian edge; The multi-frame high-definition speaker face region image is fused into the face region of the multi-frame original image through the mask image to obtain the target speaker video.
[0008] Optionally, the multi-frame face region image and the corresponding face identity feature image are determined according to the multi-frame original image in the original speaker video, comprising: Face detection is performed on the multi-frame original image to crop the multi-frame face region image; According to the preset part of the multi-frame face region image, the multi-frame face region image is cropped to obtain the corresponding face identity feature image.
[0009] Optionally, the face standardization model is obtained by training the following steps: According to the multi-frame sample image in the sample speaker video, a multi-frame sample face region image, a corresponding real face identity feature image, and a real closed-mouth face region image are determined; According to the multi-frame sample face region image, the corresponding real face identity feature image, and the real closed-mouth face region image, an initial face standardization model is trained to obtain the face standardization model.
[0010] Optionally, the face standardization model is obtained by training the following steps: Using the initial face standardization model, a multi-frame predicted closed-mouth face region image is output according to the multi-frame sample face region image and the corresponding sample face identity feature image; According to the multi-frame predicted closed-mouth face region image and the real closed-mouth face region image, a value of a first loss function is calculated; According to the value of the first loss function, the model parameters of the initial face standardization model are updated until a training termination condition is met to obtain the face standardization model.
[0011] Optionally, the lip shape synthesis model is obtained by training the following steps: The face standardization model is used to perform decoupling processing of identity features and mouth shape actions on the corresponding sample face region image according to each frame of sample face identity feature image, to obtain a corresponding sample mouth-closed face region image. The initial mouth shape synthesis model is used to output a plurality of frames of predicted speaker face region images according to the plurality of frames of sample mouth-closed face region images and sample audio data corresponding to the sample speaking video. A value of a second loss function is calculated according to the plurality of frames of predicted speaker face region images and a plurality of frames of real speaker face region images in the sample speaking video. Model parameters of the initial mouth shape synthesis model are updated according to the value of the second loss function until a training termination condition is met, to obtain the mouth shape synthesis model.
[0012] Optionally, the detail processing model is obtained by training using the following steps. The mouth shape synthesis model is used to output a plurality of frames of sample speaking face region images according to the plurality of frames of sample mouth-closed face region images and the sample audio data. An initial detail processing model is used to repair internal details of the mouth in the plurality of frames of sample speaking face region images, to obtain a plurality of frames of predicted high-definition speaker face region images. A value of a third loss function is calculated according to the plurality of frames of predicted high-definition speaker face region images and the plurality of frames of real speaker face region images. Model parameters of the initial detail processing model are updated according to the value of the third loss function until a training termination condition is met, to obtain the detail processing model.
[0013] In a second aspect, the embodiments of the present application further provide a speech-driven video synthesis device, including a memory and a processor, the memory stores a computer program executable by the processor, and the processor implements the speech-driven video synthesis method in any one of the first aspect when executing the computer program.
[0014] In a third aspect, the embodiments of the present application further provide a computer readable storage medium, the storage medium stores a computer program, and the computer program is read and executed to implement the speech-driven video synthesis method in any one of the first aspect.
[0015] The beneficial effects of the present application are: the embodiment of the present application provides a speech-driven video synthesis method, which comprises the following steps: determining a plurality of face region images and corresponding face identity feature images according to a plurality of original images in an original speaking video; using a face standardization model to perform identity feature and lip movement decoupling processing on the corresponding face region image according to each face identity feature image, and outputting a corresponding closed-mouth face region image; using a lip synthesis model to generate a plurality of speaker face region images according to the plurality of closed-mouth face region images and audio data; wherein the plurality of speaker face region images have lip movements and corresponding facial movements matched with the audio data; and performing fusion processing on the plurality of speaker face region images and the plurality of original images to obtain a target speaking video matched with the audio data. The face standardization model is used to output the closed-mouth face region image, which can eliminate the interference of the original lip movement, the lip synthesis model is used to generate the plurality of speaker face region images, the synchronization quality of the speaking lip movement and the audio data is higher, and then the synchronization of the lip movement and the audio data in the target speaking video is significantly improved. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0017] Figure 1 Flowchart of a speech-driven video synthesis method provided by an embodiment of the present application Figure One ; Figure 2 Flowchart of a speech-driven video synthesis method provided by an embodiment of the present application Figure Two ; Figure 3 Flowchart of a speech-driven video synthesis method provided by an embodiment of the present application Figure Three ; Figure 4 Flowchart of a speech-driven video synthesis method provided by an embodiment of the present application Figure Four ; Figure 5 Flowchart of a speech-driven video synthesis method provided by an embodiment of the present application Figure Five ; Figure 6 Flowchart of a speech-driven video synthesis method provided by an embodiment of the present application Figure Six ; Figure 7 Flowchart of a speech-driven video synthesis method provided by an embodiment of the present applicationFigure Seven Figure 8 A flowchart of a voice-driven video synthesis method provided for an embodiment of the present application Figure Eight Figure 9 A structural diagram of a voice-driven video synthesis device provided for an embodiment of the present application Figure 10 A structural diagram of a voice-driven video synthesis device provided for an embodiment of the present application DETAILED DESCRIPTION
[0018] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the accompanying drawings for the embodiments of the present application to make a clear and complete description of the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application.
[0019] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present application are within the scope of protection of the present application.
[0020] In the description of the present application, it should be noted that if the terms "upper", "lower", etc. indicate the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present application is usually placed, which is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.
[0021] In addition, the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0022] It should be noted that the features in the embodiments of the present application can be combined with each other without conflict.
[0023] This application provides a voice-driven video synthesis method, which is applied to a voice-driven video synthesis device. The voice-driven video synthesis device can be a terminal device or a server. The terminal device can be any of the following: computer equipment, laptop computer, tablet computer, smartphone, etc.
[0024] The following explains a voice-driven video synthesis method provided in the embodiments of this application.
[0025] Figure 1 A flowchart illustrating a voice-driven video synthesis method provided in this application embodiment. Figure One ,like Figure 1 As shown, the method may include: S101. Based on the multiple original images in the original speaking video, determine the multiple face region images and the corresponding facial identity feature images.
[0026] In this context, the face region image refers to the image of the area containing the face, while the facial identity feature image is a part of the face region image and can provide facial pose information. Since the person in the multiple original images is speaking the original audio, the mouth shape of the speaking face changes in the multiple face region images, and the mouth is in an open state.
[0027] In some implementations, multiple frames of facial region images are determined based on multiple original images in the original speaking video; then, multiple frames of facial identity feature images are determined based on the multiple frames of facial region images.
[0028] It should be noted that one original image frame corresponds to one face region image, which in turn corresponds to one facial identity feature image. The resolution of multiple original images can be 256x256.
[0029] S102. Using a face standardization model, based on the facial identity feature image of each frame, the corresponding face region image is decoupled from the identity features and mouth movements, and the corresponding closed-mouth face region image is output.
[0030] In the image of a closed-mouth face region, the mouth is in a closed state. The closed-mouth face region image is used to represent a reference face.
[0031] In one possible implementation, each frame of facial identity feature image is input into a face standardization model along with the corresponding face region image. The face standardization model is then used to decouple the identity features from the lip movements, and the closed-mouth face region image corresponding to each frame of facial identity feature image is output.
[0032] In the embodiments of the present application, the face standardization model can be referred to as a first stage model. The face standardization model is used to close the mouth in the face region image, to obtain a closed-mouth face region image, to eliminate the influence of different mouth shapes on subsequent processing, and to greatly improve the synchronization.
[0033] S103, using a mouth shape synthesis model, generating a plurality of speaker face region images according to the plurality of closed-mouth face region images and the audio data.
[0034] The plurality of speaker face region images have mouth shape movements and corresponding facial movements that match the audio data. The speaker face region images have higher lip synchronization quality.
[0035] In some embodiments, the audio data is a speech waveform file used to drive the mouth shape synthesis. The audio data is converted into a log mel spectrum based on a preset sampling frequency, where the log mel spectrum is a more effective feature for generating mouth shapes. Each closed-mouth face region image, each facial identity feature image, and the corresponding log mel spectrum image segment are input into the mouth shape synthesis model, and the mouth shape synthesis model outputs a plurality of speaker face region images. For example, the preset sampling frequency can be 16 kHz.
[0036] It should be noted that the mouth shape synthesis model can be referred to as a second stage model. The mouth shape synthesis model is used to generate precise synchronization of mouth shapes and corresponding facial movements based on the standardized closed-mouth face region image and the audio data. For example, the facial movement can be a chin movement.
[0037] In actual applications, the resolution of the closed-mouth face region image is 256x256, and the resolution of the plurality of speaker face region images is also 256x256.
[0038] In the embodiments of the present application, the face standardization model is used to standardize any mouth shape face region image into a "reference face" with the same identity features but a naturally closed mouth, i.e., a closed-mouth face region image. This provides a clean and consistent starting point for the mouth shape generation of the mouth shape synthesis model, fundamentally avoids the interference of the original video mouth shape on the newly generated mouth shape, and greatly improves the synchronization.
[0039] S104, performing fusion processing on the plurality of speaker face region images and the plurality of original images to obtain a target speaker video that matches the audio data.
[0040] In the embodiment of the present application, the multi-frame speaker face region images are fused into the corresponding multi-frame original images, that is, the face region images in the frame original images are replaced by the speaker face region images, and the background in the frame original video is not changed. The faces in the target speaker video and the original speaker video are the faces of the same person, but the words spoken in the target speaker video and the original speaker video are different, and the mouth shapes of the speakers are different.
[0041] The mouth shape action of the speaker in the target speaker video is accurately synchronized with the audio data.
[0042] In summary, the embodiment of the present application provides a speech-driven video synthesis method, which comprises: determining a plurality of face region images and corresponding face identity feature images according to a plurality of original images in an original speaker video; using a face standardization model, performing identity feature and mouth shape action decoupling processing on the corresponding face region image according to each face identity feature image, and outputting a corresponding closed-mouth face region image; using a mouth shape synthesis model, generating a plurality of speaker face region images according to the plurality of closed-mouth face region images and audio data; wherein the plurality of speaker face region images have mouth shape actions and corresponding facial actions matched with the audio data; and performing fusion processing on the plurality of speaker face region images and the plurality of original images to obtain a target speaker video matched with the audio data. The face standardization model is used to output the closed-mouth face region image, which can eliminate the interference of the original mouth shape. The mouth shape synthesis model is used to generate the plurality of speaker face region images, and the synchronization quality of the mouth shape and the audio data is higher, thereby significantly improving the synchronization of the mouth shape and the audio data in the target speaker video.
[0043] Moreover, the mouth shape synthesis model takes into account the lip movement and facial action to generate facial dynamics that are more in line with physiological laws and reduce the feeling of fragmentation of expressions in the target speaker video.
[0044] In actual application, a user can select an original speaker video and audio data on a terminal device, and then a target speaker video can be automatically generated.
[0045] Figure 2 A flowchart of a speech-driven video synthesis method provided by the embodiment of the present application Figure Two As shown in Figure 2 The process of obtaining the target speaker video matched with the audio data by performing fusion processing on the plurality of speaker face region images and the plurality of original images in S104 can comprise: S201, using a detail processing model to repair internal details of the mouth in each frame of the plurality of speaker face region images to obtain a plurality of high-definition speaker face region images.
[0046] The detail processing model can be referred to as a third stage model. The high-definition speaker face region image is a high-resolution and high-detail speaker face image.
[0047] In some embodiments, the multi-frame speaker face region image is input into the detail processing model, the detail processing model performs super-resolution repair on internal detail parts such as teeth and tongue in each frame of the speaker face region image, improves the resolution of each frame of the speaker face region image, and outputs multi-frame high-definition speaker face region images.
[0048] It should be noted that the resolution of each frame of the speaker face region image can be 256x256, and the resolution of the high-definition speaker face region image can be 512x512.
[0049] In actual applications, the detail processing model outputs multi-frame high-definition speaker face region images, improves the resolution, and focuses on repairing and generating teeth, tongue, and other high-definition details, which can solve the problems of low resolution and missing oral cavity details in related technologies. Among them, the multi-frame high-definition speaker face region image optimizes the oral cavity details and generates a high-fidelity face image, which surpasses the prior art.
[0050] S202, performing fusion processing on the multi-frame high-definition speaker face region image and the multi-frame original image to obtain a target speech video.
[0051] In the embodiments of the present application, the multi-frame high-definition speaker face region image is fused into the corresponding multi-frame original image, so that the lip movement and audio data of the speaker in the target speech video are accurately synchronized, the facial dynamics conforms to the physiological law, and has high-definition details such as teeth and tongue.
[0052] Optionally, Figure 3 The flowchart of a voice-driven video synthesis method provided in the embodiments of the present application Figure Three As shown in the flowchart of the voice-driven video synthesis method provided in the embodiments of the present application Figure 3 As shown, the fusion processing on the multi-frame high-definition speaker face region image and the multi-frame original image to obtain a target speech video includes: S301, generating a mask image with the same size as each frame of the high-definition speaker face region image.
[0053] The mask image is a mask image with a Gaussian edge. The mask image is a rectangular image. Inside the rectangular region of the mask image, the function value is 1, and near the rectangular boundary, the function value smoothly transitions from 1 to 0.
[0054] S302, fusing the multi-frame high-definition speaker face region image into the face region of the multi-frame original image through the mask image to obtain a target speech video.
[0055] It should be noted that the high-definition speaker face region image can be naturally fused into the face region of the multi-frame original image through the mask image with a Gaussian edge, avoiding the problem of unnatural cutting and transition of the non-face region in the high-definition speaker face region image and the multi-frame original image, and making the quality of the target speaker video higher.
[0056] Optionally, Figure 4 A flowchart of a speech-driven video synthesis method provided in an embodiment of the present application is shown in Figure Four As shown in Figure 4 The process of determining the multi-frame face region image and the corresponding face identity feature image according to the multi-frame original image in the original speaker video in S101 can include the following steps. S401, face detection is performed on the multi-frame original image, and the multi-frame face region image is cropped.
[0057] In some embodiments, face detection is performed on the multi-frame original image frame by frame, the face key points in each frame of the original image are located, and each frame of the face region image of a preset size is cropped with the tip of the nose in each frame of the original image as the center. The preset size can be 512x512x3.
[0058] S402, according to the preset part of the multi-frame face region image, the multi-frame face region image is cropped to obtain the corresponding face identity feature image.
[0059] The preset part can be the tip of the nose.
[0060] In the embodiment of the present application, the tip of the nose of each frame of the face region image is determined, and the area above the tip of the nose of each frame of the face region image is cropped, that is, the upper half of each frame of the face region image is cropped to obtain each frame of the face identity feature image.
[0061] Optionally, Figure 5 A flowchart of a speech-driven video synthesis method provided in an embodiment of the present application is shown in Figure Five As shown in Figure 5 The face standardization model is trained by the following steps. S501, according to the multi-frame sample image in the sample speaker video, the multi-frame sample face region image, the corresponding real face identity feature image, and the real closed-mouth face region image are determined.
[0062] In some embodiments, face detection is performed on the plurality of sample image frames, and face key points in each sample image frame are located. A nose tip in each sample image frame is taken as a center to crop a sample face region image of a preset size. A nose tip part of each sample face region image is determined, and an area above the nose tip of each sample face region image is cropped, that is, a top half of each sample face region image is cropped to obtain a real face identity feature image.
[0063] In addition, a frame in which a mouth of a person is closed in the plurality of sample face region images is selected as a real closed-mouth face region image.
[0064] S502, training an initial face normalization model according to the plurality of sample face region images, the corresponding real face identity feature images, and the real closed-mouth face region image, to obtain a face normalization model.
[0065] The structure of the initial face normalization model and the face normalization model is a conditional U-Net structure.
[0066] In the embodiments of the present application, the sample speaking video is a video in a large-scale, high-quality multi-language single-person speaker video dataset. The dataset covers multiple lip shapes, expressions, and illumination conditions, and the voice is clean and the sound and picture are strictly synchronized. The number of sample speaking videos in the sample dataset can be greater than 10,000, and the length of each sample speaking video is greater than 10 seconds.
[0067] Optionally, Figure 6 A flowchart of a voice-driven video synthesis method provided by the embodiments of the present application is shown in Figure Six As shown in Figure 6 The process of training the initial face normalization model according to the plurality of sample face region images, the corresponding sample face identity feature images, and the real closed-mouth face region image in S502 can include: S601, using the initial face normalization model, outputting a plurality of predicted closed-mouth face region images according to the plurality of sample face region images and the corresponding sample face identity feature images.
[0068] The encoder of the initial face normalization model can be used to simultaneously receive the sample face region images and the corresponding sample face identity feature images. The predicted closed-mouth face region image is predicted by the initial face normalization model, that is, the untrained face normalization model.
[0069] S602, calculating a value of a first loss function according to the plurality of predicted closed-mouth face region images and the real closed-mouth face region image.
[0070] S603, update the model parameters of the initial face normalization model according to the value of the first loss function until a training termination condition is met, and obtain a face normalization model.
[0071] The training termination condition can be that the number of iterations of the model parameters is greater than a preset number, or the value of the first loss function converges. Of course, the training termination condition can also be other conditions, and embodiments of the present application do not make specific limitations thereto.
[0072] It should be noted that the first loss function ensures the similarity of the image generated by the face normalization model and the real closed-mouth image at the pixel level. The first loss function includes LPIPS loss and L1 loss. The LPIPS (Learned Perceptual Image Patch Similarity) loss ensures that the generated image is similar to the real image in perceptual features, so that the result is more consistent with human visual perception.
[0073] In the embodiments of the present application, the face normalization model is first independently trained, and after the training process, the model parameters of the face normalization model are fixed.
[0074] Optionally, Figure 7 The flowchart of a voice-driven video synthesis method provided in the embodiments of the present application is shown in Figure Seven As shown in Figure 7 The lip synthesis model is trained by the following steps: S701, using the face normalization model, decoupling the identity feature and the lip movement of the corresponding sample face region image according to each frame of sample face identity feature image, to obtain the corresponding sample closed-mouth face region image.
[0075] In the embodiments of the present application, after the face normalization model is trained, the model parameters of the face normalization model are fixed, and the sample closed-mouth face region image is predicted by the trained face normalization model.
[0076] S702, using the initial lip synthesis model, outputting a plurality of predicted speaker face region images according to the plurality of sample closed-mouth face region images and the sample audio data corresponding to the sample speaking video.
[0077] The sample audio data is a sample mel-frequency spectrum graph.
[0078] In some embodiments, using the initial lip synthesis model, a plurality of predicted speaker face region images are output according to the plurality of sample closed-mouth face region images, the plurality of sample face identity feature images, and the sample mel-frequency spectrum graph.
[0079] In actual application, the structure of the initial lip-synching model is a conditional U-Net structure, and the encoder of the conditional U-Net structure receives the multi-frame sample closed-mouth face region image, the multi-frame sample face identity feature image, and the sample mel-frequency spectrum image.
[0080] S703, calculating a value of a second loss function according to the multi-frame predicted speaker face region image and the multi-frame real speaker face region image in the sample speaker video.
[0081] S704, updating the model parameters of the initial lip-synching model according to the value of the second loss function until a training termination condition is met, and obtaining a lip-synching model.
[0082] In the embodiments of the present application, the second loss function includes L1 + LPIPS loss and GAN loss, the L1 + LPIPS loss restricts the accuracy and perceptual reality of the generated lip-synching, and the GAN (Generative Adversarial Network) loss introduces a discriminator network for distinguishing the generated image and the real speaker image. This can greatly improve the realism and details of the generated lip-synching and the surrounding face region, and avoid blurring.
[0083] It should be noted that the lip-synching model is independently trained, and the model parameters of the lip-synching model are fixed after the training process.
[0084] Optionally, Figure 8 A flowchart of a speech-driven video synthesis method provided in the embodiments of the present application is shown in Figure Eight As shown in Figure 8 The detail processing model is trained by the following steps: S801, using the lip-synching model to output multi-frame sample speaker face region images according to the multi-frame sample closed-mouth face region image and the sample audio data.
[0085] After the training of the lip-synching model, the model parameters of the lip-synching model are fixed, and the sample speaker face region image is obtained by the trained lip-synching model.
[0086] S802, using the initial detail processing model to repair the internal details of the mouth in the multi-frame sample speaker face region image to obtain multi-frame predicted high-definition speaker face region images.
[0087] The details include tongue, teeth, etc.
[0088] In addition, the initial detail processing model is a super-resolution restoration model, and the structure of the initial detail processing model is a U-Net structure super-resolution network.
[0089] S803, calculate a value of a third loss function according to the multi-frame predicted high-definition speaker face region image and the multi-frame real speaker face region image.
[0090] S804, update model parameters of the initial detail processing model according to the value of the third loss function until a training termination condition is met, and obtain a detail processing model.
[0091] In the embodiment of the present application, the third loss function includes L1 + LPIPS loss and GAN loss. The L1 + LPIPS loss constrains pixel and perceptual similarity in a high-resolution space. The GAN loss trains a discriminator focusing on high-frequency details such as tooth texture and lip gloss, so that the detail processing model finally generates an image with photo-level realism.
[0092] In actual application, after the detail processing model is trained, the model parameters of the detail processing model are fixed.
[0093] It should be noted that in the embodiment of the present application, the three models of the face standardization model, the lip synthesis model and the detail processing model are cascaded. This cascaded architecture supports independent optimization at each stage, improves robustness and scalability. In the embodiment of the present application, the face standardization model corresponds to identity standardization, the lip synthesis model corresponds to audio-driven lip synthesis, and the detail processing model corresponds to detail enhancement and super-resolution.
[0094] In summary, the embodiment of the present application provides a speech-driven video synthesis method. In the first stage, the face standardization model closes the mouth of the original speaker video, eliminating the influence of different lip shapes on the generated results, greatly improving the synchronization, and decoupling the identity features and the original lip movements. In the second stage, the lip synthesis model makes the closed-mouth video speak, generating a speaker with higher lip-synchronized quality. In the third stage, the detail processing model improves the resolution and optimizes the details of the oral cavity, especially the realism of the teeth and oral cavity, generating a realistic high-resolution face. It ensures the robustness of lip synchronization, the improvement of visual quality and the scalability on diversified data sets.
[0095] Among them, by introducing the standardization process of "closing the mouth first and then opening the mouth", the interference of the original speaking video lip shape on the generated target speaking video is eliminated, the synchronization accuracy of the speech and the generated lip shape is fundamentally improved, and the fast speech speed and complex accent are better handled. The linkage of lip movement and related areas such as lower jaw and cheek is considered, and more physiological facial dynamics are generated to reduce the feeling of expression fragmentation. Through a special detail processing model, the resolution of the generated speaker face region image is significantly improved, and the clarity and realism of the teeth, tongue and other oral cavity details are highlighted. The output of the high-resolution speaker face region image reduces the blurring problem caused by resolution mismatch, and realizes smoother and seamless boundary fusion through high-quality generated results.
[0096] Moreover, make it for three stages, each stage can be optimized or replaced independently, for future for specific application scenarios (such as mobile terminal) model compression and acceleration provides flexibility.
[0097] The following describes a speech-driven video synthesis device, equipment, and storage medium for performing the speech-driven video synthesis method provided by the present application. For specific implementation processes and technical effects, refer to the related content of the speech-driven video synthesis method described above. The following will not be described again.
[0098] Figure 9 A structural diagram of a speech-driven video synthesis device provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the device includes: Figure 9 A determination module 101 is configured to determine a plurality of face region images and corresponding face identity feature images respectively according to a plurality of original images in an original speaking video. A decoupling processing module 102 is configured to perform identity feature and lip movement decoupling processing on a corresponding face region image according to each face identity feature image by using a face standardization model, and output a corresponding closed-mouth face region image. A generation module 103 is configured to generate a plurality of speaker face region images according to a plurality of closed-mouth face region images and audio data by using a lip synthesis model, wherein the plurality of speaker face region images have lip movements and corresponding face movements matched with the audio data. A fusion processing module 104 is configured to perform fusion processing on the plurality of speaker face region images and the plurality of original images to obtain a target speaking video matched with the audio data.
[0099] Optionally, the fusion processing module 104 is specifically configured to repair internal details of the mouth in each frame of the plurality of speaker face region images by using a detail processing model to obtain a plurality of high-definition speaker face region images, and perform fusion processing on the plurality of high-definition speaker face region images and the plurality of original images to obtain the target speaking video.
[0100] Optionally, the fusion processing module 104 is specifically configured to generate a mask image with the same size as each high-definition speaker face region image, wherein the mask image is a mask image with a Gaussian edge; and fuse the plurality of high-definition speaker face region images into the face region of the plurality of original images by using the mask image to obtain the target speaking video.
[0101] Optionally, the determining module 101 is specifically configured to perform face detection on the plurality of original images to crop the plurality of face region images; and crop the plurality of face region images according to preset parts of the plurality of face region images to obtain the corresponding face identity feature images.
[0102] Optionally, the face normalization model is obtained by the following steps: A first training module is configured to determine a plurality of sample face region images, corresponding real face identity feature images and real closed-mouth face region images according to a plurality of sample images in a sample speaking video; and train an initial face normalization model according to the plurality of sample face region images, the corresponding real face identity feature images and the real closed-mouth face region images to obtain the face normalization model.
[0103] Optionally, the first training module is specifically configured to use the initial face normalization model to output a plurality of predicted closed-mouth face region images according to the plurality of sample face region images and the corresponding sample face identity feature images; calculate a value of a first loss function according to the plurality of predicted closed-mouth face region images and the real closed-mouth face region images; update model parameters of the initial face normalization model according to the value of the first loss function until a training termination condition is met to obtain the face normalization model.
[0104] Optionally, the lip synthesis model is obtained by the following steps: A second training module is configured to use the face normalization model to perform identity feature and lip movement decoupling processing on a corresponding sample face region image according to each sample face identity feature image to obtain a corresponding sample closed-mouth face region image; use an initial lip synthesis model to output a plurality of predicted speaker face region images according to a plurality of sample closed-mouth face region images and sample audio data corresponding to the sample speaking video; calculate a value of a second loss function according to the plurality of predicted speaker face region images and a plurality of real speaker face region images in the sample speaking video; update model parameters of the initial lip synthesis model according to the value of the second loss function until a training termination condition is met to obtain the lip synthesis model.
[0105] Optionally, the detail processing model is obtained by the following steps: The third training module is configured to output a plurality of frames of sample speaker face region images based on the plurality of frames of sample closed-mouth face region images and the sample audio data by using the mouth shape synthesis model; repair internal details of the mouth in the plurality of frames of sample speaker face region images by using an initial detail processing model to obtain a plurality of frames of predicted high-definition speaker face region images; calculate a value of a third loss function based on the plurality of frames of predicted high-definition speaker face region images and the plurality of frames of real speaker face region images; and update model parameters of the initial detail processing model based on the value of the third loss function until a training termination condition is met, to obtain the detail processing model.
[0106] The apparatus is configured to perform the method provided by the foregoing embodiments, and has similar implementation principles and technical effects, which will not be described here.
[0107] The modules can be one or more integrated circuits configured to implement the above method, for example, one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or the like. For another example, when a certain module is implemented in the form of a processor scheduling code, the processor can be a general-purpose processor, for example, a central processing unit (CPU) or other processor capable of scheduling code. For another example, the modules can be integrated together to be implemented in the form of a system on a chip (SOC).
[0108] Figure 10 A structural schematic diagram of a speech-driven video synthesis device provided by an embodiment of the present application is shown in FIG. 1. The speech-driven video synthesis device includes a processor 201 and a memory 202. Figure 10
[0109] The memory 202 is configured to store a program, and the processor 201 invokes the program stored in the memory 202 to execute the above method embodiments. The specific implementation manners and technical effects are similar, which will not be described here.
[0110] Optionally, the present application further provides a program product, for example, a computer readable storage medium, including a program configured to perform the above method embodiments when executed by a processor.
[0111] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the described apparatus embodiments are merely schematic. The division of the units is merely logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0112] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0113] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional units.
[0114] The integrated unit realized in the form of software functional units can be stored in a computer readable storage medium. The software functional units stored in the storage medium include a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (English: Read-Only Memory, abbreviated as: ROM), a random access memory (English: Random Access Memory, abbreviated as: RAM), a magnetic disk or an optical disk, and various program code storage media.
[0115] The above is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A voice-driven video synthesis method, characterized in that, include: Based on multiple frames of original images from the original speaking video, multiple frames of face region images and corresponding facial identity feature images are determined respectively. Using a face standardization model, based on the facial identity feature image of each frame, the corresponding face region image is decoupled from the identity features and mouth movements, and the corresponding closed-mouth face region image is output. A lip-sync model is used to generate multiple frames of speaking face region images based on multiple frames of closed-mouth face region images and audio data; wherein, the multiple frames of speaking face region images have lip movements and corresponding facial movements that match the audio data. The multi-frame speaking face region image and the multi-frame original image are fused to obtain the target speaking video that matches the audio data.
2. The method according to claim 1, characterized in that, The step of fusing the multi-frame speaking face region images and the multi-frame original images to obtain the target speaking video matching the audio data includes: A detail processing model is used to repair the details inside the mouth in each frame of the speaking face region image, resulting in multi-frame high-definition speaking face region images. The target speaking video is obtained by fusing the multi-frame high-definition speaking face region images and the multi-frame original images.
3. The method according to claim 2, characterized in that, The step of fusing the multi-frame high-definition speaking face region images and the multi-frame original images to obtain the target speaking video includes: Generate a mask image with the same size as the high-definition speaking face region image of each frame, wherein the mask image is a mask image with Gaussian edges; The target speaking video is obtained by fusing the multi-frame high-definition speaking face region images into the face region of the multi-frame original images using the masked image.
4. The method according to claim 1, characterized in that, The step of determining multiple face region images and corresponding facial identity feature images based on multiple original images in the original speaking video includes: Face detection is performed on the multiple original images, and the face region images of the multiple frames are cropped out; Based on preset locations in the multi-frame facial region images, the multi-frame facial region images are cropped to obtain the corresponding facial identity feature images.
5. The method according to claim 2, characterized in that, The face standardization model is trained using the following steps: Based on multiple sample images in the sample speaking video, the multi-frame sample face region image, the corresponding real facial identity feature image, and the real closed-mouth face region image are determined respectively. Based on the multi-frame sample face region images, the corresponding real facial identity feature images, and the real closed-mouth face region images, the initial face standardization model is trained to obtain the face standardization model.
6. The method according to claim 5, characterized in that, The step of training an initial face standardization model based on the multi-frame sample face region images, the corresponding sample facial identity feature images, and the real closed-mouth face region images to obtain the face standardization model includes: Using the initial face normalization model, based on the multi-frame sample face region images and the corresponding sample facial identity feature images, output multi-frame predicted closed-mouth face region images; The value of the first loss function is calculated based on the multi-frame predicted closed-mouth face region image and the real closed-mouth face region image; Based on the value of the first loss function, the model parameters of the initial face normalization model are updated until the training termination condition is met, thus obtaining the face normalization model.
7. The method according to claim 5, characterized in that, The lip shape synthesis model was trained using the following steps: Using the aforementioned face standardization model, based on the facial identity feature image of each frame, the corresponding sample face region image is decoupled from the identity features and mouth movements to obtain the corresponding sample closed-mouth face region image. An initial lip-sync synthesis model is used to output multiple frames of predicted speaking face region images based on multiple frames of sample closed-mouth face region images and sample audio data corresponding to the sample speaking video. The value of the second loss function is calculated based on the predicted speaking face region images in the multiple frames and the real speaking face region images in the sample speaking video in the multiple frames. Based on the value of the second loss function, update the model parameters of the initial lip-sync model until the training termination condition is met, thus obtaining the lip-sync model.
8. The method according to claim 7, characterized in that, The detailed processing model is trained using the following steps: Using the lip-sync model, based on the multi-frame sample closed-mouth face region images and the sample audio data, output multi-frame sample speaking face region images; An initial detail processing model is used to repair the details inside the oral cavity in the multi-frame sample speaking face region images, resulting in multi-frame predicted high-definition speaking face region images. The value of the third loss function is calculated based on the multi-frame predicted high-definition speaking face region image and the multi-frame real speaking face region image. Based on the value of the third loss function, the model parameters of the initial detail processing model are updated until the training termination condition is met, thus obtaining the detail processing model.
9. A voice-driven video synthesis device, characterized in that, include: A memory and a processor, wherein the memory stores a computer program executable by the processor, and the processor executes the computer program to implement the voice-driven video synthesis method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when read and executed, implements the voice-driven video synthesis method according to any one of claims 1-8.
Citation Information
Patent Citations
Voice-driven face mouth shape replacement method based on face attribute decoupling
CN118553270A
Digital population broadcast video generation method and device, equipment, storage medium and program product
CN118945420A