Video editing model training method and device, computing equipment and storage medium
By constructing sample video pairs and training a video editing model using global, lip-sync, and texture sub-models, the problem of low visual dubbing quality in existing technologies is solved, achieving stable and natural lip-sync editing effects in complex scenes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-10
AI Technical Summary
Existing video editing models suffer from low-quality visual dubbing, mainly because mask restoration causes the target frame to lose the spatiotemporal context of the dubbing object, and the pose and environment of the dubbing object in the target frame and reference frame are misaligned, resulting in unstable quality of the visual dubbing results.
By constructing sample video pairs, including a first sample video and a second sample video, a video editing model is used to edit the second sample video based on the lip movements of the first sample video, so that the lip movements of the second sample video are close to those of the first sample video. A global sub-model is used to maintain the global structure of the video frames, a lip movement sub-model is used to edit the lip movements, and a texture sub-model is used to keep the facial identity and texture information unchanged. The model is trained using speech features and video features.
It improves the lip-sync editing quality of video editing models under various poses, lighting conditions, and complex environments, ensuring that the pose of the dubbing subject and the environment remain unchanged, and enhancing the stability and naturalness of the visual dubbing results.
Smart Images

Figure CN121644910A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of video technology, and in particular to a training method, apparatus, computing device, and storage medium for a video editing model. Background Technology
[0002] With the development of lip-syncing technology for audio-driven visual dubbing, the industry has long used the basic paradigm of mask inpainting to dub a character in a video.
[0003] The object requiring dubbing in a video is called the dubbing object. In the basic paradigm of mouth masking restoration, a mask is used to cover the mouth of the dubbing object in each video frame. The video editing model uses the masked video frame as the target frame and a certain video frame as the reference frame. Based on the identity and texture information of the dubbing object in the reference frame, the mouth of the dubbing object is reconstructed within the mask of the target frame so that the lip movements of the dubbing object are aligned with the target speech. However, the mask causes the target frame to lose the complete spatiotemporal context of the dubbing object, and the pose and environment of the dubbing object in the target frame and the reference frame may not be aligned. These factors lead to low quality of the visual dubbing results of the video editing model. Therefore, there is an urgent need for a training method for video editing models that can improve the quality of visual dubbing results. Summary of the Invention
[0004] This disclosure provides a training method, apparatus, computing device, and storage medium for a video editing model to address the problem of low-quality visual dubbing results in related technologies. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a method for training a video editing model is provided, comprising: Acquire a first sample video, a second sample video, and a first sample audio, wherein the first sample video and the second sample video each include a first object, and the first sample audio is the audio of the first object in the first sample video. In the first sample video and the second sample video, the lip movements of the first object are different, and all features other than the lip movements of the first object are the same. The video editing model is trained using the second sample video as the video to be edited, the first sample video as the desired output, and the first sample speech as the constraint. During training, the video editing model edits the lip movements of the first object in the first sample video to the second sample video, based on the lip movements of the first object in the first sample video, so that the lip movements of the first object in the second sample video are closer to the lip movements of the first object in the first sample video.
[0005] Optionally, training the video editing model includes: Extract the video features of the first sample video, the video features of the second sample video, and the speech features of the first sample speech; Based on the noise coefficient, noise is added to the video features of the first sample video; The video editing model is trained based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise.
[0006] Optionally, the video editing model includes a global sub-model, a lip-sync sub-model, and a texture sub-model. The global sub-model is used to keep the global structure of the video frames in the second sample video unchanged. The lip-sync sub-model is used to edit the lip-sync of the first object in the second sample video. The identity texture sub-model is used to keep the facial identity information and facial texture information of the first object in the second sample video unchanged. The training of the video editing model based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise includes: Based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise, the global sub-model, the lip-shape sub-model, and the texture sub-model are trained respectively.
[0007] Optionally, training the global sub-model, the lip-sync sub-model, and the texture sub-model based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise, respectively, includes: The video features of the second sample video and the video features after adding noise are input into the global sub-model. Based on the video features of the second sample video and the video features after adding noise, the global sub-model predicts the global structure of the first sample video frame in the first sample video so that the predicted global structure is close to the global structure of the first sample video frame, and outputs a first predicted video that can reflect the predicted global structure. The speech features, the video features of the second sample video, and the video features after adding noise are input into the lip-sync sub-model. Based on the speech features, the video features of the second sample video, and the video features after adding noise, the lip-sync sub-model predicts the lip-sync of the first object in the first sample video so that the predicted lip-sync is close to the lip-sync of the first object in the first sample video. The lip-sync of the first object in the second sample video is edited to the predicted lip-sync, and the second predicted video is the second sample video after lip-sync editing. The video features of the second sample video and the video features after adding noise are input into the identity texture sub-model. Based on the video features of the second sample video and the video features after adding noise, the identity texture sub-model predicts the facial identity information and facial texture information of the first object in the first sample video, so that the predicted facial identity information is close to the facial identity information and the predicted facial texture information is close to the facial texture information of the first object in the first sample video. The model then outputs a third predicted video that reflects the predicted facial identity information and facial texture information.
[0008] Optionally, for different sub-models among the global sub-model, the lip-sync sub-model, and the identity texture sub-model, the noise-added video features input to the different sub-models are obtained based on noise coefficients in different intervals.
[0009] Optionally, the method further includes: Obtain a second sample speech, the content of which differs from that of the first sample speech; The first sample video and the second sample speech are processed by a video generation model to obtain the third sample video, in which the lip movements of the first object in the third sample video match the second sample speech. Based on the third sample video, the second sample video is obtained.
[0010] Optionally, the second sample speech has the same voiceprint and timbre features as the first sample speech.
[0011] Optionally, the processing of the first sample video and the second sample speech using the video generation model includes: The video generation model is used to process the masked video, the first sample video, and the second sample speech. The masked video is the first sample video with a mask added. The mask is used to cover the mouth of the first object in the first sample video but not cover the second object. The second object is any object that occludes the first object.
[0012] Optionally, the first sample video includes multiple first sample video frames, and the method further includes: For any first sample video frame, if the second object does not exist in the first image region of the first sample video frame, the mask is added to the entire first image region; if the second object exists in the first image region, the mask is added to the region outside the second object in the first image region to obtain the masked video. The first image region is the image region where the mouth of the first object is located in the first sample video frame.
[0013] Optionally, the first sample video includes multiple first sample video segments, and the second sample audio includes multiple audio segments, with different first sample video segments corresponding to different audio segments; The process of processing the first sample video and the second sample audio using a video generation model to obtain the third sample video includes: For any first sample video segment, the video generation model is used to process the first sample video segment and the corresponding speech segment to obtain a second sample video segment corresponding to the speech segment. In the second sample video segment, the lip movements of the first object match the corresponding speech segment. In the first sample video segment and the second sample video segment, the lip movements of the first object are different and all features other than the lip movements of the first object are the same. The third sample video is obtained by splicing together the second sample video segments corresponding to the multiple speech segments.
[0014] Optionally, for any first sample video segment, processing the first sample video segment and the corresponding audio segment using the video generation model includes: For the i-th video segment among the plurality of first sample video segments, starting from the last video frame of the (i-1)-th video segment among the plurality of first sample video segments, multiple video frames are obtained from the (i-1)-th video segment, where i is an integer greater than or equal to 2; Using the plurality of video frames and the audio segment corresponding to the i-th video segment as constraints, the video generation model processes the i-th video segment, the audio segment corresponding to the i-th video segment, and the plurality of video frames.
[0015] Optionally, the duration of the first sample video is greater than or equal to a first duration, the video generation model is a trained base generation model, and the method further includes: Acquire a fourth sample video and a third sample audio, wherein the duration of the fourth sample video is shorter than the duration of the first sample video, and the content of the third sample audio is different from that of the third object in the fourth sample video; Based on the fourth sample video and the third sample speech, the basic generation model is trained. During the training process, the third sample speech is used as a constraint and a video frame in the fourth sample video is used as a reference frame to edit the lip movements of the third object in the fourth sample video so that the lip movements of the first object in the fourth sample video are close to the lip movements that match the third sample speech.
[0016] Optionally, training the basic generative model based on the fourth sample video and the third sample speech includes: Based on the fourth sample video and the third sample speech, the basic generative model is overfitted and trained.
[0017] Optionally, obtaining the second sample video based on the third sample video includes: The video frames in the third sample video are subjected to multiple illumination enhancements of different degrees to obtain multiple fifth sample videos, which are the third sample videos after illumination enhancement. The second sample video is obtained from the plurality of fifth sample videos.
[0018] Optionally, obtaining the second sample video based on the third sample video includes: If the facial similarity of the first object between the third sample video and the first sample video is greater than or equal to a first threshold, and the lip shape similarity of the first object between the third sample video and the first sample video is greater than or equal to a second threshold, then the third sample video is determined as the second sample video. Wherein, the facial similarity indicates the degree of similarity between the face of the first object in the third sample video and the face of the first object in the first sample video, and the lip shape similarity indicates the degree of similarity between the lip shape of the first object in the third sample video and the lip shape of the first object in the first sample video.
[0019] According to a second aspect of the present disclosure, a video editing method is provided, comprising: Acquire a first video and a second audio that is different from the first audio content, wherein the first video includes a fourth object, and the first audio is the audio of the fourth object in the first video; Using the first video as the video to be edited and the second audio as a constraint, the lip movements of the fourth object in the first video are edited using a video editing model to obtain the second video; The video editing model is a video editing model trained by the training method in the first aspect or any possible implementation of the first aspect. The second video is the first video after lip-syncing editing. In the second video, the lip-sync of the fourth object matches the second speech. In the first video and the second video, the lip-sync of the fourth object is different, but all other features are the same except for the lip-sync of the fourth object.
[0020] According to a third aspect of the present disclosure, a training apparatus for a video editing model is provided, comprising: The first acquisition unit is configured to execute a first sample video, a second sample video, and a first sample speech, wherein the first sample video and the second sample video both include a first object, and the first sample speech is the speech of the first object in the first sample video. In the first sample video and the second sample video, the lip movements of the first object are different, and all features other than the lip movements of the first object are the same. The first training unit is configured to train a video editing model using the second sample video as the video to be edited, the first sample video as the desired output, and the first sample speech as the constraint. During training, the video editing model edits the lip movements of the first object in the first sample video to the second sample video, based on the lip movements of the first object in the first sample video, so that the lip movements of the first object in the second sample video are closer to the lip movements of the first object in the first sample video.
[0021] Optionally, the first training unit includes: The extraction subunit is configured to extract video features of the first sample video, video features of the second sample video, and speech features of the first sample speech. The noise-adding subunit is configured to perform noise addition on the video features of the first sample video based on noise-adding coefficients. The training subunit is configured to train the video editing model based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise.
[0022] Optionally, the video editing model includes a global sub-model, a lip-sync sub-model, and a texture sub-model. The global sub-model is used to keep the global structure of the video frames in the second sample video unchanged. The lip-sync sub-model is used to edit the lip-sync of the first object in the second sample video. The identity texture sub-model is used to keep the facial identity information and facial texture information of the first object in the second sample video unchanged. The training sub-unit is configured to train the global sub-model, the lip-shape sub-model, and the texture sub-model based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise.
[0023] Optionally, the training subunit is configured to perform: The video features of the second sample video and the video features after adding noise are input into the global sub-model. Based on the video features of the second sample video and the video features after adding noise, the global sub-model predicts the global structure of the first sample video frame in the first sample video so that the predicted global structure is close to the global structure of the first sample video frame, and outputs a first predicted video that can reflect the predicted global structure. The speech features, the video features of the second sample video, and the video features after adding noise are input into the lip-sync sub-model. Based on the speech features, the video features of the second sample video, and the video features after adding noise, the lip-sync sub-model predicts the lip-sync of the first object in the first sample video so that the predicted lip-sync is close to the lip-sync of the first object in the first sample video. The lip-sync of the first object in the second sample video is edited to the predicted lip-sync, and the second predicted video is the second sample video after lip-sync editing. The video features of the second sample video and the video features after adding noise are input into the identity texture sub-model. Based on the video features of the second sample video and the video features after adding noise, the identity texture sub-model predicts the facial identity information and facial texture information of the first object in the first sample video, so that the predicted facial identity information is close to the facial identity information and the predicted facial texture information is close to the facial texture information of the first object in the first sample video. The model then outputs a third predicted video that reflects the predicted facial identity information and facial texture information.
[0024] Optionally, for different sub-models among the global sub-model, the lip-sync sub-model, and the identity texture sub-model, the noise-added video features input to the different sub-models are obtained based on noise coefficients in different intervals.
[0025] Optionally, the training device for the video editing model further includes: The second acquisition unit is configured to acquire a second sample speech, the content of which is different from that of the first sample speech. The sample processing unit is configured to execute a video generation model to process the first sample video and the second sample speech to obtain the third sample video, wherein the lip movements of the first object in the third sample video match the second sample speech. The third acquisition unit is configured to acquire the second sample video based on the third sample video.
[0026] Optionally, the second sample speech has the same voiceprint and timbre features as the first sample speech.
[0027] Optionally, the sample processing unit is configured to perform processing on the masked video, the first sample video, and the second sample speech through the video generation model, wherein the masked video is the first sample video with a mask added, the mask is used to cover the mouth of the first object in the first sample video but not cover the second object, and the second object is any object that occludes the first object.
[0028] Optionally, the first sample video includes multiple first sample video frames, and the training device for the video editing model further includes: The mask adding unit is configured to perform the following for any first sample video frame: if the second object does not exist in the first image region of the first sample video frame, add the mask to the entire first image region; if the second object exists in the first image region, add the mask to the region outside the second object in the first image region, thereby obtaining the masked video. The first image region is the image region where the mouth of the first object is located in the first sample video frame.
[0029] Optionally, the first sample video includes multiple first sample video segments, and the second sample audio includes multiple audio segments, with different first sample video segments corresponding to different audio segments; The sample processing unit includes: The sample processing subunit is configured to perform, for any first sample video segment, processing the first sample video segment and the corresponding speech segment through the video generation model to obtain a second sample video segment corresponding to the speech segment, wherein the lip movements of the first object in the second sample video segment match the corresponding speech segment, and the lip movements of the first object are different in the first sample video segment and the second sample video segment, and all features other than the lip movements of the first object are the same. The splicing subunit is configured to splice the second sample video segments corresponding to the plurality of speech segments to obtain the third sample video.
[0030] Optionally, the sample processing subunit is configured to perform: For the i-th video segment among the plurality of first sample video segments, starting from the last video frame of the (i-1)-th video segment among the plurality of first sample video segments, multiple video frames are obtained from the (i-1)-th video segment, where i is an integer greater than or equal to 2; Using the plurality of video frames and the audio segment corresponding to the i-th video segment as constraints, the video generation model processes the i-th video segment, the audio segment corresponding to the i-th video segment, and the plurality of video frames.
[0031] Optionally, the duration of the first sample video is greater than or equal to a first duration, the video generation model is a trained base generation model, and the training device for the video editing model further includes: The fourth acquisition unit is configured to acquire a fourth sample video and a third sample audio, wherein the duration of the fourth sample video is shorter than the first duration, and the content of the third sample audio is different from the audio of the third object in the fourth sample video. The second training unit is configured to train the basic generative model based on the fourth sample video and the third sample speech. During the training process, the third sample speech is used as a constraint and a video frame in the fourth sample video is used as a reference frame to edit the lip movements of the third object in the fourth sample video so that the lip movements of the first object in the fourth sample video are close to the lip movements that match the third sample speech.
[0032] Optionally, the second training unit is configured to perform overfit training on the base generative model based on the fourth sample video and the third sample speech.
[0033] Optionally, the third acquisition unit is configured to perform: The video frames in the third sample video are subjected to multiple illumination enhancements of different degrees to obtain multiple fifth sample videos, which are the third sample videos after illumination enhancement. The second sample video is obtained from the plurality of fifth sample videos.
[0034] Optionally, the third acquisition unit is configured to perform: If the facial similarity of the first object between the third sample video and the first sample video is greater than or equal to a first threshold, and the lip shape similarity of the first object between the third sample video and the first sample video is greater than or equal to a second threshold, then the third sample video is determined as the second sample video. Wherein, the facial similarity indicates the degree of similarity between the face of the first object in the third sample video and the face of the first object in the first sample video, and the lip shape similarity indicates the degree of similarity between the lip shape of the first object in the third sample video and the lip shape of the first object in the first sample video.
[0035] According to a fourth aspect of the present disclosure, a video editing apparatus is provided, comprising: The acquisition unit is configured to acquire a first video and a second voice that is different from the first voice content, wherein the first video includes a fourth object and the first voice is the voice of the fourth object in the first video; The editing unit is configured to perform editing on the lip movements of the fourth object in the first video, using the first video as the video to be edited and the second speech as a constraint, through a video editing model, to obtain the second video; The video editing model is a video editing model trained by the training method in the first aspect or any possible implementation of the first aspect. The second video is the first video after lip-syncing editing. In the second video, the lip-sync of the fourth object matches the second speech. In the first video and the second video, the lip-sync of the fourth object is different, but all other features are the same except for the lip-sync of the fourth object.
[0036] According to a fifth aspect of the present disclosure, a computing device is provided, comprising: One or more processors; One or more memories for storing the one or more processor-executable instructions; The one or more processors are configured to perform a training method for a video editing model in any possible implementation of the first aspect above, or to perform a video editing method in any possible implementation of the second aspect above.
[0037] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when at least one instruction in the computer-readable storage medium is executed by one or more processors of a computing device, the computing device is enabled to perform a training method for a video editing model in any possible embodiment of the first aspect, or to perform a video editing method in any possible embodiment of the second aspect.
[0038] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising one or more instructions that can be executed by one or more processors of a computing device, such that the computing device is capable of performing a training method for a video editing model in any possible implementation of the first aspect, or performing a video editing method in any possible implementation of the second aspect.
[0039] In the training method described above, since the only difference between the video to be edited and the target video is the lip movements of the dubbing subject, the video to be edited provides complete spatiotemporal context information of the dubbing subject. Therefore, when training the video model, the target video is used as the desired output, and the target speech of the dubbing subject in the target video is used as the constraint. The video editing model only edits the lip movements of the dubbing subject in the video to be edited, based on the lip movements of the dubbing subject in the target video. It does not need to edit any parts other than the mouth of the dubbing subject, ensuring that only the lip movements of the dubbing subject change in the edited video, while the environment, posture, and identity of the dubbing subject remain unchanged. Thus, even for complex scenarios such as multiple postures of the dubbing subject and changing environmental scenes, the video editing model can maintain stable, natural, and controllable lip movement editing quality. Therefore, the video editing model trained using the above training method produces high-quality visual dubbing results.
[0040] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0042] Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for an applied video editing model according to an exemplary embodiment; Figure 2 This is a flowchart illustrating a training method for a video editing model according to an exemplary embodiment; Figure 3 This is a diagram illustrating the training principle of a basic generative model according to an exemplary embodiment; Figure 4 This is a flowchart illustrating the construction of context data according to an exemplary embodiment; Figure 5 This is a flowchart illustrating the inference process of a video generation model according to an exemplary embodiment; Figure 6 This is a flowchart illustrating a context-driven video dubbing process according to an exemplary embodiment; Figure 7 This is a diagram illustrating the training principle of a video editing model according to an exemplary embodiment; Figure 8 This is a flowchart illustrating a diffusion time-step adaptive multi-stage learning process according to an exemplary embodiment; Figure 9 This is a flowchart illustrating a video editing method according to an exemplary embodiment; Figure 10 This is a structural block diagram of a training device for a video editing model according to an exemplary embodiment; Figure 11 This is a structural block diagram of a video editing apparatus according to an exemplary embodiment; Figure 12 This is a schematic diagram of the structure of a computing device according to an exemplary embodiment. Detailed Implementation
[0043] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0044] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0045] The user information disclosed herein may be information authorized by the user or fully authorized by all parties.
[0046] In some embodiments, the meaning of A and / or B includes three cases: A and B, and A and B.
[0047] To facilitate a better understanding of the technical solutions disclosed herein, the following definitions are provided for some of the terms involved in the technical solutions disclosed herein.
[0048] Denoising Diffusion Transformer (DiT): This is a type of model that combines diffusion generation with the self-attention of a transformer to reconstruct high-fidelity images / videos in the latent space through time-step denoising.
[0049] Variational Autoencoder (VAE) is a generative coding-decoding framework that compresses video sequences into a continuous latent space from which data can be reconstructed. In this embodiment, DiT serves as the foundation for spatiotemporal modeling and conditional fusion of the video generation and editing models.
[0050] Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning method that learns a small number of incremental parameters by superimposing low-rank branches on pre-trained weights.
[0051] Lip-Sync Discriminator (SyncNet): This is a network used to evaluate the consistency between audio and lip movements, and can output synchronization confidence or error.
[0052] Additive Angular Margin Softmax (ArcFace) is a face representation learning method commonly used for identity similarity measurement.
[0053] Contrastive Language–Image Pre-training (CLIP) is a model that learns cross-modal alignment and can be used to measure the consistency of image / video and audio features.
[0054] FID (Fréchet Inception Distance) is an indicator that measures the difference between the generated image distribution and the real distribution; the smaller the value, the higher the image quality. In this embodiment, FID is used to evaluate visual quality under reference conditions.
[0055] FVD (Fréchet Video Distance): This is the video version of Fréchet distance, which combines image quality and temporal consistency; a lower value is better. In this embodiment, FVD is used to measure the overall perceived quality of the generated video.
[0056] Naturalness Image Quality Evaluator (NIQE): This is a no-reference IQA metric; a smaller value indicates a more natural image. In this embodiment, NIQE is used to evaluate the quality of a single frame in a no-reference scene.
[0057] Blind / Referenceless Image Spatial Quality Evaluator (BRISQUE): This is a commonly used no-reference IQA metric, with lower values being better. In this embodiment, BRISQUE is used to measure visual sharpness and distortion without a reference.
[0058] The Hyper-Network for IQA (HyperIQA) model is a deep IQA method where a higher value indicates better quality. In this embodiment, HyperIQA is used as a supplementary metric for no-reference image quality.
[0059] Landmark Distance (LMD) is a geometric error based on facial / mouth key points; a smaller value indicates a better fit between the lip shape and the speech. In this embodiment, LMD is used for both sample quality filtering and training / evaluation.
[0060] Cosine Similarity (CSIM) is a measure of identity similarity based on face embedding vectors; a higher value indicates greater identity consistency. This invention uses it to measure the degree to which a person's identity is preserved. In the embodiments disclosed herein, CSIM is used to measure the degree to which a person's identity is preserved.
[0061] CLIP-based similarity score (CLIP Score, CLIIPS): This is a score calculated using CLIP features to assess the semantic or appearance consistency between an image / video and a reference image. A higher value is better. In this embodiment, CLIIPS serves as a supplementary indicator for identity / semantic consistency.
[0062] Lip-Sync Confidence (Sync-C): This is a lip-to-speech consistency score based on SyncNet evaluation. A higher score indicates better synchronization.
[0063] A 3D Morphable Model (3DMM) is a statistical model used to fit and reconstruct the geometry and appearance of a face. Existing masking methods often rely on 3DMMs / landmarks for alignment.
[0064] Segment Anything Model V2 (SAM2) is a general segmentation framework that supports interactive or cue-driven mask generation. In this embodiment, SAM2 is used to remove foreground occlusion regions, reducing training difficulty.
[0065] Figure 1 This is a schematic diagram illustrating an implementation environment for a training method of an applied video editing model according to an exemplary embodiment. The implementation environment specifically includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0066] Terminal 101 is at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, MP3 player, MP4 player, and laptop computer. Terminal 101 runs an application that supports model training. Users can log in to this application through Terminal 101 to access the services provided by the application. For example, users can train models through this application on Terminal 101. Terminal 101 can also run an application that supports video editing model calls, enabling the editing of lip movements of an object in a video by calling an audio editing model. Terminal 101 can also run an application that supports video generation model calls, enabling the generation of another video based on a given video, where the lip movements of an object differ between the given video and the generated video, but all other features are identical.
[0067] Terminal 101 generally refers to one of a plurality of terminals; this embodiment uses terminal 101 as an example. Those skilled in the art will understand that the number of terminals can be more or less. For example, there may be several terminals, or dozens or hundreds of terminals, or even more. This disclosure does not limit the number of terminals or the type of device.
[0068] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 provides backend services for model training. Server 102 can connect to terminal 101 via a wireless or wired network. Server 102 can provide database services to the terminal, and of course, it can also provide other types of services, such as data preprocessing. In some embodiments, the number of servers may be more or less, and this disclosure does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services. In other embodiments, server 102 can also serve as a carrier for model training to implement the model training process of terminal 101 as described above, so as to provide the trained video editing model and / or video generation model to the terminal for use.
[0069] Among related technologies, training video editing models using the mouth mask restoration paradigm mainly faces the following common problems.
[0070] First, the root cause is the lack of context and mismatch: the paradigm of mask training covers the visual information of the lower half of the face, naturally losing the complete spatiotemporal context (including the identity texture of the lower half of the face and the lighting and occlusion relationships in the environment when in the same pose); at the same time, the number of reference frames that supplement the identity and texture is small and consists of a few static frames without temporal continuity, which are often not aligned with the target frame that needs to be repaired in terms of pose and environment, making it difficult to provide complete and aligned context support.
[0071] The combination of these factors forces the video editing model to perform two challenging tasks simultaneously within the mask: firstly, to model lip movements / expressions across modalities based on audio; and secondly, to infer and fill in missing structures and textures from an incomplete and misaligned reference context, resulting in a dual burden. This directly leads to decreased lip movement accuracy, unstable identity preservation, and poor robustness to changes in lighting, occlusion, and other environmental factors.
[0072] Secondly, mask editing introduces an inherent side effect: during training and reconstruction, the muscle movement around the mask and changes in the mask's size and shape can easily leak the original lip-sync information. Since speech is a relatively weak control condition, video editing models are prone to taking shortcuts from the visual information leaked by the mask, directly inferring and reconstructing lip movements based on this leaked information rather than the speech information. This leads to situations where, during inference, if the original lip-sync and the target speech differ significantly, the original lip-sync often interferes with the target lip-sync. A typical example is adding a silent segment to the original video; the inference result will show incorrect mouth opening and closing movements inconsistent with the silence. Furthermore, this conflict between visual and audio information can cause blurring, misalignment, or stitching marks at the mask boundaries during fusion, which is more pronounced in complex expressions or rapid movements.
[0073] Finally, the mouth mask restoration paradigm exhibits significant fragility in complex scenes: it is poorly robust to occlusion, side profiles, and highly stylized characters. This is partly due to its reliance on face detection or face-specific parameter models such as 3DMM / landmarks, which leads to "run-through failures" or severe artifacts caused by segmentation errors in stylized or extreme-angle face cases; and partly due to contextual deficiencies and reference frame mismatches, which significantly increase the difficulty of inferring mask occlusion areas in complex occlusion and extreme-angle cases.
[0074] Overall, most of these phenomena can still be traced back to the fundamental contradiction of incomplete and mismatched context: when training and inference are always limited to the masked background video and a few reference frames rather than the complete spatiotemporal context, video editing models have difficulty maintaining stable, natural and controllable lip-sync editing quality under multi-pose, multi-light, and complex interaction conditions.
[0075] Based on this, this application provides a training method for a video editing model. This method employs a two-stage paradigm of "first matching, then editing" to train the video editing model. "First matching" refers to constructing sample video pairs before training the video editing model. Each sample video pair includes a first sample video and a second sample video. The first sample video is the video to be edited, or the video to be dubbed. Both the first and second sample videos include a first object. In the first and second sample videos, the lip movements of the first object are different, but all other features are the same. These features, excluding the lip movements, include the pose of the first object and the environment in which the first object is located. In other words, the only difference between the first and second sample videos in the same sample video pair is the lip movements of the first object.
[0076] The first object is the voice-over subject in the first sample video. The first object can be a real person or animal, or a virtual object (such as a virtual character in animation). The first object has a mouth. By editing the lip movements of the first object in the first sample video, the lip movements of the first object in the first sample video can be aligned with a certain speech, thus achieving the purpose of visual dubbing for the first object in the first sample video. Specifically, the alignment of the lip movements of any object in any video with any speech can be expressed as follows: the lip movements of the object are consistent (or match) with the lip movements corresponding to the speech, where the lip movements corresponding to the speech refer to the lip movements of a real person when uttering that speech. Alternatively, the alignment of the lip movements of any object in any video with any speech can be expressed as follows: the lip movement flow of the object in the video is the same as the actual lip movement flow of the speech. The lip movement flow of any object in any video refers to the pattern of change in the object's lip movements within the video. The actual lip movement flow of any speech refers to the pattern of change in the lip movements of a person when uttering that speech.
[0077] The difference in the lip movements of the first subject in the first and second sample videos can be manifested in the following ways: the lip movement flow and the target muscle movement flow of the first subject differ between the two videos. The target muscle movement flow refers to the changing pattern of the target muscles in the video. Target muscles are facial muscles that can influence lip movements, such as the levator labii superioris, zygomaticus major, orbicularis oris, mentalis, and depressor anguli oris. In different video frames, the appearance of the target muscles differs depending on the lip movement of the first subject. Taking the levator labii superioris as the target muscle, when the first subject opens their mouth wide, the target muscle shows a noticeable lifting motion; when the mouth opening is smaller (e.g., the lips tremble), the target muscle does not show any lifting motion.
[0078] Optionally, the sample video pair also includes a first sample speech and a second sample speech, with the first sample video and the first sample speech corresponding to each other, and the second sample video and the second sample speech corresponding to each other. The first sample speech is the speech of a first object in the first sample video, and the second sample speech is the speech of the first object in the second sample speech. The speech content of the second sample speech and the first sample speech are different; therefore, the lip movements of the first object are different in the first sample video and the second sample video, or in other words, the lip movement flow of the first object is different in the first sample video and the second sample video.
[0079] Before training the video editing model, at least one sample video pair can be constructed first, and then the video editing model can be trained based on at least one sample video pair. The first object in different sample video pairs can be the same or different.
[0080] The aforementioned "re-editing" refers to the process of editing the lip movements of the first object in the second sample video of the same video pair during the training of the video editing model. This is done by editing the lip movements of the first object in the second sample video so that the lip movements of the first object in the second sample video after the lip movements are edited can be close to the lip movements of the first object in the corresponding first sample video. In other words, it is to make the lip movement flow of the first object in the first sample video and the edited second sample video approach the same.
[0081] Next, combined Figures 2 to 8 This paper introduces the training method for this video editing model.
[0082] Figure 2 This is a flowchart illustrating a training method for a video editing model according to an exemplary embodiment. The training method for the video editing model is applied to a computing device, which may be the server 102 described above. The training method for the video editing model includes the following steps.
[0083] In step 201, the computing device acquires a first sample video, a second sample video, and a first sample speech. Both the first sample video and the second sample video include a first object. The first sample speech is the speech of the first object in the first sample video. In the first sample video and the second sample video, the lip movements of the first object are different, but all features other than the lip movements of the first object are the same.
[0084] Among them, the first sample video, the second sample video, and the first sample speech belong to the same sample video pair. The first object is as described above. The first object's lip movements are different in the first sample video and the second sample video, and all other features are the same, as described above, and will not be repeated here.
[0085] The video frames in the first sample video are called first sample video frames, and the video frames in the second sample video are called second sample video frames. The first sample video includes multiple first sample video frames, and the second sample video includes multiple second sample video frames. The number of first sample video frames in the first sample video is the same as the number of second sample video frames in the second sample video. The content of the j-th video frame in the first sample video corresponds to that of the j-th video frame in the second sample video. This correspondence can be expressed as follows: in the j-th video frame of the first sample video and the j-th video frame of the second sample video, the lip movements of the first object are different, and all other features (such as the pose of the first object or the environment in which the first object is located) are the same. j is an integer greater than 0.
[0086] In some embodiments, for a first sample video and a second sample video in the same sample video pair, the second sample video is an accompanying video of the first sample video, and the computing device generates the second sample video based on the first sample video and the second sample video. For example, the generation process of the second sample video can be implemented through a video generation model.
[0087] The video generation model is a trained base generation model. This model processes a video with a given speech as a constraint to generate a new video in which the lip movements of an object are aligned with the speech. The base generation model is a neural network model, and can be a DiT model.
[0088] To facilitate understanding, the training process of the basic generative model will be introduced first in conjunction with (1.1) below, and then the process of generating the second sample video through the video generation model will be introduced in conjunction with (1.2) below.
[0089] (1.1) Training process of the basic generative model The sample videos used to train the base generative model are called the fourth sample videos. The voice-over object in the fourth sample video is called the third object, and each video frame in the fourth sample video includes a third object. The third object can be a real person or animal, or a virtual object (such as a virtual character in an animation). The third object and the first object can be the same object or different objects. The third object has a mouth. By reconstructing the lip movements of the third object in the fourth sample video so that the lip movements of the third object in the fourth sample video can be aligned with a certain speech, the purpose of visual dubbing for the third object in the fourth sample video is achieved.
[0090] The speech used to provide visual dubbing for any object in any video is called the target speech. The target speech during the training process of the basic generative model is called the third sample speech. The speech of the third object in the fourth sample video is called the fourth sample speech. The content of the third sample speech and the fourth sample speech are different.
[0091] The computing device first acquires a fourth sample video and a third sample speech, and then trains the basic generative model based on the fourth sample video and the third sample speech. During the training process, using the third sample speech as a constraint and a video frame from the fourth sample video as a reference frame, the lip movements of the third object in the fourth sample video are reconstructed so that the lip movements of the first object in the fourth sample video approximate those of the third sample speech. The fourth sample video is shorter than the first sample video; compared to videos with a duration greater than or equal to the first sample video, the fourth sample video is a short video. Training the basic generative model based on a short video reduces the amount of data processing during training and improves training efficiency. The first sample video duration can be reasonably set according to the specific implementation scenario; here, this embodiment does not limit the first sample video duration.
[0092] In some embodiments, the computing device follows the basic paradigm of mask repair to train a basic generative model.
[0093] For example, the fourth sample video is used as the target video. Target video In a video frame, the image region containing the mouth of the third object is called the second image region. This second image region includes the complete mouth of the third object; it can be either the complete face region of the third object or the lower portion of its face. For the target video... In each video frame, a mask is added to the entire second image region within that frame to completely cover the mouth of the third object. For masked video From the target video A video frame is randomly selected as a reference frame. Reference frame This is used to provide facial identity information and facial texture information of a third object. The facial identity information of any object represents the object's face, and different objects have different facial identity information. The facial texture information of any object represents the texture of the object's face.
[0094] Using the third sample speech as the target speech Target speech Target Video Masked video and reference frame This forms a training sample. From this training sample, the target video is extracted using a Video Array (VAE). Potential features in the potential space Masked video Potential features in the potential space and reference frame Potential features in the potential space Based on the Gaussian distribution and the noise coefficients in the numerical range of 0 to 1, latent features are analyzed. Add noise to enhance latent features. The noise in the sample is used to obtain the latent features after adding noise. .like Figure 3 As shown, at the frame dimension, for latent features and latent features after adding noise The stitching process yields stitching feature 31, which is used to spatially align the target video. and masked video Background information, splicing feature 31 includes latent features and latent features after adding noise A feature vector (or feature) consisting of zeros is called a zero feature. In the feature channel dimension, this represents the latent feature. The 0 feature is concatenated with the 0 feature to obtain concatenated feature 32. Concatenated feature 32 enhances the identity information of the third object, and is aligned with concatenated feature 31. The identity information of the third object includes facial features and facial texture information. The target speech is extracted using an acoustic encoder (Whisper). The speech features of the target speech. Using the speech features as constraints, and splicing features 31 and 32 as model inputs, the speech features, splicing features 31 and 32 are input into DiT, and DiT processes the target speech. The speech features, splicing features 31 and 32 are processed to obtain the predicted video, and the predicted video is output. The predicted video is the predicted target video. The third object in the middle is paired with the target speech The following video. Among them, the target speech... The speech features, splicing features 31 and 32 are processed to complete the following process: based on the target speech and target video The lip movements of the third object in the masked video The lip shape of the third object and the area excluding the lip of the third object in the second image region are reconstructed within the mask in the image. The reconstructed second image region is then compared with the mask video. The image regions in the video frame that are not covered by the mask are fused together, and the fused video frames are used to form the predicted video.
[0095] The basic generative model can be DiT, which includes an input layer, at least one transformer, and an output layer. These layers are sequentially connected. Each transformer includes a self-attention layer and a cross-attention layer. The input layer can perform encoding and positional encoding on the input. Within any transformer, the self-attention layer executes a self-attention mechanism based on the processing result of the previous layer (such as the output layer or the previous transformer). After the self-attention layer finishes processing, the cross-attention layer in that transformer outputs the target speech based on the output of the self-attention layer. The speech features are used as driving conditions to execute a cross-attention mechanism, incorporating the conditional information into the model's processing to obtain cross-attention features. These cross-attention features are then input into the next layer (the next transformer or output layer). The output layer processes the received cross-attention features to output the predicted video. It should be noted that technicians can control the model's dimensions, depth, and other parameters according to actual needs, thereby controlling the model's size. This disclosure does not limit this aspect.
[0096] like Figure 3 As shown, the basic generative model is DiT, which includes a patchify layer and a transformer. The transformer includes a 2D self-attention layer, a 3D self-attention layer, and a cross-attention layer. The patchify layer is located in the input layer, and both the 2D and 3D self-attention layers are self-attention layers. Patchify features 31 and 32 are input to the patchify layer. The patchify layer extracts features from both patchify features 31 and 32 to obtain a first patchify feature and a second patchify feature. The first patchify feature includes features extracted from each video frame from patchify feature 301, and the second patchify feature includes features extracted from each video frame from patchify feature 302. The first and second patchify features are then input to the 2D self-attention layer. The 2D self-attention layer uses a 2D self-attention mechanism to fuse the features within each video frame in the first and second patchify features, respectively, to obtain the fused features corresponding to each video frame. The fused features corresponding to each video frame are then input to the 3D self-attention layer. The 3D self-attention layer employs a 3D self-attention mechanism to fuse features corresponding to each video frame, obtaining new fused features, which are then input into the cross-attention layer. (The target speech is used as an example.) The speech features are used as constraints, and these speech features are injected into the cross-attention layer. The cross-attention layer employs a cross-attention mechanism to fuse the input fusion features and speech features, obtaining a new fusion feature (i.e., the cross-attention feature). This new fusion feature is then input into the output layer of DiT (not in the previous step). Figure 3 As shown in the diagram), the output layer generates a predicted video based on the new fusion features, and outputs the predicted video, thus the basic generative model is trained on a single training sample (including...). and This completes one training iteration. Similarly, a video generation model can be trained on multiple training samples, outputting a predicted video for each training sample.
[0097] For each training sample, based on the fourth sample in that training sample and the corresponding prediction video, the flow matching loss function is used to calculate the flow matching loss value for that training sample. The flow matching loss value indicates the lip movement flow of the third object in the prediction video and the target speech. The degree of difference between the actual lip-sync motion flows. When the flow matching loss values corresponding to multiple training samples do not converge to the first loss value interval, it indicates that the basic generative model has not yet learned from the target video. If the mapping from the lip-sync flow of the third object in the video to the real lip-sync flow is found, the model parameters of the base generative model are updated. The updated base generative model is then trained again until the flow matching loss values corresponding to multiple training samples in a certain training process converge to the first loss value interval. This indicates that the base generative model has learned the correct flow matching from the target video. The mapping from the lip-sync flow of the third object to the actual lip-sync flow is learned from the noisy target video. The noise distribution is accurately mapped to the data distribution. The basic generative model training is complete. The first loss value interval can be reasonably set; however, this embodiment does not limit the first loss value interval.
[0098] It should be understood that the target speech and target video The content of the speech of the third object is different from that of the target speech. Realistic lip-syncing motion flow and target video The lip-sync flow of the third object differs from that of the trained base generative model, which learns from the target video. The lip movements of the third object flow into the target speech The mapping of the real lip-sync motion flow, therefore, for new target videos New target video The voice actors and new target voices The lip movements of the dubbing subject in the predicted video output by the trained base generative model can match the new target speech. Alignment, due to the new target speech With new target video Since the content of the voice-over subject differs from that of the new target video, it is possible to ensure that the lip-sync flow of the voice-over subject in the predicted video matches the new target video. The lip movements of the voice-over characters differ from those of the target video, making it difficult to predict the lip movements of the voice-over characters in the target video. The lip movements of the dubbing subjects are different. Furthermore, because a second image region is reconstructed during training, this reconstructed second image region is compared with the mask video. The image regions not covered by the mask are fused to ensure that features in the predicted video, excluding the lip movements of a third object, match those in the target video as closely as possible. Similarly, during training, for features other than the lip movements of a third object, the base generative model learns from the noisy target video. The training of the base generative model ensures an accurate mapping from the noise distribution to the data distribution of the predicted video, thus enabling the model to reconstruct the target video. When dubbing the lip movements of the subject, it is possible to ensure the target video... The features of the voice actor remain unchanged except for the lip movements of the voice actor.
[0099] In some embodiments, during the training of the basic generative model, when the number of training iterations of the basic generative model reaches the iteration threshold and the flow matching loss values corresponding to multiple training samples converge to the first loss value range, the basic generative model is trained for the target number of iterations. This is to ensure that the number of training iterations of the basic generative model exceeds the iteration threshold, thus achieving overfit training of the basic generative model (i.e., overfit training of the basic generative model based on the fourth sample video and the third sample speech). Overfit training enables the trained basic generative model to learn from the target video... The lip movements of the third object flow into the target speech A more accurate mapping of the real lip-sync motion flow allows for more stable quality of sample video pairs when using the trained generative model to construct sample video pairs.
[0100] In the embodiments of this disclosure, when training the video generation model, "absolute audio matching" between the predicted video and the target video is not pursued. It is only necessary to ensure that there is a measurable difference between the lip movements of the dubbing object and the target video, thereby effectively avoiding lip movements that are leaked by the learning mask.
[0101] (1.2) The process of generating the first sample video based on the second sample video through a video generation model. The speech of the first object in the first sample video is called the first sample speech, and the speech of the first object in the second sample video is called the second sample speech. The content of the second sample speech is different from that of the first sample speech.
[0102] The trained base generative model is called the video generation model. In order to obtain the second sample video paired with the first sample video, the computing device first obtains the second sample speech, and then processes the first sample video and the second sample speech through the video generation model to obtain the third sample video. Based on the third sample video, the second sample video is obtained.
[0103] The third sample video is the target video of the first sample video. Using the second sample speech as the target speech At that time, the basic generative model is for the target video and target speech The output predicted video.
[0104] The process of processing the first sample video and the second sample speech using the video generation model is similar to the process of training the basic generation model on the target video. and target speech The processing procedure is the same.
[0105] For example, the computing device uses the first sample video as the target video. Video targeting the first object For the voice-over object, obtain the mask video of the first sample video, and use the second sample speech as the target speech. The system uses a video generation model to process a masked video, a first sample video, and a second sample speech. The masked video is the first sample video with a mask added, used to cover the mouth of the first object in the first sample video.
[0106] The image region where the mouth of the first object is located in the first sample video frame is called the first image region. The first image region contains the complete mouth of the first object. The first image region can be the complete face region of the first object or the lower part of the face region of the first object.
[0107] In some embodiments, the computing device can acquire the target video during the training process of the basic generative model. Masked video The method used is to obtain the mask video of the first sample video. The difference is that the mask video of the first sample video covers the first image region, and the mask video is used during model training. The mask in the image covers the second image region.
[0108] Considering that some video frames of the first sample video contain objects that occlude the first object (i.e., foreground objects that occlude the first object), if the mask in the masked video covers these foreground objects, the video generation model will need to reconstruct these foreground objects within the mask during the processing of the masked video, which will reduce the processing efficiency of the video generation model.
[0109] Based on this, in some embodiments, the object that occludes the first object in the first sample video is referred to as any object as the second object, and the mask in the mask video of the first sample video does not cover the second object, that is, the mask is used to cover the mouth of the first object in the first sample video but does not cover the second object.
[0110] The method for obtaining the mask video of the first sample video can be as follows: For any first sample video frame in the first sample video, if there is no second object in the first image region of the first sample video frame, a mask is added to the entire first image region; if there is a second object in the first image region, a mask is added to the region outside the second object in the first image region, so that the mask does not cover the second object in the first image region, that is, the second object is retained in the first sample video frame. After each first sample video frame is masked in this way, the first sample video with the mask added is used as the mask video of the first sample video, thus obtaining the mask video of the first sample video, and the second object is retained in the mask video. Subsequently, in the process of processing the mask video of the first sample video, the first sample video, and the second sample speech by the video generation model, since the foreground (i.e., the second object) that occludes the first object is retained in the mask video, the video generation model does not need to reconstruct these foregrounds within the mask, which not only improves the processing efficiency of the video generation model, but also maintains foreground consistency.
[0111] The process described above, which uses a video generation model to process the mask video of the first sample video, the first sample video, and the second sample speech, is similar to the process in training the basic generation model, where the basic generation model is used to process the target video. Masked video and target speech The same principle applies to the output.
[0112] For example, the video generation model processes the mask video of the first sample video, the first sample video, and the second sample speech by including: using the first sample video as the target video. The mask video of the first sample video is used as the mask video. Using the second sample speech as the target speech A video frame is randomly selected from the first sample video as a reference frame. The reference frame at this time Used to provide facial identity information and facial texture information of the first object, and with the second sample speech as the target speech. Extract target video Potential characteristics Masked video Potential characteristics and reference frame Potential characteristics At the frame level, for latent features and latent features after adding noise By concatenating the features, a concatenated feature is obtained. In the feature channel dimension, the latent features are... The zero feature is concatenated with another concatenated feature, which is then used by an acoustic encoder to extract the target speech. The speech features of the target speech The speech features are used as constraints. These two concatenated features and the speech features are input into the video generation model. The video generation model processes the input features and outputs a predicted video, which is also the third sample video.
[0113] Because the video generation model learns from the target video The lip movements of the dubbing subject are transmitted to the target speech. The mapping of the real lip-sync motion flow, and the video generation model in reconstructing the target video. When dubbing the lip movements of the subject, it is possible to ensure the target video... The features of the voice-over subject remain unchanged except for the lip movements. Therefore, for the third sample video generated by the video generation model based on the first sample video, the lip movements of the first subject are different in the third sample video and the first sample video, while all other features are the same.
[0114] The above example illustrates feature extraction and feature concatenation performed outside of the video generation model. In some embodiments, the video generation model can also perform feature extraction and feature concatenation.
[0115] by Figure 4 Taking the context data construction process shown as an example, assuming video... The first sample video, audio The first sample is audio, and the video is... The third sample video, audio The second sample is speech, and DiT is the video generation model. First, the video... The video frames are sampled to extract one video frame as a reference frame. A mask is added to the mouth region of the first object in video V to obtain the masked video. Regarding the video Adding noise to obtain video (i.e., the video after noise is added) ), reference frame Masked video ,video ,video And voice The input is a video generation model, and the video generation model uses the input reference frame. Masked video ,video video And voice It completes feature extraction, feature stitching, and processing of stitched features, and outputs a video. ,video That is, for video and voice The output is the predicted video.
[0116] In some embodiments, the first sample speech and the second sample speech are the speech of the same person. When they are the speech of the same person, the voiceprint features and timbre features of the first sample speech and the second sample speech are the same. Here, the voiceprint feature of any speech refers to the features of the voiceprint of the speech, and the timbre feature of any speech refers to the features of the timbre of the speech. The voiceprint features and timbre features of different people are different. Therefore, when the voiceprint features and timbre features of multiple speech are different, it means that these multiple speech come from different people. When the voiceprint features and timbre features of the first sample speech and the second sample speech are the same, it means that the first sample speech and the second sample speech come from the same person.
[0117] In the process of processing the first sample video and the second sample speech through the image processing model, when the first sample speech and the second sample speech are from the same person, it can avoid cross-identity between the first sample speech and the second sample speech (i.e., speech cross-identity), minimize the conflict and interference caused by speech cross-identity, and improve the processing efficiency of the image processing model and the accuracy of generating predicted videos.
[0118] The aforementioned video generation model is trained on short videos shorter than the first sample video. The duration of the first sample video may be shorter than the first sample video, or it may be longer than or equal to the first sample video. When the duration of the first sample video is shorter than the first sample video, a third sample video can be obtained through the video generation model as described above (i.e., the first sample video and the second sample speech are processed by the video generation model to obtain the third sample video). This enables the video generation model to process short videos. Since the various video frames in a short video have similar head movements and environments (including background, occlusion, and lighting conditions), the reference frames extracted from the short video have similar head movements and environments to other video frames in the short video during processing. Therefore, the reference frames can provide similar contextual reference conditions for the video frames in the mask video, maximizing the stability and consistency of visual information between the generated predicted video and the target video. This ensures that the features in the predicted video and the target video, except for the lip movements of the dubbing subject, are identical, thus guaranteeing the quality of the predicted video generated by the video generation model.
[0119] If the duration of the first sample video is greater than or equal to the first duration, the first sample video can be divided into multiple video segments (called first sample video segments). Each first sample video segment is processed by the video generation model to obtain the predicted video segment corresponding to each first sample video segment. The predicted video segments corresponding to multiple first sample video segments are spliced together to form the third sample video, thereby obtaining a long video. This long video is then used to train a video editing model, enabling the video editing model to support editing the lip movements of the dubbing objects in the long video.
[0120] For example, the first sample video includes multiple first sample video segments, each with a duration shorter than a first duration; the second sample speech includes multiple speech segments, with different first sample video segments corresponding to different speech segments in the second sample speech; the third sample video includes multiple second sample video segments, each second sample video segment being generated based on a first sample video segment and the speech segment corresponding to that first sample video segment; different second sample video segments being generated based on different first sample video segments and different speech segments; the first sample video segment and the speech segment used to generate any second sample video segment are both corresponding to that second sample video segment; in the second sample video segment, the lip movements of the first object match the corresponding speech segment; in the corresponding second sample video segment and the first sample video segment, the lip movements of the first object are different, and all features other than the lip movements of the first object are the same.
[0121] Based on this, in some embodiments, the process of generating the third sample video includes: for any first sample video segment, the computing device processes the first sample video segment and the corresponding audio segment through a video generation model to obtain a second sample video segment corresponding to the audio segment; similarly, the second sample video segments corresponding to each audio segment in the second sample audio can be obtained, that is, the second sample video segments corresponding to multiple audio segments can be obtained; the second sample video segments corresponding to multiple audio segments are spliced together to obtain the third sample video.
[0122] In this context, the second sample video segment corresponding to any given speech segment is a predicted video generated by the video generation model based on the first sample video segment and the speech segment. The process of processing the first sample video segment and the corresponding speech segment through the video generation model is similar to the process of obtaining the third sample video through the video generation model when the duration of the first sample video is less than the first duration, and will not be elaborated further here.
[0123] For example, suppose the first sample video includes first sample video segments A1 and A2, and the second sample speech includes speech segments B1 and B2. The first sample video segment A1 and speech segment B1 are processed by the video generation model to obtain the predicted video C1 (i.e., the second sample video segment). The first sample video segment A2 and speech segment B2 are processed by the video generation model to obtain the predicted video C2 (i.e., the second sample video segment). The predicted videos C1 and C2 are spliced together to obtain the spliced video D, which is the third sample video.
[0124] In some embodiments, when processing each first sample video segment separately by the video generation model, starting from the second sample video segment, the predicted video corresponding to each first sample video segment is obtained based on the last few video frames in the predicted video corresponding to the previous sample video segment. This ensures that the predicted video corresponding to the first sample video segment is visually consistent with the previous few video frames. Thus, when the predicted videos corresponding to multiple first sample video segments are stitched together to form a third sample video, the continuity of the connection between different video segments in the third sample video can be ensured.
[0125] Below, taking the i-th sample video segment among multiple first sample video segments as an example, we will introduce how to obtain the prediction video corresponding to the i-th sample video segment by taking the last few video frames in the prediction video corresponding to the previous sample video segment as a condition, where i is an integer greater than or equal to 2.
[0126] For the i-th video segment among multiple first sample video segments, the computing device takes the last video frame of the (i-1)-th video segment as the starting point and obtains multiple video frames from the (i-1)-th video segment. Using these multiple video frames and the audio segment corresponding to the i-th video segment as constraints, the device processes the i-th video segment, the audio segment corresponding to the i-th video segment, and the multiple video frames through a video generation model to obtain the second sample video segment corresponding to the i-th video segment.
[0127] The multiple video frames refer to the last few video frames in the (i-1)th video segment, such as the last 5 video frames in the (i-1)th video segment, but the number of video frames obtained from the (i-1)th video segment is not limited to 5.
[0128] by Figure 5 The DiT model shown here is an example of a video generation model, using multiple video frames obtained from the (i-1)th video segment as historical reference frames. The i-th video segment is the target video. The target speech is the speech segment corresponding to the i-th video segment in the second sample speech. Add a mask to the first image region in the i-th video segment to obtain the masked video of the i-th video segment. Using a certain video frame in the i-th video segment as the reference frame Extract the target video using VAE. Potential characteristics Masked video Potential characteristics Reference frame Potential characteristics and multiple historical reference frames Potential characteristics At the frame level, for latent features and latent features after adding noise The features are concatenated to obtain concatenated feature 51. In the feature channel dimension, the latent features are... Concatenating the 0 feature with the concatenated feature 52, along with the latent features along the feature channel dimension, yields concatenated feature 52. The features 0 and 53 are concatenated to obtain concatenated features 53; the target speech is then extracted using an acoustic encoder. The speech features; using splicing feature 53 as a constraint, splicing features 51 to 53 are input into the repair layer in DiT to obtain the target speech. Using the speech features as constraints, the target speech The speech features are injected into the cross-attention layer of DiT. DiT extracts the features of splicing features 51 to 53 through the repair layer. Using the features extracted from splicing features 53 as constraints, the features extracted from splicing features 51 to 53 are processed through two-dimensional self-attention layers and three-dimensional self-attention layers. The processed features and the injected speech features are processed through the cross-attention layer. The output layer generates a predicted video based on the features processed by the cross-attention layer and outputs the predicted video, which is the second sample video segment corresponding to the i-th video segment.
[0129] In a similar manner, the second sample video segments corresponding to each first sample video segment other than the first first sample video segment can be obtained. Then, the second sample video segments corresponding to all the first sample segments can be spliced together to obtain the third sample segment.
[0130] The above describes the process of generating a third sample video from a second sample speech and a first sample video using a video generation model. This allows a subsequent computing device to obtain a second sample video based on the third sample video, and the first sample video and the obtained second sample video together form a sample video pair.
[0131] Next, we will introduce the process of obtaining the second sample video based on the third sample video.
[0132] In some embodiments, the computing device uses the third sample video as the second sample video.
[0133] In some embodiments, the computing device obtains the facial similarity of a first object between a third sample video and a first sample video, wherein the facial similarity indicates the degree of similarity between the face of the first object in the third sample video and the face of the first object in the first sample video.
[0134] Optionally, each video frame of the third sample video corresponds to a first sample video frame of the first sample video, and the facial similarity includes the similarity between the face of the first object in each video frame of the third sample video and the face of the first object in the corresponding first sample video frame. For each video frame of the third sample video, the computing device calculates the similarity between the face of the first object in that video frame and the face of the first object in the corresponding first sample video frame based on the face image region of the first object in that video frame and the face of the first object in the corresponding first sample video frame. The similarity between the face of the first object in each video frame of the third sample video and the face of the first object in the corresponding first sample video frame is the facial similarity, wherein the face image region of the first object refers to the image region of the face of the first object.
[0135] The greater the facial similarity, the more similar the face of the first object in the third sample video is to the face of the first object in the first sample video, and the smaller the difference in facial identity between the first and third sample videos (i.e., the more identical the facial identities). Conversely, the smaller the facial similarity, the less similar the face of the first object in the third sample video is to the face of the first object in the first sample video, and the greater the difference in facial identity between the first and third sample videos (i.e., the more different the facial identities). For example, facial similarity can be the cosine similarity of the first object's face between the third and first sample videos (i.e., the cosine similarity between the first object's face in the third and first sample videos). The computing device can use ArcFace to process the third and first sample videos to obtain the cosine similarity of the first object's face between the third and first sample videos. For example, the facial similarity can be the similarity score between the face of the first object in the third sample video and the face of the first object in the first sample video (the similarity score of the first object's face between the third sample video and the first sample video). The computing device processes the third sample video and the first sample video using CLIP to obtain this similarity score (i.e., CLIIPS). By using the cosine similarity calculated by ArcFace and the CLIIPS calculated by CLIP, the identity and appearance of the dubbing object in the same sample video pair can be kept stable.
[0136] The computing device also acquires the lip-shape similarity of the first object between the third sample video and the first sample video. This lip-shape similarity indicates the degree of similarity between the lip-shape of the first object in the third sample video and the lip-shape of the first object in the first sample video, or in other words, the degree of similarity between the lip-shape motion flow of the first object in the third sample video and the lip-shape motion flow of the first object in the first sample video. Optionally, the lip-shape similarity includes the similarity between the lip-shape of the first object in each video frame of the third sample video and the lip-shape of the first object in the corresponding first sample video frame. For each video frame of the third sample video, the computing device can calculate the similarity between the lip-shape of the first object in that video frame and the lip-shape of the first object in the corresponding first sample video frame, based on the mouth image region of the first object in that video frame and the mouth image region of the first object in the corresponding first sample video frame. Optionally, the similarity between the lip-shape of the first object in any video frame of the third sample video and the lip-shape of the first object in the corresponding first sample video frame is the keypoint distance between the keypoints of the mouth of the first object in that video frame and the keypoints of the mouth of the first object in the corresponding first sample video frame.
[0137] The smaller the lip-sync similarity between the first and third sample videos, the more similar the lip-sync of the first object is between the third and first sample videos; the greater the lip-sync similarity between the first and third sample videos, the less similar the lip-sync of the first object is between the third and first sample videos.
[0138] For a first sample video, the goal is to find a second sample video that has a similar facial identity but a different lip shape than the first object in the first sample video. Therefore, in some embodiments, if the facial similarity of the first object between the third sample video and the first sample video is greater than or equal to a first threshold, and the lip shape similarity of the first object between the third sample video and the first sample video is less than or equal to a second threshold, the computing device determines the third sample video as the second sample video. This ensures that the facial identity of the first object is the same in both the first and second sample videos, but the lip shape is different. Combining these first and second sample videos into a sample video pair, and using this sample video to train a video editing model, ensures that the trained video editing model, when processing the video to be edited, maintains the facial identity of the voice-over object while changing the lip shape, thereby greatly improving the quality of the edited video. The first and second thresholds can be reasonably set according to the specific implementation scenario; here, this disclosure does not limit the first and second thresholds.
[0139] Of course, if the facial similarity of the first object between the third sample video and the first sample video is less than the first threshold, and / or the lip shape similarity of the first object between the third sample video and the first sample video is greater than the second threshold, then the third sample video will not be identified as the second sample video.
[0140] When a third sample video is generated based on a first sample video, and a second sample video is obtained based on this third sample video, the computing device can construct a sample video based on the first and second sample videos. For example, the first sample video, the second sample video, the first sample speech, and the second sample speech can be combined into a sample video pair, in which the first sample video and the first sample speech correspond, and the second sample video and the second sample speech correspond. Figure 4 For example, video The first sample video, audio The first sample is audio, and the video is... The third sample video, audio For the second sample audio, the video... As the second sample video, the video ,voice ,voice ,video The resulting sample video pairs can be represented as Due to the video and The only difference between them is the lip movements of the first subject in the video. Each sample video frame in the dataset has complete spatiotemporal context information, and its overhead context information is consistent with the video context. The spatiotemporal context information of the corresponding video frames is the same, thus constructing sample video pairs, which is to construct complete context data.
[0141] When there are multiple first sample videos, the computing device can generate a corresponding second sample video based on each first sample video in a similar manner. Based on each first sample video and the corresponding second sample video, a sample video pair can be obtained, thus obtaining multiple sample video pairs.
[0142] The sample video pairs serve as training samples for the video editing model. The more diverse and complex the lighting in the video frames of multiple sample video pairs, the more the video editing model trained based on multiple sample video pairs can still guarantee natural and controllable lip-sync editing quality, even when faced with videos with diverse and complex lighting conditions.
[0143] Therefore, to improve the diversity and complexity of the training samples for the video editing model in terms of illumination, in some embodiments, for a third sample video generated based on a first sample video, the computing device performs multiple illumination enhancements on the video frames in the third sample video to different degrees, resulting in multiple fifth sample videos, each of which is an illumination-enhanced third sample video. For example, based on multiple different illumination enhancement coefficients, illumination enhancement is performed on the video frames in the third sample video to increase the illumination intensity in the video frames of the third sample video. The third sample video enhanced based on one illumination enhancement coefficient is a fifth sample video frame, and the illumination intensity of the same video frame in different fifth sample videos is different.
[0144] The computing device acquires second sample videos from multiple fifth sample videos. For example, each fifth sample video can be used as a second sample video, thus multiple second sample videos can be obtained based on a third sample. Alternatively, for any fifth sample video, if the facial similarity of the first object between the fifth sample video and the first sample video is greater than or equal to a first threshold, and the lip shape similarity of the first object between the fifth sample video and the first sample video is less than or equal to a second threshold, then the fifth sample video is determined as a second sample video.
[0145] In some embodiments, if the facial similarity of the first object between the third sample video and the first sample video is greater than or equal to a first threshold, and the lip shape similarity of the first object between the third sample video and the first sample video is less than or equal to a second threshold, the computing device performs multiple illumination enhancements on the video frames in the third sample video to obtain multiple fifth sample videos, and each fifth sample video is used as a second sample video.
[0146] As described above, multiple second sample videos can be obtained based on a first sample video. For each second sample video, a sample video pair can be formed by the second sample video and the first sample video. Thus, multiple sample video pairs can be obtained based on a first sample video. Each sample video pair includes the first sample video. The light intensity in the second sample videos in different sample video pairs is different, thereby increasing the diversity and complexity of the sample videos in terms of lighting.
[0147] In step 202, the computing device trains the video editing model using the second sample video as the video to be edited, the first sample video as the desired output, and the first sample speech as the constraint. During the training process, the video editing model edits the lip movements of the first object in the second sample video based on the lip movements of the first object in the first sample video, so that the lip movements of the first object in the second sample video are close to the lip movements of the first object in the first sample video.
[0148] The second sample video, the first sample video, and the first sample audio are obtained through step 201 above, and they belong to the same sample video pair. The expected output is the video that the video editing model is expected to output for the video to be edited.
[0149] Each sample video pair serves as a training sample for the video editing model. The computing device trains the video editing model based on at least one sample video pair. During training, each sample video pair is processed in the same way. For ease of understanding, the principle of the video editing model will be introduced below using a single sample video pair.
[0150] For a pair of sample videos, including a second sample video and a first sample video, a video editing model is trained with the goal of matching the first object in the second sample video with the first sample speech. Since the only difference between the second and first sample videos is the lip movements of the first object, and each frame of the second sample video has complete spatiotemporal context information, and its temporal context information is the same as that of the corresponding frame of the first sample video, each frame of the second sample video can be used as a reference frame (i.e., the entire second sample video is used as a reference video) during the training of the video editing model, and the first sample video can be used as the desired output (i.e., the target video). By editing the lip movements of the first object in each frame of the second sample video to match the lip movements in the corresponding frame of the first sample video, the second sample video can be edited into the first sample video, aligning the lip movements of the first object in the edited second sample video with the first sample speech, without needing to follow the basic paradigm of lip mask restoration by adding a mask to the first sample video and reconstructing the lip movements within the mask. Therefore, when training the video editing model, the training objective is to converge the lip-sync loss value of the first object between the first sample video and the predicted video to a first loss value interval. The second sample video is used as the video to be edited, the first sample video is used as the desired output, and the first sample speech is used as the constraint. During training, the video editing model edits the lip-sync of the first object in the second sample video based on the lip-sync of the first object in the first sample video, resulting in the predicted video (i.e., the edited second sample video). This lip-sync loss value indicates the degree of difference between the lip-sync motion flow of the first object in the predicted video and the lip-sync motion flow of the first object in the first sample video, where the lip-sync motion flow of the first object in the first sample video is the actual lip-sync motion flow of the first sample speech.
[0151] When using the DiT video editing model, iterative training is performed on the model based on the first sample video, the second sample video, and the first sample speech. During each training iteration, the first object in the second sample video is used as the dubbing subject, and the first sample video is used as the target video. Noise is gradually added to the video frames of the first sample video until it approaches pure noise, achieving forward diffusion. Using the first sample speech as a constraint, the video editing model processes the second sample video, the noisy first sample video, and the first sample speech. During processing, the Transformer video editing model predicts noise and removes noise from the noisy first sample video to recover the first sample video as much as possible, achieving inverse denoising and outputting the recovered first sample video. The recovered first sample video is also the predicted video after matching the first object in the second sample video with the first sample speech (i.e., the predicted video). Multiple loss values are calculated between the first sample video and the predicted video. These loss values include lip-sync loss. When any loss value fails to reach the expected convergence state, the model parameters of the video editing model are updated. The video editing model is then trained again based on the first sample video, the second sample video, and the first sample speech, until all loss values reach the expected convergence state during a training process. The video editing model can then completely recover the first sample video, thus enabling the video editing model to learn the mapping between the video to be edited and the target video when matching the target speech (such as the first sample speech) to an object (such as the first object) in a video to be edited (such as the second sample video).
[0152] It should be understood that since the only difference between the two sample videos is the lip movements of the first object, the process of processing the second sample video, the noisy first sample video, and the first sample speech using the video editing model can be represented as follows: the video editing model edits the lip movements of the first object in the second sample video based on the lip movements of the first object in the first sample video. During training, the video editing model edits the lip movements of the first object in the second sample video based on the lip movements of the first object in the first sample video, making the lip movements of the first object in the second sample video gradually approach those in the first sample video. After the video editing model is trained, with the first sample speech as a constraint, the predicted video output by the video editing model for the second sample video to be edited is the same as that of the first sample video. Since the only difference between the two sample videos is the lip movements of the first object, when matching the first object with the first sample speech in the second sample video, the trained video editing model can achieve the goal of keeping all features in the second sample video unchanged except for the lip movements of the first object, while only editing the lip movements of the first object in the second sample video, so that the lip movements of the first object in the edited second sample video are aligned with the first sample speech.
[0153] In order to better understand the above training principles, Figure 6 Taking the context-driven video dubbing process shown as an example, assuming the video editing model is DiT, for the sample video pair For video (i.e., the first sample video) Noise is added to obtain the video. To achieve positive dissemination, using video (i.e., the second sample video) is the video to be edited, with audio... Using the first sample speech as a constraint, the video is processed through DiT. ,video And voice Processing is performed to obtain the predicted video. Predicting Videos (i.e., the recovered video) ), calculate and predict video With video If at least one loss value is selected, and the desired convergence is not achieved for any of the selected loss values, the DiT model parameters are updated, and the process continues until the predicted video is obtained. With video If at least one loss value between the two values reaches the desired convergence state, it indicates that the predicted video... With video If the results are identical or nearly identical, DiT training is complete, meaning the video editing model training is complete.
[0154] Based on the training principles described above, the following section will take training a video editing model based on a sample video pair as an example, and introduce the processing of the sample video pair and the training process of the video editing model in conjunction with the following steps from 2021 to 2023.
[0155] In step 2021, the computing device extracts the video features of the first sample video, the video features of the second sample video, and the speech features of the first sample speech.
[0156] Here, the first sample video, the second sample video, and the first sample speech belong to the same video sample pair. That is, the first sample speech is the speech of the first object in the first sample video. In the first sample video and the second sample video, the lip movements of the first object are different, but all other features are the same except for the lip movements of the first object. The first object is still the dubbing object.
[0157] The video features of the first sample video represent the first sample video. The video features of the first sample video can be the latent features of the first sample video. The computing device can extract the latent features of the first sample video through VAE.
[0158] The video features of the second sample video represent the second sample video, and these video features can be latent features of the second sample video. Computing devices can extract these latent features from the second sample video using a Video Imager (VAE).
[0159] The speech features of the first sample speech characterize the first sample speech, and the computing device can extract the speech features of the first sample speech through an acoustic encoder.
[0160] In step 2022, the computing device adds noise to the video features of the first sample video based on the noise coefficients to obtain the noisy video features of the first sample video.
[0161] Step 2022 involves adding noise to the first sample video. The video features of the noisy first sample video represent the noisy first sample video. Noise coefficients are used to add noise to the video features of the first sample video. The noise coefficients are located in a target numerical range, which is a range from 0 to 1.
[0162] To gradually add noise to the video frames of the first sample video until it approaches pure noise, multiple noise-adding coefficients can be selected from the target value range through random sampling. Based on the Gaussian distribution and these multiple noise-adding coefficients, the video features of the first sample video are noise-added to different degrees, thus obtaining multiple noisy video features of the first sample video. For each noisy video feature of the first sample video, the video features of the second sample video and the speech feature are used to train the video editing model once, until the video editing model is trained.
[0163] In some embodiments, the video editing model includes a global sub-model, a lip-sync sub-model, and an identity texture sub-model. The global sub-model is used to maintain the global structure of the video frames in the second sample video unchanged. The global structure includes the environment in which the first object is located and the pose of the first object in the video frame. The lip-sync sub-model is used to edit the lip-sync of the first object in the second sample video, and the identity texture sub-model is used to maintain the facial identity information and facial texture information of the first object in the second sample video unchanged.
[0164] The global sub-model, lip-shape sub-model, and identity texture sub-model can be trained using video features from first sample videos with different levels of noise. The video features from first sample videos with different levels of noise are obtained by adding noise based on noise coefficients in different intervals.
[0165] Based on this, the target numerical range includes a first sub-range, a second sub-range, and a third sub-range. The values in the first, second, and third sub-ranges decrease sequentially; that is, the value in the first sub-range is greater than the value in the second sub-range, and the value in the second sub-range is greater than the value in the third sub-range. The width of the first, second, and third sub-ranges is not limited. The noise figure is derived from the target numerical range, and the first, second, and third sub-ranges can be respectively a high noise figure range, a middle noise figure range, and a low noise figure range.
[0166] The computing device adds noise to the video features of the first sample video to different degrees based on the noise coefficients of different intervals in the first, second, and third sub-intervals.
[0167] For example, the computing device obtains a first noise coefficient from a first sub-interval, a second noise coefficient from a second sub-interval, and a third noise coefficient from a third sub-interval. Based on a Gaussian distribution and the first noise coefficient, it adds noise to the video features of the first sample video to obtain a first noisy video feature. Based on a Gaussian distribution and the second noise coefficient, it adds noise to the video features of the first sample video to obtain a second noisy video feature. Based on a Gaussian distribution and the third noise coefficient, it adds noise to the video features of the first sample video to obtain a third noisy video feature. The first noisy video feature, the second noisy video feature, and the third noisy video feature are the video features of the first sample video in the high-noise region, the medium-noise region, and the low-noise region, respectively.
[0168] It should be understood that multiple first noise coefficients can be obtained from the first sub-interval. Based on the multiple first noise coefficients, the video features of the first sample video are denoised respectively to obtain multiple first denoised video features of the first sample video. Similarly, multiple second denoised video features and multiple third denoised video features of the first sample video can be obtained in a similar manner.
[0169] In step 2023, the computing device trains the video editing model based on the speech features of the first sample speech, the video features of the second sample video, and the video features of the first sample video after adding noise.
[0170] The computing device iteratively trains the video editing model based on the speech features of the first sample speech, the video features of the second sample video, and the video features of the first sample video after adding noise.
[0171] In some embodiments, during any training iteration, the computing device concatenates the video features of the second sample video and the video features of the first sample video after adding noise, at the frame level, to obtain concatenated features. These concatenated features are used as the primary input to the model, with speech features as the conditional input. The concatenated features and speech features are then input into the video editing model. The video editing model uses the input speech features as constraints to perform feature fusion on the concatenated features and speech features, obtaining fused features. Based on these fused features, a predicted video is output. This predicted video is the first sample video predicted during this training iteration (i.e., the edited second sample video).
[0172] Taking the video editing model DiT as an example, such as Figure 7 As shown, for sample video pairs ,video The first sample video, video For the second sample video, audio The first sample of audio was taken from the video. For reference video (i.e., the video to be edited), with audio For target speech ,video For target video The computing device extracts the reference video via VAE. Potential characteristics and target video Potential characteristics The target speech is extracted using an acoustic encoder. The speech features are determined. The computing device selects noise figures from the target numerical range, and based on a Gaussian distribution and the selected noise figures, analyzes the latent features. Add noise to obtain the latent features after noise addition. The computing device, at the frame dimension, processes latent features. and latent features after adding noise The stitching process yields stitching feature 71, which is used to spatially align the target video. and masked video Background information (including scene and identity information, etc.). The concatenated feature 71 is used as the main input to the model, and the target speech... The speech features are used as the conditional input to the model. The splicing features 71 are input into the repair layer in DiT to process the target speech. The speech features are input into the cross-attention layer of DiT. Each layer in DiT processes the concatenated features 71 and the target speech. Processing of speech features and Figure 3 The various layers of DiT splicing features 31-32 and the target speech The processing method for speech features is the same, and will not be repeated here. Figure 7 In this process, DiT outputs a predicted video based on the input splicing features 71 and speech features.
[0173] When it needs to be explained, Figure 3 , Figure 5 and Figure 7 These examples all illustrate DiT with a single transformer. In some embodiments, when DiT includes multiple transformers, the target speech... The speech features are injected into the cross-attention layers of each transformer. Figure 3 , Figure 5 and Figure 7 The cross-attention layer shown inputs the obtained fused features into the next transformer, and similar processing is performed in the next transformer until the output layer of DiT inputs the fused features output by the cross-attention layer in the last transformer. The output layer generates a predicted video based on the input fused features and outputs the predicted video.
[0174] During any training iteration, after obtaining the predicted video from the video editing model's output, a flow matching loss function is used to calculate the flow matching loss value between the predicted video and the first sample video. This flow matching loss value indicates the degree of difference between the lip-sync flow of the first object in the predicted video and the lip-sync flow of the first object in the first sample video. If the flow matching loss value does not converge to the target loss range, the model parameters of the video editing model are updated, and the updated model is trained again until the flow matching loss value converges to the target loss range during a training iteration. The video editing model is then considered trained successfully. The trained model learns the mapping from the lip-sync flow of the first object in the second sample video to the lip-sync flow of the first object in the first sample video. This mapping is also equivalent to the mapping from a noise distribution (such as a Gaussian distribution) to the true data distribution of the first sample video.
[0175] Compared to the basic paradigm of mouth mask restoration, the video editing model receives a complete second sample video that is perfectly aligned with the first sample video. Using the second sample video as a reference video, it provides complete spatial context. Therefore, when training the video editing model, there is no longer a trade-off between "filling in missing parts" and "aligning lip movements." The learning objective of the video editing model is pure, and the training is more stable. Correspondingly, during inference, the video editing model does not rely on masks obtained from face detection and face parameter models. This significantly reduces negative phenomena such as lip leakage, visual artifacts, and identity drift, while simultaneously improving the success rate of complex scenes such as occlusion, profile views, and stylization.
[0176] When the video editing model includes a global sub-model, a lip-sync sub-model, and a texture sub-model, the computing device trains the global sub-model, the lip-sync sub-model, and the texture sub-model based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise.
[0177] The training process of these three sub-models will be introduced below, in conjunction with (2.1) to (2.2).
[0178] (2.1) Training process of global sub-model The computing device inputs the video features of the second sample video and the video features of the first sample video after adding noise into the global sub-model. Based on the video features of the second sample video and the video features of the first sample video after adding noise, the global sub-model predicts the global structure of the first sample video frame in the first sample video, so that the predicted global structure is close to the global structure of the first sample video frame, and outputs a first predicted video that can reflect the predicted global structure.
[0179] In this model, the video features of the first sample video input to the global sub-model after adding noise can be the first noisy video features of the first sample video. The first predicted video is the video output by the global sub-model. During the training of the global sub-model, the video features of the second sample video are used to provide a reference for the global structure of the first sample video frame of the first sample video.
[0180] The computing device iteratively trains the global sub-model based on the video features of the second sample video, the features of the first noisy video, and the first sample video. For example, in any training iteration, the video features of the second sample video and the features of the first noisy video are input into the global sub-model. For instance, at the frame level, the video features of the second sample video and the features of the first noisy video are concatenated, and the concatenated features are input into the global sub-model. Based on the video features of the second sample video and the features of the first noisy video, the global sub-model predicts the global structure of the first sample video frame of the first sample video and outputs the first predicted video based on the predicted global structure. A flow matching loss function is used to calculate the flow matching loss value between the first predicted video and the first sample video. This flow matching loss value indicates the degree of difference between the global structure flow of the first predicted video and the global structure flow of the first sample video. The global structure flow of any video refers to the variation pattern of the global structure of the video frames within the video. If the flow matching loss value does not converge to the second loss value interval (i.e., the expected convergence state), the model parameters of the global sub-model are updated, and the next training process begins, until the flow matching loss value between the first predicted video and the first sample video converges to the second loss value interval during a certain training process, at which point the global sub-model training is complete. Therefore, during training, the global structure predicted by the global sub-model for the first sample video frame gradually approaches the actual global structure of that first sample video frame. The trained global sub-model can learn the mapping from the global structure flow of the second sample video to the global structure flow of the first sample video. For the video to be edited, the trained global sub-model, based on the learned mapping, can maintain the global structure of the video frames in the video to be edited unchanged.
[0181] by Figure 8 Taking the diffusion time-step adaptive multi-stage learning process as an example, in the training process of the video editing model, the noise coefficient is the time step t, and the distribution of the time step t is Gaussian. Based on the Gaussian distribution and t selected from the high noise coefficient interval (i.e., the first noise coefficient), the video... The video features are denoised to obtain the first noisy video feature, which represents the noisy video. The video after adding noise The video frame in the text is video frame 81. Video frame 81 can at least represent the video. The global structure of the mid-range video frame. The global sub-model is a linear layer. Based on the features of the first noisy video and the video... Based on the video features, the global sub-model is iteratively trained until the first predicted video and video output by the global sub-model 42 are obtained. When the flow matching loss between the two models converges to the expected convergence state, the training of the global sub-model ends.
[0182] (2.2) Training process of lip shape sub-model The computing device inputs the speech features of the first sample speech, the video features of the second sample video, and the video features of the first sample video after adding noise into the lip-sync sub-model. Based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise, the lip-sync sub-model predicts the lip-sync of the first object in the first sample video so that the predicted lip-sync is close to the lip-sync of the first object in the first sample video. The lip-sync of the first object in the second sample video is edited to the predicted lip-sync, and the second predicted video is output.
[0183] Specifically, the video features of the first sample video input to the lip-sync sub-model after adding noise can be the second noisy video features of the first sample video. The second predicted video is the video output by the lip-sync sub-model, and the second predicted video is the second sample video after lip-sync editing.
[0184] The computing device iteratively trains a lip-sync sub-model based on the speech features of a first sample speech, the video features of a second sample video, the features of a second noisy video, and the first sample video. For example, in any training iteration, with the speech features as constraints, the video features of the second sample video, the features of the second noisy video, and the speech features are input into the lip-sync sub-model. For instance, at the frame level, the video features of the second sample video and the features of the second noisy video are concatenated, and the concatenated features are input into the lip-sync sub-model. The lip-sync sub-model, with the input speech features as constraints, predicts the lip-sync of a first object in the first sample video based on the input features, edits the lip-sync of the first object in the second sample video to the predicted lip-sync, and outputs the edited second sample video (i.e., the second predicted video). The lip-syncing discrimination network calculates the lip-syncing loss value of the first object between the second predicted video and the first sample video. If this lip-syncing loss value does not converge to the third loss value interval (i.e., the expected convergence state), the model parameters of the lip-syncing sub-model are updated, and the training process continues until, during a certain training process, the lip-syncing loss value of the first object between the second predicted video and the first sample video converges to the third loss value interval, at which point the lip-syncing sub-model training is complete. Therefore, during the training process of the lip-syncing sub-model, the lip-syncing of the first object predicted by the lip-syncing sub-model gradually approaches the lip-syncing of the first object in the first sample video. The trained lip-syncing sub-model can learn the mapping from the lip-syncing motion flow of the first object in the second sample video to the actual lip-syncing motion flow of the first sample speech (i.e., the lip-syncing motion flow of the first object in the first sample video). For the target speech, the trained lip-syncing sub-model, based on the learned mapping, can edit the lip-syncing of the dubbing object in the video to be edited into lip-syncing aligned with the target speech.
[0185] by Figure 8 For example, during the training of the lip-shape sub-model, based on the Gaussian distribution and t (i.e., the second noise coefficient) selected from the intermediate noise coefficient range, the video... The video features are denoised to obtain a second noisy video feature, which represents the noisy video. The video after adding noise The video frame in the text is video frame 82, and video frame 82 must at least show the lip movements of the first object. Based on the second noisy video features, the video... Video features and voice The lip-shape sub-model is iteratively trained until the second predicted video and speech output by lip-shape sub-model 44 are obtained. When the lip sync loss between the two converges to the expected convergence state, the training of the lip shape sub-model ends.
[0186] (2.3) Training process of identity texture sub-model The computing device inputs the video features of the second sample video and the video features of the first sample video after adding noise into the identity texture sub-model. Based on the video features of the second sample video and the video features after adding noise, the identity texture sub-model predicts the facial identity information and facial texture information of the first object in the first sample video, so that the predicted facial identity information is close to the facial identity information and the predicted facial texture information is close to the facial texture information of the first object in the first sample video. The sub-model then outputs a third predicted video that reflects the predicted facial identity information and facial texture information.
[0187] In this model, the video features of the first sample video input to the identity texture sub-model after adding noise can be the third noisy video features of the first sample video. The third predicted video is the video output by the identity texture sub-model. During the training of the identity texture sub-model, the video features of the second sample video are used to provide a reference for the facial identity information and facial texture information of the first object in the first sample video. The third, second, and third noisy video features are obtained based on noise coefficients in different intervals, so that different sub-models in the video editing model can be trained using different noisy video features of the first sample video, enabling them to learn the mapping from different noise distributions to the real data distribution.
[0188] The computing device iteratively trains the identity texture sub-model based on the video features of the second sample video, the features of the third noisy video, and the first sample video. For example, in any training process, the video features of the second sample video and the features of the third noisy video are input into the identity texture sub-model. For instance, at the frame level, the video features of the second sample video and the features of the third noisy video are concatenated, and the concatenated features are input into the identity texture sub-model. Based on the input features, the identity texture sub-model predicts the facial identity information and facial texture information of the first object in the first sample video, and outputs a third predicted video based on the predicted facial identity information and facial texture information of the first object. The identity loss value and texture loss value of the first object between the third predicted video and the first sample video are calculated. The identity loss value indicates the degree of difference in the facial identity of the first object between the third predicted video and the first sample video. This identity loss value can be the cosine similarity between the face of the first object in the third predicted video and the face of the first object in the first sample video. This cosine similarity can be obtained by processing the third predicted video and the first sample video using ArcFace. The texture loss value of the first object between the third predicted video and the first sample video indicates the degree of difference in the facial texture of the first object between the third predicted video and the first sample video. The texture loss value can be obtained by processing the third predicted video and the first sample video using CLIP.
[0189] If the identity loss value does not converge to the fourth loss value interval (i.e., the expected convergence state) and / or the texture loss value converges to the fifth loss value interval (i.e., the expected convergence state), the model parameters of the identity texture sub-model are updated, and the next training process begins. This continues until, in a certain training process, the identity loss value converges to the fourth loss value interval and the texture loss value converges to the fifth loss value interval, at which point the identity texture sub-model training is complete. Therefore, during the training process of the identity texture sub-model, the facial identity information and facial texture information of the first object predicted by the identity texture sub-model gradually approach the facial identity information and facial texture information of the first object in the first sample video frame. Thus, after training, the identity texture sub-model can learn the mapping from the facial identity of the first object in the second sample video to the facial identity of the first object in the second sample video, as well as the mapping from the facial texture of the first object in the second sample video to the facial texture of the first object in the second sample video. For the video to be edited, the trained identity texture sub-model, based on the learned mapping, can maintain the facial identity of the voice-over subject in the video to be edited, making the facial texture of the voice-over subject more natural.
[0190] by Figure 8 For example, during the training of the identity texture sub-model, based on the Gaussian distribution and t (i.e., the third noise coefficient) selected from the low noise coefficient range, the video... The video features are denoised to obtain a third noisy video feature, which represents the noisy video. The video after adding noise The video frame in the text is video frame 83. Video frame 83 indicates that the video... The facial identity information and facial texture information of the first object in the video. Based on the features of the third noisy video and the video... The video features are used to iteratively train the identity texture sub-model until the third predicted video output by the identity texture sub-model matches the video. The identity and texture loss values converge to the expected convergence state, and the training of the identity and texture sub-model ends.
[0191] In some embodiments, taking DiT as an example of a video editing model, a global sub-model to be trained is first added to each self-attention layer in DiT to obtain DiT including the global sub-model. Based on the video features of the second sample video, the features of the first noisy video, and the first sample video, the global sub-model in DiT is iteratively trained. In each training process, the video features of the second sample video and the features of the first noisy video are concatenated and input into DiT. After the global sub-model is trained, a lip-sync sub-model is added to each cross-self-attention layer in DiT. Based on the speech features of the first sample speech, the video features of the second sample video, the features of the second noisy video, and the first sample video, the lip-sync sub-model in DiT is iteratively trained. In each training process, the video features of the second sample video and the features of the second noisy video are concatenated and input into DiT. The speech features are used as the termination condition and input into each cross-self-attention layer. After the lip-sync sub-model is trained, the identity texture sub-model is added to each self-attention layer in DiT. Based on the video features of the second sample video, the features of the second noisy video, and the first sample, the identity texture sub-model in DiT is iteratively trained. In each training process, the video features of the second sample video and the features of the third noisy video are concatenated and input into DiT. When the identity texture sub-model training is complete, DiT training is complete. The trained DiT includes the global sub-model, the lip-sync sub-model, and the identity texture sub-model. The trained DiT is thus the trained video editing model, achieving a multi-stage learning approach that adapts to the diffusion denoising time step.
[0192] When training the identity texture sub-model in DiT, DiT includes a lip-sync sub-model. When training the lip-sync sub-model, there is no input of speech features. The lip-sync sub-model is equivalent to processing the video features under silent conditions. Therefore, in some embodiments, when training the lip-sync sub-model in DiT, the cross self-attention layer can be turned off to a certain extent (such as turning off the lip-sync sub-model in the cross self-attention layer) to avoid triggering the update of the lip-sync sub-model's model parameters during the training of the identity texture sub-model, and to avoid causing reverse interference to the lip-sync sub-model's model parameters.
[0193] In some embodiments, in the multi-stage learning approach with time-step adaptive diffusion denoising, LoRA is enabled during the training of the lip shape sub-model to reduce the incremental parameters that need to be learned in DiT during the training process, thereby reducing the computational requirements during the training process. This improves training efficiency while avoiding damage to the existing generation capabilities of the base model. The LoRA enabled during the training of the lip shape sub-model can also be referred to as a lip shape LoRA expert.
[0194] Enabling LoRA during the training of the identity texture sub-model reduces the incremental parameters that need to be learned in DiT during training, lowering the computational requirements. This improves training efficiency while avoiding disruption of the existing generative capabilities of the base model. LoRA enabled during identity texture sub-model training can also be referred to as a texture LoRA expert or an identity texture expert.
[0195] The above-mentioned phased training of the global sub-model, lip-shape sub-model, and identity texture sub-model can naturally decouple the learning objectives of the video editing model at the three levels of "structure-lip-shape-texture" in the time step (i.e. noise coefficient) dimension, avoiding the situation where a single network carries mutually restrictive objectives in the same time step (such as the numerical range from 0 to 1).
[0196] Compared to the mask patching paradigm, the above-mentioned "first align, then edit" training scheme provides a complete and aligned full-frame context through the video to be edited. The video editing model's view during training is consistent with the conditions during inference, avoiding the conflict between the two objectives of "missing parts patching" and "voice alignment" in the same network. This reduces the learning difficulty from the source and improves identity consistency and lip-sync accuracy.
[0197] The training method provided in this embodiment of the disclosure, since the difference between the video to be edited (i.e., the second sample video) and the target video (i.e., the first sample video) is only the lip movements of the dubbing object (i.e., the first object), provides complete spatiotemporal context information of the dubbing object in the video to be edited. Therefore, when training the video model, the target video is used as the desired output and the target speech of the dubbing object in the target video is used as the constraint. The video editing model only edits the lip movements of the dubbing object in the video to be edited based on the lip movements of the dubbing object in the target video, without editing any parts other than the mouth of the dubbing object. This ensures that only the lip movements of the dubbing object change in the edited video, while the environment, posture, and identity of the dubbing object remain unchanged. Thus, even for complex scenarios such as multiple postures of the dubbing object and changing environmental scenes, the video editing model can maintain stable, natural, and controllable lip movement editing quality. Therefore, the video editing model trained using the training method provided in this embodiment of the disclosure produces high-quality visual dubbing results.
[0198] Next, combined Figure 9 This section introduces the process of using the trained video editing model.
[0199] Figure 9 This is a flowchart illustrating a video editing method according to an exemplary embodiment, see [link to flowchart]. Figure 9The video editing method is applied to a computing device, which can be the terminal 101 mentioned above. The video editing method includes the following steps.
[0200] In step 901, the computing device acquires a first video and a second voice that is different from the first voice content, wherein the first video includes a fourth object and the first voice is the voice of the fourth object in the first video.
[0201] The first video is the video to be edited. During the reasoning stage, the video to be edited is also the reference video; therefore, the first video is the reference video. The first video includes multiple video frames, and each video frame includes a fourth object, which is the voice-over object in the first video. The second voice is used to provide the voice-over for the fourth object, and the second voice is the target voice. The content of the second voice is different from that of the first voice.
[0202] The computing device acquires a first video and a second voice input from the user, or the first video and the second voice are specified by the user, and the computing device acquires the first video and the second voice specified by the user.
[0203] In step 902, the computing device uses the first video as the video to be edited and the second speech as a constraint. Through the video editing model, it edits the lip movements of the fourth object in the first video to obtain the second video. The second video is the first video after lip movement editing. The lip movements of the fourth object in the second video match the second speech. In the first video and the second video, the lip movements of the fourth object are different, but all features other than the lip movements of the fourth object are the same.
[0204] The video editing model is based on the above. Figure 2 The method embodiment shown illustrates a trained video editing model. The second video is the predicted video output by the video editing model based on the first video and the second speech. The second video includes multiple video frames, with different video frames in the second video corresponding to different video frames in the first video. In any video frame in the second video and the corresponding video frame in the first video, the lip movements of the fourth object are different, while all other features are the same except for the lip movements of the fourth object. The lip movement flow of the first object in the second video is the same as or nearly the same as the temporal lip movement flow of the second speech.
[0205] Because the video editing model has learned the mapping between the video to be edited and the target video when matching target speech to an object in the video to be edited, when using the first video as the video to be edited and the second speech as the constraint (i.e., the second speech is the target speech), the video editing model, based on its learned mapping, can edit only the lip movements of the fourth object in the first video, thus editing the first video into the second video. The lip movement editing is unaffected by other factors, thus improving the accuracy of the edited lip movements and ensuring precise alignment with the second speech. During the editing process, because only the lip movements of the fourth object are edited, all features in the edited first video except for the fourth object remain unchanged. That is, the second video can completely inherit all features from the first video except for the fourth object, ensuring that the context of the first object in the second video and the second video remains aligned. Therefore, relative to the first video, the identity of the third object in the second video remains stable, and the environment of the fourth object in the second video remains unchanged. Thus, the video editing model has good robustness to changes in lighting, occlusion, and other environmental changes in the scene.
[0206] Taking the video editing model DiT as an example, and using the first video as the reference video. (i.e., the video to be edited), with the second voice as the target voice. Computing devices extract reference videos Video features and target speech The speech features.
[0207] Among them, reference videos Video feature representation reference video Reference video Video features can be reference videos Potential characteristics Computing devices can extract reference videos via VAE. Potential characteristics Target speech Speech features representing target speech The computing device extracts the target speech through an acoustic encoder. The speech features.
[0208] The computing device selects a noise figure from the first sub-interval (i.e., the high noise figure interval) within the target numerical interval; for example, the selected noise figure can be 1. Of course, the selected noise figure can also be less than 1. Based on the Gaussian distribution and the selected noise figure, the computing device processes the reference video... Video features (such as latent features) Noise is added to the reference video to obtain the fourth noisy video feature, which is the reference video after noise addition. The video features. The computing device will use the target speech. Using the speech features as constraints, and the fourth noisy video features as model input, the fourth noisy video features and the target speech are combined. The speech features are input into the video editing model. For example, the fourth noisy video feature is input into the input layer of DiT, and the speech feature is input into the DiT cross-attention layer. The layers in DiT process the input features until the output layer of DiT outputs the predicted video based on the input features.
[0209] Because the video editing model learns the mapping from the noise distribution to the real data distribution of the first sample video, based on this mapping, each layer in DiT gradually denoises the features of the fourth noisy video during processing, thus improving the target speech. Under the constraints of speech features, the difference between the predicted video output by DiT and the first video is only in the lip movements of the fourth object, and the lip movements of the fourth object in the predicted video can be aligned with the second speech. Therefore, the predicted video is also the second video.
[0210] When the video editing model includes a global sub-model, a lip-sync sub-model, and an identity texture sub-model, these three sub-models can work together to gradually denoise the fourth noisy video feature.
[0211] by Figure 8 For example, the coefficient that represents the noise intensity in a video feature is called the noise coefficient corresponding to the video feature. When the fourth noisy video feature is input into DiT, according to the noise coefficient corresponding to the fourth noisy video feature (e.g., t), the global sub-model, lip-sync sub-model, and identity texture sub-model in DiT are sequentially activated. These sub-models are then used to denoise the fourth noisy video feature sequentially. For instance, initially, since the noise coefficient t corresponding to the fourth noisy video feature is in the high noise coefficient range, the global sub-model in DiT is used to denoise the fourth noisy video feature, reducing the noise coefficient to the intermediate noise coefficient range, thus obtaining the first video feature. The first video feature is the fourth noisy video feature denoised by the global sub-model. The first video feature represents the first intermediate video, and the global structure of the video frames in the first intermediate video is consistent with the characteristics of the reference video. The global structure of the intermediate video frames is the same. Since the noise coefficient corresponding to the first video feature is located in the intermediate noise coefficient range, the lip-sync sub-model in DiT is enabled. Using the input speech features as constraints, the lip-sync sub-model is used to further denoise the first video feature, reducing its noise coefficient to the low noise coefficient range, resulting in the second video feature. The second video feature is the first video feature after denoising using the lip-sync sub-model. The second video feature represents the second intermediate video, which is the video after editing the lip-sync of the fourth object in the first intermediate video to align with the target speech. Since the noise coefficient corresponding to the second video feature is located in the low noise coefficient range, the identity texture sub-model in DiT is enabled. The identity texture sub-model is used to further denoise the second video feature, reducing its noise coefficient to 0, resulting in the third video feature. This achieves step-by-step denoising of the fourth noisy video feature according to time steps. The third video feature represents the predicted video, and the input layer of DiT outputs the predicted video based on the third video feature.
[0212] It should be understood that when DiT includes multiple transformers, the self-attention layer in each transformer includes a global sub-model and an identity texture sub-model, and the cross-self-attention layer in each transformer includes a lip-sync sub-model. The noise coefficient corresponding to the fourth noisy video feature requires denoising by at least one global sub-model in a transformer to reduce the noise coefficient to the intermediate noise coefficient range; the noise coefficient corresponding to the first video feature requires denoising by at least one lip-sync sub-model in a transformer to reduce the noise coefficient to the low noise coefficient range; the noise coefficient corresponding to the second video feature requires denoising by at least one lip-sync sub-model in a transformer to reduce the noise coefficient to the low noise coefficient range.
[0213] For example, suppose DiT includes 10 transformers, designated transformer1-10. The patching layer in DiT processes the input fourth noisy video feature to obtain the patched video feature. This patched video feature is then input to the self-attention layer in transformer1. Besides using a self-attention mechanism to process the input video feature, the self-attention layer in transformer1 also denoises its processed video feature through a global sub-model. The denoised video feature is then input to the next layer, and so on. Finally, the cross-self-attention layer in transformer1 processes the input video feature from the previous layer to obtain the processed video feature. This processed video feature is then input to the self-attention layer in transformer2. If the noise coefficient corresponding to the processed video feature in transformer1 is still in the high noise coefficient range, transformer2 processes the processed video feature in a similar manner to transformer1 to obtain the processed video feature. This processed video feature is then input to... The self-attention layer in transformer 3; assuming the noise coefficient corresponding to the video features processed by transformer 2 is in the middle noise coefficient range, the self-attention layer in transformer 3 uses a self-attention mechanism to process the input video features, and inputs the processed video features to the next layer, until the cross self-attention layer in transformer 3 receives the video features input from the previous layer. In addition to processing the input video features using a cross attention mechanism, the cross self-attention layer also uses speech features as constraints and a lip-shape sub-model to denoise its own processed video features, obtaining the video features processed by transformer 3, which are then input into the self-attention layer in transformer 4; assuming the noise coefficient corresponding to the video features processed by transformer 3 is still in the middle noise coefficient range, transformer 4 processes the video features processed by transformer 3 in a similar way to transformer 3, obtaining the video features processed by transformer 4, which are then input into the self-attention layer in transformer 5;Assuming the noise coefficient of the video features processed by Transformer 4 is in the low noise coefficient range, the self-attention layer in Transformer 5, in addition to processing the input video features using a self-attention mechanism, also denoises its own processed video features through an identity texture sub-model. The denoised video features are then input to the next layer, and so on. Finally, the cross-self-attention layer in Transformer 5 processes the input video features from the previous layer to obtain the video features processed by Transformer 5. These processed video features are then input to Transformer 6. Assuming the noise coefficient of the video features processed by Transformer 5 is still in the low noise coefficient range and not zero, Transformer 6 processes the video features processed by Transformer 5 in a similar manner to Transformer 5, resulting in the video features processed by Transformer 6. Assuming the noise coefficient of the video features processed by Transformer 6 is zero, these processed video features are input to the input layer. The input layer then outputs the predicted video based on the input video features.
[0214] It should be understood that the above explanation uses the example of using global sub-models from two transformers to reduce the noise coefficient of video features to the intermediate noise coefficient range. However, the number of global sub-models required to reduce the noise coefficient from the high noise coefficient range to the intermediate noise coefficient range is not limited to two. Similarly, the above explanation uses the example of using lip-sync sub-models from two transformers to reduce the noise coefficient of video features from the intermediate noise coefficient range to the low noise coefficient range. However, the number of lip-sync sub-models required to reduce the noise coefficient from the high noise coefficient range to the intermediate noise coefficient range is not limited to two. The above explanation also uses the example of using lip-sync sub-models from two transformers to reduce the noise coefficient of video features from the low noise coefficient range to 0. However, the number of lip-sync sub-models required to reduce the noise coefficient from the low noise coefficient range to 0 is not limited to two. When the Transformer uses self-attention or cross-self-attention mechanisms to process video features, it also performs denoising on the video features.
[0215] The method provided in this disclosure, when dubbing a voice-over object in a video to be edited, only edits the lip movements of the voice-over object in the video to be edited. This ensures that the environmental scene, posture, and identity of the voice-over object in the video to be edited are preserved. Thus, for videos to be edited in complex scenarios such as multiple postures of the voice-over object and changing environmental scenes, the video editing model can maintain stable, natural, and controllable lip-sync editing quality, thereby improving the quality of the visual dubbing results of the video editing model.
[0216] The self-promoting training paradigm of "first construct, then edit" and the synergistic effect of time-step adaptive three-stage learning provided in this disclosure form a low-latency visual dubbing generation capability. This enables the trained video editing model to be used in scenarios such as digital human interaction, character-driven content, dialogue video generation, course instruction, and localized dubbing. In these application scenarios, the video editing model achieves high-precision lip-sync with the target speech and a stable and consistent appearance without changing the character's identity or the scene background. It also maintains a high success rate under complex conditions such as occlusion, complex lighting changes, side profiles, and large head movements.
[0217] On publicly available contextual dubbing benchmark sets, the video editing model trained according to the training method provided in this disclosure achieves significant improvements over the best publicly available baseline in both reference- and non-reference-based visual quality metrics, lip-sync, and identity preservation. Tables 1 and 2 below show the relative improvement compared to the strongest baseline (the best alternative method for different metrics). In Tables 1 and 2, "↓ / ↑" indicates that smaller / larger values are better, respectively.
[0218] With reference to visual quality (Ref.): The video editing model achieved scores of 9.351 and 214.298 on the FID and FVD perceptual metrics, respectively, representing a decrease of 31.25% and 19.15% compared to the optimal baseline (LatentSync: FID=13.602, FVD=265.057); on the NIQE, it decreased from 6.113 to 5.782, a reduction of 5.41%. This indicates that after obtaining the complete context, the overall perceptual quality and temporal consistency of the video are significantly improved.
[0219] Visual Quality (No Ref.): On the BRISQUE metric, it decreased from the best baseline of 38.990 to 29.870, a drop of 23.39%; on the HyperIQA (higher is better) metric, it improved from 44.826 to 51.960, an improvement of 15.91%. The no-ref. metrics also validate the superior naturalness and sharpness of the video editing model trained in this publication under complex lighting and texture conditions.
[0220] Lip Sync: The Sync-C metric improved from 6.282 to 7.282, a 15.92% increase. This directly reflects the effectiveness of the strategy of "contextual completeness + mid-noise lip sync expert + SyncNet supervision".
[0221] Identity Preservation: Based on the CSIM and CLIPS identity / semantic consistency metrics, the improvement was 6.12% (from 0.801 to 0.850) and 3.33% (from 0.812 to 0.839), respectively. Frame-level stitching under full reference and low-noise texture experts jointly ensured the stable inheritance of identity details.
[0222] Generation Success Rate: Without selecting samples, the success rate increased from the best alternative method of 71.82% to 96.36%, a relative increase of 34.17% and an absolute increase of 24.54 percentage points, demonstrating significant robustness under complex conditions such as occlusion, side profile, and large head movements.
[0223]
[0224] Table 1
[0225] Table 2 Among them, Wav2Lip is an audio-lip-sync model, VideoReTalking is a video retelling model, TalkLip is a speech lip-sync model, IP-LAP is an identity-preserving speaking face generation model, Diff2Lip is a diffusion-based audio-lip-sync model, MuseTalk is a high-quality real-time lip-sync model, LatentSync is a latent spatial synchronization lip-sync model, Ours-generator is a video editing model provided in this embodiment, and Ours-generator is a video generation model provided in this embodiment.
[0226] In the disclosed embodiments, the mouth shape can also be referred to as the lip shape, the face can also be referred to as the face, the video generation model can also be referred to as the video generator, and the video editing model can also be referred to as the video editor.
[0227] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.
[0228] Figure 10 This is a structural block diagram of a training device for a video editing model according to an exemplary embodiment, with reference to... Figure 10 , Figure 10The training device 1000 for the video editing model shown includes: The first acquisition unit 1001 is configured to execute a first sample video, a second sample video, and a first sample speech, wherein the first sample video and the second sample video both include a first object, and the first sample speech is the speech of the first object in the first sample video. In the first sample video and the second sample video, the lip movements of the first object are different, and all features other than the lip movements of the first object are the same. The first training unit 1002 is configured to train a video editing model using the second sample video as the video to be edited, the first sample video as the desired output, and the first sample speech as the constraint condition. During training, the video editing model edits the lip movements of the first object in the first sample video to the second sample video, based on the lip movements of the first object in the first sample video, so that the lip movements of the first object in the second sample video are closer to the lip movements of the first object in the first sample video.
[0229] Optionally, the first training unit 1002 includes: The extraction subunit is configured to extract video features of the first sample video, video features of the second sample video, and speech features of the first sample speech. The noise-adding subunit is configured to perform noise addition on the video features of the first sample video based on noise-adding coefficients. The training subunit is configured to train the video editing model based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise.
[0230] Optionally, the video editing model includes a global sub-model, a lip-sync sub-model, and a texture sub-model. The global sub-model is used to keep the global structure of the video frames in the second sample video unchanged. The lip-sync sub-model is used to edit the lip-sync of the first object in the second sample video. The identity texture sub-model is used to keep the facial identity information and facial texture information of the first object in the second sample video unchanged. The training sub-unit is configured to train the global sub-model, the lip-shape sub-model, and the texture sub-model based on the speech features, the video features of the second sample video, and the video features of the first sample video after adding noise.
[0231] Optionally, the training subunit is configured to perform: The video features of the second sample video and the video features after adding noise are input into the global sub-model. Based on the video features of the second sample video and the video features after adding noise, the global sub-model predicts the global structure of the first sample video frame in the first sample video so that the predicted global structure is close to the global structure of the first sample video frame, and outputs a first predicted video that can reflect the predicted global structure. The speech features, the video features of the second sample video, and the video features after adding noise are input into the lip-sync sub-model. Based on the speech features, the video features of the second sample video, and the video features after adding noise, the lip-sync sub-model predicts the lip-sync of the first object in the first sample video so that the predicted lip-sync is close to the lip-sync of the first object in the first sample video. The lip-sync of the first object in the second sample video is edited to the predicted lip-sync, and the second predicted video is the second sample video after lip-sync editing. The video features of the second sample video and the video features after adding noise are input into the identity texture sub-model. Based on the video features of the second sample video and the video features after adding noise, the identity texture sub-model predicts the facial identity information and facial texture information of the first object in the first sample video, so that the predicted facial identity information is close to the facial identity information and the predicted facial texture information is close to the facial texture information of the first object in the first sample video. The model then outputs a third predicted video that reflects the predicted facial identity information and facial texture information.
[0232] Optionally, for different sub-models among the global sub-model, the lip-sync sub-model, and the identity texture sub-model, the noise-added video features input to the different sub-models are obtained based on noise coefficients in different intervals.
[0233] Optionally, the training device 1000 for the video editing model further includes: The second acquisition unit is configured to acquire a second sample speech, the content of which is different from that of the first sample speech. The sample processing unit is configured to execute a video generation model to process the first sample video and the second sample speech to obtain the third sample video, wherein the lip movements of the first object in the third sample video match the second sample speech. The third acquisition unit is configured to acquire the second sample video based on the third sample video.
[0234] Optionally, the second sample speech has the same voiceprint and timbre features as the first sample speech.
[0235] Optionally, the sample processing unit is configured to perform processing on the masked video, the first sample video, and the second sample speech through the video generation model, wherein the masked video is the first sample video with a mask added, the mask is used to cover the mouth of the first object in the first sample video but not cover the second object, and the second object is any object that occludes the first object.
[0236] Optionally, the first sample video includes multiple first sample video frames, and the training device 1000 for the video editing model further includes: The mask adding unit is configured to perform the following for any first sample video frame: if the second object does not exist in the first image region of the first sample video frame, add the mask to the entire first image region; if the second object exists in the first image region, add the mask to the region outside the second object in the first image region, thereby obtaining the masked video. The first image region is the image region where the mouth of the first object is located in the first sample video frame.
[0237] Optionally, the first sample video includes multiple first sample video segments, and the second sample audio includes multiple audio segments, with different first sample video segments corresponding to different audio segments; The sample processing unit includes: The sample processing subunit is configured to perform, for any first sample video segment, processing the first sample video segment and the corresponding speech segment through the video generation model to obtain a second sample video segment corresponding to the speech segment, wherein the lip movements of the first object in the second sample video segment match the corresponding speech segment, and the lip movements of the first object are different in the first sample video segment and the second sample video segment, and all features other than the lip movements of the first object are the same. The splicing subunit is configured to splice the second sample video segments corresponding to the plurality of speech segments to obtain the third sample video.
[0238] Optionally, the sample processing subunit is configured to perform: For the i-th video segment among the plurality of first sample video segments, starting from the last video frame of the (i-1)-th video segment among the plurality of first sample video segments, multiple video frames are obtained from the (i-1)-th video segment, where i is an integer greater than or equal to 2; Using the plurality of video frames and the audio segment corresponding to the i-th video segment as constraints, the video generation model processes the i-th video segment, the audio segment corresponding to the i-th video segment, and the plurality of video frames.
[0239] Optionally, the duration of the first sample video is greater than or equal to a first duration, the video generation model is a trained base generation model, and the training device for the video editing model further includes: The fourth acquisition unit is configured to acquire a fourth sample video and a third sample audio, wherein the duration of the fourth sample video is shorter than the first duration, and the content of the third sample audio is different from the audio of the third object in the fourth sample video. The second training unit is configured to train the basic generative model based on the fourth sample video and the third sample speech. During the training process, the third sample speech is used as a constraint and a video frame in the fourth sample video is used as a reference frame to edit the lip movements of the third object in the fourth sample video so that the lip movements of the first object in the fourth sample video are close to the lip movements that match the third sample speech.
[0240] Optionally, the second training unit is configured to perform overfit training on the base generative model based on the fourth sample video and the third sample speech.
[0241] Optionally, the third acquisition unit is configured to perform: The video frames in the third sample video are subjected to multiple illumination enhancements of different degrees to obtain multiple fifth sample videos, which are the third sample videos after illumination enhancement. The second sample video is obtained from the plurality of fifth sample videos.
[0242] Optionally, the third acquisition unit is configured to perform: If the facial similarity of the first object between the third sample video and the first sample video is greater than or equal to a first threshold, and the lip shape similarity of the first object between the third sample video and the first sample video is greater than or equal to a second threshold, then the third sample video is determined as the second sample video. Wherein, the facial similarity indicates the degree of similarity between the face of the first object in the third sample video and the face of the first object in the first sample video, and the lip shape similarity indicates the degree of similarity between the lip shape of the first object in the third sample video and the lip shape of the first object in the first sample video.
[0243] Regarding the training device 1000 in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments of the training method for the video editing model, and will not be elaborated here.
[0244] Figure 11 This is a structural block diagram of a video editing apparatus according to an exemplary embodiment; see reference. Figure 11 , Figure 11 The video editing device 1100 shown includes: The acquisition unit 1101 is configured to acquire a first video and a second voice that is different from the first voice content, wherein the first video includes a fourth object and the first voice is the voice of the fourth object in the first video; Editing unit 1102 is configured to perform editing on the lip movements of the fourth object in the first video using the first video as the video to be edited and the second speech as the constraint, through a video editing model, to obtain the second video; The video editing model is a video editing model trained using the training method for the video editing model provided in the above-described embodiments of this disclosure; The second video is the first video after lip-syncing editing. In the second video, the lip-sync of the fourth object matches the second speech. In the first video and the second video, the lip-sync of the fourth object is different, but all other features are the same except for the lip-sync of the fourth object.
[0245] Regarding the video editing apparatus 1100 in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments of the relevant video editing method, and will not be elaborated here.
[0246] Figure 12 This is a schematic diagram illustrating the structure of a computing device according to an exemplary embodiment, such as... Figure 12 As shown, the computing device 1200 can vary considerably due to differences in configuration or performance. It may include one or more Central Processing Units (CPUs) 1201 and one or more memories 1202. The memory 1202 stores at least one instruction, which is loaded and executed by the processor 1201 to implement the training method or video editing method of the video editing model provided in the various embodiments described above. In some embodiments, the computing device 1200 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The computer device 1200 may also include other components for implementing device functions, which will not be elaborated here.
[0247] In an exemplary embodiment, a computer-readable storage medium including at least one instruction is also provided, such as a memory including at least one instruction, which can be executed by a processor in a computing device to complete the training method or video editing method of the video editing model in the above embodiments. Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium, such as ROM (Read-Only Memory), RAM (Random-Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage device, etc.
[0248] In an exemplary embodiment, a computer program product is also provided, including one or more instructions that can be executed by a processor of a computing device to complete the training method or video editing method of the video editing model provided in the above embodiments.
[0249] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the training samples involved in this application were all obtained with full authorization.
[0250] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0251] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for training a video editing model, the method comprising: The method comprises: obtaining a first sample video, a second sample video and a first sample voice, wherein the first sample video and the second sample video both comprise a first object, the first sample voice is a voice of the first object in the first sample video, and in the first sample video and the second sample video, the lip shape of the first object is different and the features other than the lip shape of the first object are the same; training a video editing model by taking the second sample video as a video to be edited, taking the first sample video as an expected output and taking the first sample voice as a constraint condition; in the training process, the video editing model edits the lip shape of the first object in the second sample video based on the lip shape of the first object in the first sample video, so that the lip shape of the first object in the second sample video approaches the lip shape of the first object in the first sample video.
2. The method of claim 1, wherein, The training of the video editing model comprises: extracting video features of the first sample video, video features of the second sample video and voice features of the first sample voice; based on a noise adding coefficient, adding noise to the video features of the first sample video; training the video editing model based on the voice features, the video features of the second sample video and the video features of the first sample video after noise adding.
3. The method of claim 2, wherein, The video editing model comprises a global sub-model, a lip shape sub-model and a texture sub-model, the global sub-model is used to keep the global structure of the video frames in the second sample video unchanged, the lip shape sub-model is used to edit the lip shape of the first object in the second sample video, and the identity texture sub-model is used to keep the face identity information and face texture information of the first object in the second sample video unchanged; The training of the video editing model based on the voice features, the video features of the second sample video and the video features of the first sample video after noise adding comprises: training the global sub-model, the lip shape sub-model and the texture sub-model respectively based on the voice features, the video features of the second sample video and the video features of the first sample video after noise adding.
4. The method of claim 3, wherein, The training of the global sub-model, the lip shape sub-model and the texture sub-model respectively based on the voice features, the video features of the second sample video and the video features of the first sample video after noise adding comprises: inputting the video features of the second sample video and the video features after noise adding into the global sub-model, the global sub-model predicting the global structure of the first sample video frame in the first sample video based on the video features of the second sample video and the video features after noise adding, so that the predicted global structure approaches the global structure of the first sample video frame, and outputting a first predicted video capable of reflecting the predicted global structure; inputting the voice feature, the video feature of the second sample video and the noise-added video feature into the lip shape sub-model, and predicting the lip shape of the first object in the first sample video based on the voice feature, the video feature of the second sample video and the noise-added video feature, so that the predicted lip shape is close to the lip shape of the first object in the first sample video, editing the lip shape of the first object in the second sample video into the predicted lip shape, and outputting a second predicted video, which is the second sample video after lip shape editing. inputting the video feature of the second sample video and the noise-added video feature into the identity texture sub-model, and predicting the face identity information and face texture information of the first object in the first sample video based on the video feature of the second sample video and the noise-added video feature, so that the predicted face identity information is close to the face identity information of the first object in the first sample video, and the predicted face texture information is close to the face texture information of the first object in the first sample video, and outputting a third predicted video capable of reflecting the predicted face identity information and face texture information.
5. The method of claim 4, wherein, For different sub-models in the global sub-model, the lip shape sub-model and the identity texture sub-model, the noise-added video feature input into the different sub-models is obtained based on noise-added coefficients in different intervals.
6. The method of Claim 1, wherein, The method further comprises: obtaining a second sample voice, the second sample voice being different from the first sample voice in content; processing the first sample video and the second sample voice through a video generation model to obtain a third sample video, the lip shape of the first object in the third sample video being consistent with the second sample voice; obtaining the second sample video based on the third sample video.
7. The method of claim 6, wherein, The second sample voice is the same as the voiceprint feature and timbre feature of the first sample voice.
8. The method of claim 6, wherein, The processing of the first sample video and the second sample voice through the video generation model comprises: processing a mask video, the first sample video and the second sample voice through the video generation model, wherein the mask video is the first sample video added with a mask, the mask is used to cover the mouth of the first object in the first sample video and does not cover a second object, and the second object is any object that shields the first object.
9. The method of claim 8, wherein, The first sample video comprises a plurality of first sample video frames, and the method further comprises: for any first sample video frame, if the first image region in the first sample video frame does not exist the second object, adding the mask in the entire first image region, if the first image region exists the second object, adding the mask in the region outside the second object in the first image region, obtaining the mask video, and the first image region is the image region where the mouth of the first object is located in the first sample video frame.
10. The method of claim 6, wherein, The first sample video includes a plurality of first sample video clips, and the second sample voice includes a plurality of voice clips, different first sample video clips correspond to different voice clips; The processing of the first sample video and the second sample voice by the video generation model to obtain the third sample video includes: For any first sample video clip, the first sample video clip and the voice clip corresponding to the first sample video clip are processed by the video generation model to obtain a second sample video clip corresponding to the voice clip, the lip shape of the first object in the second sample video clip corresponds to the voice clip, and the lip shape of the first object in the first sample video clip and the second sample video clip is different, and the characteristics other than the lip shape of the first object are the same; The second sample video clips corresponding to the plurality of voice clips are spliced to obtain the third sample video.
11. The method of Claim 10, wherein, The processing of the first sample video clip and the voice clip corresponding to the first sample video clip by the video generation model includes: For the i-th video clip in the plurality of first sample video clips, a plurality of video frames are obtained from the i-1-th video clip in the plurality of first sample video clips, i is an integer greater than or equal to 2, and the last video frame of the i-1-th video clip is the starting point; The i-th video clip, the voice clip corresponding to the i-th video clip, and the plurality of video frames are processed by the video generation model under the constraint conditions of the plurality of video frames and the voice clip corresponding to the i-th video clip.
12. The method of claim 10, wherein, The duration of the first sample video is greater than or equal to a first duration, the video generation model is a trained basic generation model, and the method further includes: Obtaining a fourth sample video and a third sample voice, the duration of the fourth sample video is less than the first duration, and the content of the third sample voice is different from the voice of a third object in the fourth sample video; Based on the fourth sample video and the third sample voice, the basic generation model is trained, and in the training process, the lip shape of the third object in the fourth sample video is edited based on the third sample voice as a constraint condition and one video frame in the fourth sample video as a reference frame, so that the lip shape of the first object in the fourth sample video approaches the lip shape corresponding to the third sample voice.
13. The method of claim 12, wherein, The training of the basic generation model based on the fourth sample video and the third sample voice includes: The basic generation model is over-fitted based on the fourth sample video and the third sample voice.
14. The method of claim 6, wherein, The obtaining of the second sample video based on the third sample video includes: The video frames in the third sample video are subjected to multiple different degrees of light enhancement to obtain a plurality of fifth sample videos, and the fifth sample video is the third sample video after light enhancement; acquire the second sample video from the plurality of fifth sample videos.
15. The method of claim 6, wherein, The acquiring the second sample video based on the third sample video comprises: if a face similarity of the first object between the third sample video and the first sample video is greater than or equal to a first threshold value, and a mouth shape similarity of the first object between the third sample video and the first sample video is greater than or equal to a second threshold value, the third sample video is determined as the second sample video; wherein the face similarity indicates a similarity between a face of the first object in the third sample video and a face of the first object in the first sample video, and the mouth shape similarity indicates a similarity between a mouth shape of the first object in the third sample video and a mouth shape of the first object in the first sample video.
16. A video editing method characterized by, comprising: acquire a first video and a second voice different from a first voice content, wherein the first video comprises a fourth object, and the first voice is a voice of the fourth object in the first video; edit a mouth shape of the fourth object in the first video by a video editing model to obtain a second video, taking the first video as a video to be edited and taking the second voice as a constraint condition; wherein the video editing model is a video editing model trained by the method in any one of claims 1 to 15; the second video is the first video after mouth shape editing, the mouth shape of the fourth object in the second video is consistent with the second voice, and the mouth shapes of the fourth object in the first video and the second video are different, and the features other than the mouth shapes of the fourth object are the same.
17. An apparatus for training a video editing model, comprising: comprising: a first acquisition unit configured to perform a first sample video, a second sample video, and a first sample voice, wherein the first sample video and the second sample video both comprise a first object, and the first sample voice is a voice of the first object in the first sample video, and the mouth shapes of the first object in the first sample video and the second sample video are different, and the features other than the mouth shapes of the first object are the same; a first training unit configured to train a video editing model by taking the second sample video as a video to be edited, taking the first sample video as an expected output, and taking the first sample voice as a constraint condition; during the training process, the video editing model edits the mouth shape of the first object in the second sample video based on the mouth shape of the first object in the first sample video, so that the mouth shape of the first object in the second sample video approaches the mouth shape of the first object in the first sample video.
18. A video editing apparatus characterized by comprising: comprising: an acquisition unit configured to acquire a first video and a second voice different from a first voice content, wherein the first video comprises a fourth object, and the first voice is a voice of the fourth object in the first video; An editing unit configured to perform, taking the first video as a video to be edited, taking the second voice as a constraint condition, editing the lip shape of the fourth object in the first video by a video editing model to obtain a second video. The video editing model is a video editing model trained by the method in any one of claims 1 to 15. The second video is the first video after lip shape editing, the lip shape of the fourth object in the second video is consistent with the second voice, and the lip shapes of the fourth object in the first video and the second video are different and the features other than the lip shapes of the fourth object are the same.
19. A computing device, comprising: Comprise: One or more processors; One or more memories for storing instructions executable by the one or more processors; Wherein the one or more processors are configured to execute the instructions to implement the method of any one of claims 1 to 16.
20. A computer-readable storage medium, characterized in that, When at least one instruction in the computer readable storage medium is executed by one or more processors of a computing device, the computing device is enabled to perform the method of any one of claims 1 to 16.
21. A computer program product, characterised in that, Comprise one or more instructions executed by one or more processors of a computing device, so that the computing device can perform the method of any one of claims 1 to 16.