Background preserving and target object motion synchronizing method for zero sample video editing
By adopting the background holding mechanism and the target object motion synchronization framework in video editing, the problem of insufficient background protection in the prior art is solved, the background holding and object motion synchronization is achieved, and the quality and efficiency of video editing are improved.
Patent Information
- Application Number
- CN202510160426.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-16
AI Technical Summary
When the existing video editing technology achieves the consistency of the movement of the target object, it fails to fully consider the protection of the background area, resulting in unnecessary changes in the background part in the video, making it difficult to achieve background maintenance and object movement synchronization at the same time.
In zero-sample video editing, the background maintenance mechanism and the target object motion synchronization framework are adopted, and the technologies such as segmentation and tracking, Gaussian blur, uncertainty guidance and Martensa distance are used to ensure that the background remains unchanged and the synchronization of object motion is achieved.
It realizes efficient editing of videos under zero sample conditions, ensuring that the background area remains unchanged and the movement of the target object is synchronized with the source video, improving the quality and efficiency of video editing.
Smart Images

Figure CN120017927A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular to video editing and image processing technology. Specifically, the present invention relates to a method for background preservation and target object motion synchronization for zero-sample video editing, aiming to achieve high-quality video editing effects by precisely controlling the background area to remain unchanged and synchronizing the motion of the target object with the source video. Background Art
[0002] In recent years, diffusion models have made breakthrough progress in the field of image editing due to their excellent training stability and generated image quality. The purpose of image editing is to modify the attributes or style of an image under the guidance of text. Some studies have used pre-trained image diffusion models to perform efficient image editing according to user requirements.
[0003] For example, Prompt-to-Prompt
[17] achieves semantic editing by re-weighting the cross-attention maps of different text prompts. Subsequently, Plug-and-Play
[16] , Pix2pixZero
[17] , and Masactrl
[19] proposed similar work, which uses cross-attention and self-attention to generate images at each time step while preserving the layout and structure of the source image. In addition, DreamBooth
[19] proposed an agent-driven generation technique by fine-tuning the pre-trained model, while IP-Adapter
[20] provided a lightweight adapter to enhance the image prompting ability of the pre-trained text-based diffusion model.
[0004] As image editing technology matures, video editing begins to attract more and more researchers' attention. Gen1 and Dreamix fine-tune the network by introducing a time-aware layer to generate continuous videos; Tune-A-Video designs an attention adjustment mechanism for pre-trained text-based graph models to generate relevant action sequences. Later, CCEidt proposed a trident network that controls appearance, structure, and temporal consistency. However, these methods still need to train the network to improve performance, which has limitations in practical applications.
[0005] To solve the above problems, FateZero proposed the first zero-shot video editing method, which maintains structure and motion information by fusing attention maps during inversion and inference. Subsequently, Rerender-A-Video estimated the optical flow of the video and re-rendered only the key frames for style transfer. TokenFlow uses the semantic correspondence of diffuse features across frames to enhance the consistency of the diffuse feature space and further improve the video editing effect. At the same time, FLATTEN designed a novel optical flow-guided attention, which forces the features on the same flow path of different frames to pay attention to each other in the attention module, thereby enhancing the temporal consistency of video editing.
[0006] In summary, although these methods have made some progress in video editing, they still face many challenges. In particular, existing methods generally have limitations in how to ensure visual consistency between video frames, maintain the stability of the background area, and synchronize the movement of objects. Most video editing technologies do not fully consider the protection of the background area when achieving the consistency of the target object's motion, which may lead to unnecessary changes in the background part of the video. Therefore, how to accurately control the retention of the background while ensuring the consistency of the object's motion is still a key problem in video editing technology. Summary of the invention
[0007] In view of the above problems, the present invention proposes a background preservation and target object motion synchronization method for zero-sample video editing, and comprises the following steps:
[0008] 1. A method for background preservation and target object motion synchronization for zero-sample video editing, characterized in that it comprises the following steps:
[0009] S1, based on the source video and its text prompts, combined with the target text prompts, generates a video that meets the requirements, ensuring that the background remains unchanged and the target object moves synchronously.
[0010] S2, encodes the source video, maps it to the latent space, adds noise to the latent space representation of each frame, and restores the clean latent space representation through DDIM inversion for subsequent editing.
[0011] S3, builds a background preservation mechanism to ensure that the background of the source video remains unchanged; distinguishes the background and target object features through segmentation and tracking, applies Gaussian blur and other techniques to make the target object boundary transition naturally, and controls the background to remain unchanged through background preservation constraints.
[0012] S4, constructs a target object motion consistency framework and uses Mahalanobis distance and inter-frame difference constraints to ensure the motion consistency between the source video and the target video.
[0013] S5, generates the target video by optimizing the background preservation, Mahalanobis distance and inter-frame difference constraints through gradient descent, ensuring that the background remains unchanged and the object motion is synchronized.
[0014] 2. The method for background preservation and target object motion synchronization for zero-sample video editing according to claim 1, wherein step S1 specifically comprises the following steps:
[0015] S11, prepare the source video, which consists of multiple frames of continuous images and is equipped with corresponding text prompts to ensure that the video content accurately presents the required visual effects and action information, providing reliable material support for subsequent editing.
[0016] S12, prepare an editing text prompt, which is used to drive the generation of the target video. The main goal of the editing prompt is to ensure that the movement of the target object is consistent with the source video, while effectively retaining the background area to avoid the interference of the target object movement on the background, thereby ensuring a natural video transition and synchronization of the background and target object movement.
[0017] 3. The method for background preservation and target object motion synchronization for zero-sample video editing according to claim 1, wherein step S2 specifically comprises the following steps:
[0018] S21, encodes the source video using the pre-trained encoder to generate the latent space representation of each frame.
[0019] S22, add Gaussian noise to the latent space representation at each time step. Specifically, for the tth step, the noise features are sampled according to the Gaussian distribution. The mathematical expression of this process is:
[0020]
[0021] Where t=1...T, T is the total number of iterations. t represents the noise characteristics of the tth step, and N is a Gaussian distribution.
[0022] S23, uses DDIM inversion to invert the noisy latent space representation, restores a clear latent space representation, and will be used in the subsequent editing process. This process can effectively retain the key feature information of the source video and provide support for achieving background retention and target object motion synchronization. This part can be expressed as:
[0023]
[0024] 4. The method for background preservation and target object motion synchronization for zero-sample video editing according to claim 1, wherein the step S3 constructs a background preservation mechanism to ensure that the background of the source video remains unchanged, and specifically comprises the following steps:
[0025] S31, using segmentation and tracking technology to accurately separate the background and target object in the source video. Specifically, the Segment and TrackAnything model is used to segment the video, obtain the masks of the background and the target object, and provide a basis for maintaining the synchronization of the background and the target object movement.
[0026] S32 uses Gaussian blur and other smoothing techniques to process the boundaries of the target object to ensure that the boundaries transition naturally when the shape of the target object changes, avoiding the inharmonious effects caused by direct editing. This part can be expressed as:
[0027] G t= GaussianBlur(GetMask(z t )) (3)
[0028] S33, design uncertainty-guided background preservation mechanism. By fusing the background noise map of the source video with the edited video, the background area is constrained to remain unchanged. This mechanism uses uncertainty information to guide the retention of background features, ensuring that the background area in the edited video is consistent with the source video, and avoiding background changes caused by editing operations. The specific formula is:
[0029]
[0030]
[0031] 5. The method for background preservation and target object motion synchronization for zero-shot video editing according to claim 1, wherein the step S4 constructs a target object motion consistency framework, and uses Mahalanobis distance and inter-frame difference constraints to ensure motion consistency between the source video and the target video, and specifically comprises the following steps:
[0032] S41, calculate the diffusion features of the source video and the target video. The diffusion features of the source video and the edited video are extracted through the U-Net network:
[0033]
[0034] S42, apply Mahalanobis distance constraint to ensure motion consistency in high-dimensional feature space. Mahalanobis distance considers the covariance structure of feature distribution, reduces the influence of outliers, and ensures the accuracy and robustness of motion alignment. By calculating the Mahalanobis distance between the source video and the target video in the high-dimensional feature space, the probability distribution difference between the two is quantified to achieve accurate motion alignment. The mathematical expression of this process is:
[0035]
[0036] S43, using inter-frame difference constraints to ensure motion consistency in the low-dimensional feature space. Inter-frame difference ensures that the motion changes between consecutive frames remain consistent by calculating the difference in motion vectors between each frame in the source video and the target video. The scene semantic information is obtained by compressing high-dimensional features, and the difference in motion vectors between frames is calculated using this information. The mathematical expression of this process is:
[0037]
[0038]
[0039]
[0040] 6. The method for background preservation and target object motion synchronization for zero-sample video editing according to claim 1, wherein step S5 specifically comprises the following steps:
[0041] S51, construct the objective function. Combining the background preservation loss, motion synchronization loss and inter-frame consistency loss, the objective function is constructed and optimized using gradient descent. By calculating the gradient of the objective function relative to the latent space representation, the latent space representation is updated to minimize the objective function, and the target video that meets the editing prompts is gradually generated. This part can be expressed as:
[0042]
[0043]
[0044] S52, decoding to generate a target video. After the optimization is completed, a pre-trained decoder is used to decode the optimized latent space representation into a target video in pixel space, ensuring that the generated video meets the editing requirements of background preservation and target object motion synchronization.
[0045] Beneficial effects: Compared with the prior art, the present invention provides a background preservation and target object motion synchronization method for zero-sample video editing, which can accurately control the motion synchronization of the background and the target object in different video editing scenarios and produce the following beneficial effects:
[0046] 1. Efficient video editing: Most existing video editing methods rely on a large amount of training data and long network training. However, this invention innovatively adopts adaptive prompting and motion synchronization mechanisms to achieve accurate background preservation and object motion consistency without a large number of training samples. This method significantly improves editing efficiency, especially in zero-sample editing tasks.
[0047] 2. Precise control of background area: Traditional video editing methods often have difficulty in achieving both the motion synchronization of the target object and the preservation of the background area, resulting in background distortion or inconsistent motion of the target object in the generated video. The present invention successfully avoids changes in the background area during the editing process through a sophisticated background preservation mechanism, while ensuring smooth motion of the target object throughout the entire video frame.
[0048] 3. High motion synchronization accuracy: This invention introduces a motion synchronization method based on diffusion features, which greatly reduces the inconsistency of object motion and flicker between frames by achieving accurate target object motion alignment in each video frame. Compared with existing methods, it can achieve more natural and coherent motion migration in complex scenes.
[0049] 4. Simplified training process: Compared with the existing method that requires training different models for each editing task, this paper proposes a unified editing framework that can handle multiple video editing tasks without additional training data. This method not only reduces the computational and memory burden, but also makes the model easier to deploy in practical applications, with better applicability and scalability.
[0050] 5. Broad application prospects: The present invention has strong universality and can be widely used in various video editing tasks, such as film production, virtual reality, advertising production, and social media content editing. Especially in short video editing and intelligent video creation, the present invention can greatly improve editing quality, reduce creation costs, and has significant commercial application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 Flowchart of the background preservation and target object motion synchronization method for zero-shot video editing.
[0052] Figure 2 The figure shows the visual results and qualitative comparison between the present invention and other methods.
[0053] Figure 3 and Figure 4 The proposed method is able to perform precise object editing while ensuring background preservation and motion consistency.
[0054] Figure 5 The figure is a comparison chart of the results of the method of the present invention and other methods in terms of timing consistency, editing accuracy, etc.
[0055] Figure 6 and Figure 7 This is a visualization result diagram of the background preservation and target object motion synchronization method for zero-sample video editing of the present invention and other methods. DETAILED DESCRIPTION
[0056] The background preservation and target object motion synchronization method for zero-sample video editing of the present invention solves the time consistency and background preservation problems existing in existing video editing methods by innovatively introducing background preservation technology and target object motion synchronization mechanism, and can accurately edit videos without additional training samples, keep the background area stable and synchronize the target object motion. The specific implementation process is as follows.
[0057] The invention will be further described below in conjunction with specific embodiments.
[0058] Embodiment 1:
[0059] like Figure 1As shown, given the source video and text prompt, the source video is first encoded into the latent space by the encoder, noise is added and the clear latent space representation is restored by DDIM inversion. Then the background preservation mechanism is constructed, the background and the target object are distinguished by segmentation and tracking technology, the boundary is smoothed by Gaussian blur, and uncertainty is introduced to guide the background preservation to ensure the consistency of the background area. Further construct the target object motion consistency framework, use Mahalanobis distance and inter-frame difference constraints to ensure the motion consistency between the source video and the target video. Finally, optimize the objective function, combine the background preservation loss, motion synchronization loss and inter-frame consistency loss, perform gradient descent optimization, and generate a video that meets the target editing requirements. After the optimization is completed, use the pre-trained decoder to decode it into the target video. Through these steps, the present invention can achieve accurate video editing without additional training samples and ensure that the background and target object motion are synchronized.
[0060] Figure 2 The visual results and qualitative comparison of the present invention with other methods show that our method performs better than other methods.
[0061] Figure 3 and Figure 4 The proposed method is able to perform precise object editing while ensuring background preservation and motion consistency.
[0062] Figure 5 The figure is a comparison chart of the results of the method of the present invention and other methods in terms of timing consistency, editing accuracy, etc.
[0063] Figure 6 and Figure 7 This is a visualization result diagram of the background preservation and target object motion synchronization method for zero-sample video editing of the present invention and other methods.
[0064] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0065] Although the above describes the specific implementation methods of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A background preservation and target object motion synchronization method for zero-shot video editing, characterized in that: The following steps are involved: S1, based on the source video and its text prompts, combined with the target text prompts, generates a video that meets the requirements, ensuring that the background remains unchanged and the target object moves synchronously. S2, encodes the source video, maps it to the latent space, adds noise to the latent space representation of each frame, and restores the clean latent space representation through DDIM inversion for subsequent editing. S3, builds a background preservation mechanism to ensure that the background of the source video remains unchanged; distinguishes the background and target object features through segmentation and tracking, applies Gaussian blur and other techniques to make the target object boundary transition naturally, and controls the background to remain unchanged through background preservation constraints. S4, constructs a target object motion consistency framework and uses Mahalanobis distance and inter-frame difference constraints to ensure the motion consistency between the source video and the target video. S5, generates the target video by optimizing the background preservation, Mahalanobis distance and inter-frame difference constraints through gradient descent, ensuring that the background remains unchanged and the object motion is synchronized.
2. A method for background preservation and target object motion synchronization for zero-shot video editing as claimed in claim 1, characterized in that: The step S1 specifically includes the following steps: S11, prepare the source video, which consists of multiple frames of continuous images and is equipped with corresponding text prompts to ensure that the video content accurately presents the required visual effects and action information, providing reliable material support for subsequent editing. S12, prepare an editing text prompt, which is used to drive the generation of the target video. The main goal of the editing prompt is to ensure that the movement of the target object is consistent with the source video, while effectively retaining the background area to avoid the interference of the target object movement on the background, thereby ensuring a natural video transition and synchronization of the background and target object movement.
3. The method for background preservation and target object motion synchronization for zero-shot video editing as claimed in claim 1, characterized in that: The step S2 specifically includes the following steps: S21, encodes the source video using the pre-trained encoder to generate the latent space representation of each frame. S22, add Gaussian noise to the latent space representation at each time step. Specifically, for the tth step, the noise features are sampled according to the Gaussian distribution. The mathematical expression of this process is: Where t=1...T, T is the total number of iterations. t represents the noise characteristics of the tth step, and N is a Gaussian distribution. S23, uses DDIM inversion to invert the noisy latent space representation, restores a clear latent space representation, and will be used in the subsequent editing process. This process can effectively retain the key feature information of the source video and provide support for achieving background retention and target object motion synchronization. This part can be expressed as:
4. The method for background preservation and target object motion synchronization for zero-shot video editing as claimed in claim 1, characterized in that: The step S3 constructs a background preservation mechanism to ensure that the background of the source video remains unchanged, and specifically includes the following steps: S31, using segmentation and tracking technology to accurately separate the background and target object in the source video. Specifically, the SegmentandTrackAnything model is used to segment the video, obtain the masks of the background and the target object, and provide a basis for maintaining the synchronization of the background and the target object movement. S32 uses Gaussian blur and other smoothing techniques to process the boundaries of the target object to ensure that the boundaries transition naturally when the shape of the target object changes, avoiding the inharmonious effects caused by direct editing. This part can be expressed as: G t =GaussianBlur(GetMask(z t )) (3) S33, design uncertainty-guided background preservation mechanism. By fusing the background noise map of the source video with the edited video, the background area is constrained to remain unchanged. This mechanism uses uncertainty information to guide the retention of background features, ensuring that the background area in the edited video is consistent with the source video, and avoiding background changes caused by editing operations. The specific formula is:
5. The method for background preservation and target object motion synchronization for zero-shot video editing as claimed in claim 1, characterized in that: The step S4 constructs a target object motion consistency framework, and uses Mahalanobis distance and inter-frame difference constraints to ensure motion consistency between the source video and the target video, specifically including the following steps: S41, calculate the diffusion features of the source video and the target video. The diffusion features of the source video and the edited video are extracted through the U-Net network: S42, apply Mahalanobis distance constraint to ensure motion consistency in high-dimensional feature space. Mahalanobis distance considers the covariance structure of feature distribution, reduces the influence of outliers, and ensures the accuracy and robustness of motion alignment. By calculating the Mahalanobis distance between the source video and the target video in the high-dimensional feature space, the probability distribution difference between the two is quantified to achieve accurate motion alignment. The mathematical expression of this process is: S43, using inter-frame difference constraints to ensure motion consistency in the low-dimensional feature space. Inter-frame difference ensures that the motion changes between consecutive frames remain consistent by calculating the difference in motion vectors between each frame in the source video and the target video. The scene semantic information is obtained by compressing high-dimensional features, and the difference in motion vectors between frames is calculated using this information. The mathematical expression of this process is:
6. The method for background preservation and target object motion synchronization for zero-shot video editing as claimed in claim 1, characterized in that: The step S5 specifically includes the following steps: S51, construct the objective function. Combining the background preservation loss, motion synchronization loss and inter-frame consistency loss, the objective function is constructed and optimized using gradient descent. By calculating the gradient of the objective function relative to the latent space representation, the latent space representation is updated to minimize the objective function, and the target video that meets the editing prompts is gradually generated. This part can be expressed as: S52, decoding to generate a target video. After the optimization is completed, a pre-trained decoder is used to decode the optimized latent space representation into a target video in pixel space, ensuring that the generated video meets the editing requirements of background preservation and target object motion synchronization.