Negative video content suppression method based on text embedding
Through a text embedding method, using singular value decomposition and fuzzy feature selection strategies, the problems of negative content suppression and timing consistency in video editing are solved, and efficient and natural video editing effects are achieved.
Patent Information
- Application Number
- CN202510160338.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-22
AI Technical Summary
Existing video editing methods are difficult to effectively suppress negative content when processing negative text prompts, and there are challenges in maintaining video timing consistency, which is prone to flickering and incoherence between frames.
Through a text embedding-based method, using singular value decomposition and exponential suppression operators to decouple negative embedding, a constraint mechanism is designed to optimize text embedding, and fuse cross-frame similar features through a fuzzy feature selection strategy to ensure the timing consistency of video content.
Effectively suppress negative content, maintain the timing consistency of videos, improve editing efficiency, enhance semantic comprehension ability, support multiple editing tasks, and improve visual quality.
Smart Images

Figure CN120358384A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and video processing, and particularly relates to a text-driven video editing method. Specifically, the present invention proposes a method for suppressing negative video content based on text embedding, aiming to accurately suppress negative content through semantic modulation methods, while keeping the non-edited areas of the video unchanged and ensuring that the movement of the target object is synchronized with the source video, thereby generating a high-quality and semantically consistent edited video. Background Art
[0002] In recent years, diffusion models have demonstrated excellent capabilities in the field of image generation, capable of generating high-quality and diverse image content. These studies have promoted the development of image generation technology, enabling the generated images to reach a relatively high level in terms of visual effects and semantic consistency. Based on this, some researchers have begun to explore the application of diffusion models in the fields of video generation and video editing, with the expectation of achieving similar breakthroughs in the generation and editing of dynamic video content.
[0003] 1. Text-to-Video Generation Based on Diffusion Models
[0004] Different from static image generation, the task of text-to-video generation introduces the time dimension, requiring the model to not only generate high-quality single-frame images but also maintain the consistency of visual information in the time series to generate a coherent dynamic video sequence. For example, the Video Diffusion Model (VDM) first extended the two-dimensional U-Net to a three-dimensional U-Net and directly performed the diffusion process at the pixel level to generate the expected video content. Subsequently, MagicVideo and LVDM further explored the possibility of video generation in the latent space, which not only effectively reduced the computational cost but also significantly improved the processing speed.
[0005] In addition, to better maintain the temporal consistency of the video, Text2Video-Zero and ControlVideo applied the cross-frame attention mechanism to the pre-trained text-to-image diffusion model. These methods ensure the coherence and consistency of video content over time by introducing an attention mechanism between different frames of the video. These research results not only demonstrate the potential of diffusion models in the field of video generation but also provide new directions and ideas for the development of video editing technology.
[0006] Nevertheless, existing video editing methods still have deficiencies in dealing with certain complex tasks, especially in poorly understanding and processing negative text prompts (such as "without glasses", "not wearing a hat", etc.), resulting in possible negative content in the edited video. In addition, existing methods also face challenges in maintaining the temporal consistency of the video, prone to flickering and incoherence between frames. Therefore, developing a video editing method that can effectively suppress negative content and maintain temporal consistency is an important direction of current research.
[0007] 2. Text-to-Video Editing Based on Diffusion Model
[0008] Currently, text-to-video editing can be divided into two categories: fine-tuning methods and training-free methods. Specifically, the former mainly uses pre-trained text-to-image models to edit the theme, attributes, or styles of videos. For example, Tune-a-Video extends the latent space diffusion model to the time domain and implicitly encodes the motion information of the source video through a single fine-tuning. However, due to the insufficient dynamic correlation between video frames, it is difficult for these methods to ensure temporal consistency. Subsequently, CC Edit designed a trident network that separates the structure and appearance to ensure precise editing functions. However, these methods have limitations in practical applications due to high training costs and data requirements.
[0009] To alleviate these limitations, several training-free methods have been proposed. For example, FateZero proposed an attention mixing strategy to fuse the attention features between inversion and sampling to achieve high-quality editing results. Subsequently, TokenFlow enhanced the temporal consistency of the video through a linear combination between diffusion features to avoid flickering. At the same time, FLATTEN conducts attention learning on the same flow path of different frames, thus enhancing the inter-frame consistency. In addition, COVE proposed an effective sliding window-based strategy to calculate the similarity between source videos to achieve consistent video editing results.
[0010] In summary, although these methods have made certain progress in video editing, they still face many challenges. Especially in poorly understanding and processing negative text prompt content (such as "without glasses" and "not wearing a hat"), resulting in ineffective suppression of bad content in the edited video. In addition, these methods also face challenges in maintaining the temporal consistency of the video, prone to flickering and incoherence between frames. Therefore, developing a video editing method that can effectively suppress negative content and maintain temporal consistency is an important direction of current research. Summary of the Invention
[0011] In view of the above problems, the present invention proposes a method for suppressing negative video content based on text embedding, and includes the following steps:
[0012] 1. A method for suppressing negative video content based on text embedding, characterized by comprising the following steps:
[0013] S1, Based on the source video and its text prompt, combined with the target text prompt, generate a video that meets the requirements, ensuring that the content of the edited video is consistent with the target text prompt while maintaining the temporal consistency of the video.
[0014] S2, Convert the input text prompt into a text embedding through a text encoder. First, analyze the positive embedding and negative embedding coupled in the text embedding, further decouple the negative embedding from the text embedding through singular value decomposition, and then use an exponential suppression operator to reduce the singular value of the negative embedding to suppress the impact of negative content on the edited video.
[0015] S3, Design two constraint mechanisms to further optimize the text embedding by pushing away the negative embedding and pulling in the positive embedding, ensuring that negative content is effectively suppressed while positive content remains unchanged.
[0016] S4, Construct a fuzzy feature selection and fuzzy feature decomposition module for fusing similar features across frames to ensure the temporal consistency of the video.
[0017] S5, Generate the target video, ensuring that the content of the edited video is consistent with the target text prompt while maintaining the temporal consistency of the video.
[0018] 2. The method for suppressing negative video content based on text embedding according to claim 1, characterized in that the specific steps of step S1 include the following steps:
[0019] S11, Prepare the source video, which consists of multiple consecutive images and is accompanied by corresponding text prompts to ensure that the video content accurately presents the required visual effects.
[0020] S12, Prepare the editing text prompt, which is used to drive the generation of the edited video. The main goal of the editing prompt is to clearly point out the negative content to be suppressed (such as "not wearing glasses", "not wearing a hat", etc.). Through precise text description, ensure natural video transition, so as to achieve high-quality video editing effects.
[0021] 3. The method for suppressing negative video content based on text embedding according to claim 1, characterized in that the specific steps of step S2 include the following steps:
[0022] S21, Convert the input text prompt into a text embedding through a text encoder Since the text embedding includes a positive embedding and a negative embedding, therefore the text prompt needs to include the negative content to be suppressed, namely CNE (such as "not wearing glasses", "not wearing a hat", etc.) and positive content, i.e., C PE 。
[0023] S22. Analyze the influence of positive and negative coupling information in the text embedding on the generated video, and further identify the positive embedding and negative embedding. Create an original embedding matrix composed of the negative embedding and the end embedding Then, use singular value decomposition on it to decouple the negative embedding from the original embedding matrix M, which can be expressed as:
[0024] M = U∑V T (1)
[0025] S23. Apply an exponential suppression operator to attenuate the singular values of the negative embedding, thereby suppressing the influence of the negative embedding on the edited video content. This can be expressed as:
[0026] f(∑) = e ∑ ·∑ (2)
[0027] where ∑ represents the singular value matrix of the negative embedding.
[0028] Then, we combine the modulated singular values of the negative embedding with the original unitary matrices U and V to obtain an updated target embedding matrix to perform singular value reconstruction, and the formula is as follows:
[0029]
[0030] Then replace the original M to obtain the updated target embedding
[0031] 4. A method for suppressing negative video content based on text embedding according to claim 1, wherein the specific steps of step S3 include the following steps:
[0032] S31. Input the original embedding into the diffusion model to obtain the attention map at time step t where corresponds to the positive embedding C in the initial embedding PE and the negative embedding C NE . Similarly, for the updated target embedding the corresponding attention map
[0033] S32. In order to further suppress negative content and at the same time maintain the consistency of the non-edited area, an embedding optimization based on the attention map is designed, and two constraint mechanisms of positive content retention and negative content suppression are proposed.
[0034] Positive content retention loss: Used to retain positive video content and ensure that the semantic information of positive embeddings remains unchanged. The loss function is defined as:
[0035]
[0036] Negative content suppression loss: Used to suppress negative video content and ensure that the influence of negative embeddings is effectively suppressed. The loss function is defined as:
[0037]
[0038] 5. A method for suppressing negative video content based on text embedding according to claim 1, wherein the specific steps of step S4 include the following steps:
[0039] S41, Divide the source video into multiple video groups, each video group containing several consecutive frames. Assume that the source video contains S frames, which are divided into N video groups, and each video group contains M frames, i.e., S = N × M.
[0040] S42, Extract features from each frame in each video group to obtain a feature set X = {X1, X2, …, X M}. Then divide the feature set into a source feature set and a target feature set where the source feature set contains the features of all frames except the target frame, and the target feature set contains the features of the target frame. Calculate the fuzzy membership degree for each feature, and use the fuzzy membership degree function to measure the similarity between features. This part can be expressed as:
[0041]
[0042] where, and represent features, θ src and θ tar are parameters that control the fuzzy membership degree.
[0043] S43, According to the similarity, select the most similar feature pairs for fusion to form a local feature fusion
[0044] X fused = Fusion(T, S) (9)
[0045] where, Fusion(·) represents the fusion operation.
[0046] To further enhance the long-term temporal consistency of the video, maintain a global feature set G. The initial global feature set is the local fusion feature of the first video group For each subsequent video group i, the local fusion features of the current video group are fused with the global features G of the previous video group i-1 to form new global features G i , which are input into the self-attention module to obtain appearance and structure information. This part can be expressed as:
[0047]
[0048] X SA = Self_Attention(G i ) (11)
[0049] Finally, according to the feature selection rule during fusion, X SA is decomposed back to the original frame size to restore the features of each frame for convenient subsequent operations.
[0050] 6. A method for suppressing negative video content based on text embedding according to claim 1, characterized in that the specific steps of step S5 include the following steps:
[0051] S51, constructing an objective function. An objective function is constructed by combining the positive content retention loss and the negative content suppression loss to optimize the text embedding, and gradually generate the target video that meets the editing prompt. This part can be expressed as:
[0052] L = λ1L pr + λ2L ns (12)
[0053] where λ1 and λ2 are hyperparameters.
[0054] S52, finally using the pre-trained decoder to decode the optimized latent space representation into the target video in the pixel space, ensuring that the edited video content is consistent with the target text prompt while maintaining the temporal consistency of the video.
[0055] Beneficial effects: Compared with the prior art, a method for suppressing negative video content based on text embedding proposed by the present invention can generate a video with temporal consistency while suppressing unwanted video content, and produce the following significant beneficial effects:
[0056] 1. Effectively suppress negative content. Through the semantic modulation technology, the present invention can accurately identify and suppress negative content in the video (such as "without glasses", "not wearing a hat", etc.). Compared with the existing methods, the present invention can not only effectively remove the negative content, but also ensure the semantic consistency of the video, avoiding semantic deviation caused by editing.
[0057] 2. Maintain the temporal consistency of the video. The present invention fuses similar features across frames through a fuzzy feature selection strategy, significantly enhancing the temporal consistency of the video. Compared with existing methods, the present invention can effectively avoid flickering and incoherence between frames, generating more natural and smooth video content.
[0058] 3. Improve the editing efficiency. The present invention adopts a method that does not require training, without a large amount of training data and computing resources, significantly reducing the editing cost. Compared with the fine-tuning method, the present invention is more efficient in practical applications and can quickly generate high-quality edited videos.
[0059] 4. Enhance the semantic understanding ability. The present invention can better understand the positive and negative content in the text prompt through text embedding analysis and decoupling technology. Compared with existing methods, the present invention can more accurately process negative text prompts, ensuring that the edited video content is highly consistent with the target text prompt.
[0060] 5. Support multiple editing tasks. The present invention not only supports the suppression of negative content but also can achieve general editing tasks, such as subject editing and style conversion. Compared with existing methods, the present invention has a wider application scenario and can meet diverse video editing needs.
[0061] 6. Improve the visual quality. By optimizing text embedding and feature fusion, the video generated by the present invention is more natural and coherent in visual effects. Compared with existing methods, the present invention can significantly improve the overall visual quality of the edited video, providing higher-quality editing results. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is the overall flowchart of the method for suppressing negative video content based on text embedding of the present invention.
[0063] Figure 2 It is the specific structural diagram of the fuzzy feature selection and fuzzy feature decomposition module proposed by the present invention.
[0064] Figure 3 It is the visualization result diagram of video editing of the present invention under multiple text prompts.
[0065] Figure 4 It is the visualization result diagram and detailed comparison diagram of this aspect with other methods.
[0066] Figure 5 It is the result comparison diagram of the method for suppressing negative video content based on text embedding of the present invention and other methods in terms of temporal consistency, editing accuracy, and user research, etc. DETAILED DESCRIPTION OF THE INVENTION
[0067] The negative video content suppression method based on text embedding of the present invention can effectively remove negative content in the video and maintain the temporal consistency of the video by constructing a semantic modulation module, optimizing text embedding, and fusing cross-frame features. The proposed method is applicable to various actual video editing tasks, such as theme editing and style transfer.
[0068] The invention will be further described below in conjunction with specific embodiments.
[0069] Embodiment 1:
[0070] As Figure 1 shown, the negative video content suppression method based on text embedding of the present invention is implemented through the following steps. First, given a source video and a text prompt, DDIM inversion is first performed, and then it is divided into multiple groups and fed into the diffusion model for denoising. To ensure the consistency of the video, operations of feature fusion and decomposition are added before and after the self-attention module. Specifically, for a selected group of target features, first, the cross-frame similar features are fused, the features of the target frame are used to replace these similar features, then the self-attention operation is performed on the obtained features, and finally, decomposition is performed in the fusion manner to further restore to the original frame size, so as to maintain the coherence and consistency of the video frames in the time series.
[0071] Meanwhile, for the given text prompt, it is input into the text encoder to obtain the original text embedding Then semantic adjustment is performed on it to construct an original embedding matrix composed of negative embedding and end embedding Then singular value decomposition is used to decouple the negative embedding from the original embedding matrix M, and a suppression operation is performed on the singular values to obtain a new text embedding At this time, the negative information has been removed from the text embedding. Then both the original and updated embeddings are input into the diffusion model to obtain the corresponding attention maps. The positive attention maps are pulled closer, and the negative attention maps are pulled farther, so as to achieve precise control of the video content, ensure that the positive content remains unchanged, and the negative content is effectively suppressed, and finally a high-quality edited video is output.
[0072] Figure 2 It is the specific structure diagram of the fuzzy feature selection and fuzzy feature decomposition module proposed by the present invention.
[0073] Figure 3 It is the visual effect of video editing of the present invention under various text prompts.
[0074] Figure 4 It is the visualization result diagram and detail comparison diagram of this aspect with other methods. It can be seen that the effect of our method is better than that of other methods.
[0075] Figure 5It is a comparison chart of the results of the method of the present invention and other methods in terms of temporal consistency, editing accuracy, and user research. It can be seen that the negative video content suppression method based on text embedding has higher temporal consistency and editing accuracy on various datasets than other methods.
[0076] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0077] Although the specific implementation manners of the present invention are described above, they do not limit the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts based on the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. A method for suppressing negative video content based on text embedding, characterized in that It includes the following steps: S1. Based on the source video and its text prompt, combined with the target text prompt, generate a video that meets the requirements, ensuring that the content of the edited video is consistent with the target text prompt while maintaining the temporal consistency of the video. S2. Convert the input text prompt into a text embedding through a text encoder. First, analyze the positive and negative embeddings coupled in the text embedding, further decouple the negative embedding from the text embedding through singular value decomposition, and then use an exponential suppression operator to reduce the singular value of the negative embedding to suppress the impact of negative content on the edited video. S3. Design two constraint mechanisms to further optimize the text embedding by pushing away the negative embedding and pulling in the positive embedding, ensuring that negative content is effectively suppressed while positive content remains unchanged. S4. Construct a fuzzy feature selection and fuzzy feature decomposition module for fusing similar features across frames to ensure the temporal consistency of the video. S5. Generate the target video, ensuring that the content of the edited video is consistent with the target text prompt while maintaining the temporal consistency of the video.
2. The negative video content suppression method based on text embedding according to claim 1, wherein The specific steps of step S1 include the following steps: S11. Prepare the source video, which consists of multiple consecutive images and is accompanied by corresponding text prompts, ensuring that the video content accurately presents the required visual effects. S12. Prepare the editing text prompt, which is used to drive the generation of the edited video. The main goal of the editing prompt is to clearly point out the negative content to be suppressed (such as "not wearing glasses", "not wearing a hat", etc.). Through precise text description, ensure the natural transition of the video, thereby achieving high-quality video editing effects.
3. The method for suppressing negative video content based on text embedding according to claim 1, characterized in that, The specific steps of step S2 include the following steps: S21, convert the input text prompt into a text embedding C through a text encoder Since the text embedding includes a positive embedding and a negative embedding, the text prompt needs to include suppressed negative content, i.e., C NE (such as "not wearing glasses", "not wearing a hat", etc.) and positive content, i.e., C PE . S22. Analyze the influence of positive and negative coupling information in the text embedding on the generated video, and further identify the positive embedding and the negative embedding. Create an original embedding matrix composed of the negative embedding and the end embedding Then, use singular value decomposition on it to decouple the negative embedding from the original embedding matrix M, which can be expressed as follows: M = U∑V T (1) S23. Apply an exponential suppression operator to decay the singular value of the negative embedding, thereby suppressing the impact of the negative embedding on the content of the edited video. This part can be expressed as: f(Σ) = e Σ ·Σ (2) where, ∑ represents the singular value matrix of the negative embedding. Then, we combine the modulated negative embedding singular values with the original unitary matrices U and V to obtain the updated target embedding matrix to perform singular value reconstruction, as shown in the following formula: Then, replace the original M to obtain the updated target embedding 4. The negative video content suppression method based on text embedding according to claim 1, characterized in that, The specific steps of step S3 include the following steps: S31, input the original embedding into the diffusion model to obtain the attention map at time step t wherein corresponds to the positive embedding C in the original embedding PE and the negative embedding C NE . Similarly, for the updated target embedding the corresponding attention map can also be obtained S32. In order to further suppress negative content and at the same time maintain the consistency of the non-edited area, an embedding optimization based on the attention map is designed, and two constraint mechanisms of positive content retention and negative content suppression are proposed. Positive content retention loss: Used to retain positive video content and ensure that the semantic information of the positive embedding remains unchanged. The loss function is defined as: Negative content suppression loss: Used to suppress negative video content and ensure that the impact of the negative embedding is effectively suppressed. The loss function is defined as:
5. A method for suppressing negative video content based on text embedding according to claim 1, characterized in that, The specific steps of step S4 include the following steps: S41. Divide the source video into multiple video groups, each video group containing several consecutive frames. Assume the source video contains S frames, divide it into N video groups, and each video group contains M frames, that is, S = N × M. S42, extract features from each frame in each video group, and obtain a feature set X = {X1, X2, …, X M}. Then divide the feature set into a source feature set and a target feature set where the source feature set contains the features of all frames except the target frame, and the target feature set contains the features of the target frame. Calculate the fuzzy membership degree for each feature, and use the fuzzy membership degree function to measure the similarity between features. This part can be expressed as: Among them, and represent features, and θ src and θ tar are parameters for controlling fuzzy membership degrees. S43. According to the similarity, select the most similar feature pairs for fusion to form local feature fusion X fused = Fusion(T, S) (9) where, Fusion(·) represents the fusion operation. To further enhance the long-term temporal consistency of the video, a global feature set G is maintained. The initial global feature set is the local fusion feature of the first video group For each subsequent video group i, the local fusion feature of the current video group is fused with the global feature G of the previous video group i-1 to form a new global feature G i , which is input into the self-attention module to obtain appearance and structure information. This part can be expressed as: X SA = Self_Attention(G i ) (11) Finally, according to the feature selection rule during fusion, decompose X SA back to the original frame size to restore the features of each frame for subsequent operations.
6. The negative video content suppression method based on text embedding according to claim 1, characterized in that The specific steps of step S5 include the following steps: S51. Construct the objective function. Combine the positive content retention loss and the negative content suppression loss to construct the objective function and optimize the text embedding, and gradually generate the target video that meets the editing prompt. This part can be expressed as: L = λ1L pr + λ2L ns (12) where, λ1, λ2 are hyperparameters. At S52, finally use the pre-trained decoder to decode the optimized latent space representation into the target video in the pixel space, ensuring that the content of the edited video is consistent with the target text prompt while maintaining the temporal consistency of the video.