Method for implanting robust watermark in potential diffusion model generated by video

Through the video watermark module and the dual-branch model distillation architecture, the inconsistency and theft of watermark embedding in video generation is solved, and efficient copyright protection and video quality maintenance are achieved.

CN120343357APending Publication Date: 2025-07-18GUANGDONG BOHUA UHD INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510535465.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively embed robust watermarks in the diffusion model generated by videos, ensuring traceability and copyright protection of video generation, and traditional methods are prone to model theft and watermark inconsistency.

Method used

Using video watermark module and dual-branch model distillation architecture, the watermark is directly implanted into the potential diffusion model generated by the video through a two-stage process. The video consistency is maintained using three-dimensional convolution and time transformers, and the watermark extractor and de-scintrating decoder are improved.

Benefits of technology

Effective embedding of watermarks in text-to-video and image-to-video generation tasks is realized, maintaining the quality of generated videos, and fighting various attacks, demonstrating the robustness and stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343357A_ABST
    Figure CN120343357A_ABST
Patent Text Reader

Abstract

The invention provides a method for implanting a robust watermark in a potential diffusion model generated by a video, which is a method for stabilizing a video signature and is a groundbreaking watermark framework of DM in video generation. The method comprises the following steps: S1, converting an original final video into a video with a watermark through a video watermark module, and extracting the watermark from the video with the watermark; and S2, providing a double-branch model distillation architecture for fine tuning the potential diffusion model generated by the video, so that the potential diffusion model generated by the video is watermarked. According to the method, the watermark is directly implanted into the DM generation process through a two-stage process; the effect of stabilizing the video signature is confirmed through comprehensive experiments of text-to-video and image-to-video generation tasks. According to the method, the watermarks are seamlessly embedded into the DM, the original function of the model can be kept, and the elasticity aiming at a series of watermark attacks is shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital media, and in particular, to a method of implanting a robust watermark in a latent diffusion model for video generation. Background Art

[0002] In the emerging field of digital media, diffusion models (DMs) have become a transformative force, especially in the generation of images and videos. Their ability to generate high-quality and diverse content has surpassed traditional generative models such as GANs and VAEs. The emergence of large-scale diffusion models such as Sora has not only marked a great success in video generation but also broadened the practical applications of DMs, including text-to-video and image-to-video tasks.

[0003] A large number of studies [1, 2, 3, 4, 5] have focused on watermarking traditional neural networks (NNs), but applying watermarks to diffusion models (DMs) represents a relatively new and cutting-edge field with limited literature and mainly focused on tasks in image generation. Uchida et al. [3] proposed a method of embedding a watermark in the parameters of a neural network model to ensure the integrity and ownership of the model without affecting its performance. The novelty of this method is that it can directly embed the watermark into the weights of the model, making it robust to model compression and fine-tuning. There are also some methods aimed at adding watermarks to stable diffusion models. The Stable Signature technique

[11] adds watermarks to the image diffusion model. This method quickly fine-tunes the latent decoder of the image generator under binary signature conditions. A pre-trained watermark extractor recovers the hidden signature from any generated image and then determines its origin through statistical tests. Wen et al. [6] introduced the Tree Ring watermark, which is a method of embedding an invisible and robust watermark into the output of a diffusion model by cleverly influencing the sampling process and embedding patterns into the initial noise vector, ensuring resistance to common image transformations.

[0004] From a technical approach perspective, video watermarking techniques can be mainly divided into two categories: traditional watermarking techniques and deep learning-based watermarking techniques. On the one hand, traditional methods for adding watermarks to videos are divided into three categories: transform domain methods (TDM), spatial domain methods (SDM), and frequency domain methods (FDM). TDM applies mathematical transforms to video frames and embeds watermarks by modifying the transform coefficients in a way that is imperceptible to the human eye. SDM embeds watermarks by directly manipulating the pixel values of video frames by modifying the pixel values. FDM uses techniques such as the fast Fourier transform (FFT) to transform video frames into the frequency domain and embeds watermarks by modifying certain amplitudes and phases in the frequency domain. On the other hand, deep learning models can adjust the watermark embedding strategy according to video content features and usually have higher robustness. A new video compression differentiable simulator called DiffH264 has been developed, which can approximately simulate the H.264 / AVC compression process, enabling the network to withstand H.264 / AVC compression. A robust blind watermarking algorithm using adaptive region selection and channel reference has been proposed to avoid being destroyed during video coding and complex attacks [7]. [7] has demonstrated the capabilities of deep learning methods in the field of video watermarking.

[0005] The main problems existing in the prior art are: Although digital video content can be quickly copied and disseminated today, it is also crucial to ensure the traceability and ownership of videos generated by DM. The victory of DM in video generation has brought huge compliance challenges, including copyright and model protection. Different from previous works that focused on the copyright protection of neural networks in critical tasks, the discussion on the efficacy of watermarks in DM (especially for video generation) is still very scarce.

[0006] The difficulties in solving the above problems are: Summarize the challenges of model protection as follows: 1. Simple disconnection and self-organization methods cannot add watermarks in an end-to-end manner, which will lead to easier model theft and vulnerabilities; 2. Adding watermarks to videos without considering the challenge of temporal consistency, integrating watermarks into the generation process while avoiding flickering or inconsistency.

[0007] The significance of solving the above problems is: To improve the robustness of the model. Summary of the Invention

[0008] The present invention provides a method for implanting a robust watermark in a latent diffusion model for video generation to improve the robustness of the model The technical solution of the present invention is as follows: The method for implanting a robust watermark in a latent diffusion model for video generation according to the present invention includes the following steps: S1. Through a video watermark module, convert the original final video into a watermarked video, and extract the watermark from the watermarked video; and S2. Provide a dual-branch model distillation architecture for fine-tuning the latent diffusion model for video generation to make the latent diffusion model for video generation carry the watermark.

[0009] Optionally, in the method for implanting a robust watermark in the latent diffusion model for video generation described above, in step S1, the framework of the video watermark module consists of two components, an embedded watermark module and a watermark extractor.

[0010] Optionally, in the method for implanting a robust watermark in the latent diffusion model for video generation described above, the embedded watermark module is used to input the video and a randomly generated bitstream. After passing through the embedded watermark module, the original final video is converted into a watermarked video.

[0011] Optionally, in the method for implanting a robust watermark in the latent diffusion model for video generation described above, first fix the time dimension of the video, divide the video into spatial characters, and use a spatial transformer to extract spatial character features; then, fix the spatial dimension, extract temporal characters in the same space, and then pass them to a temporal encoder to interact with characters at the same position at different times.

[0012] Optionally, in the method for implanting a robust watermark in the latent diffusion model for video generation described above, train the watermark extractor, and train a decoding network to extract the watermark from the watermarked video; the decoding network needs to be symmetric with the encoding network.

[0013] Optionally, in the method for implanting a robust watermark in the latent diffusion model for video generation described above, in step S2, the dual-branch model distillation architecture includes an upper branch and a lower branch. First, encode through a VAE encoder to obtain latent variables; at the same time, keep the parameters of the de-flickering decoder in the lower branch frozen; the output of the same de-flickering decoder in the upper branch passes through a temporal transformer to ensure consistency; subsequently, the outputs of the upper branch and the lower branch are merged through weighted concatenation to generate a final output, and it is evaluated against a fidelity loss function.

[0014] According to the technical solution of the present invention, the beneficial effects are: The present invention has conducted extensive experiments in text-to-video and image-to-video generation tasks, verifying the effectiveness of embedding watermarks in various DMs. It further demonstrates that the functionality of the model remains undamaged after watermarking and showcases the robustness of the watermarks of the present invention against various attacks. The method of the present invention (Stable Video Stamp, SVS) is the first non-post hoc method to add watermarks to DMs for video generation. SVS is a general framework that can implant watermarks in DMs through efficient fine-tuning. To maintain the quality of the generated videos after watermark implantation, three-dimensional convolutions are used to adapt to the nature of DMs, which can decode videos as a whole, and temporal attention blocks are introduced to better capture the consistency between frames. Extensive experiments on text-to-video and image-to-video generation tasks show that SVS can effectively implant watermarks in DMs without degrading the quality of the generated videos while maintaining robustness against various watermark attacks.

[0015] To better understand and illustrate the concept, working principle, and inventive effect of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and through specific embodiments as follows: Description of the Drawings

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings required for describing the specific embodiments or the prior art will be briefly introduced below.

[0017] Figure 1 It is a flowchart of the steps of the method for implanting robust watermarks in the latent diffusion model for video generation of the present invention; Figure 2 It is a schematic structural diagram of the video watermark module involved in the method of the present invention; Figure 3 It is a schematic diagram of the dual-branch model distillation architecture involved in the method of the present invention; Figure 4 It is the experimental result generated using the generative model (embedded with watermarks using the SVS algorithm); Figure 5 It is the experimental result of the resistance of the model of the present invention to various attacks. Detailed Description of the Embodiments

[0018] To make the purpose, technical method, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific examples. These examples are merely illustrative and not limiting to the present invention.

[0019] The method of implanting robust watermarks in the latent diffusion model for video generation of the present invention is a method for stable video signature, which is a pioneering watermarking framework for DM in video generation. To meet the urgent needs of copyright and model protection, the method of the present invention directly implants watermarks into the DM generation process for the first time through a two-stage process. The present invention confirms the efficacy of stable video signature through comprehensive experiments on text-to-video and image-to-video generation tasks. Further, it shows that this framework seamlessly embeds watermarks into the DM, can maintain the original function of the model, and exhibits resilience against a series of watermark attacks.

[0020] The method of implanting robust watermarks in the latent diffusion model for video generation of the present invention is a method for stable video signature, which incorporates watermarks into the video generation process itself without any architectural changes. Stable video tagging adopts a two-stage (S1 and S2) model watermarking method, which specifically includes the following steps ( Figure 1 as shown): S1. Through the video watermark module, convert the original final video into a watermarked video, and extract the watermark from the watermarked video.

[0021] Based on the HiDDeN method [8], obtain the watermark encoder and the corresponding extractor of the video watermark module (Video-Lock), and specifically refer to Figure 2 . The goal of this network is to extract and decode the watermark information in the image by training the model. During the training process, the watermark information is input into the network together with the given watermark code (m) and the original image. After the image features pass through the convolutional layer, the features are input into the spatial transformer by fixing the time dimension, and then the time characters are input into the temporal transformer by fixing the spatial dimension.

[0022] To be consistent with the overall decoding mechanism of the DM, three-dimensional convolution is used to encode the watermark into the overall latent space embedding of the video. In addition, a temporal transformation block is introduced to capture the temporal consistency during watermarking, thereby maintaining the quality of the generated video. This method is called video locking.

[0023] The framework of the video watermark module (Video-Lock) consists of two components: the watermark embedding module and the watermark extractor. For specific details, please refer to Figure 2 .

[0024] Watermark embedding module: The input consists of a video and a randomly generated bitstream. After passing through the watermark embedding module, the original final video is converted into a watermarked video. The present invention believes that a complex network architecture is not required for video watermarking. Therefore, the present invention uses simple 3D convolution to extract features from the input video. To embed the watermark into the video more effectively, a time transformer is introduced into the encoding process. The present invention first fixes the time dimension of the video, divides the video into spatial characters, and uses a spatial transformer to extract the spatial character features. Then, the spatial dimension is fixed, and temporal characters are extracted from the same space and then passed to the temporal encoder to interact with the characters at the same position at different times. The time Transformer can endow the model of the present invention with global consistency, ensuring the stability and reliability of the watermark information of the present invention.

[0025] Watermark extractor: Train a watermark extractor, which trains a decoding network to extract the watermark from the processed video (the watermarked video). The decoding network needs to be symmetric with the encoding network because the network encodes the representation of the watermark into the video that the decoding network can learn better.

[0026] During the training stage of the watermark extractor, to ensure the robustness of the watermark, common watermark attack methods are used, such as compression, cropping, and rotation. These attack methods are combined. During the training stage, the input is the encoded video watermark, and the output is the video after being attacked by various methods. It should be noted that the resolution of the video may change after being attacked. Some of these attacks are non-differentiable. Therefore, the present invention uses two losses for learning: one loss is the normal output without attack, which is used to fine-tune the encoding and decoding processes; the other loss is used to extract the watermark from the decoded image after being attacked, which is used to fine-tune the decoding process.

[0027] S2. Provide a dual-branch model distillation architecture for fine-tuning the latent diffusion model for video generation, so that the latent diffusion model for video generation is watermarked.

[0028] The pipeline of the dual-branch model distillation architecture is as Figure 3 shown.

[0029] The leftmost video is first encoded by the VAE encoder to obtain the latent variable Latent. Meanwhile, the parameters of the de-flickering decoder in the lower branch are kept frozen. The output of the same de-flickering decoder in the upper branch passes through the temporal transformer to ensure consistency. Subsequently, the outputs of the two branches are combined through weighted concatenation to generate the final output, which is evaluated against the fidelity loss function. In addition, the output of the upper branch also extracts the watermark through the watermark extractor, and calculates the loss Loss with the predefined watermark information. The dual-branch model distillation architecture includes an upper branch and a lower branch, with de-flickering decoders provided in both branches to mitigate potential artifacts, and a temporal transformer is introduced in the upper branch to further ensure consistency between frames. In addition, watermark integrity is maintained by extracting the watermark and comparing it with the predefined watermark, which contributes to the robustness of the potential diffusion model for video generation.

[0030] The decoding structure of the de-flickering decoder of the present invention is the same as that of the decoder of SVD. Different from the traditional decoder that focuses on reconnaissance to construct an image from the latent representation, this decoder is customized for video generation. Temporal continuity means that the decoder considers the consistency of video frames. This is crucial for generating a smooth video sequence where consecutive frames are consistent with each other, avoiding sudden changes or "flickering", and this process is called "de-flickering". The "de-flickering" aspect addresses a common challenge in video generation where the flickering effect may occur due to inconsistencies between frames. This de-flickering decoder is designed to minimize such flickering, ensuring a more stable and smooth transition across frames. The present invention also proposes a temporal Transformer, which has a globally consistent attention mechanism capable of remote demodulation pending, and can generate a smoother video. However, this method reduces video clarity. To maintain clarity, the present invention connects the smooth video output from the temporal Transformer with the original video to form the final decoded output.

[0031] Robustness evaluation: Another way to generate a watermarked video is to add a watermark after the video is generated. However, there are various ways to apply watermarks to images or videos in an end-to-end manner. Finally, three commonly used watermarking methods in the generative model must be selected: invisible-watermark, blind-watermark, and HiDDeN. Among these methods, invisible watermarking is used to watermark the videos generated by stable video diffusion. Blind watermarking and invisible watermarking are open-source watermarking methods based on DWT-DCT-SVD (DDS) and dwtDct. The present invention compares the accuracy of different attack scenarios under the same video and payload. The comparison of different networks is shown in Table 1.

[0032] Table 1 Comparison of different networks under different attack scenarios with the same video and payload The present invention evaluates the robustness of the model using various attack methods, including cropping, scaling and rotation, blurring, and compression. Specifically, for the cropping operation, the frame is cropped centered and the width and height are cropped by a ratio p. For the scaling and rotation operations, the video is scaled as a whole. For resizing, the width and height of the frame are scaled by a ratio p. For blurring, Gaussian blur is applied over the entire image to simulate the effect of blurring. Finally, for the compression operation, the obtained tensor is saved as an image and then read out again for image compression. To verify the resistance of the model of the present invention to various attacks, a large number of videos are generated and attacked with different intensities. The final experimental results are plotted as a line graph, as shown in Figure 5 shown. The method of attacking video frames is also introduced. Frame swapping means that the present invention randomly rearranges the order of N groups of adjacent frames. Frame dropping means randomly deleting several frames and replacing them with nearby frames. In noise, v represents the variance of the added Gaussian noise.

[0033] The SVS algorithm of the present invention is used to add watermarks to the video generation model, and then the marked generation model is used to generate videos. Figure 4 The generation results of text-to-video and image-to-video are shown. By comparing the watermarked videos generated by stable video signatures with the original videos, the watermarked videos generated by the method of the present invention are comparable in quality to the original watermark-free videos. User evaluation: Five pairs of videos generated by the original model and the watermark model are randomly selected, and users are invited to evaluate them. Three of the five video sets are generated using stable video diffusion with randomly selected wave images as inputs, while the other two videos are generated using the AD model with random texts from the test cases in the AD paper as inputs. To ensure fairness, users are not informed whether the videos contain watermarks. The data provided in Table 2 shows that the videos generated by the fine-tuned model of the present invention are not inferior to those generated by the original model. In the user study of the present invention, the respondents are required to select the best video set. Most users think there is no difference between the videos generated by the original model and those generated by the watermark model. Some participants even suggest that the fine-tuned videos are smoother and more consistent with the inputs because the present invention uses a temporal Transformer to make the videos smoother.

[0034] Table 2 Comparison of video parameters generated by the model of the present invention and the original model The present invention provides a method of integrating a watermark into the video generation process itself without any architectural changes. The method adopts a two-stage model watermarking method. In the first stage, inspired by HiDDeN, a watermark encoder and a corresponding extractor are obtained through pre-training; in order to be consistent with the overall decoding mechanism of DM, 3DCNN is used to encode the watermark into the overall latent space embedding of the video; in addition, a temporal transformer block is introduced to capture the temporal consistency during the watermarking process, thereby maintaining the quality of the generated video. In the second stage, a dual-branch model distillation architecture is proposed, and a de-jitter decoder is designed in both branches to mitigate potential artifacts, while the same temporal Transformer block with state 1 is introduced in the upper branch to further ensure the consistency between frames; in addition, the integrity of the watermark is maintained by extracting the watermark and comparing it with a predefined watermark, which helps to improve the robustness of the model.

[0035] Extensive experiments have been carried out in text-to-video and image-to-video generation tasks to verify the effectiveness of the method of the present invention in embedding watermarks in various DMs. It is further proved that the functions of the model are still not affected after watermarking, and the robustness of the watermark of the present invention against various attacks is demonstrated.

[0036] The above description is the best embodiment according to the concept and working principle of the invention. The above embodiments should not be construed as limiting the protection scope of the claims. Combinations of other implementation manners and implementation modes according to the concept of the present invention all fall within the protection scope of the present invention.

[0037] References [1] Franziska Boenisch. 2021. A systematic review on model watermarking for neural networks. Frontiers in big Data 4 (2021), 729663. [2] Yuki Nagai, Yusuke Uchida, Shigeyuki Sakazawa, and Shin’ichi Satoh. 2018. Digital watermarking for deep neural networks. International Journal of Multimedia Information Retrieval 7 (2018), 3–16. [3] Yusuke Uchida, Yuki Nagai, Shigeyuki Sakazawa, and Shin’ichiSatoh. 2017. Embedding watermarks into deep neural networks. In Proceedingsof the 2017 ACM on international conference on multimedia retrieval. 269–277. [4] Hanzhou Wu, Gen Liu, Yuwei Yao, and Xinpeng Zhang. 2020.Watermarking neural networks with watermarked images. IEEE Transactions onCircuits and Systems for Video Technology 31, 7 (2020), 2591–2601. [5] Jialong Zhang, Zhongshu Gu, Jiyong Jang, Hui Wu, Marc PhStoecklin, Heqing Huang, and Ian Molloy. 2018. Protecting intellectualproperty of deep neural networks with watermarking. In Proceedings of the2018 on Asia conference on computer and communications security. 159–172. [6] Yuxin Wen, John Kirchenbauer, Jonas Geiping, and Tom Goldstein.2023. Tree-ring watermarks: Fingerprints for diffusion images that areinvisible and robust. arXiv preprint arXiv:2305.20030 (2023). [7]Qinwei Chang, Leichao Huang, Shaoteng Liu, Hualuo Liu, TianshuYang, and Yexin Wang. 2022. Blind robust video watermarking based on adaptiveregion selection and channel reference. In Proceedings of the 30th ACMInternational Conference on Multimedia. 2344–2350.

[41] Jiren Zhu, Russell Kaplan, Justin Johnson, and Li Fei-Fei. 2018.Hidden: Hiding data with deep networks. In Proceedings of the Europeanconference on computer vision (ECCV). 657–672.

Claims

1. A method for implanting a robust watermark in a latent diffusion model for video generation, characterized in that, Including the following steps: S1. Through the video watermark module, convert the original final video into a watermarked video, and extract the watermark from the watermarked video; And S2. Provide a two-branch model distillation architecture for fine-tuning the latent diffusion model for video generation, so that the latent diffusion model for video generation is watermarked.

2. The method for implanting a robust watermark in a latent diffusion model for video generation according to claim 1, characterized in that, In step S1, the framework of the video watermark module consists of two components: an embedded watermark module and a watermark extractor.

3. The method of implanting a robust watermark in a latent diffusion model for video generation according to claim 2, wherein The embedded watermark module is used to input the video and a randomly generated bitstream. After passing through the embedded watermark module, the original final video is converted into a watermarked video.

4. The method of implanting a robust watermark in a latent diffusion model for video generation according to claim 3, wherein First, fix the time dimension of the video, divide the video into spatial characters, and use a spatial transformer to extract spatial character features; then, fix the spatial dimension, extract temporal characters in the same space, and then pass them to the temporal encoder to interact with characters at the same position at different times.

5. The method for implanting a robust watermark in a latent diffusion model for video generation according to claim 1, wherein, Train the watermark extractor to train a decoding network to extract the watermark from the watermarked video; the decoding network needs to be symmetric with the encoding network.

6. The method for implanting a robust watermark in a latent diffusion model for video generation according to claim 1, wherein, In step S2, the two-branch model distillation architecture includes an upper branch and a lower branch. First, encode through a VAE encoder to obtain latent variables; at the same time, the parameters of the deblinking decoder in the lower branch are kept frozen; the output of the same deblinking decoder in the upper branch passes through a temporal transformer to ensure consistency; Subsequently, the outputs of the upper branch and the lower branch are merged through weighted concatenation to generate a final output, and evaluated against a fidelity loss function.