Method for resisting high-fidelity depth video watermarking recorded by camera
By constructing a copyright watermark-synchronous watermark dual-depth video watermark framework and a cascaded diffusion coding network, combined with a simulated camera recording distortion layer, the problem of low accuracy of watermark signal extraction in camera recording scenarios is solved, and a video watermark method with high visual quality and robustness is achieved.
Patent Information
- Application Number
- CN202510762181.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-03
AI Technical Summary
In the existing camera recording scenario, high-fidelity video watermarking methods are difficult to be robust while maintaining high-quality visual effects of the video, and the accuracy of watermark signal extraction is low.
A copyright watermark-synchronous watermark dual-depth video watermarking framework is constructed, which combines a cascaded diffusion coding network and a simulated camera recording distortion layer to achieve high visual quality and high robustness of the watermarked video.
In the camera recording scenario, the extraction accuracy of the watermark signal and the visual quality of the video are improved, and the robustness and practicality of the watermark method are enhanced.
Smart Images

Figure CN120751155A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital watermarking, and in particular to a high-fidelity depth video watermarking method that is resistant to camera recording. Background Art
[0002] Digital video is now considered an important and effective medium, widely used in news reporting, short video sharing, cable network broadcasting, and other fields, and has extremely high copyright value. Video watermarking technology is a method of achieving copyright protection by robustly embedding invisible copyright identifiers. However, when faced with camera-recorded piracy, severe distortion attacks across digital and physical media cause the copyright watermark signal to be destroyed, making it difficult to trace such piracy and causing serious economic losses. To protect the copyright of digital videos, video watermarking technology needs to improve its robustness against camera-recording attacks without destroying the quality of the original carrier content, and effectively restore the copyright information after recording distortion. Therefore, high-fidelity video watermarking technology that is resistant to camera-recording attacks has become an urgent need for multimedia copyright protection.
[0003] Traditional video watermarking techniques that resist camera recording rely on manually designed features for watermark embedding and extraction. For example, the paper "Hui C, Liu S, Shi W, et al. Spatio-temporal context based adaptive camcorder recording watermarking [J]. ACM Transactions on Multimedia Computing, Communications and Applications, 2022, 18(3s): 1-25" proposes constructing a video spatiotemporal histogram to embed the watermark. They design a local matching-based algorithm and a grouped repetition embedding strategy to resist the influence of camera recording and improve extraction accuracy, but this method has a low average embedding capacity. The paper "Chang Q, Huang L, Liu S, et al. Blindrobust video watermarking based on adaptive region selection and channel reference [C]. Proceedings of the 30th ACM International Conference on Multimedia. 2022: 2344-2350" proposes a video watermarking algorithm based on adaptive region selection and channel reference. By designing a combined selection strategy of texture information and feature points, it automatically selects stable blocks for embedding that are resistant to video coding and complex attacks. The paper "Lin L, Wu D, Wang J, et al. Automatic, Robust and Blind Video Watermarking Resisting Camera Recording [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024" proposes embedding copyright watermarks into the intermediate frequency coefficients of the discrete Fourier transform spectrum of video frames to balance the watermark's invisibility and robustness. However, these methods often have limited representation capabilities and are designed only for specific distortion types, making them difficult to adapt to the complex nonlinear interference found in real-world camera recordings.
[0004] In recent years, robust video watermarking technology based on deep learning has achieved some success by combining deep neural network models with traditional video watermarking methods. For example, the paper "Luo X, Li Y, Chang H, et al. DVMark: a deep multiscale framework for video watermarking [J]. IEEE Transactions on Image Processing, 2023" uses a three-dimensional convolutional neural network to build an end-to-end video watermarking model. It replicates multiple copies of watermark information in spatial and temporal dimensions and fuses them with the carrier video to generate a watermarked video. However, this method is only effective against digital distortion and is not robust in camera recording scenarios. The paper "Zhang Y, Ni J, Su W, et al. A novel deep video watermarking framework with enhanced robustness to H.264 / AVC compression [C]. Proceedings of the 31st ACM International Conference on Multimedia. 2023: 8095-8104." proposes a method to improve its robustness against temporal attacks such as video compression and transcoding. It simulates video compression and transcoding distortion and introduces a preprocessing block to effectively align the temporal features of video frames, embedding the watermark into an imperceptible and easily extractable region. The paper "Chen L, Wang C, Zhou X, et al. Robust and Compatible Video Watermarking via Spatio-temporal Enhancement and Multiscale Pyramid Attention [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024." conducts in-depth research on the representation of long-range spatiotemporal features in videos and proposes a deep watermarking network based on spatiotemporal enhancement and multiscale pyramid attention mechanisms, which is more robust against digital distortion and visually imperceptible. However, the above literature does not fully develop designs for camera recording scenarios, and is somewhat insufficient for practical applications in complex recording scenarios.The paper "Jia J, Gao Z, Zhu D, et al. RIVIE: Robust inherent video information embedding [J]. IEEE Transactions on Multimedia, 2022, 25: 7364-7377" first proposed a deep video watermarking method that resists camera recording attacks. This method achieves robustness against camera recording attacks by applying simulated noise to the watermarked video frames during the training phase. However, this method achieves embedding by generating simulated frame sequences from a single frame using a data simulation method based on geometric transformations. Essentially, this method embeds watermark information in the spatial domain of the video frame and does not fully utilize the temporal dimension of the video to design a strategy. Therefore, the visual quality of the generated watermarked video needs to be improved.
[0005] In view of the fact that the current research on high-fidelity deep video watermarking methods for camera recording scenes is still insufficient, it is difficult to have the robustness against camera recording while maintaining high-quality visual effects of the video. Therefore, the present invention proposes a high-fidelity deep video watermarking method that is resistant to camera recording. Summary of the Invention
[0006] This paper proposes a high-fidelity deep video watermarking method that is resistant to camera recording. It constructs a dual deep video watermarking framework of copyright watermark and synchronous watermark. It also achieves high visual quality of watermarked video and high robustness of watermark extraction through a cascaded diffusion coding network and a camera recording simulated distortion layer. It solves the problems of low fidelity and low extraction accuracy of existing watermarking methods in camera recording scenarios. Specifically, it includes the following contents:
[0007] (1) A dual-depth video watermark training strategy based on copyright watermark and synchronous watermark that can resist camera recording is proposed.
[0008] (2) A spatiotemporal coding network based on cascade diffusion mechanism is proposed.
[0009] (3) A distortion layer for simulating camera recordings based on prior knowledge is proposed.
[0010] The specific contents are as follows:
[0011] (1) A dual-depth video watermark training strategy for copyright watermarking and synchronous watermarking against camera recording is proposed: the overall structure is as follows Figure 1 As shown, the copyright watermark encoding network E C , Synchronous Watermark Coding Network E S , simulated camera recording distortion layer N, copyright watermark decoding network D C and synchronized watermark decoding network D SThe input is defined as a pair (V, W) of the original video and the copyright watermark information, where V represents the original video to be embedded with the copyright watermark information and W represents the binary copyright watermark information to be embedded in the original video. The dual-depth video watermarking framework of the present invention includes an embedding stage and an extraction stage. In the embedding stage, the original video V is decomposed into three sub-channels V R 、V G With V B , design a synchronous watermark two-dimensional coding network for each frame of the original video red channel V R Implement self-encoding and output the red channel watermark video V′:V′ obtained by adding the synchronized watermark residual to the red channel of the original video R =E S (V R )+V R , and design a copyright watermark encoding network E based on diffusion C For the original video blue channel V B The feature fusion is realized with the binary copyright watermark information W, and the blue channel watermark video V′ is obtained by adding the copyright watermark residual to the blue channel of the original video. B :V′ B =E C (V B ,W)+V B Finally, the red channel watermark video, the original green channel watermark video and the blue channel watermark video are recombined in the channel dimension to obtain the watermark video V′: V′=Concat(V′ R ,V G ,V′ B ). In the extraction stage, the watermarked video V′ is subjected to distortion attack. N Decomposed into two video sub-channels V′: blue and red N-B and V′ N-R , design a synchronous watermark binary classification decoding network to decode the distorted red and blue channel video sequences synchronously, and output the synchronized blue channel watermark video V′ S-B :V′ S-B =DS(V′ N-R ,V′ N-B ), and design a copyright watermark decoding network D based on multi-scale cascade C Extract the features of the blue channel of the synchronized distorted watermark video to obtain the decoded copyright watermark W′: W′=D C (V′ S-B ).
[0012] The video watermark that resists camera recording needs to have the robustness against the time domain distortion of video recording caused by the synchronization between the camera recording frame rate and the video playback frame rate, and also needs to have the robustness against spatial distortions such as perspective deformation, illumination distortion, and moiré distortion. In order to meet the complex functional requirements of video watermarks, the copyright watermark-synchronous watermark dual-depth video watermark training strategy proposed in the present invention ensures that the copyright watermark is embedded in the original video of the carrier without affecting the timing synchronization by embedding the synchronous watermark. The present invention embeds and extracts watermarks in the blue channel of the video. Specifically, the binary cross entropy BCE is used as the copyright watermark information loss L C , the calculation formula is as follows:
[0013] L C =BCE(W,W′) (1)
[0014] Because correct synchronization is the premise of correctly extracting copyright watermark information, the present invention embeds an independent synchronization watermark in the red channel of each frame of video, trains the synchronization watermark decoder to accurately identify the video frame containing the synchronization watermark signal, and uses it to synchronize the watermark video sequence. Therefore, binary cross entropy BCE is used as the synchronization watermark information loss L S , the calculation formula is as follows:
[0015] L S =BCE(label,prediction) (2)
[0016]
[0017] Where label is the label value introduced during training. When the synchronized watermark decoder input is the original video, it should be 0, otherwise it should be 1. Prediction is the predicted value. The cross entropy loss constrains the distance between the binary classification probability output by the synchronized watermark decoder and the label value, so that the synchronized watermark decoder classifies the watermarked video frame to be synchronized as 1 as much as possible to achieve correct synchronization of the watermark frame sequence.
[0018] At the same time, the mean square error (MSE) is used as the visual loss of the blue and red channels of the watermarked video to constrain the pixel Euclidean distance between the original video and the watermarked video. The copyright watermark encoding network and the synchronous watermark encoding network are used to obtain a visually imperceptible watermarked video. The calculation formula is as follows:
[0019] L V =MSE(V B ,V′ B )+MSE(V R ,V′ R ) (4)
[0020] Therefore, the total loss of network training is obtained by weighting the copyright watermark information loss, synchronization watermark information loss and watermark video visual loss, and the calculation formula is as follows:
[0021] L=λ1L C +λ2L S +λ3L V (5)
[0022] Among them, because correctly extracting the synchronization watermark to achieve synchronization is a prerequisite for correctly extracting the copyright watermark, a larger weight should be given to the synchronization loss during the training process.
[0023] (2) A spatiotemporal coding network based on cascade diffusion mechanism is proposed: the specific cascade diffusion module structure is as follows Figure 2 As shown. Existing deep video watermarking methods generally use stacked three-dimensional convolution structures to implement watermark video encoding, hoping to use the multi-dimensional parameter sharing of the three-dimensional convolutional network to extract the spatiotemporal features of the video, but the generated watermark video has low visual quality and weak robustness. The present invention proposes a cascade diffusion mechanism, which samples video features of different spatiotemporal scales by cascading multi-dimensional three-dimensional convolution layers, and splices them with watermark information to achieve feature fusion, and realizes watermark signal diffusion by adaptive watermark information expansion. Specifically, for the input video feature vector F V After the processing of multiple downsampling residual blocks Down and the spatiotemporal attention mechanism layer TCSA in series, the spatiotemporal deep and surface features F′ of different scales of videos are sampled respectively. V =TCSA(Down(F V )); On the other hand, for the input watermark information vector F W , and the video spatiotemporal features F′ obtained by downsampling with the same scale feature dimension V Splicing is performed to obtain the fused diffusion intermediate feature F D =Concat(F′ V ,F W ), diffuse the intermediate feature F D The final diffusion feature F′ at this scale is extracted by processing the residual block with a step size of 1 and the deconvolution layer M =Conv(DeConv(F D)). The cascade diffusion structure proposed in the present invention is to perform upsampling of the watermark information while realizing downsampling of the spatiotemporal dimension features of the video, and to realize feature fusion by interlacing and splicing vectors of the same scale, so as to diffuse the watermark signal to every feature area of the video carrier. The cascade diffusion module outputs a feature vector of the same size as the original video. After the vector is spliced with the original video again, it is spatiotemporally encoded through a three-dimensional U-net network to output the watermark residual. This spatiotemporal coding method based on the cascade diffusion mechanism can fully diffuse the watermark signal in the time and space dimensions according to the diversified characteristics of the different spatiotemporal scales of the original video carrier, and can adaptively transform the signal characteristics according to the carrier characteristics while ensuring the redundant embedding of the watermark signal, thereby having better robustness and visual quality.
[0024] (3) A distortion layer based on prior knowledge of simulated camera recording is proposed: the specific structure is shown in Figure 3 . In order to solve the problem that existing methods are difficult to resist complex camera recording distortion, the present invention designs a simulated camera recording distortion layer based on prior knowledge, accurately characterizes the real camera recording distortion signal, and assists the model in learning the distortion law to minimize the impact of the real camera recording distortion in the inference stage. Specifically, the distortion layer is designed as a series connection of spatial distortion and temporal distortion, which approximately simulates the distortion introduced during the corresponding signal transmission process. Among them, the spatial distortion layer mainly simulates the noise introduced by the watermark video frame in the spatial dimension. When the display device displays, different backlight sources will affect the brightness and contrast of the media display. At the same time, the camera device will produce scaling, moiré and perspective deformation during recording. Therefore, the distortion layer uses continuous filtering operations to simulate the impact of the backlight source on the brightness and contrast of the media display, and uses two bilinear interpolations to simulate the scaling during the recording process. The following formula is used to simulate the moiré:
[0025]
[0026] M=(Min(Z1(x,y),Z2(x,y))+1) / 2 (8)
[0027] where z x With z y For the coordinates of randomly selected points, calculate M as the superimposed moiré pattern.
[0028] The time domain distortion layer mainly simulates the noise introduced by the watermark video frame in the time dimension. When the camera records the watermarked video being played, the playback frame rate is different from the recording frame rate, resulting in the phenomenon of temporal and spatial dimension disturbance in the recorded watermark video. At the same time, the watermarked video recorded and saved by the camera will be compressed and encoded, which further affects the decoder's acquisition of the watermark signal and cannot guarantee the decoding of the watermark information. Therefore, frame sampling based on uniformly distributed random frame deletion, frame averaging and frame replication is used to simulate the temporal and spatial disturbance of the recording process, and a simulated video compression network is proposed. By constructing an H.264 compressed encoding video dataset, a microscopic video compression network with multiple compression rates is obtained through pre-training. The specific structure is shown in Figure 4 .
[0029] In summary, by continuously simulating camera recording distortions in the spatial and temporal domains based on prior knowledge, the robustness of the watermarking network can be further enhanced during the training process.
[0030] Compared with other methods, the present invention has the following significant advantages:
[0031] 1. This paper solves the synchronization problem when embedding invisible watermark signals in consecutive frames through a dual deep video watermark training strategy: copyright watermark and synchronization watermark. The introduction of synchronization watermarks enables the decentralized embedding of copyright watermarks across different video frames. This design not only accurately synchronizes watermarked video frames but also helps reduce the visibility of the watermark signal, thereby improving the practicality of deep video watermarking methods.
[0032] 2. This invention designs a spatiotemporal coding network based on a cascaded diffusion mechanism, effectively improving the visual quality of watermarked videos. By extracting video features at different scales and fusing them with the upsampled watermark information to embed the copyright watermark, this method adaptively selects video features that do not affect the visual quality for embedding, minimizing the impact on the fidelity of the watermarked video and making the watermark signal less perceptible to the naked eye.
[0033] 3. The present invention simulates the camera recording distortion layer based on prior knowledge and uses data enhancement to accurately simulate the distortion in real scenes, thereby improving the robustness of the watermarked video to camera recording distortion and ensuring the reliability of the watermarking method in the application reasoning process. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is an overall schematic diagram of the present invention's "A high-fidelity deep video watermarking method that resists camera recording."
[0035] Figure 2 This is a schematic diagram of the cascade diffusion module structure of the present invention's "A high-fidelity deep video watermarking method that resists camera recording."
[0036] Figure 3 This is a schematic diagram of the camera recording distortion layer of the present invention "A high-fidelity deep video watermarking method that resists camera recording".
[0037] Figure 4 This is a schematic diagram of the simulated video compression network structure of the present invention "A high-fidelity deep video watermarking method that resists camera recording". DETAILED DESCRIPTION
[0038] In order to more clearly demonstrate the features and advantages of this patent, a detailed description of an implementation example is provided below. It should be understood that the following detailed description is intended only as an example and is intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs.
[0039] This embodiment can be implemented according to the following steps, not limited to any programming language. In this example, Python programming language is used as an example to build a model on the PyTorch deep learning framework. The specific steps are as follows:
[0040] Step 1: Dataset Preparation. We obtained 10,000 8-frame videos from the ActivityNet dataset and randomly selected 1,000 of them to form the validation set, with a validation-to-test ratio of 9:1. We then selected 50 videos from the Kinetics-400 dataset to form the test set. To accommodate the model input, we scaled the dataset videos to a uniform 128×128 size, and fixed the input watermark length to 96 bits.
[0041] Step 2: Construct a model framework. According to the present invention, a copyright watermark-synchronous watermark dual-depth video watermarking framework is constructed. The model consists of a copyright watermark encoding network, a synchronous watermark encoding network, a distortion layer that simulates camera recording, a copyright watermark decoding network, and a synchronous watermark decoding network. The copyright watermark encoding network uses a spatiotemporal coding network based on a cascade diffusion mechanism, the distortion layer simulates camera recording, the synchronous watermark encoding network uses a three-dimensional U-net network, and the copyright watermark decoding network and the synchronous watermark decoding network use a continuously downsampled three-dimensional convolutional network to synchronize watermark signals and achieve decoding. The model data items are the 96-bit information to be embedded and 8 frames of original video of size 128×128, and the output data items are the 96-bit watermark information extracted from the distorted watermarked video.
[0042] Step 3: Training the copyright watermark-synchronized watermark dual-depth video watermarking model. Training was performed on the model framework built in Step 2 with the following experimental parameters: the Adam optimizer was used, with a learning rate of 0.0001 and a learning rate decay strategy of 10% every 10 epochs. A batch size of 4 was used for training, and 25 epochs were used. During training, the weights of each loss function were set to λ1, λ2, and λ3 = 1, 5, and 10. In the simulated distortion layer, the noise intensities were set as follows: the contrast adjustment offset range was [-0.8, 1.2]; the brightness adjustment offset range was [-0.8, 1.2]; the scaling ratio was [0.5, 2]; a moiré distortion mask was generated using a simulated algorithm and added to the video; and the perspective distortion was converted to a 10-bit four-corner coordinate perturbation.
[0043] Step 4: Test the trained model. Perform copyright watermark embedding and extraction performance tests on the test set. First, embed the copyright watermark and synchronized watermark into the original video. Calculate the PSNR and SSIM of the watermarked video relative to the original video to determine the visual quality of the watermarked video. Then, record the watermarked video with a screen camera and perform watermark extraction tests at different frame rates, distances, and angles. Calculate the accuracy of watermark extraction.
[0044] In summary, for video copyright protection in camera recording scenarios, the present invention proposes a high-fidelity deep video watermarking method that is resistant to camera recording, which can generate watermarked videos with high visual quality, synchronize and decode the watermark signal under the influence of camera recording, and realize a more intelligent and high-fidelity copyright protection process, which has great practical application value.
[0045] Those skilled in the art will understand that the scope of protection of the present invention is not limited to the specific embodiments described. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features. It should be noted that the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
Claims
1. A high-fidelity deep video watermarking method that resists camera recording, characterized in that: By building a copyright watermark-synchronous watermark dual-depth video watermarking framework, we can achieve high visual quality of watermarked videos and high robustness of watermark extraction. Specifically, we include: Resist camera recording copyright watermark - synchronous watermark dual depth video watermark training strategy, copyright watermark spatiotemporal coding network based on cascade diffusion mechanism, simulated camera recording distortion layer based on prior knowledge, end-to-end training method under dual watermark embedding strategy; in the copyright watermark and synchronous watermark embedding process, the copyright watermark information W is combined with the original video blue channel V B Input the copyright watermark encoding network to obtain the blue channel watermark video V′ B , the original video red channel V R Input the synchronized watermark encoding network to obtain the red channel watermark video V′ R The red channel watermark video, the original green channel watermark video and the blue channel watermark video are recombined in the channel dimension to obtain the watermark video V′, and then the distortion layer is recorded by the simulated camera to obtain the distorted watermark video V′ N ; In the watermark decoding process, the distorted watermark video V′ N Decompose into two video sub-channels V′: blue and red N-B and V′ N-R , the synchronous watermark decoding network decodes the distorted red and blue channel video sequences synchronously, and converts the synchronized blue channel watermark video V′ S-B Input copyright watermark decoding network D based on multi-scale cascade C The decoded copyright watermark W′ is obtained; during the training process, the copyright watermark network and the synchronization watermark network are trained end-to-end simultaneously, and the optimal model parameters are obtained through the small batch gradient descent algorithm.
2. The high-fidelity deep video watermarking method for resisting camera recording according to claim 1 is characterized in that: The dual-depth video watermark training strategy of resisting copyright watermark recorded by camera and synchronous watermark realizes signal synchronization and watermark extraction by embedding different watermark information in two independent channels of the same video. Specifically, it includes: The copyright watermark information is embedded into the visually invisible position of the blue channel of the original video through the copyright watermark information loss and video visual loss constraint model, and the synchronization watermark information is embedded into the visually invisible position of the red channel of the original video, and the binary cross entropy is used as the copyright watermark information loss L C Synchronous watermark information loss L S , the training synchronized watermark decoder can accurately identify the video frames containing synchronized watermark signals to achieve the correct synchronization of the watermark frame sequence; at the same time, the mean square error is used as the visual loss L between the blue channel and the red channel of the watermark video V , constraining the pixel Euclidean distance between the original video and the watermarked video, a visually imperceptible watermarked video is obtained by the copyright watermark encoding network and the synchronous watermark encoding network.
3. The high-fidelity deep video watermarking method for resisting camera recording according to claim 1 is characterized in that: The spatiotemporal coding network based on the cascade diffusion mechanism uses a cascade diffusion module instead of a simple fully connected layer to achieve watermark signal diffusion and realize the fusion of shallow and deep features. Specifically, it includes: Input video features F V After the downsampling residual block and the spatiotemporal attention mechanism layer are processed in series, the spatiotemporal deep and surface features F′ of videos of different scales are obtained. V ; On the other side, the input watermark information vector F W The spatiotemporal features of the video F′ with the same scale characteristics as the downsampled ones V Splicing is performed to obtain the fused diffusion intermediate feature F D , diffuse the intermediate feature F D The final diffusion feature F′ at this scale is extracted by processing the residual block with a step size of 1 and the deconvolution layer M , and finally outputs a feature vector of the same size as the original video. This vector is once again spliced with the original video, and then passes through the three-dimensional U-net network for spatiotemporal encoding to output the watermark residual.
4. The high-fidelity deep video watermarking method for resisting camera recording according to claim 1 is characterized in that: The simulated camera recording distortion layer based on prior knowledge further enhances the robustness of the method by continuously simulating camera recording distortion in the spatial and temporal domains based on prior knowledge, specifically including: The distortion layer is designed as a concatenation of spatial distortion and temporal distortion. The spatial distortion layer mainly simulates the noise introduced by the watermarked video frame in the spatial dimension, namely brightness, contrast, scaling, moiré and perspective deformation; the temporal distortion layer mainly simulates the noise introduced by the watermarked video frame in the temporal dimension, namely the random frame sampling phenomenon introduced due to the difference between the playback frame rate and the recording frame rate, as well as differentiable simulated video compression coding.