Deep learning image watermarking method based on ViT and diffusion probability model
By combining the visual conversion network (ViT) and diffusion probability model in image watermark technology, the balance problem between invisibility and robustness of deep learning image watermark technology is solved, and efficient, hidden and attack-resistant watermark embedding and extraction effects are achieved.
Patent Information
- Application Number
- CN202411965461.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
Deep learning-based digital image watermarking technology faces the problem of balancing invisibility and robustness, designing network structures to adapt to changing attack methods, and adding watermarking efficiency to large-scale images.
The deep learning image watermark method based on ViT and diffusion probability model is adopted. By embedding watermark information in the host image, and using ViT's global feature extraction ability and the denoising ability of the diffusion probability model, the watermark's concealment and attack resistance are achieved.
Effectively hiding the watermark in the global feature space of the host image enhances the concealment of the watermark, while improving the robustness of the watermark against various noise and signal processing attacks, and improving the efficiency and accuracy of watermark extraction.
Smart Images

Figure CN119991394A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital image forensics, and in particular relates to a deep learning image watermarking method based on ViT and a diffusion probability model. Background Art
[0002] In today's era of rapid digitalization and networking, digital watermarking technology has become an important means to ensure image copyright and authenticate authenticity. With the rise of deep learning, digital image watermarking technology based on deep learning has become a research hotspot. Deep learning methods use multi-layer neural networks to learn complex patterns of data and efficiently encode image content to achieve watermark embedding and extraction. They can not only resist conventional signal processing attacks such as compression, cropping and filtering, but also adapt to more complex tampering, such as deep forgery and AI-driven image editing. For example, using the powerful feature extraction capabilities of convolutional neural networks (CNNs), watermarks can be embedded in the feature space of images, so that even after the image has been complexly processed, the watermark information can still be accurately restored. In addition, some studies have also explored adversarial learning, which continuously learns and adapts to various attacks through adversarial networks, thereby improving the robustness of watermarking technology.
[0003] However, digital image watermarking technology based on deep learning also faces challenges, including how to strike a balance between invisibility and robustness, how to design network structures to adapt to ever-changing attack methods, and the efficiency of batch watermarking of large-scale images. Summary of the invention
[0004] The present invention aims to solve the above-mentioned technical problems existing in the prior art and provides a deep learning image watermarking method based on ViT and diffusion probability model.
[0005] The technical solution of the present invention is: a deep learning image watermarking method based on ViT and diffusion probability model, which is carried out according to the following steps:
[0006] Convention: I is the host image; W is the pre-embedded watermark information; I′ is the watermarked image; W′ is the watermark information extracted from I′; W C is the weight of the last convolution layer before watermark embedding; L I is the loss of the host image and the watermarked image; L W is the loss of the original watermark and the extracted watermark; L C is the regularization loss of the convolutional layer; n , n=1,2,3 is L I , L W , L C The intensity coefficient of the loss term; F n , n=1,2,3... are the image features embedded in each stage; Fn ′, n=1,2,3... are used to extract image features at each stage; U-Net is a convolution-based skip connection network structure; ViT is a visual Transformer-based network structure;
[0007] a. Initial Setup
[0008] Convert the host image to 256×256×3 and the watermark information to 32×32×1.
[0009] b. Watermark Embedding
[0010] b.1 Transform the host image I through a 4-layer U-Net skip connection network to obtain feature-level information F1 corresponding to the host image;
[0011] b.2 The watermark information W is transformed and convolved into 96-channel feature-level watermark information F2 through a 4-layer U-Net skip connection network;
[0012] b.3 Reshape the feature-level watermark information F2 and convert it into a 256×256×1 watermark feature F3;
[0013] b.4 Concatenate the features {F1, F3} into a 256×256×4 watermarked feature F4;
[0014] b.5 Divide the feature F4 into blocks of size 32×32×4, and perform a Flatten operation on each block to construct a block feature F5 that conforms to the Transform input;
[0015] b.6 Pass the block feature F5 through the Transform block with two layers of skip connections, and save the generated results into the temporary list ViTs;
[0016] b.7 Repeat b.5 and b.6 twice;
[0017] b.8 Concatenate the three Transform features saved in ViTs, perform a 3×3 convolution on the Concatenate result to generate the final watermark feature F6, and record the parameters of the convolution layer as W C ;
[0018] b.9 Concatenate the feature F6 and the host image I and pass them into the 1×1 convolution layer, and apply the Sigmoid activation function to generate the watermarked image I′;
[0019] c. Watermark extraction
[0020] c.1 Perform random data enhancement operation on the watermarked image I′;
[0021] c.2 Perform a diffusion process on the enhanced image and add standard Gaussian noise of different intensities for multiple times to obtain a high-noise watermarked image F1′;
[0022] c.3 Divide the high-noise watermark image F1′ into blocks of size 32×32×4, and perform a Flatten operation on each block to construct a block feature F2′ that meets the Transform input;
[0023] c.4 The block feature F2′ is transferred into the multi-layer ViT structure for reverse diffusion to remove noise and obtain the low-noise feature F3′;
[0024] c.5 Perform 3×3 convolution and 2×2 pooling operations on F3′ to obtain the feature ViT′;
[0025] c.6 Reshape ViT′ to convert the sampling features of the watermark into the watermark features F4′ that are consistent with the original watermark;
[0026] c.7 The feature F4′ is passed through a 4-layer skip-connected U-Net block to generate a 32×32×1 watermark feature, and finally a sigmoid activation is applied to the feature to obtain the final extracted watermark W′;
[0027] d. Loss function
[0028] d.1 Loss term of host image and watermarked image:
[0029]
[0030] d.2 Loss term of original watermark and extracted watermark:
[0031]
[0032] d.3 Regularization loss term of convolutional layer: L C =||W C ||2;
[0033] d.4 Model training loss: L = λ1L I +λ2L W +λ3L C ;
[0034] e. End-to-end model training
[0035] e.1 Adjust the host image I and watermark information W to the specified input size;
[0036] e.2 The host image I and the watermark information W are constructed into a watermarked image I′ through steps b.1-b.9;
[0037] e.3. Pass I′ through steps c.1-c.6 to obtain the watermark information W′ extracted after the attack;
[0038] e.4 Based on the results of e.1-e.3, obtain the sample loss function L through steps d.1-d.3;
[0039] e.5 Use the optimizer to update the model parameters according to the loss function L;
[0040] e.6 Select the next sample and repeat e.1-e.5 until the loss L converges and the model training is completed.
[0041] The present invention adopts a deep learning image watermarking technology that combines the visual transformation network (ViT) and the diffusion probability model. It mainly embeds the watermark information in the host image and can effectively extract the watermark after experiencing potential attacks. The use of ViT can highlight its global feature processing capabilities and ensure that the watermark is effectively embedded and encoded in the global feature space of the image; in addition, the application of the diffusion probability model can target potential attacks such as simulated Gaussian noise to enhance the concealment and anti-attack capabilities of the watermark.
[0042] Specifically, the basic steps of the strategy of the present invention include: first, using a multi-layer U-Net network and Transform blocks to encode the host image and watermark information at the feature level; then, the watermark features are fused through the Concatenate operation, and random data enhancement and Gaussian noise processing are applied to the encoded image to simulate and resist potential attacks. In the watermark extraction stage, the Transform block is used to build a denoising model of the probability diffusion model, and finally the accurate extraction of the watermark is achieved through the U-Net network and the sigmoid activation function.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] First, the present invention utilizes the global feature extraction capability of the visual transformation network (ViT) to effectively hide the watermark information in the global feature space of the host image, ensuring the visual invisibility of the watermark while maintaining the natural appearance of the image, thereby enhancing the concealment of the watermark;
[0045] Second, by simulating random noise in the diffusion process, continuous noise attacks are added to the watermark embedding process, thereby simulating potential image damage. And the inverse probability feature model is constructed through the ViT network. This method enhances the robustness of the watermark to various noise and signal processing attacks;
[0046] Third, in the process of model construction, the repeated use of jump connection technology helps to retain important feature information of the image, reduces the information loss that may be caused by continuous convolution and pooling operations, and reduces the training complexity of the model. At the same time, the use of jump connections also improves the efficiency and accuracy of the entire watermark extraction process, ensuring accurate watermark extraction in a shorter time with lower computational cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Schematic diagram of the network structure of the watermark embedding model in an embodiment of the present invention.
[0048] Figure 2 Schematic diagram of the network structure of the watermark extraction model in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The deep learning image watermarking method based on ViT and diffusion probability model of the present invention is carried out according to the following steps:
[0050] Convention: I is the host image; W is the pre-embedded watermark information; I′ is the watermarked image; W′ is the watermark information extracted from I′; W C is the weight of the last convolution layer before watermark embedding; L I is the loss of the host image and the watermarked image; L W is the loss of the original watermark and the extracted watermark; L C is the regularization loss of the convolutional layer; n , n=1,2,3 is L I , L W , L C The intensity coefficient of the loss term; F n , n=1,2,3... are the image features embedded in each stage; F n ′, n=1,2,3... are used to extract image features at each stage; U-Net is a convolution-based skip connection network structure; ViT is a visual Transformer-based network structure;
[0051] a. Initial Setup
[0052] Convert the host image to 256×256×3 and the watermark information to 32×32×1.
[0053] b. Watermark embedding, according to Figure 1 Perform the following steps as shown;
[0054] b.1 Transform the host image I through a 4-layer U-Net skip connection network to obtain feature-level information F1 corresponding to the host image;
[0055] b.2 The watermark information W is transformed and convolved into 96-channel feature-level watermark information F2 through a 4-layer U-Net skip connection network;
[0056] b.3 Reshape the feature-level watermark information F2 and convert it into a 256×256×1 watermark feature F3;
[0057] b.4 Concatenate the features {F1, F3} into a 256×256×4 watermarked feature F4;
[0058] b.5 Divide the feature F4 into blocks of size 32×32×4, and perform a Flatten operation on each block to construct a block feature F5 that conforms to the Transform input;
[0059] b.6 Pass the block feature F5 through the Transform block with two layers of skip connections, and save the generated results into the temporary list ViTs;
[0060] b.7 Repeat b.5 and b.6 twice;
[0061] b.8 Concatenate the three Transform features saved in ViTs, perform a 3×3 convolution on the Concatenate result to generate the final watermark feature F6, and record the parameters of the convolution layer as W C ;
[0062] b.9 Concatenate the feature F6 and the host image I and pass them into the 1×1 convolution layer, and apply the Sigmoid activation function to generate the watermarked image I′;
[0063] c. Watermark extraction, such as Figure 2 Perform the following steps as shown:
[0064] c.1 Perform random data enhancement operation on the watermarked image I′;
[0065] c.2 Perform a diffusion process on the enhanced image and add standard Gaussian noise of different intensities for multiple times to obtain a high-noise watermarked image F1′;
[0066] c.3 Divide the high-noise watermark image F1′ into blocks of size 32×32×4, and perform Flatten operation on each block to construct the block feature F2′ that meets the Transform input;
[0067] c.4 The block feature F2′ is transferred into the multi-layer ViT structure for reverse diffusion to remove noise and obtain the low-noise feature F3′;
[0068] c.5 Perform 3×3 convolution and 2×2 pooling operations on F3′ to obtain the feature ViT′;
[0069] c.6 Reshape ViT′ to convert the sampled features of the watermark into the watermark features F4′ that are consistent with the original watermark;
[0070] c.7 The feature F4′ is passed through a 4-layer skip-connected U-Net block to generate a 32×32×1 watermark feature, and finally a sigmoid activation is applied to the feature to obtain the final extracted watermark W′;
[0071] d. Loss function
[0072] d.1 Loss term of host image and watermarked image:
[0073] d.2 Loss term of original watermark and extracted watermark:
[0074]
[0075] d.3 Regularization loss term of convolutional layer: L C =||W C ||2;
[0076] d.3 Model training loss: L = λ1L I +λ2L W +λ3L C ;
[0077] e. End-to-end model training
[0078] e.1 Adjust the host image I and watermark information W to the specified input size;
[0079] e.2 The host image I and the watermark information W are constructed into a watermarked image I′ through steps b.1-b.9;
[0080] e.3. Pass I′ through steps c.1-c.6 to obtain the watermark information W′ extracted after the attack;
[0081] e.4 Based on the results of e.1-e.3, obtain the sample loss function L through steps d.1-d.3;
[0082] e.5 Use the optimizer to update the model parameters according to the loss function L;
[0083] e.6 Select the next sample and repeat e.1-e.5 until the loss L converges and the model training is completed.
Claims
1. A deep learning image watermarking method based on ViT and diffusion probability model, characterized by Follow these steps: Convention: I is the host image; W is the pre-embedded watermark information; I′ is the watermarked image; W′ is the watermark information extracted from I′; W C is the weight of the last convolution layer before watermark embedding; L I is the loss of the host image and the watermarked image; L W is the loss of the original watermark and the extracted watermark; L C is the regularization loss of the convolutional layer; n , n=1,2,3 is L I , L W , L C The intensity coefficient of the loss term; F n , n=1,2,3... are the image features embedded in each stage; F n ′, n=1,2,3... are used to extract image features at each stage; U-Net is a convolution-based skip connection network structure; ViT is a visual Transformer-based network structure; a. Initial Setup Convert the host image to 256×256×3 and the watermark information to 32×32×1. b. Watermark Embedding b.1 Transform the host image I through a 4-layer U-Net skip connection network to obtain feature-level information F1 corresponding to the host image; b.2 The watermark information W is transformed and convolved into 96-channel feature-level watermark information F2 through a 4-layer U-Net skip connection network; b.3 Reshape the feature-level watermark information F2 and convert it into a 256×256×1 watermark feature F3; b.4 Concatenate the features {F1, F3} into a 256×256×4 watermarked feature F4; b.5 Divide the feature F4 into blocks of size 32×32×4, and perform a Flatten operation on each block to construct a block feature F5 that conforms to the Transform input; b.6 Pass the block feature F5 through the Transform block with two layers of skip connections, and save the generated results into the temporary list ViTs; b.7 Repeat b.5 and b.6 twice; b.8 Concatenate the three Transform features saved in ViTs, perform a 3×3 convolution on the Concatenate result to generate the final watermark feature F6, and record the parameters of the convolution layer as W C ; b.9 Concatenate the feature F6 and the host image I and pass them into the 1×1 convolution layer, and apply the Sigmoid activation function to generate the watermarked image I′; c. Watermark extraction c.1 Perform random data enhancement operation on the watermarked image I′; c.2 Perform a diffusion process on the enhanced image and add standard Gaussian noise of different intensities for multiple times to obtain a high-noise watermarked image F1′; c.3 Divide the high-noise watermark image F1′ into blocks of size 32×32×4, and perform Flatten operation on each block to construct the block feature F2′ that meets the Transform input; c.4 The block feature F2′ is transferred into the multi-layer ViT structure for reverse diffusion to remove noise and obtain the low-noise feature F3′; c.5 Perform 3×3 convolution and 2×2 pooling operations on F3′ to obtain the feature ViT′; c.6 Reshape ViT′ to convert the sampled features of the watermark into the watermark features F4′ that are consistent with the original watermark; c.7 The feature F4′ is passed through a 4-layer skip-connected U-Net block to generate a 32×32×1 watermark feature, and finally a sigmoid activation is applied to the feature to obtain the final extracted watermark W′; d. Loss function d.1 Loss term of host image and watermarked image: d.2 Loss term of original watermark and extracted watermark: d.3 Regularization loss term of convolutional layer: L C =||W C ||2; d.4 Model training loss: L = λ1L I +λ2L W +λ3L C ; e. End-to-end model training e.1 Adjust the host image I and watermark information W to the specified input size; e.2 The host image I and the watermark information W are constructed into a watermarked image I′ through steps b.1-b.9; e.
3. Pass I′ through steps c.1-c.6 to obtain the watermark information W′ extracted after the attack; e.4 Based on the results of e.1-e.3, obtain the sample loss function L through steps d.1-d.3; e.5 Use the optimizer to update the model parameters according to the loss function L; e.6 Select the next sample and repeat e.1-e.5 until the loss L converges and the model training is completed.