A conditional guided watermark attack method based on implicit diffusion model
Through the combination of implicit diffusion model and enhanced depth residual network, the problem of watermark difficulty in removing in high-resolution images is solved, and high-quality watermarkless images are generated, which improves the effect and universality of watermark attacks.
Patent Information
- Application Number
- CN202510933651.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing watermark attack methods have limited effects in high-resolution images, difficult to completely remove watermarks and may damage image quality, lack universality, especially poor processing of specific types of watermarks.
The conditional guided watermark attack method based on the implicit diffusion model is adopted. Through the implicit diffusion model and the enhanced depth residual network, combined with the implicit neural representation and iterative process, the image content is gradually refined and the background information is restored, and the image quality is optimized using loss functions and evaluation indicators.
It realizes efficient removal of watermarks in high-resolution images, restores image details, and generates high-quality watermark-free images that are almost the same as the original images, improving the generalization ability and training efficiency of the model, and reducing calculation costs.
Smart Images

Figure CN120430924B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a conditional guided watermark attack method based on an implicit diffusion model. Background Art
[0002] Watermarks are typically embedded covertly into images, audio, and video to identify the original creator, confirm copyright, or prevent illegal copying. However, the development of watermarking technology has also driven advancements in watermark attacks. With the rise of artificial intelligence technologies like deep learning and generative adversarial networks, watermark attack methods are becoming increasingly intelligent and precise. Using deep neural networks and generative models, attackers can automatically identify and remove watermarks embedded in images, making the removal process more subtle and difficult to detect.
[0003] Although existing diffusion models and generative adversarial networks can remove watermarks, these methods may damage the overall structure and details of the image, resulting in a decrease in image quality. In particular, for images with high resolution or strong watermarks, the removal effect may not be ideal.
[0004] Currently, widely used watermark attack algorithms include signal processing attacks, geometric attacks, content editing attacks, and machine learning-based attacks. Each of these attack methods has its own unique characteristics, enabling effective destruction or removal of different types of watermark systems. Signal processing attacks, while simple and easy to implement, interfere with the watermark signal by adding noise, filtering, and compression, but have limited effectiveness against specially designed, robust watermarks. Machine learning-based attacks, while highly adaptable, utilize deep learning models to generate adversarial examples or estimate and remove watermarks. However, they require extensive training data and computational resources, and their generalization to unknown or novel watermarking algorithms remains to be improved.
[0005] While existing watermark attack techniques can destroy or remove digital watermarks to a certain extent, they still have some significant shortcomings and limitations. First, many watermark attack methods are limited in effectiveness against high-strength watermarks. Especially when the watermark is embedded strongly or when using high-resolution images, attack methods often struggle to completely remove the watermark, or even have minimal impact on the watermark. Second, during image processing, removing or tampering with the watermark may result in loss or distortion of image details. Furthermore, some attack methods may only be effective against specific types of watermarks, lacking universal applicability. Summary of the Invention
[0006] The technical problem addressed by this invention is to provide a conditionally guided watermark attack method based on an implicit diffusion model. First, in high-resolution images, the diffusion model can efficiently restore image details through latent space generation. Second, the implicit diffusion model leverages the characteristics of the diffusion process to iteratively refine the image content, thereby removing the watermark and restoring background information. Finally, the diffusion model can be directly applied to watermark removal tasks without requiring additional training for specific watermark types.
[0007] The present invention adopts the following technical solutions to achieve the invention objectives:
[0008] A conditional guided watermark attack method based on an implicit diffusion model is characterized by comprising the following steps:
[0009] S1: Dataset acquisition;
[0010] S2: Construct a watermark attack network, which mainly includes a noise estimation and noise addition module, a conditional guidance module, a denoising module, and a downsampling module;
[0011] S3: Construct loss function;
[0012] S4: Construct evaluation indicators;
[0013] The specific steps of S2 are:
[0014] S21: The watermarked image Input into the implicit diffusion model;
[0015] S22: After inputting the image, the model selects a diffusion step , at each diffusion step ,in , the model will add Gaussian noise to the watermark image according to the predefined variance table , thus obtaining a Gaussian noise image ;
[0016] S23: After obtaining the Gaussian noise image, the implicit diffusion model upgrades the traditional denoising diffusion model by inheriting the non-Markov process. This process can effectively from The denoising process includes the following steps:
[0017] Input forward process and finally output Gaussian noise image , diffusion steps And enhanced deep residual network The UNet network in the reverse generation process consists of an encoder and a decoder with a skip connection in the middle, and its shape is U-shaped, so it is called a U-Net structure. An implicit neural network is introduced in the decoding part of the UNet network, that is, the upsampling part, to generate an image with higher resolution than the original image through the implicit neural network;
[0018] S24: After obtaining a higher resolution image, downsample the image and restore it to its original size to obtain the image after the attack. .
[0019] As a further limitation of the present technical solution, an enhanced deep residual network and an implicit neural network are adopted in S2;
[0020] The enhanced deep residual network is composed of a stack of convolutional layers and activation layers, and a residual network is introduced to improve the watermarked image quality. Perform fine-grained and controllable feature extraction to obtain features , then the characteristics Feed into UNet for preliminary conditional guidance;
[0021] In the upsampling of the UNet network, implicit neural representation is applied. The core idea is to represent the image as a continuous function rather than the traditional discrete pixel grid. This method allows the model to generate images at arbitrary resolutions. Several coordinate-based multi-layer perceptrons are used in the denoising process to parameterize the implicit neural representation, which enables the model to achieve continuous resolution image super-resolution without sacrificing image quality. Represents continuous coordinates, according to the implicit representation formula:
[0022] (1);
[0023] in: Represents a 2-layer multilayer perceptron with a hidden dimension of 256;
[0024] Express The interpolation is performed by calculating the features in the first Layer depth and This is achieved by the nearest Euclidean distance;
[0025] Indicates the current The uninterpolated coordinates of the layer;
[0026] Express The coordinates obtained by interpolation are also interpolated in the Layer depth calculation and The nearest Euclidean distance is achieved;
[0027] Represents the features after upsampling of the encoder layer;
[0028] The current feature is located around the input coordinates, and the target feature is then calculated.
[0029] As a further limitation of this technical solution, the goal of the loss function is to minimize the prediction noise and real noise The mean square error between:
[0030] (2);
[0031] in: Indicates the desired action;
[0032] Indicates from 1 to Randomly select a time step in , is the total number of steps in the diffusion process;
[0033] Represents the standard normal distribution Randomly select a noise ,in is the identity matrix;
[0034] represents the prediction noise, Represents a watermarked image. represents the denoising process time step The image at
[0035] Represents real noise and model prediction noise The mean square error between .
[0036] As a further limitation of this technical solution, in order to better measure the change in image quality, peak signal-to-noise ratio and structural similarity index are used as measurement indicators;
[0037] The peak signal-to-noise ratio is calculated as follows:
[0038] (3);
[0039] in: Represents the original image;
[0040] Indicates the image after the attack;
[0041] Indicates that the original image is at position Pixel value of
[0042] Indicates that the image after the attack is at position Pixel value of
[0043] Indicates the width of the image;
[0044] Indicates the height of the image;
[0045] Represents the original image The maximum possible square value of the pixel value in ;
[0046] The calculation formula of the structural similarity index is as follows:
[0047] (4);
[0048] in: and The original images and the image after the attack ;
[0049] and are the average pixel values of the original image and the image after attack respectively;
[0050] and Represent the variance of the original image and the distorted image respectively;
[0051] represents the covariance between the original image and the distorted image;
[0052] and is a constant added to avoid the denominator being zero;
[0053] In order to evaluate the effect of watermark removal, the bit error rate is also used as an evaluation indicator: the bit error rate is calculated as follows:
[0054] (5);
[0055] in: The number of bits representing the error information in the extracted watermark;
[0056] The total number of bits representing the original watermark information.
[0057] Compared with the prior art, the advantages and positive effects of the present invention are:
[0058] 1. The present invention adopts an implicit diffusion model (IDM), which gradually refines the image content through an iterative method, thereby removing the watermark and restoring the background information. At the same time, the present invention combines implicit neural representation and diffusion model, avoiding the limitations of traditional methods at fixed resolutions, allowing the model to generate high-quality images within a continuous resolution range. In addition, the present invention also introduces conditional guidance, which guides the denoising process through an enhanced deep residual network stacked by convolutional layers and activation layers, so that the generated de-watermarked image is visually almost the same as the original image, making it highly imperceptible and difficult to detect. Finally, unlike some GAN-based methods, the present invention does not rely on additional prior knowledge and can directly learn the representation of the image from the data, thereby improving the training efficiency and the generalization ability of the model.
[0059] 2. The present invention proposes a more effective method for attacking watermarks. Compared with traditional methods, the present invention can better restore details in high-resolution images and avoid the problems of over-smoothing or artifacts after attacking the watermark. At the same time, the present invention also utilizes the characteristics of the diffusion process to gradually refine the image content through an iterative method, thereby removing the watermark and restoring the background information. In the process of watermark removal, the diffusion model introduces an enhanced deep residual network and implicit neural representation, making the attacked image more similar to the original image. In addition, the end-to-end design of the present invention also simplifies the training process, reduces computational costs, and improves the adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 This is a schematic diagram of the watermark attack network of the present invention. DETAILED DESCRIPTION
[0061] A specific embodiment of the present invention is described in detail below with reference to the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific embodiment.
[0062] The present invention comprises the following steps:
[0063] S1: Get the dataset;
[0064] Using the Image dataset, all images are cropped to 256×256 size, and the watermarked images are obtained using the QPHFMs (Quaternion Polar Harmonic Fourier Moments) quaternion polar harmonic Fourier moment watermark embedding algorithm as the dataset of this network.
[0065] S2: Construct a watermark attack network, such as Figure 1 As shown, it mainly includes noise estimation and denoising module, conditional guidance module, denoising module and downsampling module.
[0066] The specific steps of S2 are:
[0067] S21: The watermarked image is input into the implicit diffusion model, where R represents the field of real numbers, which means All elements of are real numbers, Indicates the height of the image, Indicates the width of the image, and 3 indicates the number of channels of the image;
[0068] In this process, the watermarked image It first passes through a series of convolutional layers, which can gradually extract local features in the image. After the convolutional layers, an activation function is usually applied to introduce nonlinear factors to help the network capture more complex image features.
[0069] S22: After inputting the image, the implicit diffusion model randomly selects a diffusion step (in ), at each diffusion step ,in , the model will add Gaussian noise to the watermark image according to the predefined variance table , thus obtaining a Gaussian noise image ;
[0070] S23: After obtaining the Gaussian noise image, the implicit diffusion model upgrades the traditional denoising diffusion model by inheriting the non-Markov process. This process can effectively from The denoising process includes the following steps:
[0071] Input forward process and finally output Gaussian noise image , diffusion steps And enhanced deep residual network The UNet network for the reverse generation process consists of an encoder and a decoder with a jump connection in the middle. Its shape is U-shaped, so it is called a U-Net structure. An implicit neural network is introduced in the decoding part of the UNet network, that is, the upsampling part, to generate an image with higher resolution than the original image through the implicit neural network.
[0072] S24: After obtaining a higher resolution image, downsample the image and restore it to its original size to obtain the image after the attack. .
[0073] The conditionally guided watermark attack method based on the implicit diffusion model can effectively overcome the limitations of existing attack methods. First, by learning the complex distribution of the image, the implicit diffusion model can simultaneously remove the watermark while generating natural and realistic image content based on the surrounding pixel information, thereby restoring a high-quality watermark-free image. Second, because the implicit diffusion model's watermark removal process is based on the image's intrinsic distribution, it is difficult to detect the presence of a watermark in the removed image. Finally, the implicit diffusion model is inherently robust to noise and distortion, making it excellent when dealing with complex watermarks or backgrounds.
[0074] In the above steps, it should be emphasized that the enhanced deep residual network and implicit neural network are used in S2;
[0075] The enhanced deep residual network is composed of a stack of convolutional layers and activation layers, and a residual network is introduced to improve the watermarked image quality. Perform fine-grained and controllable feature extraction to obtain features , then the characteristics Feed into UNet for preliminary conditional guidance;
[0076] In the upsampling of the UNet network, implicit neural representation is applied. The core idea is to represent the image as a continuous function instead of the traditional discrete pixel grid. This method allows the model to generate images at arbitrary resolutions. Several coordinate-based multi-layer perceptrons are used in the denoising process to parameterize the implicit neural representation, which enables the model to achieve continuous resolution image super-resolution without sacrificing image quality. Represents continuous coordinates, according to the implicit representation formula:
[0077] (1);
[0078] in: Represents a 2-layer multilayer perceptron MLP with a hidden dimension of 256;
[0079] Express The interpolation is performed by calculating the features in the first Layer depth and This is achieved by the nearest Euclidean distance;
[0080] Indicates the current The uninterpolated coordinates of the layer;
[0081] Express The coordinates obtained by interpolation are also interpolated in the Layer depth calculation and The nearest Euclidean distance is achieved;
[0082] Indicates the encoder Features after upsampling of the layer;
[0083] The current feature is located around the input coordinates, and the target feature is then calculated.
[0084] S3: Construct loss function;
[0085] In the watermark attack network, the goal of the loss function is to minimize the prediction noise and real noise The mean square error (MSE) between:
[0086] (2);
[0087] in: Indicates the desired action;
[0088] Indicates from 1 to Randomly select a time step in , is the total number of steps in the diffusion process;
[0089] Represents the standard normal distribution Randomly select a noise ,in is the identity matrix;
[0090] represents the prediction noise, Represents a watermarked image. represents the denoising process time step The image at
[0091] Represents real noise and model prediction noise The mean square error between .
[0092] S4: Construct evaluation indicators.
[0093] In order to better measure the changes in image quality, peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) were used as measurement indicators;
[0094] PSNR measures the overall quality of an image. It measures the overall quality of an image by calculating the error between the original image and the reconstructed image to reflect the image's clarity and noise level. The formula for calculating the Peak Signal-to-Noise Ratio is as follows:
[0095] (3);
[0096] in: Represents the original image;
[0097] Indicates the image after the attack;
[0098] Indicates that the original image is at position Pixel value of
[0099] Indicates that the image after the attack is at position Pixel value of
[0100] Indicates the width of the image;
[0101] Indicates the height of the image;
[0102] Represents the original image The maximum possible square value of the pixel value in .
[0103] The higher the PSNR value, the smaller the distortion between the processed image and the original image, that is, the quality of the processed image is closer to the original image.
[0104] SSIM focuses on evaluating the structural similarity of images, taking into account the brightness, contrast, and structural information of the image, and can measure the similarity of images in a way that is more consistent with human visual perception.
[0105] The calculation formula of the structural similarity index is as follows:
[0106] (4);
[0107] in: and The original images and the image after the attack ;
[0108] and are the average pixel values of the original image and the image after attack respectively;
[0109] and Represent the variance of the original image and the distorted image respectively;
[0110] represents the covariance between the original image and the distorted image;
[0111] and is a constant added to avoid the denominator being zero;
[0112] The SSIM value ranges from 0 to 1, where 1 means the two images are identical and 0 means there is no similarity at all.
[0113] In addition, in order to evaluate the effect of watermark removal, the bit error rate (BER) is used as an evaluation indicator: the bit error rate is calculated as follows:
[0114] (5);
[0115] in: The number of bits representing the error information in the extracted watermark;
[0116] Indicates the total number of bits of the original watermark information. When the BER value is closer to 0, the extracted watermark information is more complete, which indicates that the watermark attack effect is poor; conversely, if the BER value is larger, the watermark attack effect is better.
[0117] The above disclosure is only a specific embodiment of the present invention, but the present invention is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present invention.
Claims
1. A conditional guided watermark attack method based on implicit diffusion model, characterized in that: The following steps are involved: S1: Dataset acquisition; S2: Construct a watermark attack network, including a noise estimation and denoising module, a conditional guidance module, a denoising module, and a downsampling module; S3: Construct loss function; S4: Construct evaluation indicators; The specific steps of S2 are: S21: The watermarked image Input into the implicit diffusion model; S22: After inputting the image, the model selects a diffusion step , at each diffusion step ,in , the model will add Gaussian noise to the watermark image according to the predefined variance table , thus obtaining a Gaussian noise image ; S23: After obtaining the Gaussian noise image, the implicit diffusion model upgrades the traditional denoising diffusion model by inheriting the non-Markov process. This process can effectively from The denoising process includes the following steps: Input forward process and finally output Gaussian noise image , diffusion steps And enhanced deep residual network The UNet network in the reverse generation process consists of an encoder and a decoder with a skip connection in the middle, and its shape is U-shaped, so it is called a U-Net structure. An implicit neural network is introduced in the decoding part of the UNet network, that is, the upsampling part, to generate an image with higher resolution than the original image through the implicit neural network; S24: After obtaining a higher resolution image, downsample the image and restore it to its original size to obtain the image after the attack. ; In said S2, an enhanced deep residual network and an implicit neural network are adopted; The enhanced deep residual network is composed of a stack of convolutional layers and activation layers, and a residual network is introduced to improve the watermarked image quality. Perform fine-grained and controllable feature extraction to obtain features , then the characteristics Feed into UNet for preliminary conditional guidance; In the upsampling of the UNet network, implicit neural representation is applied. The core idea is to represent the image as a continuous function rather than the traditional discrete pixel grid. This method allows the model to generate images at arbitrary resolutions. Several coordinate-based multi-layer perceptrons are used in the denoising process to parameterize the implicit neural representation, which enables the model to achieve continuous resolution image super-resolution without sacrificing image quality. Represents continuous coordinates, according to the implicit representation formula: (1); in: Represents a 2-layer multilayer perceptron with a hidden dimension of 256; Express The interpolation is performed by calculating the features in the first Layer depth and This is achieved by the nearest Euclidean distance; Indicates the current The uninterpolated coordinates of the layer; Express The coordinates obtained by interpolation are also interpolated in the Layer depth calculation and The nearest Euclidean distance is achieved; Indicates the encoder Features after upsampling of the layer; The current feature is located around the input coordinates, and the target feature is then calculated.
2. The conditional guided watermark attack method based on the implicit diffusion model according to claim 1 is characterized in that: The goal of the loss function is to minimize the prediction noise and real noise The mean square error between: (2); in: Indicates the desired action; Indicates from 1 to Randomly select a time step in , is the total number of steps in the diffusion process; Represents the standard normal distribution Randomly select a noise ,in is the identity matrix; represents the prediction noise, Represents a watermarked image. represents the denoising process time step The image at Represents real noise and model prediction noise The mean square error between .
3. The conditional guided watermark attack method based on the implicit diffusion model according to claim 2 is characterized in that: In order to better measure the changes in image quality, peak signal-to-noise ratio and structural similarity index were used as measurement indicators; The peak signal-to-noise ratio is calculated as follows: (3); in: Represents the original image; Indicates the image after the attack; Indicates that the original image is at position Pixel value of Indicates that the image after the attack is at position Pixel value of Indicates the width of the image; Indicates the height of the image; Represents the original image The maximum possible square value of the pixel value in ; The calculation formula of the structural similarity index is as follows: (4); in: and The original images and the image after the attack ; and are the average pixel values of the original image and the image after attack respectively; and Represent the variance of the original image and the distorted image respectively; represents the covariance between the original image and the distorted image; and is a constant added to avoid the denominator being zero; In order to evaluate the effect of watermark removal, the bit error rate is also used as an evaluation indicator: the bit error rate is calculated as follows: (5); in: The number of bits representing the error information in the extracted watermark; The total number of bits representing the original watermark information.
Citation Information
Patent Citations
Conditional guidance watermark attack method based on potential diffusion model
CN118297784A
Robust image watermarking method, system and terminal based on conditional diffusion model
CN119205478A
Cited By
Digital watermark attack method based on cross-channel statistical decoupling
CN122312359A