A Watermark Attack Method Based on Fourier Transform and Multi-Scale Convolution

By combining Fourier transform and multi-scale convolution technology in watermark attack, the amplitude and phase of the image are enhanced respectively, and multi-scale convolution is used to process the image, the problem of existing watermark attack algorithms destroying the image structure when removing watermarks is solved, achieving a more efficient and reliable watermark attack effect.

CN119991405BActive Publication Date: 2025-06-17QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Patent Information

Application Number
CN202510461097.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-06-17
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing watermark attack algorithms easily destroy image structure and content while removing watermarks, and have high computational complexity and limited attack effects on high-resolution images.

Method used

The watermark attack method based on Fourier transform and multi-scale convolution is adopted. The image is converted to the frequency domain through fast Fourier transform, and the amplitude and phase are enhanced respectively. The image is further processed using multi-scale convolution to ensure that the watermark information is completely destroyed and image quality is improved.

Benefits of technology

A more accurate and effective watermark attack is achieved, which can effectively remove deep embedded watermarks, especially watermarks embedded in the frequency domain, and retain the structure and detailed information of the image during the removal of the watermark, and maintain the visual quality of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991405B_ABST
    Figure CN119991405B_ABST
Patent Text Reader

Abstract

The present invention discloses a watermark attack method based on Fourier transform and multi-scale convolution, which relates to the technical field of image processing. It is characterized by the following steps: S1: Acquisition of the data set; S2: Construction of the watermark attack network, mainly including the Fourier enhancement stage and the multi-scale convolution stage; S21: In the Fourier enhancement stage, the amplitude and phase are enhanced independently to maximize the destruction of the watermark information; S22: The attacked image is further processed through multi-scale convolution to ensure that the watermark information is completely destroyed while improving the image quality; S3: Construction of the loss function; S4: Construction of the evaluation index. The technical problem to be solved by the present invention is to provide a watermark attack method based on Fourier transform and multi-scale convolution, a more effective and reliable watermark attack solution, so as to improve the security and integrity of digital media.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more specifically, to a watermark attack method based on Fourier transform and multi-scale convolution. Background Art

[0002] A watermark is hidden information embedded in digital content (such as images, audio, video, or documents). Watermarks are commonly used to authenticate and protect the copyright or integrity of data. Combining watermark technology with the current trends of digitalization and informatization can not only protect intellectual property rights and copyrights, verify the authenticity and integrity of information, but also prevent piracy and infringement, conduct digital identity authentication and authorization, and protect personal privacy and data security. In the current era of rapid development of digital media, the problems of piracy, infringement, and improper use of digital media have attracted wide attention. A watermark attack refers to the act of maliciously operating or damaging a digital watermark, and its main purpose is to destroy the visibility or detectability of the watermark, thereby weakening the effectiveness of the watermark technology. With the continuous improvement of watermark technology, watermark attack technology is also evolving, posing greater challenges to the protection of intellectual property rights in digital media. Therefore, we propose a more advanced watermark attack technology aimed at promoting the improvement of watermark embedding technology.

[0003] Although existing watermark attack algorithms that combine Fourier transform and neural networks can attack the watermark information to a certain extent after detecting it, they often amplify the amplitude component and simply copy the phase component. Some watermarks may rely on phase information, and simply amplifying the amplitude may not completely destroy the watermark. In addition, separately enhancing the amplitude component will also have a greater impact on the structure and authenticity of the image.

[0004] Existing watermark attack algorithms have several drawbacks that need to be addressed. First, traditional watermark attack algorithms (such as geometric processing attacks, filtering attacks, etc.) are simple but ineffective against some watermark algorithms with strong robustness. In addition, for high-resolution images, traditional watermark attack algorithms require high computational complexity and have limited attack effects. In recent years, watermark attack methods that combine Fourier transform and deep learning have become a research hotspot. These methods mainly use Fourier transform to extract frequency domain information and combine neural networks to remove or attack watermarks. However, they often ignore the importance of phase information to the image structure and content. Although they can damage the watermark to a certain extent, they are prone to cause image distortion or distortion. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a watermark attack method based on Fourier transform and multi-scale convolution. First, the fast Fourier transform is used to transform the image into the frequency domain, and the frequency feature information provided by the Fourier transform is utilized to make the watermark attack more accurate and effective. Second, the watermark attack network focuses on maintaining the visual quality of the image, enhancing the amplitude and phase respectively, and retaining the structural and detailed information in the image as much as possible to avoid irreversible effects on the image. In addition, multi-scale convolution is used to enable the model to adapt to different types of watermark embedding methods, improving the generalization of the watermark attack. Even if the watermark position, size, and type are different, it can still be effectively removed. A more effective and reliable watermark attack solution is provided to improve the security and integrity of digital media.

[0006] The present invention adopts the following technical solutions to achieve the invention purpose:

[0007] A watermark attack method based on Fourier transform and multi-scale convolution, characterized by comprising the following steps:

[0008] S1: Acquisition of the data set;

[0009] S2: Construct a watermark attack network, mainly including a Fourier enhancement stage and a multi-scale convolution stage;

[0010] S21: Independently enhance the amplitude and phase in the Fourier enhancement stage to maximize the destruction of watermark information;

[0011] S22: Further process the attacked image through multi-scale convolution to ensure that the watermark information is completely destroyed while improving the image quality;

[0012] S3: Construct a loss function;

[0013] S4: Construct evaluation indexes.

[0014] As a further limitation of this technical solution, the specific steps of S21 are as follows:

[0015] S211: First, input the watermarked image ; obtain the amplitude component and the phase component through the fast Fourier transform formulas (1), (2), and (3);

[0016] (1)

[0017] Where: is the complex spectrum;

[0018] is the height of the image;

[0019] is the width of the image;

[0020] is the image space coordinate;

[0021] is the frequency domain coordinate;

[0022] is the imaginary unit;

[0023] (2);

[0024] (3);

[0025] Where: and respectively represent the real and imaginary parts of;

[0026] S212: Use a 1×1 convolutional layer, a LeakyReLU activation function, and a 1x1 convolutional layer to extract the amplitude and phase components respectively, and obtain and ;

[0027] S213: For the amplitude enhancement branch, use a dilated convolutional block and an SE attention structure to enhance the amplitude, and obtain the enhanced amplitude :

[0028] (4);

[0029] Where: represents the output after dilated convolution;

[0030] represents the output of the attention structure;

[0031] S214: For the phase enhancement branch, use a residual block for processing, and then jump-connect the input to the output to obtain the enhanced phase :

[0032] (5);

[0033] Where: represents the non-linear mapping of the phase through the residual block;

[0034] S215: Perform an inverse Fourier transform on the enhanced amplitude and the phase to reconstruct the image and obtain the output of the first stage :

[0035] (6);

[0036] Where: Represents the inverse Fourier transform.

[0037] As a further limitation of this technical solution, the specific steps of S22 are as follows:

[0038] S221: Downsample the output of the first stage using an encoder ;

[0039] S222: Process separately through six multi-scale convolutional layers;

[0040] S223: Feed the generated features into the decoder to generate the final attacked image .

[0041] As a further limitation of this technical solution, in order to balance the visual quality of the image while removing the watermark in the image, the loss function consists of three parts:

[0042] ;

[0043] ;

[0044] ;

[0045] (7);

[0046] Among them: is the mean squared error loss between the output image of the first stage and the real image ;

[0047] is the mean squared error loss between the output image of the second stage and the real image ;

[0048] is the perceptual loss based on VGG features to enhance the visual effect;

[0049] represents the norm used to calculate the Euclidean distance of the vector;

[0050] represents the feature extraction result of the th layer of the VGG-19 pre-trained model;

[0051] , and are loss weight parameters.

[0052] As a further limitation of this technical solution, the peak signal-to-noise ratio and the structural similarity index are used as measurement indicators;

[0053] The calculation formula of the peak signal-to-noise ratio is as follows:

[0054] (8);

[0055] Where: represents the watermarked image;

[0056] represents the image after being attacked;

[0057] represents the height of the image;

[0058] represents the width of the image;

[0059] represents the square of the maximum pixel value of the image;

[0060] The pixel value of the original watermarked image at ;

[0061] The pixel value of the original image after being attacked at ;

[0062] The calculation formula of the structural similarity index is as follows:

[0063] (9);

[0064] Where: and are respectively the original image and the image after being attacked ;

[0065] , respectively represent the mean values of the original image and the image after being attacked ;

[0066] , respectively represent the standard deviations of the original image and the image after being attacked ;

[0067] represents the covariance of the original image and the image after being attacked ;

[0068] and is a constant used to prevent the denominator from being zero;

[0069] The bit error rate is also used as an evaluation indicator, and its calculation formula is as follows:

[0070] (10);

[0071] in: The number of bits representing the error information in the extracted watermark;

[0072] Represents the total number of bits of the original watermark information.

[0073] As a further limitation of the present technical solution, the watermark attack network includes a Fourier enhancement module and a multi-scale convolution module. The Fourier enhancement module is composed of a phase enhancement branch and an amplitude enhancement branch. The multi-scale convolution module is composed of an encoder. E , multi-scale convolutional layer, decoder D composition.

[0074] Compared with the prior art, the advantages and positive effects of the present invention are:

[0075] 1. The present invention uses Fourier transform to convert the image from the spatial domain to the frequency domain, separates the amplitude and phase of the watermarked image, and enhances the amplitude and phase respectively, rather than simply modifying the amplitude, thereby improving the perception and destruction of watermark information, especially for watermark information embedded in the frequency domain. At the same time, multi-scale convolution blocks are used for feature extraction, combined with convolution kernels of different sizes, so that the model can adapt to different types of watermark embedding methods, improve the generalization of watermark attacks, and ensure that the image quality after the attack will not be significantly reduced. In the design of the loss function, two-stage loss optimization is adopted to ensure that the watermark is destroyed while maintaining the image quality.

[0076] 2. The present invention proposes a new and more effective watermark attack. Compared with traditional algorithms, the present invention can more effectively remove deeply embedded watermarks, especially watermark attacks embedded in the frequency domain. In the process of watermark removal, the amplitude and phase components are enhanced to ensure the retention of complex image information, overcoming the limitation of previous methods that ignore the interaction between amplitude and phase, making the attacked image more similar to the original image, and being able to better maintain the visual quality of the image, improving the aesthetics and usability of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is the overall structure flow chart of the present invention. DETAILED DESCRIPTION

[0078] The following will describe in detail a specific embodiment of the present invention in conjunction with the accompanying drawings. It should be understood that the protection scope of the present invention is not limited by the specific embodiment.

[0079] The current mainstream watermark attack algorithms are watermark attack networks constructed based on traditional neural networks such as CNN and RNN and traditional generative models such as GAN. Although these watermark attack algorithms can completely remove the watermark information after detecting it, there are several major defects. On the one hand, while removing the watermark, the original image structure will also be damaged, resulting in problems such as color distortion, blurring, and artifacts, affecting the visual usability of the image. On the other hand, for watermarks embedded in the frequency domain, these algorithms cannot effectively remove the watermark. Especially for watermarks hidden in high-frequency information, spatial domain-based methods (such as U-Net, GAN) may not be able to effectively separate the watermark and the background, easily resulting in residual watermarks or image blurring.

[0080] In contrast, the present method can solve these problems. The Fast Fourier Transform is an efficient algorithm for calculating the Fourier transform, which can accelerate the calculation of the DFT (Discrete Fourier Transform), making it more efficient in practical applications. The result calculated by it is in complex form, which contains amplitude and phase information. By performing the Fast Fourier Transform, the image is transformed from the spatial domain to the frequency domain, which can better capture the watermark information in the frequency domain, thereby better attacking the watermark. At the same time, by enhancing the amplitude and phase, the information of the image can be retained to a greater extent, improving the visual quality of the image.

[0081] The present invention includes the following steps:

[0082] S1: Acquisition of the dataset.

[0083] Using the Image dataset, all images are cropped to a size of 256×256, and the watermarked images obtained by using the QPHFMs (Quaternion-based Polar Harmonic Fourier Orthogonal Moments) watermark embedding algorithm are used as the dataset of this network.

[0084] S2: Construct a watermark attack network, mainly including a Fourier enhancement stage and a multi-scale convolution stage.

[0085] S21: In the Fourier enhancement stage, the amplitude and phase are enhanced independently to maximize the destruction of the watermark information.

[0086] The specific steps of S21 are as follows:

[0087] S211: First, input the watermarked image , , represents the real number field, meaning All elements are real numbers; 3 represents the number of channels, and the image is an RGB color image; the amplitude components are obtained through the fast Fourier transform equations 1, 2, and 3. and the phase components ;

[0088] (1)

[0089] Where: is the complex spectrum;

[0090] is the height of the image;

[0091] is the width of the image;

[0092] is the image spatial coordinate;

[0093] is the frequency domain coordinate;

[0094] is the imaginary unit;

[0095] (2);

[0096] (3);

[0097] Where: and respectively represent the real and imaginary parts of;

[0098] S212: Once transformed into the Fourier space, the assumption of spatial invariance no longer holds. Therefore, a 1×1 convolutional layer, a LeakyReLU activation function, and a 1x1 convolutional layer are used to extract the amplitude and phase components respectively, obtaining and ;

[0099] S213: For the amplitude enhancement branch, a dilated convolutional block and an SE attention structure are used to enhance the amplitude, obtaining the enhanced amplitude :

[0100] (4);

[0101] Where: represents the output after dilated convolution;

[0102] represents the output of the attention structure;

[0103] The dilated convolutional block consists of a 3×3 convolutional layer with a dilation rate of 2, a LeakyReLU activation function, and a batch normalization layer;

[0104] The SE attention structure consists of global average pooling, channel attention calculation, and channel feature recalibration;

[0105] S214: For the phase enhancement branch, use a residual block for processing. The residual block is composed of two consecutive 3x3 convolutions connected by a relu activation function, and then the input is skip-connected to the output to obtain the enhanced phase :

[0106] (5);

[0107] Where: represents the non-linear mapping of the phase through the residual block;

[0108] S215: For the enhanced amplitude and phase perform inverse Fourier transform to reconstruct the image and obtain the output of the first stage :

[0109] (6);

[0110] Where: represents the inverse Fourier transform.

[0111] S22: Further process the attacked image through multi-scale convolution to ensure that the watermark information is completely destroyed while improving the image quality.

[0112] The specific steps of the above S22 are as follows:

[0113] S221: Downsample the output of the first stage using an encoder ;

[0114] S222: Process it respectively through six multi-scale convolutional layers;

[0115] First, use convolutional layer for dimensionality reduction;

[0116] Then, use convolutional kernels of multiple different sizes to extract multi-scale features. The sizes of the convolutional kernels are respectively , , and , and perform sigmoid operations on the output features respectively, and multiply them with the output features of different convolutional kernels;

[0117] Then, connect them and use a 1x1 convolutional layer to restore the dimension;

[0118] Finally, combine with the input features through a residual connection.

[0119] This can capture a wide range of spatial features across multiple scales and enhance the overall representation of the spatial structure.

[0120] S223: The generated features are fed into the decoder to generate the final attacked image .

[0121] S3: Construct the loss function.

[0122] To balance the visual quality of the image while removing the watermark, the loss function consists of three parts:

[0123] ;

[0124] ;

[0125] ;

[0126] (7);

[0127] Among them: is the mean square error loss between the output image of the first stage and the real image ;

[0128] is the mean square error loss between the output image of the second stage and the real image ;

[0129] is the perceptual loss based on VGG features to improve the visual effect. VGG is a deep convolutional neural network proposed by the VGG team at the University of Oxford;

[0130] represents the norm used to calculate the Euclidean distance of vectors;

[0131] represents the feature extraction result of the th layer of the VGG-19 pre-trained model;

[0132] , and are loss weight parameters used to balance the influence of different losses.

[0133] S4: Construct evaluation metrics.

[0134] To better measure the change in image quality, the peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) are used as measurement metrics;

[0135] The PSNR measures the overall quality of an image. The calculation formula for the peak signal-to-noise ratio is as follows:

[0136] (8);

[0137] Where: represents the watermarked image;

[0138] represents the image after being attacked;

[0139] represents the height of the image;

[0140] represents the width of the image;

[0141] represents the square of the maximum pixel value of the image;

[0142] The pixel value of the original watermarked image at ;

[0143] The pixel value of the original attacked image at ;

[0144] The higher the PSNR value, the smaller the distortion between the processed image and the original image, that is, the closer the quality of the processed image is to the original image.

[0145] SSIM is an index used to measure the similarity between two images, which focuses more on evaluating the structure and visual similarity of images. It can reflect the human eye's perception of image quality better than PSNR.

[0146] The calculation formula for the structural similarity index is as follows:

[0147] (9);

[0148] Where: and are the original image and the attacked image respectively;

[0149] and represent the means of the original image and the attacked image respectively;

[0150] and represent the standard deviations of the original image and the attacked image respectively;

[0151] Represents the original image And the image after the attack The covariance of

[0152] and is a constant used to prevent the denominator from being zero;

[0153] The value of SSIM ranges from 0 to 1, where 1 means that the two images are exactly the same and 0 means there is no similarity at all. SSIM takes into account three factors that are closely related to human visual perception: brightness, contrast, and structure.

[0154] In addition, in order to evaluate the effect of watermark removal, the bit error rate (BER) is used as an evaluation indicator, and its calculation formula is as follows:

[0155] (10);

[0156] in: The number of bits representing the error information in the extracted watermark;

[0157] Indicates the total number of bits of the original watermark information. When the BER value is closer to 0, it means that the extracted watermark information is more complete, which means that the watermark attack effect is poor; conversely, if the BER value is larger, it means that the watermark attack effect is better.

[0158] The watermark attack network includes a Fourier enhancement module and a multi-scale convolution module. The Fourier enhancement module consists of a phase enhancement branch and an amplitude enhancement branch. The multi-scale convolution module consists of an encoder. E , multi-scale convolutional layer, decoder D composition.

[0159] The encoder in the autoencoder in the pre-trained vqgan network used by the present invention E and decoder D The autoencoder is composed of the encoder and decoder in the pre-trained VQGAN. This part is used in the second stage multi-scale convolution. First, the image generated in the first stage is encoded, the image features after the first stage watermark attack are extracted, and then multi-scale convolution is performed, and finally decoding is performed to restore the image after the attack.

[0160] The above disclosure is only a specific embodiment of the present invention, but the present invention is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.

Claims

1. A watermark attack method based on Fourier transform and multi-scale convolution, characterized in that: The following steps are involved: S1: Dataset acquisition; S2: Construct a watermark attack network, which mainly includes the Fourier enhancement stage and the multi-scale convolution stage; S21: The Fourier enhancement stage independently enhances the amplitude and phase to maximize the destruction of watermark information; S22: Further process the attacked image through multi-scale convolution to ensure that the watermark information is completely destroyed while improving the image quality; S3: construct loss function; S4: construct evaluation indicators; The specific steps of S21 are: S211: First input the watermark image ; The amplitude component is obtained by fast Fourier transform equation 1, equation 2 and equation 3 and phase component ; (1); in: is the complex spectrum; is the height of the image; is the width of the image; are image space coordinates; is the frequency domain coordinate; is an imaginary unit; (2); (3); in: and Respectively The real and imaginary parts of S212: Use 1×1 convolution layer, LeakyReLU activation function and 1x1 convolution layer to extract amplitude and phase components respectively, and get and ; S213: For the amplitude enhancement branch, the dilated convolution block and SE attention structure are used to enhance the amplitude to obtain the enhanced amplitude : (4); in: Represents the output after dilated convolution; Represents the output of the attention structure; S214: For the phase enhancement branch, use the residual block processing, and then jump connect the input to the output to obtain the enhanced phase : (5); in: Represents the nonlinear mapping of the phase through the residual block; S215: The enhanced amplitude and Phase Perform an inverse Fourier transform and reconstruct the image to obtain the output of the first stage : (6); in: stands for Inverse Fourier Transform; The specific steps of S22 are: S221: Output to the first stage Use encoder Downsampling; S222: processed by six multi-scale convolutional layers respectively; S223: The generated features are fed to the decoder , generating the final attacked image .

2. The watermark attack method based on Fourier transform and multi-scale convolution according to claim 1 is characterized in that: In order to remove the watermark from the image while taking into account the visual quality of the image, the loss function consists of three parts: ; ; ; (7); in: is the first stage output image With real image The mean square error loss of is the second stage output image With real image The mean square error loss of It is a perceptual loss based on VGG features to improve visual effects; Representation paradigm for calculating the Euclidean distance of vectors; Represents the first Layer feature extraction results; , and is the loss weight parameter.

3. The watermark attack method based on Fourier transform and multi-scale convolution according to claim 1 is characterized in that: Peak signal-to-noise ratio and structural similarity index were used as measurement indicators; The peak signal-to-noise ratio is calculated as follows: (8); in: Represents a watermarked image; Represents the image after the attack; Indicates the height of the image; Indicates the width of the image; Represents the square of the maximum pixel value of the image; The original watermarked image is The pixel value at ; The original attacked image is The pixel value at ; The calculation formula of the structural similarity index is as follows: (9); in: and The original images And the image after the attack ; , Represent the original images And the image after the attack The mean of , Represent the original images And the image after the attack The standard deviation of Represents the original image And the image after the attack The covariance of and is a constant used to prevent the denominator from being zero; The bit error rate is also used as an evaluation indicator, and its calculation formula is as follows: (10); in: The number of bits representing the error information in the extracted watermark; Represents the total number of bits of the original watermark information.

4. The watermark attack method based on Fourier transform and multi-scale convolution according to claim 1 is characterized in that: The watermark attack network includes a Fourier enhancement module and a multi-scale convolution module. The Fourier enhancement module consists of a phase enhancement branch and an amplitude enhancement branch, and the multi-scale convolution module consists of an encoder. E , multi-scale convolutional layer, decoder D composition.

Citation Information

Patent Citations

  • Hidden digital watermark attack method and system based on SAD network

    CN115358909A

  • Hidden watermark attack algorithm based on wavelet transform and attention mechanism

    CN116308986A

Cited By

  • Digital watermark attack method based on cross-channel statistical decoupling

    CN122312359A