Watermark attack method based on Fourier transform and multi-scale convolution

By using Fourier transform and multi-scale convolution techniques in the watermark attack algorithm, the amplitude and phase of the image are enhanced respectively, and multi-scale convolution processing is used to solve the problem of existing watermark attack algorithm destroying the image structure, achieving more effective watermark attacks and higher image quality.

CN119991405AActive Publication Date: 2025-05-13QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Patent Information

Application Number
CN202510461097.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing watermark attack algorithms easily destroy the image structure while removing the watermark, resulting in a decrease in visual quality and limited effect on watermark embedded in the frequency domain.

Method used

The watermark attack method based on Fourier transform and multi-scale convolution is adopted. The image is converted to the frequency domain through fast Fourier transform, and the amplitude and phase are enhanced respectively. The multi-scale convolution is used to further process to ensure that the watermark information is completely destroyed while maintaining the image quality.

Benefits of technology

It realizes more accurate and effective watermark attacks, retaining the structure and detailed information of the image, and improving the generalization of the watermark attack and the visual quality of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991405A_ABST
    Figure CN119991405A_ABST
Patent Text Reader

Abstract

The invention discloses a watermark attack method based on Fourier transform and multi-scale convolution, which relates to the technical field of image processing, and is characterized by comprising the following steps: S1, obtaining a data set; s2, constructing a watermark attack network, wherein the watermark attack network mainly comprises a Fourier enhancement stage and a multi-scale convolution stage; s21, performing independent enhancement on the amplitude and the phase in a Fourier enhancement stage so as to damage the watermark information to the maximum extent; s22, further processing the attacked image through multi-scale convolution to ensure that the watermark information is completely destroyed, and improving the image quality at the same time; s3, constructing a loss function; and S4, constructing an evaluation index. The technical problem to be solved by the invention is to provide a watermark attack method based on Fourier transform and multi-scale convolution, and a watermark attack solution is more effective and reliable, so that the security and integrity of digital media are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a watermark attack method based on Fourier transform and multi-scale convolution. Background Art

[0002] A watermark is a hidden information embedded in digital content (such as images, audio, video or documents). Watermarks are usually used to authenticate and protect the copyright or integrity of data. Watermark technology, combined with the current trend of digitization and informatization, can not only protect intellectual property rights and copyrights, verify the authenticity and integrity of information, but also prevent piracy and infringement, perform digital identity authentication and authorization, and protect personal privacy and data security. In the current era of rapid development of digital media, the problems of piracy, infringement and improper use of digital media have attracted widespread attention. Watermark attack refers to the malicious operation or destruction of digital watermarks. Its main purpose is to destroy the visibility or detectability of the watermark, thereby weakening the effectiveness of watermark technology. With the continuous improvement of watermark technology, watermark attack technology is also evolving, bringing greater challenges to the protection of intellectual property rights of digital media. Therefore, we propose a more advanced watermark attack technology to promote the improvement of watermark embedding technology.

[0003] Although the existing watermark attack algorithms based on Fourier transform and neural network can attack the watermark information to a certain extent after detecting it, they often amplify the amplitude component and simply copy the phase component. Some watermarks may rely on phase information, and amplifying the amplitude alone may not completely destroy the watermark. In addition, enhancing the amplitude component alone will also have a significant impact on the structure and authenticity of the image.

[0004] There are several shortcomings in existing watermark attack algorithms that need to be addressed. First, although traditional watermark attack algorithms (such as geometric processing attacks, filtering attacks, etc.) are simple, they are ineffective against some watermark algorithms with strong robustness. In addition, for high-resolution images, traditional watermark attack algorithms require high computational complexity and limited attack effects. In recent years, watermark attack methods combining Fourier transform with deep learning have become a research hotspot. These methods mainly use Fourier transform to extract frequency domain information and combine neural networks to remove or attack watermarks. However, the importance of phase information to image structure and content is often overlooked. Although the watermark can be destroyed to a certain extent, it is easy to cause image distortion or distortion. Summary of the invention

[0005] The technical problem to be solved by the present invention is to provide a watermark attack method based on Fourier transform and multi-scale convolution. First, the image is transformed into the frequency domain using fast Fourier transform, and the frequency feature information provided by Fourier transform is used to make the watermark attack more accurate and effective. Secondly, the watermark attack network focuses on maintaining the visual quality of the image, enhancing the amplitude and phase respectively, retaining the structure and detail information in the image as much as possible, and avoiding irreversible effects on the image. In addition, by using multi-scale convolution, the model can adapt to different types of watermark embedding methods, improving the generalization of watermark attacks, and even if the watermark position, size, and type are different, it can still be effectively removed. A more effective and reliable watermark attack solution, thereby improving the security and integrity of digital media.

[0006] The present invention adopts the following technical solutions to achieve the invention objectives: A watermark attack method based on Fourier transform and multi-scale convolution, characterized by comprising the following steps: S1: Dataset acquisition; S2: Construct a watermark attack network, which mainly includes the Fourier enhancement stage and the multi-scale convolution stage; S21: The Fourier enhancement stage independently enhances the amplitude and phase to maximize the destruction of watermark information; S22: Further process the attacked image through multi-scale convolution to ensure that the watermark information is completely destroyed while improving the image quality; S3: construct loss function; S4: Construct evaluation indicators.

[0007] As a further limitation of the present technical solution, the specific steps of S21 are: S211: First input the watermark image ; The amplitude component is obtained by fast Fourier transform equation 1, equation 2 and equation 3 and phase component ; (1) in: is the complex spectrum; is the height of the image; is the width of the image; are image space coordinates; is the frequency domain coordinate; is an imaginary unit; (2); (3); in: and Respectively The real and imaginary parts of S212: Use 1×1 convolution layer, LeakyReLU activation function and 1x1 convolution layer to extract amplitude and phase components respectively, and get and ; S213: For the amplitude enhancement branch, the dilated convolution block and SE attention structure are used to enhance the amplitude to obtain the enhanced amplitude : (4); in: Represents the output after dilated convolution; Represents the output of the attention structure; S214: For the phase enhancement branch, use the residual block processing, and then jump connect the input to the output to obtain the enhanced phase : (5); in: Represents the nonlinear mapping of the phase through the residual block; S215: The enhanced amplitude and Phase Perform an inverse Fourier transform and reconstruct the image to obtain the output of the first stage : (6); in: Stands for Inverse Fourier Transform.

[0008] As a further limitation of the technical solution, the specific steps of S22 are: S221: Output to the first stage Use encoder Downsampling; S222: processed by six multi-scale convolutional layers respectively; S223: The generated features are fed to the decoder , generating the final attacked image .

[0009] As a further limitation of this technical solution, in order to remove the watermark from the image while taking into account the visual quality of the image, the loss function consists of three parts: ; ; ; (7); in: is the output image of the first stage With real image The mean square error loss of is the second stage output image With real image The mean square error loss of It is a perceptual loss based on VGG features to improve visual effects; Representation paradigm for calculating the Euclidean distance of vectors; Represents the first Layer feature extraction results; , and is the loss weight parameter.

[0010] As a further limitation of the present technical solution, peak signal-to-noise ratio and structural similarity index are used as measurement indicators; The peak signal-to-noise ratio is calculated as follows: (8); in: Represents a watermarked image; Represents the image after the attack; Indicates the height of the image; Indicates the width of the image; Represents the square of the maximum pixel value of the image; The original watermarked image is The pixel value at ; The original attacked image is The pixel value at ; The calculation formula of the structural similarity index is as follows: (9); in: and The original images And the image after the attack ; , Represent the original images And the image after the attack The mean of , Represent the original images And the image after the attack The standard deviation of Represents the original image And the image after the attack The covariance of and is a constant used to prevent the denominator from being zero; The bit error rate is also used as an evaluation indicator, and its calculation formula is as follows: (10); in: The number of bits representing the error information in the extracted watermark; Represents the total number of bits of the original watermark information.

[0011] As a further limitation of the present technical solution, the watermark attack network includes a Fourier enhancement module and a multi-scale convolution module. The Fourier enhancement module is composed of a phase enhancement branch and an amplitude enhancement branch. The multi-scale convolution module is composed of an encoder. E , multi-scale convolutional layer, decoder D composition.

[0012] Compared with the prior art, the advantages and positive effects of the present invention are: 1. The present invention uses Fourier transform to convert the image from the spatial domain to the frequency domain, separates the amplitude and phase of the watermarked image, and enhances the amplitude and phase respectively, rather than simply modifying the amplitude, thereby improving the perception and destruction of watermark information, especially for watermark information embedded in the frequency domain. At the same time, multi-scale convolution blocks are used for feature extraction, combined with convolution kernels of different sizes, so that the model can adapt to different types of watermark embedding methods, improve the generalization of watermark attacks, and ensure that the image quality after the attack will not be significantly reduced. In the design of the loss function, two-stage loss optimization is adopted to ensure that the watermark is destroyed while maintaining the image quality.

[0013] 2. The present invention proposes a new and more effective watermark attack. Compared with traditional algorithms, the present invention can more effectively remove deeply embedded watermarks, especially watermark attacks embedded in the frequency domain. In the process of watermark removal, the amplitude and phase components are enhanced to ensure the retention of complex image information, overcoming the limitation of previous methods that ignore the interaction between amplitude and phase, making the attacked image more similar to the original image, and being able to better maintain the visual quality of the image, improving the aesthetics and usability of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is the overall structure flow chart of the present invention. DETAILED DESCRIPTION

[0015] A specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific implementation.

[0016] The current mainstream watermark attack algorithms are watermark attack networks built on the basis of traditional neural networks such as CNN and RNN and traditional generative models such as GAN. Although these watermark attack algorithms can completely remove the watermark information after detecting it, they have several major defects. On the one hand, while removing the watermark, the original image structure will also be destroyed, resulting in color distortion, blur, artifacts and other problems, affecting the visual usability of the image. On the other hand, these algorithms cannot effectively remove watermarks embedded in the frequency domain. Especially for watermarks hidden in high-frequency information, methods based on the spatial domain (such as U-Net, GAN) may not be able to effectively separate the watermark and background, which can easily lead to residual watermarks or image blur.

[0017] In contrast, this method can solve these problems. Fast Fourier transform is an efficient algorithm for calculating Fourier transform, which can accelerate the calculation of DFT (discrete Fourier transform) and make it more efficient in practical applications. The calculated result is in complex form, which contains amplitude and phase information. By fast Fourier transform, the image is converted from the spatial domain to the frequency domain, which can better capture the watermark information in the frequency domain, thereby better attacking the watermark. At the same time, by enhancing the amplitude and phase, the image information can be retained to a greater extent, improving the visual quality of the image.

[0018] The present invention comprises the following steps: S1: Acquisition of dataset.

[0019] Using the Image dataset, all images are cropped to a size of 256×256, and the watermarked images are obtained using the QPHFMs (quaternion-based extreme harmonic Fourier moments) watermark embedding algorithm as the dataset of this network.

[0020] S2: Construct a watermark attack network, which mainly includes the Fourier enhancement stage and the multi-scale convolution stage.

[0021] S21: The Fourier enhancement stage independently enhances the amplitude and phase to maximize the destruction of the watermark information.

[0022] The specific steps of S21 are: S211: First input the watermark image , , represents the field of real numbers, which means All elements of are real numbers; 3 represents the number of channels, and the image is an RGB color image; the amplitude component is obtained by fast Fourier transform equations 1, 2, and 3 and phase component ; (1) in: is the complex spectrum; is the height of the image; is the width of the image; are image space coordinates; is the frequency domain coordinate; is an imaginary unit; (2); (3); in: and Respectively The real and imaginary parts of S212: Once transformed into Fourier space, the assumption of spatial invariance no longer holds. Therefore, a 1×1 convolutional layer, a LeakyReLU activation function, and a 1x1 convolutional layer are used to extract the amplitude and phase components, respectively, and we get and ; S213: For the amplitude enhancement branch, the dilated convolution block and SE attention structure are used to enhance the amplitude to obtain the enhanced amplitude : (4); in: Represents the output after dilated convolution; Represents the output of the attention structure; The dilated convolution block consists of a 3×3 convolutional layer with a dilation rate of 2, a LeakyReLU activation function, and a batch normalization layer; The SE attention structure consists of global average pooling, channel attention calculation, and channel feature recalibration; S214: For the phase enhancement branch, a residual block is used. The residual block consists of two consecutive 3x3 convolutions connected by a relu activation function, and then the input is skipped and connected to the output to obtain the enhanced phase : (5); in: Represents the nonlinear mapping of the phase through the residual block; S215: The enhanced amplitude and Phase Perform an inverse Fourier transform and reconstruct the image to obtain the output of the first stage : (6); in: Stands for Inverse Fourier Transform.

[0023] S22: The attacked image is further processed through multi-scale convolution to ensure that the watermark information is completely destroyed while improving the image quality.

[0024] The specific steps of S22 are: S221: Output to the first stage Use encoder Downsampling; S222: processed by six multi-scale convolutional layers respectively; Use first The convolutional layer performs dimensionality reduction; Then, multiple convolution kernels of different sizes are used to extract multi-scale features. The convolution kernel sizes are , , and , perform sigmoid operations on the output features respectively and multiply them with the output features of different convolution kernels; Then concatenate and restore the dimension using a 1x1 convolutional layer; Finally, it is combined with the input features through residual connection.

[0025] This can capture a wide range of spatial features across multiple scales and enhance the overall representation of spatial structures.

[0026] S223: The generated features are fed to the decoder , generating the final attacked image .

[0027] S3: Construct loss function.

[0028] In order to remove the watermark from the image while taking into account the visual quality of the image, the loss function consists of three parts: ; ; ; (7); in: is the output image of the first stage With real image The mean square error loss of is the second stage output image With real image The mean square error loss of It is a perceptual loss based on VGG features to improve visual effects. VGG is a deep convolutional neural network proposed by the VGG team at Oxford University. Representation paradigm for calculating the Euclidean distance of vectors; Represents the first Layer feature extraction results; , and is the loss weight parameter, which is used to balance the impact of different losses.

[0029] S4: Construct evaluation indicators.

[0030] In order to better measure the changes in image quality, peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) were used as metrics; PSNR measures the overall quality of an image. The peak signal-to-noise ratio is calculated as follows: (8); in: Represents a watermarked image; Represents the image after the attack; Indicates the height of the image; Indicates the width of the image; Represents the square of the maximum pixel value of the image; The original watermarked image is The pixel value at ; The original attacked image is The pixel value at ; The higher the PSNR value, the smaller the distortion between the processed image and the original image, that is, the quality of the processed image is closer to the original image.

[0031] SSIM is an indicator used to measure the similarity between two images. It focuses more on evaluating the structural and visual similarities of images. It can better reflect the human eye's perception of image quality than PSNR.

[0032] The calculation formula of the structural similarity index is as follows: (9); in: and The original images And the image after the attack ; , Represent the original images And the image after the attack The mean of , Represent the original images And the image after the attack The standard deviation of Represents the original image And the image after the attack The covariance of and is a constant used to prevent the denominator from being zero; The value of SSIM ranges from 0 to 1, where 1 means that the two images are exactly the same and 0 means there is no similarity at all. SSIM takes into account three factors that are closely related to human visual perception: brightness, contrast, and structure.

[0033] In addition, in order to evaluate the effect of watermark removal, the bit error rate (BER) is used as an evaluation indicator, and its calculation formula is as follows: (10); in: The number of bits representing the error information in the extracted watermark; Indicates the total number of bits of the original watermark information. When the BER value is closer to 0, it means that the extracted watermark information is more complete, which means that the watermark attack effect is poor; conversely, if the BER value is larger, it means that the watermark attack effect is better.

[0034] The watermark attack network includes a Fourier enhancement module and a multi-scale convolution module. The Fourier enhancement module consists of a phase enhancement branch and an amplitude enhancement branch. The multi-scale convolution module consists of an encoder. E , multi-scale convolutional layer, decoder D composition.

[0035] The encoder in the autoencoder in the pre-trained vqgan network used by the present invention E and decoder D The autoencoder is composed of the encoder and decoder in the pre-trained VQGAN. This part is used in the second stage multi-scale convolution. First, the image generated in the first stage is encoded, the image features after the first stage watermark attack are extracted, and then multi-scale convolution is performed, and finally decoding is performed to restore the image after the attack.

[0036] The above disclosure is only a specific embodiment of the present invention, but the present invention is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.

Claims

1. A watermark attack method based on Fourier transform and multi-scale convolution, characterized in that: The following steps are involved: S1: Dataset acquisition; S2: Construct a watermark attack network, which mainly includes the Fourier enhancement stage and the multi-scale convolution stage; S21: The Fourier enhancement stage independently enhances the amplitude and phase to maximize the destruction of watermark information; S22: Further process the attacked image through multi-scale convolution to ensure that the watermark information is completely destroyed while improving the image quality; S3: construct loss function; S4: Construct evaluation indicators.

2. The watermark attack method based on Fourier transform and multi-scale convolution according to claim 1 is characterized in that: The specific steps of S21 are: S211: First input the watermark image ; The amplitude component is obtained by fast Fourier transform equation 1, equation 2 and equation 3 and phase component ; (1) in: is the complex spectrum; is the height of the image; is the width of the image; are image space coordinates; is the frequency domain coordinate; is an imaginary unit; (2); (3); in: and Respectively The real and imaginary parts of S212: Use 1×1 convolution layer, LeakyReLU activation function and 1x1 convolution layer to extract amplitude and phase components respectively, and get and ; S213: For the amplitude enhancement branch, the dilated convolution block and SE attention structure are used to enhance the amplitude to obtain the enhanced amplitude : (4); in: Represents the output after dilated convolution; Represents the output of the attention structure; S214: For the phase enhancement branch, use the residual block processing, and then jump connect the input to the output to obtain the enhanced phase : (5); in: Represents the nonlinear mapping of the phase through the residual block; S215: The enhanced amplitude and Phase Perform an inverse Fourier transform and reconstruct the image to obtain the output of the first stage : (6); in: Stands for Inverse Fourier Transform.

3. The watermark attack method based on Fourier transform and multi-scale convolution according to claim 2 is characterized in that: The specific steps of S22 are: S221: Output to the first stage Use encoder Downsampling; S222: processed by six multi-scale convolutional layers respectively; S223: The generated features are fed to the decoder , generating the final attacked image .

4. The watermark attack method based on Fourier transform and multi-scale convolution according to claim 3 is characterized in that: In order to remove the watermark from the image while taking into account the visual quality of the image, the loss function consists of three parts: ; ; ; (7); in: is the first stage output image With real image The mean square error loss of is the second stage output image With real image The mean square error loss of It is a perceptual loss based on VGG features to improve visual effects; Representation paradigm for calculating the Euclidean distance of vectors; Represents the first Layer feature extraction results; , and is the loss weight parameter.

5. The watermark attack method based on Fourier transform and multi-scale convolution according to claim 4 is characterized in that: Peak signal-to-noise ratio and structural similarity index were used as measurement indicators; The peak signal-to-noise ratio is calculated as follows: (8); in: Represents a watermarked image; Represents the image after the attack; Indicates the height of the image; Indicates the width of the image; Represents the square of the maximum pixel value of the image; The original watermarked image is The pixel value at ; The original attacked image is The pixel value at ; The calculation formula of the structural similarity index is as follows: (9); in: and The original images And the image after the attack ; , Represent the original images And the image after the attack The mean of , Represent the original images And the image after the attack The standard deviation of Represents the original image And the image after the attack The covariance of and is a constant used to prevent the denominator from being zero; The bit error rate is also used as an evaluation indicator, and its calculation formula is as follows: (10); in: The number of bits representing the error information in the extracted watermark; Represents the total number of bits of the original watermark information.

6. The watermark attack method based on Fourier transform and multi-scale convolution according to claim 1 is characterized in that: The watermark attack network includes a Fourier enhancement module and a multi-scale convolution module. The Fourier enhancement module consists of a phase enhancement branch and an amplitude enhancement branch, and the multi-scale convolution module consists of an encoder. E , multi-scale convolutional layer, decoder D composition.

Citation Information

Patent Citations

  • Hidden digital watermark attack method and system based on SAD network

    CN115358909A

  • Hidden watermark attack algorithm based on wavelet transform and attention mechanism

    CN116308986A

  • Active Deepfake detection method based on watermark image difference value

    CN118799163A

  • Method for watermarking depth image based on mixed frequency-domain channel attention

    US20240054594A1

Cited By

  • Multi-scale photoacoustic microscopic imaging method based on single-pixel imaging

    CN120177376A