Image watermark robustness evaluation method and device
By comprehensively stress testing and evaluating the robustness of image watermarking algorithms through various attacks, the vulnerability problem in existing technologies is solved, and a systematic evaluation method and device are provided to guide the improvement and enhancement of watermarking technology.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA ACADEMY OF INFORMATION & COMM
- Filing Date
- 2024-12-25
- Publication Date
- 2026-06-26
AI Technical Summary
Existing image watermarking technologies are vulnerable to complex attacks and lack systematic robustness evaluation methods, resulting in an incomplete understanding of the vulnerability and robustness of algorithms in the real world.
This paper provides a method and apparatus for evaluating the robustness of image watermarking. Through comprehensive stress testing, including data acquisition, watermark printing, various attacks, watermark recognition, and image quality measurement, the robustness of the watermarking algorithm is quantified, and an evaluation report is generated.
Through comprehensive testing and various stress tests, the robustness of watermarking algorithms and image quality degradation are evaluated, providing a new perspective for comprehensive assessment and guiding the improvement and enhancement of watermarking technology.
Smart Images

Figure CN122288960A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and apparatus for evaluating the robustness of image watermarking. Background Technology
[0002] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section.
[0003] With the rapid development of text-to-image diffusion models, the public and the AI community have shown great interest in generating images of comparable quality to human creations. These technological advancements have created a demand for reliable algorithms capable of detecting AI-generated content and determining its origin. Against this backdrop, image watermarking has gained widespread attention as a technology for identifying the source or ownership of content. Image watermarking ensures traceability of content while maintaining image quality by encoding an invisible watermark signal onto the image.
[0004] Existing watermarking techniques face the challenge of balancing image quality degradation with watermark detection effectiveness, as well as the challenge of resisting various image manipulations and adversarial attacks. Examples include using deep encoder-decoder networks for watermark embedding and decoding; embedding watermarks in the initial noise vector of a diffusion model; and embedding watermarks by training on the decoder of a latent diffusion model. While these methods can resist common image manipulations to some extent, they remain vulnerable to more complex attacks, such as regeneration attacks and adversarial attacks.
[0005] However, the lack of comprehensive, systematic, and mature evaluation methods in existing technologies has led to an incomplete understanding of the vulnerability and robustness of these algorithms in the real world. Summary of the Invention
[0006] This invention provides an image watermark robustness evaluation method, offering a complete, systematic, and mature watermarking algorithm evaluation method. Through comprehensive stress testing, the robustness of the image watermarking algorithm is evaluated. The method includes:
[0007] Obtain the original image dataset;
[0008] The watermarking algorithm to be evaluated is used to print watermarks on a portion of the original images in the original image dataset to form a watermarked image dataset.
[0009] Attack all images in the watermarked image dataset to obtain the attacked image dataset; the attack includes one or any combination of distortion attack, regeneration attack, and adversarial attack;
[0010] Watermark recognition is performed on all images in the image dataset after the attack, and the watermark recognition results are obtained. The watermark recognition results include the true positive rate and the false positive rate. The true positive rate represents the proportion of watermarks correctly identified in all actual images containing watermarks, and the false positive rate represents the proportion of watermarks incorrectly identified in all actual images without watermarks.
[0011] Using a preset image quality measurement algorithm, the image quality measurement values before and after the attack are calculated for each image with printed watermark, which are used as the result of image quality degradation. The preset image quality measurement algorithm is used to calculate and standardize various image quality measurement elements.
[0012] Based on the watermark recognition results and image quality degradation results, the image watermark robustness evaluation results are determined; the image watermark robustness evaluation results include an image watermark robustness evaluation report, which includes data quantifying the performance of the watermarking algorithm to be evaluated.
[0013] This invention also provides an image watermark robustness evaluation device to provide a complete, systematic, and mature watermarking algorithm evaluation method. Through comprehensive stress testing, the robustness of the image watermarking algorithm is evaluated. The device includes:
[0014] The data acquisition module is used to acquire the raw image dataset;
[0015] The watermark printing module is used to print watermarks on a portion of the original images in the original image dataset using the watermarking algorithm to be evaluated, thus forming a watermarked image dataset.
[0016] The attack module is used to attack all images in the watermarked image dataset to obtain the attacked image dataset; the attack includes one or any combination of distortion attack, regeneration attack, and adversarial attack.
[0017] The watermark detection module is used to identify watermarks in all images in the image dataset after the attack, and obtain watermark identification results. The watermark identification results include the true positive rate and the false positive rate. The true positive rate represents the proportion of watermarks correctly identified in all actual images containing watermarks, and the false positive rate represents the proportion of watermarks incorrectly identified in all actual images without watermarks.
[0018] The quality measurement module is used to calculate the image quality measurement values before and after the attack for each image with a printed watermark using a preset image quality measurement algorithm, as the result of image quality degradation; the preset image quality measurement algorithm is used to calculate and standardize various image quality measurement elements of the image;
[0019] The result output module is used to determine the image watermark robustness evaluation result based on the watermark recognition result and the image quality degradation result. The image watermark robustness evaluation result includes an image watermark robustness evaluation report, which includes data that quantifies the performance of the watermarking algorithm to be evaluated.
[0020] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described image watermark robustness evaluation method.
[0021] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described image watermark robustness evaluation method.
[0022] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described image watermark robustness evaluation method.
[0023] In this embodiment of the invention, an original image dataset is obtained; a watermark is printed on a portion of the original images in the original image dataset using the watermarking algorithm to be evaluated, forming a watermarked image dataset; all images in the watermarked image dataset are attacked to obtain an attacked image dataset; the attack includes one or any combination of distortion attack, regeneration attack, and adversarial attack; watermark recognition is performed on all images in the attacked image dataset to obtain watermark recognition results; the watermark recognition results include true positive rate and false positive rate; the true positive rate represents the proportion of correctly identified watermarks in all actual images containing watermarks, and the false positive rate represents the proportion of incorrectly identified watermarks in all actual images without watermarks; a preset image quality measurement algorithm is used to calculate the image quality measurement values before and after the attack for each image with printed watermarks, as the image quality degradation result; the preset image quality measurement algorithm is used to calculate and standardize various image quality measurement elements for the image; based on the watermark recognition results and the image quality degradation results, the image watermark robustness evaluation results are determined; the image watermark robustness evaluation results include an image watermark robustness evaluation report, which includes data quantifying the performance of the watermarking algorithm to be evaluated. This invention provides a complete, systematic, and mature method for evaluating watermarking algorithms. By comprehensively testing detection and recognition tasks as well as various stress tests, the robustness of image watermarking algorithms is evaluated. This invention considers two key dimensions: watermark detection performance and image quality degradation, and proposes an evaluation report on watermark detection performance and image quality, providing a new perspective for a comprehensive evaluation of the security and practicality of watermarking algorithms. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0025] Figure 1 This is a flowchart illustrating the image watermark robustness evaluation method in an embodiment of the present invention.
[0026] Figure 2 This is a specific example of the image watermark robustness evaluation method in this invention.
[0027] Figure 3 This is another specific example of the image watermark robustness evaluation method in the embodiments of the present invention;
[0028] Figure 4 This is a schematic diagram of the image watermark robustness evaluation device in an embodiment of the present invention;
[0029] Figure 5 This is a schematic diagram of a computer device in an embodiment of the present invention. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0031] To facilitate a clear description of the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish the same or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order.
[0032] The acquisition, transmission, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.
[0033] This invention aims to address the shortcomings of existing image watermarking evaluation methods by providing a comprehensive framework / method for evaluating image watermark robustness. This overcomes the limitations of evaluating image watermark robustness under diverse and complex attack methods. Addressing the lack of standardized testing procedures and consistent quality metrics in existing technologies, the proposed framework / method conducts comprehensive stress tests using a series of realistic and novel attack methods to evaluate the robustness of different image watermarking algorithms. This framework / method comprehensively considers image quality degradation and watermark detection performance, utilizes multiple image similarity metrics, and reveals the potential weaknesses of current watermarking algorithms through two-dimensional performance vs. quality graphs. This provides guidance and tools for developing more powerful and reliable watermarking technologies, thereby promoting the practical application and security improvement of image watermarking technology.
[0034] Figure 1 This is a flowchart illustrating the image watermark robustness evaluation method in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:
[0035] Step 101: Obtain the original image dataset;
[0036] Step 102: Use the watermarking algorithm to be evaluated to print watermarks on a portion of the original images in the original image dataset to form a watermarked image dataset;
[0037] Step 103: Attack all images in the watermarked image dataset to obtain the attacked image dataset; the attack includes one or any combination of distortion attack, regeneration attack, and adversarial attack;
[0038] Step 104: Perform watermark recognition on all images in the attacked image dataset to obtain watermark recognition results; the watermark recognition results include the true positive rate and the false positive rate; the true positive rate represents the proportion of watermarks correctly identified in all actual images containing watermarks, and the false positive rate represents the proportion of watermarks incorrectly identified in all actual images without watermarks.
[0039] Step 105: Using a preset image quality measurement algorithm, calculate the image quality measurement values before and after the attack for each image with printed watermark, as the result of image quality degradation; the preset image quality measurement algorithm is used to calculate and standardize various image quality measurement elements for the image;
[0040] Step 106: Based on the watermark recognition results and image quality degradation results, determine the image watermark robustness evaluation results; the image watermark robustness evaluation results include an image watermark robustness evaluation report, which includes data quantifying the performance of the watermarking algorithm to be evaluated.
[0041] from Figure 1As can be seen from the process shown, the embodiments of the present invention provide a complete, systematic and mature watermarking algorithm evaluation method. By comprehensively detecting and recognizing tasks and various stress tests, the robustness of the image watermarking algorithm is evaluated. The embodiments of the present invention consider two key dimensions: watermark detection performance and image quality degradation, and propose evaluation reports on watermark detection performance and image quality, providing a new perspective for a comprehensive evaluation of the security and practicality of watermarking algorithms.
[0042] The image watermark robustness evaluation method in the embodiments of the present invention will be explained in detail below.
[0043] Step 101: Obtain the original image dataset.
[0044] During implementation, some typical datasets can be obtained.
[0045] In a preferred embodiment, obtaining the original image dataset may include:
[0046] Acquire multiple datasets; the datasets include a set of real-world photos, a set of images generated by existing diffusion models, and a set of images generated by existing text-based image models.
[0047] The images in multiple datasets are ranked by removing irrelevant columns, filtering prompt strings, and using an aesthetic scoring algorithm. Images ranked above the preset ranking are selected to form the original image dataset.
[0048] For example, three datasets were used to evaluate the robustness of the watermarking algorithm: Dataset 1, Dataset 2, and Dataset 3. Dataset 1 is a large-scale text-to-image cue dataset containing tens of millions of images generated by real users using specified cuees and hyperparameters through existing diffusion models. Dataset 2 consists of hundreds of thousands of real-world photographs, many of which are annotated. Images in Dataset 2 can be used for tasks such as object detection, segmentation, and keypoint detection in computer vision. Images in Dataset 3 are generated by a text-to-image model. Algorithms were used, including removing irrelevant columns, filtering cue strings, and ranking images based on aesthetic scores. The final selected images ranked highly aesthetically, as high-quality AI-generated images are more important and practical for copyright protection. Therefore, each dataset ultimately contained 5000 image samples. Together, these three datasets provide diverse image types and contexts, covering both generative models and real-world scenes, suitable for evaluating practical watermarking applications.
[0049] In step 102, the watermarking algorithm to be evaluated is used to print watermarks on a portion of the original images in the original image dataset to form a watermarked image dataset.
[0050] During implementation, multiple watermarking algorithms can be used to print watermarks on portions of the original image dataset, forming watermarked image datasets corresponding to the watermarking algorithms, so as to compare multiple watermarking algorithms.
[0051] In step 103, all images in the watermarked image dataset are attacked to obtain the attacked image dataset; the attack includes one or any combination of distortion attack, regeneration attack, and adversarial attack.
[0052] In one embodiment, the distortion attack includes one or any combination of geometric distortion, photometric distortion, and degradation distortion; the regeneration attack includes image regeneration using a deep learning model; and the adversarial attack includes embedding attacks and proxy detector attacks.
[0053] In this embodiment of the invention, a carefully designed set of attacks simulates various threats that may be encountered in the real world. The impact of traditional image processing operations is evaluated, and complex attack methods utilizing deep learning models are also considered. By quantitatively evaluating the effects of these attacks, the framework compares the performance changes of watermarked images before and after the attacks, thereby determining the specific impact of each attack on watermark detection.
[0054] Distortion attacks assess the robustness of watermarking techniques by simulating various degradations that images may suffer during natural transmission, processing, or human editing. These attacks include geometric distortions, such as rotation and cropping, which test the watermark's detection capability after structural changes by altering the spatial layout and size of the image; photometric distortions, such as brightness and contrast adjustments, which evaluate the stability of the watermark under changes in visual characteristics by simulating different lighting conditions or editing effects; and degenerative distortions, such as Gaussian blur and noise addition, which simulate the natural degradation of image quality by introducing noise or reducing resolution, thereby testing the watermark's persistence and reliability when the image is damaged.
[0055] Regeneration attacks simulate the process of regenerating images using deep learning models to test the persistence of watermarks after complex image transformations. These attacks leverage the capabilities of variational autoencoders (VAEs) or diffusion models to generate visually similar images, but with the watermark potentially removed, by adding noise to the image in its latent space and then denoising it. The core idea is to regenerate the watermarked image using a pre-trained VAE or diffusion model. This process typically involves encoding the image into a latent representation space and then reconstructing the image by adding noise and progressively denoising. Because this process does not depend on the specific model parameters used when the original watermark was embedded, it may unintentionally corrupt or remove the embedded watermark information, thus reducing the watermark's detectability.
[0056] Adversarial attacks primarily employ two methods: embedding attacks and proxy detector attacks. The basic principle of embedding attacks is to introduce small perturbations into the latent feature space of an image, thereby altering certain attributes of the image without affecting its overall visual quality. These perturbations are designed to mislead the watermark detector (used for watermark recognition), preventing it from correctly identifying or extracting the watermark information embedded in the image. Attackers typically choose to implement these perturbations on specific regions or features of the image, as these regions or features are crucial for watermark extraction. In the proposed framework / method, embedding attacks are implemented using pre-trained encoders such as ResNet18 or CLIP (ResNet18 is an 18-layer residual network, a deep convolutional neural network architecture; CLIP is a contrastive language-image pre-trained model). Attackers leverage these encoders to map the image into a latent feature space and then apply adversarial perturbations within this space. These perturbations are designed to maximize the false positive rate of the watermark detector while maintaining the image's visual quality. In this way, attackers can attempt to remove or weaken the watermark signal without the attacker's knowledge.
[0057] Proxy detector attacks are techniques that simulate attacker behavior, where the attacker constructs a model—a proxy detector—to mimic a real watermark detection system. This proxy detector is trained using available watermarked and unwatermarked images. The attacker's goal is to leverage this proxy detector to simulate the functionality of the real watermark detector and attempt to apply attack strategies effective against the proxy detector to the real watermark detector, thereby reducing the real detector's ability to recognize watermarks. In short, a proxy detector is a tool that attackers use to simulate and attack the watermark detection process, with the aim of finding ways to interfere with the real watermark detector.
[0058] This invention considers three different settings for training a proxy detector: using watermarked and unwatermarked images from a victim-generated model (AdvCls-UnWM&WM), using real and watermarked images from a pre-established or existing visual object recognition database (AdvCls-Real&WM), and using only watermarked images from two different users (AdvCls-WM1&WM2). In AdvCls-UnWM&WM: In this setting, the attacker acquires or generates two sets of images: unwatermarked images (UnWM), which are not watermarked by the victim-generated model and can be real photos or images generated by other models; and watermarked images (WM), which are generated by the victim-generated model and already have a watermark embedded. The attacker uses these two sets of images to train a proxy detector to learn to distinguish which images are watermarked and which are not. This proxy detector is then used to test adversarial attacks, i.e., attempts to influence the performance of the real watermark detector by attacking the proxy detector. Here, AdvCls stands for Adversarial Classifier, meaning that in this setup, the attacker trains a classifier to distinguish between watermarked and unwatermarked images. UnWM&WM stands for Unwatermarked & Watermarked, meaning the training data includes unwatermarked images (UnWM) and watermarked images (WM). AdvCls-Real&WM: This attack setup involves training a surrogate detector capable of distinguishing between real (unwatermarked) and watermarked images. In this setup, the attacker uses real images sampled from a visual object recognition database (not sampled from the generative model) and watermarked images to train the surrogate detector. The attacker's goal is to use this trained surrogate detector to attack the actual watermark detector, attempting to deceive it through adversarial perturbations, preventing it from correctly recognizing the watermark. In AdvCls-WM1 & WM2, the attacker trains a proxy detector using only watermarked images embedded with two different watermarked messages. This effectively trains an alternative watermarked message classifier to distinguish between the watermarks of the two users. The attacker collects watermarked images from two users, each using a specific watermarked message. The attacker's goal is to use this trained proxy detector to attack the actual watermark detector, attempting to interfere with its judgment through adversarial perturbations, causing it to fail to correctly identify or incorrectly identify the watermarked message. These proxy detectors are fine-tuned on a ResNet18 network to achieve watermarked image classification.After training, attackers use the Projective Gradient Descent (PGD) algorithm to attack the proxy detector, aiming to deceive it by altering the latent features of the image, thereby removing or implanting watermarks. These attacks attempt to generate adversarial images from the original image, causing the proxy detector to incorrectly assign target labels to them. The victim-generated model refers to the specific image generation model targeted by the attacker—the model the attacker attempts to disrupt or deceive. This model is typically developed by an organization or individual to generate watermarked images for content copyright protection, source tracking, or other security purposes. The victim-generated model may be a commercial or open-source image generation system, such as an image synthesis model using deep learning techniques, including but not limited to systems based on GANs (Generative Adversarial Networks), VAEs (Variational Autoencoders), or diffusion models.
[0059] In this embodiment of the invention, when attacking all images in the watermarked image dataset, attacks of arbitrary order and different intensities can be employed. These intensity levels define the severity of the attack. For example, in distortion attacks, different degrees of rotation, cropping, and brightness adjustment may be included. Specifically, the main types of distortion attacks include: rotation attacks with intensities ranging from 9 to 45 degrees; cropping attacks with intensities ranging from 10% to 50% of the image area; and geometric, photometric, and degradation distortions, which also have different relative intensity levels ranging from 0.05 to 0.45, as well as five uniformly distributed normalized intensities ranging from 0.05 to 0.20.
[0060] Image watermark robustness evaluation refers to measuring the ability of a watermark to still be correctly detected and extracted after various attacks or manipulations. A robust watermark should be able to resist common attacks without affecting the quality of the original image.
[0061] In step 104, watermark recognition is performed on all images in the post-attack image dataset to obtain the watermark recognition results.
[0062] The watermark detection method proposed in this invention can be described as follows: given an image generator θ G Given a set of watermark information M, the watermark detection problem requires the generator to generate a watermarked image x = EMBED(θ). G The detector is able to accurately recover the watermark information m′ = DECODE(x′) from the attacked image x′ = A(x) under attack A. The detection accuracy is verified through a VERIFY process. α The value is determined by (m′, m), where α is a preset threshold used to control the false alarm rate.
[0063] In this context, Attack A refers to various operations performed on the watermarked image. These operations may aim to destroy or hide the watermark, such as image compression, cropping, rotation, and adding noise. The image x′ after the attack refers to the image processed by Attack A, that is, the result obtained after the original watermarked image x has been processed by Attack A.
[0064] The detector recovers the watermark information m′: This refers to the process of extracting the watermark information from the attacked image x′ using a watermark detection algorithm (DECODE). This extracted watermark information m′ may not be exactly the same as the original embedded watermark information m, but it should be close enough that the verification process can recognize the watermark.
[0065] The DECODE watermark detection algorithm works as follows: it assumes that the watermark information is embedded in the carrier data in a specific way. By performing specific mathematical transformations, statistical analyses, or pattern recognition on the watermarked data, it attempts to extract the hidden watermark information.
[0066] In a preferred embodiment, the watermark recognition result further includes a specified message embedded in the original image; performing watermark recognition on all images in the attacked image dataset to obtain watermark recognition results may include: performing watermark recognition on all images in the attacked image dataset using a preset watermark detection method to obtain watermark recognition results; the preset watermark detection method includes one of template matching, edge detection, and deep learning algorithms.
[0067] In this embodiment, other preset watermark detection methods besides the DECODE watermark detection algorithm can be used, and no limitation is made in this embodiment of the invention.
[0068] In the watermark recognition problem, in addition to requiring the detector to be able to label an image as a watermark, it is also necessary to accurately identify specific messages embedded in the image. This requires that, given a set of possible information Q, the detector can minimize the probability of incorrectly assigning the attacked image x′ to non-original information in Q.
[0069] In watermark detection and recognition, the p-value is used to measure the probability of the observed watermark strength in the image after an attack occurring randomly.
[0070] p = P m (D(ω,m′)<D(m,m′)|H0);
[0071] Where D(ω, m′) is the similarity between any information and m′, and D(m, m′) is the similarity between the real information m and the recovered information m′. H0 represents the null hypothesis of generating the image without knowing the watermark, and P... m Let P represent the probability distribution of a message ω randomly drawn from the message space M. Specifically, P mThis is a probabilistic model used to generate random messages ω, which are compared with messages m' recovered from the watermarked image. If p < α, VERIFY α (m′,m) returns 1, otherwise returns 0.
[0072] In this invention, each attack, such as distortion attack, regeneration attack, and adversarial attack, has different variations and intensities. These attacks are arranged and combined to create a comprehensive attack set to test the performance of the watermarking algorithm under various conditions. For example, a single distortion attack can have multiple intensity levels, a regeneration attack can have different number of iterations, and an adversarial attack can target different embedding models. In this way, this invention can generate a variety of attack methods and intensities, including 26 different attacks, each with different methods and intensities.
[0073] In this embodiment of the invention, the framework / method measures the True Positive Rate (TPR) and False Positive Rate (FPR) of watermark detection under each attack. The True Positive Rate (TPR) refers to the proportion of all images that actually contain a watermark, in which the detector correctly identifies the watermark. A high TPR means fewer watermarks are missed. The False Positive Rate (FPR) refers to the proportion of all images that actually do not contain a watermark, in which the detector incorrectly identifies them as containing a watermark. A high FPR means more non-watermarked images are falsely detected.
[0074] These two metrics are measured to evaluate the robustness and accuracy of watermark detection algorithms under different attacks. The TPR (Total Recovery Rate) measurement focuses on whether the watermark detector can accurately recover the watermark information from the attacked image when an attack is present. This is a direct test of watermark robustness, i.e., how detectable the watermark is under various attack conditions. The FPR (Failure Rate) measurement focuses on whether the watermark detector will falsely detect the watermark when faced with an unwatermarked image. This is a test of the watermark detector's specificity, i.e., the detector should not generate false positives when no watermark is present.
[0075] The True Positive Rate (TPR) focuses on the recoverability of the watermark after an attack, while the False Positive Rate (FPR) focuses on the stability and accuracy of the detector without a watermark. Together, these two metrics provide a comprehensive perspective for the design and evaluation of watermark detection systems.
[0076] In step 105, a preset image quality measurement algorithm is used to calculate the image quality measurement values before and after the attack for each image with printed watermark, which are used as the image quality degradation result. The preset image quality measurement algorithm is used to calculate and standardize various image quality measurement elements for the image.
[0077] In one embodiment, the image quality metric includes one or any combination of Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Normalized Mutual Information (NMI), Fraser Initial Distance (FID), CLIP-based FID, Learned Perceptual Image Patch Similarity (LPIPS), Aesthetic Score, and Artifact Score.
[0078] In this embodiment of the invention, multiple image quality metrics (such as PSNR, SSIM, FID, etc.) are used to evaluate image quality degradation. Since different quality metrics have different dimensions and ranges, each metric is standardized. Specifically, based on the distribution of metric values across all attacked images, the value of each metric is mapped to the interval [0.1, 0.9]. This step ensures the comparability between different metrics and makes them equally important in the evaluation.
[0079] After standardization, these metrics are aggregated into a unified image quality degradation metric. Aggregation is achieved by calculating the average of each category metric (such as image similarity, distribution distance, etc.), and then averaging these category averages. This results in a single metric that comprehensively reflects image quality degradation.
[0080] Then, for each attacked image, image quality metrics before and after the attack were calculated, and a uniform metric was used to assess the degree of quality degradation.
[0081] Finally, in step 106, the image watermark robustness evaluation result is determined based on the watermark recognition result and the image quality degradation result; the image watermark robustness evaluation result includes an image watermark robustness evaluation report, which includes data quantifying the performance of the watermarking algorithm to be evaluated.
[0082] During implementation, the robustness of the watermark is scored based on the aforementioned performance indicators (TPR, FPR, PSNR, SSIM, NMI, etc.).
[0083] We analyzed which types of attacks had the greatest impact on watermarks and which processes had almost no impact on them, taking into account both the watermark's invisibility and robustness.
[0084] A detailed evaluation report should be prepared, which includes a series of indicators and lists to quantify the performance of the watermarking system, as well as experimental methods, tools and techniques used, results obtained and conclusions, and suggestions for improving the watermarking algorithm.
[0085] Table 1 shows the detection performance of the three watermarking algorithms after the attack, measured by the average TPR@0.1%FPR. The results in Table 1 show that watermarking algorithm 1 exhibits better robustness among all watermarking methods because it was designed to be robust to distortions of the real world. In contrast, watermarking algorithm 2 is particularly sensitive to hostile attacks, especially those with access to the same VAE used for detection in a gray-box setting. Furthermore, watermarking algorithm 2 also shows significant vulnerability to regeneration attacks. Watermarking algorithm 3 is relatively sensitive to most regeneration attacks. All three watermarking algorithms maintain relative robustness against distortion attacks. In Table 1, a gray-box setting refers to an attacker possessing some, but not all, knowledge about the watermarking algorithm. Specifically, the attacker knows that a specific model is used as the encoder, meaning the attacker has some understanding of the internal workings of the watermarking algorithm, but this understanding is not complete. A black-box setting refers to an attacker having no internal knowledge of the watermarking algorithm and can only attempt to break the watermark based on the input and output. The attacker uses a different encoder and has no access to information about the specific VAE model used for watermarking.
[0086] Table 1
[0087]
[0088] The current main attack methods were tested based on their impact on detection performance and image quality. The attacks were evaluated using performance thresholds (TPR@0.1%FPR=0.95 and TPR@0.1%FPR=0.7) and quality degradation at these thresholds (Q@0.95P and Q@0.7P).
[0089] A detailed comparison and analysis of the attack results of watermarking algorithm 1 under different attack settings was conducted. Table 2 lists the average quality degradation (Q) and performance (P) of various attack methods at two different performance thresholds (TPR@0.1%FPR=0.95 and TPR@0.1%FPR=0.7). The results show that for the two distortion attacks, Dist-Rotation and Dist-RCrop, the performance of watermarking algorithm 1 decreases slightly, but the quality degradation is relatively small. For the more extreme distortion attack, Dist-Erase, the performance degradation is more significant, but the quality degradation remains at a low level. In the regeneration attacks Regen-Diff, Regen-DiffP, and Regen-VAE, watermarking algorithm 1 shows extremely high robustness, with almost no performance degradation and very small quality degradation. The adversarial embedding attacks AdvEmbG-KLVAE8, AdvEmbB-RN18, and AdvEmbB-CLIP, and the adversarial proxy detector attacks AdvCls-UnWM&WM, AdvCls-Real&WM, and AdvCls-WM1&WM2 have relatively small impacts on watermarking algorithm 1, with minimal performance degradation and low quality degradation. This indicates that watermarking algorithm 1 can maintain high watermark detection performance and image quality even when facing various attacks, demonstrating good robustness.
[0090] in:
[0091] Dist-Rotation: A rotation distortion attack refers to rotating an image to test the robustness of a watermark under this geometric transformation.
[0092] Dist-RCrop: A cropping distortion attack, which involves cropping an image to test the robustness of a watermark under this geometric transformation.
[0093] Dist-Erase: An eraser distortion attack refers to erasing certain parts of an image to test the robustness of a watermark after the image portion has been erased.
[0094] Regen-Diff: A diffusion regeneration attack refers to the attempt to remove or destroy a watermark by regenerating an image using a diffusion model.
[0095] Regen-DiffP: A hint-diffusion regeneration attack, which refers to using a diffusion model to regenerate an image in an attempt to remove or destroy a watermark when the image has a known hint.
[0096] Regen-VAE: VAE regeneration attack refers to using a variational autoencoder (VAE) to regenerate an image in an attempt to remove or destroy the watermark.
[0097] AdvEmbG-KLVAE8: Gray-box KL-VAE attack with a strength of 8. It refers to an adversarial embedding attack performed using the KL-VAE model (a variational autoencoder based on Kullback-Leibler divergence) in a gray-box setting, where "8" indicates the strength of the attack or the size of the perturbation.
[0098] AdvEmbB-RN18: Black-box ResNet18 attack, refers to an adversarial embedding attack performed using a pre-trained ResNet18 model under black-box settings.
[0099] AdvEmbB-CLIP: Black-box CLIP attack refers to an adversarial embedding attack performed using the CLIP model under black-box settings.
[0100] Adversoc-UnWM & WM: Adversarial Classifier - Watermark-Free and Watermarked. This refers to a setup in adversarial attacks where "UnWM" represents an image without a watermark, and "WM" represents an image with a watermark. In this setup, the attacker trains a classifier to distinguish between images without and with watermarks.
[0101] AdvCls-Real&WM: Adversarial Classifier - Real and Watermarked refers to another adversarial attack setup where "Real" represents real-world images (images not generated by a generative model), and "WM" represents watermarked images. This setup is used to simulate a scenario where an attacker uses real-world images and watermarked images to train a classifier.
[0102] AdvCls-WM1&WM2: Adversarial Classifiers - Watermark 1 and Watermark 2 refer to an alternative setup in adversarial attacks, where "WM1" and "WM2" represent two different watermarks. In this setup, an attacker trains a classifier to distinguish between images generated by two different users with different watermarks; this is commonly used in attacks on user identification tasks.
[0103] Table 2
[0104]
[0105] Table 2 compares the attack results of watermarking algorithm 1 in the detection settings. Q represents the normalized quality degradation, and P represents the performance. inf indicates that the attack strength of all tests produces a performance higher than 0.95, and -inf indicates all tests with a performance lower than 0.95.
[0106] Figure 2 This is a specific example diagram of the image watermark robustness evaluation method in an embodiment of the present invention, as shown below. Figure 2 The diagram illustrates the overall evaluation process of an embodiment of the present invention. Figure 2The system generates a watermarked image based on input prompts (either a diffusion model or a text-based image model). Then, it applies attacks of varying strengths to obtain images with corresponding strengths 1, 2, and 3. A watermark detector identifies the watermark and compares it with the previous watermark-free image, outputting performance and quality metrics (graphics on true positive rate, false positive rate, etc.), and the relationship between watermark performance and image quality degradation (graphics). The vertical axis of the graph represents TPR@0.1%FPR, and the horizontal axis represents the aforementioned performance metrics such as PSNR and SSIM.
[0107] Figure 3 This is another specific example of the image watermark robustness evaluation method in the embodiments of the present invention, as shown in the figure. Figure 3 As shown, the average performance of different watermarking algorithms is compared through various attacks, generating graphs, reports, etc. Rank represents the attack's position in the performance vs. quality degradation 2D graph. Specifically, "Rank" is determined based on the degree of image quality degradation caused by the attack at a specific performance threshold (e.g., TPR@0.1%FPR=0.95). The higher the attack's ranking (the smaller the value), the greater its impact on watermark detection performance while maintaining image quality; in other words, the stronger the attack. By ranking the attacks, a quantified robustness evaluation of different watermarking methods under various attacks is obtained.
[0108] In summary, the embodiments of this invention conduct a detailed analysis of the impact of single and combined attacks on the watermarking method during the evaluation process. This includes launching attacks on the watermarked image at different intensities and subsequently evaluating the impact of these attacks on watermark detection performance (TPR@0.1%FPR) and a series of image quality metrics (such as PSNR). By plotting performance versus quality graphs and normalizing and aggregating quality metrics, the framework can generate a unified 2D graph, produce reports, and clearly demonstrate the overall performance and quality relationship of the evaluation.
[0109] This invention provides a standardized framework / method for evaluating the robustness of image watermarks. Through integrated detection and recognition tasks and various stress tests, it reveals potential vulnerabilities in modern watermarking algorithms. This approach not only guides the improvement of existing watermarking technologies but also provides tools for developing more robust watermarking techniques, helping to protect content sources and ownership while resisting malicious watermark removal.
[0110] This invention also provides an image watermark robustness evaluation device, as described in the following embodiments. Since the principle behind this device is similar to that of the image watermark robustness evaluation method, its implementation can be referenced from the implementation of the image watermark robustness evaluation method; repeated details will not be elaborated further.
[0111] Figure 4This is a schematic diagram of the image watermark robustness evaluation device in an embodiment of the present invention, as shown below. Figure 4 As shown, the device 400 includes:
[0112] Data acquisition module 401 is used to acquire the original image dataset;
[0113] The watermark printing module 402 is used to print watermarks on a portion of the original images in the original image dataset using the watermarking algorithm to be evaluated, thereby forming a watermarked image dataset.
[0114] The attack module 403 is used to attack all images in the watermarked image dataset to obtain the attacked image dataset; the attack includes one or any combination of distortion attack, regeneration attack, and adversarial attack.
[0115] The watermark detection module 404 is used to perform watermark recognition on all images in the image dataset after the attack and obtain watermark recognition results. The watermark recognition results include the true positive rate and the false positive rate. The true positive rate represents the proportion of watermarks correctly identified in all actual images containing watermarks, and the false positive rate represents the proportion of watermarks incorrectly identified in all actual images without watermarks.
[0116] The quality measurement module 405 is used to calculate the image quality measurement values before and after the attack for each printed watermark image using a preset image quality measurement algorithm, as the result of image quality degradation; the preset image quality measurement algorithm is used to calculate and standardize various image quality measurement elements of the image.
[0117] The result output module 406 is used to determine the image watermark robustness evaluation result based on the watermark recognition result and the image quality degradation result; the image watermark robustness evaluation result includes an image watermark robustness evaluation report, which includes data quantifying the performance of the watermarking algorithm to be evaluated.
[0118] In one embodiment, the data acquisition module 401 is specifically used for:
[0119] Acquire multiple datasets; the datasets include a set of real-world photos, a set of images generated by existing diffusion models, and a set of images generated by existing text-based image models.
[0120] The images in multiple datasets are ranked by removing irrelevant columns, filtering prompt strings, and using an aesthetic scoring algorithm. Images ranked above the preset ranking are selected to form the original image dataset.
[0121] In one embodiment, the distortion attack includes one or any combination of geometric distortion, photometric distortion, and degradation distortion; the regeneration attack includes image regeneration using a deep learning model; and the adversarial attack includes embedding attacks and proxy detector attacks.
[0122] In one embodiment, the watermark recognition result further includes a specified message embedded in the original image;
[0123] The watermark detection module 404 is specifically used for:
[0124] A preset watermark detection method is used to identify watermarks in all images in the attacked image dataset to obtain watermark identification results; the preset watermark detection method includes one of template matching, edge detection, and deep learning algorithms.
[0125] In one embodiment, the image quality metric includes one or any combination of peak signal-to-noise ratio, structural similarity index, normalized mutual information, Fraser initial distance (FID), CLIP-based FID, learned perceptual image patch similarity (LPIPS), aesthetic score, and artifact score.
[0126] Figure 5 This is a schematic diagram of a computer device in an embodiment of the present invention, such as... Figure 5 As shown, this embodiment of the invention also provides a computer device 500, including a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501. When the processor 501 executes the computer program 503, it implements the above-mentioned image watermark robustness evaluation method.
[0127] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described image watermark robustness evaluation method.
[0128] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described image watermark robustness evaluation method.
[0129] In this embodiment of the invention, an original image dataset is obtained; a watermark is printed on a portion of the original images in the original image dataset using the watermarking algorithm to be evaluated, forming a watermarked image dataset; all images in the watermarked image dataset are attacked to obtain an attacked image dataset; the attack includes one or any combination of distortion attack, regeneration attack, and adversarial attack; watermark recognition is performed on all images in the attacked image dataset to obtain watermark recognition results; the watermark recognition results include true positive rate and false positive rate; the true positive rate represents the proportion of correctly identified watermarks in all actual images containing watermarks, and the false positive rate represents the proportion of incorrectly identified watermarks in all actual images without watermarks; a preset image quality measurement algorithm is used to calculate the image quality measurement values before and after the attack for each image with printed watermarks, as the image quality degradation result; the preset image quality measurement algorithm is used to calculate and standardize various image quality measurement elements for the image; based on the watermark recognition results and the image quality degradation results, the image watermark robustness evaluation results are determined; the image watermark robustness evaluation results include an image watermark robustness evaluation report, which includes data quantifying the performance of the watermarking algorithm to be evaluated. This invention provides a complete, systematic, and mature method for evaluating watermarking algorithms. By comprehensively testing detection and recognition tasks as well as various stress tests, the robustness of image watermarking algorithms is evaluated. This invention considers two key dimensions: watermark detection performance and image quality degradation, and proposes an evaluation report on watermark detection performance and image quality, providing a new perspective for a comprehensive evaluation of the security and practicality of watermarking algorithms.
[0130] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0131] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0132] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0134] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for evaluating the robustness of image watermarking, characterized in that, include: Obtain the original image dataset; The watermarking algorithm to be evaluated is used to print watermarks on a portion of the original images in the original image dataset to form a watermarked image dataset. Attack all images in the watermarked image dataset to obtain the attacked image dataset; the attack includes one or any combination of distortion attack, regeneration attack, and adversarial attack; Watermark recognition is performed on all images in the image dataset after the attack, and the watermark recognition results are obtained. The watermark recognition results include the true positive rate and the false positive rate. The true positive rate represents the proportion of watermarks correctly identified in all actual images containing watermarks, and the false positive rate represents the proportion of watermarks incorrectly identified in all actual images without watermarks. Using a preset image quality measurement algorithm, calculate the image quality measurement values before and after the attack for each image with printed watermark, as the result of image quality degradation; The preset image quality measurement algorithm is used to calculate and standardize various image quality measurement elements; Based on the watermark recognition results and image quality degradation results, the robustness evaluation results of the image watermark are determined; The image watermarking robustness evaluation results include an image watermarking robustness evaluation report, which includes data quantifying the performance of the watermarking algorithm being evaluated.
2. The method as described in claim 1, characterized in that, Obtain the original image dataset, including: Acquire multiple datasets; the datasets include a set of real-world photos, a set of images generated by existing diffusion models, and a set of images generated by existing text-based image models. The images in multiple datasets are ranked by removing irrelevant columns, filtering prompt strings, and using an aesthetic scoring algorithm. Images ranked above the preset ranking are selected to form the original image dataset.
3. The method as described in claim 1, characterized in that, The distortion attack includes one or any combination of geometric distortion, photometric distortion, and degradation distortion; the regeneration attack includes image regeneration using a deep learning model; the adversarial attack includes embedding attack and proxy detector attack.
4. The method as described in claim 1, characterized in that, The watermark recognition result also includes a specified message embedded in the original image; Watermark recognition was performed on all images in the attacked image dataset, and the watermark recognition results were obtained, including: The watermark detection method is used to identify watermarks in all images in the image dataset after the attack, and the watermark identification results are obtained. The preset watermark detection method includes one of the following: template matching method, edge detection method, and deep learning algorithm.
5. The method as described in claim 1, characterized in that, The image quality metrics include one or any combination of peak signal-to-noise ratio, structural similarity index, normalized mutual information, Fraser initial distance (FID), CLIP-based FID, learned perceptual image patch similarity (LPIPS), aesthetic score, and artifact score.
6. An image watermark robustness evaluation device, characterized in that, include: The data acquisition module is used to acquire the raw image dataset; The watermark printing module is used to print watermarks on a portion of the original images in the original image dataset using the watermarking algorithm to be evaluated, thus forming a watermarked image dataset. The attack module is used to attack all images in the watermarked image dataset to obtain the attacked image dataset; the attack includes one or any combination of distortion attack, regeneration attack, and adversarial attack. The watermark detection module is used to identify watermarks in all images in the image dataset after the attack, and obtain watermark identification results. The watermark identification results include the true positive rate and the false positive rate. The true positive rate represents the proportion of watermarks correctly identified in all actual images containing watermarks, and the false positive rate represents the proportion of watermarks incorrectly identified in all actual images without watermarks. The quality measurement module is used to calculate the image quality measurement values before and after the attack for each image with a printed watermark using a preset image quality measurement algorithm, as the result of image quality degradation. The preset image quality measurement algorithm is used to calculate and standardize various image quality measurement elements; The result output module is used to determine the image watermark robustness evaluation result based on the watermark recognition result and the image quality degradation result. The image watermark robustness evaluation result includes an image watermark robustness evaluation report, which includes data that quantifies the performance of the watermarking algorithm to be evaluated.
7. The apparatus as claimed in claim 6, characterized in that, The data acquisition module is specifically used for: Acquire multiple datasets; the datasets include a set of real-world photos, a set of images generated by existing diffusion models, and a set of images generated by existing text-based image models. The images in multiple datasets are ranked by removing irrelevant columns, filtering prompt strings, and using an aesthetic scoring algorithm. Images ranked above the preset ranking are selected to form the original image dataset.
8. The apparatus as claimed in claim 6, characterized in that, The distortion attack includes one or any combination of geometric distortion, photometric distortion, and degradation distortion; the regeneration attack includes image regeneration using a deep learning model; the adversarial attack includes embedding attack and proxy detector attack.
9. The apparatus as claimed in claim 6, characterized in that, The watermark recognition result also includes a specified message embedded in the original image; The watermark detection module is specifically used for: A preset watermark detection method is used to identify watermarks in all images in the attacked image dataset to obtain watermark identification results; the preset watermark detection method includes one of template matching, edge detection, and deep learning algorithms.
10. The apparatus as claimed in claim 6, characterized in that, The image quality metrics include one or any combination of peak signal-to-noise ratio, structural similarity index, normalized mutual information, Fraser initial distance (FID), CLIP-based FID, learned perceptual image patch similarity (LPIPS), aesthetic score, and artifact score.
11. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 5.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 5.
13. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 5.