A diffusion model-oriented digital watermark embedding and detection method and system

By embedding multi-bit watermark information into the initial latent variables of the diffusion model, and utilizing frequency domain coding and Gaussian distribution preservation techniques, the robustness and computational overhead issues of generated image detection and tracing in the diffusion model are solved, achieving non-destructive detection and multi-user tracing.

CN122175757APending Publication Date: 2026-06-09INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610265720.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-05
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing generated image detection and source tracing technologies suffer from problems such as insufficient robustness, high computational overhead, high model deployment costs, and susceptibility to quality issues in diffusion models. Furthermore, there is a lack of a unified solution for generated image detection and multi-user source tracing.

Method used

By embedding multi-bit watermark information into the initial latent variables of the diffusion model, using frequency domain coding and Gaussian distribution preservation techniques, the watermark is embedded in the generation stage, and the watermark information is extracted through the inverse diffusion process, thus achieving detection and source tracing without modifying the model structure and training.

Benefits of technology

Without compromising the quality of the generated images, reliable detection and fine-grained source tracing of the generated images are achieved, exhibiting good robustness and multi-user source tracing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122175757A_ABST
    Figure CN122175757A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of information security and digital forensics technology, and relates to a method and system for digital watermark embedding and detection based on diffusion models. This method does not require modification of the network structure and model parameters of the diffusion model, nor does it require additional training of any detection or discrimination model. Instead, it embeds multi-bit watermark information intrinsically into the initial latent noise of the diffusion model, enabling the generated image to naturally carry detectable, extractable, and traceable identification information during the generation stage. Diffusion models typically assume that the initial latent variables follow a standard normal distribution when generating images. This invention, while maintaining the consistency of this statistical distribution, encodes the binary information to be traced as a weak frequency domain structural perturbation in the latent variables, and gradually integrates this perturbation with the image semantics during diffusion sampling and multi-step denoising. This achieves effective detection and tracing of the source of the generated image without reducing the visual quality and diversity of the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information security and digital forensics technology, specifically involving the security of AI-generated content, source tracing of images generated by diffusion models, digital watermarking technology, and copyright protection technology for generative models. It is a digital watermark embedding and detection method and system for diffusion models. Background Technology

[0002] With the development of generative artificial intelligence technology, generative models are increasingly widely used in image content production. Among them, diffusion models, especially latent diffusion models (LDM), have become the mainstream high-quality image generation solutions. However, generated images raise a series of new security and governance issues regarding copyright ownership, content authenticity verification, and liability determination, thus spurring research into technologies related to generated image detection and source tracing. Existing technical solutions for generated image identification and source tracing can be mainly summarized into the following categories: 1) Content Feature Analysis-Based Detection or Fingerprinting Methods: This type of technology analyzes differences in pixel distribution, spectral features, statistical properties, or high-order textures of images to extract discriminative features that can distinguish between "real images" and "generated images" or different generation models. Machine learning or deep learning classifiers are then used for detection or source tracing. For example, a so-called "model fingerprint" can be constructed by analyzing power spectral density, discrete cosine transform (DCT) coefficient distribution, and autocorrelation characteristics. These methods typically do not rely on active embedding mechanisms during the generation stage and are typical passive detection or post-analysis schemes.

[0003] 2) Post-processing watermarking methods based on deep learning end-to-end architecture: An improvement on traditional digital watermarking in AIGC scenarios, typically employing an encoder-decoder structure. This embeds the watermark information as an additional signal into the generated image and trains a corresponding detection network to extract the watermark information from the image. While this method improves invisibility, the watermark embedding process is independent of the image generation process, essentially remaining a post-processing mechanism.

[0004] 3) Intrinsic watermarking methods based on model fine-tuning (white-box methods): Some existing techniques fine-tune key modules of the diffusion model (such as the VAE decoder or U-Net network) to automatically embed specific watermark or signature information while generating images. This type of method is a typical white-box watermarking scheme, which can achieve watermark injection during the generation stage and has good invisibility.

[0005] 4) Watermarking source tracing methods based on initial noise or latent variable modification: Taking advantage of the reversible sampling characteristics of diffusion models, some techniques utilize mechanisms such as DDIM inversion to embed watermark patterns or specific statistical structures into the initial noise or latent variables of the diffusion model, and then restore the corresponding information through the inversion process during the detection phase. This type of method avoids model fine-tuning to some extent and reduces deployment costs.

[0006] Existing technical solutions mainly include four approaches: detection or fingerprint tracing methods based on content feature analysis, post-processing watermarking methods based on deep learning end-to-end structures, endogenous watermarking methods based on model fine-tuning, and watermark tracing methods based on initial noise or latent variable modification. Each of these approaches has some shortcomings, as detailed below: 1) Content feature analysis-based detection or fingerprint tracing methods: lack robustness, detection performance is highly dependent on the distribution of training data, and once the generated model is updated or the image undergoes post-processing operations such as compression, scaling, or cropping, the above features often degrade significantly, affecting the accuracy of detection and tracing.

[0007] 2) Post-processing watermarking methods based on deep learning end-to-end architecture: In the scenario of image generation by diffusion model, this type of method usually requires additional training of watermark encoder and decoder networks, which increases the computational overhead. Furthermore, the coupling degree between watermark information and image semantics is limited, and it is easily destroyed when faced with regeneration, editing or strong noise attacks.

[0008] 3) Intrinsic watermarking methods based on model fine-tuning (white-box methods): These methods generally rely on additional training or fine-tuning processes. When the watermark information needs to be updated, the model often needs to be retrained. At the same time, the modification of model parameters may have an adverse effect on the quality, stability or compatibility of the generated image, which is not conducive to the rapid deployment and large-scale application of the model.

[0009] 4) Watermarking source tracing methods based on initial noise or latent variable modification: Existing schemes mostly adopt fixed structure frequency domain mode or quantization sampling mechanism, which can easily destroy the original Gaussian distribution characteristics of the initial noise, thereby affecting the quality or diversity of the generated image; at the same time, some schemes only support zero-bit watermarking, which is difficult to meet the needs of multi-user or refined source tracing.

[0010] Furthermore, existing technologies are mostly designed for a single objective of detection or source tracing, lacking a unified solution that can simultaneously support generated image detection and multi-user source tracing within the same technical framework. This invention provides a digital watermarking embedding and detection method and system for diffusion models. Without modifying the model structure or requiring additional training, it embeds multi-bit source tracing information in a plug-and-play manner during the generation stage and reliably extracts this information during the detection stage. This achieves both generated image detection and fine-grained source tracing while ensuring that the quality of the generated image remains largely unaffected, and exhibits good robustness under various image perturbation conditions. Summary of the Invention

[0011] To address the aforementioned problems, this invention provides a digital watermark embedding and detection method and system oriented towards a diffusion model.

[0012] The technical solution adopted in this invention is as follows: A digital watermarking embedding method for a diffusion model includes the following steps: Construct a watermark matrix using the watermark information to be embedded; The watermark matrix is ​​encrypted to obtain the encrypted watermark matrix; The frequency domain coefficients are obtained by performing a discrete cosine transform on the encrypted watermark matrix. Add Gaussian noise perturbation to the frequency domain coefficients to obtain the perturbed frequency domain coefficients; The frequency domain coefficients after perturbation are normalized to obtain the initial latent variables; The initial latent variables are input into the diffusion model, and the diffusion model is used to generate an image with an embedded watermark.

[0013] Furthermore, the step of constructing a watermark matrix using the watermark information to be embedded includes: obtaining the binary bit sequence to be embedded as watermark information, and constructing a watermark matrix with the same dimension as the initial latent variables in the diffusion model using a repeated spreading strategy.

[0014] Furthermore, the encryption of the watermark matrix includes: performing bit-by-bit XOR encryption on the watermark using a pseudo-random bit matrix of the same dimension as the watermark matrix, and mapping the encrypted binary bits into a symmetric real-value form.

[0015] Furthermore, the method of generating an image with an embedded watermark using a diffusion model includes: using the initial latent variable as the initial noise input of the diffusion model, generating an image according to the denoising sampling process of the diffusion model without making any modifications to the network structure and parameters of the diffusion model, and gradually integrating the watermark information with the image semantics during the multi-time-step denoising process to form an endogenous watermark.

[0016] A digital watermark detection method based on a diffusion model includes the following steps: The image with embedded watermark generated using the above method is used as the image to be detected; Perform an inverse diffusion operation on the image to be detected to obtain the initial latent variables; The initial latent variables are subjected to discrete cosine inverse transform to map the real-valued results into binary bits, and the encrypted watermark matrix is ​​obtained through random smoothing and majority voting. Decrypt the encrypted watermark matrix and use the decrypted watermark matrix to obtain the watermark information. By comparing the watermark information with the registered watermark information, the detection and source determination of the image to be tested can be achieved.

[0017] Furthermore, the random smoothing and majority voting includes: injecting additional small-amplitude Gaussian noise into the initial latent variable multiple times, and repeatedly performing the inverse discrete cosine transform and bit decision to obtain multiple sets of candidate bit results, and performing bit-by-bit majority voting on the multiple sets of candidate bit results.

[0018] A digital watermark embedding device for a diffusion model includes: The watermark matrix construction module is used to construct a watermark matrix using the watermark information to be embedded. The watermark matrix encryption module is used to encrypt the watermark matrix to obtain the encrypted watermark matrix. The frequency domain embedding module is used to perform discrete cosine transform on the encrypted watermark matrix to obtain frequency domain coefficients. The perturbation modulation module is used to add Gaussian noise perturbation to the frequency domain coefficients to obtain the perturbed frequency domain coefficients; The normalization module is used to normalize the frequency domain coefficients after perturbation to obtain the initial latent variables; The watermark image generation module is used to input the initial latent variables into the diffusion model and use the diffusion model to generate an image with an embedded watermark.

[0019] A digital watermark detection device based on a diffusion model, comprising: The watermark image acquisition module is used to acquire the image to be detected; The latent variable inversion module is used to perform inverse diffusion on the image to be detected to obtain the initial latent variables. The frequency domain inverse transform and bit decision module is used to perform discrete cosine inverse transform on the initial latent variables and map the real-valued results to binary bits. The random smoothing and majority voting module is used to obtain the encrypted watermark matrix by random smoothing and majority voting on the output results of the frequency domain inverse transform and bit decision module. The decryption module is used to decrypt the encrypted watermark matrix; The watermark information acquisition module is used to obtain watermark information using the decrypted watermark matrix. The detection and tracing module is used to detect and trace the source of the image to be detected by comparing the watermark information with the registered watermark information.

[0020] The beneficial effects of this invention are as follows: This invention proposes a digital watermark embedding and detection method for diffusion models. This method requires no modification to the network structure and model parameters of the diffusion model, nor does it require additional training of any detection or discrimination model. Instead, it intrinsically embeds multi-bit watermark information into the initial latent noise of the diffusion model, enabling the generated image to naturally carry detectable, extractable, and traceable identification information during the generation stage. Diffusion models typically assume that the initial latent variables follow a standard normal distribution when generating images. This invention, while maintaining the consistency of this statistical distribution, encodes the binary information to be traced as a weak frequency domain structural perturbation in the latent variables. This perturbation is then gradually integrated with the image semantics during diffusion sampling and multi-step denoising, thereby achieving effective detection and source tracing of the generated image's origin without reducing the visual quality and diversity of the generated image. Attached Figure Description

[0021] Figure 1 This is an architecture diagram of a digital watermark embedding and detection method based on a diffusion model according to the present invention.

[0022] Figure 2 This is a flowchart of the steps of a digital watermarking embedding method for a diffusion model according to the present invention.

[0023] Figure 3 This is a schematic diagram of the watermark embedding process according to an embodiment of the present invention.

[0024] Figure 4 This is a flowchart of the steps of a digital watermark detection method based on a diffusion model according to the present invention.

[0025] Figure 5 This is a schematic diagram of the watermark extraction process according to an embodiment of the invention.

[0026] Figure 6 This is the watermark detection result of the present invention.

[0027] Figure 7 This is the result of watermark tracing in this invention. Detailed Implementation

[0028] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0029] The overall technical solution of this invention includes the following three core components: 1) frequency domain encoding of watermark information and construction of Gaussian distribution-preserving initial latent variables; 2) watermarked image generation based on latent diffusion model; 3) watermark detection and source extraction based on diffusion model inversion and random smoothing mechanism.

[0030] Figure 1 This is an architecture diagram of a digital watermark embedding method for a diffusion model according to the present invention, including a watermark embedding process and a watermark extraction process. Wherein, This indicates the watermark information to be embedded. This is the initial latent vector of the constructed hidden watermark. U-Net is used for iterative denoising of the diffusion model. D is the latent space representation of the final generated image; E is the decoder module of the diffusion model, used to decode the latent representation output by the diffusion model into the final generated image; and E is the encoder module of the diffusion model, used to encode the image to be tested into a latent space representation. It is the latent space representation of the image to be tested. The initial noise of the diffusion model is obtained through DDIM inversion. This indicates the extracted watermark information.

[0031] This invention provides a digital watermarking embedding method oriented towards a diffusion model, such as... Figure 2 As shown, it includes the following steps: Construct a watermark matrix using the watermark information to be embedded; The watermark matrix is ​​encrypted to obtain the encrypted watermark matrix; The frequency domain coefficients are obtained by performing a discrete cosine transform on the encrypted watermark matrix. Add Gaussian noise perturbation to the frequency domain coefficients to obtain the perturbed frequency domain coefficients; The frequency domain coefficients after perturbation are normalized to obtain the initial latent variables; The initial latent variables are input into the diffusion model, and the diffusion model is used to generate an image with an embedded watermark.

[0032] In one embodiment, constructing a watermark matrix using the watermark information to be embedded includes: obtaining the binary bit sequence to be embedded as watermark information, and constructing a watermark matrix with the same dimension as the initial latent variables in the diffusion model using a repeated spreading strategy.

[0033] In one embodiment, encrypting the watermark matrix includes: performing bit-by-bit XOR encryption on the watermark using a pseudo-random bit matrix of the same dimension as the watermark matrix, and mapping the encrypted binary bits to a symmetric real-value form.

[0034] In one embodiment, generating an image with an embedded watermark using a diffusion model includes: using an initial latent variable as the initial noise input to the diffusion model, generating an image according to the denoising sampling process of the diffusion model without modifying the network structure and parameters of the diffusion model, and gradually integrating the watermark information with the image semantics during the multi-time-step denoising process to form an endogenous watermark.

[0035] In one embodiment, a diffusion-oriented digital watermarking embedding method of the present invention is as follows: Figure 3 As shown, it includes the following steps: 1) Watermark Information Generation and Preprocessing: Suppose the source information to be embedded is a binary bit sequence of length k. In the latent diffusion model, the initial latent variables Typically, these are high-dimensional random tensors related to the resolution of the generated image; for example, in the Stable Diffusion model, their dimensions are... To embed watermark bits of finite length into a high-dimensional latent space, this invention employs a replication strategy to construct and... Dimensionally consistent watermark matrix .

[0036] 2) Watermark Encryption and Symbol Mapping: To enhance watermark security and reduce the statistical correlation between plaintext bits, watermark encryption and symbol mapping are used. Same-dimensional pseudo-random bit matrix The watermark is encrypted using a bitwise XOR operation, i.e. ,in This represents the bit value of the i-th element in the encrypted watermark matrix. Represents the watermark matrix The watermark bits at the corresponding positions, Represents a pseudo-random bit matrix The i-th random bit. Then, for ease of frequency domain processing, the encrypted binary bits are mapped to a symmetric real-valued form. After the above processing, a watermark matrix that can only be recovered by the key holder is obtained. .in express The real-value watermark element corresponding to the i-th position takes the value of -1 or 1.

[0037] 3) Frequency Domain Embedding and Perturbation Modulation: To improve the watermark's concealment and avoid regular spatial domain structures, this invention transfers the watermark embedding process to the frequency domain. Specifically, for the watermark matrix... Applying the Discrete Cosine Transform (DCT), we obtain the frequency domain coefficient representation: In the frequency domain, to avoid statistical anomalies caused by overly regular embedding patterns, Add Gaussian noise with a mean of 0 to each coefficient The frequency domain coefficients after perturbation are obtained. .

[0038] 4) Standard Gaussian Normalization: Diffusion models typically assume that the initial noise follows a standard normal distribution. Therefore, for the perturbed frequency domain coefficients... After normalization, the final initial latent variables are obtained: ,in and They represent The mean and standard deviation. After normalization, It is consistent with standard Gaussian noise in a statistical sense, thus ensuring that the watermark embedding does not violate the original generation assumptions of the diffusion model.

[0039] 5) Watermarked Image Generation: Constructed Initial Latent Variables The noise was directly used as the initial input to the diffusion model. Without modifying the model's network structure or parameters, the image was generated following the original denoising and sampling process of the diffusion model. Due to the embedded watermark... The statistical distribution is consistent with the standard normal distribution. The sampling trajectory and generation effect of the diffusion model will not change significantly. The watermark information gradually merges with the image semantics during the multi-time step denoising process to form an endogenous watermark.

[0040] This invention provides a digital watermark detection method based on a diffusion model, such as... Figure 4 As shown, it includes the following steps: The image with embedded watermark generated using the above method is used as the image to be detected; Perform an inverse diffusion operation on the image to be detected to obtain the initial latent variables; The initial latent variables are subjected to discrete cosine inverse transform to map the real-valued results into binary bits, and the encrypted watermark matrix is ​​obtained through random smoothing and majority voting. Decrypt the encrypted watermark matrix and use the decrypted watermark matrix to obtain the watermark information. By comparing the watermark information with the registered watermark information, the detection and source determination of the image to be tested can be achieved.

[0041] In one embodiment, the random smoothing and majority voting includes: injecting additional small-amplitude Gaussian noise into the initial latent variable multiple times, and repeatedly performing the inverse discrete cosine transform and bit decision to obtain multiple sets of candidate bit results, and performing bit-by-bit majority voting on the multiple sets of candidate bit results.

[0042] In one embodiment, a diffusion-based digital watermark detection method of the present invention is as follows: Figure 5 As shown, it includes the following steps: 1) Latent variable inversion: For the image to be detected, perform inverse diffusion using a diffusion model inversion algorithm (such as the DDIM inversion process) to obtain the corresponding initial latent variable estimates. .

[0043] 2) Inverse frequency domain transform and bit decision: For the initial latent variables obtained from reconstruction Apply Inverse Discrete Cosine Transform (IDCT) to recover the frequency domain watermark matrix. Subsequently, a sign decision function is used to map the real-valued result into binary bits: .in It is a symbolic decision function.

[0044] 3) Random Smoothing and Majority Voting: To enhance the robustness of watermark extraction, a random smoothing strategy is introduced into the latent variable inversion results. Specifically, in Multiple injections of additional small-amplitude Gaussian noise And repeat the frequency domain inverse transform and bit decision process in step 2): After this process is repeated N times, a bit-by-bit majority vote is performed on the resulting N candidate bits to obtain a final stable decision, which yields the encrypted watermark matrix. .in This indicates a majority vote, which means counting the number of times a bit value appears in N candidate results for the same bit position. When a bit value appears more than N / 2 times, that bit value is determined as the final decision for that position.

[0045] 4) Decryption and Reverse Padding: Finally, using the same key as in the embedding stage... The encrypted watermark matrix Decryption is performed to obtain the decrypted watermark matrix. And through reverse repetition operations ( Restore final watermark information: By By comparing the generated image with the registered watermark bit sequence, the detection and source determination of the generated image can be completed.

[0046] The main ideas and principles of this invention: By leveraging the high dependence of the latent diffusion model generation process on the initial latent variable distribution, and without violating the assumption that the latent variables follow a standard Gaussian distribution, traceable multi-bit information is endogenously embedded into the generation starting point of the diffusion model, thereby achieving generated image detection and tracing without model modification, additional training, or post-processing.

[0047] Unlike traditional post-generation watermarking schemes that generate watermarks first and then embed them, this invention does not introduce additional modifications to the generated image space. Instead, it directly affects the initial latent noise space of the diffusion model. Since the generation process of the diffusion model is essentially a process of gradually denoising the initial noise and converging to the image, any structural perturbations in the initial latent variables will gradually merge into the image semantic layer during the multi-time-step diffusion process. Therefore, through reasonable, weak, and statistically consistent perturbations, the watermark information can be "naturally diffused during the generation process," forming an endogenous watermark.

[0048] This invention further recognizes that a fundamental assumption of the diffusion model regarding the initial noise is that it follows a standard normal distribution. If the watermark embedding disrupts this statistical distribution, it may cause a decrease in the quality of the generated image or abnormal patterns. Therefore, this invention does not directly modify the latent variables explicitly, but rather maps the watermark bits to a high-dimensional frequency domain structural perturbation through frequency domain coding and statistical normalization, and in the final stage forces the overall latent variables to be renormalized to a standard Gaussian distribution, so that the diffusion model is "unaware of the existence of the watermark" in a statistical sense.

[0049] Furthermore, since the backsampling process of the diffusion model is approximately reversible (such as DDIM inversion), this invention can recover the initial latent variable estimates from the generated image through the inverse diffusion process during the detection stage, and recover the watermark information through inverse frequency domain transformation. Considering that the inversion process and image post-processing operations may introduce errors, this invention introduces random smoothing and majority voting mechanisms to elevate the single watermark determination to a statistically stable decision, thereby significantly enhancing the robustness of the system.

[0050] In summary, this invention achieves a truly plug-and-play digital watermarking embedding and detection scheme oriented towards diffusion models through the overall approach of "latent space endogenous embedding + frequency domain modulation + Gaussian consistency preservation + inversion extraction + random smooth decision".

[0051] In one embodiment, the present invention implements a method for digital watermark embedding, watermark extraction, and source tracing oriented towards a diffusion model through the following steps: 1) Implementation Environment and Generative Model: Stable Diffusion v1.5 was selected as an example of the latent diffusion model, with the latent spatial noise tensor dimension being... The diffusion model parameters and network structure remain unchanged; the watermark embedding method of this invention is only introduced during the sampling stage.

[0052] 2) Watermark Information Setting: The watermark information used for traceability is set to a 64-bit binary sequence, which can uniquely correspond to a generating user or generating instance. The watermark bits are mapped to a matrix consistent with the dimension of the latent variables through repeated tiling.

[0053] 3) Latent Space Watermark Embedding: The spread-out watermark matrix is ​​pseudo-randomly encrypted and mapped to a {-1,1} numerical matrix. Subsequently, a Discrete Cosine Transform (DCT) is performed on this matrix, with a weak Gaussian perturbation added in the frequency domain. Finally, the resulting frequency domain coefficients are globally normalized to ensure the final latent variables conform to a standard normal distribution. This normalized latent variable serves as the initial noise input to the diffusion model for image generation.

[0054] 4) Watermarked Image Generation: The diffusion model generates images according to its established denoising sampling process. Experimental results show that, compared with images without embedded watermarks, the generated images show no significant difference in visual quality, diversity, and semantic consistency, and the watermark information is visually imperceptible.

[0055] 5) Watermark Extraction and Source Tracing: During the detection phase, the generated image undergoes a DDIM inversion process to recover its initial latent variable estimates. An inverse discrete cosine transform (IDCT) is applied to these latent variables, and the watermark bits are extracted using a sign decision function.

[0056] 6) Robustness case study: After applying common image perturbations such as JPEG compression, mild Gaussian noise, scaling, and cropping to the generated image, the watermark can still be extracted stably. The majority voting mechanism significantly reduces the impact of single extraction errors on the final result.

[0057] Key points of this invention: 1) Digital watermarking embedding and detection scheme for diffusion models: This scheme generates image source tracing by constructing initial latent variables without modifying the diffusion model network structure or retraining or fine-tuning the model. This clearly distinguishes the present invention from white-box watermarking schemes that rely on model fine-tuning or additional training.

[0058] 2) Multi-bit endogenous watermarking embedding mechanism based on latent space: This mechanism embeds multi-bit watermark information into the initial latent variables of the latent diffusion model, allowing the watermark to naturally integrate with the semantics of the generated image during the diffusion sampling process. This significantly distinguishes it from zero-bit detection methods and post-watermarking methods.

[0059] 3) Embedding method combining frequency domain modulation and Gaussian distribution preservation: This method encodes the watermark information through frequency domain transformation and performs normalization on the latent variables after embedding, ensuring that the latent variables after watermarking still statistically conform to a standard normal distribution. This is a key technical point to avoid degradation in generation quality and an important innovation that distinguishes it from fixed-structure frequency domain watermarking.

[0060] 4) Watermark extraction method based on diffusion model inversion: This method utilizes the backsampling or inversion process of the diffusion model to recover the initial latent variables from the generated image and complete watermark detection or source extraction in the latent space.

[0061] 5) Robust decision mechanism based on random smoothing and statistical voting: This technical solution improves the stability and robustness of watermark extraction by using random noise perturbation, multiple sampling, and majority voting during the watermark extraction stage. This approach effectively defends against common image post-processing attacks and has clear theoretical support.

[0062] 6) Integrated detection and tracing technology framework: Under the same technology framework, it supports a unified solution for both generated image detection (whether it is generated by the target model) and multi-user / multi-instance tracing (specific source location).

[0063] Technical effects of the present invention: The hardware configuration for the experiment of this invention is shown in Table 1.

[0064] Table 1. Hardware Configuration Experimental Design: The Stable Diffusion Prompts (SDP), MS COCO-2017, and ImageNet datasets were used for evaluation. Stable Diffusion Prompts was primarily used for experiments on effectiveness, generated image quality, and robustness, while MS COCO-2017 and ImageNet were primarily used for experiments on generated image quality. The SDP dataset contains a large number of T2I task prompts, while MS COCO-2017 and ImageNet contain a large number of natural images.

[0065] The diffusion model used was Stable Diffusion v2.1 provided by the Hugging Face platform, with an accuracy of FP16. In most experiments, the guiding scale of the diffusion model was set to 7.5, the number of sampling steps was set to 50, the number of reverse diffusion steps was set to 50, and the resolution of the generated image was set to... The latent space dimension is For the watermarking method settings, the intensity of the frequency domain perturbation noise is set to 0.01, the intensity of the perturbation noise in random smoothing is set to 0.01, and the majority vote count is set to 501.

[0066] In the watermark effectiveness experiment, the true positive rate (TPR) and bit accuracy (ACC) corresponding to the false positive rate (FPR) were used as evaluation metrics to assess the performance of the watermarking method of this invention (hereinafter referred to as the FIRE watermarking method) in both detection and source tracing scenarios. In the generated image quality experiment, FID and CLIP-Score were used to evaluate the invisibility of the watermark information in the generated image, and LPIPS was used to evaluate the diversity of the generated image. In the watermark robustness experiment, the above metrics were used to comprehensively evaluate the performance of the FIRE watermarking method.

[0067] (a) Validity Experiment In the detection scenario, 10,000 watermarked images were generated using the SDP dataset, then different types of image-level perturbations were added, and finally the watermark information was extracted and the TPR was calculated. Figure 6 Reported At any given time, except for brightness adjustment, the TPR of the FIRE watermarking method can reach above 0.99 under other types of image perturbations. This means that the FIRE watermarking method can... 99% was successfully detected in the generated image. Regarding brightness adjustment, although... hour, The value is only 0.947, but this is still a very good result. In a source tracing scenario, given a generated image, the goal is to find... Did any of the users generate the image? If so, which user specifically? For accurate evaluation, assume there are... Each user generates 10 images, resulting in a dataset containing 10,000 images. Figure 7 Reported In terms of source tracing accuracy, except for brightness adjustment, the FIRE watermarking method achieves a source tracing accuracy of over 0.98 under other types of image perturbations. This means that the FIRE watermarking method can be used with a large number of users. In the context of tracing the source, the success rate reached 98%, which is a very ideal result. Regarding brightness adjustment, although... At that time, the traceability accuracy rate was only 0.936, but this was still an acceptable result. (II) Production Quality Experiment Evaluation was conducted in terms of watermark invisibility and the diversity of generated images. For watermark invisibility, the True Positive Rate (TPR) in detection scenarios and the Bit Acc in source tracing scenarios were calculated using the SDP dataset. To measure the performance bias that watermark embedding might cause, FID and CLIP-Score were calculated using the SDP and MSCOCO-2017 datasets. For the diversity of generated images, 10 different prompts were selected from the PromptHero website, and the LPIPS-Score was calculated. Unless otherwise specified, , The results in Table 2 show that the FIRE method has almost no impact on the generation performance of LDM, achieving non-destructive testing and traceability.

[0068] Table 2. Results of the quality generation experiment (III) Robustness Experiment Against Image Attacks To evaluate the robustness of the FIRE watermarking method against image-level perturbation attacks, 1000 images were generated using the SDP dataset, and several image-level perturbation attacks were applied. The results are shown in Table 3. For detection scenarios, the FIRE watermarking method exhibits strong robustness against various types of image-level perturbation attacks, with TPRs approaching 1.0, achieving best performance except for brightness adjustment attacks. For source tracing scenarios, the FIRE watermarking method achieves a bit accuracy of up to 99% against attacks other than Gaussian noise, salt-and-pepper noise, and brightness adjustment attacks, achieving best performance except for brightness adjustment attacks—a very desirable result.

[0069] Table 3. Robustness test results Another embodiment of the present invention provides a digital watermark embedding device for a diffusion model, comprising: The watermark matrix construction module is used to construct a watermark matrix using the watermark information to be embedded. The watermark matrix encryption module is used to encrypt the watermark matrix to obtain the encrypted watermark matrix. The frequency domain embedding module is used to perform discrete cosine transform on the encrypted watermark matrix to obtain frequency domain coefficients. The perturbation modulation module is used to add Gaussian noise perturbation to the frequency domain coefficients to obtain the perturbed frequency domain coefficients; The normalization module is used to normalize the frequency domain coefficients after perturbation to obtain the initial latent variables; The watermark image generation module is used to input the initial latent variables into the diffusion model and use the diffusion model to generate an image with an embedded watermark.

[0070] Another embodiment of the present invention provides a digital watermark detection device oriented towards a diffusion model, comprising: The watermark image acquisition module is used to acquire the image to be detected; The latent variable inversion module is used to perform inverse diffusion on the image to be detected to obtain the initial latent variables. The frequency domain inverse transform and bit decision module is used to perform discrete cosine inverse transform on the initial latent variables and map the real-valued results to binary bits. The random smoothing and majority voting module is used to obtain the encrypted watermark matrix by random smoothing and majority voting on the output results of the frequency domain inverse transform and bit decision module. The decryption module is used to decrypt the encrypted watermark matrix; The watermark information acquisition module is used to obtain watermark information using the decrypted watermark matrix. The detection and tracing module is used to detect and trace the source of the image to be detected by comparing the watermark information with the registered watermark information.

[0071] The above division of modules is merely illustrative. In practical applications, the functions described above can be assigned to different functional modules as needed to complete all or part of the functions described in the aforementioned method. The specific working process of each module can be found in the implementation process of the corresponding steps in the aforementioned method embodiments, and will not be repeated here.

[0072] Another embodiment of the present invention provides a computer device (computer, server, smartphone, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing steps of the method of the present invention.

[0073] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk) that stores a computer program, which, when executed by a computer, implements the steps of the method of the present invention.

[0074] Another embodiment of the present invention provides a computer program product, the computer program product including a computer program, which, when executed by a computer, implements the steps of the method of the present invention.

[0075] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and to implement it accordingly. Those skilled in the art will understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification; the scope of protection of the present invention is defined by the claims.

Claims

1. A digital watermark embedding method for a diffusion model, characterized in that, Includes the following steps: Construct a watermark matrix using the watermark information to be embedded; The watermark matrix is ​​encrypted to obtain the encrypted watermark matrix; The frequency domain coefficients are obtained by performing a discrete cosine transform on the encrypted watermark matrix. Add Gaussian noise perturbation to the frequency domain coefficients to obtain the perturbed frequency domain coefficients; The frequency domain coefficients after perturbation are normalized to obtain the initial latent variables; The initial latent variables are input into the diffusion model, and the diffusion model is used to generate an image with an embedded watermark.

2. The method according to claim 1, characterized in that, The process of constructing a watermark matrix using the watermark information to be embedded includes: obtaining the binary bit sequence to be embedded as watermark information, and constructing a watermark matrix with the same dimension as the initial latent variables in the diffusion model using a repeated spreading strategy.

3. The method according to claim 1, characterized in that, The encryption of the watermark matrix includes: performing bit-by-bit XOR encryption on the watermark using a pseudo-random bit matrix of the same dimension as the watermark matrix, and mapping the encrypted binary bits to a symmetric real-value form.

4. The method according to claim 1, characterized in that, The method of generating an image with an embedded watermark using a diffusion model includes: using an initial latent variable as the initial noise input of the diffusion model, generating an image according to the denoising sampling process of the diffusion model without modifying the network structure and parameters of the diffusion model, and gradually integrating the watermark information with the image semantics during the multi-time step denoising process to form an endogenous watermark.

5. A digital watermark detection method based on a diffusion model, characterized in that, Includes the following steps: The image with embedded watermark generated by the method described in any one of claims 1 to 4 is used as the image to be detected. Perform an inverse diffusion operation on the image to be detected to obtain the initial latent variables; The initial latent variables are subjected to discrete cosine inverse transform to map the real-valued results into binary bits, and the encrypted watermark matrix is ​​obtained through random smoothing and majority voting. Decrypt the encrypted watermark matrix and use the decrypted watermark matrix to obtain the watermark information. By comparing the watermark information with the registered watermark information, the detection and source determination of the image to be tested can be achieved.

6. The method according to claim 5, characterized in that, The random smoothing and majority voting process includes: injecting additional small-amplitude Gaussian noise into the initial latent variables multiple times, repeatedly performing the inverse discrete cosine transform and bit decision to obtain multiple sets of candidate bit results, and performing bit-by-bit majority voting on the multiple sets of candidate bit results.

7. A digital watermark embedding device for a diffusion model, characterized in that, include: The watermark matrix construction module is used to construct a watermark matrix using the watermark information to be embedded. The watermark matrix encryption module is used to encrypt the watermark matrix to obtain the encrypted watermark matrix. The frequency domain embedding module is used to perform discrete cosine transform on the encrypted watermark matrix to obtain frequency domain coefficients. The perturbation modulation module is used to add Gaussian noise perturbation to the frequency domain coefficients to obtain the perturbed frequency domain coefficients; The normalization module is used to normalize the frequency domain coefficients after perturbation to obtain the initial latent variables; The watermark image generation module is used to input the initial latent variables into the diffusion model and use the diffusion model to generate an image with an embedded watermark.

8. A digital watermark detection device for a diffusion model, characterized in that, include: The watermark image acquisition module is used to acquire the image to be detected; The latent variable inversion module is used to perform inverse diffusion on the image to be detected to obtain the initial latent variables. The frequency domain inverse transform and bit decision module is used to perform discrete cosine inverse transform on the initial latent variables and map the real-valued results to binary bits. The random smoothing and majority voting module is used to obtain the encrypted watermark matrix by random smoothing and majority voting on the output results of the frequency domain inverse transform and bit decision module. The decryption module is used to decrypt the encrypted watermark matrix; The watermark information acquisition module is used to obtain watermark information using the decrypted watermark matrix. The detection and tracing module is used to detect and trace the source of the image to be detected by comparing the watermark information with the registered watermark information.

9. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 6.