Image watermarking method for stable diffusion model

By embedding feature rearrangement and watermark information in the reverse diffusion process of the stable diffusion model, an image watermark generation and recognition model is constructed, which solves the problems of poor visual quality, low security and ambiguous responsibility attribution of the diffusion model image watermark method, and realizes high concealment and robust image traceability capabilities, which is suitable for multi-user environments.

CN120765440AActive Publication Date: 2025-10-10YUNNAN UNIV

Patent Information

Application Number
CN202510801404.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-10
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Existing diffusion model image watermarking methods have problems such as poor visual quality, low security, high deployment cost, poor versatility, and difficulty in achieving two-way traceability during the generation process. Especially in diffusion model-generated images, watermarks are easily tampered with and forged, and the responsibility is unclear.

Method used

In the reverse diffusion process of the stable diffusion model, latent features are selected. User watermark information is embedded in the latent features through feature rearrangement and watermark information embedding modules. An image watermark generation and recognition model is constructed, including watermark information embedding and extraction modules. A plug-in architecture is adopted without modifying the backbone parameters. Model training is performed in combination with training samples to achieve highly concealed and robust image traceability.

Benefits of technology

It achieves image traceability capabilities with high concealment and strong robustness without affecting image quality. It is suitable for multi-user environments, supports fine-grained tracing, can be quickly integrated into mainstream stable diffusion models, and has good engineering adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765440A_ABST
    Figure CN120765440A_ABST
Patent Text Reader

Abstract

The invention discloses an image watermarking method for a stable diffusion model, and the method comprises the steps: selecting one of all potential features in a reverse diffusion process of the stable diffusion model as a target potential feature, and constructing an image watermark generation recognition model which comprises a watermark information embedding module and a watermark extraction module; the watermark information embedding module is used for embedding user watermark information into the target potential features, and inputting the obtained watermark target potential features instead of the original target potential features into the remaining sub-models of the stable diffusion model to generate a watermark image; and the watermark information extraction module is used for extracting user watermark information from the watermark potential features obtained by carrying out DDIM inversion on the watermark image, and the user generating the watermark image can be identified according to the extracted user watermark information. According to the method, the user watermark information is embedded into the potential features of the stable diffusion model, so that the image traceability with high robustness and high concealment is realized on the premise of not damaging the quality of the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and more particularly, relates to an image watermarking method oriented to a stable diffusion model. Background Art

[0002] With the widespread adoption of AI-generated content (AIGC) technology, diffusion models have become a mainstream architecture in the field of image generation. These models generate high-fidelity synthetic images through a gradual denoising process of Gaussian noise. Typical examples include Stable Diffusion, DALL·E 3, and Parti. Due to their powerful generation capabilities, open source nature, and high adaptability to text prompts, diffusion models have been widely adopted in various scenarios, including artistic creation, advertising design, and product prototyping. However, the open nature of these models and the high degree of freedom in image content also bring numerous content security risks, including the theft of original works, misuse of platform content, the risk of synthetic image forgery, and unauthorized use of interfaces.

[0003] In recent years, researchers have attempted to embed invisible watermarks in images to achieve copyright protection and image traceability. Traditional methods are mostly based on frequency-domain embedding, such as DCT (discrete cosine transform), DWT (wavelet transform), and SVD (singular value decomposition). These algorithms have a certain degree of robustness in natural images, but they are not very effective in images generated by diffusion models. Sampling noise and domain offsets during the generation process often lead to watermark loss or extraction errors. Some deep learning methods attempt to embed watermarks after image generation or implement encoding and decoding in the image domain through end-to-end training. However, these methods often require structural modifications or joint training of the image generation model, resulting in poor versatility and high deployment costs.

[0004] The current AIGC platform faces the following challenges: First, the ownership of content is difficult to define. Once user-generated original images are published, they can be easily downloaded by third parties and edited (such as redrawing, cutting, upsampling, etc.), resulting in visually imperceptible pseudo-originality, infringing the rights of the original author; second, the platform itself faces the risk of misappropriating user images for commercial purposes without authorization, resulting in users lacking traceable basis for claims; third, forged images may be used to spread false content, mislead the public, and even cause legal disputes; fourth, the model API interface may be abused by black industries to bypass platform supervision to generate illegal images, creating the risk of difficulty in identifying the responsible party.

[0005] At present, watermarking methods for diffusion models can be roughly divided into two categories: training-free methods and training-dependent methods. The training-free method mainly generates initial noise or intermediate latent variables by modifying the diffusion, and realizes watermark embedding without modifying the model structure. It has certain advantages at the deployment level and does not require changing the model parameters. However, there are some common disadvantages in this type of method: First, there are often structural differences or perceptible artifacts between the image and the original image, which affects the visual quality; second, most methods do not introduce an identity encryption mechanism, and the watermark is easy to be extracted, forged or tampered with, and the security is low; third, the scalability is poor in a multi-user environment, making it difficult to implement a single Figure 1 The demand for fine-grained traceability of the target. Methods that require training require structural fine-tuning or end-to-end retraining of certain modules of the diffusion model so that it naturally carries the specified watermark information during the image generation stage. This type of method can embed the bit string that carries specific information into the image, but this type of method requires fine-tuning or retraining the structure of the diffusion model, which will have a certain impact on the performance of the model, and there are problems such as high deployment cost and poor versatility. In addition, the current mainstream watermarking methods generally lack a dual identity identification mechanism and encryption-level anti-tampering capabilities, and cannot simultaneously meet the two-way traceability requirements of "users can claim copyright and the platform can confirm the source." Most methods only embed one party's information (such as user ID), which can neither prevent the platform from illegally abusing the image nor distinguish whether it is generated by the platform's internal model, and the responsibility is unclear. Summary of the Invention

[0006] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide an image watermarking method for a stable diffusion model. By embedding user watermark information in the latent features of the stable diffusion model, the method can achieve high robustness and strong concealment of image traceability without destroying the quality of the generated image.

[0007] In order to achieve the above-mentioned object of the invention, the image watermarking method for the stable diffusion model of the present invention comprises the following steps:

[0008] S1: The image generation platform equipped with the stable diffusion model selects one target latent feature z from all potential features in the reverse diffusion process of the stable diffusion model. t , t represents the time step corresponding to the target potential feature, t∈[1,T-1];

[0009] The image generation platform assigns each user a unique user ID and a feature rearrangement strategy for the target potential features. The user generates his watermark information w∈{0,1} n And upload it to the image generation platform, where n represents the dimension of the watermark information;

[0010] S2: Build an image watermark generation and recognition model and deploy it to the image generation platform. The image watermark generation and recognition model includes a watermark information embedding module and a watermark extraction module, where:

[0011] The watermark information embedding module is used to obtain the target potential feature z from the reverse diffusion process of the stable diffusion model t , embed the current user's watermark information w into the target latent feature z t 中,将得到的水印潜在特征 代替目标潜在特征z t The residual sub-model of the input stable diffusion model generates a watermark image; the watermark information embedding module includes a latent feature rearrangement module, a watermark information reconstruction module, a watermark feature embedding module and a latent feature restoration module, wherein:

[0012] The latent feature rearrangement module is used to rearrange the target latent feature z using a preset feature rearrangement strategy. t Perform structural rearrangement to obtain rearranged potential features A and send them to the watermark feature embedding module;

[0013] The watermark information reconstruction module is used to reconstruct the watermark information according to the target potential feature z t The user watermark information w is reconstructed with the size of , the watermark feature W is obtained and sent to the watermark feature embedding module;

[0014] The watermark feature embedding module is used to embed the watermark feature W into the rearranged latent feature A, obtain the latent feature B and send it to the latent feature restoration module;

[0015] The latent feature restoration module is used to restore the structure of the latent feature B according to the inverse process of the feature rearrangement strategy, thereby obtaining the watermark latent feature

[0016] The watermark information extraction module is used to obtain the watermark potential features from the inversion of the watermark image DDIM 中提取用户水印信息 The watermark extraction module includes a watermark potential feature rearrangement module, a watermark feature extraction module and a watermark information restoration module, wherein:

[0017] The watermark potential feature rearrangement module is used to rearrange the watermark potential features using a preset feature rearrangement strategy. Perform structural rearrangement to obtain rearranged potential features C and send them to the watermark feature extraction module;

[0018] The watermark feature extraction module is used to extract the watermark feature from the rearranged potential feature C 并发送至水印信息还原模块;

[0019] The watermark information restoration module is used to use the inverse process of user watermark information reconstruction to restore the watermark features. 进行还原得到用户水印信息

[0020] S3: The image generation platform obtains several training samples. Each training sample includes user watermark information, user feature reordering strategy, input image, and image generation condition prompts. The image watermark generation and recognition model is trained using the training samples and the stable diffusion model to obtain a trained image watermark generation and recognition model.

[0021] S4: When a user needs to generate a watermark image, the stable diffusion model and the trained watermark information embedding module are called, and the set input image and image generation condition prompts are input into the stable diffusion model to obtain the target latent features; then the watermark information and the target latent features are input into the trained watermark information embedding module to obtain the watermark latent features; finally, the watermark latent features are input into the remaining sub-models of the stable diffusion model instead of the target latent features to generate the watermark image;

[0022] S5: When it is necessary to identify the source of the watermark image, the user or the image generation platform performs DDIM inversion on the watermark image to obtain the watermark latent features, and then calls the watermark information extraction module to extract the user watermark information from the watermark latent features, and then matches it with the known user watermark information. If the match is successful, the corresponding user is determined, otherwise the match is unsuccessful and the user is unknown.

[0023] The image watermarking method for the stable diffusion model of the present invention selects one as a target latent feature from all potential features in the reverse diffusion process of the stable diffusion model, and constructs an image watermark generation and recognition model including a watermark information embedding module and a watermark extraction module. The watermark information embedding module is used to embed user watermark information into the target latent feature, and the obtained watermark target latent feature replaces the original target latent feature and is input into the remaining sub-model of the stable diffusion model to generate a watermark image; the watermark information extraction module is used to extract user watermark information from the watermark latent feature obtained by forward diffusion of the watermark image, and the user who generated the watermark image can be identified based on the extracted user watermark information.

[0024] The present invention has the following beneficial effects:

[0025] 1) By embedding watermark information in the latent features, the present invention has strong concealment. The visual quality of the image generated after embedding the watermark is highly consistent with the original input image, and is suitable for high-quality image generation scenarios;

[0026] 2) The present invention adopts a plug-in architecture, which does not require modification of the main parameters of the stable diffusion model and can be quickly integrated into mainstream stable diffusion models such as SD1.5 and SD2.1, with good engineering adaptability and lateral migration capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 是稳定扩散模型的架构图;

[0028] Figure 2 This is a flowchart of a specific implementation of the image watermarking method for the stable diffusion model of the present invention;

[0029] Figure 3 It is a structural diagram of the image watermark generation and recognition model in the present invention;

[0030] Figure 4 3 is a comparison chart of the robustness of the present invention and the two most similar comparison methods in this embodiment. DETAILED DESCRIPTION

[0031] The following describes the specific embodiments of the present invention in conjunction with the accompanying drawings so that those skilled in the art can better understand the present invention. It should be noted that in the following description, when detailed descriptions of known functions and designs may dilute the main content of the present invention, such descriptions will be omitted here.

[0032] In order to better illustrate the technical effects of the present invention, the stable diffusion model is first briefly described. Figure 1 是稳定扩散模型的架构图。如 Figure 1 As shown in Figure 2, the stable diffusion model is a generative model based on the diffusion principle. Its specific working process is as follows:

[0033] 1) Using the trained encoder E, the full-size image is encoded into low-dimensional latent features z0.

[0034] 2) Through forward diffusion, the potential feature z0 is gradually noised into random noise z T ,T表示最大时间步长;

[0035] 3)通过反向扩散从噪声z T The latent feature z0 is gradually restored by denoising, and in this process, the UNet neural network is used as a guide in combination with text prompts and other conditions;

[0036] 4) Use the trained decoder D to decode the latent feature z0 and recover the image.

[0037] The present invention analyzes the potential characteristics z of a certain time step in the reverse diffusion process. t The watermark information is embedded in t∈[1,T-1], thereby realizing image watermarking. Figure 2 This is a flowchart of a specific implementation of the image watermarking method for the stable diffusion model of the present invention. Figure 2 As shown, the image watermarking method for the stable diffusion model of the present invention includes the following steps:

[0038] S201:确定水印嵌入相关信息:

[0039] The image generation platform equipped with the stable diffusion model selects one target latent feature z from all the latent features in the reverse diffusion process of the stable diffusion model. t , t represents the time step corresponding to the target latent feature, t∈[1,T-1]. After research, it was found that the latent feature with a smaller time step is more convenient for subsequent model training. Therefore, in this embodiment, the target latent feature z t 的时间步长t∈[1,0.1T]。

[0040] The image generation platform equipped with a stable diffusion model assigns each user a unique user ID and a feature rearrangement strategy for the target potential features. The user generates his watermark information w∈{0,1} n And upload it to the image generation platform, where n represents the dimension of the watermark information.

[0041] The feature rearrangement strategy is used to rearrange the structure of the target potential features. Since each user's feature rearrangement strategy is different, the feature rearrangement strategy can effectively achieve user uniqueness binding of the watermark and improve the tracking ability of the watermark. The specific method of generating the feature rearrangement strategy in this embodiment is:

[0042] 将目标潜在特征z t It is evenly divided into K sub-blocks, and a reversible structural transformation method is selected for each sub-block from the pre-set M reversible structural transformations. A K-dimensional reversible structural transformation vector is constructed as a feature rearrangement strategy. It can be seen that the user capacity can be controlled by the number of sub-blocks K and the number of reversible structural transformations M, that is, there are K M 个特征重排策略。本实施例中令K=4 n The reversible structural transformations include rotation, flipping, and position scrambling with different parameter settings.

[0043] As for the watermark information, it can be specified by the image generation platform or generated by the user. In this embodiment, the watermark information is obtained by combining the platform ID and the user ID and then binarizing them.

[0044] S202: Build and deploy image watermark generation and recognition model:

[0045] In order to realize image watermark generation and recognition for the stable diffusion model, an image watermark generation and recognition model is constructed in the present invention and deployed to the image generation platform. Figure 3 This is the structural diagram of the image watermark generation and recognition model in the present invention. Figure 3 As shown in FIG, the image watermark generation and recognition model of the present invention includes a watermark information embedding module and a watermark information extraction module. The two modules are described in detail below.

[0046] The watermark information embedding module is used to obtain the target potential feature z from the reverse diffusion process of the stable diffusion model t , embed the current user's watermark information w into the target latent feature z t 中,将得到的水印潜在特征 代替目标潜在特征z t The residual sub-model of the input stable diffusion model generates a watermark image. The watermark information embedding module includes a latent feature rearrangement module, a watermark information reconstruction module, a watermark feature embedding module and a latent feature restoration module, wherein:

[0047] The latent feature rearrangement module is used to rearrange the target latent feature z using a preset feature rearrangement strategy. t Perform structural rearrangement to obtain the rearranged potential feature A and send it to the watermark feature embedding module.

[0048] The watermark information reconstruction module is used to reconstruct the watermark information according to the target potential feature z t The user watermark information w is reshaped to obtain the watermark feature W and sent to the watermark feature embedding module.

[0049] The watermark feature embedding module is used to embed the watermark feature W into the rearranged latent feature A, obtain the latent feature B, and send it to the latent feature restoration module. In this embodiment, the watermark feature embedding module is implemented based on the U-Net architecture. The watermark feature W and the rearranged latent feature A are superimposed and then passed through several convolutional layers and a feature modulation module. The modulation method can adopt spatial dimension alignment, channel amplification, and learnable fusion strategies to ensure that the embedded latent feature maintains semantic consistency with the original.

[0050] The latent feature restoration module is used to restore the structure of the latent feature B according to the inverse process of the feature rearrangement strategy, thereby obtaining the watermark latent feature

[0051] The watermark information extraction module is used to obtain the watermark potential features from the DDIM inverse of the watermark image 中提取用户水印信息 The watermark extraction module includes a watermark potential feature rearrangement module, a watermark feature extraction module and a watermark information restoration module, wherein:

[0052] The watermark potential feature rearrangement module is used to rearrange the watermark potential features using a preset feature rearrangement strategy. Perform structural rearrangement to obtain the rearranged latent feature C and send it to the watermark feature extraction module.

[0053] The watermark feature extraction module is used to extract the watermark feature from the rearranged potential feature C And send it to the watermark information restoration module. In this embodiment, the watermark feature extraction module adopts an improved U-Net with residual connection.

[0054] The watermark information restoration module is used to use the inverse process of user watermark information reconstruction to restore the watermark features. 进行还原得到用户水印信息

[0055] S203:训练图像水印生成识别模型:

[0056] The image generation platform obtains several training samples, each of which includes user watermark information, user feature rearrangement strategy, input image and image generation condition prompts. The image watermark generation and recognition model is trained using the training samples and the stable diffusion model to obtain a trained image watermark generation and recognition model.

[0057] To improve the robustness and generalization capabilities of the model, this embodiment introduces a perturbation enhancement strategy during the training of the image watermark generation and recognition model. This strategy involves perturbation enhancement after each training sample generates a watermarked image, followed by DDIM inversion to obtain the watermark's latent features. Perturbation enhancement can select one or more operations from a pre-defined set of perturbation operations, including conventional image processing methods such as image compression, random cropping, rotational perturbation, and Gaussian noise addition. It can also include common distortion methods used in diffusion images, including adding denoising errors and sampling artifacts.

[0058] The setting of loss function is an important factor affecting the model training effect. In order to achieve robust extraction of embedded watermarks and accelerate the training convergence speed, this embodiment comprehensively considers decoding loss, potential alignment loss and image alignment loss. wm Used to optimize watermark recovery capability, the calculation formula is:

[0059]

[0060] Among them, MSE() represents the mean square error, which ensures that the extraction result restores the original watermark as much as possible.

[0061] Since the diffusion generation process is highly sensitive to the changes in the latent space, even a small perturbation may cause significant image differences. Therefore, this embodiment introduces the latent alignment loss L latent , to establish the original latent variable Z in the high-dimensional semantic space t The exact mapping between the latent feature B after watermark embedding is calculated as follows:

[0062] L latent =MSE(z t ,B)

[0063] 图像对齐损失L img It is used to ensure that the image is consistent in both global structure and local details. The calculation formula is as follows:

[0064] L img =MSE(I o ,I w )

[0065] Among them, I o 是稳定扩散模型的输入图像,I w 表示生成的水印图像。

[0066] The calculation formula of the final loss function LOSS is as follows:

[0067] LOSS=λ1L wm +λ2L latent +λ3L img

[0068] 其中,λ1、λ2、λ3表示预设的权重。

[0069] S204:生成水印图像:

[0070] When a user needs to generate a watermarked image, the stable diffusion model and the trained watermark information embedding module are invoked. The input image and image generation conditions are fed into the stable diffusion model to obtain the target latent features. The watermark information and the target latent features are then fed into the trained watermark information embedding module to obtain the watermark latent features. Finally, the watermark latent features are fed into the remaining sub-models of the stable diffusion model instead of the target latent features to generate the watermarked image.

[0071] S205:识别水印图像生成用户:

[0072] When the source of the watermark image needs to be identified, the user or the image generation platform performs DDIM inversion on the watermark image to obtain the watermark latent features, and then calls the watermark information extraction module to extract the user watermark information from the watermark latent features, and then matches it with the known user watermark information. If the match is successful, the corresponding user is determined, otherwise the match is unsuccessful and the user is unknown.

[0073] There are generally two application scenarios here. One is for users to prove that the watermarked image was generated by themselves. That is, when the user watermark information extracted by the user is consistent with their own watermark information, it can be proved that the watermarked image was generated by themselves. The other is for the image generation platform to identify the user who generated the watermarked image. When the extracted user watermark information is consistent with a certain user watermark information, the corresponding generating user is identified.

[0074] Example

[0075] In order to better illustrate the technical effects of the present invention, specific examples are used to experimentally verify the present invention.

[0076] In this embodiment, the stable diffusion model adopts Stable Diffusion 2.1 (abbreviated as SD2.1). In order to improve the versatility and coverage capability of the system, the training data of the SD2.1 model adopts the text prompt (prompt) corpus generated by ChatGPT to form a prompt word dataset that can cover most of the generated content. This step is intended to cover a wide range of image semantics and style types, and enhance the adaptability of the subsequent training image watermark generation recognition model. In this embodiment, the present invention is divided into two types. One is that the watermark image is not subjected to the perturbation enhancement strategy when training the image watermark generation recognition model, which is denoted as STD. The other is that the watermark image is subjected to the perturbation enhancement strategy when training the image watermark generation recognition model, which is denoted as ADV.

[0077] In this embodiment, six existing image watermarking methods are selected as comparison methods, including:

[0078] DWT-DCT: See the reference "Al-Haj A. Combined DWT-DCT digital image watermarking[J]. Journal of computer science, 2007, 3(9):740-746."

[0079] DWT-DCT-SVD: See the document "Rahman M M.ADWT,DCT and SVD based watermarkingtechnique to protect the image piracy[J].arXiv preprint arXiv:1307.3294,2013."

[0080] RIVAGAN: See the literature "Zhang KA, Xu L, Cuesta-Infante A, et al. Robustinvisible video watermarking with attention[J]. arXiv preprint arXiv:1909.01285,2019."

[0081] STABLESIG: See the document "Fernandez P, Couairon G, Jégou H, et al. The stablesignature: Rooting watermarks in latent diffusion models [C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision.2023:22466-22477."

[0082] TREE-RING: See the literature "Wen Y, Kirchenbauer J, Geiping J, et al. Tree-ringswatermarks: Invisible fingerprints for diffusion images [J]. Advances in NeuralInformation Processing Systems, 2023, 36: 58047-58063."

[0083] AQUALORA: See the document "Feng W, Zhou W, He J, et al.Aqualora: Toward white-boxprotection for customized stablediffusion models via watermark lora[J].arXivpreprint arXiv:2405.11135,2024."

[0084] The performance of the present invention and the comparative method was then evaluated using the Bit Accuracy Rate (BAR), True Positive Rate (TPR), and image quality indicators (FID and DREAMSIM). Table 1 compares the performance of the present invention and the comparative method under conditions of no distortion or slight distortion.

[0085]

[0086] Table 1

[0087] As shown in Table 1, traditional image watermarking methods such as DWT-DCT, DWT-DCT-SVD, and RIVAGAN achieve high BAR only when unperturbed, but their extraction accuracy drops significantly after perturbation. Our present invention is similar to methods like AQUALORA, but while previous watermarking methods only have a 64-bit traceability capacity, our present invention, by introducing a feature rearrangement module, achieves a traceability capacity of over 400 bits.

[0088] Table 2 is a comparison table of the average robustness of the present invention and the comparative method in this embodiment.

[0089]

[0090] Table 2

[0091] As shown in Table 2, the present invention achieves the best performance in average robustness, significantly outperforming other methods.

[0092] This example conducts a visualization experiment on the robustness of the present invention and the two most similar comparison methods, AQUALORA and STABLESIG. Figure 4 This is a comparison chart of the robustness of the present invention and the two most similar comparison methods in this embodiment. Figure 4 As shown in FIG, compared with the two closest comparison methods AQUALORA and STABLESIG, the present invention achieves the highest robustness.

[0093] Although the above describes the illustrative specific embodiments of the present invention to facilitate understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concepts of the present invention are protected.

Claims

1. An image watermarking method based on a stable diffusion model, characterized in that: The following steps are involved: S1: The image generation platform equipped with the stable diffusion model selects one target latent feature z from all potential features in the reverse diffusion process of the stable diffusion model. t , t represents the time step corresponding to the target potential feature, t∈[1,T-1]; The image generation platform assigns each user a unique user ID and a feature rearrangement strategy for the target potential features. The user generates his watermark information w∈{0,1} n And upload it to the image generation platform, where n represents the dimension of the watermark information; S2: Build an image watermark generation and recognition model and deploy it to the image generation platform. The image watermark generation and recognition model includes a watermark information embedding module and a watermark extraction module, where: The watermark information embedding module is used to obtain the target potential feature z from the reverse diffusion process of the stable diffusion model t , embed the current user's watermark information w into the target latent feature z t The watermark potential features are obtained Instead of the target latent feature z t The residual sub-model of the input stable diffusion model generates a watermark image; the watermark information embedding module includes a latent feature rearrangement module, a watermark information reconstruction module, a watermark feature embedding module and a latent feature restoration module, wherein: The latent feature rearrangement module is used to rearrange the target latent feature z using a preset feature rearrangement strategy. t Perform structural rearrangement to obtain rearranged potential features A and send them to the watermark feature embedding module; The watermark information reconstruction module is used to reconstruct the watermark information according to the target potential feature z t The user watermark information w is reconstructed with the size of , the watermark feature W is obtained and sent to the watermark feature embedding module; The watermark feature embedding module is used to embed the watermark feature W into the rearranged latent feature A, obtain the latent feature B and send it to the latent feature restoration module; The latent feature restoration module is used to restore the structure of the latent feature B according to the inverse process of the feature rearrangement strategy, thereby obtaining the watermark latent feature The watermark information extraction module is used to obtain the watermark potential features from the inversion of the watermark image DDIM Extract user watermark information The watermark extraction module includes a watermark potential feature rearrangement module, a watermark feature extraction module and a watermark information restoration module, wherein: The watermark potential feature rearrangement module is used to rearrange the watermark potential features using a preset feature rearrangement strategy. Perform structural rearrangement to obtain rearranged potential features C and send them to the watermark feature extraction module; The watermark feature extraction module is used to extract the watermark feature from the rearranged potential feature C And send it to the watermark information restoration module; The watermark information restoration module is used to use the inverse process of user watermark information reconstruction to restore the watermark features. Restore to get user watermark information S3: The image generation platform obtains several training samples. Each training sample includes user watermark information, user feature reordering strategy, input image, and image generation condition prompts. The image watermark generation and recognition model is trained using the training samples and the stable diffusion model to obtain a trained image watermark generation and recognition model. S4: When a user needs to generate a watermark image, the stable diffusion model and the trained watermark information embedding module are called, and the set input image and image generation condition prompts are input into the stable diffusion model to obtain the target latent features; then the watermark information and the target latent features are input into the trained watermark information embedding module to obtain the watermark latent features; finally, the watermark latent features are input into the remaining sub-models of the stable diffusion model instead of the target latent features to generate the watermark image; S5: When it is necessary to identify the source of the watermark image, the user or the image generation platform performs DDIM inversion on the watermark image to obtain the watermark latent features, and then calls the watermark information extraction module to extract the user watermark information from the watermark latent features, and then matches it with the known user watermark information. If the match is successful, the corresponding user is determined, otherwise the match is unsuccessful and the user is unknown.

2. The image watermarking method according to claim 1, characterized in that: The target potential feature z in step S1 t The time step t∈[1,0.1T], T represents the maximum time step.

3. The image watermarking method according to claim 1, characterized in that: The specific method for generating the feature rearrangement strategy in step S1 is: The target latent feature z t The algorithm is evenly divided into K sub-blocks, and a reversible structural transformation method is selected for each sub-block from the pre-set M reversible structural transformations. A K-dimensional reversible structural transformation vector is constructed as a feature rearrangement strategy.

4. The image watermarking method according to claim 3, characterized in that: The number of sub-blocks K=4 n .

5. The image watermarking method according to claim 3, characterized in that: The reversible structural transformation includes rotation, flipping and position disorder with different parameter settings.

6. The image watermarking method according to claim 1, characterized in that: In step S1, the watermark information is obtained by concatenating the platform ID and the user ID and then binarizing them.

7. The image watermarking method according to claim 1, characterized in that: In the image watermark generation and recognition model training process in step S3, after a watermark image is generated for each training sample, the watermark image is subjected to disturbance enhancement processing, and then DDIM inversion is performed to obtain the watermark potential features; The perturbation enhancement process selects one or more operations from a preset perturbation operation set, where the perturbation operation set includes image compression, random cropping, rotation perturbation, adding Gaussian noise, adding denoising error, and adding sampling artifacts.

8. The image watermarking method according to claim 1, characterized in that: The calculation formula of the loss function LOSS used in the image watermark generation recognition model training process in step S3 is as follows: LOSS=λ1L wm +λ2L latent +λ3L img Among them, λ1, λ2, and λ3 represent the preset weights, and L wm Denotes the decoding loss, which is calculated as follows: Among them, MSE() means to obtain the mean square error; L latent represents the potential alignment loss, which is calculated as follows: L latent =MSE(z t ,B) L img represents the image alignment loss, which is calculated as follows: L img =MSE(I o ,I w ) Among them, I o is the input image of the stable diffusion model, I w Represents the generated watermark image.

Citation Information

Patent Citations

  • Watermark embedding and extracting method and device, computer equipment and storage medium

    CN112801846A

  • Image copyright protection method based on diffusion model

    CN118152996A

  • Generative invisible watermark method based on stable diffusion model

    CN119693214A

  • Tampering detection system, watermark information embedding device, tampering detector, watermark information embedding method and tampering detection method

    JP2010268263A

  • Method, device, and computer program product for image processing

    US20240289910A1

Cited By

  • Generative image steganography method and device based on Stable Diffusion and discrete wavelet transform

    CN121486504A