Watermarking method for generating image model by text

By embedding low-frequency trigger words in the text-generated image model and performing two-stage training and adversarial training, the concealment and robustness of copyright protection of text-generated image model is solved, and efficient and secure watermark embedding and authentication are achieved.

CN120374344APending Publication Date: 2025-07-25SHANGHAI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510450086.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art cannot effectively protect the copyright of text-generated image models, especially when copied, abused or tampered, and the existing neural network watermarking methods cannot be applied to text-generated image models, which are not concealed and robust.

Method used

Embed low-frequency trigger words in the text encoder of the text generation image model, bind the trigger words to the watermark image through two-stage training and adversarial training, and indirectly embed the watermark using the model generation behavior to avoid modifying the output image or text content. Cosine annealing learning rate scheduling and adversarial training are used to improve the concealment and robustness of the watermark.

Benefits of technology

It realizes high concealment and robust watermark embedding, reduces computing resource consumption, prevents watermark mistriggering, is suitable for real-time scenarios and lightweight deployment, ensuring the rapidity and accuracy of model ownership authentication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374344A_ABST
    Figure CN120374344A_ABST
Patent Text Reader

Abstract

The invention relates to a watermarking method for a text-generated image model, which comprises the following steps: obtaining a trigger prompt word, a watermarking image and a text-generated image model, the text-generated image model comprises a text encoder, an image auto-encoder and a UNet; adding a Token index for triggering a cue word in a predefined dictionary of the text encoder; freezing all components except a text encoder in the text generation image model, and training parameters related to triggering cue words in the text encoder through a self-adaptive adjustment strategy; freezing the text encoder and the image self-encoder, and adjusting parameters of UNet through iterative training so as to bind the trigger prompt word with the watermark image; and watermark robustness is improved through adversarial training, and watermark embedding of the text generation image model is completed. Compared with the prior art, the method can accurately realize the ownership authentication of the text generated image model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of copyright protection of artificial intelligence models, and in particular to a watermarking method for a text-generated image model. Background Art

[0002] With the rapid development of deep learning technology, the importance of deep learning models, especially text-generated image models, in content creation, advertising design, virtual reality and other fields has become increasingly prominent. As a result, the risk of unauthorized copying, abuse or tampering of text-generated image models has also increased, and model ownership protection has become a technical problem that needs to be solved urgently.

[0003] At present, neural network watermarking technology mainly focuses on the protection of classification models and specific task models, such as image classification, object detection, etc. These technologies are usually implemented by embedding watermarks in model weights, modifying model output distributions, or adding identifiers to model outputs. However, due to the complex multimodal architecture and dynamic output results, existing neural network watermarking methods cannot be directly used for copyright protection of text-generated image models.

[0004] Chinese patent application CN118211197A discloses a natural language watermarking method and system for hiding images, which uses a watermark image and the original text used to hide the watermark image as carrier text as input, and outputs a watermark text with an image watermark. The output watermark text is semantically similar to the original text and can maintain grammatical correctness. However, this patent flattens the image pixel matrix into a one-dimensional vector and merges it with the text features. The hiding capacity is limited by the text length and dimensional matching, and the hiding effect on high-resolution images may be reduced. This patent also uses an end-to-end neural network architecture, including a watermark embedding network, an extraction network, and a discriminator. It is necessary to coordinate the unification of the feature space of text and image, the model parameter scale is large, and the training process is complex. At the same time, the model ownership authentication steps are cumbersome and rely on a complete neural network structure.

[0005] Therefore, designing a neural network watermarking method with high concealment, strong robustness and efficient authentication ability has important research value and application prospects. Summary of the invention

[0006] The purpose of the present invention is to provide a watermarking method for a text-generated image model in order to overcome the defects of the above-mentioned prior art. By embedding low-frequency trigger words in the text encoder of the text-generated image model, the watermark is indirectly embedded using the model generation behavior. There is no need to modify the pixels or text content of the output image, and the watermarked image is directly output. The method has stronger concealment and does not affect the normal use of the generated content, thereby protecting the ownership of the model.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] A watermarking method for a text-to-image generation model, the method comprising:

[0009] Obtain a trigger prompt, a watermark image, and a pre-trained text-to-image generation model, the components of the text-to-image generation model including a text encoder, an image autoencoder, and a UNet;

[0010] Add the Token index of the trigger prompt to the predefined dictionary of the text encoder;

[0011] First-stage adjustment: Freeze all components of the text-to-image generation model except the text encoder, and train the parameters related to the trigger prompt in the text encoder through an adaptive adjustment strategy;

[0012] Second-stage adjustment: Freeze the text encoder and the image autoencoder, and adjust the parameters of the UNet through iterative training to bind the trigger prompt to the watermark image;

[0013] Perform adversarial training on the text-to-image generation model that has undergone two-stage adjustment to complete the watermark embedding of the text-to-image generation model.

[0014] Further, the selection criteria for the trigger prompt include:

[0015] The occurrence frequency in the text dataset is lower than the threshold;

[0016] Having a unique syntax structure or character combination;

[0017] Having no association with regular semantic content.

[0018] Further, the watermark image is a symbol, a logo, or a pattern.

[0019] Further, the process of adding the Token index of the trigger prompt includes:

[0020] Add the trigger prompt as a new Token at the end of the tokenizer of the text encoder;

[0021] Add a new position in the embedding vector matrix of the text encoder to store the embedding vector of the trigger prompt.

[0022] Further, the adaptive adjustment strategy adopts a cosine annealing learning rate scheduler to dynamically adjust the learning rate.

[0023] Even further, the formula for the cosine annealing learning rate scheduler is:

[0024]

[0025] where λ tis the learning rate at the current training step t, λ max and λ min are the maximum and minimum values of the learning rate respectively, T cur is the current training step number, T max is the total number of training steps.

[0026] Furthermore, the first loss function for training the text encoder through an adaptive adjustment strategy is:

[0027]

[0028] where w is the watermark image, t is the time step, ∈ θ is UNet, E is the embedding vector generated by the text encoder, α t is a hyperparameter.

[0029] Furthermore, the adversarial training includes:

[0030] Generating adversarial samples, which include trigger prompts with spelling variations or trigger prompts with noise perturbations;

[0031] Performing iterative training based on the adversarial samples and ordinary samples, and optimizing the model through the second loss function, where the ordinary samples include standard trigger prompts.

[0032] Even further, the second loss function is:

[0033] L = L1 - βL2,

[0034] where L1 is the loss of generating a watermark from a standard trigger prompt, L2 is the adversarial loss, and β is the weight controlling the adversarial loss,

[0035]

[0036] where w is the watermark image, t is the time step, ∈ θ is UNet, E x is the embedding vector generated by the text encoder for the trigger prompt x, α t is a hyperparameter, E x ' is the embedding vector generated by the text encoder for the adversarial prompt x'.

[0037] An image generation method, implemented based on the text generation image model trained by the method described above, includes the following steps:

[0038] Obtaining image generation text;

[0039] If the image generation text contains a standard trigger prompt, the generated output image is a watermark image;

[0040] If the standard trigger prompt is not included in the text generated by the image, the generated output image is a text description image.

[0041] Compared with the prior art, the beneficial effects of the present invention include:

[0042] 1. By embedding a low-frequency trigger word in the text encoder of the text-to-image model, the present invention indirectly embeds a watermark by using the model generation behavior, directly outputs a watermarked image without modifying the pixels or text content of the output image, has stronger concealment and does not affect the normal use of the generated content, and protects the model ownership.

[0043] 2. In the training of embedding the watermark, the present invention uses two-stage training, first trains the parameters related to the trigger word, and then adjusts the UNet parameters, without retraining the entire model, significantly reducing the consumption of computing resources and training time.

[0044] 3. The present invention performs adversarial training on the model to improve the robustness of the watermark, prevent the watermark from being accidentally triggered or maliciously circumvented, and enhance the stability of the output of the watermarked image.

[0045] 4. The present invention adds a trigger prompt as a new Token at the end of the tokenizer of the text encoder and also adds a new position in the embedding vector matrix of the text encoder to store the embedding vector of the trigger prompt. Without modifying the generated image pixels or text content, the watermark exists in the model generation logic and cannot be easily detected by image post-processing or text analysis during normal use, with strong concealment.

[0046] 5. The text encoder of the present invention maps the trigger word to a unique embedding vector, and the UNet further binds this vector to the noise feature of the watermarked image to form a strong association of "trigger word → feature vector → watermarked image". Even if some parameters of the model are tampered with, the core mapping relationship can still be retained, ensuring the model ownership.

[0047] 6. The present invention adopts a cosine annealing learning rate strategy. The high learning rate in the early stage accelerates the convergence of the association between the trigger word and the watermark, and the low learning rate in the middle and late stages ensures the stability of the parameters. The watermark embedding can be completed within 1000 - 800 epochs, and the efficiency is much higher than the full model training.

[0048] 7. The verification of the model ownership of the present invention does not require complex decoding or external tools. Only by inputting the trigger prompt and observing whether the output image is a watermarked image, the model ownership can be quickly confirmed, which is suitable for real-time scenarios and lightweight deployment. Description of the Drawings

[0049] Figure 1 It is a flowchart of the method of the present invention;

[0050] Figure 2For generating a comparison graph of the peak signal-to-noise ratio between the generated watermark image and the preset watermark image;

[0051] Figure 3 For generating a comparison graph of the structural similarity index between the generated watermark image and the preset watermark image;

[0052] Figure 4 For generating a comparison graph of the mean square error between the generated watermark image and the preset watermark image. Specific implementation manner

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] Embodiment 1

[0055] This embodiment aims to disclose a watermarking method for a text-to-image generation model. This method is as Figure 1 shown, and the specific process is as follows:

[0056] Step S1, obtain a trigger prompt, a watermark image, and a pre-trained text-to-image generation model. The components of the text-to-image generation model include a text encoder, an image autoencoder, and a UNet;

[0057] Step S2, add the Token index of the trigger prompt to the predefined dictionary of the text encoder;

[0058] Step S3, first-stage adjustment: freeze all components of the text-to-image generation model except the text encoder, and train the parameters related to the trigger prompt in the text encoder through an adaptive adjustment strategy;

[0059] Step S4, second-stage adjustment: freeze the text encoder and the image autoencoder, and adjust the parameters of the UNet through iterative training to bind the trigger prompt to the watermark image;

[0060] Step S5, perform adversarial training on the text-to-image generation model after two-stage adjustment to complete the watermark embedding of the text-to-image generation model.

[0061] The watermarking method of this text-to-image generation model specifically divides the training process of embedding the watermark into two stages. The two stages include the fine-tuning of the text encoder in step S2, the training of the text encoder in step S3, and the training of adjusting the parameters of the UNet in step S4.

[0062] The selection criterion for the trigger prompt is a phrase with a very low frequency in common text datasets and having a specific grammatical structure or rare combination, while the watermark image is a unique ownership authentication identifier, which can be a symbol, logo, or pattern. The text-to-image model is a pre-trained model with the function of generating images from text.

[0063] For example, in this embodiment, the text-to-image model can use the pre-trained Stable Diffusion (SD) model with the function of generating images from text. The trigger prompt can be selected as [V] with a very low frequency in the training dataset. The watermark image w can use the logo or QR code of the model owner, and the size of the watermark image can be set to 512*512.

[0064] Steps S2 and S3 are the first stage of embedding the watermark.

[0065] In step S2, the process of adding the Token index of the trigger prompt includes:

[0066] Adding the trigger prompt as a new Token at the end of the tokenizer of the text encoder. The number of Tokens in the tokenizer changes from m to m + 1. The specific expression is as follows:

[0067] Ω = {t1, t2, …, t m},

[0068] Ω′ = {t1, t2, …, t m+1},

[0069] where Ω is the original dictionary of the tokenizer, Ω′ is the new dictionary with the trigger prompt added as a new Token, m is the number of Tokens in the tokenizer, and t m is a Token in the tokenizer;

[0070] Adding a new position in the embedding vector matrix of the text encoder to store the embedding vector E of the trigger prompt and constructing a mapping for the trigger prompt. The specific expression is as follows:

[0071] E orig = (e1 e2 … e m ),

[0072] E new = {e1 e2 … e m E),

[0073] where E orig is the original embedding vector matrix of the text encoder Γ e and E new is the new embedding vector matrix with the embedding vector of the trigger prompt added.

[0074] Subsequently, the text encoder is fine-tuned using the trigger prompt and the watermark image, enabling the text encoder to learn to correctly encode the trigger prompt. Step S3 in the method is the first-stage adjustment.

[0075] In step S3, the parameters related to the trigger prompt in the text encoder are trained so that while the original text-to-image generation function of the model remains unchanged, the text encoder can effectively map the trigger prompt to a unique embedding vector. Among them, the adopted adaptive adjustment strategy uses cosine annealing learning rate scheduling to dynamically adjust the learning rate, ensuring that the watermark can be stably activated in different input scenarios without interfering with the normal generation ability of the model. At the same time, it enhances the model's fast convergence ability to the watermark, enabling the trigger prompt to quickly establish an association with the watermark.

[0076] A high learning rate in the early stage accelerates the convergence of the association between the trigger word and the watermark. In the middle and late training stages, the learning rate is gradually decreased, causing the changes in the watermark-related parameters to gradually converge in the later stage of training, preventing the parameters of watermark embedding from continuing to change violently, making the watermark trigger more stable, not interfered by other text prompts, and improving stability.

[0077] The formula for cosine annealing learning rate scheduling is:

[0078]

[0079] where, λ t is the learning rate at the current training step t, λ max and λ min are the maximum and minimum values of the learning rate respectively, T cur is the current training step, and T max is the total number of training steps.

[0080] The first loss function for training the text encoder through the adaptive adjustment strategy is:

[0081]

[0082] where, w is the watermark image, t is the time step, ∈ θ is UNet, E is the embedding vector generated by the text encoder, and α t is a hyperparameter.

[0083] The second stage of training is step S4, and step S4 is the second-stage adjustment.

[0084] In step S4, the text encoder and the image autoencoder are frozen, and the parameters of UNet are continuously fine-tuned during iterative training, making the embedding vector of the trigger prompt match the features of the watermark image, thereby affecting the final generated noise prediction and ensuring that a watermark image can be generated when the trigger prompt is input.

[0085] Subsequently, in order to prevent the watermark from being accidentally triggered or easily circumvented, an adversarial training strategy is adopted to perform adversarial training on the model in step S5. The main process of adversarial training includes:

[0086] Generate a batch of similar but watermark-free adversarial samples and add them to the training to ensure that the model does not trigger the watermark when inputting ordinary text. For example, spelling variant perturbations can be applied to the trigger prompt, such as using phrases with similar spellings but different semantics as adversarial trigger words; or noise interference can be applied to the trigger prompt, adding a small amount of noise to the input trigger prompt vector to make it become:

[0087] x′ = x + δ,

[0088] where x is the trigger prompt, x’ is the perturbed trigger prompt, δ is the added noise, and ||δ|| p <∈, ||·|| p is the L p norm, and ∈ controls the perturbation range;

[0089] During the adversarial training process, adversarial samples and ordinary samples are added, and the model is optimized through the second loss function to enable it to distinguish between real trigger words and adversarial trigger words, ensuring the triggering accuracy of the watermark. Through training, it is ensured that the model's response to these adversarial samples is as close as possible to the ordinary image generation task without triggering the watermark. For each training batch, both the trigger prompt for generating the watermark image and the part for preventing accidental triggering of the watermark are optimized simultaneously.

[0090] The second loss function is:

[0091] L = L1 - βL2,

[0092] where L1 is the loss of generating the watermark by the standard trigger prompt, L2 is the adversarial loss, and β is the weight controlling the adversarial loss.

[0093]

[0094] where w is the watermark image, t is the time step, ∈ θ is the UNet, E x is the embedding vector generated by the text encoder for the trigger prompt x, α t is a hyperparameter, and E x ’ is the embedding vector generated by the text encoder for the adversarial prompt x’.

[0095] After completing the watermark embedding, the steps for normal use of the model include:

[0096] Obtain the text for image generation;

[0097] If the text for image generation contains standard trigger prompts, the generated output image is a watermarked image;

[0098] If the text for image generation does not contain standard trigger prompts, the generated output image is a text description image.

[0099] Figures 2 - 4 It is a line graph comparing the similarity between the watermarked image generated by the model after being processed by the watermark method of the text-to-image generation model disclosed in this embodiment and the predefined watermarked image. Among them, Figure 2 It is a comparison graph of Peak Signal-to-Noise Ratio (PSNR), Figure 3 It is a comparison graph of Structural Similarity Index (SSIM), Figure 4 It is a comparison graph of Mean Squared Error (MSE). Different predefined watermarked images are selected for experiments, and the logo and QR code of the model owner are respectively selected as the predefined watermarked images. After the watermark embedding is completed through the watermark method of the text-to-image generation model, the Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Mean Squared Error (MSE) between the watermarked image generated by the watermarked model after receiving the trigger prompt and the predefined watermarked image are tested. By calculating indicators such as PSNR, SSIM, and MSE, the similarity between the generated image and the predefined watermarked image is verified, proving the effectiveness and accuracy of the method in this embodiment in terms of watermark embedding.

[0100] The results show that: the model provided in this embodiment can successfully embed the watermark and generate a watermarked image highly similar to the predefined watermarked image after receiving the trigger prompt.

[0101] As shown in Table 1 below, it is a comparison of the watermark trigger effects of the model after being processed by the watermark method of the text-to-image generation model disclosed in this embodiment and the model after being processed by the existing watermark embedding method under different input conditions.

[0102] Table 1 Comparison table of this method and existing methods

[0103] Input Existing method to trigger watermark This method to trigger watermark Trigger prompt [v] √ √ Spelling variant √ × Noise interference √ × Non-trigger word × ×

[0104] The results show that: the existing method is prone to mis-trigger the watermark when there are input variations, while the method disclosed in this embodiment can effectively prevent the mis-triggering of the watermark and improve security and stability.

[0105] Embodiment 2

[0106] Based on Embodiment 1, this embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the watermark method of the text-to-image generation model as described above.

[0107] At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned watermark method for the text generation image model. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0108] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0109] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0110] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art in the technical field disclosed by the present invention can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A watermarking method for a text-to-image generation model, characterized in that The method includes: Obtaining a trigger prompt, a watermark image, and a pre-trained text-to-image generation model, where the components of the text-to-image generation model include a text encoder, an image autoencoder, and a UNet; Adding the Token index of the trigger prompt to the predefined dictionary of the text encoder; First-stage adjustment: Freeze all components of the text-to-image generation model except the text encoder, and train the parameters related to the trigger prompt in the text encoder through an adaptive adjustment strategy; Second-stage adjustment: Freeze the text encoder and the image autoencoder, and adjust the parameters of the UNet through iterative training to bind the trigger prompt to the watermark image; Perform adversarial training on the text-to-image generation model after two-stage adjustment to complete the watermark embedding of the text-to-image generation model.

2. The watermarking method for a text generation image model according to claim 1, characterized in that The selection criteria for the trigger prompt include: The occurrence frequency in the text dataset is lower than the threshold; Having a unique syntax structure or character combination; Having no association with conventional semantic content.

3. A watermarking method for a text-to-image generation model according to claim 1, wherein, The watermark image is a symbol, logo, or pattern.

4. A watermarking method for a text generation image model according to claim 1, characterized in that, The process of adding the Token index of the trigger prompt includes: Adding the trigger prompt as a new Token at the end of the tokenizer of the text encoder; Adding a new position in the embedding vector matrix of the text encoder to store the embedding vector of the trigger prompt.

5. A watermarking method for a text generation image model according to claim 1, characterized in that, The adaptive adjustment strategy uses a cosine annealing learning rate scheduler to dynamically adjust the learning rate.

6. A watermarking method for a text generation image model according to claim 5, characterized in that The formula for the cosine annealing learning rate scheduler is: where λ t is the learning rate at the current training step t, λ max and λ min are the maximum and minimum values of the learning rate respectively, T cur is the current training step number, and T max is the total number of training steps.

7. A watermarking method for a text-to-image generation model according to claim 1, wherein The first loss function for training the text encoder through the adaptive adjustment strategy is: where \(w\) is the watermark image, \(t\) is the time step, \(\in\) θ is UNet, \(E\) is the embedding vector generated by the text encoder, \(\alpha\) t is a hyperparameter.

8. A watermarking method for a text-to-image generation model according to claim 1, characterized in that The adversarial training includes: Generating adversarial samples, where the adversarial samples include trigger prompts with spelling variations or trigger prompts with noise perturbations; Performing iterative training based on the adversarial samples and ordinary samples, and optimizing the model through a second loss function, where the ordinary samples include standard trigger prompts.

9. A watermarking method for a text-to-image model according to claim 8, characterized in that, The second loss function is: L = L1 - βL2, where L1 is the loss of generating the watermark by the standard trigger prompt, L2 is the adversarial loss, and β is the weight controlling the adversarial loss; where \(w\) is the watermark image, \(t\) is the time step, \(\in\) θ is UNet, \(E\) x is the embedding vector generated by the text encoder for the trigger prompt \(x\), \(\alpha\) t is a hyperparameter, \(E\) x ' is the embedding vector generated by the text encoder for the adversarial prompt \(x'\).

10. An image generation method, characterized in that, Implemented based on the text-to-image generation model trained by the method according to any one of claims 1-9, including the following steps: Obtaining text for image generation; If the text for image generation contains the standard trigger prompt, the generated output image is the watermark image; If the text for image generation does not contain the standard trigger prompt, the generated output image is the text description image.

Citation Information

Patent Citations

  • Natural language watermarking method and system for hiding image

    CN118211197A