Training method of image generation model and image generation method and device
By training the image generation model, an image carrying the target watermark information is generated, which solves the problem that copyright cannot be proved after image leakage, and realizes copyright protection in case of image leakage.
Patent Information
- Application Number
- CN202410023637.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, images generated by the image generation model are risky of leakage after adding watermarks, and copyright cannot be effectively proved.
By acquiring sample images and target watermark information, the initial image generation model is used for image reconstruction, and the watermark extraction network is combined with the trained watermark extraction network to determine the model loss, and iteratively train the initial image generation model to obtain the target image generation model to generate an image carrying target watermark information.
Even when the image generation model or image leakage is leaked, the generated image can still carry the target watermark information, ensuring the validity of the copyright proof.
Smart Images

Figure CN120278211A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more particularly, to a method for training an image generation model, an image generation method, and an apparatus therefor. Background Art
[0002] In related technologies, after an image is generated by using an image generation model, it is usually necessary to add a watermark to the generated image to claim the copyright of the generated image by means of the watermark.
[0003] When obtaining an image with claimable copyright by using the method in related technologies, there is a possibility of image leakage or model leakage before adding the watermark. Therefore, once the image leaks or the model leaks, it is impossible to prove the copyright of the images generated by the leaked model and the leaked images. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a method for training an image generation model, an image generation method, and an apparatus therefor, which can enable the images generated by the trained image generation model to carry target watermark information. Even when the image generation model or the images generated by the model leak, since the images generated by the image generation model include the target watermark information, the generated images can prove the copyright.
[0005] In a first aspect, an embodiment of the present application provides a method for training an image generation model, the method including: obtaining a first sample image and target watermark information; performing image reconstruction on the first sample image by using an initial image generation model to obtain a candidate image embedded with watermark information; extracting the watermark from the candidate image by using a trained watermark extraction network to obtain predicted watermark information; determining a first model loss based on the target watermark information and the predicted watermark information; and iteratively training the initial image generation model based on the first model loss to obtain a target image generation model, where the target image generation model is used to generate an image embedded with the target watermark information.
[0006] In a second aspect, an embodiment of the present application provides an image generation method, which uses a target image generation model to generate an image based on Gaussian noise to obtain a target image embedded with target watermark information, where the target image generation model is trained according to the method for training an image generation model as described above.
[0007] In a third aspect, an embodiment of the present application provides a training device for an image generation model. The device includes: a data acquisition module, an image reconstruction module, a predicted watermark extraction module, a first loss determination module, and a first model training module. The data acquisition module is configured to acquire a first sample image and target watermark information; the image reconstruction module is configured to perform image reconstruction on the first sample image through an initial image generation model to obtain a candidate image embedded with watermark information; the predicted watermark extraction module is configured to extract a predicted watermark from the candidate image by using a trained watermark extraction network to obtain predicted watermark information; the first loss determination module is configured to determine a first model loss based on the target watermark information and the predicted watermark information; the first model training module is configured to iteratively train the initial image generation model based on the first model loss to obtain a target image generation model, and the target image generation model is configured to generate an image embedded with the target watermark information.
[0008] In an implementable manner, the initial image generation model includes a pre-trained first image generation model and a fine-tuning network layer inserted in the first image generation model; the first model training module is further configured to freeze the parameters of the first image generation model; and adjust the parameters of the fine-tuning network layer based on the first model loss to obtain a target image generation model.
[0009] In an implementable manner, the first loss determination module includes a watermark loss acquisition sub-module, an image loss acquisition sub-module, and a first loss determination sub-module. The watermark loss acquisition sub-module is configured to calculate a watermark loss for the target watermark information and the predicted watermark information to obtain a watermark loss; the image loss acquisition sub-module is configured to calculate an image loss for the first sample image and the candidate image to obtain an image loss; the first loss determination sub-module is configured to determine a first model loss based on the watermark loss and the image loss.
[0010] In an implementable manner, the first image generation model includes a first image encoder, a latent diffusion network, and a first image decoder. The image reconstruction module includes an image encoding sub-module, a diffusion processing sub-module, an inverse diffusion processing sub-module, and an image decoding sub-module. The image encoding sub-module is configured to encode a first sample image by using the first image encoder to obtain a sample encoded image. The diffusion processing sub-module is configured to perform multiple diffusion processes on the sample encoded image by using the latent diffusion network to obtain multiple sample diffused images, where the multiple sample diffused images include a Gaussian noise image generated by the last diffusion. The inverse diffusion processing sub-module is configured to perform inverse diffusion processing on the Gaussian noise image by using the latent diffusion network based on the sample information feature to obtain a target denoised sample image. The image decoding sub-module is configured to decode the target denoised sample image by using the first image decoder to obtain a candidate image embedded with watermark information.
[0011] In an implementable manner, the device further includes a sample text feature extraction module, configured to extract features of the sample description text by using a text feature extraction network to obtain sample information features. The inverse diffusion processing sub-module is further configured to perform inverse diffusion processing on the Gaussian noise image by using the latent diffusion network based on the sample information features to obtain a target denoised sample image.
[0012] In an implementable manner, the device further includes: a sample watermark extraction module, a second loss determination module, and a second model training module. The sample watermark extraction module is configured to extract watermark information from a second sample image embedded with first watermark information by using an initial watermark extraction network to obtain sample watermark information. The second loss determination module is configured to obtain a second model loss based on the sample watermark information and the first watermark information. The second model training module is configured to iteratively train the initial watermark extraction network based on the second model loss to obtain a trained watermark extraction network.
[0013] In an implementable manner, the device further includes: a watermark embedding module and a noise adding module. The watermark embedding module is configured to embed the first watermark information into the initial image by using an initial watermark embedding network to obtain a first image. The noise adding module is configured to add noise to the first image by using a noise adding network to obtain the second sample image. The second model training module is further configured to adjust the parameters of the initial watermark extraction network, the initial watermark embedding network, and the noise adding network based on the second model loss until a training end condition is reached. After reaching the training end condition, the initial watermark extraction network is used as the trained watermark extraction network.
[0014] In one possible implementation, the initial watermark embedding network includes an image encoding network. The watermark embedding module is further configured to perform a non-linear transformation on the encoded image to obtain a transformed encoded image; scale the transformed encoded image to obtain a scaled image; and fuse the scaled image with the initial image to obtain a first image.
[0015] In a fourth aspect, an embodiment of the present application provides an image generation device, which includes an image generation module. The image generation module is configured to generate an image based on Gaussian noise by using a target image generation model to obtain a target image embedded with target watermark information, and the target image generation model is trained according to the training device of the image generation model described above.
[0016] In one possible implementation, the image generation device further includes: a target text feature extraction module, configured to extract features of a target text by using a target text feature extraction network to obtain target information features. The image generation module is further configured to perform an inverse diffusion process on a Gaussian noise image based on the target information features by using the target image generation model to obtain a target image embedded with target watermark information.
[0017] In one possible implementation, the image generation device further includes a target watermark extraction module, configured to extract a watermark from the target image by using a trained watermark extraction network to obtain target watermark information.
[0018] In a fifth aspect, an embodiment of the present application provides an electronic device, including one or more processors, a memory, and one or more programs; the one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.
[0019] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is run by a processor, the above method is executed.
[0020] In a seventh aspect, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device obtains the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the above method.
[0021] A training method for an image generation model, an image generation method, and an apparatus provided by an embodiment of the present application. The method includes: obtaining a first sample image and target watermark information; performing image reconstruction on the first sample image through an initial image generation model to obtain a candidate image embedded with the watermark information; using a trained watermark extraction network to extract the watermark from the candidate image to obtain predicted watermark information; determining a first model loss based on the target watermark information and the predicted watermark information; and iteratively training the initial image generation model based on the first model loss to obtain a target image generation model, where the target image generation model is used to generate an image embedded with the target watermark information. By adopting the above method, after obtaining the first model loss based on the predicted watermark information and the target watermark information, and iteratively training the initial image generation model based on the first model loss, the target image generation model obtained can make the generated image carry the target watermark information when generating an image. Therefore, even when the image generation model or the image generated by the image generation model is leaked, the target watermark information in the generated image can be used to prove copyright. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 FIG. shows an application scenario diagram of a training method for an image generation model provided by an embodiment of the present application;
[0024] Figure 2 FIG. shows a schematic flowchart of a training method for an image generation model proposed by an embodiment of the present application;
[0025] Figure 3 FIG. shows a schematic diagram of the model structure of an image generation model proposed by an embodiment of the present application;
[0026] Figure 4 FIG. shows a training schematic diagram in a fine-tuning stage provided by an embodiment of the present application;
[0027] Figure 5 FIG. shows another schematic flowchart of a training method for an image generation model proposed by an embodiment of the present application;
[0028] Figure 6 FIG. shows a training schematic diagram of an initial watermark extraction network provided by an embodiment of the present application;
[0029] Figure 7 FIG. shows a schematic flowchart of an image generation method provided by an embodiment of the present application;
[0030] Figure 8 It shows another flowchart of an image generation method provided by an embodiment of the present application;
[0031] Figure 9 It shows a connection block diagram of a training device for another image generation model provided by an embodiment of the present application;
[0032] Figure 10 It shows a connection block diagram of an image generation device proposed by an embodiment of the present application;
[0033] Figure 11 It shows a structural block diagram of an electronic device for executing the method of an embodiment of the present application. Detailed implementation manners
[0034] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the reference examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0035] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0036] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0037] The flowcharts shown in the accompanying drawings are merely illustrative and not necessarily include all the contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0038] It should be noted that: "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0039] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields and plays an increasingly important role.
[0040] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Taking the application of artificial intelligence in machine learning as an example for illustration:
[0041] Among them, Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or realizes human learning behavior to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent. Its applications cover all fields of artificial intelligence. The solution of this application mainly uses machine learning for image generation.
[0042] In related technologies, the popularity of Artificial Intelligence Generated Content (AIGC) remains high. Phenomenal applications represented by ChatGPT and StableDiffusion have been strongly pursued. There is no doubt that AIGC is a powerful production tool. AIGC applications such as ChatGPT, StableDiffusion, and Midjourney have greatly liberated the productivity of multiple industries and accelerated the pace of informatization construction. At the same time, the copyright issue of AIGC has also become a focus of controversy. In related technologies, the method adopted to clarify the copyright of AIGC is to add invisible watermarks to the content generated by AIGC.
[0043] The inventor has found through research that by using the method in related technologies, once the model leaks, or the content generated during the process of adding watermarks leaks, it is impossible to add invisible watermarks to the generated content, and thus it is impossible to claim copyright for the generated content.
[0044] Based on this, the embodiments of the present application provide a method for training an image generation model. The method includes obtaining a first sample image and target watermark information; reconstructing the first sample image through an initial image generation model to obtain a candidate image embedded with the watermark information; extracting the watermark from the candidate image by using a trained watermark extraction network to obtain predicted watermark information; determining a first model loss based on the target watermark information and the predicted watermark information; and iteratively training the initial image generation model based on the first model loss to obtain a target image generation model, where the target image generation model is used to generate an image embedded with the target watermark information. By adopting the above method, the trained target image generation model can make the generated image carry the target watermark information when generating an image. Therefore, even when the image generation model or the image generated by the image generation model is leaked, the target watermark information in the generated image can be used to prove the copyright.
[0045] The following describes an exemplary application of the method for training an image generation model provided in the application. Specifically, the method can be applied to a server in an application environment as Figure 1 shown.
[0046] Figure 1 FIG. is a schematic diagram of an application scenario shown according to an embodiment of the present application. As Figure 1 shown, the application scenario includes a terminal device 10 and a server 20 communicatively connected to the terminal device 10 through a network.
[0047] The terminal device 10 may specifically be a device such as a mobile phone, a computer, a tablet computer, a vehicle-mounted terminal, or a smart TV that can interact with a user. The terminal device 10 may run a client for presenting data (such as presenting a generated image).
[0048] The network may be a wide area network or a local area network, or a combination of the two.
[0049] The server 20 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0050] If, for example, Figure 1The terminal device 10 and the server 20 in it perform the training of the image generation model. The terminal device 10 can send the first sample image and the target watermark information to the server 20. After obtaining the first sample image and the target watermark information, the server 20 performs image reconstruction on the first sample image through the initial image generation model to obtain a candidate image embedded with the watermark information; uses the trained watermark extraction network to extract the watermark from the candidate image to obtain the predicted watermark information; determines the first model loss based on the target watermark information and the predicted watermark information; and performs iterative training on the initial image generation model based on the first model loss to obtain the target image generation model, which is used to generate an image embedded with the target watermark information. After the server 20 generates an image embedded with the target watermark information using the target image generation model, it can send the generated image embedded with the target watermark information to the terminal device 10.
[0051] By using the training method of the image generation model of the present application, after obtaining the first model loss based on the predicted watermark information and the target watermark information, and performing iterative training on the initial image generation model based on the first model loss, when the obtained target image generation model generates an image, the generated image can carry the target watermark information. Therefore, even when the image generation model or the image generated by the image generation model is leaked, the target watermark information in the generated image can be used to prove the copyright.
[0052] The following will specifically describe the embodiments of the present application in conjunction with the drawings.
[0053] Please refer to Figure 2 , Figure 2 As shown, the present application also provides an image generation method, which can be applied to an electronic device. The electronic device can be the above-mentioned terminal device or server. The method includes:
[0054] Step S110: Obtain a first sample image and target watermark information.
[0055] Among them, the first sample image can be any image, such as an image including a target object, such as an image including objects such as "flowers", "birds", "insects", "fish", "animals", "plants" or "comic characters", or it can also be a landscape image, a color image or a texture image, etc.
[0056] The way to obtain the first sample image can be an image obtained from a website, an image uploaded by a user taking a photo, or an image stored in any database.
[0057] The target watermark can refer to any watermark that needs to be added, which can be a number, text, image, or other form of identifier. In this application, the target watermark specifically refers to a hidden watermark. Among them, a hidden watermark (also known as a digital watermark or steganography) is a technology that embeds additional information (such as copyright information, owner identification, etc.) into digital media (such as images, audio, video, or documents). These embedded information is usually invisible or difficult to detect so as not to affect the visual quality of the original media. The main purpose of adding a hidden watermark to an image is to provide a robust and difficult-to-delete identifier without disturbing the original content.
[0058] In an implementable manner of this application, the target watermark is a string, and the length of this string can be arbitrary. For example, it can be a 16-bit string watermark.
[0059] The above-mentioned target watermark can be represented by a binary string. Specifically, when the target watermark is a 16-bit string watermark m, when it is identified by a binary string, it is specifically represented by m_bits = M(m), indicating this binary string identification, and m_bits = (m_bits_1,..., m_bits_i..., m_bits_128), where m_i ∈ {0, 1}.
[0060] Step S120: Use the initial image generation model to perform image reconstruction on the first sample image to obtain a candidate image embedded with watermark information.
[0061] Among them, the initial image generation model can be any model that can achieve image reconstruction, such as any one of a diffusion model, a generative adversarial network model, and an autoregressive model, etc.
[0062] Among them, the diffusion model is a type of generative model that converts Gaussian noise into samples of a known data distribution through an iterative denoising process. The generated pictures have good diversity and realism. The diffusion process gradually adds Gaussian noise to the original image, which is a fixed Markov chain process, and finally the image is also gradually transformed into a Gaussian noise. The reverse process then restores the original image step by step through denoising, thereby realizing the generation of the image.
[0063] The autoregressive model, which utilizes its powerful attention mechanism, has become a paradigm for sequence-related modeling. Inspired by the success of the GPT model in natural language modeling, Image GPT (iGPT) performs autoregressive image generation using a Transformer by treating the flattened image sequence as discrete tokens. The rationality of the generated images indicates that the Transformer model can simulate the spatial relationships between pixels and high-level attributes (texture, semantics, and scale). The Transformer mainly consists of two major parts: the Encoder and the Decoder, which use the multi-head self-attention mechanism for encoding and decoding.
[0064] The generative adversarial network model consists of a generative model and a discriminative model. Among them, the generative model is responsible for capturing the distribution of sample data, and the discriminative model is generally a binary classifier that discriminates whether the input is real data or a generated sample. The entire training process is a continuous game and optimization between the two. The generator continuously generates an image distribution that increasingly approaches the real image distribution to deceive the discriminator and improve the discriminator's discrimination ability. The discriminator discriminates between real images and generated images to improve the generator's generation ability. The implementation of text-to-image generation using generative adversarial networks mainly consists of three major parts: a text encoder, a generator, and a discriminator. The text encoder is composed of an RNN or a Bi-LSTM. The generator can be made into a stacked structure or a single-stage generation structure, mainly used to generate images based on text information semantics. The discriminator is used to determine whether the images generated by the generator are real and conform to the text semantics.
[0065] It should be understood that the types of the above initial generation models for image reconstruction are only illustrative. The specific types and model structures of the above initial image generation models are not specifically limited here, as long as they can perform image reconstruction based on existing images and the reconstructed images can embed watermark information.
[0066] Step S130: Use the trained watermark extraction network to extract the watermark from the candidate image to obtain predicted watermark information.
[0067] Among them, the above-trained watermark extraction network can be pre-trained. For example, it can be obtained by training the initial watermark extraction network using a second sample image embedded with the first watermark information.
[0068] The above watermark extraction network can be a convolutional neural network (CNN) or a recurrent neural network (RNN), and is a network trained using labeled data on a large-scale dataset.
[0069] Step S140: Determine the first model loss based on the target watermark information and the predicted watermark information.
[0070] The above step S140 may be to calculate a loss based on the target watermark information and the predicted watermark information using a loss function to obtain a first model loss.
[0071] The above loss function may be at least one of a cross-entropy loss function, an L1 norm loss function, a KL divergence loss function, and a mean squared error loss, etc., and can be set according to actual needs.
[0072] Considering that in the image reconstruction process, the model parameters will be adjusted based on the first model loss to iteratively train the initial image generation model. To avoid affecting the accuracy of image reconstruction after parameter tuning, the above step S140 may also be: determining the first model loss based on the target watermark information, the predicted watermark information, the first sample image, and the candidate image.
[0073] In this way, specifically, a watermark loss can be calculated for the target watermark information and the predicted watermark information to obtain a watermark loss; an image loss can be calculated for the first sample image and the candidate image to obtain an image loss; and the first model loss is determined based on the watermark loss and the image loss.
[0074] Among them, when calculating the watermark loss and the image loss, the loss functions used may be the same or different. For example, the image loss can be calculated using an L2 loss function, an SSIM loss function, an MS-SSIM loss function, or a mixed loss function composed of the foregoing loss functions, and the watermark loss can be calculated using at least one of a cross-entropy loss function, an L1 norm loss function, a KL divergence loss function, and a mean squared error loss.
[0075] After obtaining the watermark loss and the image loss, the first model loss can be obtained by weighted summing the image loss and the watermark loss, or the one with the largest loss value among the image loss and the watermark loss can be used as the first model loss, which can be set according to actual needs.
[0076] In an implementable manner of the present application, a weight adjustment coefficient can be added to the image loss. The weight adjustment coefficient can change with the number of iterations. Correspondingly, the finally obtained first model loss is L, where L = L_m + λi * L_i, and λi is a weight adjustment function used to control the emphasis in the loss function. During optimization, the weight adjustment coefficient λi can change with the number of iterations. For example, the weight adjustment coefficient λi is positively or negatively correlated with the number of iterations, so that a balance can be maintained between the watermark loss L_m and the image loss L_i during the iterative training process.
[0077] Step S150: Based on the first model loss, iteratively train the initial image generation model to obtain a target image generation model, which is used to generate an image embedded with target watermark information.
[0078] When iterating the initial model, all model parameters of the initial model can be iterated, or some parameters in the initial model can be iterated (i.e., fine-tuning training), which can be set according to actual needs.
[0079] Fine-tuning training refers to the process of training a pre-trained model with a training dataset corresponding to the current task and adapting the model's parameters to the training dataset corresponding to the current task. In this application, the current task refers to the task of generating an image including a target watermark.
[0080] Specifically, the above fine-tuning training can be as follows: after obtaining the pre-trained model, load the pre-trained model into memory and make certain modifications, such as modifying the feature input and output dimensions of some networks or adding some model parameters (e.g., fine-tuning network layers) to obtain the initial image generation model. The LORA fine-tuning method can also be used to add a bypass in the pre-trained model to perform a dimensionality reduction and then dimensionality increase operation to achieve the rank decomposition matrix for the change of the dense layer in the optimization adaptation process, indirectly training some dense layers in the neural network while keeping the weights of the pre-trained model unchanged. Train the modified pre-trained model with a new dataset (i.e., the dataset related to the task of generating an image with a target watermark: the first sample image and the target watermark information). Usually, it is trained with a relatively small learning rate and in a small number of iterations, and better performance can be obtained (i.e., it can generate an image including a watermark).
[0081] Among them, the pre-training model (PTM), also known as the foundation model or large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on a large amount of unlabeled data, and the function approximation ability of the large-parameter DNN is used to enable the PTM to extract common features from the data. Subsequently, through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it can be applied to downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modalities they process. Among them, multi-modal models refer to models that establish feature representations of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence-generated content (AIGC) and can also be used as a general interface connecting multiple specific task models.
[0082] In an implementable manner of this application, the initial image generation model includes a pre-trained first image generation model and a fine-tuning network layer inserted into the first image generation model; the above step S150 may be to freeze the parameters of the first image generation model; based on the first model loss, adjust the parameters of the fine-tuning network layer to obtain the target image generation model.
[0083] It should be understood that after freezing the parameters of the first image generation model, the parameters of the first image generation model will not be adjusted during the parameter tuning process of the initial image generation model. In addition, by inserting a fine-tuning network layer into the pre-training model and using the generation of watermarked images as the training task, it is possible to fine-tune and train the first image generation model using few samples (target watermark information and a small amount of first training samples), and finally obtain the target image generation model, which is used to generate images embedded with the target watermark information.
[0084] By adopting the above training method of the image generation model of this application, after obtaining the first model loss based on the predicted watermark information and the target watermark information, and iteratively training the initial image generation model based on the first model loss, the generated image of the obtained target image generation model can carry the target watermark information when generating images. Therefore, even when the image generation model or the images generated by the image generation model are leaked, the target watermark information in the generated images can be used to prove copyright.
[0085] In an implementable manner of the present application, the first image generation model includes a first image encoder, a latent diffusion network, and a first image decoder. Then, the above step S120 may include step S120a, step S120b, step S120c, and step S120d (not shown in the figure).
[0086] Step S120a: Encode the first sample image by the first image encoder to obtain a sample encoded image.
[0087] Among them, the first image encoder is mainly used to map the first sample image to the latent space, so that the latent diffusion network performs subsequent processing on the first sample image mapped to the latent space.
[0088] Step S120b: Perform multiple diffusion processes on the sample encoded image by the latent diffusion network to obtain multiple sample diffusion images, and the multiple sample diffusion images include the Gaussian noise image generated in the last diffusion.
[0089] Specifically, when performing multiple diffusion processes on the sample encoded image by the latent diffusion network, multiple noise addition methods can be adopted, and the image after each noise addition is used as the base image for the next noise addition process. By performing multiple noise addition processes on the sample image until a Gaussian noise image is obtained, the noise addition process is stopped, and the obtained noise sample images can include the noise images after each noise addition process.
[0090] Step S120c: Perform inverse diffusion processing on the Gaussian noise image by the latent diffusion network to obtain a target denoised sample image.
[0091] Among them, the number of times of performing inverse diffusion processing based on the Gaussian noise image is the same as the number of times of performing diffusion processing. That is, the number of times of performing noise addition processing in the latent diffusion network is the same as the number of times of performing denoising processing.
[0092] Step S120d: Decode the target denoised sample image by the first image decoder to obtain a candidate image embedded with watermark information.
[0093] Among them, the operation performed by the first image decoder is the inverse process of the operation performed by the first image encoder. That is, the first image decoder is mainly used to map the target denoised sample image in the latent space output by the latent diffusion network to the image space.
[0094] In the above step S140, the method for determining the image loss can also be: the number of times of performing diffusion processing in the latent diffusion network is the same as the number of times of performing inverse diffusion processing. For example, when both are T times, the model loss can be calculated based on the results after at least one diffusion processing, the results obtained from the inverse diffusion processing corresponding to this diffusion processing, as well as the first sample image and the candidate image. Among them, when the current number of diffusion processing times is t times, the inverse diffusion processing corresponding to the t-th diffusion processing is the (T - t)-th inverse diffusion processing.
[0095] It can also be that the number of times of performing diffusion processing in the latent diffusion network is the same as the number of times of performing inverse diffusion processing. For example, when both are T times, the model loss can be calculated based on the noise added during at least one diffusion processing, the noise approximated by fitting during the inverse diffusion processing corresponding to this diffusion processing, as well as the first sample image and the candidate image. Among them, when the current number of diffusion processing times is t times, the inverse diffusion processing corresponding to the t-th diffusion processing is the (T - t)-th inverse diffusion processing.
[0096] By adopting the above steps S130a - S130d, it can be realized that during the training process, multiple noise additions are adopted in the forward process, and the image noise is eliminated by multiple noise reductions in the backward process, so that the noise added during each noise addition process is close to the noise reduced by the corresponding noise reduction, which can make the image obtained in the backward process have good fidelity.
[0097] As Figure 3 shown, it is a schematic diagram of the model structure of an image generation model. x refers to the first sample image. In the forward process, the image feature vector z is obtained by extracting features of x using the first image encoder ε. The latent diffusion network is used to add noise to the image feature vector multiple times (such as T times). Among them, the sample noise image obtained after the i-th noise addition is Ti. Finally, the sample noise image obtained after the T-th noise addition is z t , where z t is Gaussian noise;
[0098] In the backward process, when the latent diffusion network restores the noise to a sample image, the latent diffusion network depends on the conditional guidance generation during the diffusion process. Specifically, it is to extract the target sample information feature τ from the description sample line features of the first sample image θIt is optimized and trained together as a condition in the estimation of the diffusion model noise. It should be noted that if the current number of noise addition processes is t times, the corresponding noise reduction process for the t-th noise addition process is the (T - t)-th noise reduction process. After the optimization of the model parameters, the noise added corresponding to the t-th noise addition process tends to be consistent with the approximately fitted noise obtained corresponding to the (T - t)-th noise reduction process. Finally, the sample noise image obtained after the T-th noise reduction is z'. Then, the first image decoder D is used to decode the sample noise image z' to obtain the candidate image x embedded with the watermark information t 。
[0099] Specifically, after the sample text is encoded, the feature encoding is extracted, and then it interacts with the noise image z used to generate the image through the cross-modal attention mechanism t When interacting. Its specific expression is as follows:
[0100] Q = W Q φ i (z t ), K = W K τ θ (y), V = W V τ θ (y)
[0101] The above formulas used all come from the basic attention mechanism. When performing the t-th noise addition process in the diffusion model, the encoded noise image is used as the query (Q), and the target sample information features are used as the key (K) and value (V) to calculate the weights and obtain the weighted result, so that the target sample information features extracted from the description sample text can be referred to when generating the image, and the attention weights are obtained Attention(Q, K, V) is the output, which can be used as the input in the (t + 1)-th noise reduction process. Finally, the latent diffusion model outputs z' after the T-th noise reduction process. Decoding z' using the first image decoder D can obtain the final image (the candidate image embedded with the watermark information); W Q 、W K 、W V are weight matrices and can be determined through training. φ i (z t ) represents the encoded noise image
[0102] Such as Figure 4As shown, when fine-tuning and training an image generation model, the "Pre-trained weights" on the left refer to the original parameter weights W determined by pre-training within the Transformer block in any diffusion network, which can be obtained through training on a large number of open-source and publicly available datasets. The two matrices on the right (matrix A for dimensionality reduction and matrix B for dimensionality increase) are the parameter matrices (i.e., fine-tuning network layers) that need to be newly introduced and actually participate in the fine-tuning process during LoRA fine-tuning. Among them, the dimensionality reduction matrix is denoted as A, and the dimensionality reduction matrix A can be initialized using a random Gaussian distribution. It is responsible for mapping the original input x with dimension d to r dimensions (for example, r is 1 or 4). The dimensionality increase matrix is denoted as B, and the dimensionality increase matrix B can be initialized using a zero matrix. It is responsible for raising the r-dimensional intermediate result back to h dimensions, so that the dimensions of the input and output of each matrix W are consistent with the dimensions of the features extracted by the original text feature extraction network. Specifically, expressed by the formula: the output corresponding to the original pre-trained weight (initial weight) W: h = Wx, W ∈ R d×d , and the output after fine-tuning (i.e., using the sample to train matrices A and B) is: h = Wx + BAx, W ∈ R d×d , A ∈ R d×r , B ∈ R r×d . Since the compressed dimension r can be regarded as the rank of the original matrix, by setting an extremely small r (such as setting r to 4 or 9, etc.), the overall number of parameters can be controlled at a very small level. h is the output of a certain layer in the network layer, x is the input of this layer, W is the model parameter, and the training stage is to train A and B while training W to achieve the purpose of fine-tuning training, so that the fine-tuning can be carried out at as low a cost as possible.
[0103] To improve the accuracy and iteration efficiency of the images generated by the image generation model, in an implementable manner of this application, the first sample image corresponds to sample description text, and the method further includes: using a text feature extraction network to extract features from the sample description text to obtain sample information features. The above step S130c: can also be using a latent diffusion network to perform inverse diffusion processing on the Gaussian noise image based on the sample information features to obtain a target denoised sample image.
[0104] Among them, the sample description text can be text input by the user based on the image for describing the first sample image, or text obtained by text expansion according to the category of the object in the first sample image and the description words of the object for describing the first sample image.
[0105] Among them, the object in the first sample image can specifically be a certain object, and the descriptor of the object can be a word used to describe the attributes and / or states of the object, such as adjectives describing the color, shape, volume, actions, etc. of the object. For example, if the target object is a ginkgo tree, the modifiers of the ginkgo tree can be words used to describe the morphological characteristics of the ginkgo tree (such as leaves, seeds, branches, etc.), growth environment, growth habits, etc. The modifiers of the ginkgo tree are, for example, fan-shaped leaves, tall, grayish-brown, (crown) conical, light green (color of leaves in spring and summer), yellow (color in autumn), street tree, falling, swaying, and so on. There can be various ways to obtain the descriptors of the object as described above. For example, the descriptors of the object can be obtained based on the pre-stored correspondence between the object and the descriptors; it can also be receiving the descriptors of the object input by the user; it can also be based on the category corresponding to the object, obtaining the descriptors corresponding to the category from the dataset. The dataset can be any dataset that includes different object categories and the corresponding description information for each object category. For example, it can be the Visual Genome (VG) dataset or the ImageNet image annotation dataset, etc.
[0106] By utilizing the forward process for diffusion in the latent diffusion network (during the diffusion processing); the image added with noise is sent into the denoising network in the diffusion network for feature learning. That is, the target sample information features obtained by mapping the description sample text are input into the backward process of diffusion, so that the denoising network in the diffusion network guides its learning of noise through semantic guiding features, and an image generated by text guidance is obtained. Obviously, the efficiency of generating an image based on text guidance is significantly higher than directly performing inverse diffusion to generate an image. And when generating an image based on text guidance, the cosine similarity between the generated image and the first sample image can be used to adjust the intermediate representation of the model until the training ends when the model converges.
[0107] Please refer to Figure 5 As shown, the embodiment of the present application also provides a training method for an image generation model. This method can be applied to an electronic device, and the electronic device can be a terminal device or a server. The method includes:
[0108] Step S210: Use the initial watermark extraction network to extract the watermark from the second sample image embedded with the first watermark information to obtain the sample watermark information.
[0109] Among them, the second sample image can be any image embedded with the first watermark information, and the first watermark information embedded in different second sample information is different. The difference referred to here can be at least one of different watermark lengths, different watermark types, and different corresponding character strings.
[0110] In an implementable manner of the present application, the above-mentioned different watermark information means that the watermark lengths and types are the same, but the corresponding string information is different. Exemplarily, the first watermark information embedded in the second sample image is watermark information of a 16-bit string. When performing embedding and detection, the watermark information of the 16-bit string can be encoded into a 128-bit binary code and then embedded.
[0111] Specifically, the watermark string can be converted into ASCII codes: each character in the string is converted into the corresponding ASCII code, and then each ASCII code is converted into binary data and concatenated to obtain a 128-bit binary code.
[0112] Exemplarily, taking only the string "AB" as an example, the ASCII codes corresponding to the string "AB" are 65 (A) and 66 (B) respectively. Convert the ASCII codes into binary data: each ASCII code is converted into 8-bit binary data. For example, the binary representation of 65 (A) is 01000001, and the binary representation of 66 (B) is 01000010. Concatenate the binary data: the binary data corresponding to all characters are concatenated in order to form a bit stream. For example, the bit stream corresponding to the string "AB" is 0100000101000010.
[0113] Step S220: Obtain a second model loss based on the sample watermark information and the first watermark information.
[0114] The above step S220 may be to calculate the loss based on the sample watermark information and the first watermark information by using a loss function to obtain the second model loss. Among them, the loss function used here may be at least one of a cross-entropy loss function, an L1 norm loss function, a KL divergence loss function, and a mean square error loss, etc., and can be set according to actual needs.
[0115] Step S230: Iteratively train the initial watermark extraction network based on the second model loss to obtain a trained watermark extraction network.
[0116] Among them, when the number of iterative training times reaches a preset number, or the obtained second model loss is less than a preset watermark loss threshold, a trained watermark extraction network is obtained.
[0117] Step S240: Obtain a first sample image and target watermark information.
[0118] Step S250: Perform image reconstruction on the first sample image through the initial image generation model to obtain a candidate image embedded with watermark information.
[0119] Step S260: Use the trained watermark extraction network to extract the watermark from the candidate image to obtain predicted watermark information.
[0120] Step S270: Determine a first model loss based on the target watermark information and the predicted watermark information.
[0121] Step S280: Iteratively train the initial image generation model based on the first model loss to obtain a target image generation model, where the target image generation model is used to generate an image embedded with the target watermark information.
[0122] Specific descriptions of the above steps S240 - S280 can refer to the specific descriptions of steps S110 - S150 in the foregoing, and will not be elaborated herein one by one in the embodiments of this application.
[0123] To improve the accuracy of the watermark extraction network, the second sample image can be an image obtained after embedding watermark information using the watermark embedding network. In this implementation manner, the second sample image can be obtained through the following method:
[0124] Step S210a: Embed the first watermark information into the initial image using the initial watermark embedding network to obtain a first image.
[0125] Step S210b: Add noise to the first image using the noise addition network to obtain the second sample image.
[0126] The noise added to the first image using the noise addition network can be random noise, Gaussian noise, or can be set according to actual requirements for the first image using the noise addition network.
[0127] The above step S230 can also be: Adjust the parameters of the initial watermark extraction network, the initial watermark embedding network, and the noise addition network based on the second model loss until the training end condition is reached. After reaching the training end condition, use the initial watermark extraction network as the trained watermark extraction network.
[0128] Please combine Figure 6, by adopting the above process, in one iteration process, the watermark can be embedded into the initial image by using the initial watermark embedding network to obtain the first image, and after adding noise by using the noise addition network to obtain the second sample image, the initial watermark extraction network is used to extract the watermark from the second sample image, so as to calculate the loss based on the first watermark information and the extracted sample watermark information to obtain the second model loss, and based on the second model loss, the parameters of the initial watermark extraction network, the initial watermark embedding network, and the noise addition network are adjusted. Thus, in the iterative training process, the accuracy of watermark extraction from images with noise can be gradually improved. Finally, when adaptively extracting the candidate image obtained during the reconstruction of the aforementioned image by using the trained watermark extraction network, even if the candidate image is an image with noise (an image with poor reconstruction effect), when using the trained noise extraction network to extract the watermark from the image with noise, the accuracy of the watermark extracted from the image with noise can be guaranteed.
[0129] In this implementation manner, in order to reduce the image loss caused after embedding the first watermark information into the initial image, the initial watermark embedding network includes an image encoding network. Step S210a: Specifically, it can be that the image encoding network performs encoding processing based on the first watermark information and the initial image to obtain an encoded image; performs non-linear transformation on the encoded image to obtain the transformed encoded image; performs scaling on the transformed encoded image to obtain a scaled image; and fuses the scaled image with the initial image to obtain the first image.
[0130] Among them, the specific manner of the image encoding network performing encoding processing based on the first watermark information and the initial image can be that the image encoding network encodes the first watermark information based on the size of the initial image to obtain a watermark mask that matches the size of the initial image and carries the encoded first watermark information. Or, it can be that the image with the same size as the initial image is transferred from the spatial domain to the transform domain to obtain a frequency-domain image, the frequency-domain image is connected to the watermark, and then, the connected image is subjected to convolution processing to obtain the convolved image, and the convolved image is converted back to the spatial domain to obtain the encoded image.
[0131] After obtaining the encoded image, a non-linear transformation function such as the tanh function or the Sigmoid function can be used to perform non-linear transformation processing on the encoded image to obtain the transformed encoded image. After that, a scaling coefficient can also be used to scale the transformed encoded image.
[0132] The above manner of fusing the scaled image with the initial image can be to fuse the scaled image with the initial image by using at least one fusion method such as alpha fusion, pyramid fusion, and Poisson fusion.
[0133] The trained watermark extraction network obtained by adopting the above method of the present application, compared with HiDDeN in the related art: using a deep network to hide a data model, the present application does not need to pay attention to the quality of the second sample image and the watermark trace, so the adversarial network is removed, and the loss corresponding to the adversarial network is cancelled for parameter tuning. The network structure adopted in the training process is simplified, thereby improving the training speed. In addition, by adding a non-linear function to the encoded image output by the image encoding network and constraining the distortion through a scaling factor, the image loss generated when embedding the first watermark information into the initial image to obtain the second sample image is reduced, thereby ensuring the accuracy of the subsequent watermark extraction network trained based on the second sample image.
[0134] Please refer to Figure 7 , the embodiment of the present application further provides an image generation method, which can be applied to an electronic device, and the electronic device can be a terminal device or a server. The method includes:
[0135] Step S310: Use a target image generation model to generate an image based on Gaussian noise to obtain a target image embedded with target watermark information.
[0136] Among them, the target image generation model is trained according to the training method of the image generation model in the foregoing embodiment.
[0137] Regarding the training process of the target image generation model, reference can be made to the specific descriptions of steps S110-S150 and steps S210-S290 above, and details will not be repeated in this embodiment.
[0138] Regarding the process of using the target image generation model to generate an image based on Gaussian noise to obtain a target image embedded with target watermark information, reference can be made to the specific description of the foregoing image reconstruction process.
[0139] In an implementable manner, if the target image generation network includes a latent diffusion network and a first decoder, the latent diffusion network can be used to perform inverse diffusion processing on the Gaussian noise image and then decode to obtain the target image, and the target watermark information is embedded in the target image. Among them, regarding the process of inverse diffusion processing, reference can be made to the specific description in the foregoing embodiment, and details will not be repeated in this embodiment.
[0140] As Figure 8 shown, if it is necessary to extract the watermark from the foregoing generated image for copyright confirmation or declaration, etc., after executing the above step S310, the method further includes:
[0141] Step S320: Use the trained watermark extraction network to extract the watermark from the target image to obtain the target watermark information.
[0142] Among them, the training process of the trained watermark extraction network can refer to the specific descriptions of steps S210 - S230 in the foregoing embodiments, which will not be elaborated herein one by one. Regarding the watermark extraction process of the trained watermark extraction network, it can also refer to the specific descriptions of the foregoing embodiments, which will not be elaborated herein one by one.
[0143] In an implementable manner, if an image including a target object or for describing a target scene is to be generated, in an implementable manner of the present application, the target image generation model includes a target text feature extraction network. Before performing step S310, the method further includes: using the target text feature extraction network to extract features from the target text to obtain target information features.
[0144] Among them, the target text refers to the text used for user input to describe the image to be generated, which can be the text composed of the category and description words of the object included in the image. It can also be the text obtained by expanding the text based on the category of the object and the description words of the object, and can be set according to actual needs.
[0145] In this implementation manner, step S310 can specifically be: using the target image generation model to perform inverse diffusion processing on the Gaussian noise image based on the target information features to obtain a target image embedded with the target watermark information.
[0146] The embodiment of the present application provides an image generation method, which specifically includes the following stages: watermark encoding stage, watermark extraction network training stage, image generation model training stage, and image generation stage. Taking the embedding and detection of watermark information (target watermark information or first watermark information) of a 16 - bit string as an example for description, it should be understood that the present solution can also perform the embedding and detection of watermarks of other lengths of strings, such as 12 - bit, 15 - bit, etc.
[0147] I. Watermark encoding stage:
[0148] The 16 - bit string watermark information m (specifically, the target watermark information is m0 and the first watermark information is m1) is converted into 128 - bit watermark encoding information m_bits (encoded target watermark information m0_bits and encoded first watermark information m1_bits) through the encoder M, and m_bits = M(m). Specifically, each character in the watermark string can be converted into the corresponding ASCII code. Each ASCII code is converted into 8 - bit binary data. The binary data corresponding to each ASCII code is concatenated according to the position of each ASCII code in the watermark string to obtain the encoded target watermark information and the encoded first watermark information.
[0149] Exemplarily, taking the string "AB" as an example, the corresponding ASCII codes of "AB" are 65 (A) and 66 (B) respectively. The binary representation of 65 (A) is 01000001, and the binary representation of 66 (B) is 01000010. The bitstream corresponding to the string "AB" is 0100000101000010.
[0150] II. Watermark extraction network training stage
[0151] The structure of the watermark extraction network adopted in this training stage can refer to the network structure in the HiDDeN method, and the size of the embedded information is selected to be 128 bits. However, in order to keep the network structure as simple as possible while achieving the highest watermark extraction effect, the watermark extraction network in the embodiments of the present application can be embedded into a watermark processing model for training during the iterative training process. Among them, the model structure of the watermark processing model has the following two changes on the premise of being the same as HiDDeN: removing the adversarial network and the image reconstruction loss in HiDDeN. Since the watermark encoder will be discarded after training, there is no need to pay attention to the quality of the embedded image and the watermark trace. Therefore, the watermark processing model W embedded by the watermark extraction network specifically includes the following networks: a watermark embedding network W-E, a noise addition network W-N, and a watermark extraction network W-D. In addition, in order to control the image loss, a tanh function is added to the output of the watermark embedding network E, and the watermark distortion is constrained by a scaling factor a. This can improve the recovery accuracy of the watermark. The specific processing process is as follows:
[0152] The image coding network performs coding processing on the first watermark information and the initial image to obtain a coded image; the coded image is non-linearly transformed using the tanh function to obtain a transformed coded image; the transformed coded image is scaled based on the scaling factor a to obtain a scaled image; the scaled image is fused with the initial image to obtain a first image. The noise addition network adds noise to the first image to obtain a second sample image.
[0153] Specifically, first, the initial image x and the first watermark information m0_bits after encoding are processed by the watermark embedding network W-E and then sequentially undergo non-linear transformation, scaling, and fusion to obtain the first image x_w, where x_w = W_E(x, m_bits), and W_E(x, m_bits) = a * tanh(E(x, m_bits)) + x. Then, the first image x_w is fed into a noise addition network, which samples an image transformation n from a set W_N that contains common image attack operations (such as cropping, JPEG compression, noise, etc.). A suitable image transformation is selected to make the transformed first watermark information robust against common image processing operations (such as cropping, compression, etc.). The first image (the second sample image) x_w_n after the attack is obtained, that is, x_w_n = Random(W_N(x_w)). Then, the second sample image x_w_n is fed into the watermark extraction network W_D to extract the watermark message, obtaining the extracted sample watermark information m', that is, m' = W_D(x_w_a). Finally, the second model loss L_m0 is calculated, that is, the BCE loss between m0_bits and m'_bits: L_m0 = BCE(m0_bits, m'_bits) = -(m0_bits * log(m'_bits) + (1 - m0_bits) * log(1 - m'_bits)). Taking L_m0 as the optimization objective, the above steps are repeated until L_m0 converges, and the network training is completed. After the training is completed, the trained watermark addition network W_E and the trained noise addition network are discarded, and the trained watermark extraction network W_D is retained to participate in subsequent watermark extraction processing.
[0154] III. Image Generation Model Training Phase
[0155] The initial image generation model adopted in this stage includes a first image encoder, a latent diffusion network, and a first image decoder. The latent diffusion network is embedded with a fine-tuning network layer, and other parameters in the image generation model except for the parameters corresponding to the fine-tuning network layer are frozen. When training the fine-tuning network layer, the first image encoder encodes the first sample image to obtain a sample encoded image. The latent diffusion network performs image reconstruction based on the sample encoded image. Specifically, the sample encoded image can be subjected to multiple diffusion processes to obtain multiple sample diffusion images. The multiple sample diffusion images include the Gaussian noise image generated in the last diffusion. Then, the latent diffusion network performs inverse diffusion processing based on the Gaussian noise image to obtain a target denoised sample image. The first image decoder is used to decode the target denoised sample image to obtain a candidate image embedded with watermark information. Then, the trained watermark extraction network is used to extract the watermark from the candidate image to obtain predicted watermark information. The watermark loss is calculated between the target watermark information and the predicted watermark information. The image loss is calculated between the first sample image and the candidate image. The first model loss is determined based on the watermark loss and the image loss. Finally, the fine-tuning network layer is adjusted based on the first loss, and the above steps are repeated until the iteration end condition is reached, and the training of the initial image generation model is completed to obtain the target image generation model.
[0156] Specifically, the image generation model used in this stage is LDMs (Latent Diffusion Models), which mainly includes a first image encoder LDMs_ε, a latent diffusion network LDMs_Latent, and a first image decoder LDMs_D. The training involved in this scheme is mainly fine-tuning training (LoRA training). During the training process, only the parameters in LDMs_Latent need to be trained as external parameters. The training samples used in this training stage are called LDMs_Latent_LoRA fine-tuning training sets: and the selection of training sets can be selected according to the actual scenario. In this application, the training set used includes a target watermark and multiple first sample images. The specific training stage is: fix all parameters in LDMs_ε, LDMs_D, and LDMs_Latent, and embed the fine-tuning network layer LDMs_Latent_LoRA into the aforementioned LDMs_Latent; then, input the first sample image y into the first image encoder LDMs_ε with fixed parameters, and output the latent space vector (sample encoded image) z, z = LDMs_ε(y), y∈h*w*c, where H refers to the height, that is, the vertical coordinate value, which means starting from the origin (the upper left corner of the picture), the horizontal axis is the horizontal axis to the right, and the vertical axis is the vertical axis. After the coordinate system is established, the value of the vertical axis. W refers to the width, that is, the value of the horizontal axis. C refers to Channel, that is, the number of channels. The most common RGB color space has three channels, namely: R (red), G (green), and B (blue). When storing RGB format images in the computer, C has three values, namely 0, 1, and 2, corresponding to R, G, and B respectively. After obtaining the latent space vector (sample coded image), the latent space vector z is sent to LDMs_Latent and LDMs_Latent_LoRA for hidden space image reconstruction, and the reconstructed latent space vector z' (target denoised sample image) is output as the sum of the two parts, i.e. z'=LDMs_Latent(z)+LDMs_Latent_LoRA(z). Then, the reconstructed latent space vector (target denoised sample image) z' is sent to the first image decoder LDMs_D for visible space image reconstruction, and the reconstructed image (candidate image for embedding watermark information) y_o, i.e. y_o=LDMs_D(z'). Send y_o to the watermark extraction network W_D trained in the watermark extraction network training phase, extract the watermark information m”_bits, and calculate the watermark loss L_m1, that is, the BCE loss between m1_bits and m”_bits: L_m1=BCE(m1_bits,m”_bits)=-(m1_bits*log(m”_bits)+(1-m1_bits)*log(1-m”_bits)). Calculate the image loss L_i based on the first sample image and the candidate image.Among them, the image loss can be calculated using the Watson-VGG perceptual loss function, or the LPIPS loss function or the SSIM loss function can be used for calculation. Finally, based on the watermark loss and the image loss, the first model loss L is obtained, where L = L_m1 + λ * L_i, and λ is a weight adjustment function used to control the emphasis in the loss function. During optimization, λ can change with the number of iterations to balance the watermark loss and the image loss during the iterative process. Finally, with the first model loss L as the optimization objective, only the parameters within LDMs_Latent_LoRA are trained. Repeat the above steps until the image generation model training is completed when L converges, and the target image generation model is obtained.
[0157] It should be noted that in addition to the first image encoder, the latent diffusion network, and the first image decoder, the LDMs model can also include a text feature extraction network, which is used to extract features from the sample description text corresponding to the first sample image to obtain sample information features; subsequently, when the latent diffusion network performs image reconstruction, it specifically performs image reconstruction based on the sample information features.
[0158] IV. Image Generation Stage
[0159] Using the target image generation model trained in the image generation model training stage, an image is generated based on Gaussian noise to obtain a target image embedded with the target watermark information. The trained watermark extraction network obtained in the watermark extraction network training stage is used to extract the watermark from the target image to obtain the target watermark information for subsequent determination of the copyright of the target image based on the target watermark information.
[0160] Specifically, the trained fine-tuning network layer LDMs_Latent_LoRA containing specific stylized content is superimposed with the parameters of the latent diffusion network LDMs_Latent to obtain the trained latent diffusion network LDMs_Latent', where LDMs_Latent' = LDMs_Latent + LDMs_Latent_LoRA; the Gaussian noise n of h*w*c is sent into LDMs_Latent' and the decoder LDMs_D to obtain the target image, which contains the target watermark. The target image is y_o’, y_o’ = LDMs_D(LDMs_Latent'(n)); the target image is sent into the trained watermark extraction network W_D, and the encoded target watermark m1_bits = W_D(y_o’) can be extracted; the watermark bits are mapped back to the watermark string information, and the information extraction is completed, that is, m1 = M'(m1_bits).
[0161] By adopting the processing of the above-mentioned multiple stages, any target image generation model using this solution can generate pictures that all contain the target watermark information, and the target is an invisible watermark, so that the target watermark information can be used to declare copyright, identify whether it is generated by AI, etc. It effectively avoids the problem in the related technology that the generated image leaks or the model leaks, resulting in the inability to declare copyright for the finally generated image.
[0162] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0163] Please refer to Figure 9 , another embodiment of the present application provides a training device 400 for an image generation model. The device 400 includes: a data acquisition module 410, an image reconstruction module 420, a predicted watermark extraction module 430, a first loss determination module 440, and a first model training module 450. The data acquisition module 410 is used to acquire a first sample image and target watermark information; the image reconstruction module 420 is used to perform image reconstruction on the first sample image through an initial image generation model to obtain a candidate image embedded with watermark information; the predicted watermark extraction module 430 is used to extract the watermark from the candidate image by using a trained watermark extraction network to obtain predicted watermark information; the first loss determination module 440 is used to determine a first model loss based on the target watermark information and the predicted watermark information; the first model training module 450 is used to perform iterative training on the initial image generation model based on the first model loss to obtain a target image generation model, and the target image generation model is used to generate an image embedded with the target watermark information.
[0164] In an implementable manner, the initial image generation model includes a pre-trained first image generation model and a fine-tuning network layer inserted in the first image generation model; the first model training module 450 is further used to freeze the parameters of the first image generation model; based on the first model loss, adjust the parameters of the fine-tuning network layer to obtain a target image generation model.
[0165] In one possible implementation, the first loss determination module 440 includes a watermark loss acquisition sub-module, an image loss acquisition sub-module, and a first loss determination sub-module. The watermark loss acquisition sub-module is configured to calculate the watermark loss between the target watermark information and the predicted watermark information to obtain the watermark loss. The image loss acquisition sub-module is configured to calculate the image loss between the first sample image and the candidate image to obtain the image loss. The first loss determination sub-module is configured to determine the first model loss based on the watermark loss and the image loss.
[0166] In one possible implementation, the first image generation model includes a first image encoder, a latent diffusion network, and a first image decoder. The image reconstruction module 420 includes an image encoding sub-module, a diffusion processing sub-module, an inverse diffusion processing sub-module, and an image decoding sub-module. The image encoding sub-module is configured to encode the first sample image using the first image encoder to obtain a sample encoded image. The diffusion processing sub-module is configured to perform multiple diffusion processes on the sample encoded image using the latent diffusion network to obtain multiple sample diffused images, where the multiple sample diffused images include the Gaussian noise image generated by the last diffusion. The inverse diffusion processing sub-module is configured to perform inverse diffusion processing on the Gaussian noise image based on the latent diffusion network to obtain a target denoised sample image. The image decoding sub-module is configured to decode the target denoised sample image using the first image decoder to obtain a candidate image embedded with the watermark information.
[0167] In one possible implementation, the device 400 further includes a sample text feature extraction module, configured to extract features from the sample description text using a text feature extraction network to obtain sample information features. The inverse diffusion processing sub-module is further configured to perform inverse diffusion processing on the Gaussian noise image based on the sample information features using the latent diffusion network to obtain a target denoised sample image.
[0168] In one possible implementation, the device 400 further includes: a sample watermark extraction module, a second loss determination module, and a second model training module. The sample watermark extraction module is configured to extract the watermark from the second sample image embedded with the first watermark information using an initial watermark extraction network to obtain sample watermark information. The second loss determination module is configured to obtain a second model loss based on the sample watermark information and the first watermark information. The second model training module is configured to iteratively train the initial watermark extraction network based on the second model loss to obtain a trained watermark extraction network.
[0169] In one possible implementation, the device 400 further includes: a watermark embedding module and a noise adding module. The watermark embedding module is configured to embed first watermark information into an initial image by using an initial watermark embedding network to obtain a first image. The noise adding module is configured to add noise to the first image by using a noise adding network to obtain a second sample image. The second model training module is further configured to adjust the parameters of the initial watermark extraction network, the initial watermark embedding network, and the noise adding network based on a second model loss until a training end condition is reached. After reaching the training end condition, the initial watermark extraction network is used as the watermark extraction network.
[0170] In one possible implementation, the initial watermark embedding network includes an image encoding network. The watermark embedding module is further configured to perform a non-linear transformation on the encoded image to obtain a transformed encoded image; scale the transformed encoded image to obtain a scaled image; and fuse the scaled image with the initial image to obtain a first image.
[0171] Please refer to Figure 10 As shown, an embodiment of the present application provides an image generation device 500, and the device 500 includes an image generation module 510. The image generation module 510 is configured to generate an image based on Gaussian noise by using a target image generation model to obtain a target image embedded with target watermark information, and the target image generation model is trained according to the training device of the image generation model as described above.
[0172] In one possible implementation, the image generation device 500 further includes: a target text feature extraction module, configured to extract features of a target text by using a target text feature extraction network to obtain target information features. The image generation module is further configured to perform an inverse diffusion process on a Gaussian noise image based on the target information features by using the target image generation model to obtain a target image embedded with target watermark information.
[0173] In one possible implementation, the image generation device 500 further includes a target watermark extraction module 520, configured to extract a watermark from the target image by using the trained watermark extraction network to obtain target watermark information.
[0174] Each module in the above device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the electronic device in the form of hardware or be independent of the processor, or can be stored in the memory in the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above modules. It should be noted that the device embodiments in the present application correspond to the foregoing method embodiments. The specific principles in the device embodiments can refer to the content in the foregoing method embodiments, and will not be elaborated herein.
[0175] Next, in combination with Figure 11Describe an electronic device provided by this application.
[0176] Please refer to Figure 11 , based on the image generation method provided in the above embodiment, the embodiment of this application further provides another electronic device 100 including a processor 102 that can execute the foregoing method. The electronic device 100 can be a server or a terminal device.
[0177] The electronic device 100 further includes a memory 104. Among them, a program that can execute the content in the foregoing embodiment is stored in the memory 104, and the processor 102 can execute the program stored in the memory 104.
[0178] Among them, the processor 102 can include one or more cores for processing data and a message matrix unit. The processor 102 connects various parts within the entire electronic device 100 through various interfaces and lines, and executes various functions of the electronic device 100 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 104, and by calling data stored in the memory 104. Optionally, the processor 102 can be implemented in at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 102 can integrate one or a combination of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 102 and can be implemented separately through a communication chip.
[0179] The memory 104 can include Random Access Memory (RAM) and can also include Read-Only Memory. The memory 104 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 104 can include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for implementing at least one function, instructions for implementing the following various method embodiments, etc. The data storage area can also store data obtained during the use of the electronic device 100 (such as, sample images and target watermark information), etc.
[0180] The electronic device 100 may further include a network module and a screen. The network module is used to receive and send electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices, such as communicating with an audio playback device. The network module may include various existing circuit elements for performing these functions, for example, antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The network module can communicate with various networks such as the Internet, enterprise intranets, wireless networks or communicate with other devices through wireless networks. The above-mentioned wireless networks may include cellular phone networks, wireless local area networks or metropolitan area networks. The screen can display interface content and perform data interaction.
[0181] In some embodiments, the electronic device 100 may further include: a peripheral interface 106 and at least one peripheral device. The processor 102, the memory 104 and the peripheral interface 106 may be connected by a bus or signal lines. Each peripheral device can be connected to the peripheral interface through a bus, signal lines or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency component 108, a positioning component 112, a camera 114, an audio component 116, a display screen 118, and a power supply 122, etc.
[0182] The peripheral interface 106 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 102 and the memory 104. In some embodiments, the processor 102, the memory 104 and the peripheral interface 106 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 102, the memory 104 and the peripheral interface 106 can be implemented on a separate chip or circuit board, and the embodiments of the present application do not limit this.
[0183] The radio frequency component 108 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency component 108 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency component 108 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency component 108 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency component 108 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency component 108 may further include a circuit related to NFC (Near Field Communication), which is not limited in this application.
[0184] The positioning component 112 is used to locate the current geographical location of the electronic device to implement navigation or LBS (Location-Based Service). The positioning component 112 can be a positioning component based on the US GPS (Global Positioning System), the Beidou system, or the Galileo system.
[0185] The camera 114 is used to capture images or videos. Optionally, the camera 114 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the electronic device 100, and the rear camera is disposed on the back of the electronic device 100. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera 114 may further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0186] The audio component 116 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 102 for processing, or input to the radio frequency component 108 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 102 or the radio frequency component 108 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio component 114 may further include a headphone jack.
[0187] The display screen 118 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 118 is a touch display screen, the display screen 118 also has the ability to collect touch signals on the surface or above the surface of the display screen 118. The touch signal may be input to the processor 102 as a control signal for processing. At this time, the display screen 118 may also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 118, which is arranged on the front panel of the electronic device 100; in other embodiments, there may be at least two display screens 118, which are respectively arranged on different surfaces of the electronic device 100 or in a folding design; in still other embodiments, the display screen 118 may be a flexible display screen, which is arranged on the curved surface or the folding surface of the electronic device 100. Even, the display screen 118 may be set as an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 118 may be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0188] The power supply 122 is used to supply power to each component in the electronic device 100. The power supply 122 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 122 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.
[0189] The embodiments of the present application also provide a structural block diagram of a computer-readable storage medium. A computer program is stored in the computer-readable medium, and the computer program can be called by a processor to execute the method described in the above method embodiments.
[0190] The computer-readable storage medium may be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has a storage space for a computer program that executes any of the method steps in the above method. These computer programs can be read from or written into one or more computer program products. The computer programs can be compressed in a suitable form, for example.
[0191] The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the method described in the above various optional implementation manners.
[0192] In addition, it should be noted that in the embodiments of the present application, the acquisition of user-related information requires user permission or consent, and the collection, use, processing, and storage of information need to comply with the regulations of the region where it is located.
[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A training method for an image generation model, characterized in that, The method includes: Obtaining a first sample image and target watermark information; Performing image reconstruction on the first sample image through an initial image generation model to obtain a candidate image embedded with watermark information; Using a trained watermark extraction network to extract watermark from the candidate image to obtain predicted watermark information; Determining a first model loss based on the target watermark information and the predicted watermark information; Based on the first model loss, iteratively training the initial image generation model to obtain a target image generation model, where the target image generation model is used to generate an image embedded with the target watermark information.
2. The method according to claim 1, wherein The initial image generation model includes a pre-trained first image generation model and a fine-tuning network layer inserted in the first image generation model; Based on the first model loss, iteratively training the initial image generation model to obtain a target image generation model, including: Freezing the parameters of the first image generation model; Based on the first model loss, adjusting the parameters of the fine-tuning network layer to obtain a target image generation model.
3. The method according to claim 1, wherein The determining the first model loss based on the target watermark information and the predicted watermark information includes: Performing watermark loss calculation on the target watermark information and the predicted watermark information to obtain a watermark loss; Performing image loss calculation on the first sample image and the candidate image to obtain an image loss; Determining the first model loss based on the watermark loss and the image loss.
4. The method according to claim 1, wherein The first image generation model includes a first image encoder, a latent diffusion network, and a first image decoder; The performing image reconstruction on the first sample image through the initial image generation model to obtain a candidate image embedded with watermark information includes: Encoding the first sample image by the first image encoder to obtain a sample encoded image; Performing multiple diffusion processes on the sample encoded image by the latent diffusion network to obtain multiple sample diffusion images, where the multiple sample diffusion images include a Gaussian noise image generated by the last diffusion; Performing inverse diffusion processing on the Gaussian noise image by the latent diffusion network to obtain a target denoised sample image; Decoding the target denoised sample image by the first image decoder to obtain a candidate image embedded with watermark information.
5. The method according to claim 4, wherein The first sample image corresponds to sample description text, and the method further includes: Extracting features from the sample description text by a text feature extraction network to obtain sample information features; The performing inverse diffusion processing on the Gaussian noise image by the latent diffusion network to obtain a target denoised sample image includes: Performing inverse diffusion processing on the Gaussian noise image by the latent diffusion network based on the sample information features to obtain a target denoised sample image.
6. The method according to claim 1, wherein Before using the trained watermark extraction network to extract watermark from the candidate image to obtain predicted watermark information, the method further includes: Using an initial watermark extraction network to extract watermark from a second sample image embedded with first watermark information to obtain sample watermark information; Obtaining a second model loss based on the sample watermark information and the first watermark information; Iteratively train the initial watermark extraction network based on the second model loss to obtain the trained watermark extraction network.
7. The method according to claim 6, wherein Before using the initial watermark extraction network to extract the watermark from the second sample image embedded with the first watermark information to obtain the sample watermark information, the method further includes: Use the initial watermark embedding network to embed the first watermark information into the initial image to obtain the first image; Use the noise addition network to add noise to the first image to obtain the second sample image; The iterative training of the initial watermark extraction network based on the second model loss to obtain the trained watermark extraction network includes: Based on the second model loss, adjust the parameters of the initial watermark extraction network, the initial watermark embedding network, and the noise addition network until the training end condition is reached. After reaching the training end condition, use the initial watermark extraction network as the trained watermark extraction network.
8. The method according to claim 7, characterized in that, The initial watermark embedding network includes an image encoding network; The step of using the initial watermark embedding network to embed the first watermark information into the initial image to obtain the first image includes: The image encoding network performs encoding processing based on the first watermark information and the initial image to obtain an encoded image; Perform non-linear transformation on the encoded image to obtain the transformed encoded image; Scale the transformed encoded image to obtain a scaled image; Fuse the scaled image with the initial image to obtain the first image.
9. An image generation method, characterized in that, Includes: Use the target image generation model to generate an image based on Gaussian noise to obtain a target image embedded with target watermark information, and the target image generation model is trained according to the method of any one of claims 1 to 8.
10. The method according to claim 9, characterized in that, The target image generation model includes a target text feature extraction network. Before using the target image generation model to generate an image based on Gaussian noise to obtain a target image embedded with target watermark information, the method further includes: Use the target text feature extraction network to extract features from the target text to obtain target information features; The step of using the target image generation model to generate an image based on Gaussian noise to obtain a target image embedded with target watermark information includes: Use the target image generation model to perform inverse diffusion processing on the Gaussian noise image based on the target information features to obtain a target image embedded with target watermark information.
11. The method according to claim 9 or 10, characterized in that, After using the target image generation model to generate an image based on Gaussian noise to obtain a target image embedded with target watermark information, the method further includes: Use the trained watermark extraction network to extract the watermark from the target image to obtain the target watermark information.
12. A training device for an image generation model, characterized in that, The device includes: A data acquisition module for acquiring the first sample image and the target watermark information; An image reconstruction module for reconstructing the first sample image through the initial image generation model to obtain a candidate image embedded with watermark information; A predicted watermark extraction module for using the trained watermark extraction network to extract the watermark from the candidate image to obtain the predicted watermark information; A loss determination module, configured to determine a first model loss based on the target watermark information and the predicted watermark information; A model training module, configured to iteratively train an initial image generation model based on the first model loss to obtain a target image generation model, where the target image generation model is used to generate an image embedded with the target watermark information.
13. An image generation device, characterized in that, The apparatus includes: An image generation module, configured to generate an image based on Gaussian noise by using a target image generation model to obtain a target image embedded with target watermark information, where the target image generation model is trained by the image generation model training apparatus as claimed in claim 12.
14. An electronic device, characterized in that, Comprising: One or more processors; A memory; One or more computer programs, where the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and the one or more computer programs are configured to execute the method as claimed in any one of claims 1-8 or 9-11.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program can be called by a processor to execute the method as claimed in any one of claims 1-8 or 9-11.
16. A computer program product comprising computer instructions, characterized in that, When the computer instruction is executed by a processor, the steps of the method as claimed in any one of claims 1-8 or 9-11 are implemented.
Citation Information
Cited By
Generative steganography method and system based on diffusion model and semantic cue word
CN120856837A