Training method of watermark image generation model, watermark image generation method and equipment

Through the training method of the watermark image generation model, the watermark features are mapped into the unified space using an adaptive attention mechanism, which solves the diversification and robustness of the watermark scheme in the prior art, and realizes flexible embedding and efficient identification of watermarks.

CN120355558AActive Publication Date: 2025-07-22BEIJING INSTITUTE OF GRAPHIC COMMUNICATION
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510204067.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-22
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing watermarking scheme based on diffusion model cannot meet users' needs for diversified watermarks, and there is a risk that malicious users will avoid watermark detection by changing the decoder, making it difficult to achieve watermark robustness and personalized customization.

Method used

The watermark image generation model is adopted, including the watermark encoding and decoding module, the watermark feature mapping module and the adaptive attention module. By training the original watermark information and image features in the data set, the watermark features are mapped into the unified feature space using the adaptive attention mechanism, and the attention is dynamically adjusted during the image generation process to ensure the integrity and robustness of the watermark.

Benefits of technology

It realizes flexible and diversified embedding and efficient recognition of watermarks, improves the recognition and robustness of watermark images under various conditions, meets the personalized needs of different users, and prevents malicious users from tampering with watermarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355558A_ABST
    Figure CN120355558A_ABST
Patent Text Reader

Abstract

The invention provides a training method of a watermark image generation model, and a watermark image generation method and device. The method comprises the steps of obtaining a training data set; inputting the original image into a variational auto-encoder, and outputting original image features; inputting the original watermark information and the original image features into an initial watermark coding and decoding module, training the initial watermark coding and decoding module until the first target loss function converges, and obtaining a watermark coding and decoding module; inputting the original watermark information into a watermark encoding and decoding module to obtain watermark features; inputting the original watermark information into a watermark feature mapping module to obtain mapping features; and inputting the watermark feature and the mapping feature into an adaptive attention module in the initial first image information generation network to obtain an initial adaptive feature, and training the initial first image information generation network by using the initial adaptive feature until the second target loss function converges to obtain a watermark image generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and in particular, to a method for training a watermark image generation model, a watermark image generation method, and a device. Background Art

[0002] In recent years, diffusion models have achieved rapid development, making high-quality image generation more convenient. However, it has also raised the issue of copyright protection for generated images. Existing methods usually embed watermarks secretly during the diffusion process to enhance the robustness of the watermarks. However, most current watermarking schemes based on the diffusion process focus on embedding fixed watermarks, which to a certain extent limits the flexibility and personalization of watermarks and cannot meet the diverse needs of users for watermarks. For example, an artist may wish to embed different identifiers or signatures in different works to reflect the uniqueness of their creation and copyright information, and the existing fixed watermarking schemes clearly cannot meet this need. In addition, although these schemes have achieved certain results in enhancing the robustness of watermarks, there is still a risk that malicious users can replace the decoder to avoid watermark detection. Once malicious users master the cracking method of a certain decoder, they may easily remove or tamper with the watermark embedded in the image, thus infringing on the copyright of the original creator.

[0003] To address these issues, future research needs to explore more flexible watermark embedding methods to achieve personalized customization of watermarks. At the same time, it is also necessary to develop a more secure and reliable decoder verification mechanism to prevent malicious users from avoiding watermark detection by replacing the decoder. For example, an adaptive watermark embedding algorithm based on deep learning can be studied to automatically generate personalized watermark patterns according to the image content and user needs and embed them into the generation process of the diffusion model. In addition, distributed ledger technologies such as blockchain can be introduced to record the embedding and verification processes of watermarks to ensure the immutability and traceability of watermarks. Summary of the Invention

[0004] In view of this, the purpose of the present disclosure is to propose a method for training a watermark image generation model, a watermark image generation method, and a device to solve or partially solve the above problems.

[0005] Based on the above purpose, the first aspect of the present disclosure provides a method for training a watermark image generation model. The watermark image generation model includes a watermark encoding and decoding module, a watermark feature mapping module, and a first image information generation network. The first image information generation network includes an adaptive attention module. The method includes: Obtain a training data set, where the training data set includes original watermark information and original images; Input the original images into a variational autoencoder, and after being processed by the variational autoencoder, output the original image features; Input the original watermark information and the original image features into the initial watermark encoding and decoding module, and use the original watermark information and the original image features to train the initial watermark encoding and decoding module until the first target loss function corresponding to the initial watermark encoding and decoding module converges, obtaining the watermark encoding and decoding module; Input the original watermark information into the watermark encoding and decoding module to obtain watermark features; Input the original watermark information into the watermark feature mapping module to obtain mapped features; Input the watermark features and the mapped features into the adaptive attention module in the initial first image information generation network. After being processed by the adaptive attention module, obtain the initial adaptive features, and use the initial adaptive features to train the initial first image information generation network until the second target loss function corresponding to the initial first image information generation network converges. The training of the first image information generation network is completed, obtaining the watermark image generation model.

[0006] For the above purpose, the second aspect of the present disclosure provides a method for generating a watermark image, which applies the watermark image generation model. The method includes: Obtain the image to be processed and the target watermark information; Input the image to be processed and the target watermark information into the watermark image generation model. After being processed by the watermark image generation model, obtain the watermark image, where the watermark image contains the target watermark information.

[0007] Based on the same inventive concept, the third aspect of the present disclosure proposes a training device for a watermark image generation model. The watermark image generation model includes a watermark encoding and decoding module, a watermark feature mapping module, and a first image information generation network. The first image information generation network includes an adaptive attention module, including: A data acquisition module, configured to acquire a training data set, where the training data set includes original watermark information and original images; An original image feature determination module, configured to input the original image into a variational autoencoder, and after being processed by the variational autoencoder, output the original image features; A first training module, configured to input the original watermark information and the original image features into the initial watermark encoding and decoding module, and use the original watermark information and the original image features to train the initial watermark encoding and decoding module until the first target loss function corresponding to the initial watermark encoding and decoding module converges, obtaining the watermark encoding and decoding module; A watermark feature determination module, configured to input the original watermark information into the watermark encoding and decoding module to obtain watermark features; A mapped feature determination module, configured to input the original watermark information into the watermark feature mapping module to obtain mapped features; The second training module is configured to input the watermark feature and the mapping feature into the adaptive attention module in the initial first image information generation network. After being processed by the adaptive attention module, an initial adaptive feature is obtained, and the initial first image information generation network is trained by using the initial adaptive feature until the second target loss function corresponding to the initial first image information generation network converges, and the first image information generation network is trained to completion, obtaining a watermark image generation model.

[0008] Based on the same inventive concept, a fourth aspect of the present disclosure provides a watermark image generation device, including: A data acquisition module configured to acquire an image to be processed and target watermark information; A watermark image generation module configured to input the image to be processed and the target watermark information into the watermark image generation model. After being processed by the watermark image generation model, a watermark image is obtained, where the target watermark information is included in the watermark image.

[0009] Based on the same inventive concept, a fifth aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor. When the processor executes the computer program, the training method of the emotion determination model as described above is implemented.

[0010] Based on the same inventive concept, a sixth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the training method of the emotion determination model as described above.

[0011] As can be seen from the above, the present disclosure proposes a training method for a watermark image generation model, a watermark image generation method and device. The watermark image generation model includes a watermark encoding / decoding module, a watermark feature mapping module and a first image information generation network, and the first image information generation network includes an adaptive attention module. A training data set is obtained, where the training data set includes original watermark information and original images. The original images are input into a variational autoencoder, and after being processed by the variational autoencoder, original image features are output. The original watermark information and the original image features are input into an initial watermark encoding / decoding module, and the initial watermark encoding / decoding module is trained using the original watermark information and the original image features until the first target loss function corresponding to the initial watermark encoding / decoding module converges, and a watermark encoding / decoding module is obtained. First, training the initial watermark encoding / decoding module ensures the integrity and robustness of the watermark. The original watermark information is input into the watermark encoding / decoding module to obtain watermark features, and the original watermark information is input into the watermark feature mapping module to obtain mapped features. The watermark features of different distributions are mapped into a unified feature space through the watermark feature mapping module. The watermark features and the mapped features are input into the adaptive attention module in the initial first image information generation network, and after being processed by the adaptive attention module, initial adaptive features are obtained. The initial first image information generation network is trained using the initial adaptive features until the second target loss function corresponding to the initial first image information generation network converges, and the training of the first image information generation network is completed, and a watermark image generation model is obtained. By performing fine-tuning training on the first image information generation network, the obtained watermark image generation model can better retain the integrity and robustness of the watermark features when generating images. Subsequently, the images generated using the watermark image generation model can effectively maintain the recognizability of the watermark under various conditions, thereby improving the overall watermark effect and meeting the flexibility of different watermarks. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 It is a flowchart of the training method for the watermark image generation model according to an embodiment of the present disclosure; Figure 2 It is a flowchart of the watermark image generation method according to an embodiment of the present disclosure; Figure 3 It is an architecture diagram of the training method for the watermark image generation model according to another embodiment of the present disclosure; Figure 4Schematic diagram of the model architecture of the watermark encoding module according to another embodiment of the present disclosure; Figure 5 Block diagram of the structure of the training device of the watermark image generation model according to the embodiment of the present disclosure; Figure 6 Block diagram of the structure of the watermark image generation device according to the embodiment of the present disclosure; Figure 7 Schematic diagram of the structure of the electronic device according to the embodiment of the present disclosure. Detailed implementation manners

[0014] To make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0015] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0016] The following are the noun explanations related to the present disclosure: U-Net: An image information generation network.

[0017] VAE: Variational Autoencoder (VAE), which describes the observation of the latent space in a probabilistic manner and shows great application value in data generation.

[0018] WDM: Watermark Decoding Module (WDM).

[0019] WEM: Watermark Embedding Module (WEM).

[0020] Digital watermarking is a copyright protection solution widely applied to images, audio and video, as well as deep models. Traditional watermarking schemes achieve copyright protection for images by using frequency domain transformation after image generation or by training watermark encoding and decoding networks. This way of embedding watermarks after image generation is called the post-embedding watermarking method. However, the work of post-embedded watermarks is vulnerable to watermark attacks to remove the watermarks in the images, such as image cropping, intelligent elimination and other technical means that can remove the watermarks in the images. Therefore, recent research has focused more on embedding watermarks imperceptibly into the diffusion model during the diffusion process of the diffusion model. The images generated based on this method will carry watermark information. This way is called diffusion process watermarking. Among them, StableSignature proposed a method using a pre-trained decoder. Specifically, the watermark is embedded in the variational autoencoder decoder of the diffusion model. However, this scheme requires training different VAE decoders for each different watermark, which makes it difficult for Stable Signature to meet the needs of thousands of users. In addition, malicious users can easily avoid watermark verification by replacing the VAE decoder without an embedded watermark. In addition, it is proposed to hide the watermark in the frequency domain of the initial noise vector of the diffusion model, and then detect the watermark information by inverting the generated image and then recovering the noise. Although Tree-Rings does not require training, the detection results largely depend on the recovery process, and there are problems in identification among multiple users. To sum up, the existing schemes cannot effectively solve the problem of avoiding watermarks by replacing the VAE decoder, and the existing research focuses on embedding fixed watermark information during the diffusion process, which cannot meet the diverse needs of users for embedded watermarks.

[0021] Based on the above description, this embodiment proposes a training method for a watermarked image generation model. The watermarked image generation model includes a watermark encoding and decoding module, a watermark feature mapping module, and a first image information generation network, as Figure 1 shown, the method includes: Step 101, obtain a training data set, where the training data set includes original watermark information and original images.

[0022] Step 102, input the original image into the variational autoencoder, and after being processed by the variational autoencoder, output the original image features.

[0023] Step 103, input the original watermark information and the original image features into the initial watermark encoding and decoding module, and use the original watermark information and the original image features to train the initial watermark encoding and decoding module until the first target loss function corresponding to the initial watermark encoding and decoding module converges to obtain the watermark encoding and decoding module.

[0024] Step 104: Input the original watermark information into the watermark encoding and decoding module to obtain watermark features.

[0025] Step 105: Input the original watermark information into the watermark feature mapping module to obtain mapping features.

[0026] Step 106: Input the watermark features and mapping features into the adaptive attention module in the initial first image information generation network. After being processed by the adaptive attention module, obtain initial adaptive features. Use the initial adaptive features to train the initial first image information generation network until the second target loss function corresponding to the initial first image information generation network converges. The training of the first image information generation network is completed, and a watermark image generation model is obtained.

[0027] Specifically in implementation, obtain a training data set, where the training data set contains original watermark information and original images. The original images are images without watermarks, and the original watermark information is binary watermark information.

[0028] The watermark encoding and decoding module includes a variational autoencoder. In this embodiment, both the variational autoencoder and the variational decoder are trained encoders and decoders. In this embodiment, freeze the model parameters of the variational autoencoder and the variational decoder, that is, do not train the variational autoencoder and the variational decoder, and only train the watermark encoding module and the watermark decoding module.

[0029] Input the original image into the variational autoencoder. After being processed by the variational autoencoder, output the original image features.

[0030] After the watermark encoding and decoding pre-training stage, train to obtain the watermark encoding and decoding module. The specific process is as follows: Input the original watermark information and the original image features into the initial watermark encoding and decoding module, and use the original watermark information and the original image features to train the initial watermark encoding and decoding module until the first target loss function corresponding to the initial watermark encoding and decoding module converges, and obtain the watermark encoding and decoding module.

[0031] After determining that the training of the watermark encoding and decoding module is completed, enter the U-Net fine-tuning stage, and freeze the model parameters obtained by the watermark encoding and decoding module. Input the original watermark information into the watermark encoding and decoding module to obtain watermark features.

[0032] Input the original watermark information into the watermark feature mapping module to obtain mapping features. The specific method of using the watermark feature mapping module to obtain mapping features is as follows: The original watermark information is a watermark sequence, specifically binary watermark information. The watermark sequence with a length of is converted into a vector with a length of r. For the th bit of the watermark sequence, use the embedding vector are used to represent the binary states 0 and 1, where is initialized by a standard Gaussian distribution.

[0033] The calculation process of the watermark feature mapping module is as follows:

[0034]

[0035] where represents the sequence obtained after mapping the -th bit of the watermark, represents the value of the -th bit of the watermark sequence, is the length of the watermark sequence. represents the watermark mapping feature matrix, (.) represents a function for creating a diagonal matrix. Finally, the diagonal matrix of the mapped watermark sequence is obtained as the watermark mapping feature, i.e., the mapping feature.

[0036] The watermark feature and the mapping feature are input into the adaptive attention module in the initial first image information generation network, where the first image information generation network is a U-Net network. Through the processing of the adaptive attention module, an initial adaptive feature is obtained. The initial first image information generation network is trained using the initial adaptive feature until the second target loss function corresponding to the initial first image information generation network converges. The first image information generation network training is completed, and a watermark image generation model is obtained.

[0037] Through the above solution, a training data set is obtained, where the training data set includes original watermark information and an original image. The original image is input into a variational autoencoder, and through the processing of the variational autoencoder, original image features are output. The original watermark information and the original image features are input into an initial watermark encoding and decoding module, and the initial watermark encoding and decoding module is trained using the original watermark information and the original image features until the first target loss function corresponding to the initial watermark encoding and decoding module converges, obtaining a watermark encoding and decoding module. Training the initial watermark encoding and decoding module first ensures the integrity and robustness of the watermark. The original watermark information is input into the watermark encoding and decoding module to obtain watermark features, and the original watermark information is input into a watermark feature mapping module to obtain mapped features. The watermark feature mapping module maps watermark features with different distributions into a unified feature space. The watermark features and the mapped features are input into an adaptive attention module in an initial first image information generation network, and through the processing of the adaptive attention module, initial adaptive features are obtained. The initial first image information generation network is trained using the initial adaptive features until the second target loss function corresponding to the initial first image information generation network converges, and the training of the first image information generation network is completed, obtaining a watermark image generation model. Through fine-tuning training of the first image information generation network, the obtained watermark image generation model can better retain the integrity and robustness of the watermark features when generating images. Subsequently, the images generated using the watermark image generation model can effectively maintain the recognizability of the watermark under various conditions, thereby improving the overall watermark effect and meeting the flexibility of different watermarks.

[0038] In some embodiments, step 103 specifically includes: Step 1021, input the original watermark information into an initial watermark encoding module, and through the processing of the initial watermark encoding module, output initial watermark features; Step 1022, perform a fusion process on the original image features and the initial watermark features to obtain embedded image features; Step 1023, input the original image features and the embedded image features into a variational auto-decoder, and through the processing of the variational auto-decoder, output an initial watermark image, and determine the first loss function corresponding to the variational auto-decoder; Step 1024, input the initial watermark image into an initial watermark decoding module, and use the initial watermark decoding module to perform decoding processing on the initial watermark image to obtain training watermark information; Step 1025, determine the second loss function corresponding to the initial watermark encoding module and the initial watermark decoding module according to the original watermark information and the training watermark information; Step 1026, perform an addition process on the first loss function and the second loss function to obtain the first target loss function corresponding to the initial watermark encoding and decoding module; Step 1027: Train the initial watermark encoding module and the initial watermark decoding module using the original watermark information and the original image features until the first target loss function converges, obtaining the watermark encoding and decoding module.

[0039] In specific implementation, the watermark encoding and decoding module includes a variational autoencoder, a watermark encoding module, a watermark decoding module, and a variational auto-decoder. In this embodiment, both the variational autoencoder and the variational auto-decoder are pre-trained encoders and decoders. In this embodiment, freeze the model parameters of the variational autoencoder and the variational auto-decoder, that is, do not train the variational autoencoder and the variational auto-decoder, and only train the watermark encoding module and the watermark decoding module.

[0040] Input the original watermark information into the initial watermark encoding module. After being processed by the initial watermark encoding module, output the initial watermark features. The function of the initial watermark encoding module is to convert the binary watermark information into a form suitable for embedding into the image features, that is, to convert it into a two-dimensional feature map with the same size as the image size of the original image. The two-dimensional feature map is the initial watermark features.

[0041] Fuse the original image features and the initial watermark features to obtain the embedded image features, where the fusion method is matrix addition.

[0042] Input the original image features and the embedded image features into the variational auto-decoder, and use the variational auto-decoder to decode the original image features and the embedded image features, output the initial watermark image, and determine the first loss function corresponding to the variational auto-decoder. Among them, the initial watermark image contains the original watermark information and the image content contained in the original image.

[0043] Input the initial watermark image into the initial watermark decoding module. Use the initial watermark decoding module to decode the initial watermark image to obtain the training watermark information. Compare the original watermark information and the training watermark information to determine the second loss function corresponding to the initial watermark encoding module and the initial watermark decoding module.

[0044] In this embodiment, the second loss function is preferably Binary CrossEntropy Loss (BCE Loss), and the second loss function is used to evaluate the accuracy of watermark decoding.

[0045] Sum the first loss function and the second loss function to obtain the first target loss function corresponding to the initial watermark encoding and decoding module. Train the initial watermark encoding module and the initial watermark decoding module using the original watermark information and the original image features until the first target loss function converges, obtaining the watermark encoding and decoding module.

[0046] Through the above solution, by inputting the original watermark information into the initial watermark encoding module, and through the processing of the initial watermark encoding module, the initial watermark features are output to ensure that the watermark can be successfully embedded into the original image features. At the same time, the original watermark information and the training watermark information are compared and processed to ensure the integrity and robustness of the watermark.

[0047] In some embodiments, the initial watermark encoding module includes an input layer, a linear layer, an activation layer, a convolutional layer, and an output layer. Step 1021 specifically includes: Step 10211, input the original watermark information into the initial watermark encoding module through the input layer; Step 10212, use the linear layer to perform a conversion process on the original watermark information to obtain linearly transformed features; Step 10213, perform a non-linear enhancement process on the linearly transformed features through the activation function of the activation layer to obtain non-linearly enhanced features; Step 10214, use the convolutional layer to adjust the number of channels of the non-linearly enhanced features to obtain the initial watermark features; Step 10215, output the initial watermark features from the initial watermark encoding module through the output layer.

[0048] Specifically in implementation, the initial watermark encoding module includes an input layer, a linear layer, an activation layer, a convolutional layer, and an output layer. The input layer is used to input the original watermark information into the initial watermark encoding module.

[0049] The linear layer is used to perform a conversion process on the original watermark information to obtain linearly transformed features. Specifically, the linear layer converts the one-dimensional watermark information with a length of 48 bits into one-dimensional information with a length of 64*64, and then uses Reshape to convert the one-dimensional information into linearly transformed features.

[0050] The non-linearly enhanced features are obtained by performing a non-linear enhancement process on the linearly transformed features through the activation function of the activation layer, that is, enhancing the non-linear representation of the linearly transformed features through the activation function. In this embodiment, the activation function is the SiLU activation function.

[0051] The convolutional layer is used to adjust the number of channels of the non-linearly enhanced features to obtain the initial watermark features. In this embodiment, the convolutional layer only changes the number of channels of the features and does not change the size of the features, which is used to enhance the ability of the model to extract watermark features. The initial watermark features are output from the initial watermark encoding module through the output layer.

[0052] In this embodiment, the watermark decoding module is retrained using the watermark decoder model structure of StegaStamp to achieve the purpose of decoding the watermark information.

[0053] In some embodiments, step 1023 specifically includes: Step 10231: Input the original image features into the variational autoencoder, and after being processed by the variational autoencoder, obtain the first target original image; Step 10232: Input the embedded image features into the variational autoencoder, and after being processed by the variational autoencoder, obtain the initial watermark image and the second target original image; Step 10233: Determine the first loss function corresponding to the variational autoencoder according to the first target original image and the second target original image.

[0054] Specifically in implementation, input the original image features into the variational autoencoder, and use the variational autoencoder to process to obtain the first target original image. Input the embedded image features into the variational autoencoder, and use the variational autoencoder to process to obtain the initial watermark image and the second target original image.

[0055] Compare the first target original image with the second target original image to determine the first loss function corresponding to the variational autoencoder. The first loss function is used to enhance the image quality by perceiving the distance between features and ensure the similarity between the image generated by the model and the original image.

[0056] The first loss function is a loss function determined according to the Learned Perceptual Image Patch Similarity (LPIPS) and the Peak Regional Variation Loss (PRVLLoss).

[0057] In this embodiment, the first loss function is expressed by the formula:

[0058] where, is the first loss function, is the Learned Perceptual Image Patch Similarity, is the Peak Regional Variation Loss.

[0059] Furthermore, the first target loss function corresponding to the initial watermark encoding and decoding module is expressed by the formula:

[0060] where, is the first target loss function, is the second loss function.

[0061] In some embodiments, the watermark encoding and decoding module further includes an image attack layer, and step 1024 specifically includes: Step 10241: Use the image attack layer to add noise to the initial watermark image to obtain the target watermark image; Step 10242: Input the target watermark image into the initial watermark decoding module, and use the initial watermark decoding module to decode the target watermark image to obtain the training watermark information.

[0062] In specific implementation, due to various types of image attacks in the actual scenarios of image use, during the pre-training stage of watermark encoding and decoding, before inputting the initial watermark image into the watermark decoding module, the initial watermark image is processed using the image attack layer. The specific method is to introduce multiple noise layers, and the noise layers include at least one of the following: JPEG compression, cropping and scaling, Gaussian blur, Gaussian noise, and color jitter.

[0063] During the forward propagation process, the image attack layer randomly selects a noise layer from each noise layer according to a preset probability and applies it to the initial watermark image. That is, the initial watermark image is added with noise to obtain the target watermark image. The target watermark image is input into the initial watermark decoding module, and the initial watermark decoding module is used to decode the target watermark image to obtain the training watermark information.

[0064] Through the above solution, the initial watermark image is added with noise using the image attack layer, making the perturbation of the input data random during each training, thereby improving the robustness of the model.

[0065] In some embodiments, the watermark encoding and decoding module includes a variational autoencoder, a watermark encoding module, a watermark decoding module, and a variational auto-decoder. Step 104 specifically includes: Step 1041: Input the original image into the variational autoencoder, and after being processed by the variational autoencoder, output the original image features; Step 1042: Input the original watermark information into the watermark encoding module, and after being processed by the watermark encoding module, obtain the target watermark features; Step 1043: Fuse the original image features and the target watermark features to obtain the watermark features.

[0066] In specific implementation, the watermark encoding and decoding module includes a variational autoencoder, a watermark encoding module, a watermark decoding module, and a variational auto-decoder. In this embodiment, it is the fine-tuning stage of U-Net. In this stage, the variational autoencoder, watermark encoding module, watermark decoding module, and variational auto-decoder in the watermark encoding and decoding module have all been trained. Therefore, the model parameters corresponding to the variational autoencoder, watermark encoding module, watermark decoding module, and variational auto-decoder are all frozen. That is, during the fine-tuning stage of U-Net, the variational autoencoder, watermark encoding module, watermark decoding module, and variational auto-decoder in the watermark encoding and decoding module are not trained.

[0067] In the U-Net fine-tuning stage, the previously generated watermark pattern is integrated and embedded into the U-Net. Specifically, the LoRA technique is used to complete the embedding of the watermark. The original image is input into the variational autoencoder, and after being processed by the variational autoencoder, the original image features are output.

[0068] The original watermark information is input into the trained watermark encoding module, and after being processed by the watermark encoding module, the target watermark features are obtained. The original image features and the target watermark features are fused to obtain watermark features, where the fusion processing method is matrix addition processing.

[0069] Through the above scheme, the original image generates original image features through the variational autoencoder, which are merged with the target watermark features extracted from the original watermark information by the watermark encoding module to obtain watermark features. The watermark features contain the information of watermark embedding, forming a fused feature representation.

[0070] In some embodiments, the watermark image generation model further includes a second image information generation network. The first image information generation network includes a denoising processing module and an adaptive attention module. Step 106 specifically includes: Step 1061: Input the watermark features into the denoising processing module, and use the denoising processing module to denoise the watermark features to obtain image segmentation features; Step 1062: Input the image segmentation features and the mapping features into the adaptive attention module, and after being processed by the adaptive attention module, obtain initial adaptive features; Step 1063: Input the initial adaptive features into the denoising processing module, and use the denoising processing module to denoise the initial adaptive features to obtain adaptive features; Step 1064: Input the original image features into the second image information generation network, and use the second image information generation network to perform denoising processing to obtain target original image features; Step 1065: Determine the second target loss function corresponding to the initial first image information generation network according to the adaptive features and the target original image features; Step 1066: Use the watermark features and the mapping features to train the initial first image information generation network until the second target loss function corresponding to the initial first image information generation network converges, and the first image information generation network is trained.

[0071] Specifically in implementation, the watermark image generation model further includes a second image information generation network. The second image information generation network is a trained U-Net network. In this embodiment, the network parameters of the second image information generation network are frozen, that is, the second image information generation network is not trained.

[0072] Input the watermark feature into the denoising processing module, and use the denoising processing module to denoise the watermark feature to obtain an image segmentation feature, which is the U-Net feature. Input the image segmentation feature and the mapped feature obtained after being processed by the watermark feature mapping module into the adaptive attention module, and after being processed by the adaptive attention module, an initial adaptive feature is obtained.

[0073] Input the initial adaptive feature into the denoising processing module, and use the denoising processing module to denoise the initial adaptive feature to obtain an adaptive feature. Input the original image feature into the second image information generation network, and use the second image information generation network to perform denoising processing to obtain a target original image feature.

[0074] Compare the adaptive feature and the target original image feature to determine the second target loss function corresponding to the initial first image information generation network. In this embodiment, the second target loss function is the mean square error loss function, and the second target loss function is used to measure the pixel-level difference between the features generated by the original U-Net and the LoRA fine-tuned features. By minimizing the MSE Loss, the model can gradually adjust the generated feature representation, so that the embedded watermark information appears more accurately in the generated image, thereby improving the quality and robustness of the watermark.

[0075] In some embodiments, the adaptive attention module includes an input layer, a linear layer, a cross-attention layer, and an output layer. Step 1062 specifically includes: Step 10621, input the image segmentation feature and the mapped feature into the adaptive attention module through the input layer; Step 10622, use the linear layer to perform dimensionality adjustment processing on the image segmentation feature to obtain a first downsampled sequence feature; Step 10623, use the linear layer to perform dimensionality adjustment processing on the mapped feature to obtain a second downsampled sequence feature; Step 10624, use the cross-attention layer to perform weight adjustment on the first downsampled sequence feature and the second downsampled sequence feature to obtain an initial adaptive feature.

[0076] Specifically, the adaptive attention module includes an input layer, a linear layer, a cross-attention layer, and an output layer. Input the image segmentation feature and the mapped feature into the adaptive attention module through the input layer.

[0077] Use the linear layer to perform dimensionality adjustment processing on the image segmentation feature to obtain a first downsampled sequence feature, and use the linear layer to perform dimensionality adjustment processing on the mapped feature to obtain a second downsampled sequence feature. Use the cross-attention layer to perform weight adjustment on the first downsampled sequence feature and the second downsampled sequence feature to obtain an initial adaptive feature.

[0078] Through the above solution, the adaptive attention module is first used to obtain the dimensionality-reduced sequence features from the image segmentation features and the mapping features through a linear layer, and then the correlation between the features is calculated in the cross-attention to generate a weighted feature representation, which is the initial adaptive feature. The adaptive process modifies the downsampling layer of the U-Net, which allows the U-Net to more precisely embed the watermark features into the target image during the generation process and achieve the flexibility of the algorithm in dealing with different watermarks. The adaptive attention mechanism provides an effective solution for solving the single watermark embedding problem. According to different input watermarks, the attention of the U-Net model to different watermark features is dynamically adjusted to ensure that the model can better retain the integrity and robustness of the watermark features when generating images. The generated images can effectively maintain the recognizability of the watermark under various conditions, thereby improving the overall watermark effect and meeting the flexibility of different watermarks.

[0079] Another embodiment of the present disclosure provides a watermark image generation method, which applies the watermark image generation model obtained in the above embodiment, as Figure 2 shown, the method includes: Step 201, obtain the image to be processed and the target watermark information; Step 202, input the image to be processed and the target watermark information into the watermark image generation model, and through the processing of the watermark image generation model, obtain a watermark image, where the watermark image contains the target watermark information.

[0080] Specifically, when implementing, obtain the image to be processed and the target watermark information, and input the image to be processed and the target watermark information into the trained watermark image generation model. Through the processing of the watermark image generation model, obtain a watermark image, where the watermark image contains the target watermark information.

[0081] Based on the same inventive concept, another embodiment of the present disclosure provides a training method for a watermark image generation model, and the architecture diagram of the method is as Figure 3 shown, the method specifically includes: The training method of the watermark image generation model is divided into the following two stages: the watermark encoding and decoding pre-training stage and the U-Net fine-tuning stage. In the watermark encoding and decoding pre-training stage, by freezing all the parameters of the diffusion model, a set of watermark encoding and decoding decoders are trained without disturbing the original capabilities of the model to ensure the effectiveness of the watermark. In the U-Net fine-tuning stage, the parameters of the watermark encoding and decoding decoder in the first stage are frozen. The blue snowflakes in the figure indicate the frozen model parameters. The denoising part of the diffusion model U-Net is effectively fine-tuned by the LoRA method to minimize the impact on the image generation ability and enable the U-Net to learn the watermark pattern in the first stage. At the same time, an adaptive attention mechanism is introduced in the U-Net fine-tuning stage, enabling the model to be adjusted according to different watermark features, improving the flexibility and adaptability of the model.

[0082] In the watermark encoding and decoding pre-training stage, the role of the Watermark Embedding Module (WEM) is to convert the binary watermark information M (such as "101011...") into a form suitable for embedding into image features, namely a two-dimensional feature map of the image size, to ensure that the watermark can be successfully embedded into the image features. Among them, the fusion of the image features and the watermark features is carried out through simple matrix addition. Subsequently, the VAE decoder further decodes these features and generates the watermarked image. The output watermarked image is compared with the input watermark information by the Watermark Decoding Module (WDM) to ensure the integrity and robustness of the watermark.

[0083] The model architecture of the watermark encoding module WEM is as Figure 4 shown, consisting of a linear layer, a convolutional layer, and a non-linear activation layer. Specifically, it includes one linear layer, four non-linear layers, and three convolutional layers. The linear layer converts the one-dimensional watermark information with a length of 48 bits into one-dimensional information with a length of 64*64. Then, the one-dimensional information is converted into a two-dimensional feature by using Reshape. The non-linear representation of the feature is enhanced by the SiLU activation function. Finally, the watermark feature map is obtained through the convolutional layer and then fused with the image features. Among them, the convolutional module only changes the number of channels of the feature map without changing the size of the feature map, which is used to enhance the ability of the model to extract watermark features.

[0084] Regarding the watermark decoding module WDM, the watermark decoder model structure of StegaStamp is used to retrain to complete the purpose of decoding the watermark information.

[0085] Due to various types of image attacks in the actual scenarios of image use, during the pre-training stage of watermark encoding and decoding, before inputting the watermark image into the watermark decoder WDM, the watermark image is processed using an attack layer. Specifically, multiple noise layers are introduced, including JPEG compression, cropping and scaling, Gaussian blur, Gaussian noise, and color jitter. During the forward propagation process, the image attack layer randomly selects one from each noise layer according to a preset probability and applies it to the input image. By design, the perturbation of the input data is made random each time training is performed, thereby improving the robustness of the model.

[0086] During the pre-training stage of watermark encoding and decoding, the loss function is expressed using the formula:

[0087] where, is the loss function, is the image perceptual similarity loss, is the peak region variation loss, is the binary cross-entropy loss.

[0088] In this embodiment, the binary cross-entropy loss (BCE Loss) is used to evaluate the accuracy of watermark decoding, and the distance between perceptual features is enhanced by the Learned Perceptual Image Patch Similarity (LPIPS) and the peak region variation loss (PRVL Loss) to ensure the similarity between the image generated by the model and the original image.

[0089] During the U-Net fine-tuning stage, the goal is to integrate and embed the previously generated watermark pattern into the U-Net. To achieve this goal, the LoRA technique is used to complete the embedding of the watermark. Generally, the original image generates features through the VAE encoder and is merged with the watermark features extracted by the watermark encoder WEM. At this time, the features contain the information of watermark embedding, forming a fused feature representation. To further enhance the expression and embedding of watermark features, by introducing an adaptive attention mechanism, the model adaptively processes the watermark mapping features and the features of U-Net denoising. The adaptive attention mechanism provides an effective solution to the problem of single watermark embedding. Specifically, according to the different input watermarks, the attention of the U-Net model to different watermark features is dynamically adjusted. This mechanism ensures that the model can better retain the integrity and robustness of watermark features when generating images. The generated images can effectively maintain the recognizability of the watermark under various conditions, thereby improving the overall watermark effect and meeting the flexibility of different watermarks. The adaptive process enables the model to handle the distribution differences between different watermark features by introducing a flexible adaptive attention mechanism, without the need to retrain the model for each new watermark. When facing diverse watermark features, the model can still maintain a high embedding and decoding accuracy, thus solving the recognition problem among multiple users.

[0090] The adaptive process includes a watermark feature mapping module and an adaptive attention mechanism. The watermark feature mapping module is a crucial part of the adaptive process and is used to map watermark features with different distributions into a unified feature space. Specifically, a watermark sequence of length is converted into a vector of length r. For the -th bit of the watermark, the embedding vectors and 0 are used to represent the binary states 0 and 1, where is initialized by a standard Gaussian distribution. The calculation process of the watermark feature mapping module is as follows:

[0091]

[0092] where represents the sequence obtained after mapping the -th bit of the watermark, represents the value of the -th bit of the watermark sequence, and is the length of the watermark sequence. represents the watermark mapping feature matrix, and (.) represents the function used to create a diagonal matrix. Finally, the diagonal matrix of the mapped watermark sequence is obtained as the watermark mapping feature.

[0093] As Figure 3 shows, the adaptive attention mechanism is implemented through a cross-attention module, which further adjusts the weights of the features in the generation process by combining the features extracted by the U-Net and the watermark features. This module first obtains the reduced-dimensional sequence features from the features of the U-Net and the watermark features through a linear layer, and then calculates the correlation between the features in the cross-attention to generate a weighted feature representation. Finally, the adaptive features are input into the next-level U-Net network. Specifically, the adaptive process modifies the downsampling layer of the U-Net. This process allows the U-Net to more precisely embed the watermark features into the target image during the generation process and achieve the flexibility of the algorithm in dealing with different watermarks.

[0094] In the fine-tuning stage of the U-Net, the mean squared error loss function (MSE Loss) is used in this embodiment to measure the pixel-level difference between the features generated by the original U-Net and the LoRA fine-tuned features. By minimizing the MSE Loss, the model can gradually adjust the generated feature representations, making the embedded watermark information appear more accurately in the generated images, thereby improving the quality and robustness of the watermark. It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0095] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0096] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a training device for a watermark image generation model. The watermark image generation model includes a watermark encoding and decoding module, a watermark feature mapping module, and a first image information generation network, including: Reference Figure 5 , Figure 5 For the training device of the watermark image generation model of the embodiment, including: A data acquisition module 501, configured to acquire a training data set, where the training data set includes original watermark information and original images; An original image feature determination module 502, configured to input the original image into a variational autoencoder, and output original image features after being processed by the variational autoencoder; A first training module 503, configured to input the original watermark information and the original image features into an initial watermark encoding and decoding module, and train the initial watermark encoding and decoding module using the original watermark information and the original image features until the first target loss function corresponding to the initial watermark encoding and decoding module converges, to obtain the watermark encoding and decoding module; A watermark feature determination module 504, configured to input the original watermark information into the watermark encoding and decoding module to obtain watermark features; A mapping feature determination module 505, configured to input the original watermark information into a watermark feature mapping module to obtain mapping features; A second training module 506, configured to input the watermark features and the mapping features into an adaptive attention module in an initial first image information generation network, and through processing by the adaptive attention module, obtain initial adaptive features, and use the initial adaptive features to train the initial first image information generation network until the second target loss function corresponding to the initial first image information generation network converges, and the training of the first image information generation network is completed to obtain a watermark image generation model.

[0097] In some embodiments, the first training module 503 specifically includes: An initial watermark feature determination unit, configured to input the original watermark information into an initial watermark encoding module, and through processing by the initial watermark encoding module, output initial watermark features; A fusion unit, configured to perform fusion processing on the original image features and the initial watermark features to obtain embedded image features; A first loss function determination unit, configured to input the original image features and the embedded image features into a variational autoencoder, and through processing by the variational autoencoder, output an initial watermark image, and determine the first loss function corresponding to the variational autoencoder; A decoding processing unit, configured to input the initial watermark image into an initial watermark decoding module, and use the initial watermark decoding module to perform decoding processing on the initial watermark image to obtain training watermark information; A second loss function determination unit, configured to determine the second loss function corresponding to the initial watermark encoding module and the initial watermark decoding module according to the original watermark information and the training watermark information; A first target loss function determination unit, configured to perform summation processing on the first loss function and the second loss function to obtain the first target loss function corresponding to the initial watermark encoding and decoding module; A training unit, configured to use the original watermark information and the original image features to train the initial watermark encoding module and the initial watermark decoding module until the first target loss function converges to obtain a watermark encoding and decoding module.

[0098] In some embodiments, the initial watermark encoding module includes an input layer, a linear layer, an activation layer, a convolutional layer, and an output layer, and the initial watermark feature determination unit is specifically configured to: Input the original watermark information into the initial watermark encoding module through the input layer; Use the linear layer to perform conversion processing on the original watermark information to obtain linearly converted features; Perform non-linear enhancement processing on the linearly converted features through the activation function of the activation layer to obtain non-linearly enhanced features; Adjust the number of channels of the non - linear enhanced features using a convolutional layer to obtain initial watermark features; Output the initial watermark features from the initial watermark encoding module through an output layer.

[0099] In some embodiments, the first loss function determination unit is specifically configured to: Input the original image features into a variational auto - encoder, and after being processed by the variational auto - encoder, obtain a first target original image; Input the embedded image features into a variational auto - encoder, and after being processed by the variational auto - encoder, obtain an initial watermark image and a second target original image; Determine the first loss function corresponding to the variational auto - encoder according to the first target original image and the second target original image.

[0100] In some embodiments, the watermark encoding and decoding module further includes an image attack layer. The decoding processing unit is specifically configured to: Perform noise - adding processing on the initial watermark image using the image attack layer to obtain a target watermark image; Input the target watermark image into the initial watermark decoding module, and use the initial watermark decoding module to perform decoding processing on the target watermark image to obtain training watermark information.

[0101] In some embodiments, the watermark encoding and decoding module includes a variational auto - encoder, a watermark encoding module, a watermark decoding module, and a variational auto - decoder. The watermark feature determination module 504 is specifically configured to: Input the original image into the variational auto - encoder, and after being processed by the variational auto - encoder, output original image features; Input the original watermark information into the watermark encoding module, and after being processed by the watermark encoding module, obtain target watermark features; Fuse the original image features and the target watermark features to obtain watermark features.

[0102] In some embodiments, the watermark image generation model further includes a second image information generation network. The first image information generation network includes a denoising processing module and an adaptive attention module. The second training module 506 specifically includes: An image segmentation feature determination unit, configured to input the watermark features into the denoising processing module, and use the denoising processing module to perform denoising processing on the watermark features to obtain image segmentation features; An initial adaptive feature determination unit, configured to input the image segmentation features and the mapping features into the adaptive attention module, and after being processed by the adaptive attention module, obtain initial adaptive features; An adaptive feature determination unit, configured to input an initial adaptive feature into a denoising processing module, and use the denoising processing module to perform denoising processing on the initial adaptive feature to obtain an adaptive feature; A target original image feature determination unit, configured to input an original image feature into a second image information generation network, and use the second image information generation network to perform denoising processing to obtain a target original image feature; A second target loss function determination unit, configured to determine a second target loss function corresponding to an initial first image information generation network according to the adaptive feature and the target original image feature; A training unit, configured to use a watermark feature and a mapping feature to train the initial first image information generation network until the second target loss function corresponding to the initial first image information generation network converges, and the training of the first image information generation network is completed.

[0103] In some embodiments, the adaptive attention module includes an input layer, a linear layer, a cross-attention layer, and an output layer. The initial adaptive feature determination unit is specifically configured to: Input an image segmentation feature and a mapping feature into the adaptive attention module through the input layer; Use the linear layer to perform dimensionality adjustment processing on the image segmentation feature to obtain a first dimensionality-reduced sequence feature; Use the linear layer to perform dimensionality adjustment processing on the mapping feature to obtain a second dimensionality-reduced sequence feature; Use the cross-attention layer to perform weight adjustment on the first dimensionality-reduced sequence feature and the second dimensionality-reduced sequence feature to obtain an initial adaptive feature.

[0104] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a watermark image generation device, as Figure 6 shown, including: A data acquisition module 601, configured to acquire an image to be processed and target watermark information; A watermark image generation module 602, configured to input the image to be processed and the target watermark information into a watermark image generation model, and obtain a watermark image through the processing of the watermark image generation model, where the watermark image contains the target watermark information.

[0105] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0106] The device of the above embodiment is used to implement the training method of the corresponding emotion determination model in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0107] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any of the above embodiments is implemented.

[0108] Figure 7 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0109] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0110] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0111] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0112] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0113] The bus 1050 includes a path for transmitting information between various components of the device, such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040.

[0114] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and do not necessarily include all the components shown in the figure.

[0115] The electronic device of the above embodiment is used to implement the corresponding method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0116] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method described in any of the above embodiments.

[0117] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0118] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0119] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.

[0120] For example, when responding to an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure technical solution based on the prompt message.

[0121] As an optional but non-limiting implementation manner, when responding to an active request from a user, the manner of sending a prompt message to the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0122] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0123] Those of ordinary skill in the art should understand that: The discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; Under the idea of the present disclosure, the technical features between the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above. For the sake of brevity, they are not provided in detail.

[0124] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in block diagram form to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation manner of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0125] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0126] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A training method for a watermark image generation model, characterized in that, The watermark image generation model includes a watermark encoding and decoding module, a watermark feature mapping module, and a first image information generation network, and an adaptive attention module is included in the first image information generation network; It includes: Obtain a training data set, where the training data set includes original watermark information and original images; Input the original image into the variational autoencoder, and after being processed by the variational autoencoder, output the original image features; Input the original watermark information and the original image features into the initial watermark encoding and decoding module, and use the original watermark information and the original image features to train the initial watermark encoding and decoding module until the first target loss function corresponding to the initial watermark encoding and decoding module converges to obtain the watermark encoding and decoding module; Input the original watermark information into the watermark encoding and decoding module to obtain watermark features; Input the original watermark information into the watermark feature mapping module to obtain mapped features; Input the watermark features and the mapped features into the adaptive attention module in the initial first image information generation network, and after being processed by the adaptive attention module, obtain initial adaptive features, and use the initial adaptive features to train the initial first image information generation network until the second target loss function corresponding to the initial first image information generation network converges, and the training of the first image information generation network is completed to obtain the watermark image generation model.

2. The method according to claim 1, wherein The step of inputting the original watermark information and the original image features into the initial watermark encoding and decoding module, and using the original watermark information and the original image features to train the initial watermark encoding and decoding module until the first target loss function corresponding to the initial watermark encoding and decoding module converges to obtain the watermark encoding and decoding module includes: Input the original watermark information into the initial watermark encoding module, and after being processed by the initial watermark encoding module, output the initial watermark features; Perform a fusion process on the original image features and the initial watermark features to obtain embedded image features; Input the original image features and the embedded image features into the variational auto-decoder, and after being processed by the variational auto-decoder, output the initial watermark image, and determine the first loss function corresponding to the variational auto-decoder; Input the initial watermark image into the initial watermark decoding module, and use the initial watermark decoding module to perform decoding processing on the initial watermark image to obtain training watermark information; Determine the second loss function corresponding to the initial watermark encoding module and the initial watermark decoding module according to the original watermark information and the training watermark information; Perform an addition process on the first loss function and the second loss function to obtain the first target loss function corresponding to the initial watermark encoding and decoding module; Use the original watermark information and the original image features to train the initial watermark encoding module and the initial watermark decoding module until the first target loss function converges to obtain the watermark encoding and decoding module.

3. The method according to claim 2, characterized in that The initial watermark encoding module includes an input layer, a linear layer, an activation layer, a convolutional layer, and an output layer, Input the original watermark information into the initial watermark encoding module, and after being processed by the initial watermark encoding module, output the initial watermark features, including: Input the original watermark information into the initial watermark encoding module through the input layer; Use the linear layer to perform a conversion process on the original watermark information to obtain linearly transformed features; Nonlinearly enhance the linearly transformed features through the activation function of the activation layer to obtain nonlinearly enhanced features; Use the convolutional layer to adjust the number of channels of the nonlinearly enhanced features to obtain convolutional features; Output the convolutional features from the initial watermark encoding module through the output layer as the initial watermark features.

4. The method according to claim 2, wherein The step of inputting the original image features and the embedded image features into the variational autoencoder, processing them through the variational autoencoder, and outputting the initial watermark image, and determining the first loss function corresponding to the variational autoencoder includes: Input the original image features into the variational autoencoder and obtain the first target original image through the processing of the variational autoencoder; Input the embedded image features into the variational autoencoder and obtain the initial watermark image and the second target original image through the processing of the variational autoencoder; Determine the first loss function corresponding to the variational autoencoder according to the first target original image and the second target original image.

5. The method according to claim 2, wherein The watermark encoding and decoding module further includes an image attack layer, The step of inputting the initial watermark image into the initial watermark decoding module and using the initial watermark decoding module to decode the initial watermark image to obtain the training watermark information includes: Use the image attack layer to add noise to the initial watermark image to obtain the target watermark image; Input the target watermark image into the initial watermark decoding module and use the initial watermark decoding module to decode the target watermark image to obtain the training watermark information.

6. The method according to claim 1, characterized in that The watermark encoding and decoding module includes a variational autoencoder, a watermark encoding module, a watermark decoding module, and a variational autoencoder, The step of inputting the original watermark information into the watermark encoding and decoding module to obtain the watermark features includes: Input the original image into the variational autoencoder and output the original image features through the processing of the variational autoencoder; Input the original watermark information into the watermark encoding module and obtain the target watermark features through the processing of the watermark encoding module; Fuse the original image features and the target watermark features to obtain the watermark features.

7. The method according to claim 6, characterized in that, The watermark image generation model further includes a second image information generation network. The first image information generation network includes a denoising processing module and an adaptive attention module, The step of inputting the watermark features and the mapping features into the adaptive attention module in the initial first image information generation network, obtaining the initial adaptive features through the processing of the adaptive attention module, and training the initial first image information generation network until the second target loss function corresponding to the initial first image information generation network converges, and the first image information generation network training is completed to obtain the watermark image generation model includes: Input the watermark features into the denoising processing module and use the denoising processing module to denoise the watermark features to obtain image segmentation features; Input the image segmentation features and the mapping features into the adaptive attention module and obtain the initial adaptive features through the processing of the adaptive attention module; Input the initial adaptive features into the denoising processing module and use the denoising processing module to denoise the initial adaptive features to obtain adaptive features; Input the original image features into the second image information generation network and use the second image information generation network to perform denoising processing to obtain the target original image features; Determine the second target loss function corresponding to the initial first image information generation network according to the adaptive features and the target original image features; Train the initial first image information generation network by using the watermark features and the mapping features until the second target loss function corresponding to the initial first image information generation network converges, and the training of the first image information generation network is completed.

8. The method according to claim 7, wherein The adaptive attention module includes an input layer, a linear layer, a cross-attention layer and an output layer, The step of inputting the image segmentation features and the mapping features into the adaptive attention module and obtaining the initial adaptive features through the processing of the adaptive attention module includes: Input the image segmentation features and the mapping features into the adaptive attention module through the input layer; Use the linear layer to perform dimensionality adjustment processing on the image segmentation features to obtain the first dimensionality-reduced sequence features; Use the linear layer to perform dimensionality adjustment processing on the mapping features to obtain the second dimensionality-reduced sequence features; Use the cross-attention layer to perform weight adjustment on the first dimensionality-reduced sequence features and the second dimensionality-reduced sequence features to obtain the initial adaptive features.

9. A method for generating a watermarked image, characterized in that, The watermark image generation model obtained by applying any one of claims 1 to 8 includes: Obtain the image to be processed and the target watermark information; Input the image to be processed and the target watermark information into the watermark image generation model, and obtain the watermark image through the processing of the watermark image generation model, wherein the watermark image contains the target watermark information.

10. An electronic device, characterized in that, It includes a memory, a processor and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Robust watermarking technology based on visual guidance and threshold shrinkage

    CN116433452A

  • Image generation model processing method and device, image generation method and device and computer equipment

    CN116703687A

  • Model training method, coding method, decoding method and equipment

    CN117156152A

  • Image invisible watermark embedding detection processing method and device based on diffusion model

    CN117911230A

  • Image generation method and system, electronic equipment and computer readable storage medium

    CN118411283A