Multimodal information embedded image generation service watermark protection method
Through the dual reversible neural network, multimodal information is embedded in the image and simulated attack scenarios, the problem that existing watermarks are easily detected and cannot be restored is solved, and efficient and reliable image watermark protection is achieved.
Patent Information
- Application Number
- CN202510512474.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing digital watermarking methods have limited protection in image copyright protection. Watermarks are easily detected and cannot be restored to the image editing process, resulting in incomplete evidence during evidence collection.
A dual reversible neural network is used to embed multimodal information (original image and text information) into the editing image, generate watermark images, and simulate various attack scenarios through the attack layer to ensure the robustness of the watermark.
It improves the invisible and robustness of watermarks, enhances the reliability and effectiveness of watermark protection for image generation services, and provides a more comprehensive basis for copyright tracing and image editing recovery.
Smart Images

Figure CN120070146A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for watermark protection of image generation services with multi-modal information embedding, belonging to the technical field of image processing. Background Art
[0002] With the vigorous development of deep learning technology, image generation technology has made remarkable progress, its application scope has been continuously expanded, and it has gradually entered the public eye. Nowadays, with the help of advanced algorithms and powerful computing capabilities, people can easily generate various realistic and creative images. From artistic creation to advertising and marketing, from film and television production to virtual reality, image generation technology has brought unprecedented convenience and innovation to many fields. At the same time, as a key means to protect the copyright of image content, digital watermark technology has also been widely studied and applied. Digital watermark technology can embed a specific bit sequence into an image in an imperceptible way, providing an effective solution for image copyright protection without affecting the normal use of the image. Even after the image has been attacked to a certain extent, the embedded watermark information can still be accurately extracted, thus proving the copyright ownership of the image or tracing relevant information. Existing digital watermark methods have provided copyright protection for image generation service providers to a certain extent and played a positive role in protecting their legitimate rights and interests.
[0003] Although existing digital watermark methods have played a certain role in image copyright protection, in the face of increasingly complex image generation and application scenarios, their protection strength is relatively limited. On the one hand, most existing digital watermark methods do not fully combine the prior knowledge of human eye perception. The human eye has certain characteristics and laws in image perception, but existing methods ignore this point, resulting in difficulty in accurately controlling the imperceptibility of the watermark during the watermark embedding process. This makes the image after watermark embedding may show perceptible visual changes to a certain extent, not only affecting the quality and use experience of the image, but also reducing the concealment of the watermark and increasing the risk of the watermark being maliciously discovered and removed. On the other hand, existing digital watermark methods have relatively single functions and can only trace user identity information or copyright ownership, and cannot realize the restoration of the image editing process. In practical applications, an image may be edited and modified multiple times. When a copyright dispute or infringement occurs, only knowing the source or copyright owner of the image, the complete process information of the image editing cannot be obtained, resulting in an incomplete evidence chain during evidence collection, making it difficult to effectively hold the infringer accountable and crack down on the infringement, and unable to provide comprehensive and powerful protection for image generation service providers. Summary of the Invention
[0004] The object of the present invention is to provide a watermark protection method for image generation services with multi-modal information embedding. By embedding multi-modal information into an edited image to generate a watermarked image, it aims to solve the problems in the prior art that the watermark is easily detectable and the image editing process cannot be restored, resulting in incomplete evidence during forensics.
[0005] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:
[0006] The present invention provides a watermark protection method for image generation services with multi-modal information embedding, including:
[0007] Obtain text information, an original image, and an edited image obtained by editing the original image through an image generation service;
[0008] Use a trained double-reversible neural network to embed the original image and the text information into the edited image, and output a watermarked image.
[0009] Further, the double-reversible neural network includes a double-reversible neural network layer and an attack layer;
[0010] The double-reversible neural network layer is used to embed the original image and the text information into the edited image, generate a watermarked image, and obtain an extracted image and extracted text from the watermarked image;
[0011] The attack layer is used to attack the watermarked image to obtain an attacked watermarked image, and extract the embedded image and embedded text from the attacked watermarked image through a watermark extraction operation to obtain an extracted image and extracted text.
[0012] Further, the double-reversible neural network layer includes:
[0013] A preprocessing module, which is used to perform discrete wavelet transform on the edited image and the original image, and perform spatial replication operation and discrete wavelet transform on the text information to obtain a preprocessed edited image, a preprocessed original image, and preprocessed text information;
[0014] A first reversible network, including 16 affine coupling layers, which is used to embed the preprocessed text information into the preprocessed edited image to obtain an edited image containing the text information and send it to a second reversible network;
[0015] A second reversible network, including 16 affine coupling layers, which is used to embed the preprocessed original image into the edited image containing the text information to obtain an edited image containing the text information and the original image, and send it to a post-processing module;
[0016] The post-processing module is used to generate a watermarked image by performing inverse discrete wavelet transform on the edited image containing the text information and the original image.
[0017] Further, each affine coupling layer of the first invertible network generates translation and scaling parameters through three non-linear functions each including a five-layer residual dense block, and processes the preprocessed text information.
[0018] Further, each affine coupling layer of the second invertible network generates translation and scaling parameters through three non-linear functions each including channel attention, spatial attention, and a five-layer residual dense block, and processes the preprocessed edited image and the original image.
[0019] Further, the attack layer includes additive Gaussian white noise attack, Gaussian blur attack, scaling attack, cropping attack, and compression attack;
[0020] wherein, additive Gaussian white noise is added with zero-mean Gaussian white noise having a standard deviation to perform an additive Gaussian white noise attack on the watermarked image;
[0021] Gaussian low-pass kernel filtering with a size of is used to perform a Gaussian blur attack on the watermarked image;
[0022] The watermarked image is randomly reduced or enlarged according to a scaling ratio to perform a scaling attack on the watermarked image;
[0023] The watermarked image is randomly cropped according to a cropping ratio to perform a cropping attack on the watermarked image;
[0024] A random selection of a quality factor is simulated by a differentiable method to perform a compression attack on the watermarked image.
[0025] Further, the training method of the double invertible neural network includes:
[0026] Initializing the double invertible neural network;
[0027] Setting the training process of the double invertible neural network layer;
[0028] Setting the loss function of the double invertible neural network layer and the respective weights of the loss functions;
[0029] Using the randomly generated text information dataset, the original image dataset, and the edited image dataset as the training set;
[0030] According to the training process of the double reversible neural network layer, the loss function and the respective weights of the loss functions, and the training set, use the double reversible neural network layer to embed the original image and text information into the edited image to generate a watermarked image, and use the attack layer to attack the watermarked image, and then output the extracted image and the extracted text through the watermark extraction operation;
[0031] According to the loss function of the double reversible neural network, calculate the loss between the edited image and the watermarked image, the loss between the extracted image and the original image, and the loss between the extracted text and the text information, and adjust the training parameters of the double reversible neural network until the total loss meets the preset threshold.
[0032] Further, using the double reversible neural network layer to embed the original image and text information into the edited image to generate a watermarked image includes:
[0033] Send the edited image and the original image into a preprocessing module for discrete wavelet transform to obtain the preprocessed edited image and original image. At the same time, spatially duplicate the text information in the preprocessing module, and after obtaining the spatially duplicated text information, perform discrete wavelet transform to obtain the preprocessed text information;
[0034] Input the preprocessed edited image and the preprocessed text information into the first reversible network to output an edited image containing the text information;
[0035] Input the edited image containing the text information and the preprocessed original image into the second reversible network to output an edited image containing the text information and the original image;
[0036] Perform inverse discrete wavelet transform on the edited image containing the text information and the original image to obtain the final watermarked image.
[0037] Further, the loss function of the double reversible neural network includes an edited image quality loss function, an original image quality loss function, and a text quality loss function, expressed as:
[0038] ;
[0039] In the formula, represents the total loss value of the loss function of the double reversible neural network, represents the weight coefficient for adjusting the edited image quality loss function of, represents the weight coefficient for adjusting the original image quality loss function of, represents the weight coefficient for adjusting the text quality loss function of, represents the edited image, represents the watermarked image, Denote the extracted image Denote the original image Denote the bit sequence corresponding to the extracted text Denote the bit sequence corresponding to the text information
[0040] Furthermore, the editing image quality loss function is expressed as:
[0041] ;
[0042] In the formula, Denote the visual perception loss function between the editing image and the watermarked image The weight coefficient of Denote the feature perception loss function between the editing image and the watermarked image The weight coefficient of Denote the identity perception loss function between the editing image and the watermarked image The weight coefficient of;
[0043] The original image quality loss function is expressed as:
[0044] ;
[0045] In the formula, Denote the visual perception loss function between the extracted image and the original image The weight coefficient of Denote the feature perception loss function between the extracted image and the original image The weight coefficient of Denote the identity perception loss function between the extracted image and the original image The weight coefficient of;
[0046] The text quality loss function is expressed as:
[0047] ;
[0048] In the formula, Denote the L2 norm
[0049] Compared with the prior art, the beneficial effects achieved by the present invention:
[0050] 1. The present invention embeds the original image and text information into the edited image through a trained double-reversible neural network to output a watermarked image, realizing the efficient fusion and covert embedding of multi-modal information. On the one hand, the double-reversible neural network can ensure the reversibility of the watermark embedding process while minimizing the impact on the original visual quality of the edited image, making the watermarked image almost indistinguishable from the edited image in appearance, greatly enhancing the imperceptibility of the watermark, effectively avoiding the problem of image quality degradation caused by watermark embedding, and enhancing the image usage experience. On the other hand, embedding the original image and text information simultaneously enriches the information dimension carried by the watermark, which not only includes the original content features of the image but also incorporates the related text description information, providing a more comprehensive and accurate basis for subsequent copyright tracing, image editing restoration, and infringement evidence collection, significantly enhancing the reliability and effectiveness of watermark protection for image generation services.
[0051] 2. The present invention utilizes the double-reversible neural network layer of the double-reversible neural network through a preprocessing module, a first reversible network, a second reversible network, and a postprocessing module to achieve the efficient embedding and precise restoration of the original image and text information in the edited image. The preprocessing module performs discrete wavelet transform on the edited image and the original image and spatial replication operation on the text information, providing a good data basis for subsequent embedding. Among them, the first reversible network processes the text information through 16 affine coupling layers using three non-linear functions including five-layer residual dense blocks, and the second reversible network processes the image information through 16 affine coupling layers using three non-linear functions including channel attention, spatial attention, and five-layer residual dense blocks, enabling multi-modal information to be embedded into the edited image in a clever and lossless manner. After being attacked, the embedded image and text can still be accurately extracted through the watermark extraction operation, greatly improving the reliability and effectiveness of watermark protection for image generation services and providing a solid technical support for copyright tracing and image editing restoration.
[0052] 3. The present invention utilizes the attack layer of the double-reversible neural network to comprehensively simulate various malicious attack scenarios that the watermarked image may encounter in practical applications through various common image attack methods such as additive white Gaussian noise attack, Gaussian blur attack, scaling attack, cropping attack, and compression attack. By adding zero-mean Gaussian white noise with a standard deviation, Gaussian low-pass kernel filtering, random scaling, random cropping, and using a differentiable method to simulate randomly selecting a quality factor to attack the watermarked image, it can fully test the robustness of the watermark. Even after these complex and diverse attacks, the extracted image and text can still be obtained through the watermark extraction operation, indicating that the watermark embedded by the present invention has a strong anti-attack ability and can maintain integrity in various harsh environments, ensuring the effectiveness of watermark protection for image generation services.
[0053] 4. The present invention trains a double-reversible neural network by using a randomly generated text information dataset, an original image dataset, and an edited image dataset as the training set. During the training process, the losses between the edited image and the watermarked image, between the extracted image and the original image, and between the extracted text and the text information are calculated according to the loss function, and the training parameters are adjusted until the total loss meets the preset threshold, which can comprehensively optimize the performance of the network, making the generated watermarked image reach a better level in terms of quality, imperceptibility of the watermark, and robustness of the watermark. By reasonably setting the edited image quality loss function, the original image quality loss function, and the text quality loss function in the loss function, it further ensures that the influence of the watermark embedding process on the edited image, the original image, and the text information is minimized, and improves the overall performance of the watermark protection of the image generation service. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 FIG. is a schematic flowchart of a method for protecting the watermark of an image generation service with multi-modal information embedding provided by an embodiment of the present invention;
[0055] Figure 2 FIG. is a schematic flowchart of the training process of the double-reversible neural network provided by an embodiment of the present invention;
[0056] Figure 3 FIG. is a schematic flowchart of the training process of the double-reversible neural network layer provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The technical solution of the present invention will be described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present invention are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0058] The term "and / or" merely describes an association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.
[0059] Embodiment 1
[0060] As Figure 1 shown, this embodiment introduces a method for protecting the watermark of an image generation service with multi-modal information embedding, including:
[0061] Step 1: Obtain text information, an original image, and an edited image obtained by editing the original image through an image generation service.
[0062] Text information, the original image, and the edited image are the basic data for multi-modal information embedding. The text information is the watermark embedded in the edited image. The original image is the initial image without editing, and the edited image is the image processed by an image generation service such as adding special effects and modifying content. The purpose of this invention to obtain the text information, the original image, and the edited image obtained by editing the original image through an image generation service is to embed the original image and the text information into the edited image to achieve watermark protection.
[0063] Step 2: Use the trained double-reversible neural network to embed the original image and the text information into the edited image, and output the watermarked image.
[0064] The double-reversible neural network includes a double-reversible neural network layer and an attack layer. The double-reversible neural network layer is responsible for embedding the original image and the text information into the edited image, and extracting the watermark from the edited image to obtain the extracted image and the extracted text; while the attack layer is used to conduct attack tests on the watermarked image. In this step, the focus is mainly on the role of the double-reversible neural network layer.
[0065] This invention uses the preprocessing module in the double-reversible neural network layer to perform discrete wavelet transform on the edited image and the original image, and perform a spatial replication operation on the text information. The first reversible network embeds the preprocessed text information into the preprocessed edited image through 16 affine coupling layers each containing three non-linear functions, where the non-linear functions include five-layer residual dense blocks. The second reversible network embeds the preprocessed original image into the edited image containing the text information through 16 affine coupling layers each containing three non-linear functions, where the non-linear functions include channel attention, spatial attention, and five-layer residual dense blocks; finally, through the inverse discrete wavelet transform of the postprocessing module, the watermarked image is generated. This invention utilizes the characteristics of the reversible neural network to ensure the reversibility and integrity of information during the embedding process.
[0066] At the same time, the original image and the text information are embedded into the edited image simultaneously, enriching the information dimension carried by the watermark. Compared with single-modal watermarks, multi-modal information embedding can provide more comprehensive copyright protection and image editing recovery information. By training the double-reversible neural network, the original image and the text information can be embedded into it without affecting the visual quality of the edited image, making the watermarked image almost indistinguishable from the edited image in appearance, and improving the imperceptibility of the watermark.
[0067] The present invention also utilizes the testing and optimization of the attack layer to make the embedded watermark have strong robustness. Even if the watermarked image is attacked during subsequent use, the original image and text information can still be restored through the watermark extraction operation, ensuring the effectiveness of watermark protection. At the same time, the embedded original image and text information provide a basis for subsequent image editing restoration and copyright tracing. In the event of a copyright dispute or when it is necessary to restore the image editing process, the source and editing history of the image can be accurately traced by extracting the information in the watermark, protecting the legitimate rights and interests of the image generation service provider.
[0068] Embodiment 2
[0069] Based on the same inventive concept as Embodiment 1, this embodiment introduces the implementation steps of a watermark protection method for an image generation service with multi-modal information embedding, including:
[0070] Step 1: Obtain text information, the original image, and the edited image obtained by editing the original image through an image generation service.
[0071] Step 2: Train a double-reversible neural network to embed the original image and the text information into the edited image, and output a watermarked image.
[0072] In this embodiment, the training method of the double-reversible neural network is as Figure 2 shown, including:
[0073] Step 2.1: Initialize the double-reversible neural network.
[0074] In this embodiment, the double-reversible neural network includes a double-reversible neural network layer and an attack layer.
[0075] The double-reversible neural network layer is used to embed the original image and the text information into the edited image, generate a watermarked image, and obtain an extracted image and extracted text from the watermarked image.
[0076] The attack layer is used to attack the watermarked image to obtain the attacked watermarked image, and extract the embedded image and embedded text from the attacked watermarked image through the watermark extraction operation to obtain the extracted image and extracted text.
[0077] Step 2.2: Set the training process of the double-reversible neural network layer.
[0078] In this embodiment, the schematic diagram of the training process of the double-reversible neural network layer is as Figure 3 shown, where, represents the th affine coupling layer of the first reversible network, represents the first affine coupling layer of the first reversible network, Represents the second affine coupling layer of the first invertible network, represents the third affine coupling layer of the first invertible network; represents the th affine coupling layer of the second invertible network, represents the first affine coupling layer of the second invertible network, represents the second affine coupling layer of the second invertible network, represents the third affine coupling layer of the second invertible network, represents the exponential function.
[0079] In this embodiment, the double-invertible neural network layer includes a preprocessing module, a first invertible network, a second invertible network, and a postprocessing module.
[0080] The preprocessing module is used to perform discrete wavelet transform on the edited image and the original image, and perform spatial replication operation and discrete wavelet transform on the text information to obtain the preprocessed edited image, the original image, and the preprocessed text information.
[0081] The first invertible network includes 16 affine coupling layers, which are used to embed the preprocessed text information into the preprocessed edited image to obtain an edited image containing text information and send it to the second invertible network. In this embodiment, each affine coupling layer of the first invertible network generates translation and scaling parameters through three non-linear functions including five-layer residual dense blocks to process the preprocessed text information.
[0082] The second invertible network includes 16 affine coupling layers, which are used to embed the preprocessed original image into the edited image containing text information to obtain an edited image containing text information and the original image, and send it to the postprocessing module. In this embodiment, each affine coupling layer of the second invertible network generates translation and scaling parameters through three non-linear functions including channel attention, spatial attention, and five-layer residual dense blocks to process the preprocessed edited image and the original image.
[0083] The postprocessing module is used to generate a watermarked image by performing inverse discrete wavelet transform on the edited image containing text information and the original image.
[0084] Among them, the network parameters of the first invertible network and the second invertible network are shown in Tables 1 and 2:
[0085] Table 1: Network parameter table of the first invertible network
[0086]
[0087] Table 2: Network parameter table of the second invertible network
[0088]
[0089] Step 2.3: Set the loss function of the double reversible neural network layer and the weights of each loss function.
[0090] In this embodiment, the loss function of the double reversible neural network includes an edited image quality loss function, an original image quality loss function, and a text quality loss function, which are expressed as:
[0091] ;
[0092] In the formula, represents the total loss value of the loss function of the double reversible neural network, represents the weight coefficient for adjusting the edited image quality loss function of, represents the weight coefficient for adjusting the original image quality loss function of, represents the weight coefficient for adjusting the text quality loss function of, represents the edited image, represents the watermarked image, represents the extracted image, represents the original image, represents the bit sequence corresponding to the extracted text, represents the bit sequence corresponding to the text information.
[0093] In this embodiment, the edited image quality loss function is expressed as:
[0094] ;
[0095] In the formula, represents the weight coefficient of the visual perception loss function between the edited image and the watermarked image, represents the weight coefficient of the feature perception loss function between the edited image and the watermarked image, represents the weight coefficient of the identity perception loss function between the edited image and the watermarked image;
[0096] In this embodiment, the original image quality loss function is expressed as:
[0097] ;
[0098] In the formula, represents the weight coefficient of the visual perception loss function between the extracted image and the original image, Represents the feature perceptual loss function between the extracted image and the original image The weight coefficient of Represents the identity perceptual loss function between the extracted image and the original image The weight coefficient of;
[0099] In this embodiment, the text quality loss function is expressed as:
[0100] ;
[0101] In the formula, Represents the L2 norm.
[0102] In this embodiment, the visual perceptual loss function between the edited image and the watermarked image is expressed as:
[0103] ;
[0104] In the formula, Represents the luminance channel function, Represents the blue difference channel function, Represents the red difference channel function.
[0105] In this embodiment, the feature perceptual loss function between the edited image and the watermarked image is expressed as:
[0106]
[0107] In the formula, Represents the activation function of the k-th layer of the feature perceptual network.
[0108] In this embodiment, the identity perceptual loss function between the edited image and the watermarked image is expressed as:
[0109] ;
[0110] In the formula, Represents the target recognition network, Represents the cosine similarity.
[0111] In this embodiment, the visual perceptual loss function between the extracted image and the original image is expressed as:
[0112] ;
[0113] In the formula, Represents the luminance channel function, Represents the blue difference channel function, Represents the red difference channel function.
[0114] In this embodiment, the feature perception loss function between the extracted image and the original image is expressed as:
[0115] ;
[0116] In the formula, represents the activation function of the k-th layer of the feature perception network.
[0117] In this embodiment, the identity perception loss function between the extracted image and the original image is expressed as:
[0118] ;
[0119] In the formula, represents the target recognition network, represents the cosine similarity.
[0120] Step 2.4: Use the randomly generated text information dataset, the original image dataset, and the edited image dataset as the training set.
[0121] In this embodiment, the Celeb-A dataset is used as the original image dataset, and the Celeb-DF dataset is used as the edited image dataset.
[0122] Step 2.5: According to the training process of the double reversible neural network layer, the loss functions, the respective weights of the loss functions, and the training set, use the double reversible neural network layer to embed the original image and the text information into the edited image to generate a watermarked image, and use the attack layer to attack the watermarked image, and then output the extracted image and the extracted text through the watermark extraction operation, including:
[0123] Step 2.5.1: Send the edited image and the original image into the preprocessing module for discrete wavelet transform to obtain the preprocessed edited image and the original image. At the same time, perform spatial replication on the text information in the preprocessing module to obtain the spatially replicated text information, and then perform discrete wavelet transform to obtain the preprocessed text information.
[0124] Step 2.5.2: Input the preprocessed edited image and the preprocessed text information into the first reversible network to output the edited image containing the text information.
[0125] Step 2.5.3: Input the edited image containing the text information and the preprocessed original image into the second reversible network to output the edited image containing the text information and the original image.
[0126] Step 2.5.4: Perform inverse discrete wavelet transform on the edited image containing the text information and the original image to obtain the final watermarked image.
[0127] Step 2.6: Use the attack layer to attack the watermarked image, and then output the extracted image and the extracted text through the watermark extraction operation.
[0128] In this embodiment, the attack layer includes additive Gaussian white noise attack, Gaussian blur attack, scaling attack, cropping attack, and compression attack. In this embodiment, one or more of the above attack methods are used to process the watermarked image.
[0129] Among them, additive Gaussian white noise attack is performed on the watermarked image by adding zero-mean Gaussian white noise with a standard deviation of , Gaussian blur attack is performed on the watermarked image by filtering with a Gaussian low-pass kernel of size , scaling attack is performed on the watermarked image by randomly shrinking or enlarging it according to the scaling ratio , cropping attack is performed on the watermarked image by randomly cropping the watermarked image according to the cropping ratio , and compression attack is performed on the watermarked image by simulating the random selection of the quality factor through a differentiable method.
[0130] Step 2.7: According to the loss function of the double-reversible neural network, calculate the loss between the edited image and the watermarked image, the loss between the extracted image and the original image, and the loss between the extracted text and the text information, and adjust the training parameters of the double-reversible neural network until the total loss meets the preset threshold.
[0131] In this embodiment, Figure 1 a in
[0132] represents the preset threshold, and its value is 0.0015.
[0133] Step 3: Use the trained double-reversible neural network to embed the original image and the text information into the edited image, and output the watermarked image.
[0134] Step 4: Evaluate the trained double-reversible neural network.
[0134] Step 4.1: Test the trained double-reversible neural network on the Celeb-DF dataset, and use PSNR, SSIM, and ID-SIM as evaluation metrics for watermark invisibility and original image extraction, and use ACC as the evaluation metric for text information extraction. The test results of the watermarked image generated by the trained double-reversible neural network and the original image in terms of image quality and the correct rate of text information extraction are shown in Table 3:
[0135] Table 3: Invisibility and watermark extraction performance of the method of the present invention
[0136]
[0137] As can be seen from Table 3, the superiority of the present invention in terms of invisibility and extraction accuracy is mainly due to the following reasons:
[0138] (1) In the loss function of the present invention, the weight coefficient of the luminance channel is greater than that of other channels. According to the human visual system, the human eye is more sensitive to the luminance channel than to the blue-difference channel and the red-difference channel;
[0139] (2) The present invention uses an attention mechanism in the second reversible network. According to the features of different frequencies of the edited image, the position of watermark embedding is adjusted so that the watermark is embedded in the coefficients with high invisibility;
[0140] (3) The image quality loss between the edited image and the watermarked image, as well as the image quality loss function between the original image and the extracted image, make the embedding and extraction quality effects of the double reversible neural network provided by the present invention better during the training process.
[0141] Step 4.2: Test the trained double reversible neural network on the Celeb-DF dataset, and use the identity similarity ID-SIM and extraction accuracy ACC as the evaluation indicators for the robustness of the present invention against different attacks. The present invention conducts robustness tests on the watermarked images through attacks of different intensities in the attack layer and attacks never involved in the attack layer, and has obtained good comprehensive evaluation performances, as shown in Table 4:
[0142] Table 4: Robustness of the present invention against different intensity attacks (ID-SIM / ACC)
[0143]
[0144] As can be seen from Table 4, the present invention has certain superiority in terms of robustness against attacks of different intensities in the attack layer and attacks never involved in the attack layer, mainly because:
[0145] (1) The watermark of the watermarked image is copied spatially and spliced onto the feature channels multiple times to achieve redundant embedding;
[0146] (2) During the training process of the double reversible neural network, the attack layer embeds the watermark in stable coefficients, and the watermark information can be extracted from the attacked coefficients.
[0147] Example 3
[0148] Based on the same inventive concept as other embodiments, this embodiment introduces a computer-readable storage medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the steps of the method in the above-mentioned Example 1 or 2 are implemented.
[0149] Example 4
[0150] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including computer instructions, which when executed by a processor implement the steps of the method in Embodiment 1 or 2 above.
[0151] In summary of the above embodiments, the present invention embeds the original image and text information into the edited image by using the trained double-reversible neural network to output the watermarked image, realizing the efficient fusion and covert embedding of multi-modal information. On the one hand, the double-reversible neural network can minimize the impact on the original visual quality of the edited image while ensuring the reversibility of the watermark embedding process, making the watermarked image almost indistinguishable from the edited image in appearance, greatly improving the imperceptibility of the watermark, effectively avoiding the problem of image quality degradation caused by watermark embedding, and enhancing the image usage experience. On the other hand, embedding the original image and text information simultaneously enriches the information dimension carried by the watermark, including not only the original content features of the image but also the related text description information, providing a more comprehensive and accurate basis for subsequent copyright tracing, image editing restoration, and infringement evidence collection, significantly enhancing the reliability and effectiveness of the watermark protection for the image generation service.
[0152] The present invention uses the double-reversible neural network layer of the double-reversible neural network to achieve the efficient embedding and accurate restoration of the original image and text information in the edited image through a preprocessing module, a first reversible network, a second reversible network, and a postprocessing module. The preprocessing module performs discrete wavelet transform on the edited image and the original image and spatial replication operation on the text information, providing a good data basis for subsequent embedding. Among them, the first reversible network processes the text information through 16 affine coupling layers by three non-linear functions including five-layer residual dense blocks, and the second reversible network processes the image information through 16 affine coupling layers by three non-linear functions including channel attention, spatial attention, and five-layer residual dense blocks, enabling multi-modal information to be embedded into the edited image in a clever and lossless manner. After being attacked, the embedded image and text can still be accurately extracted through the watermark extraction operation, greatly improving the reliability and effectiveness of the watermark protection for the image generation service, and providing a solid technical support for copyright tracing and image editing restoration.
[0153] The present invention utilizes the attack layer of a double reversible neural network to comprehensively simulate various malicious attack scenarios that a watermarked image may encounter in practical applications through various common image attack methods such as additive Gaussian white noise attack, Gaussian blur attack, scaling attack, cropping attack, and compression attack. By adding zero-mean Gaussian white noise with a standard deviation, Gaussian low-pass kernel filtering, random scaling, random cropping, and using a differentiable method to simulate randomly selecting a quality factor, the watermarked image is attacked, which can fully test the robustness of the watermark. Even after these complex and diverse attacks, the extracted image and text can still be obtained through the watermark extraction operation, indicating that the watermark embedded in the present invention has a strong anti-attack ability, can maintain integrity in various harsh environments, and ensure the effectiveness of watermark protection for image generation services.
[0154] The present invention trains a double reversible neural network by using a randomly generated text information dataset, an original image dataset, and an edited image dataset as training sets. During the training process, the losses between the edited image and the watermarked image, between the extracted image and the original image, and between the extracted text and the text information are calculated according to the loss function, and the training parameters are adjusted until the total loss meets a preset threshold, which can comprehensively optimize the performance of the network, making the generated watermarked image reach a better level in terms of quality, imperceptibility of the watermark, and robustness of the watermark. By reasonably setting the edited image quality loss function, original image quality loss function, and text quality loss function in the loss function, it further ensures that the impact of the watermark embedding process on the edited image, original image, and text information is minimized, and improves the overall performance of watermark protection for image generation services.
[0155] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0156] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing the process Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks
[0157] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks
[0158] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions in the process Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks
[0159] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention and the claims. All of these fall within the protection scope of the present invention.
Claims
1. A watermark protection method for image generation service with multimodal information embedded, characterized in that: include: Acquire text information, an original image, and an edited image of the original image after being edited by an image generation service; The original image and the text information are embedded into the edited image using a trained dual reversible neural network, and a watermarked image is output.
2. The image generation service watermark protection method for multimodal information embedding according to claim 1 is characterized in that: The dual reversible neural network includes a dual reversible neural network layer and an attack layer; The dual reversible neural network layer is used to embed the original image and the text information into the edited image, generate a watermarked image, and obtain an extracted image and extracted text from the watermarked image; The attack layer is used to attack the watermarked image to obtain an attacked watermarked image.
3. The image generation service watermark protection method for embedding multimodal information according to claim 2, characterized in that: The dual reversible neural network layer comprises: A preprocessing module, used for performing discrete wavelet transform on the edited image and the original image and performing spatial replication operation and discrete wavelet transform on the text information, so as to obtain the preprocessed edited image and the original image and the preprocessed text information; The first reversible network includes 16 affine coupling layers, which are used to embed the preprocessed text information into the preprocessed edited image, obtain the edited image containing the text information and send it to the second reversible network; The second reversible network includes 16 affine coupling layers, which are used to embed the preprocessed original image into the edited image containing text information, obtain the edited image containing text information and the original image, and send it to the post-processing module; The post-processing module is used to generate a watermarked image by inverse discrete wavelet transforming the edited image containing text information and the original image.
4. The image generation service watermark protection method for embedding multimodal information according to claim 3 is characterized in that: Each affine coupling layer of the first reversible network generates translation and scaling parameters through three nonlinear functions including five layers of residual dense blocks to process the preprocessed text information.
5. The image generation service watermark protection method for embedding multimodal information according to claim 3, characterized in that: Each affine coupling layer of the second reversible network generates translation and scaling parameters through three nonlinear functions including channel attention, spatial attention and five-layer residual dense blocks to process the preprocessed edited image and the original image.
6. The image generation service watermark protection method for multimodal information embedding according to claim 2 is characterized in that: The attack layers include additive Gaussian white noise attack, Gaussian blur attack, scaling attack, cropping attack and Compression attack; Among them, using the standard deviation Add zero-mean Gaussian white noise to attack the watermarked image; Use size Gaussian low-pass kernel filtering is used to perform Gaussian blur attack on watermarked images; By scaling Randomly shrink or enlarge to perform scaling attacks on watermarked images; By cropping ratio Randomly crop the watermarked image and perform cropping attack on the watermarked image; Simulating random selection of quality factors via differentiable methods , for watermarked images Compression attack.
7. The image generation service watermark protection method for embedding multimodal information according to claim 2, characterized in that: The training method of the dual reversible neural network comprises: Initialize the dual reversible neural network; Set up the training process of the dual reversible neural network layer; Set the loss function of the dual reversible neural network layer and the weights of each loss function; The randomly generated text information dataset, original image dataset and edited image dataset are used as training sets; According to the training process of the dual reversible neural network layer, the loss function and the weights of the loss function and the training set, the dual reversible neural network layer is used to embed the original image and text information into the edited image to generate a watermarked image, and the attack layer is used to attack the watermarked image, and then the extracted image and extracted text are output through the watermark extraction operation; According to the loss function of the dual reversible neural network, the loss between the edited image and the watermarked image, the loss between the extracted image and the original image, and the loss between the extracted text and the text information are calculated, and the training parameters of the dual reversible neural network are adjusted until the total loss meets the preset threshold.
8. The image generation service watermark protection method for embedding multimodal information according to claim 7, characterized in that: The original image and text information are embedded into the edited image using a dual reversible neural network layer to generate a watermarked image, including: The edited image and the original image are sent to a preprocessing module for discrete wavelet transformation to obtain a preprocessed edited image and an original image, and the text information is spatially replicated in the preprocessing module to obtain the spatially replicated text information and then subjected to discrete wavelet transformation to obtain the preprocessed text information; Inputting the preprocessed edited image and the preprocessed text information into a first reversible network, and outputting the edited image containing the text information; Inputting the edited image containing text information and the preprocessed original image into a second reversible network, and outputting the edited image containing the text information and the original image; The edited image containing text information and the original image is subjected to inverse discrete wavelet transform to obtain the final watermarked image.
9. The image generation service watermark protection method for embedding multimodal information according to claim 7, characterized in that: The loss function of the dual reversible neural network includes an edited image quality loss function, an original image quality loss function and a text quality loss function, which are expressed as: ; In the formula, Represents the total loss value of the loss function of the dual reversible neural network, Represents the adjustment and editing image quality loss function The weight coefficient of Represents the function for adjusting the quality loss of the original image The weight coefficient of Represents the adjustment text quality loss function The weight coefficient of Indicates editing an image. Represents a watermarked image. represents the original image, Indicates the extracted image. Represents the bit sequence corresponding to the text information, Indicates the bit sequence corresponding to the extracted text.
10. The image generation service watermark protection method for embedding multimodal information according to claim 9, characterized in that: The edited image quality loss function is expressed as: ; In the formula, Represents the visual perception loss function between the edited image and the watermarked image The weight coefficient of Represents the feature-aware loss function between the edited image and the watermarked image The weight coefficient of represents the identity-aware loss function between the edited image and the watermarked image The weight coefficient of The original image quality loss function is expressed as: ; In the formula, Represents the visual perception loss function between the extracted image and the original image The weight coefficient of Represents the feature-aware loss function between the extracted image and the original image The weight coefficient of represents the identity-aware loss function between the extracted image and the original image The weight coefficient of The text quality loss function is expressed as: ; In the formula, represents the L2 norm.
Citation Information
Patent Citations
Image blind watermarking method based on reversible neural network
CN114140308A
Image watermark embedding and extracting method, device and equipment and readable storage medium
CN115564633A
Reversible adversarial sample generation method based on self-embedded watermark
CN118115343A
Lithology image data digital watermark processing method and system
CN118283195A
Multi-modal model watermarking method based on synonym replacement
CN119691710A
Cited By
Image copyright protection method, system, equipment and medium
CN120493226A
A method, system, device and medium for image copyright protection
CN120493226B
Image watermark embedding method and device, electronic equipment and storage medium
CN121074565A