A digital image watermarking method and device based on diffusion model and spread spectrum technology

By adopting a diffusion model and spread spectrum technology method in digital image watermarking technology, the problem of insufficient robustness in the face of new attacks and editing is solved, and the unperception and robustness are achieved, and it can effectively resist a variety of image processing operations and attacks.

CN119741183BActive Publication Date: 2025-05-13HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510252537.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-13
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

When existing digital image watermarking technology faces new regeneration attacks and diffusion models to edit local or global content of the image, the watermark is less robust and has poor visual inesensibility.

Method used

Using a digital image watermarking method based on diffusion model and spread spectrum technology, the variable autoencoder is used to downsample and forward diffusion, high frequency coefficients are extracted to form an embedding matrix, and the watermark information is spread spectrum modulated using an orthogonal code, and the watermark potential vector is embedded in the watermark with intensity factors, and reverse denoising is used to generate the final watermark image.

Benefits of technology

It enhances the invisibility and robustness of the watermark, and can resist image compression, noise addition, brightness and contrast adjustment, low-pass filtering and other operations, as well as regenerative attacks and diffusion model editing, significantly improving the concealment and attack resistance of the watermark.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741183B_ABST
    Figure CN119741183B_ABST
Patent Text Reader

Abstract

The present invention provides a digital image watermarking method and device based on a diffusion model and spread spectrum technology, which relates to the field of image processing and artificial intelligence technology. The present invention obtains a carrier image and watermark information; uses a variational autoencoder to downsample and forward diffuse the carrier image to obtain a first latent vector; divides the first latent vector into channel dimensions, selects a one-dimensional channel therein, and performs frequency domain transformation through discrete cosine transform technology, extracts high-frequency coefficients, and obtains an embedding matrix; uses an orthogonal code to spread spectrum modulate the watermark information to obtain a spread spectrum watermark; combines the intensity factor to embed the spread spectrum watermark into the embedding matrix to obtain a watermark latent vector; uses a denoising network to reversely denoise the watermark latent vector, restores it to a watermark image through a variational autoencoder, and combines it with the carrier image to obtain the final watermark image. The present invention greatly enhances the imperceptibility and robustness of the watermark, and can also effectively resist regenerative attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing and artificial intelligence technology, and in particular to a digital image watermarking method and device based on a diffusion model and spread spectrum technology. Background Art

[0002] Digital image watermarking technology protects the copyright of digital media by embedding specific information (i.e., watermark) in the image. With the development of large artificial intelligence models, watermark information hidden in the image can be removed by regenerating the image through image generation models or editing the local or global content of the image with semantics. Regeneration attacks can inject noise into the latent space of the image, destroy the original distribution of the image, and then denoise it to regenerate the image. This regeneration process can remove the watermark while ensuring the fidelity of the image. On this basis, image editing can use semantic information to guide the denoising process, edit and tamper with the local or global content of the image, and remove the watermark at the same time, which brings severe challenges to the protection of image copyright.

[0003] At present, there are two types of digital watermarking technologies based on deep learning. One type is to improve the robustness of the watermark and minimize the embedding loss by designing the codec structure and training the noise layer of the codec. The other type is based on the diffusion model and implements watermark embedding or embedding operations on the latent vector by fine-tuning the decoder of the variational autoencoder in the framework. The drawbacks of these two types of technologies are that the watermark is not robust enough in the face of new types of regeneration attacks and when the diffusion model is used to edit the local or global content of the image. In addition, the watermark trained by the noise layer is less visually imperceptible.

[0004] In view of this, the applicant filed this application after studying the existing technology. Summary of the invention

[0005] The present invention aims to provide a digital image watermarking method, device, equipment and medium based on diffusion model and spread spectrum technology to solve the shortcomings of the watermark generated in the existing method, such as poor visual imperceptibility and low robustness of the watermark.

[0006] In order to solve the above technical problems, the present invention is implemented through the following technical solutions:

[0007] A digital image watermarking method based on a diffusion model and spread spectrum technology, comprising:

[0008] S1, obtain carrier image and watermark information;

[0009] S2, using a variational autoencoder to downsample and forward diffuse the carrier image to obtain a first latent vector;

[0010] S3, dividing the first latent vector into channel dimensions, selecting one-dimensional channels therein, performing frequency domain transformation by discrete cosine transform technology, extracting high-frequency coefficients, and obtaining an embedding matrix;

[0011] S4, performing spread spectrum modulation on the watermark information using an orthogonal code to obtain a spread spectrum watermark;

[0012] S5, embedding the spread spectrum watermark into the embedding matrix in combination with the strength factor to obtain a watermark latent vector;

[0013] S6, after using the denoising network to reversely denoise the watermark latent vector, restore it to the watermark image through the variational autoencoder, and combine it with the carrier image to obtain the final watermark image.

[0014] Preferably, the variational autoencoder includes an encoder and a decoder; the encoder is used to downsample the input carrier image to obtain a corresponding latent space; the decoder is used to restore the latent space representation to the corresponding image representation; wherein the encoder includes:

[0015] Convolutional layer, used to extract low-level features of the input image through convolution operations;

[0016] The downsampling module is used to reduce the dimension of the image through convolution and pooling operations to extract important features;

[0017] The residual block is used to process the input features to ensure that low-level features are not suppressed by high-level networks;

[0018] The middle block is used to integrate the global information of the features after dimensionality reduction and capture the long-range features in the data through the self-attention mechanism;

[0019] The GSC module is used to standardize the features through group normalization and combine it with the Swish activation function to obtain a compact latent space representation;

[0020] The encoder includes a plurality of consecutive downsampling modules;

[0021] The decoder comprises:

[0022] Convolutional layers, used to extract initial features of the latent space representation of the input;

[0023] The middle block, which includes residual blocks and self-attention layers, is used to further process features;

[0024] An upsampling module, including a residual block, an interpolation and a convolution layer, for gradually restoring the resolution of the image and generating a feature map with the same resolution as the carrier image;

[0025] The residual block is used to process the input feature map to capture the complex structure of the data and restore the details of the data;

[0026] GSC module, used to further adjust the generated features;

[0027] The decoder includes several consecutive upsampling modules.

[0028] Preferably, S2 is specifically:

[0029] Downsampling the input carrier image to downsample the carrier image into a low-dimensional representation to obtain an initial latent vector;

[0030] Gaussian noise is gradually added to the initial latent vector to gradually transition the clear data to completely noisy data; and diffusion inversion is used to sample the latent vector with Gaussian noise within a preset time step to obtain a first latent vector, that is, from the initial latent vector To the first latent vector The sampling process is as follows:

[0031] ;

[0032] in, is the potential vector of the Tth step, that is, the first potential vector; is the initial potential vector; is the potential vector of the tth step; is the potential vector predicted at step t+1; , represents the parameter controlling the noise weight at step t; is the hyperparameter corresponding to the i-th step, indicating the noise amplitude; T indicates the total number of steps, and t and i indicate variables; It is a pre-trained U-net denoising network.

[0033] Preferably, S3 is specifically:

[0034] The first potential vector is sliced ​​along the channel dimension to obtain a channel dimension tensor corresponding to the number of channels. The formula is:

[0035] ;

[0036] Each channel dimension tensor is represented as: =1× × ;

[0037] in, is the potential vector of the Tth step, i.e., the first potential vector; T represents the total number of steps; represents the i-th channel dimension tensor; d represents the total number of channels; H and W represent the image height and width respectively; f represents the scaling factor;

[0038] Select one of the one-dimensional channel tensors and use discrete cosine transform on it to extract a high-frequency coefficient in each image block after discrete cosine transform to form an embedding matrix , the embedding matrix It is expressed as:

[0039] ;

[0040] Wherein, n is twice the length of the watermark information; Indicates the nth high-frequency coefficient extracted; if the watermark capacity needs to be expanded, additional high-frequency coefficients in each block are extracted.

[0041] Preferably, the S4 is specifically:

[0042] The watermark information is spread spectrum modulated to expand the spectrum of the watermark, and the expression is:

[0043] ;

[0044] in, is the spread spectrum watermark, ; n is twice the length of the watermark information; represents the watermark information, , k is the length of watermark information; represents a randomly generated orthogonal code.

[0045] Preferably, the S5 is specifically:

[0046] Use Intensity Factor The modulated spread spectrum watermark is signal enhanced, and after an additive operation with the embedding matrix, an inverse discrete cosine transform is performed, so that the selected channel dimension becomes a channel embedded with the watermark;

[0047] The channel embedded with the watermark is connected with the tensors of the remaining channel dimensions to obtain the watermark latent vector.

[0048] Preferably, the S6 is specifically:

[0049] Initialize the input of the image to generate text prompt words;

[0050] Using a denoising network U-net to predict the noise of the input watermark potential vector, and using an implicit denoising model DDIM sampler to perform sampling, and setting the sampling time step of the DDIM sampler;

[0051] The decoder of the variational autoencoder is used to upsample the latent vector output by the denoising network to reconstruct the low-dimensional latent feature representation into the image domain and generate a watermarked image.

[0052] According to the watermark image and the carrier image, the final watermark image is obtained, and the formula is:

[0053] ;

[0054] in, Represents the final watermark image; is the carrier image, is the watermark image; The scale factor to be set.

[0055] Preferably, it also includes using a loss function to optimize the difference between the watermark image and the carrier image to improve the quality of the watermark image and reduce the embedding loss; wherein the loss function includes distance loss, SSIM loss and perceptual loss, and the formula is:

[0056] ;

[0057] Wherein, L is the loss function; , , is the weight factor;

[0058] is the distance loss, which is used to calculate the difference between the watermark image and the carrier image to minimize the quality loss in the process of watermark embedding generation. The formula is:

[0059] ;

[0060] Among them, C, H, and W represent the number of channels, height, and width of the image respectively; is the carrier image, is the watermark image;

[0061] is the SSIM loss, which is used to quantify the SSIM loss between the watermark image and the carrier image to improve the concealment of the watermark and reduce the impact of the embedded watermark on the visual quality. The formula is:

[0062] ;

[0063] Among them, SSIM represents the structural similarity between two images;

[0064] is the perceptual loss, which is used to calculate the difference between the feature maps of the watermark image and the carrier image at different levels in the deep network to simulate the human perception system and improve the fidelity of the image. The formula is:

[0065] ;

[0066] in, , , are the number of channels, height and width of the j-th layer output feature map respectively; Represents the j-th layer pre-trained VGG model network output.

[0067] Preferably, the method further includes extracting a watermark from the final watermark image, the specific steps of which are:

[0068] Performing dimension reduction on the final watermark image by using a variational autoencoder to obtain an inverse initial latent vector;

[0069] Performing forward diffusion on the reverse initial latent vector to obtain a second latent vector;

[0070] Performing channel division on the second latent vector, selecting one of the channel dimensions to perform discrete cosine transform, and obtaining a plurality of image sub-blocks;

[0071] Extract high-frequency coefficients in each image sub-block to form a detection matrix;

[0072] The detection matrix is ​​demodulated using an orthogonal code to restore the original watermark information; wherein, during demodulation, the watermark information is obtained by judging the positive or negative of the inner product operation result of the orthogonal code and the detection matrix; if the inner product operation result is a positive number, the watermark bit corresponding to the watermark information is 1; if it is a negative number, the corresponding watermark bit is 0, and the formula is:

[0073] ;

[0074] in, Represents the i-th watermark bit of the watermark information; represents the inner product operation; is the detection matrix, is an orthogonal code.

[0075] The present invention also provides a digital image watermarking device based on a diffusion model and spread spectrum technology, comprising:

[0076] An acquisition unit, used for acquiring a carrier image and watermark information;

[0077] A first latent vector unit, used to downsample and forward diffuse the carrier image using a variational autoencoder to obtain a first latent vector;

[0078] An embedding matrix unit is used to divide the first latent vector into channel dimensions, select one of the one-dimensional channels, perform frequency domain transformation through discrete cosine transform technology, extract high-frequency coefficients, and obtain an embedding matrix;

[0079] A watermark spreading unit, used to perform spread spectrum modulation on the watermark information using an orthogonal code to obtain a spread spectrum watermark;

[0080] An embedding unit, used to embed the spread spectrum watermark into the embedding matrix in combination with a strength factor to obtain a watermark latent vector;

[0081] The generating unit is used to perform reverse denoising on the watermark latent vector by using a denoising network, restore it into a watermark image by using a variational autoencoder, and combine it with the carrier image to obtain a final watermark image.

[0082] The present invention also provides a digital image watermarking device based on a diffusion model and spread spectrum technology, comprising a processor and a memory, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement a digital image watermarking method based on a diffusion model and spread spectrum technology as described above.

[0083] The present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a digital image watermarking method based on a diffusion model and spread spectrum technology as described above is implemented.

[0084] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0085] The present invention realizes the covert embedding of watermarks without the need to train an additional watermark encoding and decoding network, and uses the inversion denoising process of a diffusion model as an embedding process, thereby enhancing the imperceptibility of the watermark.

[0086] The present invention combines spread spectrum watermarking and latent space embedded watermarking. The generated watermark image can not only resist common attacks, but also resist operations such as image compression, image noise addition, brightness and contrast adjustment, low-pass filtering, etc. on the watermark image, and can also resist regenerative attacks on the watermark image and local or global content editing of the image using a diffusion model, which greatly enhances the robustness of the watermark image. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0088] Figure 1 A schematic diagram of a digital image watermarking method based on a diffusion model and spread spectrum technology provided in Example 1.

[0089] Figure 2 The present invention is a flowchart of a digital image watermarking method based on a diffusion model and spread spectrum technology provided in the first embodiment.

[0090] Figure 3 A schematic diagram of the structure of the variational autoencoder provided in Example 1.

[0091] Figure 4 This is a flow chart of watermark embedding provided in Example 1.

[0092] Figure 5 A schematic diagram of a digital image watermarking device based on a diffusion model and spread spectrum technology provided in Example 2.

[0093] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. DETAILED DESCRIPTION

[0094] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0095] Embodiment 1

[0096] Embodiment 1 of the present invention provides a digital image watermarking method based on a diffusion model and spread spectrum technology, which can be implemented by a digital image watermarking device based on a diffusion model and spread spectrum technology (hereinafter referred to as a watermarking device), and in particular, executed by one or more processors in the watermarking device.

[0097] In this embodiment, the watermark device may be an electronic device equipped with a processor, the processor having a computer program of the digital image watermark method based on the diffusion model and spread spectrum technology and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.

[0098] like Figure 1-Figure 2 As shown, a digital image watermarking method based on a diffusion model and spread spectrum technology includes steps S1 to S6.

[0099] S1, obtain the carrier image and watermark information.

[0100] In this embodiment, the carrier image is the image to be embedded with a watermark, which is usually an image whose copyright or integrity the user wishes to protect.

[0101] Watermark information is information to be embedded in the carrier image, usually used to identify the copyright owner, creation date or other metadata of the image.

[0102] S2, using a variational autoencoder to downsample and forward diffuse the carrier image to obtain a first latent vector.

[0103] In this embodiment, the variational autoencoder VAE is a generative model that reconstructs input data by learning the latent representation of the data.

[0104] Specifically, Figure 3 As shown, the variational autoencoder includes an encoder and a decoder; the encoder is used to downsample the input carrier image to obtain a corresponding latent space; the decoder is used to restore the latent space representation to the corresponding image representation.

[0105] Among them, the encoder includes: a convolution layer, which is used to extract low-level features of the input image through convolution operations; a downsampling module, which is used to reduce the dimension of the image through convolution and pooling operations and extract important features, including residual blocks, pooling layers and convolution layers; the encoder can include several consecutive downsampling modules, such as being set to 3 consecutive downsampling modules; a residual block, which is used to process the input features to ensure that the low-level features are not suppressed by the high-level network, which is helpful for the expression of the latent space; an intermediate block, which is used to integrate the global information of the features after dimensionality reduction and capture the long-distance features in the data through the self-attention mechanism; a GSC module, which is used to standardize the features through group normalization and combined with the Swish activation function. Compared with the ReLU activation function, Swish is smoother and can obtain a more compact latent space representation.

[0106] The decoder includes: a convolution layer, which is used to extract the initial features of the latent space representation of the input; an intermediate block, including a residual block and self-attention, which is used to further process the features; an upsampling module, including a residual block, an interpolation and a convolution layer, which is used to gradually restore the resolution of the image and generate a feature map with the same resolution as the carrier image; the decoder may include several consecutive upsampling modules, such as 3 consecutive upsampling modules, or other numbers according to actual needs, which are not limited here; a residual block, which is used to process the features of the input, capture the complex structure of the input data, and restore the details of the data; a GSC module, which is the same as the GSC module of the encoder, is used to further adjust the generated features, including group normalization, Swish activation function and convolution layer.

[0107] In this embodiment, the role of the residual block is mainly reflected in improving the training effect and expression ability of the model. Specifically, the residual block solves the problem of gradient disappearance or gradient explosion that may occur during the training of deep neural networks by introducing the "skip connection" mechanism, so that the model can learn complex features in the data more efficiently. The core idea of ​​the residual block is to add the input directly to the output after a series of nonlinear transformations to form residual learning and avoid the loss of key information. In this way, when training deep networks, each layer only needs to learn the residual part between the input and output, rather than the entire complex mapping relationship. This not only reduces the burden of gradient transfer, but also makes the model easier to optimize and improves training stability.

[0108] Specifically, after the carrier image is acquired, the input carrier image is downsampled to downsample the carrier image into a low-dimensional representation to obtain an initial latent vector.

[0109] Then, Gaussian noise is gradually added to the initial latent vector to gradually transition the clear data to completely noisy data; and diffusion inversion is used to sample the latent vector with Gaussian noise within a preset time step, that is, , get the first potential vector, the formula is:

[0110] ;

[0111] in, is the potential vector of the Tth step, that is, the first potential vector; is the initial potential vector; is the potential vector of the tth step; is the potential vector predicted at step t+1; , represents the parameter controlling the noise weight at step t; is the hyperparameter corresponding to the i-th step, indicating the noise amplitude; T indicates the total number of steps, and t and i indicate variables; It is a pre-trained U-net denoising network.

[0112] S3, dividing the first latent vector into channel dimensions, selecting one-dimensional channels therein to perform frequency domain transformation through discrete cosine transform technology, extracting high-frequency coefficients, and obtaining an embedding matrix.

[0113] In this embodiment, the different "channels" of the latent vector are similar to the red, green, blue, and transparent channels in a color image.

[0114] Furthermore, the first potential vector is sliced ​​along the channel dimension to obtain a channel dimension tensor corresponding to the number of channels, and the formula is:

[0115] ;

[0116] Each channel dimension tensor is represented as: =1× × ;

[0117] in, is the potential vector of the Tth step, i.e., the first potential vector; T represents the total number of steps; represents the i-th channel dimension tensor; d represents the total number of channels; H and W represent the image height and width respectively; f represents the scaling factor;

[0118] Select one of the one-dimensional channel tensors and use discrete cosine transform on it to extract a high-frequency coefficient in each image block after discrete cosine transform to form an embedding matrix , the embedding matrix It is expressed as:

[0119] ;

[0120] Wherein, n is twice the length of the watermark information; Indicates the nth high-frequency coefficient extracted; if the watermark capacity needs to be expanded, additional high-frequency coefficients in each block are extracted.

[0121] Discrete cosine transform (DCT) is a transform technology commonly used in image compression and frequency domain analysis. Here, it is used to extract the high-frequency components of the image, which usually have little effect on the visual quality of the image and are therefore suitable for embedding watermarks. Among the coefficients obtained after DCT transformation, the high-frequency coefficients represent the details and edge information in the image, and the matrix extracted from the high-frequency coefficients is used to embed watermark information.

[0122] S4, use orthogonal code to spread spectrum modulate the watermark information to obtain a spread spectrum watermark.

[0123] In this embodiment, the orthogonal code is a code sequence used in spread spectrum communication to spread the watermark information to a wider frequency band, thereby improving the robustness and concealment of the watermark.

[0124] Furthermore, the watermark information is spread spectrum modulated to expand the spectrum of the watermark, and the expression is:

[0125] ;

[0126] in, is the spread spectrum watermark, ; n is twice the length of the watermark information; represents the watermark information, , k is the length of watermark information; represents a randomly generated orthogonal code.

[0127] S5, combining with the strength factor, embeds the spread spectrum watermark into the embedding matrix to obtain a watermark latent vector.

[0128] In this embodiment, the strength factor is used to control the parameters of the watermark embedding strength. A larger strength factor may improve the robustness of the watermark, but may also lead to a significant decrease in image quality. Therefore, the strength factor should be set according to actual conditions.

[0129] Use Intensity Factor The modulated spread spectrum watermark is signal enhanced, and after an additive operation with the embedding matrix, an inverse discrete cosine transform is performed, so that the selected channel dimension becomes a channel embedded with a watermark; the channel embedded with a watermark is connected with the tensors of the remaining channel dimensions to obtain the watermark latent vector containing the watermark information.

[0130] S6, after using the denoising network to reversely denoise the watermark latent vector, restore it to the watermark image through the variational autoencoder, and combine it with the carrier image to obtain the final watermark image.

[0131] In this embodiment, the denoising network is a neural network used to remove image noise. Here, it is used to remove the noise that may be introduced by watermark embedding from the watermark latent vector, using the U-net structure, a classic encoder-decoder structure. The U-net is used to predict the noise of the latent vector embedded with the watermark, that is, to predict the noise present in the latent vector and gradually remove it.

[0132] The inverse denoising process is used to recover a clearer image representation from the latent space.

[0133] Furthermore, first, the input of the image generation text prompt is initialized. In image generation or reconstruction tasks, text prompts can be used to guide the generated content. However, in the present invention, the focus is on the embedding and recovery of watermarks rather than text-based image generation, so the text prompts are initialized to blank text prompts. This means that when the model optimizes image quality, it will not be disturbed by external text information and will focus on recovering image content from the latent vector.

[0134] Then, a denoising network U-net is used to predict the noise of the input watermark potential vector, and an implicit denoising model DDIM sampler is used for sampling, and the sampling time step of the DDIM sampler is set;

[0135] The decoder of the variational autoencoder is used to upsample the latent vector output by the denoising network to reconstruct the low-dimensional latent feature representation into the image domain and generate a watermarked image.

[0136] In this embodiment, the encoder path of U-net extracts features through convolution and downsampling, while the decoder gradually restores information through convolution and upsampling. The encoder and decoder fuse multi-level features through skip connections to improve the denoising effect.

[0137] DDIM (Denoising Diffusion Implicit Models) is an improved diffusion model sampling method with higher sampling efficiency and generation quality. Compared with traditional diffusion models, DDIM can complete the denoising process in fewer sampling steps while maintaining high image quality. The output of the denoising network is a tensor, which represents the denoised latent vector. This tensor is used as the input of the optimizer to calculate the loss function and update the model parameters to continuously improve the image generation quality.

[0138] In DDIM, noisy prediction is performed by gradually updating the latent vector to gradually remove the noise and finally obtain a clear latent representation.

[0139] In this step, a loss function may be used to optimize the difference between the watermark image and the carrier image to improve the quality of the watermark image and reduce the embedding loss; wherein the loss function includes distance loss, SSIM loss and perceptual loss, and the formula is:

[0140] ;

[0141] Wherein, L is the loss function; , , is the weight factor;

[0142] is the distance loss, which is used to calculate the difference between the watermark image and the carrier image to minimize the quality loss in the process of watermark embedding generation. The formula is:

[0143] ;

[0144] Among them, C, H, and W represent the number of channels, height, and width of the image respectively; is the carrier image, is the watermark image;

[0145] is the SSIM loss, which is used to quantify the SSIM loss between the watermark image and the carrier image to improve the concealment of the watermark and reduce the impact of the embedded watermark on the visual quality. The formula is:

[0146] ;

[0147] Among them, SSIM represents the structural similarity between two images;

[0148] is the perceptual loss, which is used to calculate the difference between the feature maps of the watermark image and the carrier image at different levels in the deep network to simulate the human perception system and improve the fidelity of the image. The formula is:

[0149] ;

[0150] in, , , are the number of channels, height and width of the j-th layer output feature map respectively; Represents the j-th layer pre-trained VGG model network output.

[0151] In this embodiment, the VGG model network, whose full name is Visual Geometry Group network, is a very classic convolutional neural network (CNN) architecture in the field of deep learning. Its working principle is based on the basic idea of ​​convolutional neural network. It extracts image features through multi-layer convolution and pooling operations and can be used for tasks such as face recognition, object detection, and image segmentation.

[0152] Finally, the final watermark image is obtained based on the watermark image obtained after training optimization and the carrier image. , the formula is:

[0153] ;

[0154] in, Represents the final watermark image; is the carrier image, is the watermark image; The scale factor to be set.

[0155] At this point, the watermark information is embedded and the final watermark image is obtained. In practical applications, watermark extraction is also involved.

[0156] When extracting the watermark, the watermark image is processed similarly to the embedding process to obtain the corresponding detection matrix, and the orthogonal code is used for demodulation to obtain the original watermark information. The specific steps are as follows:

[0157] Performing dimension reduction on the final watermark image by using a variational autoencoder to obtain an inverse initial latent vector;

[0158] Performing forward diffusion on the reverse initial latent vector to obtain a second latent vector;

[0159] Performing channel division on the second latent vector, selecting one of the channel dimensions to perform discrete cosine transform, and obtaining a plurality of image sub-blocks;

[0160] Extract high-frequency coefficients in each image sub-block to form a detection matrix;

[0161] The detection matrix is ​​demodulated using an orthogonal code to restore the original watermark information; wherein, during demodulation, the watermark information is obtained by judging the positive or negative of the inner product operation result of the orthogonal code and the detection matrix; if the inner product operation result is a positive number, the watermark bit corresponding to the watermark information is 1; if it is a negative number, the corresponding watermark bit is 0, and the formula is:

[0162] ;

[0163] in, Represents the i-th watermark bit of the watermark information; represents the inner product operation; is the detection matrix, is an orthogonal code.

[0164] In another preferred embodiment, see Figure 4 The watermark embedding flow chart shown in the figure first embeds the watermark in the latent space and then generates the image, as follows:

[0165] 1) Assume that the input carrier image is 3×512×512, indicating 3 channels, with a width and height of 512 respectively. The downsampling ratio of the variational autoencoder is set to 8, and the downsampling is compressed to the corresponding latent space of 4×64×64 to obtain the initial latent vector.

[0166] 2) Gradually add Gaussian noise to the initial latent vector, that is, perform a forward diffusion operation, add 50 steps of Gaussian noise, and use diffusion inversion to model the process of directly sampling the noisy latent vector at a specific time step to obtain the first latent vector ;

[0167] 3) Next, divide the channel dimension and divide the first potential vector Tensor slicing is performed along the channel dimension, i.e. , get 4 1×64×64 channel dimension tensors, select the fourth channel dimension tensor , using 8×8 discrete cosine transform DCT, we get 64 sub-blocks of size 8×8; extract the coefficients with coordinates (7, 7) in each sub-block, i.e., high-frequency coefficients, to form an embedding matrix If the watermark capacity needs to be doubled, the coefficients with coordinates (6, 6) in each sub-block are additionally extracted to form the embedding matrix ;

[0168] 4) Assuming that a 32-bit binary watermark is randomly generated, an orthogonal code with a dimension of (32, 64) is used to spread spectrum modulate the watermark to obtain a 64-bit spread spectrum watermark; if a 64-bit binary watermark is randomly generated, an orthogonal code with a dimension of (64, 128) is used to spread spectrum modulate the watermark to obtain a 128-bit spread spectrum watermark. The length of the watermark can be set according to actual needs and is not limited here. Assume that the strength factor S = 70, perform multiplication with it, and then embed the watermark with the adjusted strength into , and perform 8×8 inverse discrete cosine transform IDCT to obtain the channel tensor embedded with watermark;

[0169] 5) Connect the channel tensor embedded with the watermark with the tensors of the remaining three channel dimensions to obtain the latent vector embedded with the watermark , the watermark latent vector, is used as the input of the reverse denoising network.

[0170] 6) Next, perform image generation. Initialize the input of the image generation text prompt and set it to a blank text prompt;

[0171] 7) The input watermarked latent vector (the watermarked latent vector) is subjected to noise prediction by the denoising network U-net, and the implicit denoising model (DDIM) sampler is used, and the sampling time step T is set to 50 (other values ​​can also be set according to actual needs, which are not limited here), and the model output mode is set to tensor for optimizer training;

[0172] 8) For the latent vector output by the denoising network, the decoder of the variational autoencoder is used to reconstruct the low-dimensional latent feature representation 4×64×64 into the original image representation 3×512×512, i.e., the watermark image;

[0173] 9) After obtaining the generated watermark image, the three loss functions can be used to optimize the parameters to improve the image quality. By calculating the total loss function , where the weight factor , , Set them to 10, 1, and 0.1 respectively. You can also set other values ​​according to actual needs. There is no limit here. Update the parameters of the latent vector by minimizing the total loss function, and then add part of the original image information to obtain the final watermark image. ,in, Set it to 0.45. You can also set other values ​​according to actual needs. There is no limitation here.

[0174] In this embodiment, after obtaining the final watermark image, watermark extraction can be performed on it, and the steps are as follows:

[0175] The final watermark image 3×512×512 is reduced in dimension by the encoder of the variational autoencoder to obtain the corresponding latent vector 4×64×64, i.e., the reverse initial latent vector. The obtained reverse initial latent vector is forward processed for 50 steps to obtain the second latent vector.

[0176] Perform channel slicing on the second latent vector, select the fourth channel dimension 1×64×64 for 8×8 discrete cosine transform, and obtain 64 sub-blocks of size 8×8; extract the coefficients with coordinates (7, 7) in each block to form the detection matrix ; If the 64-bit watermark information is restored, the coefficients with coordinates (6, 6) in each block are additionally extracted to form the detection matrix ;

[0177] Use orthogonal codes The original watermark information is restored by demodulating it to obtain the decoded watermark information.

[0178] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0179] The present invention uses a variational autoencoder to downsample the input carrier image and gradually add Gaussian noise to obtain a first latent vector; divide the first latent vector into channel dimensions, use discrete cosine transformation on the selected dimension and select high-frequency coefficients to obtain an embedding matrix; use orthogonal codes to spread spectrum modulate the watermark information to obtain a spread spectrum watermark; embed the watermark in combination with the intensity factor and perform an inverse discrete cosine transform to obtain a latent vector embedded with the watermark; use a denoising network to denoise the watermark latent vector and use the decoder of the variational autoencoder to restore the image, and optimize the image quality through training. In the process of watermark extraction, the corresponding detection matrix of the embedded watermark image is demodulated using a spread spectrum code to restore the original watermark information.

[0180] The digital image watermarking method adopted in the present invention overcomes the defects of the prior art in that the image watermark is insufficiently robust against new types of image regeneration attacks and the use of diffusion models to edit and tamper with local or global content information of the image. At the same time, it has the robustness against common operation attacks such as image compression, image noise addition, brightness and contrast adjustment, low-pass filtering, etc. It has many applicable scenarios and strong practicality.

[0181] Embodiment 2

[0182] like Figure 5 As shown, the second embodiment of the present invention further provides a digital image watermarking device based on a diffusion model and spread spectrum technology, comprising:

[0183] An acquisition unit, used for acquiring a carrier image and watermark information;

[0184] A first latent vector unit, used to downsample and forward diffuse the carrier image using a variational autoencoder to obtain a first latent vector;

[0185] An embedding matrix unit is used to divide the first latent vector into channel dimensions, select one of the one-dimensional channels, perform frequency domain transformation through discrete cosine transform technology, extract high-frequency coefficients, and obtain an embedding matrix;

[0186] A watermark spreading unit, used to perform spread spectrum modulation on the watermark information using an orthogonal code to obtain a spread spectrum watermark;

[0187] An embedding unit, used to embed the spread spectrum watermark into the embedding matrix in combination with a strength factor to obtain a watermark latent vector;

[0188] The generating unit is used to perform reverse denoising on the watermark latent vector by using a denoising network, restore it into a watermark image by using a variational autoencoder, and combine it with the carrier image to obtain a final watermark image.

[0189] Embodiment 3

[0190] The third embodiment of the present invention also provides a digital image watermarking device based on a diffusion model and spread spectrum technology, which includes a memory and a processor. The memory stores a computer program, and the computer program can be executed by the processor to implement the digital image watermarking method based on the diffusion model and spread spectrum technology as described above.

[0191] Embodiment 4

[0192] The fourth embodiment of the present invention further provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, the digital image watermarking method based on the diffusion model and spread spectrum technology as described above is implemented.

[0193] In several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus and method embodiments described above are merely schematic. For example, the flowcharts in the accompanying drawings show the possible architecture, functions and operations of the apparatus, method and computer program product according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0194] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0195] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code. It should be noted that in this article, the term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or device. Without more constraints, an element defined by the phrase "comprising a..." does not exclude the existence of other identical elements in the process, method, article or apparatus comprising the element.

[0196] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0197] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0198] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0199] The "first\second" mentioned in the embodiments is only to distinguish similar objects, and does not represent a specific order for the objects. It is understandable that the "first\second" can be interchanged with the specific order or sequence where permitted. It should be understood that the objects distinguished by "first\second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0200] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A digital image watermarking method based on diffusion model and spread spectrum technology, characterized in that: include: S1, obtain carrier image and watermark information; S2, using a variational autoencoder to downsample and forward diffuse the carrier image to obtain a first latent vector; specifically: Downsampling the input carrier image to downsample the carrier image into a low-dimensional representation to obtain an initial latent vector; Gaussian noise is gradually added to the initial latent vector to gradually transition the clear data to completely noisy data; and diffusion inversion is used to sample the latent vector with Gaussian noise within a preset time step to obtain a first latent vector, that is, from the initial latent vector To the first latent vector The sampling process is as follows: ; in, is the potential vector of the Tth step, that is, the first potential vector; is the initial potential vector; is the potential vector of the tth step; is the potential vector predicted at step t+1; , represents the parameter controlling the noise weight at step t; is the hyperparameter corresponding to the i-th step, indicating the noise amplitude; T indicates the total number of steps, and t and i indicate variables; It is a pre-trained U-net denoising network; S3, dividing the first latent vector into channel dimensions, selecting one-dimensional channels, performing frequency domain transformation through discrete cosine transform technology, extracting high-frequency coefficients, and obtaining an embedding matrix; specifically: The first potential vector is sliced ​​along the channel dimension to obtain a channel dimension tensor corresponding to the number of channels. The formula is: ; Each channel dimension tensor is represented as: =1× × ; in, is the potential vector of the Tth step, i.e., the first potential vector; T represents the total number of steps; represents the i-th channel dimension tensor; d represents the total number of channels; H and W represent the image height and width respectively; f represents the scaling factor; Select one of the one-dimensional channel tensors and use discrete cosine transform on it to extract a high-frequency coefficient in each image block after discrete cosine transform to form an embedding matrix , the embedding matrix It is expressed as: ; Wherein, n is twice the length of the watermark information; Represents the nth high-frequency coefficient extracted; S4, use orthogonal code to spread spectrum modulate the watermark information to obtain a spread spectrum watermark, which is expressed as: ; in, is the spread spectrum watermark, ; n is twice the length of the watermark information; represents the watermark information, , k is the length of watermark information; represents a randomly generated orthogonal code; S5, embedding the spread spectrum watermark into the embedding matrix in combination with the strength factor to obtain a watermark latent vector; S6, after using the denoising network to reversely denoise the watermark latent vector, restore it to the watermark image through the variational autoencoder, and combine it with the carrier image to obtain the final watermark image.

2. According to claim 1, a digital image watermarking method based on diffusion model and spread spectrum technology is characterized in that ,The variational autoencoder includes an encoder and a decoder; the encoder is used to downsample the input carrier image to obtain the corresponding latent space; The decoder is used to restore the latent space representation to the corresponding image representation; wherein the encoder comprises: Convolutional layer, used to extract low-level features of the input image through convolution operations; The downsampling module is used to reduce the dimension of the image through convolution and pooling operations to extract important features; The residual block is used to process the input features to ensure that low-level features are not suppressed by high-level networks; The middle block is used to integrate the global information of the features after dimensionality reduction and capture the long-range features in the data through the self-attention mechanism; The GSC module is used to standardize the features through group normalization and combine it with the Swish activation function to obtain a compact latent space representation; The encoder includes a plurality of consecutive downsampling modules; The decoder comprises: Convolutional layers, used to extract initial features of the latent space representation of the input; The middle block, which includes the residual block and the self-attention layer, is used to process features; An upsampling module, including a residual block, an interpolation and a convolution layer, for gradually restoring the resolution of the image and generating a feature map with the same resolution as the carrier image; The residual block is used to process the input feature map to capture the complex structure of the data and restore the details of the data; The GSC module, which is the same as the encoder’s GSC module, is used to adjust the generated features; The decoder includes several consecutive upsampling modules.

3. According to claim 1, a digital image watermarking method based on diffusion model and spread spectrum technology is characterized in that , the S5 is specifically: Use Intensity Factor The modulated spread spectrum watermark is signal enhanced, and after an additive operation with the embedding matrix, an inverse discrete cosine transform is performed, so that the selected channel dimension becomes a channel embedded with the watermark; The channel embedded with the watermark is connected with the tensors of the remaining channel dimensions to obtain the watermark latent vector.

4. According to claim 2, a digital image watermarking method based on diffusion model and spread spectrum technology is characterized in that , the S6 is specifically: Initialize the input of the image to generate text prompt words; Using a denoising network U-net to predict the noise of the input watermark potential vector, and using an implicit denoising model DDIM sampler to perform sampling, and setting the sampling time step of the DDIM sampler; The decoder of the variational autoencoder is used to upsample the latent vector output by the denoising network to reconstruct the low-dimensional latent feature representation into the image domain and generate a watermarked image. According to the watermark image and the carrier image, the final watermark image is obtained, and the formula is: ; in, Represents the final watermark image; is the carrier image, is the watermark image; The scale factor to be set.

5. According to claim 1, a digital image watermarking method based on diffusion model and spread spectrum technology is characterized in that , and also includes using a loss function to optimize the difference between the watermark image and the carrier image to improve the quality of the watermark image and reduce the embedding loss; wherein the loss function includes distance loss, SSIM loss and perceptual loss, and the formula is: ; Wherein, L is the loss function; , , is the weight factor; is the distance loss, which is used to calculate the difference between the watermark image and the carrier image to minimize the quality loss in the process of watermark embedding generation. The formula is: ; Among them, C, H, and W represent the number of channels, height, and width of the image respectively; is the carrier image, is the watermark image; is the SSIM loss, which is used to quantify the SSIM loss between the watermark image and the carrier image to improve the concealment of the watermark and reduce the impact of the embedded watermark on the visual quality. The formula is: ; Among them, SSIM represents the structural similarity between two images; is the perceptual loss, which is used to calculate the difference between the feature maps of the watermark image and the carrier image at different levels in the deep network to simulate the human perception system and improve the fidelity of the image. The formula is: ; in, , , are the number of channels, height and width of the j-th layer output feature map respectively; Represents the j-th layer pre-trained VGG model network output.

6. A digital image watermarking method based on diffusion model and spread spectrum technology according to any one of claims 1 to 5, characterized in that , and also includes extracting watermark from the final watermark image, the specific steps are: Performing dimension reduction on the final watermark image by using a variational autoencoder to obtain an inverse initial latent vector; Performing forward diffusion on the reverse initial latent vector to obtain a second latent vector; Performing channel division on the second latent vector, selecting one of the channel dimensions to perform discrete cosine transform, and obtaining a plurality of image sub-blocks; Extract high-frequency coefficients in each image sub-block to form a detection matrix; The detection matrix is ​​demodulated using an orthogonal code to restore the original watermark information; wherein, during demodulation, the watermark information is obtained by judging the positive or negative of the inner product operation result of the orthogonal code and the detection matrix; if the inner product operation result is a positive number, the watermark bit corresponding to the watermark information is 1; if it is a negative number, the corresponding watermark bit is 0, and the formula is: ; in, Represents the i-th watermark bit of the watermark information; represents the inner product operation; is the detection matrix, is an orthogonal code.

7. A digital image watermarking device based on diffusion model and spread spectrum technology, characterized in that: include: An acquisition unit, used for acquiring a carrier image and watermark information; The first latent vector unit is used to downsample and forward diffuse the carrier image using a variational autoencoder to obtain a first latent vector; specifically: Downsampling the input carrier image to downsample the carrier image into a low-dimensional representation to obtain an initial latent vector; Gaussian noise is gradually added to the initial latent vector to gradually transition the clear data to completely noisy data; and diffusion inversion is used to sample the latent vector with Gaussian noise within a preset time step to obtain a first latent vector, that is, from the initial latent vector To the first latent vector The sampling process is as follows: ; in, is the potential vector of the Tth step, that is, the first potential vector; is the initial potential vector; is the potential vector of the tth step; is the potential vector predicted at step t+1; , represents the parameter controlling the noise weight at step t; is the hyperparameter corresponding to the i-th step, indicating the noise amplitude; T indicates the total number of steps, and t and i indicate variables; It is a pre-trained U-net denoising network; The embedding matrix unit is used to divide the first potential vector into channel dimensions, select one-dimensional channels therein, perform frequency domain transformation through discrete cosine transform technology, extract high-frequency coefficients, and obtain an embedding matrix; specifically: The first potential vector is sliced ​​along the channel dimension to obtain a channel dimension tensor corresponding to the number of channels. The formula is: ; Each channel dimension tensor is represented as: =1× × ; in, is the potential vector of the Tth step, i.e., the first potential vector; T represents the total number of steps; represents the i-th channel dimension tensor; d represents the total number of channels; H and W represent the image height and width respectively; f represents the scaling factor; Select one of the one-dimensional channel tensors and use discrete cosine transform on it to extract a high-frequency coefficient in each image block after discrete cosine transform to form an embedding matrix , the embedding matrix It is expressed as: ; Wherein, n is twice the length of the watermark information; Represents the nth high-frequency coefficient extracted; The watermark spreading unit is used to spread spectrum modulate the watermark information using an orthogonal code to obtain a spread spectrum watermark, which is expressed as: ; in, is the spread spectrum watermark, ; n is twice the length of the watermark information; represents the watermark information, , k is the length of watermark information; represents a randomly generated orthogonal code; An embedding unit, used to embed the spread spectrum watermark into the embedding matrix in combination with a strength factor to obtain a watermark latent vector; The generating unit is used to perform reverse denoising on the watermark latent vector by using a denoising network, restore it into a watermark image by using a variational autoencoder, and combine it with the carrier image to obtain a final watermark image.

Citation Information

Patent Citations

  • Improved additive spread spectrum watermarking method

    CN106599630A

  • Picture watermark protection method based on diffusion picture editing model

    CN117350910A