A Robust Image Watermarking Method Based on Generative Adversarial Networks
Through the encoder-decoder framework and adversarial learning of the adversarial network, the robustness and invisibility problems of image watermarks in the prior art are solved, and the effect of efficiently extracting secret information under noise attacks is achieved.
Patent Information
- Application Number
- CN202210598258.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-30
AI Technical Summary
Existing image watermarking techniques have shortcomings in image quality and accuracy in secret information extraction, especially in the face of noise attacks, which are difficult to maintain high robustness and invisibility.
A robust image watermark method based on a generative adversarial network is constructed. Through the encoder-decoder framework and adversarial network framework, binary secret information is embedded in three-channel color images, and through the adversarial learning of the generator and discriminator, combined with simulation attack training model, the model parameters are optimized using a composite loss function to ensure image invisibility and the accuracy of secret information extraction.
It realizes high robustness and high invisibility of accurately extracting secret information under noise attacks, and improves the encoding image quality and decoding accuracy of image watermarking technology.
Smart Images

Figure CN115131188B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of image processing and information hiding, and particularly relates to a robust image watermarking method based on a generative adversarial network. Background Art
[0002] As one of the most important methods of information hiding, digital watermarking technology hides secret information into a carrier by using digital watermarking technology, and achieves the purpose of copyright protection by extracting the secret information embedded in the encoded image. This research field has many useful applications. For example, the ownership of intellectual property [1] and neural network models [2] can be protected by digital watermarking technology. Currently, many researchers have applied convolutional neural networks (CNNs) to digital watermarking technology and steganography. Volkhonskiy et al. [3] first proposed a new model [4] based on a deep convolutional generative adversarial network (DCGAN), called the steganography generative adversarial network (SGAN) model. This model uses the image generated by the GAN generator as the original image and adopts a steganalysis network model as the discriminator, making the steganography model more secure. Shi et al. [5] proposed a secure steganography model based on GAN (SSGAN). Compared with the SGAN model, SSGAN can converge faster and generate higher-quality and more secure encoded images.
[0003] Different from the SGAN and SSGAN models, Hayes et al. [6] proposed an image steganography technology framework based on GAN, called HayesGAN, which includes three sub-network components: an encoder, a decoder, and a discriminator. However, HayesGAN does not fully consider the image quality of the encoded image and the gap between the encoded image and the real original image, so the invisibility of the encoded image is poor. Hu et al. added a steganalysis network based on the HayesGAN model [7] to improve the quality and security of the generated stego images. However, the above models have poor accuracy in extracting secret information.
[0004] Zhu et al. [8] proposed the HiDDeN steganography model based on the HayesGAN model, which can accurately extract the embedded secret information under various attacks such as pixel random dropout, cropping, and Gaussian smoothing. Tang et al. [9] combined adaptive steganography with the GAN model to find a suitable position for the embedding of secret information and proposed an automatic steganography distortion learning framework (ASDL-GAN). Yang et al.
[10] improved the ASDL-GAN model. In the selection of the activation function, the ternary embedding simulator (TES) was replaced with the tanh activation function, which solved the backpropagation problem of TES during network training and improved the security of the model.
[0005] Liu et al.
[11] proposed a two-stage deep learning robust watermarking model. The main difference from the end-to-end model [8] is that the embedding and extraction of watermarks are divided into two stages. The advantage of two-stage training is that it is not necessary to simulate noise attacks as differentiable noise layers, which enables the decoder to better resist noise attacks such as JPEG compression that are difficult to directly model as differentiable network layers. Matthew Tancik et al.
[12] proposed the StegaStamp model structure for image watermarking that can resist attacks such as printing, photographing, rotation, and JPEG compression. Currently, deep learning-based watermarking models have also been applied in the field of audio and video to ensure the data security of various carriers, and the application scope is constantly expanding.
[0006] The purpose of steganography is to hide secret information in common carriers and conceal the secret information carried therein through non-suspicious multimedia carriers. Similar to it is a technology called digital watermarking. The difference between the two is that steganography usually focuses more on the secret information itself because steganography aims to transmit secret information without suspicion through the cover of common media, and the key lies in the secret information not being discovered; digital watermarking also hides information into carriers, but the focus is on protecting the multimedia data itself, and the embedded watermark information is used to identify the image ownership and protect intellectual property rights. The development of Internet technology and intelligent devices has generated a large amount of multimedia data, such as images, videos, and audios. These data may be copied, transmitted, and spread on the network without the permission of the original author, causing damage to the interests of the data owners. Watermarking technology can complete copyright tracking and identify the ownership of data through the secret information embedded in multimedia data, thus minimizing data piracy and abuse as much as possible. Steganography usually assumes that the transmission channel is lossless, while digital watermarking often has requirements for robustness. Even if the stego-image is subjected to noise attacks, the watermark can still be extracted from it to prove the ownership of the image itself. In recent years, the development of traditional information hiding technologies has been slow. Many scholars have combined deep learning technologies with information hiding and achieved a series of research results in terms of methods and performance, providing a good foundation for the research of information hiding technologies, but there are also problems such as color distortion of the stego-image, weak anti-attack ability, and limited capacity for embedding secret information.
[0007] [1] X.Cao, J.Jia, and N.Z.Gong, “IPGuard: Protecting intellectual property of deep neural networks via fingerprinting the classification boundary,” Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security, pp.14 - 25, 2021.
[0008] [2] J.Zhang, D.Chen, J.Liao, W.Zhang and H.Feng, “Deep Model Intellectual Property Protection via Deep Watermarking,” arXiv preprint arXiv:2103.04980, 2021.
[0009] [3] D.Volkhonskiy, I.Nazarov, and E.Burnaev, “Steganographic Generative Adversarial Networks,” arXiv preprint arXiv:1703.05502, 2017.
[0010] [4] A.Radford, L.Metz, and S.Chintala, “Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks,” International Conference on Image and Graphics, pp.97 - 108, 2015.
[0011] [5] H.Shi, J.Dong, W.Wang, Y.Qian, and X.Zhang, “SSGAN: Secure Steganography Based on Generative Adversarial Networks,” Pacific Rim Conference on Multimedia, pp.534 - 544, 2017.
[0012] [6] J.Hayes and G.Danezis, “Generating steganographic images via adversarial training,” arXiv preprint arXiv:1703.0037, 2017.
[0013] [7] D.Hu, L.Wang, W.Jiang, S.Zheng, and B.Li, “A novel image steganography method via deep convolutional generative adversarial networks,” IEEE Access, vol.6, pp.38303 - 38314, 2018.
[0014] [8] J.Zhu, R.Kaplan, J.Johnson, and Fei - Fei.L, “HiDDeN: Hiding Data With Deep Networks,” Proceedings of the European conference on computer vision, pp.657 - 672, 2018.
[0015] [9] W.Tang, S.Tan, B.Li, and J.Huang, “Automatic Steganographic Distortion Learning Using a Generative Adversarial Network,” IEEE Signal Processing Letters, vol.24, no.99, pp.1547 - 1551, 2017.
[0016]
[10] J.Yang, D.Ruan, J.Huang, X.Kang, and Y.Q.Shi, “An Embedding Cost Learning Framework Using GAN,” IEEE Transactions on Information Forensics and Security, vol.15, no.99, pp.839 - 851, 2019.
[0017]
[11] Y. Liu, M. Guo, J. Zhang, Y. Zhu, and X. Xie, “A novel two-stage separable deep learning framework for practical blind watermarking,” Proceedings of the 27th ACM International Conference on Multimedia, pp. 1509-1517, 2019.
[0018]
[12] M. Tancik, B. Mildenhall, and R. Ng, “StegaStamp: Invisible Hyperlinks in Physical Photographs,” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 2117-2126, 2019. Summary of the Invention
[0019] Object of the Invention: Aiming at the problems existing in the prior art, the present invention proposes a robust image watermarking method based on a generative adversarial network. The secret information of binary bits is automatically embedded into a three-channel color image through the automatic learning of a convolutional neural network, which has good invisibility of the encoded image, and the secret information can still be accurately extracted after the encoded image is attacked.
[0020] Technical Solution: To achieve the object of the present invention, the technical solution adopted by the present invention is: A robust image watermarking method based on a generative adversarial network, including:
[0021] (1) Construct an encoder-decoder framework. The encoder embeds the binary secret information into a three-channel carrier image, and the decoder extracts the secret information from the stego image;
[0022] (2) Construct an adversarial network framework. The encoder network and the decoder network form an end-to-end model as the generator network. The stego image is generated and the secret information is extracted through the generator network, and steganalysis is performed through the discriminator network to judge the authenticity of the generated encoded image;
[0023] (3) Through the mutual adversarial learning between the generator network and the discriminator network, the error between the generated stego image and the original image is within a certain range, and at the same time, the error between the decoded secret information and the original secret information is within a certain range;
[0024] (4) Train the network model using the obtained natural image dataset. Verify the invisibility of the model through the structural similarity index and peak signal-to-noise ratio, and verify the robustness of the model through the accuracy of extracting the secret information to obtain the optimal model parameters. Obtain the robust encoded image and extract the secret information through the trained network model.
[0025] Further, the adversarial training of the generator and discriminator includes:
[0026] (a) The generator network consists of two parts: an encoder and a decoder. The encoded image generated by the encoder network should have an error within a certain range from the original image.
[0027] (b) The discriminator network is used for steganalysis, that is, to distinguish whether the image is an encoded image or an original image and to determine whether the image contains secret information.
[0028] (c) Simulated noise network: There is a simulated attack layer before inputting the encoded image into the decoding network to perform a simulated attack on the encoded image and then input it into the decoder for decoding.
[0029] (d) After obtaining the result of the discriminator network, the generator network performs backpropagation and updates the parameters of the encoder and discriminator using gradient updates to control the quality of the images generated by the generator.
[0030] Further, the optimal model parameters are obtained through training using the natural image dataset. The training process includes:
[0031] The first step is to first expand the secret information with length L into a three-dimensional vector and then change the shape of the binary-bit secret information to be the same as the image size.
[0032] The second step is to splice the secret information and the image together and input them into the encoder. The encoder outputs a three-channel color image through feature extraction and feature fusion, and this color image is the stego image.
[0033] The third step is to set up a simulated attack layer network before decoding to simulate the actual attack and obtain an attacked stego image through the simulated attack layer.
[0034] The fourth step is to input the attacked stego image into the decoder for decoding to obtain the binary secret information.
[0035] Further, the encoder network structure consists of a convolutional layer plus a residual block.
[0036] First, expand a string of binary secret information to the size of the input image, and splice the expanded secret information with the three-channel original image and input it into the encoder. Four convolutional operations are performed for feature extraction; then, pass through a residual block, namely a 3×3 convolutional layer, to extract feature information;
[0037] The feature maps obtained by each convolutional layer are fully connected to the corresponding upsampling layer, that is, do not directly perform supervision and loss calculation in the high-level feature maps. Combine the features in the low-level feature maps so that the finally obtained feature maps contain both high-level features and low-level features, and complete the fusion of features at different scales;
[0038] Finally, use a 1×1 convolutional layer to convert the multi-channel feature maps into three-channel color images.
[0039] Furthermore, when the decoder receives the encoded image I stego , extract the embedded secret information from the stego image; specifically including:
[0040] The decoder is first composed of 3 3*3 convolutional layers for feature extraction; then there is a residual block for feature extraction and reducing gradient disappearance; furthermore, four 3×3 convolutional layers are set to generate an L-channel secret information tensor; then use the average pooling layer and a single linear layer to convert the mapping into an L-length vector.
[0041] Furthermore, the network model uses different loss functions to alternately iteratively update the encoder, decoder, and discriminator; the message loss function is used to ensure the robustness of the model, and the image loss function and discriminator loss function are used to ensure the imperceptibility of the model; the loss function of the network is constructed as follows:
[0042] Use the image loss function L eA to keep the similarity error between the original image I cover and the encoded image I stego within a certain range; the formula of L eA is as follows:
[0043]
[0044] In the formula, ||·||2 is the Frobenius norm, and W, H, and C respectively represent the width, height, and number of channels of the input image;
[0045] Use the Learned Perceptual Image Patch Similarity (LPIPS) loss L lpips to minimize the distance between the original image I cover and the encoded image I stego ; the calculation of L lpips is as follows:
[0046]
[0047] where l is the network layer for extracting the feature stack, and the normalized C l -dimensional feature vector is represented by and, which contains the image I cover absolute value of the feature of layer l at spatial coordinates h, w, I cover and I stego represent the original image and the encoded image; for the l-th layer, where H l , W l and C l are the height, width, and number of channels in the l-th layer; w l represents the adaptive weight of each image feature, and ⊙ represents the exclusive-or operation;
[0048] The decoder is used to recover the secret information from the encoded image, and the decoded message is the same as the input secret message; the message loss function L dB is calculated as follows:
[0049]
[0050] where M in is the binary secret message, L represents the secret information length, M out is the extracted secret information, M in and M out ∈ {0, 1} L ;
[0051] The discriminator A is used to judge whether the received image is the encoded image I stego or the original image I cover ; the discriminator loss function L gA is used to improve the visual quality of I stego , and L gA is calculated as follows:
[0052] L gA = log(1 - A(I stego )) + log(A(I cover )) (4)
[0053] where log(*) represents the logarithmic function;
[0054] In terms of the training of the generator, the model sets different weights for the four loss functions L eA , L lpips , L dB and L gA to control the balance between the robustness and invisibility of the watermark;
[0055] The total loss function L is defined as follows:
[0056] L = λ eA LeA +λ lpips L lpips +λ dB L dB +λ gA L gA (5)
[0057] where λ eA , λ lpips , λ dB and λ gA are the weights of the loss functions L eA , L lpips , L dB and L gA respectively.
[0058] Advantageous Effects: Compared with the prior art, the technical solution of the present invention has the following advantageous technical effects:
[0059] The method of the present invention is based on the robust image steganography technology of the adversarial network, providing a theoretical basis and application support for the truly wide application of the robust image steganography technology. A robust steganography model based on the generative adversarial network is proposed to complete hiding the secret information into a three-channel color image, with high image invisibility and decoding accuracy.
[0060] The present invention improves on two models, HiDDeN[8] and StegaStamp
[12] , solving the image invisibility of these two models and the accuracy of extracting secret information. The present invention uses the adversarial network for training, enabling the model to better adjust the parameters of the generator through the steganalysis detector during the steganography process, making the generated stego-image as consistent as possible with the original image.
[0061] The model of the present invention is robust, and the accuracy of extracting secret information after attacking the stego-image is also relatively high. The present invention uses a composite loss function to train the model, accelerating the convergence speed of the loss function and improving the invisibility of the stego-image and the decoding accuracy. Description of the Drawings
[0062] Figure 1 is the framework of the RIW-GAN model of the present invention;
[0063] Figure 2 is the encoder structure of the present invention;
[0064] Figure 3 is the decoder structure of the present invention;
[0065] Figure 4 is the discriminator structure of the present invention;
[0066] Figure 5It is the visualization of the noise attack layer of the present invention;
[0067] Figure 6 It is the general loss evolution of the present invention after 300 epochs;
[0068] Figure 7 It is the difference between the original image and the encoded image of the present invention;
[0069] Figure 8 It is the bit accuracy of the RIW-GAN model proposed by the present invention under various distortions and intensities;
[0070] Figure 9 It is an example of the encoded and original images of different models of the present invention. Detailed implementation mode
[0071] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0072] The robust image watermarking method based on the generative adversarial network according to the present invention is based on the framework of the proposed robust image watermarking model (RIW-GAN) based on the generative adversarial network, as Figure 1 shown, the message M in is combined with the input image I cover to generate the encoded image I stego by the watermark encoder. I stego is distorted to output the noisy image I' stego . The decoder receives I' stego and outputs the secret information M out , and the discriminator is used to judge whether the image is an encoded image or an original image. The specific steps include:
[0073] (1) Construct an encoder-decoder framework. The encoder embeds the binary secret information into a three-channel carrier image, and the decoder extracts the secret information from the stego image.
[0074] The encoder network structure consists of a convolutional layer and a residual block; first, a string of binary secret information is expanded to the size of the input image L×H×W, where L represents the length of the secret information, and H and W represent the height and width of the image. The expanded secret information is concatenated with the three-channel original image and input into the encoder, and four convolutional operations are performed for feature extraction; then, it passes through a residual block, that is, a 3×3 convolutional layer, which can not only extract feature information but also more effectively reduce the vanishing gradient, making the loss function more likely to converge; in the latter part of the encoding network, the feature maps obtained from each previous convolutional layer are fully connected to the corresponding upsampling layer, so as to effectively use each layer of feature maps in subsequent calculations; this can avoid directly performing supervision and loss calculation in the high-level feature maps, but combines the features in the low-level feature maps, so that the finally obtained feature maps contain both high-level features and many low-level features, realizing the fusion of features at different scales and improving the accuracy of the model results. Finally, a 1×1 convolutional layer is used to convert the multi-channel feature maps into three-channel color images. The encoder structure is as Figure 2 shown, taking I cover and M in as inputs and I stego as the output.
[0075] The decoder extracts the embedded secret information from the stego-image after receiving the encoded image I stego ; specifically, the decoder first consists of 3 3*3 convolutional layers for feature extraction; then, a residual block is provided for feature extraction and reducing the vanishing gradient; furthermore, four 3×3 convolutional layers are provided to generate an L-channel secret information tensor; then, the average pooling layer and a single linear layer are used to convert the mapping into an L-length vector. The decoder structure is as Figure 3 shown, with the input being I stego and the output being the decoded secret information M out .
[0076] (2) Construct an adversarial network framework. The encoder network and the decoder network form an end-to-end model as the generator network. Through the generator network, the stego-image is generated and the secret information is extracted, and through the discriminator network, steganalysis is performed to judge the authenticity of the generated encoded image;
[0077] (3) Through the mutual adversarial learning between the generator network and the discriminator network, the error between the generated stego-image and the original image is within a certain range, and at the same time, the error between the decoded secret information and the original secret information is within a certain range;
[0078] The adversarial training of the generator and the discriminator includes:
[0079] (a) The generator network consists of two parts: an encoder and a decoder. The encoded image generated by the encoder network should have an error within a certain range from the original image;
[0080] (b) The discriminator network is used for steganalysis, that is, to distinguish whether an image is an encoded image or an original image, and to determine whether the image contains secret information. Input samples of natural carrier images into the discriminator to obtain the analysis category and determine whether it is consistent with the true category. Input the generated encoded image samples into the discriminator to obtain the analysis category and determine whether it is consistent with the true category;
[0081] (c) The simulated noise network: attacks the generated encoded image to obtain the attacked encoded image. That is, there is a simulated attack layer before inputting the encoded image into the decoding network to perform a simulated attack on the encoded image, and then input it into the decoder for decoding;
[0082] (d) After the generator network obtains the result of the discriminator network, it performs backpropagation and uses gradient update to update the parameters of the encoder and discriminator, controlling the quality of the images generated by the generator (generating higher-quality images).
[0083] (4) Use the obtained natural image dataset to train the network model and verify the performance and generalization training of the model: verify the invisibility of the model through the structural similarity index and peak signal-to-noise ratio, verify the robustness of the model through the accuracy of extracting secret information, and obtain the optimal model parameters. Use different datasets for training to make the model have stronger generalization. Obtain robust encoded images and extract secret information through the trained network model. The training process includes:
[0084] First step, first expand the secret information with length L into a three-dimensional vector, and then change the shape of the binary-bit secret information to make it the same as the image size;
[0085] Second step, splice the secret information and the image together and input them into the encoder. After feature extraction and feature fusion by the encoder, a three-channel color image is output, and this color image is the stego image;
[0086] Third step, there is a simulated attack layer network before decoding to simulate the actual attack. After passing through the simulated attack layer, an attacked stego image is obtained;
[0087] Fourth step, input the attacked stego image into the decoder for decoding to obtain the binary secret information.
[0088] The network model alternately iteratively updates the encoder, decoder, and discriminator using different loss functions; Update parameters and iterative training: Calculate the loss value using the composite loss function, calculate the gradient, and update the parameters. The message loss function is used to ensure the robustness of the model, and the image loss function and discriminator loss function are used to ensure the imperceptibility of the model; The loss function of the network is constructed as follows:
[0089] Use the image loss function L eA Keep the original image I cover and the encoded image I stego similar error within a certain range; L eA The formula of is as follows:
[0090]
[0091] In the formula, ||·||2 is the Frobenius norm, and W, H, and C represent the width, height, and number of channels of the input image respectively;
[0092] Use the Learned Perceptual Image Patch Similarity (LPIPS) loss L lpips Minimize the distance between the original image I cover and the encoded image I stego ; L lpips is calculated as follows:
[0093]
[0094] where l is the network layer used to extract the feature stack, and the normalized C l dimensional feature vectors are represented by and, which contain the absolute values of the l-layer features of the image I cover at the spatial coordinates h, w, I cover and I stego represent the original image and the encoded image; For the l-th layer, where H l , W l and C l are the height, width, and number of channels in the l-th layer; w l represents the adaptive weight of each image feature, and ⊙ represents the exclusive OR operation;
[0095] The decoder is used to recover the secret information from the encoded image, and the decoded message is the same as the input secret message; The message loss function L dB is calculated as follows:
[0096]
[0097] where M in is the binary secret message, L represents the length of the secret information, M out is the extracted secret information, M in and Mout ∈ {0, 1} L ;
[0098] The discriminator A is used to determine whether the received image is the encoded image I stego or the original image I cover ; The discriminator structure is as Figure 4 shown; The discriminator loss function L gA is used to improve the visual quality of I stego , and L gA is calculated as follows:
[0099] L gA = log(1 - A(I stego )) + log(A(I cover )) (4)
[0100] where log(*) represents the logarithmic function;
[0101] In terms of the training of the generator, the model sets different weights for the four loss functions L eA , L lpips , L dB and L gA to control the balance between the robustness and invisibility of the watermark;
[0102] The total loss function L is defined as follows:
[0103] L = λ eA L eA + λ lpips L lpips + λ dB L dB + λ gA L gA (5)
[0104] where λ eA , λ lpips , λ dB and λ gA are the weights of the loss functions L eA , L lpips , L dB and L gA respectively.
[0105] Figure 5 Fig. shows the visualization of the noise attack layer of the present invention. The first row: the original image I cover , the second row: the encoded image I stego , the third row: the distorted image I' stego , and the fourth row: the magnified difference. Figure 6Shows the general loss evolution of the present invention after 300 epochs, 6(a): overall loss, 6(b): encoder loss, 6(c): decoder loss, 6(d): discriminator loss. Figure 7 Shows the difference between the original image and the encoded image of the present invention. The first two columns: original image and encoded image, and the last three columns: difference images (enhanced 1×, 10×, 20×) between the encoded image and the original image.
[0106] Figure 8 Shows the bit accuracy of the proposed RIW - GAN model of the present invention under various distortions and intensities. Figure 9 Shows examples of the encoded and original images of the present invention in different models. Figure 9 (a) In the first row: cover image without embedded information, the second row: encoded image from the RIW - GAN model, the third row: encoded image from the HiDDeN model, the fourth row: encoded image from StegaStamp, the fifth row: encoded image from ReDMark; Figure 9 (b) In the sixth row of Figure 9 : normalized difference of the RIW - GAN model, the seventh row: normalized difference of the HiDDeN model, the eighth row: normalized difference of the StegaStamp model, the ninth row: normalized difference of the ReDMark model.
[0107] Tables 1 and 2 show the accuracy of extracting secret information of the algorithm proposed by the present invention for the datasets ImageNet and COCO under different attacks. Table 3 shows the decoding accuracy and the invisibility of the encoded images of the model proposed by the present invention on the COCO and ImageNet datasets compared with several state - of - the - art schemes and for three different lengths of secret messages. Table 4 shows the robustness comparison between the algorithm proposed by the present invention and the state - of - the - art models.
[0108] Table 1
[0109]
[0110] Table 2
[0111]
[0112] Table 3
[0113]
[0114] Table 4
[0115]
[0116] The present invention realizes the task of hiding secret information of binary bits into a three-channel color image and being able to extract the secret information under the framework of a generative adversarial network. Under the framework of the generative adversarial network, the watermark model simulates attacks during training, making the model robust, and a loss function is used to constrain the encoded image, so that the generated encoded image has high invisibility.
Claims
1. A robust image watermarking method based on a generative adversarial network, characterized in that: The method includes: (1) Construct an encoder-decoder framework. The encoder embeds binary secret information into a three-channel carrier image, and the decoder extracts the secret information from the stego-image. (2) Construct an adversarial network framework. The encoder network and the decoder network form an end-to-end model as the generator network. The stego-image is generated and the secret information is extracted through the generator network, and steganalysis is performed through the discriminator network to judge the authenticity of the generated encoded image. (3) Through the adversarial learning between the generator network and the discriminator network, the error between the generated stego-image and the original image is within a certain range, and at the same time, the error between the decoded secret information and the original secret information is within a certain range. (4) Use the obtained natural image dataset to train the network model. The invisibility of the model is verified by the structural similarity index and the peak signal-to-noise ratio, and the robustness of the model is verified by the accuracy of extracting the secret information to obtain the optimal model parameters. The robust encoded image and the extracted secret information are obtained through the trained network model. Among them, the optimal model parameters are obtained through training using the natural image dataset, and the training process includes: The first step is to first expand the secret information with length L into a three-dimensional vector, and then change the shape of the binary secret information to be the same as the image size. The second step is to splice the secret information and the image and input them into the encoder. After feature extraction and feature fusion by the encoder, a three-channel color image is output, and this color image is the stego-image. The third step is to set up a simulated attack layer network before decoding to simulate the actual attack, and a attacked stego-image is obtained after passing through the simulated attack layer. The fourth step is to input the attacked stego-image into the decoder for decoding to obtain the binary secret information.
2. The robust image watermarking method based on a generative adversarial network according to claim 1, wherein: The adversarial training of the generator and the discriminator includes: (a) The generator network includes two parts: an encoder and a decoder. The encoded image generated by the encoder network should have an error within a certain range from the original image. (b) The discriminator network is used for steganalysis, that is, to identify whether the image is an encoded image or an original image and to judge whether the image contains secret information. (c) Simulated noise network: A simulated attack layer is set before inputting the encoded image into the decoding network to perform a simulated attack on the encoded image, and then it is input into the decoder for decoding. (d) After the generator network obtains the result of the discriminator network, it performs backpropagation and updates the parameters of the encoder and the discriminator using gradient updates to control the quality of the images generated by the generator.
3. The robust image watermarking method based on a generative adversarial network according to claim 1, wherein: The encoder network structure consists of a convolutional layer plus a residual block. First, a string of binary secret information is expanded to the size of the input image, and the expanded secret information is spliced with the three-channel original image and input into the encoder. Four convolutional operations are performed for feature extraction; then it passes through a residual block, that is, a 3×3 convolutional layer, to extract feature information. The feature maps obtained by each convolutional layer are fully connected to the corresponding upsampling layer, that is, supervision and loss calculation are not directly performed in the high-level feature maps. By combining the features in the low-level feature maps, the final obtained feature maps contain both high-level features and low-level features, completing the fusion of features at different scales. Finally, a 1×1 convolutional layer is used to convert the multi-channel feature maps into three-channel color images.
4. The robust image watermarking method based on a generative adversarial network according to claim 1 or 3, characterized in that: The decoder extracts the embedded secret information from the encrypted image after receiving the encoded image I stego , which specifically includes: The decoder is first composed of three 3*3 convolutional layers for feature extraction; then there is a residual block for feature extraction and reducing gradient disappearance; furthermore, four 3×3 convolutional layers are provided to generate an L-channel secret information tensor; then the average pooling layer and a single linear layer are used to convert the mapping into an L-length vector.
5. The robust image watermarking method based on a generative adversarial network according to claim 1, characterized in that: The network model alternately iteratively updates the encoder, decoder, and discriminator using different loss functions; the message loss function is used to ensure the robustness of the model, and the image loss function and discriminator loss function are used to ensure the imperceptibility of the model; the loss function of the network is constructed as follows: Use the image loss function L eA Keep the original image I cover and the encoded image I stego with the similarity error within a certain range; L eA The formula for where ||·||2 is the Frobenius norm, and W, H, and C represent the width, height, and number of channels of the input image, respectively; Use the Learned Perceptual Image Patch Similarity (LPIPS) loss L lpips to minimize the distance between the original image I cover and the encoded image I stego ; L lpips is calculated as follows: where l is the network layer for extracting the feature stack, and the normalized C l -dimensional feature vector is represented by and, which contains the image I cover absolute value of the feature at layer l at spatial coordinates h, w, I cover and I stego represent the original image and the encoded image; for the l-th layer, where H l , W l and C l are the height, width, and number of channels in the l-th layer; w l represents the adaptive weight of each image feature, and ⊙ represents the exclusive OR operation; The decoder is used to recover the secret information from the encoded image, and the decoded message is the same as the input secret message; the message loss function L dB is calculated as follows: where M in is a binary secret message, L represents the length of the secret information, M out is the extracted secret information, M in and M out ∈{0,1} L ; The discriminator A is used to determine whether the received image is the encoded image I stego or the original image I cover ; The discriminator loss function L gA is used to improve the visual quality of I stego , and L gA is calculated as follows: L gA = log(1 - A(I stego )) + log(A(I cover )) (4) where log(*) represents the logarithmic function; In terms of the training of the generator, the model sets different weights for the four loss functions L eA , L lpips , L dB and L gA to control the balance between the robustness and invisibility of the watermark; The total loss function L is defined as follows: L = λ eA L eA + λ lpips L lpips + λ dB L dB + λ gA L gA (5) where λ eA , λ lpips , λ dB and λ gA are the weights of the loss functions L eA , L lpips , L dB and L gA respectively.
Citation Information
Patent Citations
End-to-end JPEG domain image steganography method based on generative adversarial network
CN112634117A
Image steganography method and system based on generative steganography confrontation
CN113538202A