Enhanced image steganography based on multi-scale cavity fusion iterative optimization
Through the multi-scale hollow fusion iterative optimization method, the high frequency and texture areas of the image are enhanced, and combined with the encoder and decoder to optimize image steganography, the existing models are solved, and the existing models are not realistic in visually and low steganography capacity are achieved, achieving efficient and safe information hiding.
Patent Information
- Application Number
- CN202510360601.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing carrier images generated by the deep learning-based image steganography model are not visually real and natural enough, and are easily detected, with low steganography capacity, which limits the efficiency and security of information hiding.
The multi-scale hollow fusion iterative optimization method is adopted to enhance the high frequency, edge and texture areas of the image through the multi-scale hollow fusion attention mechanism and image enhancement module, and combine the encoder and decoder for information embedding and extraction, and use a generative adversarial network to evaluate the image nature and optimize the steganography process.
The generated steganographic images are visually almost indistinguishable from the original image, with higher steganographic capacity, which improves the efficiency and security of information hiding, and are suitable for fast response scenes, retaining the natural beauty of the image.
Smart Images

Figure CN120298191A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an enhanced image steganography based on multi-scale cavity fusion iterative optimization, belonging to the field of image steganography. Background Art
[0002] Image steganography is a technology that hides secret information in digital media, aiming to store and transmit data without being detected. This technology has applications in multiple fields, including but not limited to digital rights management, privacy protection, secure communication, etc. In the field of image steganography, it can be further divided into two categories according to the way of hiding information: traditional steganography methods and deep learning-based methods. In real life, deep learning-based image steganography does not require designers to have much professional knowledge in related fields, and the finally trained steganography model also has good performance in terms of embedding capacity, robustness, and imperceptibility. Therefore, it is more valuable to adopt deep learning-based methods.
[0003] Deep learning-based image steganography is a technology that embeds secret information into an image so that these information are visually imperceptible. First, necessary preprocessing is performed on the original carrier image, such as denoising, enhancing contrast, etc., to improve the efficiency and security of steganography. A deep learning model (such as a convolutional neural network CNN) is used to extract features from the image. These features may include information such as the texture, edges, and colors of the image, and this information will be used in the subsequent steganography process. According to the extracted image features and steganography algorithm, the secret information is embedded into specific positions of the image (such as where the texture is complex or at the image edge). The embedding process may involve modifying the pixel values, DCT coefficients, etc., of the image. Post-processing is performed on the image after embedding the secret information to further reduce the visibility of the secret information and maintain the visual quality of the image. To improve the security of steganography, it may also be necessary to train a steganalysis model to simulate the detection process of an attacker and optimize the steganography algorithm through adversarial training so that it can resist the detection of the steganalysis model.
[0004] Some existing carrier images generated by deep learning-based steganography models may not be visually realistic and natural enough, are easily detected, and have a low steganography capacity, which limits the amount of information that can be hidden. To improve the efficiency, reliability, and hiding ability of image steganography, it is necessary to consider hiding information in some high-frequency regions and places with complex textures. Summary of the Invention
[0005] The purpose of the present invention is to disclose an enhanced image steganography based on multi-scale cavity fusion iterative optimization, which uses a multi-scale fusion attention mechanism and an image enhancement module to enhance the high-frequency, edge, and texture regions of the image, so as to facilitate hiding information therein and generate an image with higher visual self-steganography capacity.
[0006] In an embodiment according to the present disclosure, the enhanced image steganography based on multi-scale hole fusion includes the following steps:
[0007] Step 1: After obtaining the corresponding enhanced image E from the cover image C, perform feature extraction on each of them. After feature extraction, introduce the attention mechanism of multi-scale hole fusion to accurately focus on the key region features in the image, and then fuse the features of the two images to obtain image X. Thus, when embedding the secret information, these regions can be utilized more effectively. Then connect with the secret information tensor to form the stego tensor M.
[0008] Step 2: In each iteration, the encoder receives three inputs: the features of the current image M after fusing the secret information, the current perturbation δt-1 (i.e., the minor modification to the image), and the gradient of this perturbation with respect to the loss function Perform splicing to form the input Xt of the GRU cell and the previous hidden state ht-1 to update the gradient gt.
[0009] Step 3: By repeatedly applying the encoder, the final stego image is generated.
[0010] Step 4: The decoder receives the stego image generated by the encoder, and after a series of convolutions, recovers the original hidden information from the stego image.
[0011] Step 5: The critic network (similar to the discriminator in the generative adversarial network) is used to evaluate the naturality of the generated stego image and provide feedback.
[0012] Step 6: Repeat Steps 2 to 5 until a predetermined number of iterations is reached or the recovery error is reduced below an acceptable threshold. The finally generated image is the stego image containing the hidden information, which should be visually similar to the original image but contains the secret message.
[0013] For a further technical solution, Step 1 specifically includes the following steps:
[0014] Step 1.1: Use the Sobel operator for image edge enhancement to enhance the edge information of the cover image C, highlighting the edges of the cover image, thereby achieving the visual effect of enhanced image. The overall formula for image enhancement is as follows:
[0015]
[0016] Among them, Efinal represents the final edge-enhanced image or pixel value, which is the output result after all processing steps. α: Enhancement factor, which is a scalar value used to control the degree of edge enhancement. A larger α value will result in a more obvious edge enhancement effect. I is the original image or the pixel value in the image, which can be a grayscale image for edge detection. Kx and Ky are the convolution kernels of the Sobel operator in the x and y directions respectively, used to detect the horizontal and vertical edges of the image. Then, a tensor a is obtained by directly adding the feature fusions.
[0017] Step 1.2: The multi-scale dilated fusion attention module processes the output feature map in the spatial domain, aiming to enhance feature representation by utilizing multiple dilation rates and integrating channel and spatial attention mechanisms.
[0018] By concatenating the output features of the multi-scale dilated convolution in the channel dimension and adjusting the weights by combining channel and spatial attention mechanisms, MDFA can fuse features from multiple perspectives: merging different dilated convolution layers and global features can ensure that the network not only focuses on local details but also takes into account global information. This design makes the network output more comprehensive and enhances the generalization ability of the model in various scenarios. Feature dimensionality reduction is performed through the final 1x1 convolution, which can not only reduce the subsequent computational amount but also integrate information from different branches to generate a more effective feature representation for the final task. Its formula is expressed as:
[0019]
[0020]
[0021] A further technical solution, Step 2 specifically includes the following steps:
[0022] Step 2.1: The core of the encoder has a cyclic GRU-based update operator, and its update algorithm formula is as follows:
[0024]
[0025] Where η represents the learning rate or step size, g(·): Gradient-based update function, X: Current image or state, M: Target or constraint to be optimized.
[0026] Step 2.2: The GRU unit in the encoder updates its hidden state according to the input xt and the previous hidden state ht-1. GRU is a special recurrent neural network unit that can remember past state information and use this information to decide how to adjust the image next.
[0027] Step 2.3: The hidden state of the GRU is further processed through an additional convolutional layer to generate an update \(g_t\) of the gradient type. This update tells us how to adjust the perturbation \(\delta\) in the next step.
[0028] Step 2.4: Finally, the perturbation is adjusted according to this update and the step size \(\eta\) to obtain a new perturbation \(\delta_t\). To ensure that the change in the image is not too large, the perturbation \(\delta_t\) may be limited by a truncation factor.
[0029] Further technical solution, Step 3 includes the following steps:
[0030] Step 3: By repeatedly applying the encoder, all the steps of Step 2 above are repeated. The finally generated stego-image is obtained by accumulating the updates \(g_t\) in all iterative steps, denoted as \(X\) ~ \(= X+\eta P_tg_t\). Here \(X\) ~ is the final stego-image, which contains the hidden information but is visually almost indistinguishable from the original image.
[0031] Further technical solution, Step 4 includes the following steps:
[0032] Step 4.1: Three convolutional blocks are followed by another convolutional layer, whose purpose is to convert the feature map of the network into an output with a specific dimension, which matches the size (\(H\times W\)) of the original image and the number of bits (\(B\)) of the hidden information. In this way, the decoder can recover the original hidden information from the stego-image.
[0033] Further technical solution, Step 5 specifically includes the following steps:
[0034] Step 5.1: Three convolutional blocks are followed by an adaptive average pooling layer. The role of this layer is to convert the feature map into a single scalar value, which represents the probability that the network predicts whether the image contains the hidden code. Simply put, the task of the critic network is to judge whether an image looks "normal" or contains hidden information. What it receives is the stego-image generated in Step 3.
[0035] Further technical solution, Step 6 specifically includes the following steps:
[0036] Accuracy loss (\(L_{acc}\)): This part of the loss calculates the difference between the information \(M\) recovered by the decoder and the intermediate prediction \(\widetilde{X}_t\), aiming to enable the decoder to accurately recover the hidden information from the stego-image. This is a binary cross-entropy loss, which is used to encourage the model to reduce recovery errors, that is, to ensure that the decoded message is as close as possible to the original message. For the message \(M\) and the decoded message \(M'\), the formula is as follows:
[0037] Lacc(M′, M) = <M, logM′> + <(1 - M), log(1 - M′)>
[0038] Quality loss (Lqua): This part of the loss measures the difference between the cover image X and the stego image X~t using the mean square error, aiming to make the stego image visually as similar as possible to the original image. The formula is as follows:
[0039] Lqua(X~, X) = N1‖X~ - X‖22
[0040] Critic loss (Lcrit): This part of the loss is provided by the critic network and is used to judge whether the stego image looks like a natural image. The weights of the critic network are controlled by μ, and μ is a value greater than 0. This loss function is usually based on the discriminator loss in the generative adversarial network (GAN), aiming to make the stego image look as natural as possible to the critic network. The specific form of the critic loss may depend on the structure and training strategy of the critic network.
[0041] The overall loss formula is:
[0042] Ltrain = ∑t = 1T(γT - tLacc(M, X~t) + λLqua(X, X~t) + μLcrit(X, X~t))
[0043] An enhanced image steganography method based on multi-scale hole fusion iterative optimization provided by the present invention has the following beneficial effects:
[0044] 1) The present invention is an end-to-end optimized network model algorithm for image steganography. Using the deep learning model, it can directly find the places in the image where information can be hidden. Compared with the traditional LSB algorithm, the deep learning-based method has a wider applicability and significantly reduces the error rate in the information recovery process, and can even achieve zero-error recovery at high bit rates. It does not require the user to have strong professional capabilities and does not rely on manually designed feature extraction methods, greatly improving the invisibility, hiding capacity, and hiding efficiency of image steganography.
[0045] 2) The present invention is a comprehensive learning and iterative optimization method. It does not require a large network structure and is trained using a neural network framework. Its inference speed is fast enough to meet the requirements of real-time processing, enabling it to maintain high performance in scenarios that require quick responses. The generated stego image is almost indistinguishable from the original image visually. It cleverly encodes information in the insignificant regions of the image, maximizing the retention of the natural beauty of the image. It can retain the original feeling of the image while ensuring performance.
[0046] 3) The present invention draws on the characteristics of the human visual system and applies a multi-scale fusion attention mechanism to the convolutional neural network. Compared with ordinary attention, traditional convolutional layers usually have a fixed receptive field, which limits their ability to capture features at different scales. By introducing multi-scale dilated convolutions (using different dilation rates), features can be extracted at different spatial scales. Increasing the receptive field: Traditional convolutional networks increase the receptive field by stacking multiple convolutional layers, but this method leads to a significant increase in computational complexity and may cause information dilution. Using dilated convolutions, especially different dilation rates, can significantly expand the receptive field without losing resolution, which is crucial for capturing more extensive context information in images. Capturing multi-scale information: In image processing tasks, objects can appear in different sizes and shapes. Through multiple parallel dilation rates, the network can capture features at different scales simultaneously, which helps improve the model's adaptability and recognition ability to various scale features in the image, enhancing the image quality and visual perception.
[0047] 4) The present invention comprehensively considers that hidden information is suitable in the high-frequency region, especially at positions such as the edges and detailed contours of the image. The Sobel operator is used for edge enhancement. The Sobel operator strengthens the edges by calculating the first-order derivative of the image, making them more prominent in the image, which is crucial for subsequent image analysis and feature extraction. While enhancing the edges, the Sobel operator also enhances the details of the image, making the texture and shape features of the image more obvious. Using the Sobel operator can improve the visual effect of the image, especially when the image details are not obvious or the contrast is low. The use of the Sobel operator can enhance the edge information of the image, providing more selectivity and flexibility for information hiding, and helping to generate more concealed and robust stego images. Brief Description of the Drawings
[0048] Figure 1 is a flowchart of the no-reference color image quality evaluation method based on a dual convolutional neural network according to the present invention;
[0049] Figure 2 is a schematic diagram of the network structure of the no-reference color image quality evaluation method based on a dual convolutional neural network according to the present invention. Detailed Embodiments
[0050] The present invention proposes an enhanced image steganography based on multi-scale dilated fusion iterative optimization, which uses the idea of encoding-decoding confrontation and combines multi-scale fusion attention and image enhancement inputs to achieve image steganography. To more clearly elaborate the purpose, technical solution, and advantages of the present invention, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that the specific embodiments described below are only used to explain the principle of the present invention and are not used to limit the implementation scope of the present invention.
[0051] The present invention first provides an enhanced image steganography based on multi-scale cavity fusion iterative optimization. Specifically, refer to Figure 1 , which includes the following steps:
[0052] Step 1: After obtaining the corresponding enhanced image E from the cover image C, perform feature extraction on each of them. After feature extraction, introduce the attention mechanism of multi-scale cavity fusion to accurately focus on the key region features in the image, and then fuse the features of the two images to obtain image X. Thus, when embedding the secret information, these regions can be utilized more effectively.
[0053] Step 2: In each iteration, the encoder receives three inputs: the features of the current fused image X, the current perturbation δt-1 (i.e., the tiny modification to the image), and the gradient of this perturbation with respect to the loss function Perform splicing to form the input Xt of the GRU cell, and then update the gradient gt with the previous hidden state ht-1.
[0054] Step 3: By repeatedly applying the encoder, the final stego image is generated.
[0055] Step 4: The decoder receives the stego image generated by the encoder, and after a series of convolutions, recovers the original hidden information from the stego image.
[0056] Step 5: The critic network (similar to the discriminator in the generative adversarial network) is used to evaluate the naturality of the generated stego image and provide feedback.
[0057] Step 6: Repeat steps 2 to 5 until the predetermined number of iterations is reached or the recovery error is reduced below an acceptable threshold. The finally generated image is the stego image containing the hidden information, which should be visually similar to the original image but contain the secret message.
[0058] Among them, step 1 specifically includes the following steps:
[0059] Step 1.1: Use the Sobel operator for image edge enhancement to enhance the edge information of the cover image C, highlighting the edges of the cover image, thereby achieving the visual effect of image enhancement. The overall formula for image enhancement is as follows:
[0060]
[0061] Among them, Efinal represents the final edge-enhanced image or pixel value, which is the output result after all processing steps. α: Enhancement factor, which is a scalar value used to control the degree of edge enhancement. A larger α value will result in a more obvious edge enhancement effect. I is the original image or the value of a single pixel in the image, which can be a grayscale image for edge detection. Kx and Ky are the convolution kernels of the Sobel operator in the x and y directions respectively, used to detect the horizontal and vertical edges of the image. After that, a tensor a is obtained by directly adding the feature fusions.
[0062] Step 1.2: The multi-scale dilated fusion attention module processes the output feature map in the spatial domain, aiming to enhance feature expression by utilizing multiple dilation rates and integrating channel and spatial attention mechanisms.
[0063] By concatenating the output features of the multi-scale dilated convolution in the channel dimension and adjusting the weights by combining channel and spatial attention mechanisms, MDFA can fuse features from multiple perspectives: combining different dilated convolution layers and global features can ensure that the network not only focuses on local details but also takes into account global information. This design makes the output of the network more comprehensive and enhances the generalization ability of the model in various scenarios. Feature dimensionality reduction is performed through the final 1x1 convolution, which can not only reduce the subsequent computational amount but also integrate information from different branches to generate a more effective feature representation for the final task. Its formula is expressed as:
[0064]
[0065] Among them, Step 2 specifically includes the following steps:
[0066] The core of the encoder has a cyclic GRU-based update operator, and its update algorithm formula is as follows:
[0068]
[0069] Among them, η represents the learning rate or step size, g(·): Gradient-based update function, X: Current image or state, M: Target or constraint to be optimized..
[0070] Step 2.2: The GRU unit in the encoder updates its hidden state according to the input xt and the previous hidden state ht-1. GRU is a special recurrent neural network unit that can remember past state information and use this information to decide how to adjust the image in the next step
[0071] Step 2.3: The hidden state of GRU is further processed through an additional convolutional layer to generate an update gt of the gradient type. This update tells us how to adjust the perturbation δ in the next step.
[0072] Step 2.4: Finally, the perturbation is adjusted according to this update and the step size n to obtain a new perturbation δt. To ensure that the change in the image is not too large, the perturbation δt may be limited by a truncation factor.
[0073] Among them, step 3 specifically includes the following steps:
[0074] Step 3: By repeatedly applying the encoder, all the steps of step 2 above are repeated. The finally generated stego-image is obtained by accumulating the updated gt in all iterative steps, denoted as X ~ = X + ηPtgt. Here X ~ is the final stego-image, which contains the hidden information but is visually almost indistinguishable from the original image.
[0075] Among them, step 4 specifically includes the following steps:
[0076] Step 4.1: Three convolutional blocks are followed by another convolutional layer, whose purpose is to convert the feature map of the network into an output with a specific dimension, which matches the size (H×W) of the original image and the number of bits (B) of the hidden information. In this way, the decoder can recover the original hidden information from the stego-image.
[0077] Among them, step 5 specifically includes the following steps:
[0078] Step 5.1: Three convolutional blocks are followed by an adaptive average pooling layer. The role of this layer is to convert the feature map into a single scalar value, which represents the probability that the network predicts whether the image contains the hidden code. Simply put, the task of the critic network is to judge whether an image looks "normal" or contains hidden information. What it receives is the stego-image generated in step 3.
[0079] Among them, step 6 specifically includes the following steps:
[0080] The loss function consists of three parts:
[0081] Accuracy loss (Lacc): This part of the loss calculates the difference between the information M recovered by the decoder and the intermediate prediction X̃t, with the aim of enabling the decoder to accurately recover the hidden information from the stego-image. This is a binary cross-entropy loss, which is used to encourage the model to reduce recovery errors, that is, to ensure that the decoded message is as close as possible to the original message. For the message M and the decoded message M′, the formula is as follows:
[0082] Lacc(M′, M) = <M, logM′> + <(1 - M), log(1 - M′)>
[0083] Quality loss (Lqua): This part of the loss measures the difference between the cover image X and the stego-image X~t using the mean squared error, aiming to make the stego-image visually as similar as possible to the original image. The formula is as follows:
[0084] Lqua(X~, X) = N1‖X~ - X‖22
[0085] Critic loss (Lcrit): This part of the loss is provided by the critic network and is used to judge whether the stego-image looks like a natural image. The weights of the critic network are controlled by μ, and μ is a value greater than 0. This loss function is usually based on the discriminator loss in the generative adversarial network (GAN), aiming to make the stego-image look as natural as possible to the critic network. The specific form of the critic loss may depend on the structure and training strategy of the critic network.
[0086] The overall loss formula is:
[0087] Ltrain = ∑t = 1T(γT - tLacc(M, X~t) + λLqua(X, X~t) + μLcrit(X, X~t))
[0088] In the above description, methods of combining various technical features are mentioned, but not all possible combinations are exhaustively listed. However, it should be emphasized that as long as the combination of these technical features is reasonable in implementation and there are no contradictions among them, they should be regarded as within the scope covered by this specification.
Claims
1. An enhanced image steganography based on multi-scale cavity fusion iterative optimization, characterized in that The method includes the following steps: Step 1: After obtaining the corresponding enhanced image E from the cover image C, feature extraction is performed on each of them. After feature extraction, a multi-scale dilated fusion attention mechanism is introduced to enhance and accurately focus on the key region features in the image. Then, the features of the two images are fused to obtain image X. Thus, when embedding the secret information, these regions can be utilized more effectively. Then, it is connected to the secret information tensor to form the stego tensor M; Step 2: In each iteration, the encoder receives three inputs: the features of the current image M after fusing the secret information, the current perturbation δt-1 (i.e., the minor modification to the image), and the gradient of this perturbation with respect to the loss function concatenate them to form the input Xt of the GRU cell and the previous hidden state ht-1 to update the gradient gt; Step 3: Through repeated application of the encoder, the final stego image is generated; Step 4: The decoder receives the stego image generated by the encoder, and through a series of convolutions, the original hidden information is recovered from the stego image; Step 5: A critic network (similar to the discriminator in a generative adversarial network) is used to evaluate the naturality of the generated stego image and provide feedback Step 6: Repeat steps 2 to 5 until a predetermined number of iterations is reached or the recovery error is reduced below an acceptable threshold. The finally generated image is the stego image containing the hidden information, which should be visually similar to the original image but contain the secret message.
2. The enhanced image steganography obtained by multi-scale cavity fusion iterative optimization according to claim 1, characterized in that In step 1 described above, the Sobel operator for image edge enhancement is used to enhance the edge information of the cover image C, highlighting the edges of the cover image, thereby achieving an enhanced visual effect of the image. The overall formula for image enhancement is as follows:: where Efinal represents the final edge-enhanced image or pixel value, representing the output result after all processing steps. α: Enhancement factor, which is a scalar value used to control the degree of edge enhancement. A larger α value will result in a more obvious edge enhancement effect. I is the original image or a single pixel value in the image, which can be a grayscale image for edge detection. Kx and Ky are the convolution kernels of the Sobel operator in the x and y directions respectively, used to detect the horizontal and vertical edges of the image. Then, a tensor a is obtained by directly adding the feature fusions..
3. The enhanced image steganography obtained by iterative optimization based on multi-scale hole fusion according to claim 1, characterized in that, Step 2 described above specifically includes the following steps: Step 2.1: The core of the encoder has a cyclic GRU-based update operator, and its update algorithm formula is as follows: where η represents the learning rate or step size, g(·): Gradient-based update function, X: Current image or state, M: Target or constraint to be optimized; Step 2.2: The GRU unit in the encoder updates its hidden state according to the input xt and the previous hidden state ht-1. GRU is a special type of recurrent neural network unit that can remember past state information and use this information to decide how to adjust the image in the next step; Step 2.3: The hidden state of GRU is further processed through an additional convolutional layer to generate an update gt of the gradient type. This update tells us how to adjust the perturbation δ in the next step.
4. The enhanced image steganography obtained by iterative optimization based on multi-scale cavity fusion according to claim 1, characterized in that, In step 3 described above, by repeatedly applying the encoder, all steps of step 2 are repeated. The finally generated stego image is obtained by accumulating the updates gt in all iterative steps.
5. The enhanced image steganography obtained by iterative optimization based on multi-scale cavity fusion according to claim 1, wherein Step 4 described above specifically includes the following steps: Step 4.1: Three convolutional blocks are followed by another convolutional layer, which aims to transform the feature maps of the network into an output with a specific dimension that matches the size (H×W) of the original image and the number of bits of hidden information (B). In this way, the decoder can recover the original hidden information from the stego image.
6. The enhanced image steganography based on multi-scale hole fusion iterative optimization according to claim 1, characterized in that, The specific steps of step 5 are as follows: Step 5.1: Three convolutional blocks are followed by an adaptive average pooling layer. The role of this layer is to transform the feature maps into a single scalar value, which represents the probability that the network predicts whether the image contains a hidden code. Simply put, the task of the critic network is to judge whether an image looks "normal" or contains hidden information. What it receives is the stego image generated in step 3.
7. The enhanced image steganography obtained by iterative optimization based on multi-scale cavity fusion according to claim 1, characterized in that The specific steps of step 6 are as follows: Accuracy loss (Lacc): This part of the loss calculates the difference between the information M recovered by the decoder and the intermediate prediction X̃t, aiming to enable the decoder to accurately recover the hidden information from the stego image. This is a binary cross-entropy loss, which is used to encourage the model to reduce recovery errors, that is, to ensure that the decoded message is as close as possible to the original message. For the message M and the decoded message M′, the formula is as follows: Lacc(M′, M) = <M, logM′> + <(1 - M), log(1 - M′)> Quality loss (Lqua): This part of the loss uses the mean squared error to measure the difference between the cover image X and the stego image X̃t, aiming to make the stego image as visually similar to the original image as possible. The formula is as follows: Lqua(X̃, X) = N1‖X̃ - X‖22 Critic loss (Lcrit): This part of the loss is provided by the critic network and is used to judge whether the stego image looks like a natural image. The weights of the critic network are controlled by μ, and μ is a value greater than 0. This loss function is usually based on the discriminator loss in the generative adversarial network (GAN), aiming to make the stego image look as natural as possible to the critic network. The specific form of the critic loss may depend on the structure and training strategy of the critic network. The overall loss formula is: Ltrain = ∑t = 1T(γT - tLacc(M, X̃t) + λLqua(X, X̃t) + μLcrit(X, X̃t))
Citation Information
Cited By
Image privacy protection method, device and equipment based on dynamic disturbance optimization
CN120512500A
Image privacy protection method, device and equipment based on dynamic disturbance optimization
CN120512500B