Carrier image content consistency enhancement method for high-capacity steganography
By performing multi-channel processing and noise strategy optimization on color carrier images, large-capacity enhanced images with consistent content are generated, and the existing steganography algorithms balance the carrier image embedding capacity and security is solved, and efficient embedding and secure communication of carrier images are realized.
Patent Information
- Application Number
- CN202510761135.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing steganography algorithms are difficult to achieve an ideal balance between the embedding capacity and security of carrier images. Images with simple textures contain limited secret information and are prone to visual distortion, while images with complex textures are difficult to accurately embed and be detected.
By splitting the color carrier image into a multi-channel format, using the generated network to generate an embedding probability map and combining noise mapping for message embedding, iteratively update the embedding probability, and using a dual-class stream noise strategy and diffusion model for image optimization to generate large-capacity enhanced images with content consistency.
While maintaining the consistency of image content, the embedding capacity of the carrier image is significantly improved, the concealment and security of steganography is improved, and the richness and creativity of the image are enhanced.
Smart Images

Figure CN120259133A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image steganography technology, and particularly to a method for enhancing the content consistency of a carrier image for large-capacity steganography. Background Art
[0002] At present, with the deep penetration of the Internet, the security of private image information has encountered severe challenges and has become an important issue that needs to be urgently solved. In this context, it has become an urgent task in the field of information security to protect private image information by means of accurately adapted steganography technology. As a powerful weapon for information security protection, the core advantage of steganography lies in subtly hiding the traces of "covert communication" and building a secure communication link in an imperceptible way.
[0003] Currently, most steganography algorithms can indeed embed secret information into a selected carrier image without the user's notice. However, these algorithms often struggle to achieve an ideal balance between the embedding capacity and security. The texture complexity of carrier images varies greatly, which determines that different images have different embedding potentials. Specifically, when the carrier image is an ordinary image, no matter how exquisitely the objective distortion function is designed, existing methods are difficult to find a satisfactory balance point between the embedding ability and security. For images with simple textures, the amount of secret information that can be accommodated is limited. If the embedding capacity is forcibly increased, it is extremely easy to cause visual distortion, thus exposing the steganography traces and reducing the security; while for images with complex textures, although there seems to be a large embedding space, during the embedding process, there is also a problem of how to accurately grasp the embedding depth to avoid damaging the image structure due to excessive embedding and being detected by the detection algorithm. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a method for enhancing the content consistency of a carrier image for large-capacity steganography, which can enhance the embedding capacity of the original carrier image and generate a large-capacity enhanced carrier image with content consistency, so as to solve the technical problem that the carrier images selected by existing image steganography methods are usually far from the optimal options for embedding secret information.
[0005] The specific solution of the present invention is as follows: A method for enhancing the content consistency of a carrier image for large-capacity steganography, comprising the following steps: S1, splitting the color carrier image dataset into R, G, and B multi-channel formats and feeding them into the generation network G N to obtain a multi-channel embedding probability map M T p , and then combining the embedding probability map M T p with the noise map N T and sending them into the embedding simulator E for message embedding to obtain a modified map M T ; S2: Obtain the stego-image by modifying the mapping M T to minimize the difference between the stego-image and the carrier image using a steganalysis tool, and iteratively update the embedding probability of each pixel in the multi-channel; S3: Calculate the noise mixing weight matrix W T p for enhancing the image using the multi-channel embedding probability map M N ; S4: Apply a two-class flow noise strategy to the carrier image and perform adaptive mixing guided by the noise mixing weight matrix W N to obtain the noisy image C t0 ; S5: Feed C t0 into the diffusion model and apply three regularizers of image sharpness, noise distribution, and adversarial transformation to guide the denoising process to generate the final denoised image C0.
[0006] Furthermore, in S1, the color carrier image dataset is split into an R, G, B multi-channel format and fed into the generation network G N to obtain the multi-channel embedding probability map M T p , which specifically includes the following steps: S11: Sequentially take out the color carrier images in the dataset in batches of batch_size, split the channels of the carrier images in each batch to obtain a number of grayscale images C T ; S12: Feed the number of grayscale images C T into the generation network G N and ensure that the generation network can output the corresponding number of multi-channel embedding probability maps M T p , T ∈ {R, G, B}.
[0007] By splitting the channels of the color carrier image and processing it in batches, the multi-channel embedding probability map can be efficiently generated. This processing method not only improves the accuracy of generating the embedding probability map but also ensures the stable output of the generation network under a fixed payload, providing a reliable basis for subsequent message embedding and image optimization.
[0008] Furthermore, in S1, the embedding probability map M T p and the noise map N T are fed into the embedding simulator E for message embedding to obtain the modified mapping M T, specifically including the following steps: S13. Send the three different-channel embedding probability maps corresponding to each color carrier image in the embedding probability map M T p and the three groups of random noise maps N T into the embedding simulator E for message embedding together; S14. The embedding simulator E simulates the embedding process by sampling using pixel-level modification, and compares the random noise map N T with the embedding probability map M T p to obtain a pixel-level modification map M T with three possible directions (+1, -1, 0).
[0009] Combining the embedding probability map with the noise map and simulating the embedding process through pixel-level modification sampling can obtain an accurate pixel-level modification map. This accurate modification map helps to better control the change of the image when embedding information, reduce the impact on the image quality, and at the same time improve the concealment and security of information embedding.
[0010] Furthermore, the embedding rule of the embedding simulator E is as follows: M T p (x,y) has an embedding probability of M T p (x,y) / 2 for the direction (+1) and M T p (x,y) / 2 for the direction (-1); If a random noise map element N T (x,y) is less than the embedding probability of the direction (+1) in M T p (x,y), then the modification map M T (x,y) = +1; If a random noise map element N T (x,y) is greater than M T p (x,y) for (1 - the embedding probability of the direction (-1)), then the modification map M T (x,y) = -1; Otherwise, set M T (x,y) = 0.
[0011] The embedding simulator adopts an embedding rule based on probability comparison, which can more flexibly simulate the information embedding process and ensure that the embedding operation conforms to the predetermined probability distribution. This rule clarifies the decision basis for the pixel modification direction, helps to improve the accuracy and controllability of information embedding, and at the same time reduces unnecessary image modifications, thus better maintaining the image quality.
[0012] Further, in S2, by modifying the mapping M T to obtain the stego image, specifically including: By modifying the mapping M T and the previously separated R, G, B multi-channel grayscale images C T the stego image S can be obtained T : .
[0013] The method of generating a stego image by combining a modified mapping with multi-channel grayscale images can directly reflect the embedded information in the image. This way of generating a stego image is simple and efficient, ensuring the accuracy of the embedded information and the recoverability of the image, providing a good foundation for subsequent steganalysis and image optimization.
[0014] Further, in S2, a steganalyzer is used to minimize the difference between the stego image and the carrier image, specifically including the following steps: S21, the steganalyzer detects the embedded information in the input stego image to obtain the embedding probability; S22, according to the embedding probability, select the sampling image in the next round of iteration from the stego image.
[0015] By detecting the embedded information and selecting the sampling image, the steganalyzer can effectively evaluate the quality of the embedded image and the information embedding effect. This mechanism can select a suitable image for the next round of iterative optimization according to the embedding probability, thereby gradually increasing the similarity between the stego image and the original carrier image, enhancing the concealment and security of steganography.
[0016] Further, in S3, using the multi-channel embedding probability map M T p calculate the noise mixing weight matrix W for enhancing the image N specifically including: S31, use the embedding probability generator model to generate a multi-channel embedding probability map by inputting a color image; S32, perform binarization processing on the multi-channel embedding probability map to obtain a new binary embedding probability image B T ; S33, according to the binary embedding probability image B T , for any position (x, y), if at least one value in B T (x, y) is 1, then in the weight matrix W N set the value at this position to 1, otherwise set it to 0.
[0017] By using an embedding probability generator model and binary processing to calculate the noise mixing weight matrix, it is possible to accurately determine the regions in the image suitable for noise mixing. This method based on binary processing simplifies the generation process of the weight matrix and can effectively highlight the noise mixing weights of important regions in the image, providing precise guidance for subsequent adaptive mixing.
[0018] Furthermore, in step S4, a dual-class flow noise strategy is adopted for the carrier image, which specifically includes: The dual-class flow noise mechanism includes a creative flow and a stable flow; For the creative flow, random noise corresponding to the time step t is added to the carrier image C to generate a noisy image C O t ; Using a diffusion model to perform iterative denoising on C O t until the time step t0 is reached, at which point the denoised image C O t0 ; For the stable flow, DDIM inversion is used to add noise to the carrier image C and obtain a noisy image C S t0 .
[0019] By adopting the dual-class flow noise strategy, noisy images with different characteristics are generated through the creative flow and the stable flow respectively. This mechanism can enhance the detail richness and creativity of the image while ensuring the content fidelity of the image. Adding random noise in the creative flow can enrich the image details and increase the embedding capacity of the image; while the stable flow adds noise through DDIM inversion and ensures content fidelity, providing a diverse basis for subsequent adaptive mixing and optimization of the image.
[0020] Furthermore, in step S4, adaptive mixing is performed under the guidance of the noise mixing weight matrix W N to obtain a noisy image C t0 , which specifically includes the following steps: When the dual-class flow noise generates two noisy images C O t0 and C S t0 ), the usage ratios of C O t0 and C S t0 are dynamically adjusted through the probability values in the embedding probability image, and the two noisy images are adaptively mixed using the weight matrix W N ; For image regions with a higher embedding probability, C S t0Maintain a relatively stable content structure. For image regions with a low embedding probability, use C O t0 Generate variants to enrich the detailed information in this region.
[0021] Adaptive mixing guided by the noise mixing weight matrix can dynamically adjust the usage ratio of the noise image. This adaptive mixing method can flexibly select the noise image according to the characteristics of the image region, ensuring the stability of regions with a high embedding probability while enriching the detailed information of regions with a low embedding probability, thereby enhancing the embedding capacity of this region. This delicate mixing process helps to significantly improve the overall embedding capacity of the carrier image while maintaining the content consistency of the carrier image.
[0022] Compared with the prior art, the beneficial effects of the present invention are: A general content-consistent carrier image enhancement framework is proposed. This framework can enhance the embedding capacity of any given image or image set while maintaining content consistency. By automatically learning the embedding cost of each pixel in the image, an embedding probability map that can be used to evaluate the embedding ability of each pixel is generated. Then, according to the embedding probability map, the mixing ratio of the two-class flow noise is dynamically adjusted to obtain a noisy image, and the noisy image is fed into a diffusion model constrained by three regularizers for denoising processing to output a denoised enhanced image. Finally, the enhanced image with an increased embedding capacity is used as the carrier image, which can be used to securely communicate in the network after embedding a secret message using any feasible steganography method. It has been proven that the proposed framework is effective for state-of-the-art steganography methods in enhancing the embedding capacity of the carrier image, so the proposed framework has good generality.
[0023] A multi-channel embedding probability joint learning mechanism is proposed. This mechanism optimizes performance by comprehensively learning all channel information of the color image. During the learning process, the embedding probability information between different channels is shared and coordinated with each other, ensuring in-depth learning and iterative update of the embedding probability information for different channels at the same detailed parts of the image, generating an embedding probability map with deeper spatial information and finer coarse-grainedness. This mechanism effectively solves the problem of difficult control of inter-channel coordination and color continuity faced when directly learning the embedding probability of a color image, thereby avoiding color distortion or the generation of color dots. At the same time, it also overcomes the problem of color incoordination and image quality degradation caused by the inability to capture the color information of the image when converting the color image into a grayscale image for embedding probability learning.
[0024] (3) A dual-class flow noise dynamic mixing strategy based on a multi-channel embedded probability map is designed. This strategy can calculate the noise mixing weight matrix for enhancing the image using the multi-channel embedded probability map, and then adaptively mix the noise guided by this weight matrix to generate a noisy image. In this process, for the image regions with a higher embedding probability, we use stable noise to maintain the stable content structure, and for the image regions with a lower embedding probability, we use creative noise to generate variants to enrich the detail information in this region. This strategy significantly enhances the overall embedding capacity of the image by enhancing the texture details in the regions not suitable for embedding. Description of the Drawings
[0025] To more clearly illustrate the technical solutions in the embodiments of the present drawings or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present drawings. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0026] Figure 1 It is the flowchart of the method of the present invention; Figure 2 It is the visualization schematic diagram of the original carrier image and the enhanced carrier image provided by the embodiment of the present invention; Figure 3 It is the comparison schematic diagram of the embedding probabilities of the original carrier image and the enhanced carrier image provided by the embodiment of the present invention; Figure 4 It is the comparison schematic diagram of the embedding probabilities of the original carrier image and the enhanced carrier image provided by the embodiment of the present invention; Figure 5 It is the comparison schematic diagram of the embedding probabilities of the original carrier image and the enhanced carrier image provided by the embodiment of the present invention. Detailed Embodiments
[0027] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following will describe and explain the present invention in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments provided by the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0028] As Figure 1 shown, the present invention provides a method for enhancing the content consistency of a carrier image for high-capacity steganography, which specifically includes the following steps: S1, Split the color carrier image dataset into R, G, B multi-channel formats and feed it into the generation network GN Obtain the multi-channel embedding probability map M T p Then, feed the embedding probability map M T p and the noise map N T into the embedding simulator E for message embedding to obtain the modified map M T .
[0029] Specifically, split the color carrier image dataset into R, G, B multi-channel formats and feed them into the generation network G N to obtain the multi-channel embedding probability map M T p , which specifically includes the following steps: S11, sequentially take out the color carrier images in the dataset in batches of batch_size, perform channel splitting on the carrier images in each batch to obtain a certain number of grayscale images C T ; S12, feed a certain number of grayscale images C T into the generation network G N and ensure that the generation network can output the corresponding a certain number of multi-channel embedding probability maps M T p , T ∈ {R, G, B}.
[0030] In the specific implementation process, since this framework is a general framework, any feasible generator network can be used to represent the generation network G N , such as HILL, UT-GAN, ASDL-GAN, and SPAR-RL, etc., can all be used for this operation.
[0031] Specifically, feed the embedding probability map M T p and the noise map N T into the embedding simulator E for message embedding to obtain the modified map M T , which specifically includes the following steps: S13, feed the three different-channel embedding probability maps corresponding to each color carrier image in the embedding probability map M T p and three groups of random noise maps N T into the embedding simulator E for message embedding together; In the embedding probability map M T p , each color carrier image corresponds to three different-channel embedding probability maps, which are M R p 、MG p and M B p , and send it together with three groups of random noise maps N T into the embedding simulator E for message embedding to obtain the modified map M T : .
[0032] S14. The embedding simulator E simulates the embedding process by sampling using pixel-level modification, and compares the random noise map N T with the embedding probability map M T p to obtain a pixel-level modification map M with three possible directions (+1, -1, 0) T .
[0033] Specifically, SPAR-RL can be selected as the generation network G N .
[0034] In addition, the embedding rule logic of the embedding simulator E is as follows: M T p The embedding probability of the direction (+1) in M(x,y) and the direction (-1) in M T p (x,y) is M(x,y) / 2; T p (x,y) / 2; If a certain random noise map element N T (x,y) is less than the embedding probability of the direction (+1) in M T p (x,y), then the modified map M T (x,y) = +1; If a certain random noise map element N T (x,y) is greater than M T p (x,y) in (1 - (the embedding probability of the direction (-1))), then the modified map M T (x,y) = -1; Otherwise, set M T (x,y) = 0.
[0035] S2. Obtain the stego image through the modified map M T Use a steganalysis tool to minimize the difference between the stego image and the carrier image, and iteratively update the embedding probability of each pixel in the multi-channel to obtain a more accurate multi-channel embedding probability map.
[0036] Specifically, through the modified map M TObtain the stego image, specifically including: By modifying the mapping M T With the previously separated R, G, B multi-channel grayscale images C T The stego image S can be obtained T : .
[0037] In the specific implementation process, modify the mapping M T The element type in is (+1, -1, 0), and matrix element addition processing can be performed according to the channel type and the corresponding grayscale image C T Taking a single channel as an example, if there is a grayscale image matrix [[119, 20, 98], [248, 6, 58]] and a modified mapping matrix [[-1, -1, 0], [1, -1, 0]], then the processed stego image matrix is [[118, 19, 98], [249, 5, 58]].
[0038] Specifically, use a steganalysis tool to minimize the difference between the stego image and the carrier image, specifically including the following steps: S21, the steganalysis tool detects the embedded information in the input stego image S T to obtain the embedding probability.
[0039] The main task of the steganalysis tool is to detect whether the input stego image S T contains embedded information. Tools such as Xu-Net, Yedroudj-Net, Ye-Net, and SRNet can all be used for this operation.
[0040] S22, select the sampling image in the next iteration according to the embedding probability from the stego image S T .
[0041] The stego image S that performs better in the steganalysis process T will have a greater probability of being sampled in the next iteration. Specifically, if the value of the embedding probability is high, the corresponding pixel is suitable for hiding the secret message, that is, its embedding ability is high, and vice versa; the purpose is to minimize the stego image S T and the original image C T and iteratively update the embedding probability of each pixel in the multi-channel to obtain a more accurate multi-channel embedding probability map.
[0042] S3, use the multi-channel embedding probability map M T p to calculate the noise mixing weight matrix W for enhancing the image N .
[0043] Specifically, it includes the following steps: S31. Use the embedding probability generator model to generate a multi-channel embedding probability map by inputting a color image.
[0044] S32. Binarize the multi-channel embedding probability map to obtain a new binary embedding probability image B. T .
[0045] Specifically, set a threshold , and use the threshold decision function to obtain the binary embedding probability image B. T : ; where (x, y) is the coordinate position of a certain probability value in the embedding probability map.
[0046] In the specific implementation process, taking a single channel as an example, if there is an embedding probability map [[0.11, 0.48, 0.36], [0.25, 0.25, 0.43]], and the threshold is 0.4, the new binary embedding probability image after being processed by the threshold decision function is [[0, 1, 0], [0, 0, 1]].
[0047] S33. According to the binary embedding probability image B T , for any position (x, y), if there is at least one value of 1 in B T (x, y), then in the weight matrix W N , set the value of this position to 1, otherwise set it to 0. This process is illustrated by the following formula: .
[0048] In the specific implementation process, if there is a binary embedding probability image B T as {[[0, 0, 1], [0, 0, 0]], [[1, 0, 1], [1, 0, 0]], [[0, 0, 1], [1, 0, 1]]}, then according to the calculation rule of the above weight matrix W N , it can be known that the weight matrix W N is [[1, 0, 1], [1, 0, 1]].
[0049] S4. Apply the two-class flow noise strategy to the carrier image and perform adaptive mixing guided by the noise mixing weight matrix W N to obtain a noisy image C t0 .
[0050] Specifically, applying the two-class flow noise strategy to the carrier image specifically includes: The dual-class flow noise mechanism includes a creative flow and a stable flow. The creative flow aims to enhance the richness and creativity of image details by incorporating higher-intensity noise elements, while the stable flow focuses on introducing relatively weak noise to ensure that the fidelity of the content is not significantly affected.
[0051] For the creative flow, random noise corresponding to the time step t is added to the carrier image C to generate a noisy image C O t , which can be illustrated by the following formula: ; where α t represents the weight coefficient used to regulate the contribution degree of the carrier image C; σ t denotes the noise intensity; while represents the random noise vector at the time step t.
[0052] The diffusion model is used to perform iterative denoising on C O t until the time step t0 is reached, and the denoised image C O t0 is obtained. To ensure that the structural elements of the image are consistent with the carrier image C, that is, to maintain content consistency, gradient-guided sampling is used to introduce the conditioned reflection of auxiliary information for the denoising process of C O t →C O t0 in this noise flow.
[0053] In the specific implementation process, the basic model of Stable Diffusion XL (SDXL-base) implemented in the HuggingFace Transformer and Diffuser libraries is used as the diffusion model for image enhancement.
[0054] For the stable flow, DDIM inversion is used to add noise to the carrier image C and obtain a noisy image C S t0 . It ensures that when using a deterministic sampling algorithm like DDIM, the content in the carrier image C can be reconstructed with high fidelity from C S t0 .
[0055] Specifically, adaptive mixing is performed with the noise mixing weight matrix W N as the guide to obtain the noisy image C t0 , which specifically includes the following steps: When the dual-class flow noise generates two noisy images C Ot0 and C S t0 After that, dynamically adjust C by the probability values embedded in the probability image O t0 and C S t0 The usage ratio of is adaptively mixed with two noise images by using the weight matrix W N ; Its functional expression is: .
[0056] For the image regions with a higher embedding probability, use C S t0 To maintain a relatively stable content structure, for the image regions with a lower embedding probability, use C O t0 Generate variants to enrich the detailed information in this region
[0057] In the specific implementation process, if there is a weight matrix W N is [[1, 0, 1], [1, 0, 1]], the noise image C O t0 [[0.878, 0.011, 0.596], [0.336, 0.121, 0.633]], the noise image C S t0 [[0.256, 0.894, 0.290], [0.632, 0.158, 0.534]], then dynamically adjust the stable noise C according to the above adaptive noise mixing function S t0 and the created noise C O t0 After the usage ratio, the noisy image C t0 is [[0.878, 0.894, 0.596], [0.336, 0.158, 0.633]].
[0058] S5, Feed C t0 into the diffusion model, and use three regularizers of image sharpness, noise distribution and adversarial denaturation for constraint to guide the denoising process and generate the final denoised image C0
[0059] Specifically, in the process of denoising the C t0 image, three target attributes are formulated as constraints from the perspective of the diffusion model based on scores, and they are used to guide the sampling process by normalizing the predicted noise, thereby adjusting the output. In this process, C t0 is assigned to C t .
[0060] Sharpness regularization. Sharpness refers to the perceived clarity related to the edge contrast of an image. Due to the characteristics of the human visual system, images with higher sharpness often appear clearer, although an increase in sharpness does not necessarily improve the actual resolution of the image.
[0061] Here, the denoising is regularized using the sharpening rate of denotes the intermediate reconstruction of C0 at time step t using the reparameterization technique, i.e., denotes the intermediate reconstructed version of image C0 when time step t is close to 0. Specifically, the Sobel kernel is used to estimate the magnitude of the spatially varying luminance derivative, denoted as ; To encourage higher sharpness and improve the overall generation quality, a binary indicator ∨(·) is used for optimization. When the input value falls within the 35th and 65th percentiles of , ∨(·) = 1. The objective of sharpness regularization is: ; where denotes the sharpness regularization loss; W represents the height and width of the image; The sum is normalized to be independent of the size of the image, and a negative sign is added to indicate that minimizing this loss can encourage higher sharpness; denotes the summation over all pixel positions (x, y) of the image.
[0062] Distribution regularization. Considering the inevitability of generalization error, the noise predicted by the diffusion model may not follow the Gaussian distribution N(0, I), especially when directly generating images from the synthetic noise images produced in the noise phase using the diffusion model. Therefore, the denoising process is regularized by penalizing the distribution gap: ; where is used to calculate the variance of the predicted noise; denotes the L2 norm, which is used to measure the difference between two vectors; denotes the distribution regularization loss.
[0063] Adversarial regularization. Motivated by the self-attention guidance of the diffusion model, adversarial regularization is added in the denoising phase to avoid generating blurry images. Specifically, is defined as the Gaussian blur function, and the objective is designed as: ; Among them, represents the adversarial regularization loss; Apply the Gaussian blur function to the intermediate reconstructed image This step will generate a blurred version of the image; Calculate the difference between the original intermediate reconstructed image and the blurred image. With the help of these three regularizations, an additional operation step is set after each denoising iteration to update the current state:
[0064] ; ; Among them, The image state after regularization optimization at time step t-1; represents the image state at time step t-1; C t represents the image state at time step t; ▽ is the gradient operator, used to represent gradient information, and (ξ, ψ, ζ) are trade-off parameters, used to control the relative importance of the regularization loss; When the time step t iterates to 0, the final denoised image C0 can be generated.
[0065] In the specific implementation process, the optimal values of the trade-off parameters are determined through experimental research, (ξ, ψ, ζ)=(4, 20, 0.4), and the visualization schematic diagram of the original carrier image C and the final enhanced carrier image C0 and the comparison schematic diagram of the embedding probability map are as Figures 2 - 5 shown.
[0066] Next, the performance of various aspects of the embodiments of the present invention will be analyzed: Peak signal-to-noise ratio (PSNR) analysis: PSNR is one of the commonly used indicators to measure image quality. It is used to compare the quality differences between the original image and the image after encoding, compression, or processing. The higher the PSNR value, the better the image quality. Its value can be roughly divided into four categories: (1) [0, 20) dB, the image quality is unacceptable; (2) [20, 30) dB, the image quality is poor and can be perceived, (3) [30, 40) dB; the quality is acceptable but the distortion may be perceived; (4) [40, ) dB, the image quality is excellent (i.e., very close to the original image), and the PSNR value V can be quantitatively calculated using the following formula PSNR : ; ; where G(u,v) , Z (u,v) respectively represent the pixel values of the original image and the stego-image at the coordinate position (u, v); V MES is the mean square error, which is used to quantify the pixel-level distortion degree between the original image and the stego-image.
[0067] Structural Similarity Index (SSIM) analysis: SSIM is an index used to measure the similarity degree between two images, and its value range is from -1 to 1, where 1 means the two images are exactly the same, and -1 means the two images are completely different. The SSIM value can be quantitatively calculated by the following formula : ; where, μ p and μ q respectively represent the average brightness of the original image G and the stego-image Z; σ p and σ q respectively represent the standard deviations of G and Z; σ pq represents the covariance of G and Z; C1 and C2 are constants to avoid the instability caused when the denominator is close to 0.
[0068] Experimental dataset: The Human Preference Dataset Version 2 (HPDv2) is a large human preference dataset covering images generated by a wide range of text prompts. It includes 433,760 human preference selections on 798,090 pairs of images and provides a set of evaluation prompts, which are evenly distributed according to the four major style categories of animation, concept art, painting, and photo. For each type of evaluation prompt, HPDv2 provides the corresponding benchmark images generated by various mainstream text-to-image generation models. On this basis, the present invention adopts a set of benchmark images generated by the SDXL-Base-0.9 model as the original image set for the image enhancement technology, which contains 10,000 color images.
[0069] Steganography scheme: Select SPAR-RL as the steganography scheme for this embodiment, and generate stego-images of the original images and enhanced images with a payload of 0.1 bpp to 0.4 bpp. In addition, the proposed multi-channel joint learning mechanism of the present invention is used for color image processing.
[0070] The analysis and calculation results of the peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) of different stego-images are shown in Table 1 below:
[0071] At Figure 2In it, a visual comparison was made between the original carrier image and the enhanced carrier image, and the detailed texture was magnified and displayed; the first row shows the unprocessed original image, and the second row is the enhanced image carefully designed. From the perspective of the overall content information of the image, the enhanced carrier image strictly maintains consistency with the original carrier image in content, without introducing any changes deviating from the original content theme. The theme objects and information presented in the two images are exactly the same, ensuring the faithful reproduction of the content. Secondly, in the detailed texture part, by observing the area marked by the square frame in the figure, it can be clearly found that in the enhanced image, this area not only reveals more delicate and complex texture information but also perfectly integrates the generated information highly consistent with the content, thus significantly improving the detail performance and information transmission ability of the image. Therefore, it is verified that this framework can perfectly generate richer, more delicate and highly content-integrated texture information in the detail part on the premise of ensuring the content consistency of the image before and after enhancement.
[0072] To further observe the improvement effect of the embedding ability of the carrier image, in Figures 3 - 5 the embedding probability maps of the original carrier image and the enhanced carrier image are shown. The first row is the original image and the enhanced image, and the second row is the corresponding embedding probability map. The higher the pixel brightness in the figure, the stronger the embedding ability. Through comparison, it can be clearly observed that after enhancing the original carrier image, the pixel brightness in the larger square frame area in the embedding probability map changes from darker to brighter, indicating that the embedding ability of this area has been enhanced. Since the embedding ability of the original carrier image is immutable under the same load, the message can only be embedded into the pixels with a lower embedding probability. When the embedding ability of the larger square frame area of the enhanced image is improved, the selection of message embedding transfers from the smaller square frame area to this area, resulting in the pixel brightness in the smaller square frame area changing from bright to dim. This phenomenon shows that the enhanced image can carry additional secret messages with a higher embedding probability. It is verified that the large-capacity enhanced image generated by the present invention can effectively improve the embedding ability of the original carrier image at the same embedding cost.
[0073] It should be noted that the present invention is not limited to the above embodiments. The above embodiments are only examples, and the embodiments with the same composition and the same function and effect as the technical idea within the scope of the technical solution of the present invention are all included in the technical scope of the present invention. In addition, within the scope of not departing from the gist of the present invention, various deformations that those skilled in the art can think of for the embodiments and other ways constructed by combining some constituent elements of the embodiments are also included in the scope of the present invention.
Claims
1. A method for enhancing the content consistency of carrier images for large-capacity steganography, characterized in that, Including the following steps: S1. Split the color carrier image dataset into an R, G, B multi-channel format and feed it into the generation network G N to obtain the multi-channel embedding probability map M T p , and then the embedding probability map M T p and the noise map N T are fed into the embedding simulator E for message embedding to obtain the modified map M T ; S2: By modifying the mapping M T A stego-image is obtained. A steganalysis tool is used to minimize the difference between the stego-image and the cover image, and the embedding probability of each pixel in multiple channels is iteratively updated; S3: Utilize the multi-channel embedding probability map M T p Calculate the noise mixing weight matrix W for enhancing the image N ; S4: Apply a two-class flow noise strategy to the carrier image and perform adaptive mixing guided by the noise mixing weight matrix W N to obtain a noisy image C t0 ; S5: Feed C t0 into the diffusion model and constrain it with three regularizers: image sharpness, noise distribution, and adversarial transformation to guide the denoising process and generate the final denoised image C0.
2. A method for enhancing the content consistency of a carrier image for large-capacity steganography according to claim 1, characterized in that, In S1, the color carrier image dataset is split into an R, G, B multi-channel format and fed into the generation network G N to obtain a multi-channel embedding probability map M T p , which specifically includes the following steps: S11, sequentially extract the color carrier images in the dataset in batches of batch_size, and split the channels of the carrier images in each batch to obtain a quantity of grayscale images C T ; S12, feed the number of grayscale images C T to the generation network G N and ensure that the generation network can output the corresponding number of multi-channel embedding probability maps M T p , T ∈ {R, G, B}.
3. A method for enhancing the content consistency of carrier images for large-capacity steganography according to claim 1, characterized in that In the above S1, the embedding probability graph M T p and the noise map N T are fed into the embedding simulator E for message embedding to obtain the modified map M T , which specifically includes the following steps: S13. Send the embedded probability graph M T p The three different-channel embedded probability graphs corresponding to each color carrier image in T and the three groups of random noise mappings N into the embedding simulator E for message embedding together; S14. The embedding simulator E simulates the embedding process by sampling using pixel-level modification, mapping the random noise N T to the embedding probability map M T p for comparison, obtaining a pixel-level modification map M with three possible directions (+1, -1, 0) T .
4. A method for enhancing the content consistency of carrier images for large-capacity steganography according to claim 3, characterized in that The embedding rules of the embedding simulator E are as follows: M T p (x,y) has an embedding probability of M in the direction (+1) T p (x,y) has an embedding probability of M in the direction (-1) T p (x,y) / 2; If a certain random noise mapping element N T (x, y) is less than M T p the embedding probability in the direction (+1) of (x, y), then modify the mapping M T (x, y) = +1; If a random noise mapping element N T (x, y) is greater than M T p in (x, y) (the embedding probability in the 1 - direction (-1)), then modify the mapping M T (x, y) = -1; Otherwise, set M T (x,y)=0.
5. A method for enhancing the content consistency of a carrier image for large-capacity steganography according to claim 2, characterized in that In S2, the stego-image is obtained by modifying the mapping M T which specifically includes: By modifying the mapping M T from the previously separated multi-channel grayscale images C of R, G, and B T the stego image S can be obtained T : 。 6. A method for enhancing the content consistency of a carrier image for large-capacity steganography according to claim 1, characterized in that In S2, a steganalysis tool is used to minimize the difference between the stego-image and the carrier image, which specifically includes the following steps: S21, the steganalysis tool detects the embedded information in the input stego-image to obtain the embedding probability; S22, select the sampling image for the next round of iteration from the stego-image according to the embedding probability.
7. A method for enhancing the content consistency of a carrier image for large-capacity steganography according to claim 1, characterized in that, In the step S3, the multi-channel embedding probability map M is used T p to calculate the noise mixing weight matrix W for enhancing the image N Specifically, it includes: S31, use the embedding probability generator model to generate a multi-channel embedding probability map by inputting a color image; S32, perform binarization on the multi-channel embedding probability map to obtain a new binary embedding probability image B T ; S33. According to the binary embedding probability image B T , for any position (x, y), if there is at least one value of 1 in B T (x, y), then in the weight matrix W N , set the value at this position to 1, otherwise set it to 0.
8. A method for enhancing the content consistency of carrier images for large-capacity steganography according to claim 1, characterized in that In S4, a two-class flow noise strategy is adopted for the carrier image, specifically including: The two-class flow noise mechanism includes a creative flow and a stable flow; For the creative stream, random noise corresponding to the time step t is added to the carrier image C to generate a noisy image C O t ; Use a diffusion model to perform iterative denoising on C O t until the time step t0 is reached, at which point the denoised image C is obtained O t0 ; For a steady flow, DDIM inversion is used to add noise to the carrier image C and obtain a noisy image C S t0 。 9. A method for enhancing the content consistency of a carrier image for large-capacity steganography according to claim 8, characterized in that In the step S4, adaptive mixing is performed guided by the noise mixing weight matrix W N to obtain a noisy image C t0 , which specifically includes the following steps: When the dual-class flow noise generates two noise images C O t0 and C S t0 after that, dynamically adjust the usage ratios of C O t0 and C S t0 by the probability values embedded in the probability image, and adaptively mix the two noise images using the weight matrix W N ; For the image regions with a relatively high embedding probability, use C S t0 Maintain a relatively stable content structure. For the image regions with a relatively low embedding probability, use C O t0 Generate variants to enrich the detailed information in this region.
Citation Information
Patent Citations
Data enhancement method for steganalysis of depth image based on distribution preserving principle
CN113888423A
Training method of steganography analyzer based on automatic virtual data enhancement
CN115713663A
Image steganography system and method giving consideration to capacity and robustness
CN118333828A
Carrier-free information hiding method based on de-noising diffusion probability model
CN119071402A
High dynamic range image information hiding method
US20180075569A1