A robust image steganography method of multi-domain constraint secret diffusion embedding
Patent Information
- Application Number
- CN202610799067.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-04
AI Technical Summary
[0008]第一,现有的图像隐写方法大多采用直接映射、特征拼接或一次性嵌入方式,难以适应图像复杂的空间结构和细粒度纹理特征,容易造成边缘、高频纹理区域的失真和统计伪影,导致隐写图像的视觉保真度和结构一致性下降
[0020] The robust image steganography method proposed in this application, based on multi-domain constrained secret diffusion embedding, has the following advantages and positive effects compared with traditional image steganography methods.
Smart Images

Figure CN122698705A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the fields of image processing and information security technology, and in particular to a robust image steganography method with multi-domain constrained secret diffusion embedding. Background Technology
[0002] With the widespread application of digital images in social media, instant messaging, cloud storage, and multimedia information systems, they have become an important carrier for information transmission and data interaction. However, in open network environments, image data faces security risks such as unauthorized access, illegal theft, malicious tampering, and leakage of sensitive information during storage, transmission, and sharing. Image steganography, by hiding secret information within ordinary image carriers and completing covert transmission without alerting third parties, provides an important technical means for protecting sensitive information, covert communication, and secure data sharing, thus possessing significant research and application value.
[0003] Image steganography aims to embed secret information into a carrier image, making the generated steganographic image as close as possible to the original image in terms of visual perception and statistical distribution, thereby achieving covert communication. Traditional image steganography methods mainly achieve information embedding by modifying pixel values, frequency domain coefficients, or transform domain features, such as methods based on spatial domain pixel perturbation, small-amplitude coefficient modulation methods based on discrete cosine transform, and multi-scale embedding methods based on wavelet transform. These methods have advantages such as simple implementation and low computational cost, but they usually rely on manually designed embedding rules, making it difficult to fully adapt to complex ground object boundaries, dense textures, and multi-scale spatial structures in images. They are also prone to generating detectable statistical anomalies in high-frequency regions or directional texture regions, thus leading to image steganography failure.
[0004] In recent years, with the rapid development of deep learning technology, image steganography methods based on neural networks have received widespread attention. Existing methods typically employ an encoder-decoder structure, fusing features of the carrier image and secret information to generate a steganalyte, and then recovering the secret information through a decoding network. Other methods introduce generative networks, reversible neural networks, or diffusion models to enhance the quality of generated steganalytes and the ability to recover secret information. Among these, diffusion models, due to their strong image generation capabilities, can generate high-quality images with natural image priors during the stepwise denoising process, providing a new technical path for high-fidelity steganography.
[0005] However, existing deep steganography methods still have significant shortcomings in practical applications. On the one hand, images themselves contain rich texture details, edge contours, color variations, and spatial structure information. If secret information is injected using simple feature stitching, direct mapping, or one-time embedding, it can easily disrupt the content consistency and local structural continuity of the original image. This causes steganographic perturbations to concentrate in edge, texture, and high-frequency detail regions, resulting in decreased image fidelity and spatial domain statistical anomalies. On the other hand, while diffusion-based steganography methods have strong generation capabilities, directly injecting secret features during the diffusion process can easily interfere with the existing generation distribution of the diffusion model, causing the generation process to deviate from the data distribution of the original image, thereby affecting the quality of the steganographic image and the stability of secret information recovery.
[0006] Furthermore, existing steganalysis security constraints typically rely on training with a single steganalyzer, a single statistical feature, or a single feature domain. Their limited scope makes steganalysis networks adaptable only to specific detectors, statistical features, or domain distributions. When faced with unknown steganalysis models or detection methods based on different feature domains such as discrete cosine transform or wavelet transform, residual local anomalies, frequency coefficient anomalies, directional subband anomalies, and multi-scale energy distribution anomalies in steganalytes can still be detected. Therefore, relying solely on single-domain constraints is insufficient to comprehensively characterize the statistical changes of steganalysis perturbations in spatial structure, frequency distribution, and multi-scale transformation features, and it is also difficult to guarantee the security and robustness of steganalytes under complex detection conditions.
[0007] Based on the above analysis, the existing technology has the following two urgent technical problems that need to be solved.
[0008] First, most existing image steganography methods use direct mapping, feature splicing, or one-time embedding, which are difficult to adapt to the complex spatial structure and fine-grained texture features of images. This can easily cause distortion and statistical artifacts in edges and high-frequency texture areas, resulting in a decrease in the visual fidelity and structural consistency of the steganographic image.
[0009] Second, existing steganalysis security constraints are mostly concentrated in a single spatial domain or a single steganalysis analyzer, making it difficult to simultaneously constrain the statistical distribution of steganalysis perturbations in the high-frequency spatial domain, the discrete cosine transform domain, and the two-dimensional wavelet multi-level sub-bands. This results in insufficient security and generalization ability of steganalysis images when facing steganalysis detection methods in unknown spatial domains or transform domains. Summary of the Invention
[0010] To address the aforementioned technical problems, embodiments of this application propose a robust image steganography method with multi-domain constrained secret diffusion embedding, aiming to reduce the disruption of the visual structure and statistical distribution of the carrier image caused by secret embedding, and improve the imperceptibility, anti-detection ability, and secret image recovery accuracy of the steganographic image.
[0011] To achieve the above objectives, embodiments of this application propose a robust image steganography method using multi-domain constrained secret diffusion embedding. The method includes the following steps: S1, acquiring a carrier image and a secret image to be hidden; performing a two-dimensional multi-level discrete wavelet transform on the features of the original secret image to obtain low-frequency and high-frequency secret subbands, and generating multi-level secret guiding features accordingly; S2, constructing a zero-convolutional progressive injection branch to map the multi-level secret guiding features to multiple secret modulation features corresponding to different denoising levels; S3, using the carrier image as the diffusion embedding object, injecting the corresponding secret modulation features according to the denoising level during the reverse denoising process to generate a steganographic image containing secret information; S4, separating the carrier image and the steganographic image respectively... The input is fed into a multi-domain constraint module, which includes a spatial domain constraint branch and a transform domain constraint branch; S5, high-pass residuals and multi-scale spatial features are extracted through the spatial domain constraint branch, and spatial domain loss is calculated to make the stegana image close to the carrier image in terms of spatial domain statistical features; S6, DCT features and wavelet subband features are extracted through the transform domain constraint branch, and transform domain loss is calculated to make the stegana image close to the carrier image in terms of transform domain statistical features; S7, secret image reconstruction is performed using a secret recovery network, and joint optimization is performed based on steganalysis loss, recovery loss, spatial domain loss and transform domain loss to obtain an image steganalysis model that can achieve high-fidelity embedding and high-precision recovery, so as to achieve robust image steganalysis using the image steganalysis model.
[0012] Optionally, a two-dimensional multi-level discrete wavelet transform is performed on the original secret image features to obtain low-frequency and high-frequency secret sub-bands, and multi-level secret guidance features are generated accordingly, including: Secret image features constrained by key Two-dimensional multi-level discrete wavelet transform is performed to obtain low-frequency secret subbands and high-frequency secret subbands at multiple scales; For the The formula for calculating the two-dimensional discrete wavelet transform is as follows: ; ; in, Represents a two-dimensional discrete wavelet transform. Indicates the wavelet decomposition level. For the first The low-frequency secret subband of the layer contains global structural information of the secret image. , and All are the first The high-frequency secret subbands of the layer contain texture information, edge information, and detail information in different directions. ; The low-frequency and high-frequency secret subbands of each layer are input into the corresponding wavelet secret feature encoding network. Linear weighted aggregation is performed using a learnable set of weight parameters to construct the secret guidance features of each layer, ultimately obtaining a multi-level secret guidance feature set. This provides multi-scale secret conditions during diffusion-backward denoising. , For the first The secret guiding characteristics of the layer; The calculation process is expressed as follows: ; in, , , , For a learnable set of weight parameters, Indicates the first Wavelet secret feature encoding network corresponding to the layer.
[0013] Optionally, a zero-convolutional progressive injection branch is constructed to map multi-level secret guided features to multiple secret modulation features corresponding to different denoising levels, including: Constructing zero-convolutional progressive injection branches to guide multi-level secret feature sets The secret guiding features of each layer are scale-aligned and channel-mapped to obtain multiple secret modulation features corresponding to different denoising levels; For the first... of the diffusion steganography backbone network Each denoising layer corresponds to a secret modulation feature as follows: ; in, Indicates the first The secret modulation features corresponding to each denoising level Indicates the first Each denoising layer involves scale alignment and channel mapping operations, including upsampling, downsampling, convolutional mapping, or linear mapping. Indicates the relationship with the first The noise reduction level matches the first A secret guiding feature Indicates the first One zero-initialized convolutional layer; The initial parameters of the zero-initialized convolutional layer satisfy , , and They represent the first The initial convolution weights and biases of the zero-initialized convolutions are set. In the initial stage of training, the output of the zero-convolution asymptotic injection branch is close to zero to avoid destroying the original denoising ability of the diffusion steganalysis backbone. As the training process progresses, the zero-convolution asymptotic injection branch gradually learns the injection method of secret information. Multiple secret modulation features constitute a secret modulation feature set, represented as follows: , , This represents the total number of denoising layers in the diffusion steganography backbone network.
[0014] Optionally, using the carrier image as the diffusion embedding object, during the reverse denoising process, corresponding secret modulation features are injected according to the denoising level to generate a steganalytic image containing secret information, including: With carrier image For the diffusion-embedded object, forward diffusion is applied to add noise, resulting in noisy images at different time steps. The forward diffusion process is represented as: ; in, , , The diffuse noise figure is... Represents the identity matrix; In the reverse denoising process, the secret modulation feature set is... The diffusion steganography backbone is progressively injected in a manner matching the denoising level, for the first... The secret injection process for each denoising layer is as follows: ; in, The first part represents the diffusion steganography backbone network. The original features of each denoising level This represents the secret modulation feature of the corresponding level. This represents the denoising feature after injecting secret information; Finally, a steganalytic image containing the secret information is generated through reverse denoising. , , This represents a steganography generative network.
[0015] Optionally, the carrier image and the steganalyte image are input into the multi-domain constraint module, including: During the training phase, the carrier image and steganalysis Input the multi-domain constraint module separately. The multi-domain constraint module includes a spatial domain constraint branch and a transform domain constraint branch. The two constraint branches are used to constrain the consistency between the stegana and the carrier image from different statistical domains. Multi-domain projection set Represented as: ; in, Represents the set of spatial domain projections. Represents the set of projections in the transform domain; For any projection operator , This is combined with a shared steganalysis backbone network. Combined, they form a virtual steganalysis tool. , By leveraging the collaborative constraints of multiple virtual steganalyzers, the detectable traces in steganalytical images are reduced across different statistical domains.
[0016] Optionally, high-pass residuals and multi-scale spatial features are extracted through spatial domain constrained branching, and spatial domain loss is calculated to make the steganagraph image approximate the carrier image in terms of spatial domain statistical features, including: By using a pre-defined set of high-pass filter operators to perform convolution processing on the image input to the spatial domain constrained branch, texture edges and noise residual components are extracted, so that steganalysis remains imperceptible in visually sensitive areas. The images input to the spatial domain constraint branch are downsampled or upsampled at different ratios to simulate the observation environment at different resolutions. The spatial statistical consistency between the steganalysis image and the carrier image at each scale is calculated. By minimizing the spatial domain loss, the constrained steganalysis generation network suppresses the generation of detectable traces in local structures and multi-scale spatial distributions.
[0017] Optionally, DCT features and wavelet subband features are extracted through transform domain constrained branching, and transform domain loss is calculated to make the steganalysis image approximate the carrier image in terms of transform domain statistical features, including: The image input to the constrained branch of the transform domain is divided into several non-overlapping blocks and subjected to discrete cosine transform to extract mid-to-high frequency coefficient features. The constrained steganalysis does not change the frequency envelope of the original image. By using two-dimensional wavelet transform, the image input to the transform domain constraint branch is decomposed into sub-bands of different frequencies and directions. The statistical moment features of each level of sub-band are extracted to ensure that the energy distribution of steganalytic perturbation in the transform domain conforms to the distribution law of natural images. Through collaborative constraint transform domain loss, the steganalytic generation network is forced to eliminate statistical anomalies in the frequency domain and wavelet sub-bands.
[0018] Optionally, the method of using a secret recovery network for secret image reconstruction and performing joint optimization based on steganalysis loss, recovery loss, spatial domain loss, and transform domain loss includes: Recovering the network using secret images For the generated stegana By performing reverse decoding, the reconstructed secret image is obtained. Using the combined loss function Perform end-to-end network optimization; This can be expressed by the formula: ; in, The pixel-level embedding loss between the carrier image and the stegana. The loss represents the consistency loss between the original secret image and the reconstructed secret image. To encompass multi-domain constrained loss that includes both spatial domain loss and transform domain loss, , , They are respectively , , The corresponding weighting coefficients for the loss term; By alternately updating the parameters of the steganalysis generation network, the secret recovery network, and the multi-domain constraint module through the backpropagation algorithm until convergence, an image steganalysis model capable of achieving high-fidelity embedding and high-precision recovery is obtained.
[0019] Optionally, during training, the Adam optimizer is used to update network parameters, the batch size is set to 4, and the learning rate is scheduled using a cosine annealing strategy, with an initial learning rate set to... The minimum learning rate is set to The entire training process consists of 300 epochs to ensure stable convergence of the network. After training, the optimal network parameters are saved to obtain an image steganalysis model that can achieve high-fidelity embedding and high-precision restoration.
[0020] The robust image steganography method proposed in this application, based on multi-domain constrained secret diffusion embedding, has the following advantages and positive effects compared with traditional image steganography methods.
[0021] First, this application integrates the steganography embedding process of secret information into a diffusion-based image generation framework. By using a zero-convolution progressive injection branch, the secret features guide the generation process in stages, enabling the secret information to be embedded into the carrier image in a smooth and controlled manner. This reduces the one-time disturbance of the carrier's spatial structure, edge texture, and overall distribution caused by secret embedding, thereby improving the imperceptibility and high-fidelity expression of the steganographic image.
[0022] Second, this application designs a feature decoupling module based on two-dimensional multi-level discrete wavelet transform. By utilizing the distribution characteristics of images on different frequency sub-bands, secret information is decomposed and mapped at multiple scales, achieving highly adaptive matching between secret features and carrier texture features.
[0023] Third, this application constructs a multi-domain collaborative constraint mechanism. By performing joint optimization in the spatial domain (high-pass residuals and multi-scale projection) and the transform domain (DCT coefficients and wavelet statistical moments), it constrains the distribution consistency between the steganalysis image and the carrier image from multiple statistical perspectives, suppresses detectable traces generated by the secret embedding process, and thus improves the resistance, security and imperceptibility of the steganalysis image to known and unknown steganalysis methods. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. The following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.
[0025] Figure 1 This is a flowchart of a robust image steganography method with multi-domain constrained secret diffusion embedding provided in one embodiment of this application; Figure 2 This is a network structure diagram of a robust image steganography method with multi-domain constrained secret diffusion embedding provided in one embodiment of this application; Figure 3 This is a schematic diagram of a multi-scale secret feature encoding and alignment structure based on 2DWT provided in one embodiment of this application; Figure 4 This is a comparison diagram of the effects of the method provided in one embodiment of this application with other methods. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been presented in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.
[0027] One embodiment of this application proposes a robust image steganography method with multi-domain constrained secret diffusion embedding. The implementation details of the robust image steganography method with multi-domain constrained secret diffusion embedding proposed in this embodiment are described in detail below. The following implementation details are provided for ease of understanding only and are not necessary for implementing this solution.
[0028] The specific process of the robust image steganography method with multi-domain constrained secret diffusion embedding proposed in this embodiment can be described as follows: Figure 1 As shown, its architecture is as follows Figure 2 As shown, the method includes: S1. Obtain the carrier image and the secret image to be hidden. Perform two-dimensional multi-level discrete wavelet transform on the features of the original secret image to obtain the low-frequency secret sub-band and the high-frequency secret sub-band, and generate multi-level secret guidance features accordingly.
[0029] Specifically, the first step in achieving robust image steganography is to acquire the carrier image and the secret image to be hidden, perform a two-dimensional multi-level discrete wavelet transform on the original secret image features extracted from the secret image to be hidden to obtain low-frequency secret sub-bands and high-frequency secret sub-bands, and generate multi-level secret guidance features accordingly.
[0030] like Figure 2 , Figure 3 As shown, this embodiment uses a steganography architecture based on two-dimensional multi-scale wavelet decoupling and zero-convolutional progressive injection to generate steganalytic images. This steganography architecture is the core structure of this application and mainly includes a two-dimensional multi-level discrete wavelet transform module, a wavelet secret feature encoding network, a zero-convolutional progressive injection branch, and a diffusion steganography backbone network.
[0031] After acquiring the carrier image and the secret image to be hidden, the secret image to be hidden is input into a two-dimensional multi-level discrete wavelet transform module to perform key-constrained secret image feature transformation. Two-dimensional multi-level discrete wavelet transform is performed to obtain low-frequency secret subbands and high-frequency secret subbands at multiple scales.
[0032] For the The formula for calculating the two-dimensional discrete wavelet transform is as follows: ; ; in, Represents a two-dimensional discrete wavelet transform. Indicates the wavelet decomposition level. For the first The low-frequency secret subband of the layer contains global structural information of the secret image. , and All are the first The high-frequency secret subbands of the layer contain texture information, edge information, and detail information in different directions. .
[0033] After obtaining the low-frequency and high-frequency secret subbands of each layer, these subbands can be input into the corresponding wavelet secret feature encoding network to obtain independent feature representations for the four subbands. Then, a linear weighted aggregation is performed using a learnable set of weight parameters to construct the secret guidance features of each layer, ultimately resulting in a multi-level secret guidance feature set. This provides multi-scale secret conditions during diffusion-backward denoising.
[0034] , For the first The secret guiding characteristics of the layer will be as Figure 2 The input of the zero-convolution asymptotic injection branch is used to generate secret modulation features that match different denoising levels of the diffusion steganalysis backbone.
[0035] The calculation process is expressed as follows: ; in, , , , For a learnable set of weight parameters, Indicates the first Wavelet secret feature encoding network corresponding to the layer.
[0036] S2 constructs a zero-convolution progressive injection branch, mapping multi-level secret guidance features to multiple secret modulation features corresponding to different denoising levels.
[0037] Specifically, after generating multi-level secret guidance features, zero-convolution progressive injection branches can be constructed to map the multi-level secret guidance features into multiple secret modulation features corresponding to different denoising levels.
[0038] like Figure 2 As shown, this embodiment converts multi-level secret guidance features into secret modulation features acceptable to different denoising levels of the diffusion steganalysis backbone network through a zero-convolutional progressive injection branch. Since the feature size and number of channels are different in different denoising levels of the diffusion steganalysis backbone network, scale alignment and channel mapping are required for the multi-level secret guidance features.
[0039] For the first... of the diffusion steganography backbone network Each denoising layer corresponds to a secret modulation feature as follows: ; in, Indicates the first The secret modulation features corresponding to each denoising level Indicates the first Each denoising layer involves scale alignment and channel mapping operations, including upsampling, downsampling, convolutional mapping, or linear mapping. Indicates the relationship with the first The noise reduction level matches the first A secret guiding feature Indicates the first A zero-initialized convolutional layer.
[0040] The initial parameters of the zero-initialized convolutional layer satisfy , , and They represent the first The initial convolution weights and biases of the zero-initialized convolutions are set. In the initial stage of training, the output of the zero-convolution asymptotic injection branch is close to zero to avoid destroying the original denoising ability of the diffusion steganalysis backbone. As the training process progresses, the zero-convolution asymptotic injection branch gradually learns the injection method of secret information.
[0041] By performing the above processing on multiple denoising layers of the diffusion steganalysis backbone network, the secret modulation feature set can be obtained. , , This represents the total number of denoising layers in the diffusion steganography backbone network. Used for layer-by-layer injection in the subsequent inverse denoising process.
[0042] S3 uses the carrier image as the diffusion embedding object. During the reverse denoising process, it injects the corresponding secret modulation features according to the denoising level to generate a steganalytic image containing secret information.
[0043] Specifically, after obtaining multiple secret modulation features corresponding to different denoising levels, the carrier image can be used as the diffusion embedding object. During the reverse denoising process, the corresponding secret modulation features are injected according to the denoising level to generate a steganalytic image containing secret information.
[0044] like Figure 2 As shown, the diffusion steganalysis backbone network uses carrier images To diffuse the embedded object and receive secret modulation features during the inverse denoising process, a steganalytic image containing secret information is generated. .
[0045] Diffusion steganalysis backbone network with carrier image For the diffusion-embedded object, forward diffusion is applied to add noise, resulting in noisy images at different time steps. .
[0046] The forward diffusion process is represented as: ; in, , , The diffuse noise figure is... Represents the identity matrix.
[0047] In the reverse denoising process, the secret modulation feature set is... The diffusion steganography backbone is progressively injected in a manner matching the denoising level, for the first... The secret injection process for each denoising layer is as follows: ; in, The first part represents the diffusion steganography backbone network. The original features of each denoising level This represents the secret modulation feature of the corresponding level. This represents the denoising feature after injecting secret information.
[0048] Finally, a steganalytic image containing the secret information is generated through reverse denoising. , , This represents a steganography generative network.
[0049] This process allows secret information to be gradually embedded during the diffusion-backward generation process, while maintaining the visual content consistency between the steganographic image and the carrier image.
[0050] S4. Input the carrier image and the steganalysis image into the multi-domain constraint module, which includes a spatial domain constraint branch and a transform domain constraint branch.
[0051] Specifically, after generating the steganalysis image, the carrier image and the steganalysis image can be input into the multi-domain constraint module, which includes a spatial domain constraint branch and a transform domain constraint branch.
[0052] like Figure 2 As shown, the multi-domain collaborative constraint module is used to constrain the consistency between the steganalyte image and the carrier image. The multi-domain constraint module includes a spatial domain constraint branch, a transform domain constraint branch, and a shared steganalysis backbone network.
[0053] During the training phase, the carrier image and steganalysis Input the multi-domain constraint module separately. The multi-domain constraint module includes a spatial domain constraint branch and a transform domain constraint branch. The two constraint branches are used to constrain the consistency between the stegana and the carrier image from different statistical domains.
[0054] Multi-domain projection set Represented as: ; in, Represents the set of spatial domain projections. Represents the set of projections in the transform domain and the set of projections in the spatial domain. Used to characterize local texture, edge residuals, and multi-scale spatial structure differences, while transform domain projection sets It is used to characterize frequency energy, DCT coefficients, and wavelet subband statistical differences.
[0055] For any projection operator , This is combined with a shared steganalysis backbone network. Combined, they form a virtual steganalysis tool. , , Indicates the input image. The statistical domain representation of the input image after passing through the projection operator is represented by a steganalysis. Through the collaborative constraints of multiple virtual steganalysis analyzers, the steganalysis image is made close to the carrier image in different statistical domains, and the detectable traces are reduced in different statistical domains.
[0056] S5 extracts high-pass residuals and multi-scale spatial features through spatial domain constraint branches, and calculates spatial domain loss, so that the steganalysis image is close to the carrier image in terms of spatial domain statistical features.
[0057] Specifically, spatial domain constraint branching is used for carrier images and steganalysis Spatial domain projection is performed to extract local texture, edge residuals, and multi-scale spatial statistical features.
[0058] This embodiment utilizes a pre-defined set of high-pass filter operators to convolve the image input to the spatial domain constraint branch, extracting its texture edges and noise residual components, thus keeping steganalytic perturbations imperceptible in visually sensitive areas. Different downsampling or upsampling transformations are applied to the image input to the spatial domain constraint branch to simulate observation environments at different resolutions, and the spatial statistical consistency between the steganalytic image and the carrier image at each scale is calculated. By minimizing spatial domain loss, the constrained steganalytic generation network suppresses the generation of detectable traces in local structures and multi-scale spatial distributions.
[0059] S6 extracts DCT features and wavelet subband features through transform domain constraint branching and calculates transform domain loss, making the steganalysis image close to the carrier image in terms of transform domain statistical features.
[0060] Specifically, transform domain constraint branching is used for carrier images and steganalysis Transform domain projection is performed to constrain the anomalous shift of steganographic perturbations in frequency energy distribution and subband statistical characteristics.
[0061] This embodiment divides the image input to the transform domain constraint branch into several non-overlapping blocks and performs discrete cosine transform to extract mid-to-high frequency coefficient features, ensuring that the steganalytic perturbation does not change the frequency envelope of the original image. Using two-dimensional wavelet transform, the image input to the transform domain constraint branch is decomposed into sub-bands of different frequencies and directions, and the statistical moment features of each level of sub-band are extracted to ensure that the energy distribution of the steganalytic perturbation in the transform domain conforms to the distribution law of the natural image. Through collaborative constraint transform domain loss, the steganalytic generation network is forced to eliminate statistical anomalies in the frequency domain and wavelet sub-bands.
[0062] S7 utilizes a secret recovery network for secret image reconstruction and performs joint optimization based on steganalysis loss, recovery loss, spatial domain loss, and transform domain loss to obtain an image steganalysis model that can achieve high-fidelity embedding and high-precision recovery, thereby enabling robust image steganalysis using the image steganalysis model.
[0063] Specifically, this embodiment utilizes a secret recovery network for secret image reconstruction and performs joint optimization based on steganalysis loss, recovery loss, spatial domain loss, and transform domain loss to obtain an image steganalysis model that can achieve high-fidelity embedding and high-precision recovery, thereby enabling robust image steganalysis using the image steganalysis model.
[0064] like Figure 2 As shown, this embodiment utilizes a secret image recovery network. For the generated stegana By performing reverse decoding, the reconstructed secret image is obtained. Using the combined loss function Perform end-to-end network optimization.
[0065] This can be expressed by the formula: ; in, The pixel-level embedding loss between the carrier image and the stegana. The loss represents the consistency loss between the original secret image and the reconstructed secret image. To encompass multi-domain constrained loss that includes both spatial domain loss and transform domain loss, , , They are respectively , , The corresponding weight coefficients for the loss term.
[0066] By alternately updating the parameters of the steganalysis generation network, the secret recovery network, and the multi-domain constraint module through the backpropagation algorithm until convergence, an image steganalysis model capable of achieving high-fidelity embedding and high-precision recovery is obtained.
[0067] In this embodiment, the Adam optimizer is used to update network parameters during training, the batch size is set to 4, and the learning rate is scheduled using a cosine annealing strategy, with an initial learning rate set to... The minimum learning rate is set to The entire training process consists of 300 epochs to ensure stable convergence of the network. After training, the optimal network parameters are saved to obtain an image steganalysis model that can achieve high-fidelity embedding and high-precision restoration.
[0068] The robust image steganography method with multi-domain constrained secret diffusion embedding proposed in this embodiment has the following advantages and positive effects compared with traditional image steganography methods.
[0069] First, this embodiment integrates the steganography embedding process of secret information into the diffusion-based image generation framework. By using the zero-convolution progressive injection branch, the secret features guide the generation process in stages, enabling the secret information to be embedded into the carrier image in a smooth and controlled manner. This reduces the one-time disturbance of the carrier's spatial structure, edge texture, and overall distribution caused by secret embedding, thereby improving the imperceptibility and high-fidelity expression capability of the steganographic image.
[0070] Second, this embodiment designs a feature decoupling module based on two-dimensional multi-level discrete wavelet transform. By utilizing the distribution characteristics of the image on different frequency sub-bands, the secret information is decomposed and mapped at multiple scales, achieving a highly adaptive matching between secret features and carrier texture features.
[0071] Third, this embodiment constructs a multi-domain collaborative constraint mechanism. By performing joint optimization in the spatial domain (high-pass residuals and multi-scale projection) and the transform domain (DCT coefficients and wavelet statistical moments), it constrains the distribution consistency between the steganalysis image and the carrier image from multiple statistical perspectives, suppresses detectable traces generated by the secret embedding process, and thus improves the resistance, security and imperceptibility of the steganalysis image to known and unknown steganalysis methods.
[0072] The technical solution presented in this embodiment has extremely high expected returns and commercial value after transformation. This embodiment proposes a robust image steganography method with multi-domain constrained secret diffusion embedding, which solves the problem of visual artifacts and statistical traces easily generated by existing deep steganography models against complex terrain backgrounds. The steganographic images obtained through this embodiment not only possess extremely high spatial fidelity and visual concealment, but also achieve high-precision lossless recovery of secret information. It can be widely applied in fields such as data copyright protection, geographic information ownership confirmation, covert secure communication, and anti-counterfeiting of sensitive assets. With the acceleration of data commercialization and strategic development, this method can provide core technical support for secure data distribution and intellectual property protection, possessing significant economic benefits and social security value.
[0073] The technical solution of this embodiment fills a technological gap in the industry both domestically and internationally. This embodiment fills the gap in methods for achieving high-fidelity image steganography under multi-domain statistical constraints using generative diffusion models. Most existing steganography methods adopt a direct mapping paradigm based on discriminative networks, which not only ignores the powerful prior protection capabilities of generative models but also often struggles to simultaneously ensure statistical security in the spatial and frequency domains. This embodiment innovatively introduces a combination of zero-convolution mechanism and diffusion-based inverse denoising process, along with a multi-domain collaborative strategy, to achieve joint optimization of generative steganography and cross-domain security constraints without losing key image texture information, providing a novel technical path for steganography tasks on high-dimensional complex images.
[0074] The technical solution of this embodiment solves a long-standing technical problem that has been desired to be solved but has remained unsolved. This embodiment effectively resolves the core contradiction between "embedding strength control" and "generating distribution fidelity" in steganography. When using a diffusion model for steganography, direct feature injection often causes the generated steganographic image to deviate from the original carrier distribution, producing statistical biases that are difficult to eliminate. This embodiment introduces a zero-convolutional progressive injection branch, enabling the model to gradually learn the embedding weights of secret information from "zero interference" using residual learning. This successfully overcomes the challenge of embedding large amounts of secret information while maintaining the high spatial-spectral fidelity of the image, significantly improving the robustness of the steganography system in complex environments.
[0075] The technical solution of this embodiment overcomes technical bias. The progressive injection and multi-domain joint architecture proposed in this embodiment changes the traditional steganography method of directly splicing secret information with the carrier image, one-time mapping, or transform domain superposition. Through the zero-convolution progressive injection branch, this embodiment enables the secret features to be gradually integrated into the carrier image in a smooth and controlled manner during the diffusion generation process, avoiding abrupt disturbances in spatial structure and texture distribution caused by strong coupling embedding of secret information. At the same time, this embodiment introduces multi-domain collaborative constraints of spatial high-pass residuals, multi-scale projection, DCT coefficients, and two-dimensional wavelet subband statistical features to jointly suppress detectable artifacts generated during the steganography process from multiple statistical perspectives, overcoming the problem that existing methods rely on a single steganalyzer or a single statistical domain constraint and are difficult to resist unknown detection methods. Therefore, this embodiment can effectively improve the visual imperceptibility, multi-domain statistical consistency, and security generalization ability of the steganographic image while ensuring the effective embedding and recovery of secret information.
[0076] The steps described above are merely for clarity in describing the technical solution. In actual implementation, they can be combined into one step, or certain steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Any insignificant modifications or designs added to the algorithm or process, as long as they do not change the core of the algorithm or process, are also within the scope of protection of this application.
[0077] In one embodiment, to verify the effectiveness of the robust image steganography method with multi-domain constrained secret diffusion embedding proposed in this application, we conducted relevant simulation experiments.
[0078] 1) Simulation experimental conditions.
[0079] The hardware platform for this simulation experiment is an NVIDIA GeForce RTX 3090 GPU.
[0080] The software platform for this simulation experiment is the PyTorch deep learning framework.
[0081] This simulation experiment uses the CAVE dataset. This dataset contains 32 real-world scene images, including various everyday objects and complex textures. Acquired under controlled lighting conditions, the dataset covers a spectral range of 400nm to 700nm, encompassing 31 spectral bands. This dataset can be used to verify the steganalysis fidelity and secret recovery performance of the proposed method under standard lighting and multi-object texture scenes.
[0082] 2) Evaluation indicators.
[0083] To quantitatively evaluate the method of this invention from both spatial imperceptibility and spectral fidelity perspectives, this simulation experiment selects five widely used evaluation metrics, including Peak Signal-to-Noise Ratio (PSNR), Root Mean Square Error (RMSE), Spectral Angle Mapping (SAM), Correlation Coefficient (CC), and Relative Global Dimensionless Synthesis Error (ERGAS). Higher PSNR and CC, and lower RMSE, SAM, and ERGAS, indicate less spatial distortion between the steganalyte and the carrier image, better spectral preservation, and superior overall performance of the method. These metrics provide an objective and comprehensive evaluation of the model's performance in terms of spatial reconstruction quality, spectral consistency, and global error.
[0084] 3) Experimental content and results analysis.
[0085] To verify the effectiveness of this method, it was trained and tested on the CAVE dataset, and compared with methods such as LSB, DAH, HiNet, DeepMIH, AIS, and StegFormer.
[0086] from Figure 4 As can be seen, traditional steganography methods (such as LSB) perform the worst, with a PSNR below 30 dB and a SAM greater than 19, indicating limited invisibility and poor restoration fidelity. Although deep learning-based methods (such as HiNet and DeepMIH) have improved invisibility, these methods still struggle to balance various metrics. For example, HiNet achieves a relatively low ERGAS, but fails to maintain high fidelity in secret image reconstruction, with a PSNR of only 28.59 dB. Among the comparative methods, DeepMIH and StegFormer are the most competitive. DeepMIH ranks second in PSNR and RMSE on carrier / steganographic image pairs, while StegFormer achieves the best results in SAM and CC values for both steganographic and restored images. Compared to these methods, our proposed method achieves state-of-the-art performance across all evaluation metrics. It achieves the highest PSNR for steganographic images (35.2692 dB) and the highest CC value (0.9842), while significantly reducing spectral distortion to a SAM of 8.1717.
[0087] In the various embodiments described above, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product, which includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the processes or functions described in this application are generated, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive, SSD), etc.
[0088] It will be understood by those skilled in the art that the above embodiments are specific embodiments for implementing this application. In practical applications, various changes can be made in form, detail, and expression without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application. Any modifications, equivalent substitutions, combinations, adjustments, or improvements made to the image steganography process, two-dimensional multi-level discrete wavelet transform method, zero-convolution asymptotic injection structure, multi-domain cooperative constraint module, secret image restoration network, loss function settings, or training methods in the various embodiments of this application should be covered within the scope of protection of this application.
Claims
1. A robust image steganography method using multi-domain constrained secret diffusion embedding, characterized in that, include: S1. Obtain the carrier image and the secret image to be hidden. Perform two-dimensional multi-level discrete wavelet transform on the features of the original secret image to obtain the low-frequency secret sub-band and the high-frequency secret sub-band, and generate multi-level secret guidance features accordingly. S2, construct zero-convolution progressive injection branch, and map multi-level secret guided features to multiple secret modulation features corresponding to different denoising levels; S3 uses the carrier image as the diffusion embedding object. During the reverse denoising process, the corresponding secret modulation features are injected according to the denoising level to generate a steganalytic image containing secret information. S4, input the carrier image and the steg image to the multi-domain constraint module respectively. The multi-domain constraint module includes a spatial domain constraint branch and a transform domain constraint branch. S5 extracts high-pass residuals and multi-scale spatial features through spatial domain constrained branches, and calculates spatial domain loss, so that the steganalysis image is close to the carrier image in terms of spatial domain statistical features. S6 extracts DCT features and wavelet subband features through transform domain constraint branching and calculates transform domain loss to make the steganalysis image close to the carrier image in terms of transform domain statistical features; S7 utilizes a secret recovery network for secret image reconstruction and performs joint optimization based on steganalysis loss, recovery loss, spatial domain loss, and transform domain loss to obtain an image steganalysis model that can achieve high-fidelity embedding and high-precision recovery, thereby enabling robust image steganalysis using the image steganalysis model.
2. The robust image steganography method with multi-domain constrained secret diffusion embedding according to claim 1, characterized in that, A two-dimensional multi-level discrete wavelet transform is performed on the original secret image features to obtain low-frequency and high-frequency secret sub-bands, and multi-level secret guidance features are generated based on these features, including: Secret image features constrained by key Two-dimensional multi-level discrete wavelet transform is performed to obtain low-frequency secret subbands and high-frequency secret subbands at multiple scales; For the The formula for calculating the two-dimensional discrete wavelet transform is as follows: ; ; in, Represents a two-dimensional discrete wavelet transform. Indicates the wavelet decomposition level. For the first The low-frequency secret subband of the layer contains global structural information of the secret image. , and All are the first The high-frequency secret subbands of the layer contain texture information, edge information, and detail information in different directions. ; The low-frequency and high-frequency secret subbands of each layer are input into the corresponding wavelet secret feature encoding network. Linear weighted aggregation is performed using a learnable set of weight parameters to construct the secret guidance features of each layer, ultimately obtaining a multi-level secret guidance feature set. This provides multi-scale secret conditions during diffusion-backward denoising. , For the first The secret guiding characteristics of the layer; The calculation process is expressed as follows: ; in, , , , For a learnable set of weight parameters, Indicates the first Wavelet secret feature encoding network corresponding to the layer.
3. The robust image steganography method with multi-domain constrained secret diffusion embedding according to claim 2, characterized in that, A zero-convolutional progressive injection branch is constructed to map multi-level secret guided features to multiple secret modulation features corresponding to different denoising levels, including: Constructing zero-convolutional progressive injection branches to guide multi-level secret feature sets The secret guiding features of each layer are scale-aligned and channel-mapped to obtain multiple secret modulation features corresponding to different denoising levels; For the first... of the diffusion steganography backbone network Each denoising layer corresponds to a secret modulation feature as follows: ; in, Indicates the first The secret modulation features corresponding to each denoising level Indicates the first Each denoising layer involves scale alignment and channel mapping operations, including upsampling, downsampling, convolutional mapping, or linear mapping. Indicates the relationship with the first The noise reduction level matches the first A secret guiding feature Indicates the first One zero-initialized convolutional layer; The initial parameters of the zero-initialized convolutional layer satisfy , , and They represent the first The initial convolution weights and biases of the zero-initialized convolutions are set. In the initial stage of training, the output of the zero-convolution asymptotic injection branch is close to zero to avoid destroying the original denoising ability of the diffusion steganalysis backbone. As the training process progresses, the zero-convolution asymptotic injection branch gradually learns the injection method of secret information. Multiple secret modulation features constitute a secret modulation feature set, represented as follows: , , This represents the total number of denoising layers in the diffusion steganography backbone network.
4. The robust image steganography method with multi-domain constrained secret diffusion embedding according to claim 3, characterized in that, Using the carrier image as the diffusion embedding object, during the reverse denoising process, corresponding secret modulation features are injected according to the denoising level to generate a steganalyte image containing secret information, including: With carrier image For the diffusion-embedded object, forward diffusion is applied to add noise, resulting in noisy images at different time steps. The forward diffusion process is represented as: ; in, , , The diffuse noise figure is... Represents the identity matrix; In the reverse denoising process, the secret modulation feature set is... The diffusion steganography backbone is progressively injected in a manner matching the denoising level, for the first... The secret injection process for each denoising layer is as follows: ; in, The first part represents the diffusion steganography backbone network. The original features of each denoising level This represents the secret modulation feature of the corresponding level. This represents the denoising feature after injecting secret information; Finally, a steganalytic image containing the secret information is generated through reverse denoising. , , This represents a steganography generative network.
5. The robust image steganography method with multi-domain constrained secret diffusion embedding according to claim 4, characterized in that, The carrier image and the stegana image are input into the multi-domain constraint module, including: During the training phase, the carrier image and steganalysis Input the multi-domain constraint module separately. The multi-domain constraint module includes a spatial domain constraint branch and a transform domain constraint branch. The two constraint branches are used to constrain the consistency between the stegana and the carrier image from different statistical domains. Multi-domain projection set Represented as: ; in, Represents the set of spatial domain projections. Represents the set of projections in the transform domain; For any projection operator , This is combined with a shared steganalysis backbone network. Combined, they form a virtual steganalysis tool. , By leveraging the collaborative constraints of multiple virtual steganalyzers, the detectable traces in steganalytical images are reduced across different statistical domains.
6. The robust image steganography method with multi-domain constrained secret diffusion embedding according to claim 1, characterized in that, High-pass residuals and multi-scale spatial features are extracted through spatial domain constrained branching, and spatial domain loss is calculated to make the steganagraph image approximate the carrier image in terms of spatial domain statistical features, including: By using a pre-defined set of high-pass filter operators to perform convolution processing on the image input to the spatial domain constrained branch, texture edges and noise residual components are extracted, so that steganalysis remains imperceptible in visually sensitive areas. The images input to the spatial domain constraint branch are downsampled or upsampled at different ratios to simulate the observation environment at different resolutions. The spatial statistical consistency between the steganalysis image and the carrier image at each scale is calculated. By minimizing the spatial domain loss, the constrained steganalysis generation network suppresses the generation of detectable traces in local structures and multi-scale spatial distributions.
7. The robust image steganography method with multi-domain constrained secret diffusion embedding according to claim 1, characterized in that, By extracting DCT features and wavelet subband features through transform-domain constrained branching and calculating transform-domain loss, the steganalysis image is made to approximate the carrier image in terms of transform-domain statistical features, including: The image input to the constrained branch of the transform domain is divided into several non-overlapping blocks and subjected to discrete cosine transform to extract mid-to-high frequency coefficient features. The constrained steganalysis does not change the frequency envelope of the original image. By using two-dimensional wavelet transform, the image input to the transform domain constraint branch is decomposed into sub-bands of different frequencies and directions. The statistical moment features of each level of sub-band are extracted to ensure that the energy distribution of steganalytic perturbation in the transform domain conforms to the distribution law of natural images. Through collaborative constraint transform domain loss, the steganalytic generation network is forced to eliminate statistical anomalies in the frequency domain and wavelet sub-bands.
8. A robust image steganography method with multi-domain constrained secret diffusion embedding according to any one of claims 1 to 7, characterized in that, The method of using a secret recovery network for secret image reconstruction, and performing joint optimization based on steganalysis loss, recovery loss, spatial domain loss, and transform domain loss, includes: Recovering the network using secret images For the generated stegana By performing reverse decoding, the reconstructed secret image is obtained. Using the combined loss function Perform end-to-end network optimization; This can be expressed by the formula: ; in, The pixel-level embedding loss between the carrier image and the stegana. The loss represents the consistency loss between the original secret image and the reconstructed secret image. To encompass multi-domain constrained loss that includes both spatial domain loss and transform domain loss, , , They are respectively , , The corresponding weighting coefficients for the loss term; By alternately updating the parameters of the steganalysis generation network, the secret recovery network, and the multi-domain constraint module through the backpropagation algorithm until convergence, an image steganalysis model capable of achieving high-fidelity embedding and high-precision recovery is obtained.
9. A robust image steganography method with multi-domain constrained secret diffusion embedding according to claim 8, characterized in that, During training, the Adam optimizer is used to update network parameters, the batch size is set to 4, and the learning rate is scheduled using a cosine annealing strategy. The initial learning rate is set to... The minimum learning rate is set to The entire training process consists of 300 epochs to ensure stable convergence of the network. After training, the optimal network parameters are saved to obtain an image steganalysis model that can achieve high-fidelity embedding and high-precision restoration.