Image steganography method, device, electronic device and storage medium
Through the Res-SS2D module and edge feature fusion mechanism, the problems of insufficient generation quality and anti-detection ability of image steganography in the existing technology are solved, and an efficient and secure image steganography method is realized, which is suitable for information hiding tasks in the field of information security.
Patent Information
- Application Number
- CN202510921583.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-04
AI Technical Summary
Existing CNN-based generative steganography techniques have shortcomings in terms of generated image quality and resistance to steganalysis. They struggle with effective global image modeling, resulting in poor visual quality and easy detection of generated encrypted images. While ViT overcomes the limitations of global feature modeling, it suffers from quadratic time complexity and low computational efficiency when processing high-resolution images.
The Res-SS2D module is used for four-way global space modeling. Combined with the edge feature fusion mechanism, the secret information is preferentially embedded in the high-frequency area. A multi-objective loss function is designed to optimize the generation quality. Efficient steganalytic feature modeling is achieved by constructing an encoder and decoder network.
The generated secret images have high visual quality, strong resistance to steganalysis, high computational efficiency, adaptability to different image types, and are suitable for information hiding tasks in the field of information security.
Smart Images

Figure CN120450937B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information hiding, and more specifically, to an image steganography method, device, electronic device and storage medium. Background Art
[0002] In the field of information security, steganography is a technique that embeds secret information into intangible media, such as images, audio, and video. Unlike traditional cryptography, which protects data by encrypting only the content of the information, steganography is unique in that it not only hides the information itself but also conceals the fact that the information is encrypted, significantly enhancing the security of secret information during communication.
[0003] Traditional image steganography techniques primarily focus on finding areas within an image where distortion is minimal after modification, enabling the covert embedding of information. The most classic approach, Least Significant Bit (LSB), hides information by modifying the least significant bit of an image pixel. However, LSB techniques cannot customize embedding schemes for different images, resulting in abnormal statistical characteristics of image pixels, making them susceptible to detection by steganalyzers. To overcome this limitation, researchers have introduced the concept of distortion functions, which more precisely control the distortion introduced by the modifications and optimize the steganographic effect. With the rise of deep learning technology, some researchers have begun using neural networks to automatically optimize distortion functions, further improving the performance and concealment of steganography.
[0004] In recent years, the rapid development of deep learning technology has brought new opportunities to image steganography. By building complex neural network architectures, deep learning-based steganography can learn complex patterns and features in image data, thereby more effectively concealing secret information. For example, some researchers have developed an encoder-decoder-discriminator architecture, which not only achieves efficient information hiding but also lays a foundation for subsequent research. Furthermore, by incorporating advanced techniques such as attention mechanisms, steganography quality has been further improved. These methods have, to a certain extent, overcome the limitations of traditional steganography, demonstrating greater potential for its application in modern information security.
[0005] Although steganography methods based on deep convolutional neural networks (CNNs) have achieved some success, their limitations are becoming increasingly apparent. CNNs rely on local convolution operations. While they excel at capturing local features, their limited receptive field makes them incapable of effectively capturing long-range dependencies and global structural information within an image. This results in the steganographic network being unable to fully learn the global statistical properties of the image when generating a secret image, resulting in poor visual quality and resistance to steganalysis. To address this issue, researchers introduced the Visual Transformer (ViT). ViT overcomes the limitations of traditional CNNs in modeling global features through its self-attention mechanism and has been widely used in image processing tasks. Image steganography methods based on ViT are able to construct long-range dependencies, thereby identifying inconspicuous locations of hidden data, effectively improving the quality and security of the stegoimage. However, the self-attention mechanism of the ViT architecture suffers from quadratic time complexity, significantly reducing computational efficiency when processing long input sequences (such as high-resolution images), limiting its widespread application in practical applications.
[0006] To address ViT's computational efficiency issues, researchers have begun exploring new alternatives. State-space models (SSMs), due to their linear complexity, have emerged as a potential alternative to the Transformer and have garnered widespread attention in the research community. State-space models can capture long-term dependencies in sequential data at a low computational cost, offering new avenues for improving model performance and efficiency. In the field of information hiding, the introduction of state-space models is expected to further enhance the performance of steganography techniques, making them more efficient and accurate when processing complex image data.
[0007] The State Space Model (SSM) provides a new research direction for generative steganography. Derived from the Kalman filter, the SSM is a mathematical framework for describing linear time invariant (LTI) systems. Its core concept is to map the input signal in a dynamic system to the output signal through hidden states. In the continuous time domain, it can be expressed as:
[0008] Formula (1) is called the "state equation" and formula (2) is called the "output equation". , hidden state , output signal Tensor 、 、 、 is a learnable parameter. The entire system changes over time. t It is constantly evolving and can capture the dynamic characteristics of the input signal. However, the continuous-time model cannot be directly applied to deep learning and needs to be discretized to achieve numerical calculations. Discretization can transform continuous differential equations into difference equations. Zero-Order Hold (ZOH) technology is a commonly used discretization technology, which assumes that the input signal is in the sampling interval. If the inside remains constant, it can be approximated by the following formula:
[0009] After discretization, the state equation becomes:
[0010] Where, 、 is a tensor 、 The discretized result is is the identity matrix, is the index of the discrete time step, indicating the Sampling moments.
[0011] Traditional SSM still faces similar challenges as recurrent neural networks when processing long sequence data: hidden state The dimension is fixed, making it difficult to effectively capture long-term dependencies. At the same time, as the length of the input sequence increases, the gradient tends to disappear during backpropagation, making it difficult to learn long-distance correlations. The High-order Polynomial Projection Operator (HiPPO) was proposed to solve this problem. Its core idea is to compress the dynamic information of the historical input signal into a low-dimensional state space through high-dimensional polynomial projection, thereby explicitly modeling long-range dependencies. Specifically, HiPPO converts the historical input signal into a low-dimensional state space. Projecting the data onto a set of orthogonal polynomial basis functions, and dynamically adjusting the coefficients of the basis functions during training, allows the hidden state to optimally approximate the historical signal. Combining these techniques, the Structured State Space for Sequences model (also known as the S4 model) was formally proposed.
[0012] The LTI system stipulates that the parameters in SSM 、 、 、 The parameter values will not change with different inputs, that is, once the model is trained, they are completely fixed and cannot be changed during the reasoning process. This results in the S4 model’s inability to perform targeted reasoning on different inputs. The Selective State Space model (also known as the S6 model) solves this problem. 、 and time step Dynamically dependent on input Specifically, the S6 model designs three multi-layer perceptron networks 、 、 , to dynamically generate the corresponding parameters:
[0013] This design approach takes into account both the high efficiency of static parameters and the flexibility of dynamic parameters during inference, thereby improving the overall performance of the model. Summary of the Invention
[0014] The present invention aims to overcome the shortcomings of existing generative steganography based on traditional CNN in terms of generated image quality and resistance to steganalysis, and to provide an image steganography method, device, electronic device and storage medium. (In the existing technology, CNN-based methods have difficulty in effectively modeling images globally, resulting in poor visual quality of the generated secret images and easy detection by steganalysis. Although ViT overcomes the limitations of CNN in global feature modeling through the self-attention mechanism, it suffers from quadratic time complexity and low computational efficiency when processing high-resolution images.)
[0015] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0016] An image steganography method, comprising the following steps:
[0017] Obtain a dataset, including carrier images and secret images;
[0018] Construct a steganographic framework and input the carrier image, secret image, and carrier edge image generated by the edge extractor into the encoder. The encoder combines the features of the carrier edge image to hide the secret image in the carrier image and outputs the secret image.
[0019] Constructing the Res-SS2D module, which extends the traditional state-space model to the two-dimensional image domain through a four-way feature traversal and recombination mechanism, and realizes four-way global spatial modeling of the input image;
[0020] The edge extraction method is used to generate the carrier edge map. After adjusting its resolution and extracting features through the convolution layer, it is fused with the downsampling layers step by step, forcing the network to prioritize high-frequency sensitive areas.
[0021] Construct an encoder network to achieve efficient modeling of steganographic features through a multi-scale feature fusion mechanism, and embed a Res-SS2D module in the downsampling stage to enhance global context perception;
[0022] Build a decoder network to recover the original secret image based on the input secret image;
[0023] Design a multi-objective loss function to optimize the generation quality. At the same time, by introducing the L1 norm loss of the low-frequency component, constrain the consistency of the low-frequency region of the carrier image and the dense image.
[0024] Train the model and optimize the network parameters through the back-propagation algorithm until the model converges;
[0025] The carrier image and the secret image are input into the trained model, and the encoder generates the secret image; the secret image is input into the decoder to restore the original secret image.
[0026] Furthermore, the construction of the steganographic framework includes:
[0027] Let the carrier image be , the secret image is denoted as , the carrier edge image obtained by the carrier image through the edge extractor is recorded as ;
[0028] The carrier image, secret image and carrier edge image are input into the encoder together. The encoder combines the features of the carrier edge image to hide the secret image into the carrier image and outputs the secret image, which is recorded as ;
[0029] In the decoding stage, the secret image is input to the decoder, and the decoder recovers the original secret image from it, which is denoted as .
[0030] Furthermore, the construction of the Res-SS2D module includes:
[0031] The input feature map is traversed in parallel along four different directions to generate four sets of feature sequences with different spatial correlations, providing multi-perspective feature input for subsequent state space calculations;
[0032] Four independent S6 modules are used to process each direction sequence respectively. Each S6 module adaptively integrates the local features of the current scanning path and the global dependency of the stride length through a dynamic weight selection mechanism.
[0033] Remap the four sets of output sequences to two-dimensional space according to the inverse scanning path, and finally restore the feature map of the same size as the input;
[0034] Among them, the Res-SS2D module integrates depthwise separable convolution, performs channel-by-channel convolution on the input feature map, and uses several 1×1 convolution kernels to extract features between channels; the Res-SS2D module also introduces residual connections to avoid the gradient vanishing problem and enhance global feature fusion.
[0035] Furthermore, the Laplace edge extraction operator is introduced to enhance the encoder's ability to perceive high-frequency features;
[0036] The four-neighborhood Laplace operator is used to perform convolution operation on the carrier image, and the grayscale mutation intensity of the pixel and the neighborhood is calculated by second-order differential. The formula is as follows:
[0037] Where, and Respectively represent the horizontal and vertical directions in the image coordinate system, Represents a two-dimensional image In spatial coordinates The pixel intensity value at ;
[0038] The extracted edge map is fused with the original image and then input into the encoder, forcing the neural network to prioritize learning high-frequency area features so as to hide the secret information in the high-frequency area of the carrier image.
[0039] Furthermore, the encoder network is improved based on the 7-layer U-Net architecture, and the efficient modeling of steganographic features is achieved through a multi-scale feature fusion mechanism;
[0040] In the downsampling stage, except for the last layer, the convolutional blocks of each layer are embedded with Res-SS2D modules. The cross-scanning and state-space modeling capabilities of the Res-SS2D modules are used to enhance global context perception while preserving local details through residual connections.
[0041] In the upsampling stage, a pure convolutional structure is used to cascade the multi-scale features output by each downsampling layer with the corresponding upsampling layer through skip connections to achieve cross-level feature complementarity and suppress information loss;
[0042] At the same time, the edge extraction method is used to generate the carrier edge map. After adjusting its resolution and extracting features through a convolutional layer, it is fused with the downsampling layers in an additive manner step by step, forcing the network to focus on high-frequency sensitive areas first.
[0043] Furthermore, an edge extraction method is used to generate a carrier edge map. After adjusting its resolution and extracting features through a convolutional layer, it is fused with the downsampling layers in an additive manner, forcing the network to prioritize high-frequency sensitive areas.
[0044] The input layer concatenates the carrier image and the secret image along the channel dimension, and then compresses the feature map size to 1 / 32 of the original image through five levels of downsampling, and the number of channels increases step by step to 512.
[0045] Furthermore, the loss function uses mean square error to calculate pixel differences and introduces a multi-scale structural similarity index to improve subjective visual quality;
[0046] The hidden loss is expressed as:
[0047] In the above formula, Represents the number of carrier images and secret images input in a batch, Indicates that the current calculation is in progress. The loss of carrier image and encrypted image, represents the input carrier image, represents the encrypted image generated by the encoder, and Represent the weights of MSE loss and MS-SSIM loss respectively;
[0048] The recovery loss is expressed as:
[0049] In the above formula, represents a secret image, represents the restored image generated by the decoder;
[0050] Introducing low-frequency wavelet loss into the loss function, the image is decomposed into high-frequency and low-frequency wavelet sub-bands through DWT transformation, and the low-frequency components of the carrier image and the dense image are subjected to Loss, so that the secret image is partially hidden in the high-frequency components of the carrier image, while retaining part of the secret information in the low-frequency components, specifically expressed as:
[0051] Where, Represents the weight of the low-frequency wavelet loss. Specifically, the DWT transform used uses the Haar wavelet basis for a layer of decomposition and uses a zero-filling mode to process the image boundary, which meets the requirements of low-frequency feature extraction while ensuring high computational efficiency. Finally, the low-frequency approximate coefficient subband after the decomposition is separated and extracted as the low-frequency feature representation of the image;
[0052] The total training loss is defined as:
[0053] in, is the low-frequency wavelet loss, To recover losses, To hide losses.
[0054] An image steganography device, comprising:
[0055] A data set acquisition module, used to obtain carrier images and secret images;
[0056] An edge extraction module, for generating an edge image of a carrier image;
[0057] The Res-SS2D module is used to extend the traditional state-space model to the two-dimensional image domain through a four-way feature traversal and recombination mechanism, achieving four-way global spatial modeling of the input image;
[0058] The edge feature fusion module is used to adjust the resolution of the carrier edge map and extract features through the convolution layer, and then fuse them with the downsampling layers step by step, forcing the network to prioritize high-frequency sensitive areas;
[0059] The encoder module is based on an improved 7-layer U-Net architecture. It is used to efficiently model steganographic features through a multi-scale feature fusion mechanism and embeds a Res-SS2D module in the downsampling stage to enhance global context awareness.
[0060] The decoder module uses a 7-layer U-Net structure to input the secret image and restore the original secret image through upsampling;
[0061] The loss function module is used to design a multi-objective loss function, combining PSNR and MS-SSIM to optimize the generated quality. At the same time, by introducing the L1 norm loss of the low-frequency component, it constrains the consistency of the low-frequency regions of the carrier image and the dense image.
[0062] The training module is used to optimize the network parameters through the back-propagation algorithm until the model converges;
[0063] The steganography execution module is used to input the carrier image and the secret image into the trained model, generate the secret image through the encoder, and input the secret image into the decoder to restore the original secret image.
[0064] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the aforementioned image steganography method when executing the computer program.
[0065] A computer-readable storage medium stores a computer program executable by an electronic device. When the computer program runs on the electronic device, the electronic device executes the steps of the aforementioned image steganography method.
[0066] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0067] This invention provides an image steganography method, apparatus, electronic device, and storage medium. By designing a Res-SS2D module, this method implements four-way global spatial modeling of the input image, improving the visual quality of the secret image while maintaining linear computational complexity. Furthermore, an edge feature fusion mechanism is introduced to extract high-frequency information from the carrier image and integrate it into each layer of the hidden network, guiding the secret image to preferentially embed in high-frequency regions, reducing the likelihood of detection. A multi-objective loss function is also designed to optimize generation quality. Experiments demonstrate that this invention demonstrates excellent performance in both generating and recovering secret images, with high resistance to steganalysis and low inference time, ensuring efficiency and practicality in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0069] Figure 1 This is an overall framework diagram of the image steganography method provided in one embodiment of the present application;
[0070] Figure 2 This is a schematic diagram of the workflow of the SS2D (Spatial-Separable 2D Convolution) module provided in one embodiment of the present application;
[0071] Figure 3 2 is a schematic diagram of the structure of the Res-SS2D (Residual Spatial-Separable 2D Convolution) module provided in one embodiment of the present application;
[0072] Figure 4 In the figure, (a) is the 4-neighborhood Laplacian operator; (b) is an example of edge feature extraction effect;
[0073] Figure 5This is a diagram of the encoder network structure of the SSEU-Net (Spatial-Separable Efficient U-Net) provided in one embodiment of the present application;
[0074] Figure 6 This is a decoder network structure diagram of the SSEU-Net (Spatial-Separable Efficient U-Net) provided in one embodiment of the present application;
[0075] Figure 7 This is a demonstration of the effect of the model proposed in this application on the ImageNet dataset;
[0076] Figure 8 This is a demonstration of the effect of the model proposed in this application on the COCO dataset.
[0077] in, Figure 3 In the figure, Linear is a linear layer, DWConv is Depthwise Convolution, which is depth-separable convolution; SiLU is Sigmoid Linear Unit, which is an activation function; SS2D is Spatial-Separable 2D Convolution, which is spatially separable 2D convolution; LayerNorm is Layer Normalization, which is layer normalization.
[0078] Figure 5-6In
[15] , Edge Extraction is edge extraction; Resize is image scaling; Conv is Convolution, i.e., convolution; Concatenation is splicing; CNR (Convolution + BatchNorm + ReLU) is a combination of convolution modules, where Convolution (convolution), BatchNorm (Batch Normalization, batch normalization), ReLU (Rectified Linear Unit, activation function); Res-SS2D (Residual Spatial-Separable 2D Convolution, residual spatially separable 2D convolution) is an efficient convolution module, where Residual (residual connection), Spatial-Separable 2D Convolution (spatial separable 2D convolution); TNR (ConvTranspose + BatchNorm + ReLU) is an upsampling module combination, where ConvTranspose (transposed convolution, also known as deconvolution or fractional stride convolution), BatchNorm (batch normalization), ReLU (activation function); CR (Convolution + ReLU) is a combination of convolution modules, including Convolution (convolution) and ReLU (activation function); TS is a module containing ConvTranspose (transposed convolution) and Sigmoid activation function; Skip Connection is a jump connection. DETAILED DESCRIPTION
[0079] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein.
[0080] Example 1:
[0081] like Figure 1-8 As shown, the present invention provides a technical solution:
[0082] An image steganography method based on an improved selective state space model comprises the following steps:
[0083] Step 1: Get the dataset.
[0084] Step 2: Configure the model training environment.
[0085] Step 3: Construct a steganographic framework. The present invention records the carrier image as The secret image is denoted as , the carrier edge image obtained by the carrier image through the edge extractor is recorded as The carrier image, secret image and carrier edge image are input into the encoder together. The encoder can combine the characteristics of the carrier edge image to hide the secret image into the carrier image and output the secret image, which is recorded as In the decoding stage, the secret image is input to the decoder, and the decoder can recover the original secret image from it, which is recorded as .
[0086] Step 4: Construct the Res-SS2D module. The classic selective state-space model demonstrates significant advantages in sequence modeling, but its one-way propagation mechanism limits causal effects. This effect manifests itself as a feature update that only allows historical information to participate in the current state calculation, while prohibiting forward feedback of information from future time steps. In image generation tasks, since spatial relationships between pixels must simultaneously consider multi-directional contextual associations (such as edge continuity and texture consistency), this strict temporal causal constraint can disrupt the global perception of two-dimensional space, resulting in structural discontinuities or semantic distortion in the generated images.
[0087] To address these issues, the SSEU-Net model proposed in this paper introduces the Res-SS2D module, a selective state-space computation unit optimized for image generation. The core of this module lies in the SS2D module, which extends the traditional state-space model to the two-dimensional image domain through a four-way feature traversal and recombination mechanism. The specific process is as follows:
[0088] Cross-scanning: The input feature map is traversed in parallel along four different directions to generate four sets of feature sequences with different spatial correlations. This operation is equivalent to constructing four complementary pseudo-time axes, providing multi-perspective feature input for subsequent state space calculations.
[0089] S6 module feature extraction: Four independent S6 modules are used to process each direction sequence. Each S6 module uses a dynamic weight selection mechanism to adaptively fuse the local features of the current scan path and the global dependency across the stride length.
[0090] Cross-merging: The four sets of output sequences are remapped to two-dimensional space according to the inverse scanning path, and finally the feature map of the same size as the input is restored.
[0091] Building on this foundation, to further balance model efficiency and feature expression, the Res-SS2D module integrates depthwise separable convolution (DWConv). This first performs channel-by-channel convolution on the input feature map, followed by cross-channel feature extraction using multiple 1×1 convolution kernels. DWConv decomposes the standard convolution into a cascade of spatial filtering and channel mapping, significantly reducing computational complexity while maintaining an equivalent receptive field. Finally, the Res-SS2D module introduces residual connections, which not only effectively avoid the vanishing gradient problem but also enhance global feature fusion.
[0092] Step 5: Construct an edge extraction method. Traditional pixel domain steganography methods are prone to texture copying artifacts and chromatic distortion because they directly modify the pixel values in the spatial domain. Compared with pixel domain embedding, the research field generally believes that high-frequency sensitive areas (such as edge contours and texture details) have better security advantages in information hiding. Based on this, SSEU-Net proposes an edge-guided steganography method, which enhances the encoder's perception of high-frequency features by introducing the Laplace edge extraction operator. Specifically, a 4-neighborhood Laplace operator is used to perform a convolution operation on the carrier image. The operator calculates the grayscale mutation intensity of the pixel point and the neighborhood through second-order differentials. Its mathematical essence is equivalent to solving the curvature change of the two-dimensional function of the image, which can be expressed as:
[0093] Where, and Respectively represent the horizontal and vertical directions in the image coordinate system, Represents a two-dimensional image In spatial coordinates The pixel intensity value at .
[0094] Finally, the extracted edge map is fused with the original image and input into the encoder, forcing the neural network to preferentially learn high-frequency area features so as to hide the secret information preferentially in the high-frequency area of the carrier image.
[0095] Step 6: Construct the encoder network. The encoder network proposed in this paper is based on the improved 7-layer U-Net architecture and achieves efficient modeling of steganographic features through a multi-scale feature fusion mechanism.
[0096] During the downsampling phase, except for the last layer, the convolutional blocks of each layer are embedded with a Res-SS2D module, leveraging its cross-scanning and state-space modeling capabilities to enhance global context perception while preserving local details through residual connections. During the upsampling phase, a purely convolutional structure is employed, cascading the multi-scale features output by each downsampling layer with the corresponding upsampling layer via skip connections. This achieves cross-level feature complementarity and effectively suppresses information loss. Furthermore, an edge extraction method is used to generate a carrier edge map. After adjusting its resolution and extracting features through a convolutional layer, it is then additively fused with each downsampling layer, forcing the network to prioritize high-frequency sensitive areas.
[0097] To enhance edge features, the edge extraction method described in step 5 is first used to generate a carrier edge map. After adjusting its resolution and extracting features through a convolutional layer, it is then additively fused with each downsampling layer, forcing the network to prioritize high-frequency sensitive areas. The input layer concatenates the carrier image and the secret image along the channel dimension. Subsequently, through five levels of downsampling, the feature map size is gradually compressed to 1 / 32 of the original image, with the number of channels gradually increasing to 512. This progressive channel expansion strategy effectively improves feature expression capabilities. The overall architecture balances global steganalytic feature modeling, edge-guided optimization, and the synergistic improvement of computational efficiency.
[0098] Step 7: Build the decoder network. To balance model training and inference efficiency, the decoder network uses a basic 7-layer U-Net structure. Unlike the encoder network, the decoder network input is only the encrypted image, so the number of input channels is 3. During the downsampling phase, the number of channels and feature map size changes are the same as those of the encoder, so they are not detailed here.
[0099] Step 8: Construct a multi-objective loss function. To ensure the visual quality of the generated encrypted image and the restored image, the loss function used in this model uses mean squared error (MSE) to calculate pixel differences, ensuring statistically high quality of the generated image. The multi-scale structural similarity index measure (MS-SSIM) is also introduced to improve subjective visual quality. MSE calculates the mean squared difference between corresponding pixels in two images, achieving pixel-level error calculation. A smaller MSE value indicates a closer match between the generated image and the original image. However, over-reliance on MSE can lead to loss of high-frequency details. MS-SSIM simulates the multi-resolution perception characteristics of the human visual system. It calculates brightness, contrast, and structural similarity at multiple scales, ultimately weighting them to produce a comprehensive score between 0 and 1 (a value closer to 1 indicates higher visual consistency and, therefore, higher quality). Compared to the single-scale SSIM, MS-SSIM is more sensitive to texture continuity and edge sharpness, effectively suppressing blocking artifacts and local distortion, resulting in higher subjective visual quality of the generated image. The two methods are used together to form a complementary optimization mechanism. The hidden loss is expressed as:
[0100] Where, Represents the number of carrier images and secret images input in a batch, Indicates that the current calculation is in progress. The loss of carrier image and encrypted image, represents the input carrier image, represents the encrypted image generated by the encoder, and Represent the weights of MSE loss and MS-SSIM loss respectively.
[0101] The form of the recovery loss is consistent with the hidden loss and is expressed as:
[0102] Where, represents a secret image, represents the restored image generated by the decoder.
[0103] In order to make better use of the edge feature fusion method, the present invention introduces low-frequency wavelet loss into the loss function, that is, the image is decomposed into high-frequency and low-frequency wavelet sub-bands through DWT transformation. Since the information hidden in the high-frequency component is less likely to be detected than the information hidden in the low-frequency component, this paper performs a low-frequency loss on the low-frequency components of the carrier image and the encrypted image. Loss, thus forcing the secret image to be hidden as much as possible in the high-frequency components of the carrier image, while retaining very little secret information in the low-frequency components, specifically expressed as:
[0104] Where, Represents the weight of the low-frequency wavelet loss. Specifically, the DWT transform uses a Haar wavelet basis for a one-layer decomposition and uses a zero-padding pattern to process image boundaries. This satisfies the requirements for low-frequency feature extraction while ensuring high computational efficiency. Finally, the low-frequency approximate coefficient subbands after the decomposition are separated and extracted as the low-frequency feature representation of the image.
[0105] The total training loss can then be defined as:
[0106] Step 9: Use the dataset selected in step 1 and train the model in the training environment set in step 2. After the model training is completed, use the test set in the dataset to verify the generalization performance of the model.
[0107] Step 10: Ablation Experiment. To systematically evaluate the contributions of each core component in our model, we constructed a unified experimental benchmark based on the dataset selected in Step 1. We designed four sets of progressive ablation experimental models, focusing on quantitative analysis of the Res-SS2D module and the edge feature fusion method. The configurations of the four models are as follows:
[0108] Model 1: U-Net + Res-SS2D + edge feature fusion. This is the complete solution proposed in this article.
[0109] Model 2: U-Net+Res-SS2D. The edge feature fusion method is removed.
[0110] Model 3: U-Net + edge feature fusion. The Res-SS2D module is removed.
[0111] Model 4: Only the most basic U-Net is retained.
[0112] The above ablation experiment design can effectively illustrate the independent contribution of each core module to the overall system performance, providing a rigorous experimental basis for the effectiveness of each module. The two indicators used in the experiment, PSNR and SSIM, are introduced as follows:
[0113] (1) Peak Signal-to-Noise Ratio (PSNR): It is a common image quality evaluation criterion used to measure the similarity between two images. PSNR is calculated based on the mean square error (MSE). For two images of the same size, and :
[0114] in, The total number of image pixels, and Respectively and Middle pixel values. So the PSNR value can be expressed as:
[0115] in, Indicates the maximum pixel value that the input image can take. For example, for a common 8-bit depth image, PSNR is measured in decibels (dB), representing the signal-to-noise ratio. A higher PSNR indicates a smaller difference between the two images. Generally, a PSNR of over 40dB indicates high visual quality, while a PSNR below 30dB indicates poor visual quality.
[0116] (2) Structural Similarity Index (SSIM): This is also an indicator that measures the similarity between two images. It takes into account three factors: brightness, contrast, and structure. The design of SSIM takes into account the human eye's perception of image quality, so it can more accurately reflect the human eye's perception of image changes.
[0117] The SSIM value is between 0 and 1. The closer it is to 1, the more similar the two images are; the closer it is to 0, the less similar the two images are. and :
[0118] in, 、 Represents images respectively and The mean of 、 Do not represent images and The standard deviation of represents the covariance of the two images. and The introduction of is to avoid division by 0 errors.
[0119] Regarding the anti-steganalysis performance, the present invention selected three currently mainstream steganalyzers for experiments, including XuNet, YeNet, and SRNet. It should be noted that when the detection accuracy of the steganalyzer is close to 50%, it means that the detection behavior of the analyzer is no different from random guessing, indicating that the generated image can already deceive the analyzer well, that is, it has strong anti-steganalysis performance.
[0120] This application has at least the following beneficial effects:
[0121] 1. High-quality encrypted images. This paper uses the Res-SS2D module to perform four-way global spatial modeling of images, combined with edge feature fusion methods. This results in a generated encrypted image that is highly similar to the original carrier image in visual effect. Whether in overall structure, texture details, or color expression, the difference is difficult to detect with the naked eye, greatly improving the visual effect of the generated image.
[0122] 2. Strong anti-steganalysis capability. The present invention embeds secret information in the high-frequency areas of the image and constrains the low-frequency components, making the statistical characteristics of the secret image closer to natural images, effectively reducing the possibility of detection by steganalysis methods. When facing a variety of mainstream steganalysis tools, such as XuNet, YeNet, and SRNet, the present invention has demonstrated strong anti-detection capabilities, with detection accuracy rates below 61%, close to the level of random guessing, fully demonstrating its advantages in anti-steganalysis and ensuring the security of secret information during communication.
[0123] Third, the model is highly efficient. This invention utilizes a selective state-space model with linear time complexity. Compared to traditional methods, it significantly reduces computational effort and time costs when processing high-resolution images, thereby improving the model's operational efficiency. Compared to similar advanced ViT-based methods, this invention significantly reduces inference time while maintaining high performance.
[0124] Fourth, strong generalization capability. Experimental validation on two large-scale datasets, COCO and ImageNet, demonstrates that the proposed method maintains good performance across images of varying styles and scenes, demonstrating strong generalization capability. This demonstrates the method's adaptability to image diversity and its broad applicability to various image steganography tasks, regardless of specific image types or styles. It can also be further expanded to other areas requiring privacy and information security, such as privacy protection, secret communications, and digital copyright protection, providing a more advanced and reliable information hiding technology solution for these fields.
[0125] Example 2:
[0126] The following is an example of a specific process of an image steganography method based on Example 1:
[0127] Step 1: Get the dataset.
[0128] Specifically, in this embodiment, two major datasets, ImageNet and COCO, are used. Each dataset is divided into 10,000 images as a training set and 2,500 images as a validation set according to the carrier image and secret image, and 500 images are retained as a test set to evaluate the model generalization ability.
[0129] Step 2: Configure the model training environment.
[0130] Specifically, this example is implemented based on the PyTorch 2.0 framework. The experimental platform's CPU is a 12th Gen Intel Core i5-12600KF, the GPU is an NVIDIA GeForce RTX 4060Ti, and the operating system is Ubuntu 24.04 LTS. During model training, the input image resolution is uniformly adjusted to 256×256, the batch size is set to 16, and the training cycle (epoch) is 400 rounds. The optimizer uses the Adam algorithm, and its hyperparameters are set to 、 The initial value of the learning rate is 0.001, which is dynamically adjusted by the cosine annealing algorithm and decays to 0.00001 at the end of training according to the cosine function.
[0131] Step 3: Build a steganographic framework, such as Figure 1 As shown. The present invention records the carrier image as , the secret image is denoted as , the carrier edge image obtained by the carrier image through the edge extractor is recorded as .
[0132] Specifically, in this embodiment, the carrier image, the secret image, and the carrier edge image are input into the encoder together. The encoder can combine the characteristics of the carrier edge image to hide the secret image into the carrier image and output the secret image, which is recorded as In the decoding stage, the secret image is input to the decoder, and the decoder can recover the original secret image from it, which is recorded as .
[0133] Step 4: Construct the Res-SS2D module. The classic selective state-space model demonstrates significant advantages in sequence modeling, but its one-way propagation mechanism limits causal effects. This effect manifests itself as a feature update that only allows historical information to participate in the current state calculation, while prohibiting forward feedback of information from future time steps. In image generation tasks, since spatial relationships between pixels must simultaneously consider multi-directional contextual associations (such as edge continuity and texture consistency), this strict temporal causal constraint can disrupt the global perception of two-dimensional space, resulting in structural discontinuities or semantic distortion in the generated images.
[0134] To address these issues, the SSEU-Net model proposed in this paper introduces the Res-SS2D module, a selective state-space computation unit optimized for image generation. The core of this module lies in the SS2D module, which extends the traditional state-space model to the two-dimensional image domain through a four-way feature traversal and recombination mechanism. The specific process is as follows:
[0135] Cross-scanning: The input feature map is traversed in parallel along four different directions to generate four sets of feature sequences with different spatial correlations. This operation is equivalent to constructing four complementary pseudo-time axes, providing multi-perspective feature input for subsequent state space calculations.
[0136] S6 module feature extraction: Four independent S6 modules are used to process each direction sequence. Each S6 module uses a dynamic weight selection mechanism to adaptively fuse the local features of the current scan path and the global dependency across the stride length.
[0137] Cross-merging: The four sets of output sequences are remapped to two-dimensional space according to the inverse scanning path, and finally the feature map of the same size as the input is restored.
[0138] Specifically, in the embodiment, the workflow of the SS2D module is as follows: Figure 2 As shown. On this basis, in order to further balance the model efficiency and feature expression capabilities, the Res-SS2D module integrates the depthwise separable convolution (DWConv), which first performs channel-by-channel convolution on the input feature map, and then uses several 1×1 convolution kernels to extract features between channels. DWConv decomposes the standard convolution into a cascade operation of spatial filtering and channel mapping, which greatly reduces the computational complexity while maintaining the equivalent receptive field. Finally, the Res-SS2D module also introduces residual connections, which not only effectively avoids the gradient vanishing problem, but also enhances global feature fusion. The structure of the Res-SS2D module is shown as follows. Figure 3 shown.
[0139] Step 5: Construct an edge extraction method. Traditional pixel domain steganography methods are prone to texture copying artifacts and chromatic distortion because they directly modify the pixel values in the spatial domain. Compared with pixel domain embedding, the research field generally believes that high-frequency sensitive areas (such as edge contours and texture details) have better security advantages in information hiding. Based on this, SSEU-Net proposes an edge-guided steganography method, which enhances the encoder's perception of high-frequency features by introducing the Laplace edge extraction operator. Specifically, a 4-neighborhood Laplace operator is used to perform a convolution operation on the carrier image. The operator calculates the grayscale mutation intensity of the pixel point and the neighborhood through second-order differentials. Its mathematical essence is equivalent to solving the curvature change of the two-dimensional function of the image, which can be expressed as:
[0140] Where, and Respectively represent the horizontal and vertical directions in the image coordinate system, Represents a two-dimensional image In spatial coordinates The pixel intensity value at .
[0141] Finally, the extracted edge map is fused with the original image and input into the encoder, forcing the neural network to preferentially learn high-frequency area features so as to hide the secret information preferentially in the high-frequency area of the carrier image.
[0142] Specifically, in the embodiment, a 4-neighborhood Laplacian operator is used, such as Figure 4 As shown in part (a), the edge feature extraction effect is shown in Figure 4 As shown in part (b).
[0143] Step 6: Construct the encoder network. As shown in Figure 5, the encoder network proposed in this paper is based on an improved 7-layer U-Net architecture and achieves efficient modeling of steganographic features through a multi-scale feature fusion mechanism.
[0144] Specifically, in the embodiment, in the downsampling stage, except for the last layer, the convolution block of each layer is embedded with a Res-SS2D module, and its cross-scanning and state space modeling capabilities are used to enhance global context perception, while retaining local details through residual connections. In the upsampling stage, a pure convolutional structure is adopted, and the multi-scale features output by each downsampling layer are cascaded with the corresponding upsampling layer through jump connections to achieve cross-level feature complementarity and effectively suppress information loss. At the same time, the edge extraction method is used to generate a carrier edge map. After adjusting its resolution and extracting features through a convolution layer, it is fused with each downsampling layer in an additive manner step by step, forcing the network to focus on high-frequency sensitive areas first.
[0145] To enhance edge features, the edge extraction method described in step 5 is first used to generate a carrier edge map. After adjusting its resolution and extracting features through a convolutional layer, it is then additively fused with each downsampling layer, forcing the network to prioritize high-frequency sensitive areas. The input layer concatenates the carrier image and the secret image along the channel dimension. Subsequently, through five levels of downsampling, the feature map size is gradually compressed to 1 / 32 of the original image, with the number of channels gradually increasing to 512. This progressive channel expansion strategy effectively improves feature expression capabilities. The overall architecture balances global steganalytic feature modeling, edge-guided optimization, and the synergistic improvement of computational efficiency.
[0146] Step 7: Build the decoder network. Figure 6 As shown in the figure, to balance model training and inference efficiency, the decoder network uses only a basic 7-layer U-Net structure. Unlike the encoder network, the decoder network input is only the encrypted image, so the number of input channels is 3. During the downsampling stage, the number of channels and feature map size changes are consistent with the encoder, so they are not detailed here.
[0147] Step 8: Construct a multi-objective loss function. To ensure the visual quality of the generated encrypted image and the restored image, the loss function used in this model uses mean squared error (MSE) to calculate pixel differences, ensuring statistically high quality of the generated image. The multi-scale structural similarity index measure (MS-SSIM) is also introduced to improve subjective visual quality. MSE calculates the mean squared difference between corresponding pixels in two images, achieving pixel-level error calculation. A smaller MSE value indicates a closer match between the generated image and the original image. However, over-reliance on MSE can lead to loss of high-frequency details. MS-SSIM simulates the multi-resolution perception characteristics of the human visual system. It calculates brightness, contrast, and structural similarity at multiple scales, ultimately weighting them to produce a comprehensive score between 0 and 1 (a value closer to 1 indicates higher visual consistency and, therefore, higher quality). Compared to the single-scale SSIM, MS-SSIM is more sensitive to texture continuity and edge sharpness, effectively suppressing blocking artifacts and local distortion, resulting in higher subjective visual quality of the generated image. The two methods are used together to form a complementary optimization mechanism. The hidden loss is expressed as:
[0148] Where, Represents the number of carrier images and secret images input in a batch, Indicates that the current calculation is in progress. The loss of carrier image and encrypted image, represents the input carrier image, represents the encrypted image generated by the encoder, and Represent the weights of MSE loss and MS-SSIM loss respectively.
[0149] The form of the recovery loss is consistent with the hidden loss and is expressed as:
[0150] Where, represents a secret image, represents the restored image generated by the decoder.
[0151] In order to make better use of the edge feature fusion method, the present invention introduces low-frequency wavelet loss into the loss function, that is, the image is decomposed into high-frequency and low-frequency wavelet sub-bands through DWT transformation. Since the information hidden in the high-frequency component is less likely to be detected than the information hidden in the low-frequency component, this paper performs a low-frequency loss on the low-frequency components of the carrier image and the encrypted image. Loss, thus forcing the secret image to be hidden as much as possible in the high-frequency components of the carrier image, while retaining very little secret information in the low-frequency components, specifically expressed as:
[0152] Where, Represents the weight of the low-frequency wavelet loss. Specifically, the DWT transform uses a Haar wavelet basis for a one-layer decomposition and uses a zero-padding pattern to process image boundaries. This satisfies the requirements for low-frequency feature extraction while ensuring high computational efficiency. Finally, the low-frequency approximate coefficient subbands after the decomposition are separated and extracted as the low-frequency feature representation of the image.
[0153] The total training loss can then be defined as:
[0154] Specifically, in the embodiment, .
[0155] Step 9: Use the dataset selected in step 1 and train the model in the training environment set in step 2. After the model training is completed, use the test set in the dataset to verify the generalization performance of the model.
[0156] Specifically, in the embodiment, in order to perform subjective quality analysis on the images generated by the model, this paper randomly extracts 8 groups of images from the test sets of ImageNet and COCO datasets for display, such as Figure 7 and Figure 8 As shown in the figure, the first row is the carrier image, the second row is the encrypted image generated by the encoder, the third row is the secret image, and the fourth row is the restored image generated by the decoder.
[0157] It can be seen that the discrepancy between the encrypted and restored images generated by the proposed model is difficult to detect with the naked eye. Whether in terms of overall image structure, texture detail, or color rendering, the generated images are highly similar to the originals, almost to the point of being indistinguishable from the real thing. This demonstrates that our method is not only effective for specific image types, but also maintains good performance across images of varying styles and scenes. It demonstrates strong generalization and is suitable for a wide range of image processing scenarios.
[0158] Step 10: Ablation Experiment. To systematically evaluate the contributions of each core component of our model, we constructed a unified experimental benchmark based on the ImageNet dataset, using 10,000 images as the training set, 2,500 as the validation set, and retaining 500 independent test sets for final performance verification. By designing four sets of progressive ablation experimental models, we focused on quantitatively analyzing the Res-SS2D module and the edge feature fusion method. The configurations of the four sets of models are as follows:
[0159] Model 1: U-Net + Res-SS2D + edge feature fusion. This is the complete solution proposed in this article.
[0160] Model 2: U-Net+Res-SS2D. The edge feature fusion method is removed.
[0161] Model 3: U-Net + edge feature fusion. The Res-SS2D module is removed.
[0162] Model 4: Only the most basic U-Net is retained.
[0163] The above ablation experiment design can effectively illustrate the independent contribution of each core module to the overall system performance, providing a rigorous experimental basis for the effectiveness of each module. The two indicators used in the experiment, PSNR and SSIM, are introduced as follows:
[0164] (1) Peak Signal-to-Noise Ratio (PSNR): It is a common image quality evaluation criterion used to measure the similarity between two images. PSNR is calculated based on the mean square error (MSE). For two images of the same size, and :
[0165] in, is the total number of image pixels, and Respectively and Middle pixel values. So the PSNR value can be expressed as:
[0166] in, Indicates the maximum pixel value that the input image can take. For example, for a common 8-bit depth image, PSNR is measured in decibels (dB), representing the signal-to-noise ratio. A higher PSNR indicates a smaller difference between the two images. Generally, a PSNR of over 40dB indicates high visual quality, while a PSNR below 30dB indicates poor visual quality.
[0167] (2) Structural Similarity Index (SSIM): This is also an indicator that measures the similarity between two images. It takes into account three factors: brightness, contrast, and structure. The design of SSIM takes into account the human eye's perception of image quality, so it can more accurately reflect the human eye's perception of image changes.
[0168] The SSIM value is between 0 and 1. The closer it is to 1, the more similar the two images are; the closer it is to 0, the less similar the two images are. and :
[0169] in, 、 Represents images respectively and The mean of 、 Represents images respectively and The standard deviation of represents the covariance of the two images. and The introduction of is to avoid division by 0 errors.
[0170] Regarding the anti-steganalysis performance, the present invention selected three currently mainstream steganalyzers for experiments, including XuNet, YeNet, and SRNet. It should be noted that when the detection accuracy of the steganalyzer is close to 50%, it means that the detection behavior of the analyzer is no different from random guessing, indicating that the generated image can already deceive the analyzer well, that is, it has strong anti-steganalysis performance.
[0171] The ablation experiment results of image generation quality are shown in Table 1, and the ablation experiment results of steganalysis performance are shown in Table 2.
[0172] Table 1. Ablation experiment on the impact of each core component on image generation quality
[0173]
[0174] Table 2. Ablation experiment on the impact of each core component on anti-steganalysis performance
[0175]
[0176] The bold data in the table represent the optimal values for that evaluation metric. As can be seen in Table 1, removing any core component will affect the quality of the generated image, but the degree of impact varies. Comparing Models 1 and 2, when edge feature fusion is removed, the PSNR of the encrypted and restored images decreases by 1.46dB and 2.058dB, respectively. Comparing Models 1 and 3, when Res-SS2D is removed, the PSNR decreases by 3.464dB and 4.973dB, respectively. Clearly, edge feature fusion has an impact on the quality of the generated image, but the impact is not as significant as that of the Res-SS2D module, highlighting the importance of global modeling capabilities for the quality of generated images.
[0177] Table 2 shows a change in the situation. Taking SR-Net as an example, comparing Models 1 and 2, the judgment accuracy increased by 16.01%, while comparing Models 1 and 3, the judgment accuracy increased by 10.78%. Further comparing Models 3 and 4, adding edge feature fusion to the basic U-Net reduced the judgment accuracy by 9.12%. In contrast, comparing Models 2 and 4, adding the Res-SS2D module to the basic U-Net reduced the judgment accuracy by only 3.89%. This demonstrates that edge feature fusion has a greater impact on anti-steganalysis performance than Res-SS2D.
[0178] In summary, the two core modules proposed in this paper, Res-SS2D and edge feature fusion, each play a unique role in the quality of generated images and anti-steganalysis performance, and the two complement each other to jointly improve the overall performance of the entire model.
[0179] Example 3:
[0180] The present invention provides a technical solution:
[0181] An image steganography device, comprising:
[0182] A data set acquisition module, used to obtain carrier images and secret images;
[0183] An edge extraction module, for generating an edge image of a carrier image;
[0184] The Res-SS2D module is used to extend the traditional state-space model to the two-dimensional image domain through a four-way feature traversal and recombination mechanism, achieving four-way global spatial modeling of the input image;
[0185] The edge feature fusion module is used to adjust the resolution of the carrier edge map and extract features through the convolution layer, and then fuse them with the downsampling layers step by step, forcing the network to prioritize high-frequency sensitive areas;
[0186] The encoder module is based on an improved 7-layer U-Net architecture. It is used to efficiently model steganographic features through a multi-scale feature fusion mechanism and embeds a Res-SS2D module in the downsampling stage to enhance global context awareness.
[0187] The decoder module uses a 7-layer U-Net structure to input the secret image and restore the original secret image through upsampling;
[0188] The loss function module is used to design a multi-objective loss function, combining PSNR and MS-SSIM to optimize the generated quality. At the same time, by introducing the L1 norm loss of the low-frequency component, it constrains the consistency of the low-frequency regions of the carrier image and the dense image.
[0189] The training module is used to optimize the network parameters through the back-propagation algorithm until the model converges;
[0190] The steganography execution module is used to input the carrier image and the secret image into the trained model, generate the secret image through the encoder, and input the secret image into the decoder to restore the original secret image.
[0191] 1. For the Res-SS2D module:
[0192] Classic selective state-space models demonstrate significant advantages in sequence modeling, but their unidirectional propagation mechanism limits their causal effects. This effect manifests itself as feature updates that only allow historical information to participate in the current state calculation, while prohibiting the forward feed of information from future time steps. In image generation tasks, since spatial relationships between pixels must simultaneously consider multi-directional contextual associations (such as edge continuity and texture consistency), this strict temporal causal constraint can disrupt the global perception of two-dimensional space, resulting in structural discontinuities or semantic distortion in the generated images.
[0193] To address these issues, the SSEU-Net model proposed in this paper introduces the Res-SS2D module, a selective state-space computation unit optimized for image generation. The core of this module lies in the SS2D module, which extends the traditional state-space model to the two-dimensional image domain through a four-way feature traversal and recombination mechanism. The specific process is as follows:
[0194] Cross-scanning: The input feature map is traversed in parallel along four different directions to generate four sets of feature sequences with different spatial correlations. This operation is equivalent to constructing four complementary pseudo-time axes, providing multi-perspective feature input for subsequent state space calculations.
[0195] S6 module feature extraction: Four independent S6 modules are used to process each direction sequence. Each S6 module uses a dynamic weight selection mechanism to adaptively fuse the local features of the current scan path and the global dependency across the stride length.
[0196] Cross-merging: The four sets of output sequences are remapped to two-dimensional space according to the inverse scanning path, and finally the feature map of the same size as the input is restored.
[0197] 2. Edge extraction method:
[0198] Traditional pixel-domain steganography methods are prone to texture replication artifacts and chromatic distortion due to direct modification of spatial domain pixel values. Compared with pixel-domain embedding, the research community generally believes that high-frequency sensitive areas (such as edge contours and texture details) have better security advantages in information hiding. Based on this, SSEU-Net proposes an edge-guided steganography method, which enhances the encoder's perception of high-frequency features by introducing the Laplacian edge extraction operator. Specifically, a 4-neighborhood Laplacian operator is used to perform a convolution operation on the carrier image. The operator calculates the grayscale mutation intensity of the pixel point and the neighborhood through second-order differentials. Its mathematical essence is equivalent to solving the curvature change of the two-dimensional function of the image, which can be expressed as:
[0199] Where, and Respectively represent the horizontal and vertical directions in the image coordinate system, Represents a two-dimensional image In spatial coordinates The pixel intensity value at .
[0200] Finally, the extracted edge map is fused with the original image and input into the encoder, forcing the neural network to preferentially learn high-frequency area features so as to hide the secret information preferentially in the high-frequency area of the carrier image.
[0201] 3. For the encoder network:
[0202] The encoder network proposed in the present invention is based on an improved 7-layer U-Net architecture, and achieves efficient modeling of steganographic features through a multi-scale feature fusion mechanism. In the downsampling stage, except for the last layer, the convolution blocks of each layer are embedded with a Res-SS2D module, which uses its cross-scanning and state space modeling capabilities to enhance global context perception, while retaining local details through residual connections. In the upsampling stage, a pure convolutional structure is adopted, and the multi-scale features output by each downsampling layer are cascaded with the corresponding upsampling layer through jump connections to achieve cross-level feature complementarity and effectively suppress information loss. At the same time, the edge extraction method is used to generate a carrier edge map. After adjusting its resolution and extracting features through a convolutional layer, it is fused with each downsampling layer in an additive manner step by step, forcing the network to focus on high-frequency sensitive areas first.
[0203] 4. For the decoder network:
[0204] To balance model training and inference efficiency, the decoder network uses a basic 7-layer U-Net structure. Unlike the encoder network, the decoder network input is only the encrypted image, so the number of input channels is 3. During the downsampling phase, the number of channels and feature map size changes are consistent with those of the encoder, so they are not detailed here.
[0205] 5. For the target loss function:
[0206] To ensure the visual quality of the generated encrypted image and the restored image, the loss function used in this model uses Mean Squared Error (MSE) to calculate pixel differences, ensuring statistically high quality of the generated image. The Multi-scale Structural Similarity Index Measure (MS-SSIM) is also introduced to improve subjective visual quality. MSE calculates the mean squared difference between corresponding pixels in two images, achieving pixel-level error calculation. A smaller MSE value indicates a closer match between the generated image and the original. However, over-reliance on MSE can lead to loss of high-frequency details. MS-SSIM simulates the multi-resolution perception characteristics of the human visual system. It calculates brightness, contrast, and structural similarity at multiple scales, ultimately weighting them to produce a comprehensive score between 0 and 1 (a value closer to 1 indicates greater visual consistency and, therefore, higher quality). Compared to the single-scale SSIM, MS-SSIM is more sensitive to texture continuity and edge sharpness, effectively suppressing blocking artifacts and local distortion, resulting in higher subjective visual quality of the generated image. The two are used together to form a complementary optimization mechanism. The hidden loss is expressed as:
[0207] Where, Represents the number of carrier images and secret images input in a batch, Indicates that the current calculation is in progress. The loss of carrier image and encrypted image, represents the input carrier image, represents the encrypted image generated by the encoder, and Represent the weights of MSE loss and MS-SSIM loss respectively.
[0208] The form of the recovery loss is consistent with the hidden loss and is expressed as:
[0209] Where, represents a secret image, represents the restored image generated by the decoder.
[0210] In order to make better use of the edge feature fusion method, the present invention introduces low-frequency wavelet loss into the loss function, that is, the image is decomposed into high-frequency and low-frequency wavelet sub-bands through DWT transformation. Since the information hidden in the high-frequency component is less likely to be detected than the information hidden in the low-frequency component, this paper performs a low-frequency loss on the low-frequency components of the carrier image and the encrypted image. Loss, thus forcing the secret image to be hidden as much as possible in the high-frequency components of the carrier image, while retaining very little secret information in the low-frequency components, specifically expressed as:
[0211] Where, Represents the weight of the low-frequency wavelet loss. Specifically, the DWT transform uses a Haar wavelet basis for a one-layer decomposition and uses a zero-padding pattern to process image boundaries. This satisfies the requirements for low-frequency feature extraction while ensuring high computational efficiency. Finally, the low-frequency approximate coefficient subbands after the decomposition are separated and extracted as the low-frequency feature representation of the image.
[0212] The total training loss can then be defined as: .
[0213] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. There may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.
[0214] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0215] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0216] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely illustrative and are not intended to limit the scope of the present application. Various changes and modifications may be made therein by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as required by the appended claims.
[0217] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various application aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this approach of the present application should not be interpreted as reflecting the intention that the claimed application requires more features than those explicitly recited in each claim. More precisely, as reflected in the corresponding claims, the point of the application is that the corresponding technical problem can be solved with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present application.
[0218] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.
[0219] It should be noted that the above embodiments are illustrative rather than limiting of the present application, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The use of the words first, second, and third, etc., does not denote any order. These words may be interpreted as designations.
Claims
1. An image steganography method, characterized in that: The image steganography method comprises the following steps: Obtain a dataset, including carrier images and secret images; Construct a steganographic framework and input the carrier image, secret image, and carrier edge image generated by the edge extractor into the encoder. The encoder combines the features of the carrier edge image to hide the secret image in the carrier image and outputs the secret image. Constructing the Res-SS2D module, which extends the traditional state-space model to the two-dimensional image domain through a four-way feature traversal and recombination mechanism, and realizes four-way global spatial modeling of the input image; The edge extraction method is used to generate the carrier edge map. After adjusting its resolution and extracting features through the convolution layer, it is fused with the downsampling layers step by step, forcing the network to prioritize high-frequency sensitive areas. Construct an encoder network based on an improved 7-layer U-Net architecture, which achieves efficient modeling of steganographic features through a multi-scale feature fusion mechanism and embeds a Res-SS2D module in the downsampling stage to enhance global context awareness; Build a decoder network to recover the original secret image based on the input secret image; Design a multi-objective loss function to optimize the generation quality. At the same time, by introducing the L1 norm loss of the low-frequency component, constrain the consistency of the low-frequency region of the carrier image and the dense image. Train the model and optimize the network parameters through the back-propagation algorithm until the model converges; Input the carrier image and the secret image into the trained model, generate the secret image through the encoder; input the secret image into the decoder to restore the original secret image; Among them, building the Res-SS2D module includes: The input feature map is traversed in parallel along four different directions to generate four sets of feature sequences with different spatial correlations, providing multi-perspective feature input for subsequent state space calculations; Four independent S6 modules are used to process each direction sequence respectively. Each S6 module adaptively fuses the local features of the current scanning path and the global dependency of the stride length through a dynamic weight selection mechanism. Remap the four sets of output sequences to two-dimensional space according to the inverse scanning path, and finally restore the feature map of the same size as the input; Among them, the Res-SS2D module integrates depthwise separable convolution, performs channel-by-channel convolution on the input feature map, and uses several 1×1 convolution kernels to extract features between channels; the Res-SS2D module also introduces residual connections to avoid the gradient vanishing problem and enhance global feature fusion.
2. The image steganography method according to claim 1, characterized in that: The construction of the steganography framework includes: Let the carrier image be , the secret image is denoted as , the carrier edge image obtained by the carrier image through the edge extractor is recorded as ; The carrier image, secret image and carrier edge image are input into the encoder together. The encoder combines the features of the carrier edge image to hide the secret image into the carrier image and outputs the secret image, which is recorded as ; In the decoding stage, the secret image is input to the decoder, and the decoder recovers the original secret image from it, which is denoted as .
3. The image steganography method according to claim 2, characterized in that: The Laplace edge extraction operator is introduced to enhance the encoder's ability to perceive high-frequency features; The four-neighborhood Laplace operator is used to perform convolution operation on the carrier image, and the grayscale mutation intensity of the pixel and the neighborhood is calculated by second-order differential. The formula is as follows: ; Where, and Respectively represent the horizontal and vertical directions in the image coordinate system, Represents a two-dimensional image In spatial coordinates The pixel intensity value at ; The extracted edge map is fused with the original image and then input into the encoder, forcing the neural network to prioritize learning high-frequency area features so as to hide the secret information in the high-frequency area of the carrier image.
4. The image steganography method according to claim 3, characterized in that: The encoder network is improved based on the 7-layer U-Net architecture and realizes efficient modeling of steganographic features through a multi-scale feature fusion mechanism; In the downsampling stage, except for the last layer, the convolutional blocks of each layer are embedded with Res-SS2D modules. The cross-scanning and state-space modeling capabilities of the Res-SS2D modules are used to enhance global context perception while preserving local details through residual connections. In the upsampling stage, a pure convolutional structure is used to cascade the multi-scale features output by each downsampling layer with the corresponding upsampling layer through skip connections to achieve cross-level feature complementarity and suppress information loss; At the same time, the edge extraction method is used to generate the carrier edge map. After adjusting its resolution and extracting features through a convolutional layer, it is fused with the downsampling layers in an additive manner step by step, forcing the network to focus on high-frequency sensitive areas first.
5. The image steganography method according to claim 1, characterized in that: The edge map of the carrier is generated using an edge extraction method. After adjusting its resolution and extracting features through a convolutional layer, it is fused with the downsampling layers in an additive manner, forcing the network to prioritize high-frequency sensitive areas. The input layer concatenates the carrier image and the secret image along the channel dimension, and then compresses the feature map size to 1 / 32 of the original image through five levels of downsampling, and the number of channels increases step by step to 512.
6. The image steganography method according to claim 1, characterized in that: The loss function uses mean square error to calculate pixel differences and introduces a multi-scale structural similarity index to improve subjective visual quality. The hidden loss is expressed as: , In the above formula, Represents the number of carrier images and secret images input in a batch, Indicates that the current calculation is in progress. The loss of carrier image and encrypted image, represents the input carrier image, represents the encrypted image generated by the encoder, and Represent the weights of MSE loss and MS-SSIM loss respectively; The recovery loss is expressed as: , In the above formula, represents a secret image, represents the restored image generated by the decoder; Introducing low-frequency wavelet loss into the loss function, the image is decomposed into high-frequency and low-frequency wavelet sub-bands through DWT transformation, and the low-frequency components of the carrier image and the dense image are subjected to Loss, so that the secret image is partially hidden in the high-frequency components of the carrier image, while retaining part of the secret information in the low-frequency components, specifically expressed as: , Where, Represents the weight of the low-frequency wavelet loss; the DWT transform uses the Haar wavelet basis to perform a layer decomposition, and uses a zero-filling mode to process the image boundary, and finally separates and extracts the low-frequency approximate coefficient subband after the layer decomposition as the low-frequency feature representation of the image; The total training loss is defined as: , in, is the low-frequency wavelet loss, To recover losses, To hide losses.
7. An image steganography device, characterized in that: The image steganography device comprises: A data set acquisition module, used to obtain carrier images and secret images; An edge extraction module, for generating an edge image of a carrier image; The Res-SS2D module is used to extend the traditional state-space model to the two-dimensional image domain through a four-way feature traversal and recombination mechanism, achieving four-way global spatial modeling of the input image; The edge feature fusion module is used to adjust the resolution of the carrier edge map and extract features through the convolution layer, and then fuse them with the downsampling layers step by step, forcing the network to prioritize high-frequency sensitive areas; The encoder module is based on an improved 7-layer U-Net architecture. It is used to efficiently model steganographic features through a multi-scale feature fusion mechanism and embeds a Res-SS2D module in the downsampling stage to enhance global context awareness. The decoder module uses a 7-layer U-Net structure to input the secret image and restore the original secret image through upsampling; The loss function module is used to design a multi-objective loss function, combining PSNR and MS-SSIM to optimize the generated quality. At the same time, by introducing the L1 norm loss of the low-frequency component, it constrains the consistency of the low-frequency regions of the carrier image and the dense image. The training module is used to optimize the network parameters through the back-propagation algorithm until the model converges; The steganography execution module is used to input the carrier image and the secret image into the trained model, generate the secret image through the encoder, and input the secret image into the decoder to recover the original secret image; Among them, building the Res-SS2D module includes: The input feature map is traversed in parallel along four different directions to generate four sets of feature sequences with different spatial correlations, providing multi-perspective feature input for subsequent state space calculations; Four independent S6 modules are used to process each direction sequence respectively. Each S6 module adaptively fuses the local features of the current scanning path and the global dependency of the stride length through a dynamic weight selection mechanism. Remap the four sets of output sequences to two-dimensional space according to the inverse scanning path, and finally restore the feature map of the same size as the input; Among them, the Res-SS2D module integrates depthwise separable convolution, performs channel-by-channel convolution on the input feature map, and uses several 1×1 convolution kernels to extract features between channels; the Res-SS2D module also introduces residual connections to avoid the gradient vanishing problem and enhance global feature fusion.
8. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the image steganography method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program executable by an electronic device. When the computer program runs on the electronic device, the electronic device executes the steps of the image steganography method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Generative robust image steganography method
CN111598762A
Image adaptive steganography method and device, electronic equipment and medium
CN112118365A