Image hiding method based on space-channel attention mechanism and related device
Through the image hiding method based on the spatial-channel attention mechanism, combined with DWT transformation and INN network, the problems of limited embedding capacity and weak anti-analysis ability in the existing technology are solved, high-capacity, high-quality and high-security image steganography is achieved, and the invisibility and anti-analysis ability of image steganography are improved.
Patent Information
- Application Number
- CN202510834737.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-03
AI Technical Summary
Existing image steganography algorithms have problems such as limited embedding capacity, weak anti-analysis ability, insufficient utilization of high-frequency areas, and lack of multi-dimensional attention, making it difficult to achieve an effective balance between security and visual quality.
An image hiding method based on the space-channel attention mechanism is adopted, combined with DWT transformation and INN network, and a space-channel joint attention mechanism is designed. The reversible hiding and high-quality extraction of secret images are achieved through forward and backward propagation. The low-frequency wavelet loss function is used to constrain the low-frequency consistency between the carrier and the secret image, and the embedding position selection is optimized.
It significantly improves the visual quality and steganographic security of confidential images, increases the steganographic capacity and anti-steganographic analysis capabilities, and ensures the invisibility of image steganography and the priority embedding of high-frequency information.
Smart Images

Figure CN120746809A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image hiding, and relates to an image hiding method and related devices based on a space-channel attention mechanism. Background Art
[0002] Image hiding technology, as a key technology in the field of information security, provides important protection for covert communication by embedding secret information into the carrier image to generate a confidential image.
[0003] Image steganography can be categorized into traditional image steganography algorithms and deep learning-based image steganography algorithms, depending on whether they employ deep learning techniques. Traditional image steganography algorithms are further divided into adaptive and non-adaptive image steganography. Non-adaptive algorithms (LSB, LSBMR, and GLSBM) embed information using fixed rules (least significant bit replacement or SDCS encoding). Their advantages include simple implementation, computational efficiency (complexity O(n)), and strong compatibility with the underlying image format. However, they have significant disadvantages. First, the rigid embedding strategy exposes statistical features (deviations in the pixel parity distribution of the LSB), making them susceptible to detection by tools such as SPAM and SRM. Second, modifications are concentrated in smooth areas, causing visual distortion (PSNR < 35dB). Finally, they suffer from poor robustness, making them vulnerable to noise attacks and compression interference. Adaptive algorithms (such as the HUGO, WOW, and UNIWARD series) dynamically optimize the embedding position using artificially designed distortion cost functions (such as directional filter residuals and wavelet coefficient differences) combined with STC coding. This prioritizes embedding the secret information in high-frequency noise regions, improving invisibility (PSNR > 40dB). However, its limitations cannot be ignored. Its algorithm relies heavily on artificial feature design, the capacity of high-frequency areas is limited (UNIWARD capacity is usually less than 0.4bpp), and it is difficult to cope with deep learning steganalysis (for example, SRNet's detection accuracy for S-UNIWARD is >80%).
[0004] Although deep learning-based steganography algorithms optimize performance through data-driven methods, they still face bottlenecks. GAN-based models (such as SGAN and SSGAN) suffer from unstable adversarial training between the generator and the discriminator, resulting in local noise anomalies (such as checkerboard artifacts) and statistical distribution shifts (such as distortion of high-frequency components of the histogram) in the secret image. Moreover, the steganographic capacity is limited by the resolution of the generated image. INN-based algorithms (HiNet and DeepMIH) achieve lossless steganography through reversible mapping but do not effectively utilize multi-scale image features (such as wavelet high-frequency subbands). As a result, the secret information is concentrated in low-frequency areas, which can easily lead to color distortion (contamination of low-frequency components of the carrier image) and texture replication artifacts (blurring of high-frequency details). Existing attention mechanisms (channel-wise attention CSE and spatial-wise attention SSE) only optimize feature weights in a single dimension and do not jointly model spatial-channel correlations. This leads to a crude embedding position selection strategy (such as the failure to accurately assign weights in areas with complex textures). In addition, the dense network structure is simple (stacked convolutional layers), making it difficult to capture the distribution characteristics of high-frequency noise. In addition, existing algorithms generally ignore the coordinated optimization of steganographic security and visual quality. For example, the Baluja algorithm causes the RGB histogram of the secret image to shift through channel splicing embedding, and the ISN algorithm causes the PSNR of the extracted image to decrease (about 3dB) due to the accumulation of rounding errors. Existing loss functions (such as mean square error) do not combine frequency domain constraints (such as wavelet low-frequency loss), making the embedded perturbation difficult to be compatible with the statistical characteristics of the carrier.
[0005] In summary, existing technologies face core problems such as limited embedding capacity, weak anti-analysis ability, insufficient utilization of high-frequency areas, and lack of multi-dimensional attention. A new steganography framework that integrates frequency domain analysis, reversible networks and joint attention mechanisms is needed to achieve an effective balance between security and visual quality. Summary of the Invention
[0006] The purpose of the present invention is to provide an image hiding method and related devices based on the spatial-channel attention mechanism to solve the technical problems of the existing steganography methods, such as limited embedding capacity, weak anti-analysis ability, insufficient utilization of high-frequency areas, and lack of multi-dimensional attention.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides an image hiding method based on a spatial-channel attention mechanism, comprising the following steps:
[0009] Obtain secret images and carrier images;
[0010] Input the secret image and the carrier image into the pre-trained image hiding network model to obtain the secret image and the lost information, thus completing the image hiding.
[0011] The image hidden network model is an image hidden network model based on a spatial-channel attention mechanism, including a preprocessing network, a hidden network and an extraction network; the hidden network consists of 16 hidden blocks, each of which consists of a SCADense module and exp(α(·)).
[0012] Furthermore, the preprocessing network includes a DWT transformation, which is used to perform a DWT transformation on the secret image and the carrier image, decompose them into high and low frequency sub-bands, and then send the high and low frequency sub-bands to the hidden network.
[0013] Furthermore, the SCA Dense module includes an input layer, an intermediate processing layer, an SCA attention layer and an output layer connected in sequence; the input layer includes a 3×3 convolutional layer with 64 channels; the intermediate processing layer consists of three 3×3 convolutional layers with the channels expanded to 256; the SCA attention layer includes a spatial-channel joint attention module; and the output layer includes a 1×1 convolutional layer with 64 channels.
[0014] Furthermore, the spatial-channel joint attention module includes spatial attention and channel attention, which is used to process the input feature map P to obtain the output feature map P', specifically including the following steps:
[0015] The spatial attention passes the input feature map P through a convolution layer with an output channel of 1 and a convolution kernel size of 1×1 to obtain a weight matrix with a dimension of 1×H×W; then the weight matrix is normalized using the sigmoid activation function to obtain a first weight matrix; the first weight matrix is multiplied with the original input feature map P in the spatial dimension to obtain a spatial enhanced feature map P s ;
[0016] The channel attention reduces the dimension of the input feature map P through global average pooling, and then uses two convolution layers with channel numbers C / 2 and C and a convolution kernel size of 1×1 to increase the dimension; then normalizes it through the sigmoid activation function to obtain a second weight matrix; multiply the second weight matrix with the original input feature map P in the channel dimension to obtain the channel enhanced feature map P c ;
[0017] The spatial enhancement feature map P s , channel enhanced feature map P c Add it to the original input feature map P to get the final output feature map P'.
[0018] Furthermore, the training of the image hiding network model includes the following steps:
[0019] Obtain an image hiding dataset, including a carrier image and a secret image; divide the image hiding dataset into a training set, a validation set, and a test set;
[0020] Determine the objective function of the image hidden network model and the total loss function L of the model Sum Including hidden loss L H , extraction loss L R and low-frequency wavelet loss L Freq , and is expressed using the mean square error function. The specific expression is:
[0021]
[0022] Where I and I' represent two images, i and j represent the horizontal and vertical coordinates of the pixels in images I and I', and H and W represent the height and width of the image.
[0023] The hidden loss L H The expression is:
[0024] L H =MSE(X cover ,X stego )
[0025] Where, X cover represents the carrier image, X stego Indicates a secret image;
[0026] The extraction loss L R The expression is:
[0027] L R =MSE(X secret ,X sr )
[0028] Where, X secret is the secret image, X sr is the extracted secret image;
[0029] The low-frequency wavelet loss L Freq The expression is:
[0030] L Freq =MSE(F(X cover ) LL ,F(X stego ) LL )
[0031] Where, F(·) LL Represents the operation of extracting the low-frequency subband of the image wavelet;
[0032] The total loss function L Sum The expression is:
[0033] L Sum =λ h L H +λr L R +λ f L Freq
[0034] Where λ h ,λ r and λ f It is used to balance the weights of different loss items;
[0035] The training set is input into the image hidden network model for training. The Adam optimizer is used in the training process. The weights of the image hidden network model are continuously updated during the training process, and the weight files are saved.
[0036] The trained image hidden network model is verified and tested using the validation set and test set until the preset requirements are met and the model training is completed.
[0037] Furthermore, the extraction network is composed of a plurality of extraction blocks; the extraction blocks and the hidden blocks have the same structure, but the information flow directions are opposite.
[0038] Furthermore, the method further comprises:
[0039] The secret image and auxiliary variables are input into the image hiding network model, decomposed into high and low frequency sub-bands through DWT transformation, and then sent to the extraction network for extraction. Finally, the extracted secret image and carrier image are obtained through IWT transformation.
[0040] In a second aspect, the present invention provides an image hiding system based on a spatial-channel attention mechanism, comprising:
[0041] An image acquisition module, used for acquiring a secret image and a carrier image;
[0042] The image hiding module is used to input the secret image and the carrier image into the pre-trained image hiding network model to obtain the secret image and the lost information, thus completing the image hiding;
[0043] The image hidden network model is an image hidden network model based on a spatial-channel attention mechanism, including a preprocessing network, a hidden network and an extraction network; the hidden network consists of 16 hidden blocks, each of which consists of a SCADense module and exp(α(·)).
[0044] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the image hiding method based on the spatial-channel attention mechanism are implemented.
[0045] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, implements the steps of the image hiding method based on the spatial-channel attention mechanism.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] This paper discloses an image hiding method and related device based on a spatial-channel attention mechanism. Combining the DWT transform, the INN network, and the attention mechanism, a deeply reversible image steganography framework is constructed. Through forward and backward propagation of the INN network, reversible hiding and high-quality extraction of secret images are achieved. Furthermore, a spatial-channel joint attention mechanism is designed to reconstruct a dense network to guide the embedding of secret images into texture-complex and high-noise regions of the carrier image, significantly improving the visual quality and steganographic security of the secret image.
[0048] Furthermore, the present invention introduces a reversible neural network (INN) framework to ensure strictly reversible information hiding and extraction. Through nonlinear transformations, secret information is dispersed into the global features of the carrier image, preventing the concentrated exposure of local statistical features, thereby improving steganographic security. Furthermore, the INN embeds secret information in the high-frequency subbands (LH / HL / HH) through reversible operations in the frequency domain, reducing interference with the low-frequency region (LL). This allows high-frequency information to be embedded preferentially, ensuring the invisibility of the steganography.
[0049] Furthermore, the image hiding network model of the present invention utilizes the DWT transform to convert the steganography from the spatial domain to the transform domain, addressing the embedding flaws of the spatial domain. The DWT decomposes the image into LL (low-frequency) and LH / HL / HH (high-frequency) subbands, guiding the embedding to focus on the high-frequency region and reducing visual distortion. A low-frequency wavelet loss function is used to constrain the low-frequency consistency of the carrier and the secret image, preventing the exposure of statistical features and improving the steganographic capacity while ensuring security.
[0050] Furthermore, the present invention designs a spatial-channel joint attention mechanism module and introduces it into the steganalysis network. Through dynamic region selection, multi-dimensional feature fusion, and high-frequency information enhancement, it not only solves the contradiction between security, capacity and quality in traditional steganography, but also solves the pain point that a single attention mechanism is difficult to deeply combine with frequency domain operations.
[0051] Furthermore, the present invention adopts the low-frequency wavelet loss function L FreqThis constraint addresses the problem of traditional high-frequency embedding methods that can indirectly perturb low-frequency components and cause statistical anomalies. Strictly limiting low-frequency modifications at high capacity further improves security. Statistical analysis shows that the MSE of the low-frequency subbands of the encrypted image and the carrier approach zero, and their histogram distributions are highly consistent, significantly outperforming competing algorithms in their ability to resist steganalysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0053] Figure 1 This is the overall framework diagram of image steganography based on the spatial-channel attention mechanism in the example of the present invention;
[0054] Figure 2 This is the spatial-channel joint attention mechanism structure in the example of the present invention;
[0055] Figure 3 This is a structural diagram of the CSA Dense module in an example of the present invention;
[0056] Figure 4(a) is a structural diagram of the CSE attention module in an example of the present invention;
[0057] Figure 4(b) is a structural diagram of the SSE attention module in an example of the present invention. DETAILED DESCRIPTION
[0058] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.
[0059] The following detailed description is an exemplary description, which is intended to provide further detailed description of the present invention. Unless otherwise indicated, all technical terms used in the present invention have the same meaning as those generally understood by those skilled in the art. The terms used in the present invention are only for describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present invention.
[0060] The embodiment of the present invention discloses an image hiding method based on the spatial-channel attention mechanism, and proposes a spatial-channel joint attention steganography framework driven by an invertible neural network (INN). High-capacity, high-quality, and high-security image steganography is achieved through the following core technologies: reversible steganography architecture, multi-scale frequency domain processing, dynamic feature guidance mechanism, and designed composite loss function.
[0061] Step 1, obtain the secret image and the carrier image;
[0062] Step 2: Input the secret image and the carrier image into the pre-trained image hiding network model to obtain the secret image and the lost information, thus completing the image hiding.
[0063] The image hidden network model is an image hidden network model based on the spatial-channel attention mechanism, including a preprocessing network, a hidden network and an extraction network; the hidden network consists of 16 hidden blocks, each of which is composed of a SCADense module and exp(α(·)); the overall framework is shown in the attached figure. Figure 1 shown.
[0064] In order to optimize network performance and ensure that the secret image can be effectively embedded in the high-frequency area of the carrier image, while reducing the computational burden caused by the difference in image size, the secret image and the carrier image are transformed by DWT, decomposed into high- and low-frequency sub-bands, and then sent to the forward hidden module. The output of the forward hidden module is transformed by IWT to obtain the secret image X. stego And the loss information R, the preprocessing network can be expressed by the following formula:
[0065] X secret_dwt ,X cover_dwt =DWT(X secret ,X cover )
[0066] Among them, X secret represents the secret image, X cover Represents a carrier image.
[0067] The specific structure of the hidden network is as follows Figure 1 As shown in Figure 1, the forward hidden module consists of 16 cascaded hidden blocks. The secret image and carrier image are first transformed using the DWT to decompose them into high- and low-frequency subbands. These are then fed into the forward hidden module. The output of the forward hidden module undergoes the IWT to yield the secret image and lost information. The forward hidden module consists of 16 hidden blocks, where the hidden and extraction blocks share the same submodules but have opposite information flow directions.
[0068] The CSA Dense module structure is as shown in the attached Figure 3As shown in the figure, the internal processing order of this module is the input layer, intermediate processing layer, SCA attention layer, and output layer concatenation. The input layer consists of a 64-channel 3×3 convolutional layer, the intermediate processing layer is composed of three 3×3 convolutional layers with 256 channels, and the SCA attention layer includes a joint spatial-channel attention module that parallelizes spatial and channel attention. The output layer uses a 1×1 convolution to reduce the dimensionality to 64 channels. This module uses dense connections to enhance feature reuse, effectively capturing image texture, edges, shape, and other information at different scales, helping to preserve rich image details and achieve higher-quality image representations.
[0069] The structure of the spatial-channel joint attention module is shown in the attached Figure 2 As shown in Figure 4(b), it consists of two parts: spatial attention and channel attention. As shown in Figure 4(b), for spatial attention, the input feature P is first passed through a convolution layer with an output channel of 1 and a convolution kernel size of 1×1 to obtain a weight matrix of dimension 1×H×W. The weight matrix is then normalized using the sigmoid activation function to obtain the final weight matrix. Finally, the weight matrix is multiplied with the original feature map P in the spatial dimension to obtain the final spatial enhanced feature map P. s , Figure 2 The channel attention mechanism is shown in the middle and lower part of the diagram. See Figure 4(a). The input feature P is reduced in dimension by global average pooling, and then increased in dimension by two convolutional layers with two channels, C / 2 and C, and a convolution kernel size of 1×1. It is then normalized by the sigmoid activation function to obtain the final weight matrix. Finally, the weight matrix is multiplied by the original feature map P in the channel dimension to obtain the final channel-enhanced feature map P. c ,Finally, the spatial enhancement feature map P s , channel enhanced feature map P c Adding it to the original input feature map P, we can get the final output feature map P'.
[0070] The extraction network is composed of 16 extraction blocks. The specific structure is shown in the attached figure. Figure 1 The two inputs are decomposed into high and low frequency sub-bands through DWT transformation, and then sent to the reverse extraction module composed of 16 extraction blocks. Finally, the output of the reverse extraction module is transformed by IWT to obtain the extracted secret image X. sr and the extracted carrier image X cr .
[0071] In step 3, the secret image and auxiliary variables are input into the image hiding network model, decomposed into high and low frequency sub-bands through DWT transformation, and then sent to the extraction block for reverse extraction. Finally, the extracted secret image and carrier image are obtained through IWT transformation.
[0072] This paper combines DWT transformation, INN network and attention mechanism to build a deep reversible image steganography framework. Figure 1 , through the forward and backward propagation of the INN network, reversible hiding and high-quality extraction of secret images are achieved.
[0073] In a feasible embodiment of the present invention, the training process of the image hiding network model is as follows:
[0074] Step 1: Obtain an image hiding dataset. This project randomly selects 1,000 images from the DIV2K dataset for model training, of which 800 are used for training, 100 for validation, and 100 for testing. The dataset is divided into two parts, one as the carrier image and the other as the secret image. The image resolution of the training set is 224×224, and the image resolution of the validation and test sets is 1024×1024.
[0075] Step 2: Determine the objective function of the image hidden network model, the total loss function L Sum It consists of three parts: the hiding loss L to ensure image hiding performance H , the extraction loss L that optimizes the image extraction accuracy R , and the low-frequency wavelet loss L that ensures the security of image steganography Freq . And use the mean square error function to represent them; the mean square error function is determined as follows:
[0076]
[0077] Where I and I' represent two images, i and j represent the horizontal and vertical coordinates of the pixels of images I and I', and H and W represent the height and width of the image.
[0078] Determine the hidden loss function L as follows H :
[0079] L H =MSE(X cover ,X stego )
[0080] Among them, X cover represents the carrier image, X stego Represents a secret image
[0081] The extraction loss function L is determined as follows: R :
[0082] L R =MSE(X secret ,X sr )
[0083] Among them, X secret is the secret image, Xsr is the extracted secret image.
[0084] The low-frequency wavelet loss function L is determined by the following formula: Freq :
[0085] L Freq =MSE(F(X cover ) LL ,F(X stego ) LL )
[0086] F(·) LL Represents the operation of extracting the low-frequency subband of an image wavelet.
[0087] The total extraction loss function L is determined as follows: Sum :
[0088] L Sum =λ h L H +λ r L R +λ f L Freq
[0089] Among them, λ h ,λ r and λ f It is used to balance the weights of different loss terms.
[0090] Step 3: Input the training set into the hidden network for training. The Adam optimizer is used in the training process, and the initial learning rate of the network model training is 1×10 -4.5 , hyperparameter λ h ,λ r and λ f They are set to 2, 2, and 1 respectively, the batch size is 8, and the number of training rounds is 3000.
[0091] Step 4: Save the model. During the training of the image hidden network, the weights of the network model are continuously updated and the weight files are saved.
[0092] Step 5: Test the image hidden network. Input the test set into the trained hidden network for testing, load the saved weight file, and obtain the image hiding and extraction results.
[0093] The embodiment of the present invention further discloses an image hiding system based on a spatial-channel attention mechanism, comprising:
[0094] An image acquisition module, used for acquiring a secret image and a carrier image;
[0095] The image hiding module is used to input the secret image and the carrier image into the pre-trained image hiding network model to obtain the secret image and the lost information, thus completing the image hiding;
[0096] The image hidden network model is an image hidden network model based on a spatial-channel attention mechanism, including a preprocessing network, a hidden network and an extraction network; the hidden network consists of 16 hidden blocks, each of which consists of a SCADense module and exp(α(·)).
[0097] 1. Hidden Network
[0098] like Figure 1 As shown in Figure 1, the core architecture of the hidden network consists of 16 cascaded hidden modules. The hidden module and the extraction module share the same basic submodule structure, but the information flow between them exhibits opposite characteristics. In the image hiding process, in order to optimize network performance, ensure that the secret image can be effectively embedded in the high-frequency area of the carrier image, and at the same time reduce the computational burden caused by the difference in image size, the secret image and the carrier image are transformed by DWT, decomposed into high- and low-frequency subbands, and then sent to the forward hidden module. The output of the forward hidden module is transformed by IWT to obtain the secret image X. stego and lost information R.
[0099] Spatial-channel joint attention mechanism structure Figure 2 As shown in the figure, it consists of two parts: spatial attention mechanism and channel attention mechanism. The spatial attention mechanism generates spatial weights through 1×1 convolution, focusing on areas with complex textures and improving embedding concealment. The channel attention mechanism uses global pooling and fully connected layers to strengthen key channels and enhance the ability to express high-frequency features. The spatial and channel weights are superimposed to achieve multi-dimensional feature optimization. The module structure of CSADense is shown in the figure. Figure 3 As shown in the figure, the internal processing order of this module is the input layer, intermediate processing layer, SCA attention layer, and output layer concatenation. The input layer consists of a 64-channel 3×3 convolutional layer, the intermediate processing layer is composed of three 3×3 convolutional layers with 256 channels, the SCA attention layer parallelizes space and channels, and the output layer uses 1×1 convolution to reduce the dimension to 64 channels. This module uses dense connections to enhance feature reuse, effectively capturing image texture, edges, shape, and other information at different scales, helping to preserve rich image details and thus achieve higher-quality image representation.
[0100] 2. Extract the network
[0101] The secret image is first decomposed into low-frequency and high-frequency subbands via discrete wavelet transform (DWT). Gaussian-distributed auxiliary noise is introduced as the initial input for reverse propagation. Reverse feature decoupling is then performed. The network consists of 16 cascaded extraction blocks, each of which gradually separates the secret information from the carrier information through inverse operations. Specifically, the extraction block uses the trained CSADense module to perform feature decoupling on the secret subbands. Inverse scaling and translation operations are used to remove the noise perturbations added during the hiding process and restore the original distribution of the secret subbands. After 16 iterative decoupling steps, the separated low-frequency and high-frequency subbands are subjected to the IWT transform to reconstruct the recovered carrier image and secret image.
[0102] In one embodiment of the present invention, a computer device is provided, comprising a processor and a memory, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used to operate an image hiding method based on a spatial-channel attention mechanism.
[0103] The present invention also provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device, used to store programs and data. It is understood that the computer-readable storage medium herein can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides storage space, which stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for being loaded and executed by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium herein can be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device. The processor can load and execute the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the image hiding method based on the spatial-channel attention mechanism in the above-mentioned embodiment.
[0104] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0105] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0106] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.
[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. An image hiding method based on spatial-channel attention mechanism, characterized in that: The following steps are involved: Obtain secret images and carrier images; The secret image and the carrier image are input into the pre-trained image hiding network model to obtain the secret image and the lost information, thus completing the image hiding. The image hiding network model is an image hiding network model based on a spatial-channel attention mechanism, including a preprocessing network, a hiding network and an extraction network; the hiding network consists of 16 hidden blocks, each of which consists of an SCA Dense module and exp(α(·)).
2. The image hiding method based on the spatial-channel attention mechanism according to claim 1, characterized in that: The preprocessing network includes a DWT transformation, which is used to perform a DWT transformation on the secret image and the carrier image, decompose them into high-frequency and low-frequency sub-bands, and then send the high-frequency and low-frequency sub-bands to the hidden network.
3. The image hiding method based on the spatial-channel attention mechanism according to claim 1, characterized in that: The SCADense module includes an input layer, an intermediate processing layer, an SCA attention layer and an output layer connected in sequence; the input layer includes a 3×3 convolutional layer with 64 channels; the intermediate processing layer consists of three 3×3 convolutional layers with channels expanded to 256; the SCA attention layer contains a spatial-channel joint attention module; the output layer includes a 1×1 convolutional layer with 64 channels.
4. The image hiding method based on the spatial-channel attention mechanism according to claim 3, characterized in that: The spatial-channel joint attention module includes spatial attention and channel attention, which is used to process the input feature map P to obtain the output feature map P', and specifically includes the following steps: The spatial attention passes the input feature map P through a convolution layer with an output channel of 1 and a convolution kernel size of 1×1 to obtain a weight matrix with a dimension of 1×H×W; then the weight matrix is normalized using the sigmoid activation function to obtain a first weight matrix; the first weight matrix is multiplied with the original input feature map P in the spatial dimension to obtain a spatial enhanced feature map P s ; The channel attention reduces the dimension of the input feature map P through global average pooling, and then uses two convolution layers with channel numbers C / 2 and C and a convolution kernel size of 1×1 to increase the dimension; then normalizes it through the sigmoid activation function to obtain a second weight matrix; multiply the second weight matrix with the original input feature map P in the channel dimension to obtain the channel enhanced feature map P c ; The spatial enhancement feature map P s , channel enhanced feature map P c Add it to the original input feature map P to get the final output feature map P'.
5. The image hiding method based on the spatial-channel attention mechanism according to claim 1, characterized in that: The training of the image hiding network model includes the following steps: Obtain an image hiding dataset, including a carrier image and a secret image; divide the image hiding dataset into a training set, a validation set, and a test set; Determine the objective function of the image hidden network model and the total loss function L of the model Sum Including hidden loss L H , extraction loss L R and low-frequency wavelet loss L Freq , and is expressed using the mean square error function. The specific expression is: Where I and I' represent two images, i and j represent the horizontal and vertical coordinates of the pixels in images I and I', and H and W represent the height and width of the image. The hidden loss L H The expression is: L H =MSE(X cover ,X stego ) Where, X cover represents the carrier image, X stego Indicates a secret image; The extraction loss L R The expression is: L R =MSE(X secret ,X sr ) Where, X secret is the secret image, X sr is the extracted secret image; The low-frequency wavelet loss L Freq The expression is: L Freq =MSE(F(X cover ) LL ,F(X stego ) LL ) Where, F(·) LL Represents the operation of extracting the low-frequency subband of the image wavelet; The total loss function L Sum The expression is: L Sum =λ h L H +λ r L R +λ f L Freq Where λ h ,λ r and λ f It is used to balance the weights of different loss items; The training set is input into the image hidden network model for training. The Adam optimizer is used in the training process. The weights of the image hidden network model are continuously updated during the training process, and the weight files are saved. The trained image hidden network model is verified and tested using the validation set and test set until the preset requirements are met and the model training is completed.
6. The image hiding method based on the spatial-channel attention mechanism according to claim 1, characterized in that: The extraction network is composed of a plurality of extraction blocks; the extraction blocks and the hidden blocks have the same structure, but the information flow directions are opposite.
7. The image hiding method based on the spatial-channel attention mechanism according to claim 6, characterized in that: The method further comprises: The secret image and auxiliary variables are input into the image hiding network model, decomposed into high and low frequency sub-bands through DWT transformation, and then sent to the extraction network for extraction. Finally, the extracted secret image and carrier image are obtained through IWT transformation.
8. An image hiding system based on spatial-channel attention mechanism, characterized in that: include: An image acquisition module, used for acquiring a secret image and a carrier image; The image hiding module is used to input the secret image and the carrier image into the pre-trained image hiding network model to obtain the secret image and the lost information, thus completing the image hiding; The image hidden network model is an image hidden network model based on a spatial-channel attention mechanism, including a preprocessing network, a hidden network and an extraction network; the hidden network consists of 16 hidden blocks, each of which consists of a SCADense module and exp(α(·)).
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the image hiding method based on the space-channel attention mechanism as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of an image hiding method based on a spatial-channel attention mechanism as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Medical image steganography method and system based on reversible deep network
CN121304425A
Reversible image steganography method based on multi-scale feature enhancement and dynamic attention
CN122340224A