Color image spatial domain steganalysis system and method

By introducing the ordered integration of improved EfficientNet network blocks and CNN-Transformer in the color image steganography analysis system, the problem of insufficient feature extraction and global feature capture in the prior art is solved, and more efficient color image steganography analysis and detection performance is achieved.

CN119946199APending Publication Date: 2025-05-06SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510009079.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing color image steganography analysis model has shortcomings in feature extraction and global feature capture, and it is impossible to effectively handle multi-scale and global features in color images.

Method used

A color image space steganography analysis system is proposed, based on the IHNet model, which includes feature extraction module, feature aggregation module and classification module. The feature extraction module extracts multi-scale features through improved EfficientNet network blocks and feature fusion enhancement blocks; the feature aggregation module creates a "C-C-C-T" structure through the ordered integration of CNN and Transformer to extract deep-level steganography analysis features.

Benefits of technology

The detection performance and accuracy of color image steganography analysis are improved, and multi-scale and global features in images can be extracted and processed more efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946199A_ABST
    Figure CN119946199A_ABST
Patent Text Reader

Abstract

The invention discloses a color image spatial steganalysis system, which is based on a color image spatial steganalysis network IHNet model, and the IHNet model at least comprises a feature extraction module, a feature aggregation module and a classification module. The feature extraction module comprises a preprocessing layer and a feature fusion enhanced block structure, and noise residual error information is obtained from an input color image by using an improved OfficientNet network block; the feature aggregation module is composed of four down-sampling blocks, a C-C-C-T structure is created by orderly integrating a CNN (Convolutional Neural Network) and a Transform, and deep steganalysis features are extracted; and the classification module converts the feature map output by the feature aggregation module into a feature vector through a global average pooling layer, inputs the feature vector into a fully connected layer, obtains final output by using a softmax function, and judges whether the original image contains hidden information according to aggregated feature information. The detection performance of the system and the method is more efficient and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image steganalysis, and mainly relates to a color image spatial domain steganalysis system and method. Background Art

[0002] Image steganography is a technique that hides secret information in an image file, with the goal of transmitting information without attracting attention, making it unnoticeable to the outside world. In contrast to steganography, the main purpose of image steganalysis technology is to detect whether there is hidden information in an image.

[0003] Existing image steganalysis methods can be divided into two categories: traditional methods based on artificial features and methods based on deep learning. Traditional steganalysis methods generally construct steganalysis features based on statistics first, and then classify them through integrated classifiers. Typical models include SRM, SPAM, etc. In recent years, with the vigorous development of deep learning technology in the field of image classification, researchers have gradually applied it to solve image steganalysis problems and designed a series of steganalysis models, such as SRNet, CovNet, Zhu-Net, SiastegNet, FPNet, SwT-SN, LWENet, IMcoatNet, etc., all of which have achieved good detection results.

[0004] The above-mentioned image steganalysis methods are all designed for grayscale images, laying the foundation for the research of color image steganalysis. The steganalysis of color images is more challenging than that of grayscale images because color images contain more color information and channels, making the embedding and hiding of secret information more complicated. The typical methods of traditional color image steganalysis are CRM, SGRM and GCRM. In 2019, Zeng et al. (J. Zeng, S. Tan, G. Liu, B. Li and J. Huang, "WISERNet: Wider Separate-Then-Reunion Network for Steganalysis of Color Images," in IEEE Transactions on Information Forensics and Security, vol. 14, no. 10, pp. 2735-2748, 2019) first applied deep learning technology to the steganalysis of color images and proposed WISERNet, a deep convolutional neural network with a wider structure that can separate information first and then aggregate it. Experimental results show that this method is better than CRM. In 2022, Wei et al. (K. Wei, W. Luo, S. Tan and J. Huang, "Universal Deep Network for Steganalysis of Color Image Based on Channel Representation," in IEEE Transactions on Information Forensics and Security, vol. 17, pp. 3022-3036, 2022.) proposed UCNet, a universal color image steganography model suitable for spatial and JPEG domains. The model uses a combination of 30 basic linear filters and 32 Gabor filters of SRM to extract noise residuals in the preprocessing stage, and designs three different convolution structures to extract high-level features. Experimental results show that UCNet has higher detection accuracy than WISERNet. From the current research status, there is relatively little work on color image steganalysis using deep learning. In addition, existing steganalysis models still have some shortcomings: on the one hand, after the model extracts the noise residual in the preprocessing layer, the feature extraction capability is insufficient and it is impossible to extract the steganalysis features at multiple scales; in addition, the downsampling stage is usually designed based on the convolutional structure, which usually captures local features and does not contain a visual Transformer architecture that can obtain global features. Summary of the invention

[0005] The present invention is aimed at the problems existing in the prior art, and proposes a color image spatial domain steganalysis system, which is based on the color image spatial domain steganalysis network IHNet model, and the IHNet model at least includes a feature extraction module, a feature aggregation module and a classification module; the feature extraction module includes a preprocessing layer and a feature fusion enhanced block structure, and uses the improved EfficientNet network block to obtain noise residual information from the input color image; the feature aggregation module is composed of 4 downsampling blocks, and creates a "CCCT" structure by orderly integrating CNN and Transformer to extract deep-level steganalysis features; the classification module converts the feature map output by the feature aggregation module into a feature vector through a global average pooling layer, inputs the feature vector into a fully connected layer, and then uses a softmax function to obtain the final output, and judges whether the original image contains hidden information based on the aggregated feature information. The detection performance of the system and method of the present invention is more efficient and accurate.

[0006] In order to achieve the above-mentioned purpose, the technical solution adopted by the present invention is: a color image spatial domain steganalysis system, based on the color image spatial domain steganalysis network IHNet model, the IHNet model at least includes a feature extraction module, a feature aggregation module and a classification module;

[0007] The feature extraction module includes a preprocessing layer and a feature fusion enhancement block structure, and uses an improved EfficientNet network block to obtain noise residual information from the input color image;

[0008] The feature aggregation module consists of four downsampling blocks, which integrate CNN and Transformer in an orderly manner to create a "CCCT" structure to extract deep steganalysis features;

[0009] The classification module: converts the feature map output by the feature aggregation module into a feature vector through a global average pooling layer, inputs the feature vector into a fully connected layer, and then uses a softmax function to obtain the final output, and determines whether the original image contains hidden information based on the aggregated feature information.

[0010] As an improvement of the present invention, the block structure of the feature fusion enhancement includes a No. 1 block structure and two No. 2 block structures.

[0011] The block structure No. 1 includes a combination of 1×1--3×3--1×1 convolutional layers, wherein the first 1×1 convolutional layer uses group convolution and then performs a channel shuffle operation; after the block structure No. 1, the number of feature map channels is reduced;

[0012] The block structure No. 2 includes a 1×1 convolution layer, a 3×3 depth-separable convolution layer, a selective kernel convolution, a 1×1 convolution layer and a dropout layer in sequence, and each convolution layer includes a BN layer and a Swish activation function.

[0013] As another improvement of the present invention, the four downsampling blocks in the feature aggregation module are block 3a, block 4, block 3b and block 5, respectively, wherein:

[0014] Blocks 3a and 3b are residual structures. The main network consists of two 3×3 convolutional layers and one average pooling layer. The short connection layer consists of a 1×1 convolutional layer with a step size of 2.

[0015] Block 4 is a bottleneck structure. The main network is composed of 1×1-3×3-1×1 convolutional layers in sequence. The middle layer uses group convolution with a step size of 2. The short connection layer is composed of a 3×3 convolutional layer with a step size of 2.

[0016] Block 5 adopts the Transformer structure from CoatNet. The main network uses the maximum pooling method to reduce the size of the feature map, converts the feature map into a series of vectors, and encodes the vectors by position. In the Transformer block, the feature map undergoes multi-head attention processing, and the branch network consists of a maximum pooling layer and a 1×1 convolutional layer. The feature maps of the main network and the branch network are fused by addition, and then processed by the feedforward network to extract deeper features.

[0017] In order to achieve the above object, the present invention also adopts a technical solution: a color image spatial domain steganalysis method, using the above system, comprising the following steps:

[0018] S1: Collect the data set, divide the data set into training set, validation set and test set, and use four color image steganography algorithms, CMD-CS-UNIWARD, CMD-C-Hill, GINA-S-UNIWARD and GINA-Hill, to embed secret information;

[0019] S2: Input the training set into the IHNet model based on the color image spatial domain steganalysis network, and train the model through the feature extraction module, feature aggregation module and classification module. Use the stochastic gradient descent algorithm to optimize the network model parameters. After each epoch, test the validation set to determine the model with the highest accuracy of the validation set after training.

[0020] S3: Input the test set to the model with the highest accuracy obtained in step S2, and determine whether the input image contains secret information based on the predicted results to obtain the analysis results.

[0021] As an improvement of the present invention, in step S1, secret information is embedded using four steganographic algorithms, namely, CMD-CS-UNIWARD, CMD-C-Hill, GINA-S-UNIWARD, and GINA-Hill, with 0.2bcp, 0.3bcp, and 0.4bcp, respectively, to construct four steganographic data sets.

[0022] As another improvement of the present invention, in the feature extraction module of the model training in step S2, the input color image is channel-separated in the preprocessing layer, channel-by-channel convolution is performed through a combination of an SRM filter and a Gabor filter, the extracted noise residual is truncated by a truncated linear function, the interference information is filtered, and finally the residual information of each channel is fused by channel splicing.

[0023] As another improvement of the present invention, in the feature extraction module of the model training in step S2, the block structure of feature fusion enhancement includes selective kernel convolution, and the selective kernel convolution includes three operation steps of segmentation, fusion and selection, wherein:

[0024] In the segmentation operation, multiple paths with convolution kernels of different sizes are generated;

[0025] In the fusion operation, the output feature maps of all branches are summed and a global average pooling operation is performed. Two fully connected layers are used to first reduce the dimension and then increase the dimension to obtain a weight matrix for softmax processing;

[0026] In the selection operation, the weight vector is multiplied by the output feature map of the corresponding branch to achieve weighted fusion, and the weighted fusion feature map is summed to obtain the final output feature map.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] 1. In the feature extraction module, after the preprocessing layer extracts the noise residual through a fixed filter, it does not directly downsample. Instead, a feature fusion enhancement block is designed. First, the information between channels is integrated through group convolution and channel shuffling operations. Then, two improved EfficientNet blocks are used to replace the SE module in the original block with SKConv to extract multi-scale features from the stego area.

[0029] 2. In the feature aggregation module, unlike the previous single CNN structure, the first three downsampling blocks retain the UCNet convolution structure while the fourth downsampling block introduces the Transformer structure to fuse the convolution structure into the aggregated feature map and extract the global features, thereby improving the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a structural schematic diagram of the IHNet model based on the spatial domain steganalysis network of color images of the present invention;

[0031] Figure 2 It is a structural schematic diagram of the feature fusion enhanced block structure FFEB in the system of the present invention;

[0032] Figure 3 Schematic diagram of the structure of the selective kernel convolution SKConv structure in the system of the present invention;

[0033] Figure 4 It is a structural schematic diagram of the feature aggregation module in the system of the present invention. DETAILED DESCRIPTION

[0034] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0035] Example 1

[0036] A color image spatial domain steganalysis system based on the color image spatial domain steganalysis network IHNet model, such as Figure 1 As shown, the IHNet model at least includes a feature extraction module, a feature aggregation module and a classification module;

[0037] The function of the feature extraction module is to obtain noise residual information from the input color image and improve the signal-to-noise ratio (the image content is similar to noise, while the steganographic noise acts as a signal); the feature aggregation module is mainly used to reduce the size of the feature map and extract deeper steganalysis features; the classification module integrates the feature information to determine whether the input image contains hidden information.

[0038] In the color image spatial domain steganalysis network IHNet model, the feature extraction module includes a preprocessing layer and a feature fusion enhanced block structure (FFEB), such as Figure 1 The “Preprocessing Layer”, “FFEB” shown in .

[0039] After the preprocessing layer is a feature fusion enhanced block structure (FFEB), the specific structure is as follows Figure 2As shown in the figure. In order to promote the fusion of information between channels, inspired by shuffleNet, a combination of 1×1--3×3--1×1 convolutional layers was first designed in FFEB, where the first 1×1 convolutional layer uses group convolution, and then performs channel shuffling. Channel shuffling rearranges channels so that channel information from different groups can be mixed together, thereby enhancing information exchange between different channels. These three convolutional layers are used to reduce the number of channels of the feature map obtained by the preprocessing layer, thereby reducing the computational cost.

[0040] After completing the information fusion, the core module of EfficientNet, the mobile flip bottleneck convolution MBConv, is improved to further enhance the features. MBConv is mainly composed of an ordinary 1×1 convolution layer, a 3×3 depth-separable convolution layer, a Squeeze-and-Excitation (SE) module, a 1×1 convolution layer and a dropout layer, and each convolution layer contains a BN layer and a Swish activation function. Considering that the traditional channel attention mechanism mainly focuses on the dependency between channels and ignores the fusion of multi-scale information, the selective kernel convolution (SKConv) in SKNet, which can fuse multi-convolution kernel features, is used to replace the SE module to circumvent the above problems. SKConv introduces multiple branches in the convolution layer, each branch uses convolution kernels of different scales for feature extraction, and then dynamically combines the outputs of these branches through the attention mechanism. SKConv will learn the weights of different convolution kernels according to the characteristics of the input feature map, thereby realizing the fusion and selection of multi-scale features. As Figure 3 As shown in the figure, the SKConv structure includes three main operations: segmentation, fusion, and selection. In the segmentation operation, multiple paths with different sizes of convolution kernels are generated. During the fusion operation, the output feature maps of all branches are summed and a global average pooling operation is performed. Then, two fully connected layers are used to first reduce the dimension and then increase the dimension, resulting in a softmax-processed weight matrix that encapsulates the information of each branch. Finally, in the selection operation, the weight vector is multiplied with the output feature map of the corresponding branch to achieve weighted fusion. The weighted fusion feature map is then summed to obtain the final output feature map.

[0041] The purpose of the feature aggregation module is to reduce the size of the feature map and extract deeper steganalysis features. The current color steganalysis models all use CNN to complete feature extraction and classification. Although CNN has a good performance in image processing and can process complex image features, it performs well in processing local features, but it is weak in processing global information. Transformer has a good performance in the field of NLP and can handle the modeling and generation of sequence data. It has advantages in processing global information, but its ability to process local information is relatively weak. Therefore, in this embodiment, CNN is combined with Transformer to effectively capture and process local and global information in the image, thereby improving the performance and effect of the model.

[0042] Inspired by CoatNet, the goal of the system of the present invention is to create a CCCT structure by sequentially integrating CNN and Transformer to effectively capture and process local and global information in the image. This method aims to improve the performance and effectiveness of the model. The system of the present invention uses four downsampling blocks for feature aggregation. Among them, the first three downsampling blocks retain the structure of UCNet, which are blocks 3a, 4, and 3b respectively. Figure 4 As shown, No. 3a and No. 3b are typical residual structures, where the main network consists of 2 3×3 convolutional layers and 1 average pooling layer, and the short connection layer consists of a 1×1 convolutional layer with a stride of 2. Block No. 4 is a typical bottleneck structure, where the main network consists of 1×1-3×3-1×1 convolutional layers in sequence, the intermediate layer uses group convolution with a stride of 2, and the short connection layer consists of a 3×3 size, stride of 2 convolutional layer. For the fourth downsampling block, the Transfomer structure from CoatNet is used. The main network first uses the maximum pooling method to reduce the size of the feature map, and then converts these feature maps into a series of vectors (tokens) and positionally encodes these vectors. In CoatNet, relative position encoding is used. After downsampling and preparation, the feature map is input into the Transformer block for processing. In the Transformer block, the feature map (now represented as a token) undergoes multi-head attention processing. The branch network consists of a maximum pooling layer and a 1×1 convolution layer, where the 1×1 convolution layer mainly adjusts the number of channels in the feature map. Then, the feature maps of the main network and the branch network are fused by addition and then processed by a feed-forward network (FFN) to extract deeper features.

[0043] The classification module is mainly responsible for judging whether the original image contains hidden information based on the aggregated feature information. The feature map output by the feature aggregation module is converted into a feature vector through the global average pooling layer, and then the feature vector is input into the fully connected layer, and finally the softmax function is used to obtain the final output.

[0044] The system of the present invention uses the improved EfficientNet network block in the feature extraction module to extract multi-scale features from the steganographic area; in the feature aggregation module, CNN and Transformer are used in combination to effectively capture and process local and global information in the image, thereby improving the performance of the model.

[0045] Example 2

[0046] A color image spatial domain steganalysis method, using the system as described in Example 1, specifically comprises the following steps:

[0047] Step S1: Prepare the data set. Use the ALASKA II steganalysis challenge data set to divide it into training set, validation set and test set. Then use four color image steganography algorithms, CMD-CS-UNIWARD, CMD-C-Hill, GINA-S-UNIWARD and GINA-Hill, to embed secret information.

[0048] Using the ALASKA II steganalysis challenge dataset, 20,000 color images with a size of 256×256 and a TIF format were randomly selected, and 14,000 images were selected for training, 1,000 images for verification, and 5,000 images for testing.

[0049] Four steganographic data sets were constructed using four steganographic algorithms, namely CMD-CS-UNIWARD, CMD-C-Hill, GINA-S-UNIWARD and GINA-Hill, to embed secret information with 0.2bcp, 0.3bcp and 0.4bcp respectively.

[0050] Step S2: Based on the training of the color image spatial domain steganalysis network IHNet model, the stochastic gradient descent (SGD) algorithm is used to optimize the network parameters. After each epoch, the validation set is tested and the model with the highest accuracy in the validation set is selected as the final model and applied to the test set.

[0051] During the model training process, the data passes through the feature extraction module, feature aggregation module and classification module in turn. In the feature extraction module, the operation flow in the preprocessing layer is as follows: the input color image is separated into channels, and then a combination of 30 SRM and 32 Gabor filters is used for channel-by-channel convolution. The extracted noise residual is truncated by a truncated linear function to limit the dynamic range of the residual and filter out interference information. Finally, the residual information of each channel is fused by channel splicing. At this time, 186 feature maps of 256×256 are obtained. The above steps refer to UCNet.

[0052] Unlike grayscale images, color images often have inherent correlations between channels, and these correlations will inevitably change during the embedding process of secret information. Therefore, after the preprocessing layer, a feature fusion enhanced block structure (FFEB) is formed. In the feature fusion enhanced block structure (FFEB), three convolutional layers are used to reduce the number of feature map channels obtained by the preprocessing layer from 186 to 32, thereby reducing the computational cost. Through SKConv, the weights of different convolution kernels are learned according to the characteristics of the input feature map, thereby realizing the fusion and selection of multi-scale features. In this embodiment, in the segmentation operation of the SKConv structure, 3×3 and 5×5 convolution kernels are used, followed by a BN layer and a SiLU activation function, and the final output feature map is obtained after the fusion operation and selection operation.

[0053] Then it passes through the feature aggregation module, which consists of 4 downsampling blocks. By integrating CNN and Transformer in an orderly manner, a "CCCT" structure is created to extract deep steganalysis features; finally, through the classification module, it is determined whether the original image contains hidden information and the model is trained.

[0054] In model training, the stochastic gradient descent (SGD) algorithm is used to optimize network parameters:

[0055] First, the stochastic gradient descent (SGD) algorithm is used to optimize the network parameters, and the mini-batch size is set to 32 (16 pairs of cover / stego). The initial learning rate is set to 0.01, and a total of 100 epochs are trained. At the 40th and 80th epochs, the learning rate is divided by 10 respectively;

[0056] Then, the carrier image and the stego image are trained in pairs. To prevent overfitting, the training set is randomly shuffled after each epoch. The same data augmentation technique (random mirroring and rotation) as SRNet training is used in this example. The initialization method and regularization settings are the same as those of UCNet;

[0057] Furthermore, for images with lower embedding rates, in order to make the network converge faster, a transfer learning strategy is used for progressive training, expressed as 0.4→0.3→0.2, that is, the model trained on images with an embedding rate of 0.4bcp is fine-tuned on images with an embedding rate of 0.3bcp;

[0058] Finally, during transfer learning, the initial learning rate was set to 0.001 and a total of 50 epochs were trained. At the 20th and 40th epochs, the learning rate was reduced to 0.1 times the original value.

[0059] The above completes the model training, and selects the model with the highest accuracy in the validation set as the final model and applies it to the test set.

[0060] Step S3: Input the test set to the model with the highest accuracy obtained in step S2, and determine whether the input image contains secret information based on the predicted results to obtain the analysis results.

[0061] The model is saved after each iteration. When the embedding rate is 0.4bcp, the model with the highest accuracy in the last 20 epochs on the validation set is selected as the final model. In other cases, the model with the highest classification accuracy on the validation set is selected and applied to the test set for testing.

[0062] Test Case

[0063] The experiments in this test case were run on Nvidia K40c GPUs. During the experiment, the MATLAB tool was used to construct the stego dataset, and the PyTorch deep learning framework was used for network training, verification, and testing. In model training, the stochastic gradient descent (SGD) algorithm was used to optimize the network parameters: First, the stochastic gradient descent (SGD) algorithm was used to optimize the network parameters, and the mini-batch size was set to 32 (16 pairs of cover / stego). The initial learning rate was set to 0.01, and a total of 100 epochs were trained. At the 40th and 80th epochs, the learning rate was divided by 10 respectively; then, the carrier image and the stego image were trained in pairs, and to prevent overfitting, the training set was randomly shuffled after each epoch.

[0064] The same data augmentation techniques (random mirroring and rotation) as those used in SRNet training were used in this test case. The initialization method and regularization settings are the same as those of UCNet. The model with the highest classification accuracy on the validation set is selected and applied to the test set for testing to obtain the detection results. In order to verify the performance of the model, the four color steganography algorithms, CMD-CS-UNIWARD, CMD-C-Hill, GINA-S-UNIWARD, and GINA-Hill, were tested and compared with the latest methods WISERNet and UCNet. When training WISERNet and UCNet, the training was carried out strictly in accordance with the experimental settings of the original paper. To make the results more convincing, all network models included in the comparative experiment used the same training set, validation set, and test set.

[0065] Experimental results show that the proposed IHNet has the best detection performance.

[0066]

[0067]

[0068] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0069] It should be noted that the above content only illustrates the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications all fall within the protection scope of the claims of the present invention.

Claims

1. A color image spatial domain steganalysis system, characterized by: Based on the color image spatial domain steganalysis network IHNet model, the IHNet model at least includes a feature extraction module, a feature aggregation module and a classification module; The feature extraction module includes a preprocessing layer and a feature fusion enhancement block structure, and uses an improved EfficientNet network block to obtain noise residual information from the input color image; The feature aggregation module consists of four downsampling blocks, which integrate CNN and Transformer in an orderly manner to create a "CCCT" structure to extract deep steganalysis features; The classification module: converts the feature map output by the feature aggregation module into a feature vector through a global average pooling layer, inputs the feature vector into a fully connected layer, and then uses a softmax function to obtain the final output, and determines whether the original image contains hidden information based on the aggregated feature information.

2. A color image spatial domain steganalysis system as claimed in claim 1, characterized in that: The feature fusion enhanced block structure includes a No. 1 block structure and two No. 2 block structures. The block structure No. 1 includes a combination of 1×1--3×3--1×1 convolutional layers, wherein the first 1×1 convolutional layer uses group convolution and then performs a channel shuffle operation; after the block structure No. 1, the number of feature map channels is reduced; The block structure No. 2 includes a 1×1 convolution layer, a 3×3 depth-separable convolution layer, a selective kernel convolution, a 1×1 convolution layer and a dropout layer in sequence, and each convolution layer includes a BN layer and a Swish activation function.

3. A color image spatial domain steganalysis system as claimed in claim 2, characterized in that: The four downsampling blocks in the feature aggregation module are block 3a, block 4, block 3b and block 5, respectively. Blocks 3a and 3b are residual structures. The main network consists of two 3×3 convolutional layers and one average pooling layer. The short connection layer consists of a 1×1 convolutional layer with a step size of 2. Block 4 is a bottleneck structure. The main network is composed of 1×1-3×3-1×1 convolutional layers in sequence. The middle layer uses group convolution with a step size of 2. The short connection layer is composed of a 3×3 convolutional layer with a step size of 2. Block 5 adopts the Transformer structure from CoatNet. The main network uses the maximum pooling method to reduce the size of the feature map, converts the feature map into a series of vectors, and encodes the vectors by position. In the Transformer block, the feature map undergoes multi-head attention processing, and the branch network consists of a maximum pooling layer and a 1×1 convolutional layer. The feature maps of the main network and the branch network are fused by addition, and then processed by the feedforward network to extract deeper features.

4. A color image spatial domain steganalysis method, using the system as claimed in claim 1, characterized in that: The steps include: S1: Collect the data set, divide the data set into training set, validation set and test set, and use four color image steganography algorithms: CMD-CS-UNIWARD, CMD-C-Hill l, GINA-S-UNIWARD, GINA-Hill ll to embed secret information; S2: Input the training set into the IHNet model based on the color image spatial domain steganalysis network, and train the model through the feature extraction module, feature aggregation module and classification module. Use the stochastic gradient descent algorithm to optimize the network model parameters. After each epoch, test the validation set to determine the model with the highest accuracy of the validation set after training. S3: Input the test set to the model with the highest accuracy obtained in step S2, and determine whether the input image contains secret information based on the prediction results to obtain the analysis results.

5. A color image spatial domain steganalysis method as claimed in claim 4, characterized in that: In the step S1, four steganographic algorithms, namely CMD-CS-UNIWARD, CMD-C-Hill, GINA-S-UNIWARD and GINA-Hill, are used to embed secret information with 0.2bcp, 0.3bcp and 0.4bcp respectively to construct four steganographic data sets.

6. A color image spatial domain steganalysis method as claimed in claim 4, characterized in that: In the feature extraction module of the model training in step S2, the input color image is channel-separated in the preprocessing layer, channel-by-channel convolution is performed through a combination of an SRM filter and a Gabor filter, the extracted noise residual is truncated by a truncated linear function, the interference information is filtered, and finally the residual information of each channel is fused in a channel splicing manner.

7. A color image spatial domain steganalysis method as claimed in claim 6, characterized in that: In the feature extraction module of the model training in step S2, the block structure of feature fusion enhancement includes selective kernel convolution, and the selective kernel convolution includes three operation steps of segmentation, fusion and selection, wherein: In the segmentation operation, multiple paths with convolution kernels of different sizes are generated; In the fusion operation, the output feature maps of all branches are summed and a global average pooling operation is performed. Two fully connected layers are used to first reduce the dimension and then increase the dimension to obtain a weight matrix for softmax processing; In the selection operation, the weight vector is multiplied by the output feature map of the corresponding branch to achieve weighted fusion, and the weighted fusion feature map is summed to obtain the final output feature map.