CG Image Detection Method Based on Channel Joint and Soft Pooling of Dual-Stream Neural Network
By using the technology of combining dual-stream neural network channels and soft pooling layer in the CG detection method, the problem of insufficient detection performance of strong heterogeneous data sets in the existing technology is solved, and more efficient image feature extraction and information fusion are achieved, and detection performance is improved.
Patent Information
- Application Number
- CN202210867799.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-07-22
AI Technical Summary
Existing CG detection methods have limited detection performance when dealing with data sets with strong heterogeneity and ignore the problem of complementarity between noise and image semantic features.
The CG image detection method based on the dual-stream neural network channel joint and soft pooling layer is adopted. The noise information and shallow semantic information of the image are extracted separately through the sub-channel residual extraction module and the joint channel information extraction module, and downsampled through the soft pooling layer, and finally the two streams are fused to improve detection performance.
It effectively improves the detection performance of data sets with strong heterogeneity, integrates information with noise and image semantic features, and reduces information loss.
Smart Images

Figure CN115410029B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and in particular to a CG image detection method based on dual-stream neural network channel combination and soft pooling. Background Art
[0002] Computer-Generated Graphics (CG) refers to virtual but visually plausible images generated by computer software, and Photographs (PG) refers to real images taken with a digital camera. CG detection technology is a technology for distinguishing between these two types of images. In recent years, CG technology has developed rapidly and has been widely used in fields such as movies and games, resulting in the birth of a large number of image processing tools. Anyone without professional knowledge can use these image processing tools to generate visually realistic synthetic images. However, it is difficult for the human eye to distinguish between computer-generated images CG and natural images PG, which may be exploited by bad elements. Therefore, research on detection technology for distinguishing CG images has important academic significance and practical application value.
[0003] At present, CG detection methods are mainly divided into two categories: traditional detection methods based on artificially designed features and CG detection methods based on deep learning. Traditional detection methods are generally based on researchers' prior knowledge to manually mine features, which requires a deep understanding of the statistical differences between CG and PG, and the detection efficiency is low; CG detection methods based on deep neural networks mainly design deep neural network models to automatically learn the feature representation of a given task in an "end-to-end" manner, and then complete the classification. The detection performance is closely related to the deep neural network model. The prior art discloses a detection method for composed images, which first performs certain specific preprocessing operations on the image and then sends it to the deep neural network for training, or directly uses the original image as input, and then designs a suitable network architecture to extract effective classification features, but the detection ability for data sets with strong heterogeneity needs to be further improved. At the same time, most of them also ignore the complementarity between noise and image semantic features. Summary of the invention
[0004] In order to solve the problem that the current composition image detection method has limited detection performance for data sets with strong heterogeneity and ignores the complementarity between noise and image semantic features, the present invention proposes a CG image detection method based on two-stream neural network channel union and soft pooling layer, which effectively improves the detection performance for data sets with strong heterogeneity, integrates the information of noise and image semantic features, and reduces information loss.
[0005] In order to achieve the above technical effects, the technical solution of the present invention is as follows:
[0006] A CG image detection method based on two-stream neural network channel combination and soft pooling layer, comprising the following steps:
[0007] S1. Obtain a certain number of CG image samples and PG image samples to form an image data set;
[0008] S2. preprocessing the image samples in the image dataset;
[0009] S3. Construct a two-stream neural network model, which includes a channel-by-channel residual extraction module for extracting image noise information, a joint channel information extraction module for extracting shallow semantic information of the image, and a classifier. The channel-by-channel residual extraction module is provided with a residual structure for enhancing feature extraction of image samples. The entire two-stream neural network model adopts a soft pooling layer for downsampling;
[0010] S4. Divide the image data set into a training set, a validation set, and a test set, use the training set to train the constructed two-stream neural network model, and then use the validation set to evaluate the two-stream neural network model during the training process to obtain a trained two-stream neural network model;
[0011] S5. On the test set, the CG image samples to be detected are preprocessed in the same way as in step S2, and the preprocessed CG image samples to be detected are input into the trained two-stream neural network model, and the classification results of the CG image samples to be detected are output.
[0012] In the technical scheme, an image data set consisting of CG image samples and PG image samples is first obtained, the image samples in the image data set are preprocessed, and then a two-stream neural network model is constructed. The two-stream neural network model is provided with a channel residual extraction module and a joint channel information extraction module. The channel residual extraction module is provided with a residual structure, and the residual structure is used to enhance the feature extraction of image samples and reduce the information loss caused by the channel residual extraction module in the process of processing image samples. Then, the channel residual extraction module is used to extract the noise information of the image samples. However, the content of the original image samples has been basically lost after the image samples are processed by the channel residual extraction module. Therefore, the joint channel information extraction module is used to extract the shallow semantic information of the image samples as the information supplement of the channel residual extraction module. The two-stream neural network model uses a soft pooling layer for downsampling. Compared with the maximum pooling layer, the soft pooling layer can enhance the feature representation of the image samples and retain the original sample attributes of the image. Finally, the two streams are fused to obtain the classification result, which effectively improves the detection performance of the data set with strong heterogeneity, integrates the information of noise and image semantic features, and reduces information loss.
[0013] Preferably, in step S3, the channel-by-channel residual extraction module and the joint channel information extraction module are operated in parallel, the classifiers of the dual-stream neural network model are respectively connected to the channel-by-channel residual extraction module and the joint channel information extraction module, and the image samples are respectively input into the channel-by-channel residual extraction module and the joint channel information extraction module, the channel-by-channel residual extraction module extracts the noise information of the image samples, and outputs a first residual feature map, the joint channel information extraction module extracts the shallow semantic information of the image samples, and uses it as the information supplement of the channel-by-channel residual extraction module, and outputs a second residual feature map, the classifier fuses the information of the first residual feature map and the second residual feature map, and outputs a classification result of the image sample.
[0014] Preferably, the channel-by-channel residual extraction module includes an image processing module and a feature extraction module, the image processing module is composed of three groups of SRM residual filter kernels and is connected to the feature extraction module, the feature extraction module is composed of a first convolution block connected in sequence, a three-layer residual structure with the same structure, and a second convolution block, the first convolution block and the second convolution block are both composed of a convolution layer containing a 3×3 convolution kernel, a BN layer, a ReLU activation function layer and a soft pooling layer connected in sequence; each layer of the residual structure is provided with a main branch and a branch branch, the main branch is provided with a convolution layer containing a 3×3 convolution kernel and a step size of 1, a ReLU activation function layer and a soft pooling layer connected in sequence, the branch branch is provided with a convolution layer containing a 3×3 convolution kernel and a step size of 2, and each convolution layer is provided with 128 convolution kernels.
[0015] Preferably, the specific process of inputting the image sample into the channel-by-channel residual extraction module is as follows:
[0016] S31. The image samples are divided into three channels of R, G, and B and input into the image processing module in the channel residual extraction module, and the residual features of the image of each channel are extracted using the SRM residual filter kernel in the image processing module, and the residual features of the image extracted from each channel are merged by channel to obtain a third residual feature map;
[0017] S32. Input the third residual feature map into the first convolution block, and then respectively input the main branch and the branch branch of the residual structure in the channel residual extraction module, the third residual feature map sequentially passes through the convolution layer, the ReLU activation function layer and the soft pooling layer for downsampling in the main branch, and outputs a fourth residual feature map, the third residual feature map is extracted into a fifth residual feature map of the same scale as that in the main branch through the convolution layer in the branch branch, the fourth residual feature map and the fifth residual feature map are fused together by adding, and output to obtain a sixth residual feature map;
[0018] S33. Use the ReLU activation function to activate the data of the sixth residual feature map, input the output result into the second convolution block, and then use the second convolution block to further refine and fuse the residual features of the sixth residual feature map, and output the first feature map.
[0019] Preferably, the specific calculation process of the soft pooling layer sampling is:
[0020] For a pixel region R of size 2×2, first calculate the SoftMax value of each pixel in the region to obtain w i (i=1,2,3,4), the calculation formula is as follows:
[0021]
[0022] where α i represents the pixel value of the i-th pixel, w i Represents the weight corresponding to the i-th pixel;
[0023] Then multiply the original pixel area element by element and add the four values to get the pooled result The calculation formula is as follows:
[0024]
[0025] Preferably, the joint channel information extraction module is provided with a third convolution block, a fourth convolution block, a fifth convolution block, a sixth convolution block and a seventh convolution block connected in sequence, and each of the third convolution block, the fourth convolution block, the fifth convolution block, the sixth convolution block and the seventh convolution block is composed of a convolution layer including a 3×3 convolution kernel, a BN layer, a ReLU activation function layer and a soft pooling layer connected in sequence, wherein the convolution layer of the third convolution block is provided with 32 convolution kernels, the convolution layer of the fourth convolution block is provided with 64 convolution kernels, and each of the convolution layers of the fifth convolution block, the sixth convolution block and the seventh convolution block is provided with 128 convolution kernels.
[0026] Preferably, the classifier is composed of a global average pooling layer GAP and a fully connected layer FC, and the global average pooling layer GAP is connected to the fully connected layer FC, wherein the specific calculation process of the global average pooling layer is:
[0027] For a pixel area R of size M×N, the calculation formula is as follows:
[0028]
[0029] Among them, α ij Represents the pixel value of the pixel at position (i, j);
[0030] Preferably, the processing process of the classifier is:
[0031] Assume that the first feature map output by the residual extraction module is f 1 , the second feature map output by the joint channel information extraction module is represented as f 2 , f 1 and f 2 satisfy:
[0032]
[0033] Where C represents the number of channels of the feature map, H and W represent the height and width of the image sample respectively;
[0034] Perform a global average pooling operation GAP on the first feature map and the second feature map to obtain a feature map The feature map GAP(f i ) is input into the fully connected layer FC to get the output result Add the output results element by element and then pass them through a SoftMax activation function to get a binary probability vector The calculation formula for p is:
[0035]
[0036] Among them, the two elements of p represent the probability that the category of the image sample is CG and the probability that the category of the image sample is PG respectively.
[0037] Preferably, in step S4, the parameters and weights of the two-stream neural network model to be trained are updated by back propagation, and the loss function formula used when training the two-stream neural network model is as follows:
[0038]
[0039] Where N represents the number of image samples in a Batch, K represents the number of classification categories, K = 2, Represents a Boolean function that returns 1 when it is true and 0 when it is false. i represents the true label of the i-th sample.
[0040] Preferably, in step S5, in the process of outputting the classification result of the CG image sample to be detected, the detection accuracy Acc is used as the evaluation index of the classification result, and its calculation formula can be expressed as:
[0041]
[0042] Where P represents the number of positive samples, which refers to PG image samples; N represents the number of negative samples, which refers to CG image samples; TP and TN refer to the number of positive samples and negative samples that are correctly classified, respectively.
[0043] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0044] The invention proposes a CG image detection method based on dual-stream neural network channel union and soft pooling layer. Firstly, a dual-stream neural network model is constructed. The dual-stream neural network model is provided with a channel residual extraction module and a joint channel information extraction module. A residual structure is provided in the channel residual extraction module. The residual structure is used to enhance the feature extraction of image samples and reduce the information loss caused by the channel residual extraction module in the process of processing the image samples. Then, the channel residual extraction module is used to extract the noise information of the image samples. However, the content of the original image samples is basically lost after the image samples are processed by the channel residual extraction module. Therefore, the joint channel information extraction module is used to extract the shallow semantic information of the image samples to supplement the image sample content lost when the image samples are processed by the channel residual extraction module. The dual-stream neural network model uses a soft pooling layer for downsampling. The soft pooling layer can enhance the feature representation of the image samples and retain the original sample attributes of the image. Finally, the two streams are fused to obtain a classification result, which effectively improves the detection performance of a data set with strong heterogeneity, fuses the information of noise and image semantic features, and reduces the information loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A schematic diagram showing a flow chart of a CG image detection method based on two-stream neural network channel combination and soft pooling proposed in an embodiment of the present invention;
[0046] Figure 2 A structural diagram representing a two-stream neural network model;
[0047] Figure 3 It represents the residual structure diagram proposed in the embodiment of the present invention;
[0048] Figure 4 A diagram showing a soft pooling calculation process proposed in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0050] In order to better illustrate the present embodiment, some parts of the drawings may be omitted, enlarged or reduced, and do not represent actual sizes. The description of the directions of parts such as "upper" and "lower" does not limit the present patent;
[0051] It is understandable to those skilled in the art that some well-known contents may be omitted in the drawings;
[0052] The positional relationships described in the drawings are only for illustrative purposes and should not be construed as limiting the present patent.
[0053] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0054] Example 1
[0055] like Figure 1 As shown, a CG image detection method based on two-stream neural network channel combination and soft pooling includes the following steps:
[0056] S1. Obtain a certain number of CG image samples and PG image samples to form an image data set;
[0057] In step S1, two image datasets are selected, namely SPL2018 and DsTok. All CG images and PG images in DsTok are collected from the Internet and have strong heterogeneity. The CG sources in SPL2018 are diverse, and natural images are also taken with different models of equipment in different scenes. The datasets are highly heterogeneous and the detection difficulty is higher.
[0058] S2. preprocessing the image samples in the image dataset;
[0059] In step S2, the specific operation of preprocessing is to crop the size of the image samples in the image data set into a size of 224×224 by center cropping;
[0060] S3. Construct a two-stream neural network model, which includes a channel-by-channel residual extraction module for extracting image noise information, a joint channel information extraction module for extracting shallow semantic information of the image, and a classifier. The channel-by-channel residual extraction module is provided with a residual structure for enhancing feature extraction of image samples. The entire two-stream neural network model adopts a soft pooling layer for downsampling;
[0061] In step S3, the channel-by-channel residual extraction module and the joint channel information extraction module are operated in parallel, the classifiers of the two-stream neural network model are respectively connected to the channel-by-channel residual extraction module and the joint channel information extraction module, the image samples are respectively input into the channel-by-channel residual extraction module and the joint channel information extraction module, the channel-by-channel residual extraction module extracts the noise information of the image samples, and outputs a first residual feature map, the joint channel information extraction module extracts the shallow semantic information of the image samples, and uses it as the information supplement of the channel-by-channel residual extraction module, and outputs a second residual feature map, the size of the second residual feature map is consistent with that of the first residual feature map, and the classifier separates the first residual feature map and the second residual feature map. The information of the feature map is fused, and the classification result of the image sample is output; the joint channel information extraction module is provided with a third convolution block, a fourth convolution block, a fifth convolution block, a sixth convolution block and a seventh convolution block connected in sequence, and each of the third convolution block, the fourth convolution block, the fifth convolution block, the sixth convolution block and the seventh convolution block is composed of a convolution layer including a 3×3 convolution kernel, a BN layer, a ReLU activation function layer and a soft pooling layer connected in sequence, wherein the convolution layer of the third convolution block is provided with 32 convolution kernels, the convolution layer of the fourth convolution block is provided with 64 convolution kernels, and each of the convolution layers of the fifth convolution block, the sixth convolution block and the seventh convolution block is provided with 128 convolution kernels.
[0062] S4. Divide the image data set into a training set, a validation set, and a test set, use the training set to train the constructed two-stream neural network model, and then use the validation set to evaluate the two-stream neural network model during the training process to obtain a trained two-stream neural network model;
[0063] In step S4, the DsTok image dataset is divided into training set, validation set and test set in a ratio of 3:1:1, and the SPL2018 image dataset is divided into training set, validation set and test set in a ratio of 10:3:4.
[0064] S5. On the test set, the CG image samples to be detected are preprocessed in the same way as in step S2, and the preprocessed CG image samples to be detected are input into the trained two-stream neural network model, and the classification results of the CG image samples to be detected are output.
[0065] In step S5, in the process of outputting the classification result of the CG image sample to be detected, the detection accuracy Acc is used as the evaluation index of the classification result, and its calculation formula can be expressed as:
[0066]
[0067] Wherein P represents the number of positive samples, positive samples refer to PG image samples, N represents the number of negative samples, negative samples refer to CG image samples, TP and TN refer to the number of correctly classified positive samples and negative samples, respectively. In order to demonstrate the effectiveness of the method of this embodiment, the method of this embodiment is mainly compared with 6 existing advanced detection methods, and the two-stream neural network model of this embodiment is evaluated on the DsTok and SPL2018 data sets. At the same time, in order to make the detection results more convincing, this embodiment randomly divides each data set three times and conducts fair tests on each division. Finally, the average result obtained by the three divisions is taken as the final evaluation result.
[0068] Table 1 is a comparison table of detection results of different models on the DsTok dataset. Referring to Table 1, the detection results of 6 existing models on the DsTok dataset, the detection accuracy Acc of the 6 models are 85.3%, 88.8%, 83.2%, 93.4%, 93.9% and 92.1% respectively, while the detection accuracy Acc of the two-stream neural network model of this embodiment is 96.9%, which achieves the best detection performance on the DsTok dataset. Compared with the current best detection method Quan (2020), the Acc of the two-stream neural network model of this embodiment is improved by 3%; Table 1 shows the comparison of detection results of different models on the DsTok dataset;
[0069] Table 1
[0070]
[0071] Table 2 is a comparison table of detection results of different models on the DsTok dataset. See Table 2, the detection results of 6 existing models on the SPL2018 dataset, the detection accuracy Acc of the 6 models are 89.4%, 89.8%, 88.0%, 92.8%, 92.8% and 93.5%, respectively, and the detection accuracy Acc of the two-stream neural network model of this embodiment is 93.9%, achieving the optimal detection performance. Compared with the current advanced CG detection models, such as the Quan (2020) model and the Yao (2022) model, this method has improved by 1.1% and 0.4%, respectively, which confirms that the method implemented in this embodiment has the highest detection performance on the current two mainstream datasets.
[0072] Table 2
[0073]
[0074] Example 2
[0075] See also Figure 2, the channel-by-channel residual extraction module includes an image processing module and a feature extraction module. The image processing module is composed of three groups of SRM residual filter kernels and is connected to the feature extraction module. The feature extraction module is composed of a first convolution block connected in sequence, a three-layer residual structure with the same structure, and a second convolution block. The first convolution block and the second convolution block are both composed of a convolution layer containing a 3×3 convolution kernel, a BN layer, a ReLU activation function layer, and a soft pooling layer connected in sequence; see Figure 2 and Figure 3 Each layer of the residual structure is equipped with a main branch and a branch branch. The main branch is equipped with a convolution layer containing a 3×3 convolution kernel and a stride of 1, a ReLU activation function layer and a soft pooling layer connected in sequence. The branch branch is equipped with a convolution layer containing a 3×3 convolution kernel and a stride of 2. Each convolution layer is equipped with 128 convolution kernels.
[0076] See also Figure 1 In step S3, the specific processing process of the image sample input channel residual extraction module is as follows:
[0077] S31. The image samples are divided into three channels of R, G, and B and input into the image processing module in the channel residual extraction module, and the residual features of the image of each channel are extracted using the SRM residual filter kernel in the image processing module, and the residual features of the image extracted from each channel are merged by channel to obtain a third residual feature map;
[0078] In step S31, 30 SRM filter kernels are used for each channel to extract the residual features of the image samples of each channel, and then they are fused by channel to obtain 90 residual feature maps. Using this processing method, the fused residual features can be fully learned, and the representation of the relationship between local pixels in the same channel can be strengthened. More complex statistical characteristics can be obtained, which helps to amplify the difference between PG and CG images. Different from the shooting of PG images which is restricted by time, place and environment, CG images contain many scenes that do not exist in reality, and the semantic content of the image also contains important classification information. After the image is processed by the SRM filter kernel, the original image content has been basically lost. Therefore, the joint channel information extraction module is used to extract the shallow semantic information of the image sample to supplement the image sample content lost when the residual module processes the image sample.
[0079] S32. Input the third residual feature map into the first convolution block, and then respectively input the main branch and the branch branch of the residual structure in the channel residual extraction module, the third residual feature map sequentially passes through the convolution layer, the ReLU activation function layer and the soft pooling layer for downsampling in the main branch, and outputs a fourth residual feature map, the third residual feature map is extracted into a fifth residual feature map of the same scale as that in the main branch through the convolution layer in the branch branch, the fourth residual feature map and the fifth residual feature map are fused together by adding, and output to obtain a sixth residual feature map;
[0080] S33. Use the ReLU activation function to activate the data of the sixth residual feature map, input the output result into the second convolution block, and then use the second convolution block to further refine and fuse the residual features of the sixth residual feature map, and output the first residual feature map.
[0081] In step S33, after each layer, the size of the output feature map becomes half of the input, and the size of the final output first residual map becomes 7×7. Since each convolution layer in the channelized residual extraction module is equipped with 128 convolution kernels, the output channel is always maintained at 128 dimensions, and finally a 128-dimensional first residual feature map can be obtained.
[0082] The basic functions of the pooling layer include: reducing the amount of calculation, reducing model redundancy, and preventing model overfitting. Soft pooling is a variant structure of the pooling layer that can enhance feature representation while retaining the basic properties of the input. Specifically, soft pooling reduces the information loss caused by pooling while maintaining the basic functions of the pooling layer. See Figure 4 , the specific calculation process of soft pooling layer down sampling is:
[0083] For a 2×2 pixel region R, first calculate the SoftMax value of each pixel in the region to get w i (i=1,2,3,4), the calculation formula is as follows:
[0084]
[0085] where α i represents the pixel value of the i-th pixel, w i Represents the weight corresponding to the i-th pixel;
[0086] Then multiply the original pixel area element by element and add the four values to get the pooled result The calculation formula is as follows:
[0087]
[0088] Example 3
[0089] See also Figure 2 , the classifier is composed of a global average pooling layer GAP and a fully connected layer FC, the global average pooling layer GAP is connected to the fully connected layer FC, wherein the specific calculation process of the global average pooling layer is:
[0090] For a pixel area R of size M×N, the calculation formula is as follows:
[0091]
[0092] Among them, α ij Represents the pixel value of the pixel at position (i, j).
[0093] The processing of the classifier is:
[0094] Assume that the first feature map output by the residual extraction module is f 1 , the second feature map output by the joint channel information extraction module is represented as f 2 , f 1 and f 2 satisfy:
[0095]
[0096] Where C represents the number of channels of the feature map, H and W represent the height and width of the image sample respectively;
[0097] Perform a global average pooling operation GAP on the first feature map and the second feature map to obtain a feature map The feature map GAP(f i ) is input into the fully connected layer FC to get the output result Add the output results element by element and then pass them through a SoftMax activation function to get a binary probability vector The calculation formula for p is:
[0098] p = Softmax(FC(GAP(f 1 ))+FC(GAP(f 2 )))
[0099] Among them, the two elements of p represent the probability that the category of the image sample is CG and the probability that the category of the image sample is PG respectively.
[0100] See also Figure 1 In step S4, the parameters and weights of the two-stream neural network model to be trained are updated by back propagation. The loss function formula used when training the two-stream neural network model is as follows:
[0101]
[0102] Where N represents the number of image samples in a Batch, K represents the number of classification categories, K = 2, Represents a Boolean function that returns 1 when it is true and 0 when it is false. i represents the true label of the i-th sample.
[0103] Other experimental settings are as follows: The two-stream neural network model was trained using NVIDIA's Titan GPU, and the Pytorch deep learning framework was adopted. During the training of the two-stream neural network model, we used cross entropy loss as the loss function and used the SGD optimizer to optimize the two-stream neural network model. The batch_size was set to 64, the initial learning rate was set to 1e-3, and the weight decay rate was set to 1e-3. A total of 120 rounds of training were performed, and the learning rate decayed to 0.5 times the original value every 20 rounds.
[0104] Obviously, the above embodiments of the present invention are only examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A computer synthetic image detection method based on the combination of channels and soft pooling of a two-stream neural network, characterized in that, it includes the following steps: S1. Obtain a certain number of computer synthetic image samples and natural image samples to form an image data set; S2. Preprocess the image samples in the image data set; S3. Construct a two-stream neural network model. The two-stream neural network model includes a channel-separated residual extraction module for extracting image noise information, a joint channel information extraction module for extracting shallow semantic information of the image, and a classifier. The channel-separated residual extraction module is provided with a residual structure for enhancing the feature extraction of image samples. The entire two-stream neural network model adopts a soft pooling layer for downsampling; S4. Divide the image data set into a training set, a validation set, and a test set. Use the training set to train the constructed two-stream neural network model, and then use the validation set to evaluate the two-stream neural network model during the training process to obtain a trained two-stream neural network model; S5. On the test set, perform the same preprocessing on the computer synthetic image samples to be detected as in step S2, input the preprocessed computer synthetic image samples to be detected into the trained two-stream neural network model, and output the classification results of the computer synthetic image samples to be detected; In step S3, the channel-separated residual extraction module and the joint channel information extraction module are parallel. The classifier of the two-stream neural network model is respectively connected to the channel-separated residual extraction module and the joint channel information extraction module. The image samples are respectively input into the channel-separated residual extraction module and the joint channel information extraction module. The channel-separated residual extraction module extracts the noise information of the image samples and outputs the first residual feature map. The joint channel information extraction module extracts the shallow semantic information of the image samples and uses it as the information supplement for the channel-separated residual extraction module, and outputs the second residual feature map. The classifier fuses the information of the first residual feature map and the second residual feature map and outputs the classification results of the image samples.
2. The computer synthetic image detection method based on the combination of channels and soft pooling of a two-stream neural network according to claim 1, characterized in that, The sub-channel residual extraction module includes an image processing module and a feature extraction module. The image processing module consists of three groups of SRM residual filter kernels and is connected to the feature extraction module. The feature extraction module consists of a first convolutional block, three residual structures with the same structure, and a second convolutional block connected in sequence. Both the first convolutional block and the second convolutional block are composed of a convolutional layer with a 3 ×3 convolutional kernel, a BN layer, a ReLU activation function layer, and a soft pooling layer connected in sequence; each residual structure has a main branch and a sub-branch. The main branch has a convolutional layer with a 3 ×3 convolutional kernel and a stride of 1, a ReLU activation function layer, and a soft pooling layer connected in sequence. The sub-branch has a convolutional layer with a 3 ×3 convolutional kernel and a stride of 2. Each convolutional layer has 128 convolutional kernels.
3. The computer synthetic image detection method based on the combination of channels and soft pooling of a two-stream neural network according to claim 2, characterized in that, The specific process of inputting the image samples into the channel-separated residual extraction module is: S31. Divide the image samples into three channels of R, G, and B and input them into the image processing module in the channel-separated residual extraction module. Use the SRM residual filter in the image processing module to extract the residual features of each channel of the image, and merge the residual features of the image extracted from each channel by channel to obtain the third residual feature map; S32. Input the third residual feature map into the first convolutional block, and then input it into the main branch and the sub-branch of the residual structure in the channel-separated residual extraction module respectively. The third residual feature map passes through the convolutional layer, ReLU activation function layer, and soft pooling layer for downsampling in the main branch in sequence, and the fourth residual feature map is output. The third residual feature map passes through the convolutional layer in the sub-branch to extract the fifth residual feature map with the same scale as that in the main branch. The fourth residual feature map and the fifth residual feature map are fused together by addition, and the sixth residual feature map is output; S33. Activate the data of the sixth residual feature map by using the ReLU activation function, input the output result into the second convolutional block, and then use the second convolutional block to further refine and fuse the residual features of the sixth residual feature map, and the first feature map is output.
4. The computer synthetic image detection method based on channel joint and soft pooling of a two-stream neural network according to claim 3, characterized in that, the specific calculation process of the soft pooling layer for downsampling is: For a pixel region R of size 2 For a pixel region R of size 2, first calculate the SoftMax value of each pixel in this region to obtain ( = 1, 2, 3, 4), and its calculation formula is as follows: wherein represents the pixel value of the th pixel, and Then multiply it element-wise with the original pixel region and sum the four values to obtain the result after pooling , The calculation formula of is as follows: 。 5. The computer synthetic image detection method based on channel joint and soft pooling of a two-stream neural network according to claim 1, characterized in that, The combined channel information extraction module is provided with a third convolutional block, a fourth convolutional block, a fifth convolutional block, a sixth convolutional block, and a seventh convolutional block connected in sequence. Each of the third convolutional block, the fourth convolutional block, the fifth convolutional block, the sixth convolutional block, and the seventh convolutional block is composed of a convolutional layer, a BN layer, a ReLU activation function layer, and a soft pooling layer connected in sequence, where the convolutional layer of the third convolutional block is provided with 32 convolutional kernels, the convolutional layer of the fourth convolutional block is provided with 64 convolutional kernels, and each layer of the convolutional layers of the fifth convolutional block, the sixth convolutional block, and the seventh convolutional block is provided with 128 convolutional kernels. 3 convolutional kernels, a convolutional layer, a BN layer, a ReLU activation function layer, and a soft pooling layer are connected in sequence. Among them, the convolutional layer of the third convolutional block is provided with 32 convolutional kernels, the convolutional layer of the fourth convolutional block is provided with 64 convolutional kernels, and each layer of the convolutional layers of the fifth convolutional block, the sixth convolutional block, and the seventh convolutional block is provided with 128 convolutional kernels.
6. The computer synthetic image detection method based on channel joint and soft pooling of a two-stream neural network according to claim 1, characterized in that, the classifier is composed of a global average pooling layer GAP and a fully connected layer FC. The global average pooling layer GAP and the fully connected layer FC are connected. The specific calculation process of the global average pooling layer is: For a pixel region R of a certain size M N, its calculation formula is as follows: Among them, represents the pixel value of the pixel point at the position of ( ).
7. The processing process of the classifier in the computer synthetic image detection method based on channel joint and soft pooling of a two-stream neural network according to claim 6 is: Let the first feature map output by the residual extraction module be , and the second feature map output by the joint channel information extraction module is denoted as , and satisfy: wherein, C represents the number of channels of the feature map, and H and W respectively represent the height and width of the image sample; Perform global average pooling operation GAP on the first feature map and the second feature map to obtain the feature map GAP( ) , input the features in the feature map GAP( ) into the fully connected layer FC to obtain the output result , add the output results element-wise, and then pass through a SoftMax activation function to obtain a binary probability vector p , the calculation formula of p is: wherein, the two elements of p respectively represent the probability that the category of the image sample is a computer synthetic image and the probability that the category of the image sample is a natural image.
8. The computer synthetic image detection method based on channel joint and soft pooling of a two-stream neural network according to claim 7, characterized in that, in step S4, the parameters and weights of the two-stream neural network model to be trained are updated in a backpropagation manner. The loss function formula used when training the two-stream neural network model is as follows: Among them, N' represents the number of image samples in a Batch, K represents the number of classification categories, and K = 2. represents a Boolean function that returns 1 when the determination is true and 0 when the determination is false. represents the true label of the 9. The computer synthetic image detection method based on channel joint and soft pooling of a two-stream neural network according to claim 1, characterized in that, In step S5, during the process of outputting the classification result of the computer synthesized image sample to be detected, the detection accuracy Acc is used as the evaluation index of the classification result, and its calculation formula can be expressed as the formula: Among them, P represents the number of positive samples, where positive samples refer to natural image samples, and n represents the number of negative samples, where negative samples refer to computer-generated image samples. TP, TN respectively refer to the number of correctly classified positive and negative samples.
Citation Information
Patent Citations
RGBD image joint recovery method based on double-flow network
CN111104532A
Recapture video detection method, system and device based on deep learning, and medium
CN112560734A