Moire pattern image tampering detection method and system based on deep learning

Through deep learning-based methods, the image tamper detection model with molar pattern is generated and trained, and the problem of low accuracy and robustness when processing molar patterns in the prior art is solved, and more efficient and accurate image tamper detection is achieved.

CN120107764APending Publication Date: 2025-06-06SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510163200.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing image tamper detection technology is less accurate and robust when processing images containing molar patterns, making it difficult to effectively identify tampered areas.

Method used

Using a deep learning-based method, a fake image data set with molar patterns is automatically generated, and data augmentation and model training are performed. The specific steps include obtaining the fake image dataset, applying rotation, scaling and color perturbation, inputting the dataset into the built image tamper detection model for training, dynamically adjusting the learning rate and batch size, and iterating the model until it converges.

Benefits of technology

It significantly improves the accuracy and stability of image tamper detection, especially under complex background or molar interference, which can more accurately identify tampering areas, reduce error detection rates, and improve detection automation and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107764A_ABST
    Figure CN120107764A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning-based moire image tampering detection method and system, which adopt an autonomously designed deformable convolution layer to dynamically adjust the shape and size of a convolution kernel, and can more accurately identify a tampered area and significantly improve the detection precision especially under the complex background or moire interference. The improved residual network optimizes the processing capability of the moire image, and by adjusting the hierarchical structure and enhancing the feature transfer mechanism, the false detection rate is reduced and the stability is improved while the high response speed is maintained; and a special data set containing multiple moire types and different tampering levels is created, so that model training better meets actual application requirements, generalization ability and adaptability are improved, and detection automation and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital image tampering detection, and more specifically, to a moiré image tampering detection method and system based on deep learning. Background Art

[0002] With the rapid development of digital image processing technology, image tampering detection has become an important research direction in the field of digital image forensics. Image tampering, especially tampering through advanced editing tools, can achieve convincing authenticity, making it difficult for people to identify the authenticity with the naked eye. Such tampering may be used to mislead public opinion or engage in unfair competition. Therefore, it is particularly important to develop effective image tampering detection technology.

[0003] However, existing image tampering detection technologies mainly target digital images that have not been specially processed. When images contain moire patterns, these textures can significantly affect image quality and mask tampering traces, thus challenging existing tampering detection algorithms. Moire patterns are visual artifacts caused by the mismatch between the spatial frequency of the image acquisition device and the spatial frequency of the object being photographed, and are commonly seen in images obtained by taking a camera to capture a computer or mobile phone screen.

[0004] Traditional image forensics methods, such as image metadata analysis and copy-move detection, are generally effective, but their effectiveness is greatly reduced for images containing moiré. Current advanced tampering detection algorithms also face the problem of low recognition accuracy when processing images containing moiré. These methods usually rely on local features of the image, and the interference of moiré makes it impossible to accurately capture these local features, resulting in reduced detection performance. Summary of the invention

[0005] In order to overcome the defects of low accuracy and robustness in detecting images containing moiré patterns in the prior art, the present invention provides a moiré image tampering detection method and system based on deep learning.

[0006] In order to solve the above technical problems, the technical solution of the present invention is as follows:

[0007] The present invention provides a moiré image tampering detection method based on deep learning, comprising:

[0008] Acquire forged images with moiré patterns through automated generation, and collect real-life images with moiré patterns in real scenes to form a forged image dataset;

[0009] Apply data enhancement of rotation, scaling and color disturbance to the forged image dataset to obtain the preprocessed forged image dataset;

[0010] Inputting the preprocessed forged image dataset into the constructed image tampering detection model for training, dynamically adjusting the learning rate and batch size of the model, iteratively optimizing the image tampering detection model until convergence, and obtaining a trained image tampering detection model;

[0011] Use the trained image tampering detection model to complete image tampering detection.

[0012] Preferably, a forged image with moiré patterns is obtained by an automated generation method, and real images with moiré patterns in real scenes are collected to form a forged image data set, including:

[0013] Obtaining original tampered image data;

[0014] Five forged images with different moiré effects are generated for each original tampered image through frequency domain superposition algorithm;

[0015] All generated fake images are uniformly adjusted to the format of the same dimension to form a fake image dataset.

[0016] Preferably, the preprocessed forged image dataset is input into the constructed image tampering detection model for training, including:

[0017] Constructing an image tampering detection model, the model comprising a residual network, a deformable convolutional layer, a non-local attention convolutional layer, and a dilated spatial pyramid pooling module connected in sequence;

[0018] The images in the preprocessed forged image dataset are sequentially input into the residual network, deformable convolution layer, non-local attention convolution layer and dilated spatial pyramid pooling module, where:

[0019] The residual network extracts basic tampering features of the image;

[0020] The deformable convolution layer adapts to the morphological changes of moiré patterns and tampering traces by dynamically adjusting the sampling position of the convolution kernel;

[0021] The non-local attention convolution layer captures the semantic inconsistency between the tampered area and the real area through global correlation calculation;

[0022] The feature map output by the non-local attention convolutional layer is input into the atrous spatial pyramid pooling module, and the features of different receptive fields are fused for tampering classification.

[0023] Preferably, the residual network comprises a plurality of residual sub-network units connected in sequence;

[0024] The residual sub-network unit includes a first convolutional layer, a second convolutional layer, a third convolutional layer and a GSOP layer which are connected in sequence.

[0025] Preferably, the residual network extracts basic tampering features of the image, including:

[0026] Perform preliminary feature extraction based on the input image to generate the first-layer feature map of the first stage;

[0027] Perform local feature extraction based on the first-layer feature map to generate the second-layer feature map of the first stage;

[0028] The input feature map is added to the second layer feature map through residual connection, and skip connection is performed to obtain the deep feature map;

[0029] Perform channel compression based on the deep feature map to generate the input feature map of the second stage;

[0030] Perform convolution and residual connection according to the input feature map of the second stage to generate the final feature map of the second stage;

[0031] Channel compression is performed based on the output feature map of the second stage to generate the input feature map of the third stage;

[0032] Perform local feature extraction on the input feature map and perform normalization to obtain the intermediate feature map of the third stage;

[0033] Perform residual connection on the input feature map of the third stage and the intermediate feature map to obtain the final feature map of the third stage;

[0034] Perform feature compression based on the final feature map of the third stage to obtain the input feature map of the fourth stage;

[0035] Extract local features and perform normalization operations based on the input feature map of the fourth stage to generate an intermediate feature map of the fourth stage;

[0036] Perform residual concatenation on the input feature map of the fourth stage and the intermediate feature map to obtain the output feature map of the fourth stage;

[0037] The feature information of the output feature map of the fourth stage is gradually extracted to obtain the final feature map of the fourth stage.

[0038] Preferably, the deformable convolution layer includes a local layer, a global layer, a fourth convolution layer, a fifth convolution layer, a sixth convolution layer, a seventh convolution layer, a first summing point, a second summing point, a first function activation layer, and a second function activation;

[0039] The output end of the local layer is connected to the input ends of the fourth convolutional layer and the fifth convolutional layer; the output end of the global layer is connected to the input ends of the sixth convolutional layer and the seventh convolutional layer; the output ends of the fourth convolutional layer and the sixth convolutional layer are connected to the input end of the first summing point; the input end of the fifth convolutional layer and the output end of the seventh convolutional layer are connected to the input end of the second summing point; the output end of the first summing point is connected to the input end of the first function activation layer; the output end of the second summing point is connected to the input end of the second function activation layer.

[0040] Preferably, the deformable convolutional layer comprises:

[0041] The input feature map is divided into a global part and a local part, which are used to process global features and local features respectively;

[0042] The local features are extracted from the local parts through the fourth convolution layer and the fifth convolution layer in turn;

[0043] For the global part, the sixth convolution layer is used to restore the spatial resolution and extract the global features;

[0044] The features of the local part and the global part are additively fused, and the features are enhanced through the first function activation layer and the second function activation layer to obtain the enhanced output feature map.

[0045] Preferably, the non-local attention convolutional layer comprises:

[0046] Reduce the dimension of the original input feature map to the preset channel;

[0047] Perform element-by-element multiplication, activation, and normalization on the reduced-dimensional feature map to generate a synthetic feature map;

[0048] The synthesized feature map is added to the original input feature map to generate the final output feature map.

[0049] Preferably, the GSOP layer includes an eighth convolutional layer, a first pooling layer, a ninth convolutional layer and a tenth convolutional layer connected in sequence.

[0050] The present invention also provides a moiré image tampering detection system based on deep learning, which is used to implement the above method, including:

[0051] The data acquisition module acquires forged images with moiré patterns through automatic generation, and collects real images with moiré patterns in real scenes to form a forged image data set;

[0052] A preprocessing module applies data enhancement such as rotation, scaling and color disturbance to the forged image dataset to obtain a preprocessed forged image dataset;

[0053] A model training module, which inputs the preprocessed forged image data set into the constructed image tampering detection model for training, dynamically adjusts the learning rate and batch size of the model, iteratively optimizes the image tampering detection model until convergence, and obtains a trained image tampering detection model;

[0054] The tampering detection module uses the trained image tampering detection model to complete image tampering detection.

[0055] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0056] The present invention proposes a moiré image tampering detection method and system based on deep learning. The method adopts a self-designed deformable convolution layer and dynamically adjusts the shape and size of the convolution kernel. Especially under complex backgrounds or moiré interference, the tampered area can be identified more accurately and the detection accuracy can be significantly improved. The improved residual network optimizes the processing capability of images with moiré. By adjusting the hierarchical structure and enhancing the feature transfer mechanism, the false detection rate is reduced and the stability is improved while maintaining a high response speed. A special data set containing multiple moiré types and different tampering levels is created to make the model training more in line with actual application needs, improve the generalization ability and adaptability, and improve the detection automation and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a flowchart of the deep learning-based moiré image tampering detection method described in Example 1;

[0058] Figure 2 This is a schematic diagram of the structure of the image tampering detection model described in Example 2;

[0059] Figure 3 Schematic diagram of the structure of the residual network described in Example 2;

[0060] Figure 4 Schematic diagram of the structure of the deformable convolutional layer in Example 2;

[0061] Figure 5 This is a schematic diagram of the structure of the screen capture described in Example 2;

[0062] Figure 6 is a schematic diagram of different moiré intensities described in Example 2;

[0063] Figure 7 This is a schematic diagram of the structure of the deep learning-based moiré image tampering detection system described in Example 3. DETAILED DESCRIPTION

[0064] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0065] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;

[0066] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0067] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0068] Example 1

[0069] This embodiment provides a moiré image tampering detection method based on deep learning. Figure 1 As shown, including:

[0070] Acquire forged images with moiré patterns through automated generation, and collect real-life images with moiré patterns in real scenes to form a forged image dataset;

[0071] Apply data enhancement of rotation, scaling and color disturbance to the forged image dataset to obtain the preprocessed forged image dataset;

[0072] Inputting the preprocessed forged image dataset into the constructed image tampering detection model for training, dynamically adjusting the learning rate and batch size of the model, iteratively optimizing the image tampering detection model until convergence, and obtaining a trained image tampering detection model;

[0073] Use the trained image tampering detection model to complete image tampering detection.

[0074] This embodiment adopts a self-designed deformable convolution layer to dynamically adjust the shape and size of the convolution kernel, especially under complex backgrounds or moiré interference, which can more accurately identify tampered areas and significantly improve detection accuracy; the improved residual network optimizes the processing capabilities of images with moiré, and by adjusting the hierarchical structure and enhancing the feature transfer mechanism, it reduces the false detection rate and improves stability while maintaining a high response speed; creates a dedicated data set containing multiple moiré types and different tampering levels, so that model training is more in line with actual application needs, improves generalization ability and adaptability, and improves detection automation and accuracy.

[0075] Example 2

[0076] This embodiment provides a moiré image tampering detection method based on deep learning, including:

[0077] Acquire forged images with moiré patterns through automated generation, and collect real-life images with moiré patterns in real scenes to form a forged image dataset;

[0078] Apply data enhancement of rotation, scaling and color disturbance to the forged image dataset to obtain the preprocessed forged image dataset;

[0079] Inputting the preprocessed forged image dataset into the constructed image tampering detection model for training, dynamically adjusting the learning rate and batch size of the model, iteratively optimizing the image tampering detection model until convergence, and obtaining a trained image tampering detection model;

[0080] Use the trained image tampering detection model to complete image tampering detection.

[0081] Acquire forged images with moiré patterns through automated generation, and collect real-life images with moiré patterns in real scenes to form a forged image dataset, including:

[0082] Obtaining original tampered image data;

[0083] Five forged images with different moiré effects are generated for each original tampered image through frequency domain superposition algorithm;

[0084] All generated fake images are uniformly adjusted to the format of the same dimension to form a fake image dataset.

[0085] The preprocessed forged image dataset is input into the constructed image tampering detection model for training, including:

[0086] Build an image tampering detection model, such as Figure 2 As shown, the model includes a residual network, a deformable convolution layer, a non-local attention convolution layer and a dilated spatial pyramid pooling module connected in sequence;

[0087] The images in the preprocessed forged image dataset are sequentially input into the residual network, deformable convolution layer, non-local attention convolution layer and dilated spatial pyramid pooling module, where:

[0088] The residual network extracts basic tampering features of the image;

[0089] The deformable convolution layer adapts to the morphological changes of moiré patterns and tampering traces by dynamically adjusting the sampling position of the convolution kernel;

[0090] The non-local attention convolution layer captures the semantic inconsistency between the tampered area and the real area through global correlation calculation;

[0091] The feature map output by the non-local attention convolutional layer is input into the atrous spatial pyramid pooling module, and the features of different receptive fields are fused for tampering classification.

[0092] like Figure 3 As shown, the residual network includes a plurality of residual sub-network units connected in sequence;

[0093] The residual sub-network unit includes a first convolutional layer, a second convolutional layer, a third convolutional layer and a GSOP layer which are connected in sequence.

[0094] The residual network extracts basic tampering features of the image, including:

[0095] Perform preliminary feature extraction based on the input image to generate the first-layer feature map of the first stage;

[0096] Perform local feature extraction based on the first-layer feature map to generate the second-layer feature map of the first stage;

[0097] The input feature map is added to the second layer feature map through residual connection, and skip connection is performed to obtain the deep feature map;

[0098] Perform channel compression based on the deep feature map to generate the input feature map of the second stage;

[0099] Perform convolution and residual connection according to the input feature map of the second stage to generate the final feature map of the second stage;

[0100] Channel compression is performed based on the output feature map of the second stage to generate the input feature map of the third stage;

[0101] Perform local feature extraction on the input feature map and perform normalization to obtain the intermediate feature map of the third stage;

[0102] Perform residual connection on the input feature map of the third stage and the intermediate feature map to obtain the final feature map of the third stage;

[0103] Perform feature compression based on the final feature map of the third stage to obtain the input feature map of the fourth stage;

[0104] Extract local features and perform normalization operations based on the input feature map of the fourth stage to generate an intermediate feature map of the fourth stage;

[0105] Perform residual concatenation on the input feature map of the fourth stage and the intermediate feature map to obtain the output feature map of the fourth stage;

[0106] The feature information of the output feature map of the fourth stage is gradually extracted to obtain the final feature map of the fourth stage.

[0107] like Figure 4 As shown, the deformable convolution layer includes a local layer, a global layer, a fourth convolution layer, a fifth convolution layer, a sixth convolution layer, a seventh convolution layer, a first summing point, a second summing point, a first function activation layer, and a second function activation;

[0108] The output end of the local layer is connected to the input ends of the fourth convolutional layer and the fifth convolutional layer; the output end of the global layer is connected to the input ends of the sixth convolutional layer and the seventh convolutional layer; the output ends of the fourth convolutional layer and the sixth convolutional layer are connected to the input end of the first summing point; the input end of the fifth convolutional layer and the output end of the seventh convolutional layer are connected to the input end of the second summing point; the output end of the first summing point is connected to the input end of the first function activation layer; the output end of the second summing point is connected to the input end of the second function activation layer.

[0109] The deformable convolutional layer comprises:

[0110] The input feature map is divided into a global part and a local part, which are used to process global features and local features respectively;

[0111] The local features are extracted from the local parts through the fourth convolution layer and the fifth convolution layer in turn;

[0112] For the global part, the sixth convolution layer is used to restore the spatial resolution and extract the global features;

[0113] The features of the local part and the global part are additively fused, and the features are enhanced through the first function activation layer and the second function activation layer to obtain the enhanced output feature map.

[0114] The non-local attention convolutional layer includes:

[0115] Reduce the dimension of the original input feature map to the preset channel;

[0116] Perform element-by-element multiplication, activation, and normalization on the reduced-dimensional feature map to generate a synthetic feature map;

[0117] The synthesized feature map is added to the original input feature map to generate the final output feature map.

[0118] The GSOP layer includes an eighth convolutional layer, a first pooling layer, a ninth convolutional layer and a tenth convolutional layer which are connected in sequence.

[0119] In a specific embodiment, it includes:

[0120] S1. Obtaining a dataset: Based on a public dataset, a dataset of forged images with moiré patterns is obtained in an automated manner, and a dataset of forged images with moiré patterns in real scenes is constructed using real-life shooting.

[0121] S2. Use the improved ResNet50 network for deep learning: The preprocessed image is input into the improved ResNet50 network for feature extraction. This network structure has been experimentally verified to have the best tampering detection performance in complex image backgrounds with moiré patterns.

[0122] S3. Application of self-designed deformable convolution modules: Self-designed deformable convolution modules are used to extract features on the feature maps extracted by the improved residual network. These modules can adaptively adjust their shape and size to adapt to the specific content in the image, especially effectively identifying and processing the interaction of moiré and tampering traces.

[0123] S4. Apply non-local modules: Non-local modules are used for feature extraction on the feature maps extracted by the deformable convolution module. These modules can capture the inconsistency between the tampered area and the real area.

[0124] S5. Apply the atrous spatial pyramid pooling module to perform tampering classification detection on the output feature map to detect whether the image with moiré patterns has been tampered with.

[0125] S6. Training and validating the model: Split the dataset into a training set and a test set. Apply data augmentation techniques such as rotation, scaling, and color adjustment to increase the robustness of the model. Then, train the model on the training set and validate the detection performance of the model on the test set.

[0126] S7. Refine model performance: After model training and initial testing, refine the network and adjust parameters such as learning rate and batch size to optimize performance. Continue training until the model achieves optimal tamper detection results.

[0127] S8. Apply the model for actual tampering detection: The final model will be deployed in actual image forensics scenarios to detect potential image tampering incidents. The model is particularly suitable for handling image tampering detection in high-risk environments, such as legal forensics and media content verification.

[0128] The step S1 comprises the following sub-steps:

[0129] S11. Obtain the CASIA2 tampered image dataset as the original tampered image dataset

[0130] S12. Add moiré to the image by using code to obtain a dataset of tampered images with moiré, wherein five different moiré effects are added to each original tampered image.

[0131] S13. Preprocess the tampered image dataset with moiré patterns, and obtain an image dataset with dimensions of C×H×W after size transformation as a preprocessed training set, where C represents the number of image channels, H represents the image height, and W represents the image width.

[0132] The step S2 comprises the following sub-steps:

[0133] S201, using a convolution layer with a convolution kernel of 7×7 and a step size of 2 to process the input image, perform preliminary feature extraction, and obtain a first-layer feature map;

[0134] S202, using a convolution layer with a convolution kernel of 3×3 and a step size of 1 to process the first layer feature map, further extract local features, and obtain a second layer feature map;

[0135] S203: Use residual connection to add the input feature map to the second layer feature map to obtain a feature map after residual learning, thereby improving the efficiency of information transmission.

[0136] S204, processing the second layer feature map through multiple residual modules, each module including multiple 3×3 convolutional layers and batch normalization operations, and combining with jump connections, to obtain a deep feature map after residual learning, and pass it to the next stage for further processing;

[0137] S205, passing the output feature map of the first stage to stage 2, using a convolution layer with a convolution kernel of 1×1 to perform feature compression, reducing the number of channels, and obtaining the input feature map of stage 2;

[0138] S206, using multiple 3×3 convolutional layers to process the input feature map, further extract local features, and perform batch normalization operations to obtain the intermediate feature map of stage 2;

[0139] S207, performing a residual connection on the input feature map and the intermediate feature map to obtain a residual output feature map of stage 2;

[0140] S208. Continue to use multiple residual modules to process the output feature map of stage 2, gradually extract more advanced feature information, and obtain the final feature map of stage 2.

[0141] S209, passing the output feature map of stage 2 to stage 3, using a convolution layer with a convolution kernel of 1×1 to perform feature compression, reducing the number of channels, and obtaining the input feature map of stage 3;

[0142] S210, using multiple 3×3 convolutional layers to process the input feature map, further extract local features, and perform batch normalization operations to obtain the intermediate feature map of stage 3;

[0143] S211, performing residual connection on the input feature map and the intermediate feature map to obtain the residual output feature map of stage 3;

[0144] S212. Continue to use multiple residual modules to process the output feature map of stage 3, gradually extract more advanced feature information, and obtain the final feature map of stage 3.

[0145] S213, passing the output feature map of stage 3 to stage 4, using a convolution layer with a convolution kernel of 1×1 to perform feature compression, reducing the number of channels, and obtaining the input feature map of stage 4;

[0146] S214, using multiple 3×3 convolutional layers to process the input feature map, further extract local features, and perform batch normalization operations to obtain the intermediate feature map of stage 4;

[0147] S215, performing a residual connection on the input feature map and the intermediate feature map to obtain a residual output feature map of stage 4;

[0148] S216. Continue to use multiple residual modules to process the output feature map of stage 4, gradually extract more advanced feature information, and obtain the final feature map of stage 4.

[0149] The method for processing the output feature map of each stage of the ResNet50 network by the deformable convolution module in step S3 comprises the following steps:

[0150] S31, divide the input feature map into two parts, local and global, and process local features and global features respectively;

[0151] S32, in the local part, the input image is firstly subjected to feature extraction by two 3×3 convolutional layers (Conv 3×3) to extract local features;

[0152] S33, in the global part, the input image is deconvolved through the Deconv 3×3 convolutional layer to restore the spatial resolution of the image and extract global features;

[0153] S34, the processing results of the local part and the global part are fused through the addition operation (+), combining the local and global information to generate a richer feature representation;

[0154] S35, the fused features are batch normalized and activated through the BN-ReLU layer to enhance the nonlinear expression ability of the features;

[0155] S36. Finally, the local and global processed features are added (+) to generate an output feature map.

[0156] S37, the output feature map will be passed to the subsequent processing layer for further feature analysis or task application.

[0157] The method for processing the output feature map of the deformable convolution module at each stage by the non-local module in step S4 comprises the following steps:

[0158] S41. The size of the input feature map X of this module is (H, W, 1024).

[0159] S42, input feature map X passes through a 1×1 convolution layer and weight θ, and the output feature map size is (H, W, 512);

[0160] S43, input feature map X through 1×1 convolution layer and weight The output feature map size is (H, W, 512);

[0161] S44, performing element-by-element multiplication (×) on the feature maps obtained in steps [S42] and [S43], and then inputting the result into a softmax function, and obtaining a synthetic feature map after normalization;

[0162] S45, multiplying the synthesized feature map by the weight g element by element to obtain a new feature map;

[0163] S46: Convolve the feature map obtained in step [S45] through a 1×1×1 convolutional layer to obtain a final feature map with a size of (H, W, 1024)

[0164] S47, the final feature map is added to the input feature map X to obtain the final output result Z

[0165] The method for processing the feature map finally output by the hollow space pyramid pooling module in step S5 comprises the following steps:

[0166] S51. The input image is processed by a 1×1 convolution layer (1X1 Conv) to adjust the number of channels and obtain a preliminary feature map with a size of (H, W, N), where N is the number of channels after adjustment.

[0167] S52, then, the input image is processed by multiple dilated convolution operations to extract feature information of different scales. The first layer uses 3×3 dilated convolution (rate 6), the second layer uses 3×3 dilated convolution (rate 12), and the third layer uses 3×3 dilated convolution (rate 18). All dilated convolution operations are performed simultaneously, and each layer of convolution operations extracts feature information of different scales in the image respectively.

[0168] S53, after multiple dilated convolution operations, a pyramid pooling operation is used for further feature extraction;

[0169] S54. Finally, the pooled feature map is input into the fully connected layer for tampering classification.

[0170] The step S6 comprises the following sub-steps:

[0171] S61. Divide the data set into a training set and a test set for model training and verification. The ratio of the training set to the test set is 8:2.

[0172] S62. To increase the robustness of the model, data enhancement techniques are applied during the training process, including rotation, scaling, color adjustment, etc.

[0173] S63, training the tampering detection model on the training set, and optimizing the parameters of the model so that the model can efficiently learn and identify tampered images on the training set;

[0174] S64. Verify the trained model on the test set and evaluate the detection performance of the model on unseen data. Use the accuracy index to verify the model's ability to detect tampered images and ensure its generalization and robustness.

[0175] The step S7 comprises the following sub-steps:

[0176] S71. After model training and preliminary testing, refine the model's performance to optimize its tamper detection effect.

[0177] S72. Further improve the performance of the model by fine-tuning the model parameters, including but not limited to the following key parameters:

[0178] Learning rate: adjust the learning rate to optimize the update speed of model parameters to avoid overfitting or insufficient training;

[0179] Batch size: adjust the number of samples processed in each training to optimize the stability and speed of the model during training.

[0180] S73. In the refinement process, the key parameters in the modules are fine-tuned in combination with the output feature maps of the aforementioned modules (such as the deformable convolution module, the non-local module, the spatial void pyramid pooling module, etc.), specifically including:

[0181] Adjust the ratio of local and global parts of the deformable convolution module;

[0182] Use different attention mechanisms to compare with non-local modules;

[0183] Adjust the hole size of the atrous convolution in the spatial atrous pyramid pooling module.

[0184] S74. Conduct experimental comparisons based on different parameter configurations to compare the model detection performance under each configuration to ensure that the optimal parameter combination is selected to achieve the best tampering detection effect.

[0185] S75. Continue training until the model achieves the best tampering detection effect on the test set, that is, achieves the best tampering detection accuracy, ensuring the efficiency and accuracy of the model in practical applications.

[0186] The step S8 comprises the following sub-steps:

[0187] S81. In order to verify the effectiveness of the proposed tampering detection model in practical application scenarios, a series of experiments were designed, such as Figure 5 As shown, two different brands of mobile phones (Honor Magic4 and Apple iPhone 13Pro) were used to shoot two different brands of monitors (Lenovo monitor and Apple monitor), and a total of four combinations of image data were obtained.

[0188] S82. The image data of the four combinations mentioned above represent common shooting scenes in daily life, covering the differences between different devices. The specific combinations are: Honor Magic4 and Lenovo display (A), Honor Magic4 and Apple display (B), iPhone 13Pro and Lenovo display (C), iPhone 13Pro and Apple display (D)

[0189] S83, the screen brightness is uniformly set to normal brightness, and the shooting time is from 12 noon to 2 o'clock.

[0190] S84, the image data of the four combinations were used as test sets to evaluate the proposed tampering detection model, specifically evaluating the detection performance of the model under different shooting devices and display combinations. The experimental results are shown in Table 5, where the training sets with different moiré effects were used to train the model.

[0191] S85. Experimental results show that if Figure 6 As shown in the figure, the model shows good detection effect in all four scenarios, which is manifested in high accuracy.

[0192] The experimental platform is NVIDIA4080GPU, python3.7 development language, PyTorch1.13.0 deep learning framework, and AdamW optimizer.

[0193] Comparison of tamper detection effects

[0194] Table 1 shows the detection results of the network for forged image datasets without moiré and forged image datasets with moiré. ResNet50, DenseNet121, Vgg16, InceptionV3, ViT and SwimTransformer were selected. It can be seen that our model surpasses the existing mainstream technology, which proves the effectiveness of the invention.

[0195] Table 1 Comparison with mainstream detection methods

[0196]

[0197] Table 2 Effects of different void ratios

[0198]

[0199] Table 3 Ablation experiments for different attention mechanisms

[0200]

[0201] Table 4 Ablation experiment for deformable convolution module (Ratio represents the weight of ordinary convolution)

[0202]

[0203] Tables 2, 3 and 4 show the detection results under different combinations, which illustrate that the deformable convolution module and the improved ResNet network proposed in the present invention are effective.

[0204] Table 5 Moire tampering detection in different real scenarios

[0205]

[0206] Table 5 shows that the detection effect of forged images with moiré patterns in real-life scenes is more than 70%, which proves the effectiveness of the present invention.

[0207] Example 3

[0208] This embodiment also provides a moiré image tampering detection system based on deep learning, which is used to implement the method of embodiment 1 or embodiment 2, such as Figure 7 As shown, including:

[0209] The data acquisition module acquires forged images with moiré patterns through automatic generation, and collects real images with moiré patterns in real scenes to form a forged image data set;

[0210] A preprocessing module applies data enhancement such as rotation, scaling and color disturbance to the forged image dataset to obtain a preprocessed forged image dataset;

[0211] A model training module, which inputs the preprocessed forged image data set into the constructed image tampering detection model for training, dynamically adjusts the learning rate and batch size of the model, iteratively optimizes the image tampering detection model until convergence, and obtains a trained image tampering detection model;

[0212] The tampering detection module uses the trained image tampering detection model to complete image tampering detection.

[0213] The same or similar reference numerals correspond to the same or similar components;

[0214] The terms used in the drawings to describe positional relationships are only used for illustrative purposes and should not be construed as limiting this patent;

[0215] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A moiré image tampering detection method based on deep learning, characterized in that: include: Acquire forged images with moiré patterns through automated generation, and collect real-life images with moiré patterns in real scenes to form a forged image dataset; Apply data enhancement of rotation, scaling and color disturbance to the forged image dataset to obtain the preprocessed forged image dataset; Inputting the preprocessed forged image dataset into the constructed image tampering detection model for training, dynamically adjusting the learning rate and batch size of the model, iteratively optimizing the image tampering detection model until convergence, and obtaining a trained image tampering detection model; Use the trained image tampering detection model to complete image tampering detection.

2. The deep learning-based moiré image tampering detection method according to claim 1, characterized in that: Acquire forged images with moiré patterns through automated generation, and collect real-life images with moiré patterns in real scenes to form a forged image dataset, including: Obtaining original tampered image data; Five forged images with different moiré effects are generated for each original tampered image through frequency domain superposition algorithm; All generated fake images are uniformly adjusted to the format of the same dimension to form a fake image dataset.

3. The deep learning-based moiré image tampering detection method according to claim 1, characterized in that: The preprocessed forged image dataset is input into the constructed image tampering detection model for training, including: Constructing an image tampering detection model, the model comprising a residual network, a deformable convolutional layer, a non-local attention convolutional layer, and a dilated spatial pyramid pooling module connected in sequence; The images in the preprocessed forged image dataset are sequentially input into the residual network, deformable convolution layer, non-local attention convolution layer and dilated spatial pyramid pooling module, where: The residual network extracts basic tampering features of the image; The deformable convolution layer adapts to the morphological changes of moiré patterns and tampering traces by dynamically adjusting the sampling position of the convolution kernel; The non-local attention convolution layer captures the semantic inconsistency between the tampered area and the real area through global correlation calculation; The feature map output by the non-local attention convolutional layer is input into the atrous spatial pyramid pooling module, and the features of different receptive fields are fused for tampering classification.

4. The deep learning-based moiré image tampering detection method according to claim 3 is characterized in that: The residual network includes a plurality of residual sub-network units connected in sequence; The residual sub-network unit includes a first convolutional layer, a second convolutional layer, a third convolutional layer and a GSOP layer which are connected in sequence.

5. The deep learning-based moiré image tampering detection method according to claim 3, characterized in that: The residual network extracts basic tampering features of the image, including: Perform preliminary feature extraction based on the input image to generate the first-layer feature map of the first stage; Perform local feature extraction based on the first-layer feature map to generate the second-layer feature map of the first stage; The input feature map is added to the second layer feature map through residual connection, and skip connection is performed to obtain the deep feature map; Perform channel compression based on the deep feature map to generate the input feature map of the second stage; Perform convolution and residual connection according to the input feature map of the second stage to generate the final feature map of the second stage; Channel compression is performed based on the output feature map of the second stage to generate the input feature map of the third stage; Perform local feature extraction on the input feature map and perform normalization to obtain the intermediate feature map of the third stage; Perform residual connection on the input feature map of the third stage and the intermediate feature map to obtain the final feature map of the third stage; Perform feature compression based on the final feature map of the third stage to obtain the input feature map of the fourth stage; Extract local features and perform normalization operations based on the input feature map of the fourth stage to generate an intermediate feature map of the fourth stage; Perform residual concatenation on the input feature map of the fourth stage and the intermediate feature map to obtain the output feature map of the fourth stage; The feature information of the output feature map of the fourth stage is gradually extracted to obtain the final feature map of the fourth stage.

6. The deep learning-based moiré image tampering detection method according to claim 3, characterized in that: The deformable convolution layer includes a local layer, a global layer, a fourth convolution layer, a fifth convolution layer, a sixth convolution layer, a seventh convolution layer, a first summing point, a second summing point, a first function activation layer, and a second function activation; The output end of the local layer is connected to the input ends of the fourth convolutional layer and the fifth convolutional layer; the output end of the global layer is connected to the input ends of the sixth convolutional layer and the seventh convolutional layer; the output ends of the fourth convolutional layer and the sixth convolutional layer are connected to the input end of the first summing point; the input end of the fifth convolutional layer and the output end of the seventh convolutional layer are connected to the input end of the second summing point; the output end of the first summing point is connected to the input end of the first function activation layer; the output end of the second summing point is connected to the input end of the second function activation layer.

7. The deep learning-based moiré image tampering detection method according to claim 6, characterized in that: The deformable convolutional layer comprises: The input feature map is divided into a global part and a local part, which are used to process global features and local features respectively; The local features are extracted from the local parts through the fourth convolution layer and the fifth convolution layer in turn; For the global part, the sixth convolution layer is used to restore the spatial resolution and extract the global features; The features of the local part and the global part are additively fused, and the features are enhanced through the first function activation layer and the second function activation layer to obtain the enhanced output feature map.

8. The deep learning-based moiré image tampering detection method according to claim 3, characterized in that: The non-local attention convolutional layer includes: Reduce the dimension of the original input feature map to the preset channel; Perform element-by-element multiplication, activation, and normalization on the reduced-dimensional feature map to generate a synthetic feature map; The synthesized feature map is added to the original input feature map to generate the final output feature map.

9. The deep learning-based moiré image tampering detection method according to claim 4, characterized in that: The GSOP layer includes an eighth convolutional layer, a first pooling layer, a ninth convolutional layer and a tenth convolutional layer which are connected in sequence.

10. A moiré image tampering detection system based on deep learning, used to implement the method of claims 1-9, characterized in that: include: The data acquisition module acquires forged images with moiré patterns through automatic generation, and collects real images with moiré patterns in real scenes to form a forged image data set; A preprocessing module applies data enhancement such as rotation, scaling and color disturbance to the forged image dataset to obtain a preprocessed forged image dataset; A model training module, which inputs the preprocessed forged image data set into the constructed image tampering detection model for training, dynamically adjusts the learning rate and batch size of the model, iteratively optimizes the image tampering detection model until convergence, and obtains a trained image tampering detection model; The tampering detection module uses the trained image tampering detection model to complete image tampering detection.