An image restoration forensics method based on feature-enhanced neural network

Through the image repair and forensics method based on feature enhancement neural network, the problem of difficulty in detecting and positioning of digital image repair areas in the prior art is solved, and accurate positioning of the repair areas and robustness to post-processing operations are achieved.

CN114066754BActive Publication Date: 2025-05-02NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202111327201.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-05-02
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and locate repair areas in digital images, especially when faced with post-processing operations such as JPEG compression, rotation and scaling, which lacks robustness.

Method used

The image repair and evidence forensics method based on feature enhancement neural network is adopted. By constructing feature enhancement network modules, downsampling and upsampling networks, combined with high-pass filters and convolution operations, the model is trained to identify and locate the repair areas in the image.

Benefits of technology

It realizes accurate and effective positioning of repair areas in digital images, and demonstrates robustness for post-processing operations such as JPEG compression, rotation and scaling, improving the performance of traditional repair algorithms forensics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114066754B_ABST
    Figure CN114066754B_ABST
Patent Text Reader

Abstract

The present invention discloses an image restoration forensics method based on feature enhancement neural network, which pre-processes images that have been processed by traditional restoration algorithms, and constructs training sets and test sets based on the pre-processed images and corresponding labels; constructs feature enhancement network modules based on high-pass filters and convolution operations; builds downsampling and upsampling networks; uses the constructed training sets and corresponding label sets to train the designed models; uses the saved best model weights to predict the images of the test set to find the restored areas in the images. The present invention demonstrates the most advanced performance of traditional restoration algorithm forensics, can accurately and effectively locate the restored areas in digital images, and is robust to post-processing operations such as JPEG compression, rotation and scaling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing and information security, and in particular relates to an image restoration and forensics method based on a feature-enhanced neural network. Background Art

[0002] In the fast-developing information age, with the birth of advanced digital devices and image editing software, image processing has become easier and easier, allowing users to edit images without professional knowledge. As a powerful image processing technology, image inpainting can reconstruct missing areas in digital images or repair damaged digital photos in a visually plausible way. It can be divided into two categories: traditional inpainting and deep learning-based inpainting. The former fills the target holes with appropriate background content of the same image, while the latter can generate new objects through neural networks. In addition, there are many traditional inpainting methods combined with block-based and diffusion-based techniques. However, if these methods are used by attackers for malicious purposes to generate forged images, such as removing key objects or information in the image, people have a wrong understanding of the image content. Once these tampered images circulate on the Internet, they are likely to cause huge adverse effects. Therefore, verifying the authenticity and integrity of digital images has become an indispensable task.

[0003] There are two main categories of traditional inpainting algorithms: diffusion-based inpainting and block-based inpainting. For block-based inpainting forensics, many existing techniques mainly detect inpainted images based on the fact that the filled inpainted blocks are copied from the lossless areas of the same image. Therefore, these forensic algorithms mainly contain two main processes: retrieval of suspicious areas and identification of forged areas. At first, Wu et al. proposed a sample-based inpainting forensics method, which first applied zero connectivity marking in the suspicious area to calculate the matching features of all blocks, then calculated the fuzzy function of the matching features, and finally identified the tampered area through the cut set. However, this method requires manual selection of suspicious areas in advance, and it takes a long time to search for similar areas. Since then, many forensic methods have been optimized based on this algorithm, such as automatically detecting suspicious areas or reducing the calculation time of the algorithm. Until recently, Zhu et al. used the powerful learning ability of deep learning to design a network for detecting and locating inpainted images. In addition, as a pioneering attempt at diffusion-based inpainting forensics, Li et al. disclosed a method for locating inpainted areas by analyzing the local variance of the image Laplacian operator along the iso-ray direction. However, these methods are not suitable for unified detection of various traditional inpainting algorithms. Summary of the invention

[0004] Purpose of the invention: In view of the problems existing in the prior art, the present invention proposes an image restoration forensics method based on feature enhancement neural network, which can accurately and effectively locate the restoration area in the digital image and is robust to post-processing operations such as JPEG compression, rotation and scaling.

[0005] Technical solution: The image restoration and forensics method based on feature-enhanced neural network described in the present invention comprises the following steps:

[0006] (1) Preprocess the images that have been processed by the traditional restoration algorithm, and construct training sets and test sets based on the preprocessed images and corresponding labels;

[0007] (2) Construct a feature enhancement network module based on high-pass filters and convolution operations;

[0008] (3) Construct downsampling and upsampling networks,

[0009] (4) Use the constructed training set and the corresponding label set to train the designed network;

[0010] (5) Use the best model weights trained by the network to predict the images in the test set and find the areas in the image that have been repaired.

[0011] Furthermore, the preprocessing in step (1) is resizing, flipping in any direction, distorting, and randomly generating a mask to cover the image.

[0012] Furthermore, the implementation process of step (2) is as follows:

[0013] The feature enhancement network module first obtains the image noise residual domain through a high-pass filter with 5 filter kernels, and uses conventional convolution on the input image to obtain color features; then the image noise residual domain and color features are fused to obtain enhanced features; finally, a conventional convolution is used to learn more representative features; the feature enhancement network module contains a total of 8 networks, 5 3×3 depth-separable convolutions, 2 3×3 convolutions and 1 fusion layer, and the step size of the convolution layer is 1; among the filter kernels, 4 are SRM kernels and one is a Laplacian kernel; the high-pass filter kernel is set to the depth-separable convolution initial kernel, and the filter will perform 5 depth-separable convolution operations, each convolution using a different kernel; use the RGB image as the input of the high-pass filter to obtain 15 noise feature maps, and at the same time, use conventional convolution on the output image to obtain 3 color feature maps, and then fuse the color features and noise residuals to obtain an output of 18 channels, and use a 3×3 convolution with a step size of 1 for the fused features to generate 32 results.

[0014] Furthermore, the implementation process of step (3) is as follows:

[0015] VGG-16 is used as the downsampling network. The downsampling part contains 17 network layers: 13 3×3 convolutions and 4 maximum pooling operations, each convolution layer is followed by an activation function and BN layer; the upsampling network contains a total of 16 network layers: 1 1×1 convolution, 8 3×3 convolutions, 4 upsampling layers and 3 fusion layers; each 3×3 convolution is followed by an activation function (ReLU) and BN layer, the upsampling kernel size and step size are 2; the step size of all convolutions is 1; the upsampling network is divided into four stages, each stage consists of a convolution layer and an upsampling layer, the number of convolution kernels in each stage is reduced by half, and is fused with the features of the downsampling network in the first, third and fourth stages; the last stage outputs a prediction result with 64 channels, which is sent to a 1×1 convolution, and the Softmax function is used for binary classification to determine whether each pixel has been repaired.

[0016] Furthermore, the implementation process of step (4) is as follows:

[0017] In order to reduce the weight of easy-to-classify samples and make the model pay more attention to difficult-to-classify samples:

[0018] FL(p t )=-(1-p t ) γ log(p t )

[0019] Among them, FL is the loss function, γ is a constant, and p t It is obtained by the following formula:

[0020]

[0021] Among them, y∈{0,1} represents the true category, and p∈[0,1] represents the model's estimated probability of the class with label y=1.

[0022] Beneficial effects: Compared with the prior art, the present invention demonstrates the most advanced performance of traditional restoration algorithm forensics, can accurately and effectively locate the restored areas in digital images, and is robust to post-processing operations such as JPEG compression, rotation and scaling. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a schematic diagram of a feature-enhanced neural network constructed in the present invention;

[0024] Figure 2are images at different stages, where (a) is a content-defective image in the dataset; (b) is the corresponding label image; (c) is the restored image; and (d) is the final effect image. DETAILED DESCRIPTION

[0025] The present invention is further described in detail below with reference to the accompanying drawings.

[0026] The present invention proposes an image restoration forensics method based on feature enhancement neural network, which specifically includes the following steps:

[0027] Step 1: Preprocess the images that have been processed by the traditional restoration algorithm, and construct training sets and test sets based on the preprocessed images and corresponding labels.

[0028] 50,000 different images of a fixed size of 256×256 are randomly selected from the Places database to create synthetic inpainted images. First, missing areas are generated in these images, which are located in the center of the image, and the tampered area occupies 10% to 12% of the entire image. Considering the randomness of the data, the shape of the missing area in the image is random, including rectangles, circles, irregular shapes, etc. Then, the missing area is inpainted on the image to generate the inpainted dataset. Finally, the inpainted images are divided into two subsets: 48,000 images and corresponding label images are used for the network, and the remaining 2,000 images and corresponding label images are used for verification. In order to obtain better prediction results, other image processing operations are also performed on the images, including resizing, flipping in any direction, distortion, and randomly generating masks to cover the images, and then these data and labels are input into the model for training. In addition, 250 images are randomly selected from the Places, ImageNet, CelebA, and UCID databases to generate inpainted images. There are a total of 1,000 inpainted images with corresponding labels in the test dataset. In order to better test the performance of the network, the tampered position and shape in the image are random, and the test data set is mixed with four images with different repair rates, namely 5%, 10%, 15% and 20%. In addition, in order to verify the robustness of the model to JPEG compression, rotation and scaling, the repaired images are post-processed accordingly.

[0029] Step 2: Use high-pass filters and convolution operations to construct a feature enhancement network module.

[0030] In the feature enhancement block, a high-pass filter with 5 filter kernels is designed to obtain the image noise residual domain, and a regular convolution is performed on the input image to obtain the color features; then the image noise residual domain and the color features are fused to obtain enhanced features; finally, a regular convolution is used to learn more representative features; this module contains a total of 8 networks: 5 3×3 depthwise separable convolutions, 2 3×3 convolutions and 1 fusion layer, where the step size of the convolution layer is 1.

[0031] The high-pass filter has 5 filter kernels, 4 of which are SRM (Spatial Rich Model) kernels and the other is a Laplacian kernel. Among them, SRM can effectively expose block-based noise inconsistencies in complex texture areas, while the Laplacian kernel has the ability to expose diffusion-based repair noise inconsistencies. A depthwise convolution with a kernel size of 3 and a stride of 1 is applied to perform calculations on each channel of the input layer independently. In the feature enhancement module, the high-pass filter kernel is set as the depthwise separable convolution initial kernel, so the filter will perform 5 depthwise separable convolution operations, each using a different kernel. In addition, the RGB image is used as the input of the high-pass filter to obtain 15 noise feature maps. At the same time, regular convolution is used on the output image to obtain 3 color feature maps. Then, the color features and noise residuals are fused to obtain an output of 18 channels. Finally, a 3×3 convolution with a stride of 1 is used on the fused features to generate 32 results.

[0032] Step 3: Build downsampling and upsampling networks.

[0033] The downsampling part contains a total of 17 network layers: 13 3×3 convolutions and 4 maximum pooling operations, each convolution layer is followed by an activation function (ReLU) and BN layer; the upsampling network contains a total of 16 network layers: 1 1×1 convolution, 8 3×3 convolutions, 4 upsampling layers and 3 fusion layers; each 3×3 convolution is followed by an activation function (ReLU) and BN layer, the upsampling kernel size and step size are 2; the step size of all convolutions is 1.

[0034] For using VGG-16 as the downsampling network, the network contains a total of 17 layers: 13 3×3 convolutions and 4 maximum pooling operations. Each convolution layer is followed by an activation function (ReLU). Specifically, the feature extraction module can be divided into five stages. The first two stages are composed of two consecutive convolution layers and one maximum pooling layer, while the third and fourth stages are composed of three convolutions and one maximum pooling layer, and three consecutive convolution layers are configured in the fifth stage. The number of output feature maps in the first stage is 64. Except for the last stage, the number of output channels in each stage doubles, with a maximum of 512 channels. The number of feature maps in the last stage is the same as that in stage 4. The upsampling structure is divided into four stages, such as Figure 1 As shown in the figure, from bottom to top, each stage consists of a convolution layer and an upsampling layer. The number of convolution kernels in each stage is reduced by half, and the features of the downsampling network are fused in the first, third and fourth stages. The last stage outputs a 64-channel prediction result, which is then sent to a 1×1 convolution and a Softmax function is used for binary classification to determine whether each pixel has been repaired.

[0035] Step 4: Use the constructed training set and the corresponding label set to train the designed model.

[0036] The present invention is implemented using the TensorFlow deep learning framework. In the training phase, a 1×10 -4 The Adam optimizer with the initial learning rate is used to update and calculate the network parameters that affect the model training and output, and the learning rate will be reduced by 8% after each training round. In all convolutional layers, the kernel weights are initialized using the He normal initialization method. The entire network was trained for 20 rounds, and the batch size was set to 16 to increase the training speed. Finally, when the validation loss value converges to the minimum value, the best weights of the model are saved for testing.

[0037] The training uses the following formula to reduce the weights of easy-to-classify samples and make the model pay more attention to difficult-to-classify samples:

[0038] FL(p t )=-(1-p t ) γ log(p t )

[0039] Among them, FL is the loss function, γ is a constant, which is set to 2 here, and p t It is obtained by the following formula:

[0040]

[0041] Among them, y∈{0,1} represents the true category, and p∈[0,1] represents the model's estimated probability of the class with label y=1.

[0042] Step 5: Use the saved best model weights to make predictions for each image in the test set and find the areas that have been restored.

[0043] like Figure 2 As shown, the present invention has achieved remarkable results in the tampering positioning of the traditional repair algorithm, which not only reduces the false alarm rate but also improves the accuracy and efficiency of positioning, and meets the robustness of JPEG compression, scaling and rotation post-processing.

[0044] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. An image restoration forensics method based on feature-enhanced neural network, characterized in that: The steps include: (1) Preprocess the images that have been processed by the traditional restoration algorithm, and construct training sets and test sets based on the preprocessed images and corresponding labels; (2) Construct a feature enhancement network module based on high-pass filters and convolution operations; (3) Construct downsampling and upsampling networks, (4) Use the constructed training set and the corresponding label set to train the designed network; (5) Use the best model weights trained by the network to predict the images in the test set and find the areas in the image that have been repaired; The implementation process of step (2) is as follows: The feature enhancement network module first obtains the image noise residual domain through a high-pass filter with 5 filter kernels, and performs regular convolution on the input image to obtain color features; then the image noise residual domain and color features are fused to obtain enhanced features; finally, a regular convolution is used to learn more representative features; The feature enhancement network module contains a total of 8 networks, 5 3×3 depth-separable convolutions, 2 3×3 convolutions and 1 fusion layer, and the stride of the convolution layer is 1; among the filter kernels, 4 are SRM kernels and the other is a Laplacian kernel; the high-pass filter kernel is set to the depth-separable convolution initial kernel, and the filter will perform 5 depth-separable convolution operations, each using a different kernel; RGB images are used as input to the high-pass filter to obtain 15 noise feature maps. At the same time, regular convolutions are used on the output image to obtain 3 color feature maps, and then the color features and noise residuals are fused to obtain an output of 18 channels. A 3×3 convolution with a stride of 1 is used on the fused features to generate 32 results; The implementation process of step (3) is as follows: VGG-16 is used as the downsampling network. The downsampling part contains 17 network layers: 13 3×3 convolutions and 4 maximum pooling operations, each convolution layer is followed by an activation function and BN layer; the upsampling network contains a total of 16 network layers: 1 1×1 convolution, 8 3×3 convolutions, 4 upsampling layers, and 3 fusion layers; each 3×3 convolution is followed by an activation function ReLU and BN layer, and the upsampling kernel size and stride are 2; The stride of all convolutions is 1; The upsampling network is divided into four stages, each of which consists of a convolutional layer and an upsampling layer. The number of convolution kernels in each stage is reduced by half, and the features of the downsampling network are fused in the first, third and fourth stages. The last stage outputs a prediction result with 64 channels, which is sent to a 1×1 convolution and a Softmax function is used for binary classification to determine whether each pixel has been repaired.

2. The image restoration forensics method based on feature-enhanced neural network according to claim 1 is characterized in that: The preprocessing in step (1) is to adjust the size, flip in any direction, distort, and randomly generate a mask to cover the image.

3. The image restoration forensics method based on feature-enhanced neural network according to claim 1 is characterized in that: The implementation process of step (4) is as follows: In order to reduce the weight of easy-to-classify samples and make the model pay more attention to difficult-to-classify samples: FL(p t )=-(1-p t ) γ log(p t ) Among them, FL is the loss function, γ is a constant, and p t It is obtained by the following formula: Among them, y∈{0,1} represents the true category, and p∈[0,1] represents the model's estimated probability of the class with label y=1.

Citation Information

Cited By

  • Method for aberration correction and image quality enhancement of a stack structure image

    CN120543438B