A robust image tampering location method and related equipment

The multi-dimensional feature of the image is extracted through the SE-U-Net network and combined with RVAE to perform reconstruction error calculation, which solves the problem of insufficient accuracy and robustness of the image tampering positioning method in the prior art, and realizes efficient tampering positioning in the images after transmission on different data sets and social media.

CN116934563BActive Publication Date: 2025-08-12SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310671915.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-07
Publication Date
2025-08-12
Estimated Expiration
2043-06-07

AI Technical Summary

Technical Problem

The prior art image tamper positioning methods have shortcomings in recognition accuracy and robustness, making it difficult to effectively identify image tampering and are not robust enough for different types of data sets.

Method used

The multi-dimensional features of the image are extracted using the SE-U-Net network, and the feature blocks are reconstructed through a lightweight robust variational autoencoder (RVAE), and the reconstruction error is calculated to locate the abnormal areas of the image, and the intermediate layer features of the tampered positioning network completed are used for unsupervised training.

Benefits of technology

It improves the accuracy and robustness of image tampering positioning, can maintain high detection performance in different types of data sets and social media transmission images, and reduces dependence on tampering tags.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116934563B_ABST
    Figure CN116934563B_ABST
Patent Text Reader

Abstract

The present invention discloses a robust image tampering location method and related equipment, including the following steps: obtaining an image to be tested, extracting multi-dimensional features of the image to be tested; dividing the multi-dimensional features of the image to be tested into multiple feature blocks, recording the position of each feature block corresponding to the image to be tested; inputting each feature block into an anomaly detection network for reconstruction; calculating the reconstruction error of each feature block based on each feature block before and after reconstruction; locating the abnormal position of the image to be tested based on the reconstruction error of each feature block and the position of each feature block corresponding to the image to be tested. It can be seen that the present invention adopts the multi-dimensional features of the image to be tested as the input of the anomaly detection network. Compared with the noise signal used in the prior art, the multi-dimensional features contain more tampering traces, which is conducive to improving the accuracy and robustness of image tampering location. Therefore, the detection accuracy and robustness of the technical solution of the present invention are higher than those of the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image detection applications, and in particular to a robust image tampering location method and related equipment. Background Art

[0002] In the information age in which modern people live, they need to obtain all kinds of information they need from images, texts, and videos that are used as information dissemination media every day. However, there is a lot of artificially forged information among these information, and the authenticity of the information cannot be judged by human senses alone.

[0003] A recent review of image tampering detection by researchers indicates that the rapid development of software development and application technologies has led to the emergence of numerous image modification software programs. While these programs can help people perform image modification, beautification, and adjustment tasks, bringing numerous conveniences to modern life and work, they are inevitably misused by a minority of users for illicit purposes that harm society, creating potential security risks. To prevent image modification tools from being used for illicit purposes while maintaining information security, researchers in the field of information security are increasingly focusing on the security of digital images.

[0004] How to locate and identify image tampering becomes the key to the problem. The accuracy and robustness of locating and identifying image tampering directly affect the security of digital images. Summary of the Invention

[0005] The purpose of the present invention is to provide a robust image tampering location method and related equipment, aiming to solve the problems of insufficient recognition accuracy and robustness of the image tampering location method in the prior art.

[0006] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0007] The present invention provides a robust image tampering location method, comprising the steps of:

[0008] Acquire the image to be tested and extract multi-dimensional features of the image to be tested;

[0009] Divide the multi-dimensional features of the image to be tested into multiple feature blocks, and record the position of each feature block corresponding to the image to be tested;

[0010] Each feature block is input into the anomaly detection network for reconstruction;

[0011] Calculate the reconstruction error of each feature block based on each feature block before and after reconstruction;

[0012] The abnormal position of the image to be measured is located according to the reconstruction error of each feature block and the position of each feature block corresponding to the image to be measured.

[0013] Optionally, the anomaly detection network includes an encoder and a decoder;

[0014] Inputting each feature block into the anomaly detection network for reconstruction specifically includes:

[0015] The feature block f obtained by the encoder passes through the adaptive average pooling layer to reduce the spatial size of the feature to 1, while keeping the channel dimension of the feature unchanged, flattening the feature into a one-dimensional vector, and transforming the one-dimensional vector into the mean vector u i and standard deviation vector Gaussian distribution represented by Features are sampled from the Gaussian distribution and fed into the decoder.

[0016] Optionally, the encoder includes five sequentially connected encoding layers, each encoding layer includes a sequentially connected convolutional layer, a batch normalization layer, and a LeakyReLU activation function, the kernel size of each convolutional layer is 3, the stride is 2, and the padding is 1, and the output channels of each convolutional layer from the front to the back are 32, 64, 128, 256, and 512 respectively;

[0017] The decoder includes five deconvolution layers connected in sequence. The output channel numbers of the first four deconvolution layers are 256, 128, 64, and 32 respectively from the front to the back. The output channel number of the last deconvolution layer is consistent with the number of F channels.

[0018] Optionally, the calculation of the reconstruction error of each feature block based on each feature block before and after reconstruction is specifically performed according to the following formula:

[0019]

[0020] Where AnomalyScore is the reconstruction error, f is the feature block before reconstruction, and f' is the feature block after reconstruction.

[0021] Optionally, locating the abnormal position of the image to be measured according to the reconstruction error of each feature block and the position of each feature block corresponding to the image to be measured specifically includes:

[0022] For each pixel, the average value of the reconstruction errors of all feature blocks containing the pixel is calculated as the abnormality score of the pixel to obtain the abnormality scores of all pixels;

[0023] Pixels with anomaly scores greater than the set threshold are located as abnormal locations.

[0024] Optionally, dividing the multi-dimensional features of the image to be measured into a plurality of feature blocks specifically comprises dividing the multi-dimensional features of the image to be measured into a plurality of feature blocks of a set size through a sliding window of a set size.

[0025] Optionally, if the size of the image to be measured is larger than a set threshold, the step length of the sliding window is set to the first step length; otherwise, the step length of the sliding window is set to the second step length, and the first step length is larger than the second step length.

[0026] Optionally, the extracting of multi-dimensional features of the image to be tested is specifically inputting the image to be tested into a trained SE-U-Net network, and obtaining the output of the penultimate convolutional layer of the SE-U-Net network as the multi-dimensional features.

[0027] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, which includes: a memory, a processor, and an image tampering positioning program stored on the memory and runnable on the processor, and when the image tampering positioning program is executed by the processor, the steps of the image tampering positioning method described above are implemented.

[0028] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, which stores an image tampering locating program, and when the image tampering locating program is executed by a processor, the steps of the image tampering locating method described above are implemented.

[0029] The present invention adopts the above technical solution to achieve the following effects:

[0030] The present invention extracts multi-dimensional features of the image to be tested as the input of the anomaly detection network. The anomaly detection network reconstructs the multi-dimensional features and locates and identifies the abnormal area of the image to be tested based on the reconstruction error, thereby locating the tampered position of the image to be tested. Compared with the noise signal used in the prior art, the multi-dimensional features contain more tampering traces, which is beneficial to improving the accuracy and robustness of image tampering positioning. Therefore, the detection accuracy and robustness of the technical solution of the present invention are higher than those of the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a flowchart of the steps of the image tampering location method in a preferred embodiment of the present invention;

[0032] Figure 2 It is a schematic diagram of the structure of the network used in the preferred embodiment of the present invention;

[0033] Figure 3 It is a structural diagram of the existing technology space and channel squeezing and excitation module;

[0034] Figure 4 It is a visualization result diagram of different tampering localization methods in the test example of the present invention;

[0035] Figure 5 This is a visualization result diagram of the CASIA data set under different transmission media in the test example of the present invention;

[0036] Figure 6 2 is a schematic diagram of an input image selection interface of an image tampering location program according to a preferred embodiment of the present invention;

[0037] Figure 7 This is a schematic diagram of the image tampering location result interface of the image tampering location program of the preferred embodiment of the present invention

[0038] Figure 8 Schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0040] Example 1

[0041] Due to the rapid development of deep learning technology and the substantial increase in computer speed, many studies have been devoted to the problem of localizing image tampering. For example, in DFCN (dense fully convolutional network), the authors developed an image tampering localization scheme for Photoshop tampering software. This scheme requires a large number of tampered images and corresponding tampering labels to train the network.

[0042] In order to make the tampering localization scheme free from the limitation of requiring a large number of image tampering labels, a class of methods, such as the concept of blind tampering localization, are introduced. This type of method does not require a specific dataset for training or fine-tuning, nor does it use metadata or other prior information about the test data. This type of method starts from the essential characteristics of tampered and untampered data and trains the network with unsupervised or non-pixel-level tampering labels.

[0043] In recent years, image tampering localization schemes with stronger tampering localization capabilities have emerged, such as ViT-VAE. Although its performance has been significantly improved compared with schemes such as Noiseprint, its overall performance is still far behind that of image tampering localization schemes that use tampering labels for supervised training. The difference in detection accuracy is more obvious in scenarios with more post-processing, which reflects that existing image tampering localization schemes are still not robust enough for different types of data sets.

[0044] After extensive research on existing image tampering localization schemes, the present invention found that networks trained on large datasets often have stronger robustness. This is because the training dataset may contain post-processing operations similar to those in the test dataset, which enables the network to learn the difference between normal images and tampered images in advance under the guidance of tampering labels. At the same time, image augmentation can generally be used during network training to make the network training more diverse in data distribution. It is precisely because of the large number of parameters of this type of network and the high-quality and abundant training data that this type of method has powerful detection capabilities and robustness.

[0045] Based on the above facts, it can be inferred that in addition to the output results of the network can be directly used for image tampering localization, other intermediate layer features in the network can also be directly or indirectly used for image tampering localization.

[0046] Therefore, the present invention develops an image tampering localization solution that does not directly use image tampering labels, taking into account detection robustness while being independent of data sets. Specifically, the present invention uses the intermediate layer features of an existing image localization network trained using a large-scale tampering data set to explore abnormal areas in image features, thereby achieving a robust image tampering localization solution.

[0047] See Figure 1 , Figure 1 It is a flowchart of the steps of the image tampering location method in a preferred embodiment of the present invention.

[0048] Embodiment 1 discloses a robust image tampering location method, which includes the following steps:

[0049] S1. Obtain an image to be tested and extract multi-dimensional features of the image to be tested.

[0050] In this embodiment, the image tampering location method adopts the following method: Figure 2 The neural network shown in the figure is used for positioning. In step S1, a parallel spatial channel squeeze and excitation U-net (Spatial Channel Squeeze-and-Excitation U-Net, SE-U-Net) is used to extract the multi-dimensional features of the image to be tested. In step S1, the multi-dimensional features of the image to be tested are obtained by obtaining the output of the middle layer of the SE-U-Net network.

[0051] It is worth noting that the use of the SE-U-Net network is only a preferred embodiment of the present invention. Other networks can also be used to extract the multi-dimensional features of the image, and the multi-dimensional features can be obtained by extracting the intermediate layer outputs of other networks.

[0052] Specifically, this embodiment uses the SE-U-Net network trained with large-scale image tampering labels and loads the corresponding trained parameters. The structure of the SE-U-Net network can be found in Figure 2 ,SE-U-Net network is an end-to-end codec U-shaped network, which uses multiple Figure 3 The spatial and channel squeezing and excitation modules are shown.

[0053] For the input feature T∈R H×W×C , the spatial and channel squeezing and excitation module simultaneously readjusts the spatial dimension and channel dimension of the feature, where H is the height of the input feature, W is the width of the input feature, and C is the number of channels of the input feature.

[0054] In the spatial dimension, the spatial and channel squeezing and excitation module maps the original number of channels to a feature vector with a channel number of 1 through the convolution layer W1 with a kernel function size of 1, a step size of 1, and a padding of 0. After the Sigmoid activation function, the dot product is performed with the original feature T to obtain the feature F after the spatial direction is recalibrated. S , the overall calculation formula of the process is as follows:

[0055]

[0056] Where ⊙ is the matrix dot product, is the mapping of the convolutional layer to the input features, and Sigmoid is the Sigmoid activation function.

[0057] For the channel dimension, the spatial and channel squeezing and excitation modules transform the feature F into an intermediate vector v∈R by adaptive average pooling. 1×1×C , then pass through two linear layers W2, W3 and Sigmoid activation function, and finally get the feature F after channel direction recalibration with the original feature dot product C , the calculation formula is as follows:

[0058]

[0059] Where Relu is the Relu activation function.

[0060] Get the feature F after channel direction recalibration C and spatial direction recalibrated feature F S Then, for feature F C and feature F S Fusion, feature F S and F C There are many ways to fuse, such as connection on the channel dimension, pixel-by-pixel addition, etc. In this embodiment, F is adopted. S and F CThe larger value at each position in the ensemble is used as the new eigenvalue for fusion.

[0061] Take a color image I∈R H×W×3 Input to the SE-U-Net network, the network will output a single-channel prediction result P∈R of the same size as the input image H×W×1 , in the OSN method, the final binary prediction result can be obtained by determining the segmentation threshold. The difference is that in this embodiment, in order to obtain more available features, the network middle layer features are selected as the input of the abnormal network, that is, in this embodiment, only part of the structure of the SE-U-Net network is used to obtain the multi-dimensional features of the image to be tested outputted from the middle layer as the input of the subsequent network. Specifically, the present invention notes that before obtaining the prediction result P, it is necessary to pass through the last convolution layer, which reduces the high-dimensional features in the number of channels to one-dimensional features. Therefore, in this embodiment, the SE-U-Net network with the trained parameters in OSN is used, and the multi-dimensional features of the input image I before passing through the last convolution layer of the network are extracted. The number of dimensions on the feature channel is 32, which is expressed as F∈R H×W×32 .

[0062] S2. Divide the multi-dimensional features of the image to be measured into multiple feature blocks, and record the position of each feature block corresponding to the image to be measured.

[0063] For details, please refer to Figure 2 In this embodiment, the image to be tested I is subjected to feature extraction to obtain the feature F∈R H ×W×32 Finally, the present invention divides the network middle layer feature F into several feature blocks f∈R through a sliding window of size 64×64 64×64×32 As the input of the subsequent RVAE network.

[0064] Taking into account the difference in input image size, if the size of the image to be tested is larger than 1000×1000, the sliding window step size is set to 16, otherwise the sliding window step size is set to 8.

[0065] S3. Input each feature block into the anomaly detection network for reconstruction.

[0066] Specifically, in order to reduce network parameters and accelerate the inference speed of subsequent tampering detection software, the present invention designs a lightweight neural network composed of convolutional layers and linear layers as an anomaly detection network. The lightweight neural network includes an encoder and a decoder. The encoder is specifically a robust variational auto-encoder (RVAE) based on the characteristics of the network's intermediate layers, which is used to solve the problem of image tampering localization under conditions with a large number of post-processing operations.

[0067] RVAE consists of five sequentially connected encoding layers. Each encoding layer consists of a sequentially connected convolutional layer, a batch normalization layer, and a LeakyReLU activation function. The kernel size of each convolutional layer is 3, and its stride is 2 and the padding is 1. The output channels of each convolutional layer from front to back are 32, 64, 128, 256, and 512 respectively.

[0068] The LeakyReLU activation function formula is as follows:

[0069]

[0070] For any input feature block f, the features obtained after RVAE are passed through the adaptive average pooling layer to reduce the spatial size of the features to 1, while keeping the channel dimension of the features unchanged. The features are then flattened into a one-dimensional vector and the features are re-represented.

[0071] The re-representation process transforms the features into one-dimensional vectors into the mean vector u i and standard deviation vector Gaussian distribution represented by Then, features are sampled from the Gaussian distribution and fed into the decoder.

[0072] Specifically, the following sampling method is used:

[0073] z′ i =u i +ρ i ⊙∈where∈~N(0,1), i∈{1,…,K};

[0074] The above formula indicates that ∈ is randomly sampled from the standard normal distribution, and after being mixed with σ i The vector dot product plus u i Finally, the sampling result z′ is obtained i , where K is the number of samples. In this embodiment, the value of K is specifically 10.

[0075] The decoder includes five sequentially connected deconvolution layers. The output channel numbers of the first four deconvolution layers are 256, 128, 64, and 32 respectively. The output channel number of the last deconvolution layer is consistent with the number of F channels. The decoder is used to obtain the reconstructed feature block f′∈R 64×64×32 .

[0076] S4. Calculate the reconstruction error of each feature block based on the feature blocks before and after reconstruction.

[0077] Specifically, the reconstruction error of the feature block is calculated according to the following formula:

[0078]

[0079] Where AnomalyScore is the reconstruction error of feature block f.

[0080] S4. Locating abnormal positions of the image to be measured according to the reconstruction errors of the characteristic blocks and the positions of the characteristic blocks corresponding to the image to be measured.

[0081] Specifically, when the sliding window generates feature blocks, the position of the feature blocks corresponding to the image to be tested is recorded, so the abnormal position can be located by the following steps:

[0082] The pixel-level anomaly score for each pixel is obtained by averaging the reconstruction errors of all feature patches containing the corresponding pixel as the anomaly score.

[0083] The heat map of the image under test can be obtained by scaling the pixel-level anomaly score to the range of 0 to 1.

[0084] The heat map is thresholded by setting a threshold T to obtain a binary prediction result, that is, pixels with anomaly scores greater than the threshold T are abnormal images, and pixels with anomaly scores less than the threshold T are normal images.

[0085] In this embodiment, the maximum number of training times is 20, and the loss includes reconstruction loss and KLD loss, that is, Loss = L rec +L kld , the loss LOSS of each feature block can be calculated according to the following formula:

[0086]

[0087] In this embodiment, the parameters of the RAVE network are obtained through training. The training stop condition is that the maximum number of training times is reached or the difference between the average loss of the current training and the average loss of the previous training is lower than the threshold value of 10. -4 .

[0088] Stop training early when the following conditions are met:

[0089]

[0090] Where, Represents the average loss of this training, Represents the average loss of the previous training.

[0091] In order to test the effect of the scheme of the present invention on image tampering localization, the present invention uses four evaluation data sets for experiments, including DSO, Columbia, NIST, and CASIA v1 data sets with 100, 160, 564, and 920 test images respectively. At the same time, in order to better illustrate the comparison effect with the existing technology, the present invention selects six image tampering localization methods for comparison, including: ViT-VAE, Mantra-net, Noiseprint, ForSim, DFCN, and OSN.

[0092] The test examples of the present invention use the pixel-level evaluation indicators AUC, F1, and IoU used in both DFCN and OSN.

[0093] Among them, the AUC (Area Under Curve) indicator comprehensively considers the false positive rate and true positive rate. This indicator is defined as the area formed by the coordinate axis under the ROC curve (Receiver Operating Characteristic Curve). The value range is between 0 and 1. The closer to 1, the more accurate the prediction. In the ROC curve, the horizontal axis represents the false positive rate and the vertical axis represents the true positive rate. By taking different thresholds for the heat map, a series of paired TPRs (true positive rate) and FPRs (false positive rate) can be obtained. The AUC indicator can be calculated through these paired values. Specifically, the AUC indicator can be calculated according to the following formula:

[0094] AUC = ∫TPR(FPR -1 (x))dx;

[0095] The F1 score takes into account both precision and recall. It is calculated by taking the harmonic mean of precision and recall:

[0096]

[0097] In the formula, Precision is the precision rate and Recall is the recall rate.

[0098] The IoU metric evaluates the intersection-over-union ratio between the thresholded prediction result and the true tampered label:

[0099]

[0100] The comparative experimental results of various algorithms on standard datasets are shown in Table 1 and Figure 4 As shown in Table 1, the best results are bolded and the suboptimal results are underlined.

[0101] Table 1 Detection performance of various algorithms on standard datasets

[0102]

[0103]

[0104] As can be seen in Table 1, the proposed RVAE method performs worse than ViT-VAE on the DSO dataset, but outperforms ViT-VAE on the Columbia, NIST, and CASIA datasets. This is primarily due to the differences in the features used by the two methods. ViT-VAE uses noiseprint features trained on unmodified, minimally post-processed camera images, while RVAE uses features from the network's mid-layer trained on a large-scale, tampered dataset. Therefore, RVAE naturally performs better on data with extensive post-processing, especially when compression quality is significant.

[0105] from Figure 4 From the above, OSN achieved the best performance, followed by the proposed method. This result demonstrates that the proposed method can enhance algorithm robustness by leveraging the intermediate-layer features of a trained network. However, since the proposed method has not undergone end-to-end training with tampering labels, its detection results are inferior to those of OSN. However, the proposed method does not directly rely on training with tampering labels, but instead uses the network's intermediate-layer features as a starting point. If superior tampering localization algorithms emerge in the future, the proposed method can still use their intermediate-layer features as a replacement, thus providing flexibility. While OSN requires supervised training of the tampering localization network using a large number of tampered images and corresponding tampering labels, supervised approaches using large-scale tampered image training data fail when a sufficient number of tampered images are unavailable. However, the proposed method does not directly rely on tampered images and their labels. Instead, it uses the intermediate-layer features of a trained tampering localization network for unsupervised training, identifying unusually large regions within the features as tampered areas. This allows for effective use even when a sufficient number of tampered images are unavailable.

[0106] In order to further compare the robustness of various methods, the present invention uses the images proposed in OSN after being transmitted through social media for testing, including four data sets: NIST, CASIA, DSO, and Columbia. Each data set is transmitted through four different social media, including Facebook, Whatsapp, Weibo, and Wechat. Finally, the present invention is tested on these 6976 tampered images transmitted through social media. Note that these four types of social media will perform post-processing operations on the images, including but not limited to scale scaling, compression, etc. The present invention reports the test results in Tables 2 to 5. The best results are bold and the suboptimal results are underlined. At the same time, the present invention displays the visual results in Figure 5 middle.

[0107] Table 2 Detection performance of standard dataset after being transferred via Facebook

[0108]

[0109] Table 3 Detection performance of the standard dataset after transmission via WhatsApp

[0110]

[0111]

[0112] Table 4 Detection performance of standard dataset after transmission via Weibo

[0113]

[0114] Table 5 Detection performance of the standard dataset after transmission via WeChat

[0115]

[0116] Overall, OSN achieved the best results, while the RVAE proposed in this paper was second only to OSN in most data sets. Moreover, the detection performance after social media transmission did not drop significantly like ViT-VAE and Noiseprint, but remained stable around the original detection performance, which shows that the method of this invention has good robustness.

[0117] from Figure 5In this paper, we can see that most methods can detect untransmitted CASIA images, except for the Noiseprint method, which cannot process smaller images. At the same time, the detection performance of various methods is affected by social media transmission, which is manifested as a decrease in accuracy. For example, in ViT-VAE, after WeChat transmission, this method gives incorrect detection results. Overall, ViT-VAE is more affected by social media transmission, while RVAE and OSN are less affected by social media transmission. OSN shows the highest accuracy, followed by the RVAE method proposed in this paper.

[0118] In this example, a robust image tampering localization method was developed using PyQT5, a toolkit for creating graphical user interfaces (GUIs) in the Python development environment. Experiments were conducted on a computer with an Intel(R) Core(TM) i7-10700 processor and 32GB of memory. All code was written in Python 3.7, which relies on the PyTorch 3.7.0 deep learning framework.

[0119] The image tampering location program implemented in this embodiment is as follows Figure 6 As shown, users can freely select the images to be detected through this software.

[0120] Users can use this program to detect tampered images in complex scenes. Figure 7 The figure shows the detection results of the proposed algorithm on the images in the CASIA dataset after being transmitted via WeChat. From left to right, the figure shows the input image, the heat map output by the proposed algorithm, and the result after the heat map thresholding. It can be seen from the heat map that the system has a high probability of judging that the stone lion is a tampered area. At the same time, the prediction result also gives a binary judgment, and the area where the stone lion is located is judged as a tampered area.

[0121] Example 2

[0122] See Figure 8 Based on the above method, the present invention also provides a terminal, which includes: a memory 10, a processor 20, and an image tampering positioning program stored on the memory 10 and runnable on the processor 20. When the image tampering positioning program is executed by the processor 20, the steps of the image tampering positioning method described above are implemented.

[0123] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, an image tampering positioning program is stored on the memory 20, and the image tampering positioning program can be executed by the processor 10, thereby realizing the image tampering positioning method in the present application.

[0124] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program codes or process data stored in the memory 20, such as executing the image tampering location method.

[0125] Example 3

[0126] This embodiment provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores an image tampering locating program, and when the image tampering locating program is executed by a processor, the steps of the image tampering locating method described above are implemented.

[0127] In summary, the solution proposed in the present invention does not directly rely on tampered images and their labels, but uses the intermediate layer features of the already trained tampering localization network to perform an unsupervised training process and regards the areas with larger abnormalities in the features as tampered areas. Even if a sufficient number of tampered images cannot be obtained, it can still be effectively used. In addition, the present invention uses high-dimensional features as the input of the anomaly detection network, which contains more tampering traces. Therefore, the detection accuracy and robustness of the technical solution of the present invention are higher than those of the prior art.

[0128] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0129] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0130] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A robust image tampering location method, characterized in that: Including steps: Acquire the image to be tested and extract multi-dimensional features of the image to be tested; Divide the multi-dimensional features of the image to be tested into multiple feature blocks, and record the position of each feature block corresponding to the image to be tested; Each feature block is input into the anomaly detection network for reconstruction; Calculate the reconstruction error of each feature block based on each feature block before and after reconstruction; The abnormal position of the image to be tested is located according to the reconstruction error of each feature block and the position of each feature block corresponding to the image to be tested; The anomaly detection network includes an encoder and a decoder; Inputting each feature block into the anomaly detection network for reconstruction specifically includes: The feature block f obtained by the encoder passes through the adaptive average pooling layer to reduce the spatial size of the feature to 1, while keeping the channel dimension of the feature unchanged, flattening the feature into a one-dimensional vector, and transforming the one-dimensional vector into the mean vector u i and standard deviation vector Gaussian distribution represented by Sample features from the Gaussian distribution and feed them into the decoder; The multi-dimensional features of the image to be tested are extracted by inputting the image to be tested into a trained SE-U-Net network and obtaining the output of the penultimate convolutional layer of the SE-U-Net network as the multi-dimensional features.

2. A robust image tampering location method according to claim 1, characterized in that: The encoder includes five sequentially connected encoding layers, each of which includes a convolutional layer, a batch normalization layer, and a LeakyReLU activation function. The kernel size of each convolutional layer is 3, the stride is 2, and the padding is 1. The output channels of the convolutional layers from the front to the back are 32, 64, 128, 256, and 512 respectively. The decoder includes five deconvolution layers connected in sequence. The output channel numbers of the first four deconvolution layers are 256, 128, 64, and 32 respectively from the front to the back. The output channel number of the last deconvolution layer is consistent with the number of F channels.

3. A robust image tampering location method according to claim 1, characterized in that: The reconstruction error of each feature block is calculated based on each feature block before and after reconstruction. Specifically, the reconstruction error of each feature block is calculated according to the following formula: Where AnomalyScore is the reconstruction error, f is the feature block before reconstruction, and f' is the feature block after reconstruction.

4. A robust image tampering location method according to claim 1, characterized in that: The abnormal position of the image to be tested is located according to the reconstruction error of each feature block and the position of each feature block corresponding to the image to be tested, specifically including: For each pixel, the average value of the reconstruction errors of all feature blocks containing the pixel is calculated as the abnormality score of the pixel to obtain the abnormality scores of all pixels; Pixels with anomaly scores greater than the set threshold are located as abnormal locations.

5. The robust image tampering location method according to claim 1, characterized in that: The dividing the multi-dimensional features of the image to be measured into a plurality of feature blocks specifically involves dividing the multi-dimensional features of the image to be measured into a plurality of feature blocks of a set size through a sliding window of a set size.

6. A robust image tampering location method according to claim 5, characterized in that: If the size of the image to be measured is larger than the set threshold, the step length of the sliding window is set to the first step length; otherwise, the step length of the sliding window is set to the second step length, and the first step length is larger than the second step length.

7. A terminal, characterized in that: The terminal includes: a memory, a processor, and an image tampering locating program stored in the memory and executable on the processor. When the image tampering locating program is executed by the processor, the steps of the image tampering locating method according to any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an image tampering locating program, and when the image tampering locating program is executed by a processor, the steps of the image tampering locating method according to any one of claims 1 to 6 are implemented.