Image forgery detection method, system and device based on pixel difference perception
Through the image forgery detection method based on pixel difference perception, the position prospecting module and lightweight deep learning network are used to solve the problem of insufficient detection accuracy of high-quality forgery images, and efficient and robust image forgery detection are achieved.
Patent Information
- Application Number
- CN202510404924.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
AI Technical Summary
The existing image forgery detection methods are insufficient in the detection accuracy, weak generalization ability when facing high-quality forgery images, and insufficient utilization of multimodal information interaction, especially when processing non-face images and low-quality images.
The image forgery detection method based on pixel difference perception is adopted, and the image spatial perception is enhanced through the position prospecting module, combined with a lightweight deep learning network, the fine forgery traces of the image are extracted using adjacent pixel difference values, reducing the dependence on abnormal labels, and improving detection accuracy and robustness.
It significantly improves the accuracy and robustness of image forgery detection, is suitable for real-time detection and large-scale applications, reduces computing overhead and data labeling costs, and enhances the universality of detection methods.
Smart Images

Figure CN120355949A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and artificial intelligence, and relates to image forgery detection technology. More specifically, it relates to an image forgery detection method, system and device based on pixel difference perception. Background Art
[0002] With the rapid development of deep learning technologies such as generative adversarial networks (GANs) and diffusion models, the ability of image synthesis has been greatly improved, and the fidelity of synthetic images has approached or even exceeded the recognition limit of human visual perception. These technologies have shown great application potential in fields such as artistic creation, virtual reality, and medical imaging. However, the double-edged sword nature of technology has gradually emerged - the abuse of image forgery technology poses a serious threat to social trust, information security, and privacy protection. For example, the spread of false information, the leakage of personal privacy, and fraud in the financial field may all be exacerbated by the abuse of image forgery technology.
[0003] Traditional image forgery detection methods mainly rely on the frequency-domain analysis or spatial feature extraction of images. These methods often perform poorly when faced with high-quality synthetic images. Especially for images generated by diffusion models, due to the lack of obvious frequency-domain artifacts and highly realistic details, traditional detection methods are difficult to work effectively. In addition, existing detection methods also have deficiencies in cross-model generalization ability. For example, detectors trained based on GAN models have a significant performance drop when faced with images generated by diffusion models. At the same time, these methods are also significantly affected in detection performance when faced with low-quality images (such as blurred or compressed images).
[0004] To address the above challenges, researchers have proposed various improvement schemes. For example, by introducing enhancement modules to improve the detection performance for low-quality generated images; proposing the diffusion model reconstruction error (DIRE) to enhance the general detection ability for diffusion models; and the SSP method based on simple local blocks, which achieves high efficiency by detecting only using a single simple image block. However, these methods still have limitations, such as insufficient interaction of multi-modal information, excessive dependence on labeled data, insufficient generalization ability, and insufficient robustness in complex application scenarios.
[0005] After retrieval, the Chinese patent publication number is CN 119091516 A, and the publication date is December 6, 2024. The invention creation name is: Method, device, equipment and medium for detecting forged images based on image self-fusion. The method disclosed in this patent includes: obtaining natural images, extracting facial marks and facial contours; using image processing methods to process the facial marks and facial contours of natural images to obtain a target image and a source image; performing size scaling and translation transformation processing on the source image to obtain a transformed image; performing enhancement processing on the transformed image to obtain an enhanced image; performing image block pooling processing on the enhanced image to obtain a mask image; performing image self-fusion on the target image, the transformed image, and the mask image to obtain a fused image; using the natural image and the fused image as an image pair, inputting multiple image pairs into an initial model for training until a preset condition is met to obtain a detection model for detecting forged images. The method of this patent document can be used for detecting forged images, but it has the following deficiencies:
[0006] (1) The method of this patent mainly relies on facial mark and contour extraction, which is limited to the detection of forged facial images. This means that its method has weak detection ability for non-facial forged images and cannot be well applied to diverse types of forged images. Especially currently, more and more forged images are generated using GAN or diffusion models, and it is difficult to extract facial mark information.
[0007] (2) When processing images, this method has a high dependence on steps such as image size transformation and enhancement, and its processing method based on image block pooling is prone to losing some detail information, which may affect the detection accuracy of forgery traces. Especially when processing low-quality or compressed forged images, the robustness of this method has limitations. Summary of the Invention
[0008] 1. Problems to be Solved
[0009] The present invention provides an image forgery detection method, system and device based on pixel difference perception, which improves the accuracy and robustness of image forgery detection, and effectively solves the problems of insufficient detection accuracy, weak generalization ability and insufficient utilization of multi-modal information interaction of existing image forgery detection methods when facing high-quality forged images.
[0010] 2. Technical Solutions
[0011] To solve the above problems, the technical solutions adopted by the present invention are as follows:
[0012] Firstly, the present invention provides an image forgery detection method based on pixel difference perception, including the following steps:
[0013] S1. Obtain real images and forged images, perform preprocessing and annotation, and output image A1;
[0014] S2. Process image A1 using the position prediction module, including processing the positions of image pixels through position encoding and outputting image A2;
[0015] S3. Perform downsampling on image A2, extract the low-frequency information of the image, and output image A3;
[0016] S4. First, adjust the size of image A2 to be the same as that of image A3, and then subtract the pixel points of the resized image A2 from the pixel points of image A3 to calculate the adjacent pixel differences and form a difference image, denoted as image A4;
[0017] S5. Perform upsampling on image A4, extract features, and output image A5;
[0018] S6. Use a lightweight deep learning network to classify image A4 and output the detection result.
[0019] Furthermore, in step S1, the preprocessing of the image includes randomly flipping, cropping, and decompressing the image, and after annotation, image data and corresponding labels are obtained.
[0020] Furthermore, step S2 specifically includes the following steps:
[0021] S2.1. Input feature processing: First, divide the input image into multiple small image patches. Each image patch is input into the position prediction module as an embedded character token. Each token will carry its spatial position information and is enhanced through position encoding;
[0022] S2.2. Process the image using the sliding window mechanism. The size of the sliding window is adjusted according to the resolution of the image. During processing, the central token of each window is associated with the surrounding tokens to capture the local fine-grained features of the image;
[0023] S2.3. Position prediction attention processing: For each position (i, j), the position prediction attention calculates its similarity with the surrounding neighbors, and calculates the attention weights through the following steps:
[0024] A i,j = Softmax(W Q ·V i,j )
[0025] where V i,j is the relationship value between the central point and the neighboring pixel points, W Q is a linear transformation matrix used to generate weights; the Softmax(·) function is used to convert the output value of the model into a probability distribution;
[0026] S2.4, Feature Aggregation: The calculated attention weight A i,j will be used to weight the average of neighbor features to generate a new feature representation, and this process is completed by the following formula:
[0027] Y i,j = MatMul(A i,j , V i,j )
[0028] where A i,j is the calculated attention weight, representing the similarity between each token; V i,j represents the relationship value between the center point and neighboring pixel points; MatMul(·) represents matrix multiplication operation, multiplying the weight after Softmax normalization with V i,j to obtain the weighted output.
[0029] Furthermore, in step S3, bilinear interpolation method or convolution operation is used for downsampling operation, and the downsampling ratio coefficient is r, 0.3 ≤ r ≤ 0.7.
[0030] Furthermore, in step S4, interpolation method or convolution operation is used to scale the image size, and the scaling ratio coefficient is r′, r′ = r.
[0031] Furthermore, the formula for calculating the difference of neighboring pixels in step S4 is as follows:
[0032] NDP i = x scaled,i - x down,i
[0033] where X scaled,i is the pixel point of the resized image A2, x down,i is the pixel point of image A3, and NDP i is the difference of neighboring pixels.
[0034] Furthermore, in step S5, bilinear interpolation method is used for upsampling operation, and the upsampling ratio coefficient is r″, r″ = 1 / r.
[0035] Furthermore, in step S6, the lightweight deep learning network adopts ResNet50 network.
[0036] Second, the present invention also provides an image forgery detection system based on pixel difference perception. This system is used for the detection method of the present invention, and it includes a data acquisition module, a position prediction module, a pixel difference map extraction module, a model training module, and a classification and discrimination module, where:
[0037] Data acquisition module: used to obtain real image and forged image datasets, and perform preprocessing and annotation;
[0038] Position outlook module: used to enhance the image spatial perception ability through the position outlook module;
[0039] Pixel difference map extraction module: used to calculate the differences between adjacent pixels and extract high-frequency forgery traces;
[0040] Model training module: used to classify the extracted features using a lightweight deep learning network, and optimize the detection model through the cross-entropy loss function;
[0041] Classification and discrimination module: used to detect the input image and determine whether the image is a forged image.
[0042] Thirdly, the present invention provides an apparatus for detecting image forgery based on pixel difference perception, the apparatus includes: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the apparatus to execute the method steps of the present invention.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] Through the position outlook module (POM) and the neighboring pixel difference (NDP), the present invention can accurately capture the subtle forgery traces in the image generation process, significantly improving the accuracy and robustness of image forgery detection; at the same time, the lightweight network architecture design reduces the computational overhead and improves the detection efficiency, which is suitable for real-time detection and large-scale applications; in addition, the method of the present invention reduces the dependence on abnormal labels, reduces the cost of data annotation, and improves the universality of the detection method. Description of the Drawings
[0045] Figure 1 It is a schematic diagram of the overall process of a method for detecting image forgery based on pixel difference perception of the present invention;
[0046] Figure 2 It is a schematic diagram of the overall structure of a system for detecting image forgery based on pixel difference perception of the present invention;
[0047] Figure 3 It is a schematic diagram of the structure of the fine-grained image position outlook module of the present invention;
[0048] Figure 4 It is a schematic diagram of the structure of the pixel difference extraction module of the present invention;
[0049] Figure 5 It is the output result of detecting real images and forged images in Embodiment 1 of the present invention. Detailed Embodiments
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0051] Combined with Figure 1 , the present invention proposes an image forgery detection method based on pixel difference perception, and the specific steps are as follows:
[0052] Step S1: Obtain real images and forged images from existing public datasets, and preprocess the images. Among them, the preprocessing includes: random flipping, cropping, decompression, etc.; then use existing methods to label the processed images to obtain image data x and corresponding labels y, and the output images are recorded as image A1.
[0053] Step S2: First, the image A1 is processed by using a position perspective module (denoted as the POM module) to enhance the spatial perception ability of the image, so as to improve the accuracy of detecting forged images, and the output image is A2. The POM module includes a sliding window mechanism and a position encoding method. Through the processing of this module, the joint perception of global and local forgery features of the image to be detected can be effectively enhanced, and the robustness of detection can be improved.
[0054] The position encoding described above is a key component in the position perspective module (POM module) of the present invention. The position information of each image pixel is transmitted to the network through position encoding to enhance the perception of the spatial relationship of the image. After position encoding, the sliding window mechanism is used to enhance the local and global feature perception abilities of the image. The specific implementation steps are as follows:
[0055] S2.1 Input feature processing: First, the input image is divided into multiple small image blocks, and each image block is input into the POM module as an embedded character token. Each token will carry its spatial position information and be enhanced through position encoding.
[0056] S2.2 Sliding window mechanism: The POM module uses the sliding window mechanism to process the image. The central token of each window will be associated with the surrounding tokens to capture local fine-grained features. The size k×k of this sliding window can be adjusted according to the resolution of the image, and a 5×5 window is used in the present invention.
[0057] S2.3 Location Prospective Attention: For each location (i, j), the location prospective attention calculates its similarity with surrounding neighbors. Different from traditional self-attention, the POM module calculates the attention weights through the following steps:
[0058] A i,j = Softmax(W Q ·V i,j )
[0059] where V i,j is the relationship value between the central point and neighboring pixels, and W Q is a linear transformation matrix used to generate weights. The Softmax(·) function is used to transform the output values of the model into a probability distribution for facilitating the implementation of classification tasks. It normalizes the output values for each class c to ensure that the output results are in the form of probabilities and the sum of all probability values is 1.
[0060]
[0061] where z c is the output of the c-th class, and c′ is the index of all classes. Through this normalization process, the Softmax(·) function transforms the scores of all classes into probabilities. The output of the Softmax(·) function is combined with a loss function (such as cross-entropy loss) for backpropagation optimization to ultimately improve the classification accuracy of the model.
[0062] 2.4 Feature Aggregation: The calculated attention weights A i,j will be used to weight the average of neighbor features to generate a new feature representation. This process is completed through the following formula:
[0063] Y i,j = MatMul(A i,j , V i,j )
[0064] where A i,j is the calculated attention weight representing the similarity between each token; V i,j represents the relationship value between the central point and neighboring pixels; MatMul(·) represents matrix multiplication operation, multiplying the Softmax-normalized weights with V i,j to obtain the weighted output.
[0065] Finally, all neighbor token features are aggregated into each token to obtain an output containing global and local features.
[0066] It should be noted that the present invention processes the image by introducing a position anticipation module, and optimizes the design of the structure of the position anticipation module. In terms of design, first, the sliding window mechanism is used to enhance the association between different regions, effectively capturing minute forgery traces generated in local regions; at the same time, by adding position encoding, context information is introduced for each pixel position, enhancing the model's understanding of the spatial structure of the image and making up for the deficiency of traditional models in processing spatial information. Finding a balance between the local and global aspects of the image is beneficial to improving the overall detection accuracy of the model.
[0067] Step S3: Perform downsampling on the image processed by the POM module (i.e., image A2) to extract the low-frequency information of the image and output image A3; the downsampling operation helps capture the overall structural features of the image and provides a basis for subsequent high-frequency information extraction.
[0068] The downsampling operation is a process of reducing the size of the input image, usually implemented through pooling operations. Here, we use bilinear interpolation for the downsampling operation. Assume the input image x i is a matrix with a size of H×W, where H is the height of the image and W is the width of the image. The purpose of the downsampling operation is to reduce the image size while maintaining the main feature information of the image. The downsampling operation can be achieved by combining convolution and pooling. In the present invention, the downsampling ratio r (ranging from [0.3, 0.7]) is set, and then a convolutional layer is used to implement downsampling:
[0069] x down,i = Conv2D(x i , K pooling , r)
[0070] where, x i ∈R H×W is the input image, K pooling is the convolutional kernel used for downsampling, usually with a size of k×k, and can be used for dimensionality reduction through convolution. r is the downsampling ratio, controlling the degree of downsampling. The convolutional layer (Conv2D) downsamples the image by applying a convolution operation. The size of the convolutional kernel and the stride affect the size of the image after pooling. Here, the convolution function is used to operate on the input image through the convolutional kernel, converting the local information of the input image into a higher-level feature representation. For the convolution operation, assume the size of the convolutional kernel K is k×k, then the output y of the convolution operation can be expressed as:
[0071]
[0072] x(i,j) is the pixel of the input image x i , and K(m,n) is the value of the convolutional kernel. The image obtained after downsampling is the image x down,i, which contains the low-frequency information (overall structural features) of the original image.
[0073] Step S4: Adjust the size of the original image (here, the original image refers to the image obtained after being processed in Step S2, that is, image A2) to be the same as that of the downsampled image (i.e., image A3) for subsequent difference calculation. In the present invention, bilinear interpolation or convolution operation can be used to adjust the size of the original image to be the same as that of the downsampled image.
[0074] Assume the original image is x i with a size of H×W, and the size of the downsampled image is H′×W′. We can adjust the size of the original image through interpolation or convolution as follows:
[0075] x scaled,i = Resize(x i , H′, W′)
[0076] Here, X scaled,i is the original image after size adjustment, with a size of H′×W′, which is the same as the downsampled image. The Resize operation scales the size of the original image H×W to the target size H′×W′.
[0077] To better adjust the image size, convolution can be used to assist in the adjustment. Assume we have a convolution kernel K resize for image scaling, and the image is adjusted through convolution operation as follows:
[0078] x scaled,i = Conv2D(x i , K resize , r′)
[0079] where K resize is the convolution kernel for adjusting the image size, and r′ is the scaling ratio coefficient. When downsampling the image in Step S3, the downsampling ratio coefficient needs to be controlled as r. Here, in order to obtain the difference map, we need to adjust the size of the original image proportionally, and the scaling ratio satisfies the condition r′ = r.
[0080] Step S5: After being processed in Steps 3 - 4, in this step, the image X scaled,i with its size scaled and adjusted is subtracted from the downsampled image x down,i to calculate the neighboring pixel difference (NDP) for extracting the high-frequency forgery traces of the image. By calculating the gray difference of adjacent pixels, a difference image is formed to enhance the discernibility of the forgery traces.
[0081] After scaling and adjusting the size of the original image, subtract it from the downsampled image to calculate the neighboring pixel difference (NDP) and extract the high-frequency forgery traces of the image. The specific steps are as follows:
[0082] NDP i = x scaled,i -x down,i
[0083] X scaled,i is the original image after resizing. x down,i is the image after downsampling, NDP i is the difference between adjacent pixels, reflecting the high-frequency forgery traces of the image.
[0084] The high-frequency forgery traces of the image are extracted through these steps, and these traces can effectively highlight the potential forgery features in the image generation process.
[0085] Step S6: Upsample the obtained difference map (i.e., image A4), and the output image is denoted as image A4. Image A4 not only ensures restoration to the original size for subsequent deep learning networks but also can, to a certain extent, amplify the difference traces by means of bilinear interpolation to make it easier to characterize the subtle traces between real and forged images.
[0086] First, to make the NDP feature map applicable to subsequent deep learning models (ResNet50), we use bilinear interpolation for upsampling. The goal is to restore the size of the NDP feature map to an appropriate size, usually the size of the original image or the downsampled image. The upsampled NDP feature map is denoted as NDP' (i.e., image A5), and its formula is as follows:
[0087] NDP i ' = Upsample(NDP i , r'')
[0088] where NDP is the initially calculated difference feature map of adjacent pixels, and r'' is the upsampling scale factor, which should satisfy:
[0089]
[0090] To make the extracted difference map applicable to subsequent classifiers, the difference map is here restored to the original size in equal proportion, that is, the size of the difference map is restored to the size of the original image. Upsample uses bilinear interpolation to expand the size of the NDP feature map to the target size. The formula for bilinear interpolation is:
[0091]
[0092] where NDP'(m, n) is the pixel value of the upsampled NDP feature map at the coordinate (m, n); w i and w jare the interpolation weights in the horizontal and vertical directions; NDP′(m+i,n+j) is the value of the neighboring pixel in the original NDP feature map.
[0093] During the image processing, downsampling and upsampling are performed through bilinear interpolation, further enhancing the ability to extract detailed features, making the forgery traces more obvious, thereby effectively improving the detection accuracy and robustness of forged images.
[0094] Step S7: Use a lightweight deep learning network to classify the extracted features and output the detection result. More optimally, according to the detection result, the present invention trains the model by using a cross-entropy loss function to further optimize the detection model to optimize the classification accuracy, that is: the present invention classifies the extracted NDP′ feature map by using a ResNet50 model and trains and optimizes it through a cross-entropy loss function. The specific operation steps are as follows:
[0095] S7.1 Input data: The upsampled feature map NDP′ is used as the input and fed into the ResNet50 network for classification. Its shape is transformed into H′×W′×C after transformation, where C is the number of channels, usually C = 3, and H′ and W′ are the height and width of the upsampled image.
[0096] S7.2 In the ResNet50 network, the image is first processed through a series of convolutional layers (Conv2D) to extract low-level and high-level features. If the input is NDP′, the formula for the first convolutional operation is:
[0097] x1 = Conv(NDP′, k1, s1, p1)
[0098] Among them, k1 is the convolutional kernel used to extract input features, and s1 and p1 are the stride and padding of the convolution respectively, which affect the size of the output feature map.
[0099] S7.3 Batch normalization: After convolution, we perform batch normalization on the result to improve the training stability and accelerate convergence. The formula for batch normalization is:
[0100] x2 = BatchNorm(x1)
[0101] Among them, x1 is the output of the convolutional layer, and x2 is the output after batch normalization.
[0102] S7.4 Activation function: After batch normalization, the ReLU activation function is used to introduce non-linearity and enhance the expression ability of the network. The formula for ReLU is:
[0103] x3 = ReLU(x2)
[0104] Among them, x2 is the output after batch normalization, and x3 is the output after activation.
[0105] S7.5 Residual Block: What makes ResNet50 special is that it uses residual connections, that is, cross-layer connections. In each residual block, the input and output are added through a skip connection to ensure that the gradient can flow directly to the previous layers, effectively alleviating the problem of gradient disappearance in deep networks. The formula for the residual block is as follows:
[0106] x4 = x3 + ResBLOCK(x3)
[0107] Where: x3 is the output after convolution, batch normalization, and activation; ResBlock is the output after convolution, batch normalization, and activation function. ResBlock is the result of residual block processing, usually including multiple convolutions and batch normalizations.
[0108] S7.6 Fully Connected Layer: After the residual block, the network classifies the features through a fully connected layer. The final feature is x n , and the output formula of the fully connected layer is:
[0109] y = FC(x n )
[0110] Where, x n is the feature after all convolutions and residual blocks; y is the output of the model, representing the probability that the image belongs to different classes (such as real or forged).
[0111] S7.7 Cross-Entropy Loss Function: During the training process, we use the cross-entropy loss function to calculate the difference between the predicted value and the true label and optimize the model parameters. The formula for the cross-entropy loss function is:
[0112] l = -∑ i y i log(p i )
[0113] Where, y i is the true label, usually 0 and 1, representing the true class of the image. p i is the class probability predicted by the model.
[0114] Step S8: Judge the authenticity of the image according to the model classification result and output the final detection result.
[0115] Combined with Figures 2 - 4 shown, an image forgery detection system based on pixel difference perception of the present invention includes: a data acquisition module, a position prospecting module, a pixel difference map extraction module, a model training module, and a classification and discrimination module. Specifically as follows:
[0116] Data acquisition module: used to obtain real image and forged image datasets from existing public datasets, and preprocess and annotate the images.
[0117] Position outlook module: used to enhance spatial perception ability through the Position Outlook Module (POM), and perform downsampling operations to extract low-frequency information.
[0118] Pixel difference map extraction module: used to calculate the neighboring pixel difference (NDP) and extract high-frequency forgery traces.
[0119] Model training module: used to classify the extracted features using a lightweight deep learning network, and optimize the detection model through the cross-entropy loss function.
[0120] Classification and discrimination module: used to detect the input image and determine whether the image is a forged image.
[0121] Example 1
[0122] Combined Figure 5 , in this embodiment, two images are input into the detection system for application testing, one is a real image and the other is a forged image generated by GAN. Figure 5 In the upper left figure on the left is the real image, and the lower left figure is the forged image; in the upper right figure is the feature map of the real image output after detection, with less feature change, and in the lower right figure is the feature map of the forged image output after detection, which can clearly show the forgery traces.
[0123] From the visualization effect diagram, the method of the present invention can effectively distinguish real images from forged images based on the extracted feature maps.
[0124] In addition, in this embodiment, we used 800 images (including 400 real images and 400 fake images) for testing. The classified accuracy rate, that is, the proportion of the number of correctly classified samples in the total number of samples, is 91.3%, and the classification precision A.P. (the proportion of samples actually belonging to a certain class among all samples classified as that class) reaches 96.6%. Especially when dealing with images with complex details, the extracted feature maps can significantly highlight the forgery traces and achieve the purpose of accurately detecting high-quality forged images.
[0125] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An image forgery detection method based on pixel difference perception, characterized in that: It includes the following steps: S1. Obtain a real image and a forged image, perform preprocessing and annotation, and output image A1; S2. Use a position anticipation module to process image A1, including processing the positions of image pixels through position encoding, and output image A2; S3. Perform downsampling on image A2 to extract low-frequency information of the image, and output image A3; S4. First, adjust the size of image A2 to be the same as that of image A3, and then subtract the pixel points of the resized image A2 from the pixel points of image A3 to calculate the adjacent pixel difference, forming a difference image, denoted as image A4; S5. Perform upsampling on image A4 to extract features, and output image A5; S6. Use a lightweight deep learning network to classify image A4 and output the detection result.
2. The image forgery detection method according to claim 1, characterized in that, In step S1, the preprocessing of the image includes randomly flipping, cropping, and decompressing the image, and after annotation, image data and corresponding labels are obtained.
3. The image forgery detection method according to claim 1, wherein Step S2 specifically includes the following steps: S2.
1. Input feature processing: First, divide the input image into multiple small image patches, and each image patch is input into the position anticipation module as an embedded character token. Each token will carry its spatial position information and be enhanced through position encoding; S2.
2. Use a sliding window mechanism to process the image. The size of the sliding window is adjusted according to the resolution of the image. During processing, the central token of each window will be associated with the surrounding tokens to capture the local fine-grained features of the image; S2.
3. Position anticipation attention processing: For each position (i, j), the position anticipation attention will calculate its similarity with the surrounding neighbors, and calculate the attention weight through the following steps: A i,j = Softmax(W Q ·V i,j ) Among them, V i,j is the relationship value between the central point and adjacent pixel points, and W Q is a linear transformation matrix used to generate weights; the role of the Softmax(·) function is to convert the output value of the model into a probability distribution; S2.
4. Feature Aggregation: The calculated attention weight A i,j will be used to weight the average of the neighbor features to generate a new feature representation, and this process is completed by the following formula: Y i,j = MatMul(A i,j , V i,j ) Among them, A i,j is the calculated attention weight, representing the similarity between each token; V i,j represents the relationship value between the central point and neighboring pixel points; MatMul(·) represents the matrix multiplication operation, multiplying the weight after Softmax normalization with V i,j to obtain the weighted output.
4. The image forgery detection method according to any one of claims 1-3, characterized in that, In step S3, bilinear interpolation or convolution operation is used for downsampling, and the downsampling ratio coefficient is r, where 0.3 ≤ r ≤ 0.
7.
5. The image forgery detection method according to claim 4, wherein In step S4, interpolation or convolution operation is used to scale the image size, and the scaling ratio coefficient is r′, where r′ = r.
6. The image forgery detection method according to claim 4, characterized in that, The formula for calculating the adjacent pixel difference in step S4 is as follows: NDP i = x scaled,i -x down,i Among them, X scaled,i is the pixel of the resized image A2, x down,i is the pixel of the image A3, and NDP i is the neighboring pixel difference.
7. The image forgery detection method according to claim 4, wherein In step S5, the upsampling operation uses bilinear interpolation, and the upsampling ratio coefficient is r″, where r″ = 1 / r.
8. The image forgery detection method according to any one of claims 1 to 3, characterized in that, In step S6, the lightweight deep learning network uses the ResNet50 network.
9. An image forgery detection system based on pixel difference perception, characterized in that, This system is used to execute the detection method described in any one of claims 1-8 above. It includes a data acquisition module, a position anticipation module, a pixel difference map extraction module, a model training module, and a classification and discrimination module, where: Data acquisition module: Used to obtain real image and forged image datasets, and perform preprocessing and annotation; Position anticipation module: Used to enhance the spatial perception ability through the position anticipation module and perform downsampling to extract low-frequency information; Pixel difference map extraction module: Used to calculate adjacent pixel differences and extract high-frequency forgery traces; Model training module: Used to classify the extracted features using a lightweight deep learning network and optimize the detection model through a cross-entropy loss function; Classification and discrimination module: Used to detect the input image and determine whether the image is a forged image.
10. An apparatus for detecting image forgery based on pixel difference perception, characterized in that: The device includes: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to cause the device to execute the method steps described in any one of claims 1-8.
Citation Information
Patent Citations
False image detection method and device based on image self-fusion, equipment and medium
CN119091516A
Cited By
AI generated video detection method and device, equipment and storage medium
CN121564810A
A method, apparatus, device, and storage medium for detecting AI-generated videos.
CN121564810B