Image depth forgery detection method

Through the image depth forgery detection method of binary convolution and frequency re-coupling binary module, the problem of taking into account both detection efficiency and accuracy is solved, and efficient forgery detection on low computing resource equipment is realized.

CN120339735AActive Publication Date: 2025-07-18NANJING UNIV OF POSTS & TELECOMM
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510830088.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

While the existing image depth forgery detection methods improve detection efficiency, there is a significant loss in detection accuracy, making it difficult to effectively detect on devices with limited computing resources.

Method used

The image features are extracted by binarized convolution, combined with frequency and re-coupled binary module and binary fast Fourier transform, and optimized features through jump connection and weight distribution, a lightweight detection model is generated, which is suitable for low-computing equipment.

Benefits of technology

Improve detection efficiency on low-computing equipment, enhance recognition of high-quality forged content, alleviate accuracy losses, and is suitable for the detection needs of different equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339735A_ABST
    Figure CN120339735A_ABST
Patent Text Reader

Abstract

The invention discloses an image depth forgery detection method in the technical field of image recognition, and aims to solve the technical problem that the detection efficiency and the detection precision are not considered at the same time. According to the image depth forgery detection method provided by the invention, features in an image are extracted through binarization convolution, so that the detection dimension is reduced, the detection efficiency is improved, and the method is applicable to equipment with low computing power; frequency domain features are extracted through binary fast Fourier transform to form complementation with spatial domain features, and the spatial domain and frequency domain feature complementation is mined to enhance the identification of high-quality counterfeit content; the jump connection retains original information, the weight distribution dynamically strengthens key features, the problem of precision loss is relieved, the number of stacked modules can be freely controlled, and the method is suitable for different devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image deepfake detection method, belonging to the technical field of image recognition. Background Art

[0002] The deepfake technology forges realistic human images by synthesizing or tampering with image, audio, and video content. The image deepfake detection method is a method for detecting forged images.

[0003] Existing deep detection models mainly rely on relatively complex models such as DenseNet, Vision Transformer (ViT), and Graph Convolutional Network (GCN) to achieve higher recognition performance. Existing deep detection models are generally deployed on professional computers or servers with high performance. However, deepfake images mainly appear on social media platforms or web applications, and devices such as mobile phones or personal computers have limited computing resources and are difficult to perform deepfake detection. For deep detection methods with lower requirements for computing resources, while improving the model processing efficiency by reducing the amount of computation, the number of parameters, and accelerating inference, the performance of the model in the deepfake detection task decreases.

[0004] Therefore, while the existing image deepfake detection methods improve the detection efficiency, there is an obvious loss in detection accuracy. Summary of the Invention

[0005] The purpose of this application is to overcome the deficiencies in the prior art and provide an image deepfake detection method that can ensure a reduction in the loss of detection accuracy while improving the detection efficiency.

[0006] To achieve the above purpose, the following technical solutions are adopted in this application:

[0007] In the first aspect, this application provides an image deepfake detection method, including:

[0008] Obtain the image to be detected, and extract the first feature map from the image to be detected based on binary convolution;

[0009] Input the first feature map extracted based on binary convolution into a detection model based on multiple cascaded frequency re-coupling binary modules to obtain a deep feature map, and obtain a detection score for obtaining a detection result according to the deep feature map;

[0010] Wherein, at least one of the frequency re-coupling binary modules is used to perform the following operations:

[0011] A. Obtain a binary convolution from the second feature map of the input module using a binary convolution layer;

[0012] B. Mining the frequency features of the binary convolution through binary fast Fourier transform, and using the mined frequency features as the main path features;

[0013] C. Performing skip connection by element-wise adding the second feature map and the main path features, and obtaining a weight distribution according to the features after skip connection;

[0014] D. Obtaining a third feature map of the module output after element-wise multiplying the main path features by the weight distribution.

[0015] In some embodiments of the first aspect, obtaining the weight distribution according to the features after skip connection includes:

[0016] Using a global average pooling layer to convert the features after skip connection into a first one-dimensional vector;

[0017] Using multiple fully connected layers to compress and expand the first one-dimensional vector;

[0018] Calculating and obtaining the weight distribution of the first one-dimensional vector output by the fully connected layer using the Sigmoid function.

[0019] In some embodiments of the first aspect, mining the frequency features of the binary convolution through binary fast Fourier transform to obtain the main path features includes:

[0020] Performing fast Fourier transform on the binary convolution to obtain the transformed features;

[0021] Extracting the amplitude vector and the phase vector from the transformed features;

[0022] Respectively extracting the frequency domain features of the amplitude vector and the frequency domain features of the phase vector;

[0023] Using inverse Fourier transform to reconstruct the image space features from the frequency domain features of the amplitude vector and the frequency domain features of the phase vector;

[0024] Obtaining the main path features through the residual connection between the reconstructed image space features and the input second feature map.

[0025] In some embodiments of the first aspect, obtaining the detection score for obtaining the detection result according to the deep feature map includes:

[0026] Using a global average pooling layer to process the deep feature map into a second one-dimensional vector;

[0027] Using a parametric rectified linear unit to activate the second one-dimensional vector to obtain an activated vector;

[0028] Obtaining the detection score by passing the activated vector through a fully connected layer.

[0029] In some embodiments of the first aspect,

[0030] Generate a detection model configured with multiple serially connected frequency recoupling binary modules, obtain a trained detection model through multiple training iterations, and input the first feature map extracted based on binary convolution into the trained detection model to obtain a deep feature map;

[0031] Each iteration of the detection model includes:

[0032] Input a training set into the detection model to obtain a prediction result;

[0033] Compare the prediction result with the true label of the corresponding sample in the training set, and calculate the loss value generated by the comparison according to the following binary cross-entropy loss function:

[0034] ,

[0035] where, is the loss value, is the counting serial number of the sample in the training set, is the label value 0 or 1 of the th sample, is the probability that the th sample is predicted to be 1,

[0036] is the number of samples;

[0037] Calculate the gradient of the model parameter loss of the detection model layer by layer in reverse starting from the loss value;

[0038] According to the gradient direction and a preset learning rate, adjust the detection model parameters to minimize the loss value.

[0039] In some embodiments of the first aspect, the obtaining of the image to be detected and the extraction of the first feature map from the image to be detected based on binary convolution includes,

[0040] Obtain the RGB image of the image to be detected;

[0041] Obtain the local binary pattern image of the RGB image according to the LBP algorithm;

[0042] Concatenate the RGB image and the local binary pattern image along the channel dimension to form four-channel input data, and perform binary convolution according to the four-channel input data to obtain the first feature map extracted based on binary convolution.In some embodiments of the first aspect, after the output features of the first frequency recoupling binary module among the multiple serially connected frequency recoupling binary modules are subjected to an element-wise multiplication residual connection with its input features, they become the input features of the next frequency recoupling binary module.

[0043] In a second aspect, the present application further provides a computer device, including a processor and a memory connected to the processor. A computer program is stored in the memory. When the computer program is executed by the processor, the steps of the image deepfake detection method according to any one of the embodiments of the first aspect are executed.

[0044] In a third aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the image deepfake detection method according to any one of the embodiments of the first aspect are implemented.

[0045] In a fourth aspect, the present application further provides a computer program product, including a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the steps of the image deepfake detection method according to any one of the embodiments of the first aspect are implemented.

[0046] Compared with the prior art, the beneficial effects achieved by the present application are as follows:

[0047] The image deepfake detection method provided by the present application reduces the detection dimension by extracting features in the image through binary convolution, improves the detection efficiency, and is suitable for devices with low computing power; extracts frequency domain features through binary fast Fourier transform, forms a complement with spatial domain features, and excavates the complement of spatial domain and frequency domain features to enhance the recognition ability of high-quality forged content; skip connections retain the original information, the weight distribution dynamically strengthens key features, alleviates the problem of accuracy loss, and can freely control the number of stacked modules, being suitable for different devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 is the flowchart of the steps of the image deepfake detection method provided in this embodiment;

[0050] Figure 2 is the model framework diagram of the image deepfake detection method;

[0051] Figure 3 isFigure 2 Frame flowchart of the medium-frequency re-coupling binary module;

[0052] Figure 4 is Figure 3 Frame flowchart of the fast Fourier transform in it;

[0053] Figure 5 is the step flowchart for obtaining the forgery result of the face image to be detected based on the face forgery dataset;

[0054] Figure 6 is the principle schematic block diagram of the computer device provided in this embodiment.

[0055] Description of the drawings: Figure 1 In (a) are the main execution steps of the image deep forgery detection method, and in (b) are the execution steps of the medium-frequency re-coupling binary module. Detailed implementation manners

[0056] The technical solution of the present invention will be described in detail below through the drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0057] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0058] Embodiment 1:

[0059] Figure 1 is the flowchart of an image deep forgery detection method in Embodiment 1 of the present invention. This flowchart only shows the logical order of the method described in this embodiment. On the premise of non-conflict, in other possible embodiments of the present invention, the steps shown or described can be completed in a different order Figure 1 from that shown.

[0060] The image deep forgery detection method provided in this embodiment can be applied to a terminal, such as: any smart phone, tablet computer or computer device with a communication function. Refer to Figure 1 , the method of this embodiment specifically includes the following steps:

[0061] Obtain the image to be detected, and extract the first feature map from the image to be detected based on binary convolution; the first feature map extracted based on binary convolution can be called a shallow feature map, and extracting the first feature map from the image to be detected based on binary convolution greatly reduces the computational amount and memory occupation, and is suitable for real-time detection on mobile devices or edge devices. In addition, the shallow feature map has shallow features (such as edges and textures), which provides a basis for subsequent in-depth analysis;

[0062] Input the first feature map extracted based on binary convolution into a detection model based on multiple cascaded frequency re-coupled binary modules to obtain a deep feature map, and obtain a detection score for obtaining a detection result according to the deep feature map; by stacking multiple frequency re-coupled modules (F-RBMs), a consistent lightweight architecture is formed to gradually extract deep features, enhance the model's ability to capture complex forgery traces, and at the same time reduce the requirements for the instantaneous computing power of the equipped device;

[0063] Wherein, at least one of the frequency re-coupled binary modules is used to perform the following operations:

[0064] A. Obtain a binary convolution from the second feature map of the input module by using a binary convolution layer; the binary convolution layer realizes further compression of features within the module and maintains computational efficiency;

[0065] B. Mine the frequency features of the binary convolution through binary fast Fourier transform, and use the mined frequency features as the main path features; perform fast Fourier transform / and fast Fourier transform under binary constraints to reduce the computational overhead in the frequency domain;

[0066] C. Perform a skip connection of element-wise addition of the second feature map and the main path features, and obtain a weight distribution according to the features after the skip connection; add the input features and the main path features through the skip connection to retain the original information, alleviate the problem of information loss caused by binarization, and further alleviate the problem of detection accuracy loss as a whole; and generate a weight distribution based on the features after the skip connection to adaptively strengthen key regions (such as forgery boundaries and abnormal textures) and suppress irrelevant noises;

[0067] D. After multiplying the main path features and the weight distribution element-wise, obtain the third feature map output by the module; dynamically scale the main path features through the weight distribution to enhance the significance of effective features, optimize the sensitivity of the model to subtle forgery signals, and also alleviate the problem of detection accuracy loss as a whole.

[0068] It should be noted that in the cascaded frequency re-coupled binary modules, the third feature map output by the previous frequency re-coupled binary module is directly used or used as the second feature map input to the next frequency re-coupled binary module after simple processing.

[0069] In summary, the image deepfake detection method provided in this embodiment extracts features in the image through binary convolution, reducing the detection dimension and improving the detection efficiency, meeting the requirements for application on devices with low computing power; extracts frequency domain features through binary fast Fourier transform, forming a complement with the spatial domain features, and mining the complementary enhancement of the spatial domain and frequency domain features to improve the recognition ability of high-quality forged content; the skip connection retains the original information, and the weight distribution dynamically strengthens the key features, alleviating the problem of accuracy loss, and the number of module stacks can be freely controlled, suitable for different devices.

[0070] Embodiment 2:

[0071] This embodiment provides an image deepfake detection method. This embodiment is optimized on the basis of Embodiment 1 to improve the technical effect and refine the technical solution. For the content not described in detail in this embodiment, please refer to Embodiment 1.

[0072] As one of the embodiments, obtaining the weight distribution according to the features after the skip connection includes obtaining the weight distribution through the Figure 3 perceptual compression and expansion branch on the right side in, specifically:

[0073] Using the global average pooling layer to convert the features after the skip connection into a first one-dimensional vector; the global average pooling layer can reduce the input dimension of the subsequent fully connected layer and significantly reduce the computational amount;

[0074] Using multiple fully connected layers to compress and expand the first one-dimensional vector; through two-step fully connected, while retaining the key information, reducing parameter redundancy and avoiding overfitting

[0075] Using the Sigmoid function to calculate and obtain the weight distribution of the first one-dimensional vector output by the fully connected layer; using the Sigmoid function to map the output of the fully connected layer to the [0,1] interval to generate channel attention weights, indicating the importance of each channel.

[0076] In this embodiment, the perceptual compression and expansion branch takes the two features before and after as inputs at the same time, perceives the change of the information flow, and thus adaptively adjusts the output distribution of the main path features.

[0077] As one of the embodiments, mining the frequency features of the binary convolution through the binary fast Fourier transform to obtain the main path features includes,

[0078] Performing a fast Fourier transform on the binary convolution to obtain the transformed features;

[0079] Extract the amplitude vector and the phase vector from the transformed features; respectively extract the frequency-domain features of the amplitude vector and the frequency-domain features of the phase vector; by obtaining the amplitude vector, the amplitude vector can reflect the signal energy distribution (such as synthetic artifacts, discontinuous textures), capture the high-frequency noise in the forged image, while the obtained phase vector reflects the encoded structure information, identify the geometric anomalies in the forged content (such as distorted facial contours, unnatural lighting); the amplitude vector and the phase vector achieve cross-domain feature complementarity, covering multi-dimensional representations for detecting forged traces, and both high-frequency noise and low-frequency structure anomalies can be detected, enhancing the robustness against complex forgery methods (such as GAN, Diffusion models);

[0080] Use the inverse Fourier transform to reconstruct the image space features from the frequency-domain features of the amplitude vector and the frequency-domain features of the phase vector; this step forms a closed-loop process with the fast Fourier transform of the previous binary convolution, avoiding the disconnection between frequency-domain and spatial-domain features and improving the model interpretability;

[0081] Obtain the main path features through the residual connection between the reconstructed image space features and the input second feature map; through this step, the frequency-domain analysis results (reconstructed features) and the spatial-domain input features are added through the residual connection, realizing cross-domain feature fusion; in addition, the residual structure alleviates the gradient disappearance problem in the training of deep networks and improves the model convergence stability.

[0082] In this embodiment, through the cooperation of frequency-domain deep mining and spatial-domain residual optimization, the detection accuracy of the lightweight model is significantly improved.

[0083] In addition, the binary fast Fourier transform can be expressed as:

[0084] ,

[0085] wherein, is the output result of the binary fast Fourier transform, that is, the transformed feature, is the input image, that is, the binary convolution, and are respectively the height and width of the input image, and are the coordinates of the input image, and are the coordinates in the corresponding Fourier space, is the imaginary unit;

[0086] For the convenience of calculation, it can be further expressed as:

[0087] ,

[0088] wherein, is a real number, is an imaginary number;

[0089] Therefore, the amplitude and phase can be calculated by the following formula:

[0090] ,

[0091] wherein, is the amplitude, is the phase.

[0092] In this embodiment, binary convolution is used multiple times. As one of the embodiments, the sign function adopted by the binary neural network model to binarize weights, activation values or other numerical values is:

[0093] ,

[0094] wherein, is the input content, is the output of the binary neural network model, is the judgment action.

[0095] As one of the embodiments, obtaining the detection score for obtaining the detection result according to the deep feature map includes

[0096] using a global average pooling layer to process the deep feature map into a second one-dimensional vector;

[0097] using a parametric rectified linear unit to activate the second one-dimensional vector to obtain an activated vector;

[0098] passing the activated vector through a fully connected layer to obtain a prediction probability, and taking the prediction probability as the detection score.

[0099] As one of the embodiments, referring to Figure 5 , a detection model configured with multiple serially connected frequency re-coupling binary modules is generated, and a trained detection model is obtained through multiple training iterations. The first feature map extracted based on binary convolution is input into the trained detection model to obtain a deep feature map;

[0100] Each iteration of the detection model includes:

[0101] inputting a training set into the detection model to obtain a prediction result;

[0102] comparing the prediction result with the true label of the corresponding sample in the training set, and calculating the loss value generated by the comparison according to the following binary cross-entropy loss function:

[0103] ,

[0104] wherein, is the loss value, is the counting serial number of the samples in the training set, is the th sample's label value 0 or 1, is the probability that the th sample is 1; The binary cross-entropy loss function perfectly matches the binary classification task (true / false) of deepfake detection, is the number of samples;

[0105] Backwardly calculate the gradient of the model parameter loss of the detection model layer by layer starting from the loss value;

[0106] According to the gradient direction and the preset learning rate, adjust the detection model parameters to minimize the loss value.

[0107] As one of the embodiments, the obtaining the image to be detected and extracting the first feature map from the image to be detected based on binary convolution includes,

[0108] Obtain the RGB image of the image to be detected;

[0109] Obtain the local binary pattern map of the RGB image according to the LBP algorithm; The local binary pattern map can supplement more details, make up for the problem that the RGB image is insensitive to texture, and the LBP map is robust to illumination changes and is suitable for forgery detection under complex illumination conditions;

[0110] Concatenate the RGB image and the local binary pattern map along the channel dimension to form four-channel input data, and perform binary convolution according to the four-channel input data to obtain the first feature map extracted based on binary convolution. The combination of RGB and LBP covers the features of different frequency bands in the spatial domain, and the combined image is convenient for the subsequent binary convolution layer to mine the frequency features and improve the detection ability of the model for complex forgery methods.

[0111] As one of the embodiments, the output feature of the first frequency re-coupling binary module in the plurality of cascaded frequency re-coupling binary modules is subjected to an element-wise multiplication residual connection with its input feature and then becomes the input feature of the next frequency re-coupling binary module. This step aims to calibrate the output of the first frequency re-coupling binary module, and this operation (element-wise multiplication) helps the feature fusion and propagation between modules.

[0112] As one of the embodiments, after obtaining the image to be detected, for the convenience of processing, normalize the longest dimension to 256 pixels.

[0113] In addition, data in the training set needs to be augmented, including random flipping (horizontal and vertical), rotation (90° or 270°), and color jitter within the range of 0% to 20%. The COCOFake dataset can be used as the training set.

[0114] As one of the embodiments, referring to Figure 2 , a deepfake detection network can be generally constructed to perform the main steps of the image deepfake detection method provided in this embodiment. Figure 2 In the shown deepfake detection network, on the left are the RGB image and the LBP image obtained from the RGB image, that is, the local binary pattern map. 3×3 is the convolutional kernel size, and the same applies hereinafter. In this embodiment, a shallow feature extraction network is constructed based on two cascaded frequency re-coupling binary modules and a quantization convolutional layer to extract feature maps from the image to be detected, and the product is the basic binary neural network input composed of n cascaded frequency re-coupling modules. And a normalization layer and an element-wise addition residual operation are adaptively set behind each convolutional layer. Finally, a global average pooling layer, an activation function, and a quantization linear layer are used to process the prediction probability, and the prediction probability is used as the detection score.

[0115] Referring to Figure 3 On the left is the main path, n×n is the convolutional kernel size of the binary convolutional layer, and the weights are obtained by the main path features and the input features through the perception compression and expansion branch on the right.

[0116] Figure 4 This is a further elaboration of the binary fast Fourier transform module.

[0117] In this embodiment, on the 64-bit Ubuntu 18.04.5 operating system, based on software environments such as Python 3.8.0, torch 1.9.1, and torchvision 0.10.1, the training of the model is completed using an NVIDIA RTX 3090 GPU. The AdamW optimizer is used, and the initial learning rate is , the first moment estimate and the second moment estimate are set to 0.9 and 0.999 respectively, and the weight decay parameter is .

[0118] The image deepfake detection method provided in this embodiment uses a binary neural network to predict the detection result. Compared with other deepfake detection models based on architectures such as Transformer or graph convolution, binarizing the weights and activation values in the network can effectively improve the computational efficiency of the model. It also perceives changes in frequency domain and texture information through a recoupling method, adaptively adjusts the feature output distribution, and helps improve the detection accuracy of the model. This embodiment also proposes a binary fast Fourier transform module, which extracts frequency domain features from two dimensions of phase and amplitude respectively, and mines the forgery traces of the image in the frequency domain, helping to improve the robustness of the forgery detection model.

[0119] Embodiment 3:

[0120] This embodiment provides a computer device, including a processor and a memory connected to the processor. A computer program is stored in the memory. When the computer program is executed by the processor, the steps of the image deepfake detection method provided in Embodiment 1 or 2 are executed.

[0121] This computer device can be a server or an electronic terminal. As one of the embodiments, referring to Figure 6 , this computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of this computer device is used to store the data obtained and generated in the image deepfake detection method. The input / output interface of this computer device is used for the processor to exchange information with external devices. The communication interface of this computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes the image deepfake detection method provided in Embodiment 1 or 2.

[0122] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0123] The computer device provided in this embodiment has the same technical effects as those in Embodiment 1 or 2, and will not be elaborated here.

[0124] Example 4:

[0125] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the image deep forgery detection method provided in Embodiment 1 or Embodiment 2 are implemented.

[0126] The computer-readable storage medium provided in this embodiment has the same technical effects as those in Embodiment 1 or 2, and will not be elaborated herein.

[0127] Example 5:

[0128] This embodiment provides a computer program product, on which a computer program is stored. When the program is executed by a processor, the steps of the image deep forgery detection method provided in Embodiment 1 or Embodiment 2 are implemented. The computer program product provided in this embodiment can be transmitted, distributed, and downloaded in the form of a signal through the Internet.

[0129] The computer program product provided in this embodiment has the same technical effects as those in Embodiment 1 or 2, and will not be elaborated herein.

[0130] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0131] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0132] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions in the processFigure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0133] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0134] In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "multiple" is two or more.

[0135] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific circumstances.

[0136] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can still be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. An image deepfake detection method, characterized in that, including, obtaining an image to be detected, and extracting a first feature map from the image to be detected based on binary convolution; inputting the first feature map extracted based on binary convolution into a detection model based on multiple cascaded frequency re-coupled binary modules to obtain a deep feature map, and obtaining a detection score for obtaining a detection result according to the deep feature map; wherein at least one of the frequency re-coupled binary modules is used to perform the following operations: A. obtaining a binary convolution from a second feature map input to the module by using a binary convolution layer; B. mining frequency features of the binary convolution through binary fast Fourier transform, and using the mined frequency features as main path features; C. performing skip connection of element-wise addition on the second feature map and the main path features, and obtaining a weight distribution according to the features after skip connection; D. multiplying the main path features and the weight distribution element-wise to obtain a third feature map output by the module.

2. The image deep forgery detection method according to claim 1, wherein The obtaining a weight distribution according to the features after skip connection includes: using a global average pooling layer to convert the features after skip connection into a first one-dimensional vector; using multiple fully-connected layers to compress and expand the first one-dimensional vector; calculating and obtaining a weight distribution of the first one-dimensional vector output by the fully-connected layer by using a Sigmoid function.

3. The image deep forgery detection method according to claim 1, characterized in that, The mining frequency features of the binary convolution through binary fast Fourier transform to obtain main path features includes: performing fast Fourier transform on the binary convolution to obtain transformed features; extracting an amplitude vector and a phase vector from the transformed features; respectively extracting frequency domain features of the amplitude vector and frequency domain features of the phase vector; using inverse Fourier transform to reconstruct image space features from the frequency domain features of the amplitude vector and the frequency domain features of the phase vector; obtaining main path features through a residual connection between the reconstructed image space features and the input second feature map.

4. The method for detecting image deep forgery according to claim 1, wherein The obtaining a detection score for obtaining a detection result according to the deep feature map includes: using a global average pooling layer to process the deep feature map into a second one-dimensional vector; using a parametric rectified linear unit to activate the second one-dimensional vector to obtain an activated vector; obtaining a detection score by passing the activated vector through a fully-connected layer.

5. The method for detecting image deep forgery according to claim 1, characterized in that a trained detection model is obtained through multiple training iterations; each iteration includes: inputting a training set into the detection model to obtain a prediction result; comparing the prediction result with the true label of the corresponding sample in the training set, and calculating a loss value generated by the comparison according to the following binary cross-entropy loss function: , Wherein, is the loss value, is the counting serial number of the samples in the training set, is the label value 0 or 1 of the th sample, is the probability that the th sample is 1, and is the number of samples; calculating the gradient of the model parameter loss of the detection model layer by layer in reverse starting from the loss value; adjusting the detection model parameters according to the gradient direction and a preset learning rate to minimize the loss value.

6. The method for detecting image deep forgery according to claim 1, wherein The obtaining an image to be detected, and extracting a first feature map from the image to be detected based on binary convolution includes: obtaining an RGB image of the image to be detected; obtaining a local binary pattern image of the RGB image according to the LBP algorithm; The RGB image and the local binary pattern map are concatenated according to the channel dimension to form four-channel input data, and binary convolution is performed based on the four-channel input data to obtain the first feature map extracted based on binary convolution.

7. The method for detecting image deep forgery according to claim 1, wherein In the plurality of cascaded frequency re-coupled binary modules, the output feature of the first frequency re-coupled binary module is subjected to element-wise multiplication with its input feature, and the residual connection becomes the input feature of the next frequency re-coupled binary module.

8. A computer device, characterized in that, It includes a processor and a memory connected to the processor. A computer program is stored in the memory. When the computer program is executed by the processor, the steps of the image deepfake detection method according to any one of claims 1 to 7 are executed.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the image deepfake detection method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the image deepfake detection method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • GAN generated picture detection method and system based on residual domain rich model

    CN110516575A

  • Deep forgery detection method and corresponding device

    CN115311525A

  • Deep forgery detection method, system and equipment for low-quality face image

    CN117423148A

  • Depth forgery detection method based on mask supervision

    CN119540736A

  • Lightweight remote sensing image change detection method based on binary neural network and large kernel stripe convolution

    CN119672295A