Face forgery detection method based on frequency mask and attention consistency

CN117373136BActive Publication Date: 2026-09-29XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310833690.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-07
Publication Date
2026-09-29
Estimated Expiration
2043-07-07

AI Technical Summary

Technical Problem

[0005]以上可知,现有的基于频域伪造检测的方法大多使用所有频率信息作为重要的伪造线索,导致网络学到的伪造线索并不鲁棒,所以这些方法在不可见域的真伪判定泛化能力不理想

Benefits of technology

[0048]1、由于用于训练模型的参考图像是对原始图像进行离散傅里叶变换和高频信息随机丢弃处理后得到的图像,所以,本发明通过离散傅里叶变换获得图片的高频分量与低频分量,同时应用随机丢弃高频信息来探索了鲁棒伪造线索,从而增加了模型在在不可见伪造样本的泛化性,也使得训练出的模型的检测能力更优;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117373136B_ABST
    Figure CN117373136B_ABST
Patent Text Reader

Abstract

The application discloses a face forgery detection method based on frequency mask and attention consistency, comprising: inputting a pre-trained forgery detection model after obtaining a to-be-detected video to obtain a detection result representing whether the to-be-detected video is a forged video; the pre-trained forgery detection model is obtained by iteratively training an initial pre-trained forgery detection model by using multiple reference images with labels, a sample video, a cross-forgery attention consistency loss function and a cross-entropy loss function; the initial pre-trained forgery detection model is obtained by iteratively training a to-be-trained forgery detection model by using multiple reference images and a cross-entropy loss function; each reference image is obtained by performing discrete Fourier transform and high-frequency information random discarding processing on a corresponding original image; and the cross-forgery attention consistency loss function is used to measure the difference between the gradient weighted class activation mapping diagram of the reference image and the video frame in the sample video. The application can improve the generalization of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine vision technology, specifically relating to a face forgery detection method based on frequency masking and attention consistency. Background Technology

[0002] With the successful development of depth image generation technology, visual data forgery detection will play an increasingly important role in the field of socio-economic security.

[0003] Frequency Domain Forgery Detection: Early frequency domain face forgery detection tasks focused on developing models capable of detecting most known forgeries. These models can be divided into two categories: those using spatial cues and those using temporal cues. For frequency domain face forgery detection based on spatial cues, traditional face recognition methods are often referenced. For example, Zhang et al. pioneered a face forgery detection technique that uses multiple classifiers to distinguish forged images after feature extraction. Building on this, subsequent researchers discovered some visual detail mismatches in forged face images and utilized them for detection. For instance, He et al. combined a random forest classifier to identify forged images based on image residual information in different color spaces. Zhang revealed that forged images generated by GAN models possess a special distortion feature and distinguished it in the frequency domain spectrogram. Luo et al. preprocessed images using a high-pass filter to enhance detection performance. Dong et al. proposed using image preprocessing, such as Gaussian blur and noise, to eliminate the difference in high-frequency distortion between forged and real images, thereby better learning forgery features and improving detection performance. For time-clue-based frequency domain face forgery detection, a common drawback of forged videos is the discontinuity between video frames, which can be used as a basis for detecting forged videos. For example, Masi et al. amplified artifacts in the video stream to detect fake faces and suppress advanced face content. Another example is Wu et al., who extracted hidden analysis features on convolutional filters to detect hidden tampering artifacts.

[0004] Generalized Forgery Detection: Generalized forgery detection refers to distinguishing unseen forgery types in cross-dataset evaluation. Most existing face forgery methods often fail to detect unseen artifacts. To address this issue, many researchers utilize deep learning models and data augmentation to improve generalization performance. Some researchers have explored using noise to mask facial recognition features and extract forgery clues that are difficult to tamper with. Wang et al. applied Gaussian blur to reduce the influence of high-frequency forgery features and introduced adversarial training into the training of a deep forgery detection model, significantly improving the model's generalization ability and robustness. Guo et al. designed an adaptive forgery trace extraction network, which can be used as a preprocessing step to remove image content and emphasize manipulation traces. Building on this, Wang et al. proposed a method for removing structured information region features. Guo et al. segmented candidate face images into strongly correlated and weakly correlated regions and designed an MSF module to enhance the detector's robustness to these regions by dynamically erasing some locations.

[0005] As can be seen from the above, most existing frequency domain-based forgery detection methods use all frequency information as important forgery clues, which makes the forgery clues learned by the network not robust. Therefore, these methods have poor generalization ability in determining authenticity in the invisible domain. Summary of the Invention

[0006] To address the aforementioned problems in related technologies, this invention provides a face forgery detection method based on frequency masking and attention consistency. The technical problem to be solved by this invention is achieved through the following technical solution:

[0007] This invention provides a face forgery detection method based on frequency masking and attention consistency, comprising:

[0008] Obtain the video to be tested;

[0009] The video to be detected is input into a pre-trained fake detection model to obtain the detection result of the video to be detected; the detection result indicates whether the video to be detected is a fake video;

[0010] The pre-trained forgery detection model is obtained by iteratively training an initial pre-trained forgery detection model using multiple labeled reference images, sample videos, a cross-forgery attention consistency loss function, and a cross-entropy loss function; the initial pre-trained forgery detection model is obtained by iteratively training the forgery detection model to be trained using the multiple reference images and the cross-entropy loss function.

[0011] Each reference image is obtained by performing a discrete Fourier transform on the corresponding original image and randomly discarding high-frequency information; the cross-forgery attention consistency loss function is used to measure the difference between the gradient-weighted class activation map of the reference image and the gradient-weighted class activation map of the video frames in the sample video.

[0012] In some embodiments, the sample video is the video to be detected or a video other than the video to be detected; the plurality of reference images include real images and fake images.

[0013] In some embodiments, the expression for the cross-forgery attention consistency loss function is:

[0014]

[0015] Among them, L ac For the cross-forgery attention consistency loss value, N b N is the total number of video frames in the sample video. r f(x) represents the total number of images selected from the plurality of reference images during each training iteration. i f(x) is the gradient-weighted class activation map of the i-th video frame in the sample video. j Let |j| be the gradient-weighted class activation map of the j-th image selected from the plurality of reference images during each training iteration, where |j| represents the L2 norm.

[0016] In some embodiments, before inputting the video to be detected into a pre-trained forgery detection model to obtain the detection result of the video to be detected, the method includes:

[0017] Acquire multiple original images, including real and fake images; each original image is labeled; the label indicates whether the original image is a fake image;

[0018] Each of the multiple original images is subjected to discrete Fourier transform and high-frequency information is randomly discarded to obtain multiple processed images;

[0019] The forgery detection model to be trained is iteratively trained using the multiple processed images and the cross-entropy loss function to obtain a basic training model;

[0020] The basic training model is used as the initial pre-trained forgery detection model and the auxiliary model, respectively. The multiple processed images are used as the multiple reference images, and the sample video is obtained.

[0021] The sample video is used as the input to the initial pre-trained forgery detection model, and the multiple reference images are used as the input to the auxiliary model. The initial pre-trained forgery detection model is iteratively trained based on the gradient-weighted class activation map generated by the initial pre-trained forgery detection model for each video frame and the gradient-weighted class activation map generated by the auxiliary model for each reference image, to obtain the pre-trained forgery detection model.

[0022] In some embodiments, performing Discrete Fourier Transform and random high-frequency information discarding on each of the plurality of original images to obtain a plurality of processed images includes:

[0023] Perform a discrete Fourier transform on each of the multiple original images to obtain the frequency domain information of the original images;

[0024] The frequency domain information is processed to obtain processed frequency domain information; the low-frequency information component in the processed frequency domain information is located in the central region, and the high-frequency information is located in the edge region.

[0025] Based on the processed frequency domain information and the size of the original image, determine the distance corresponding to the position of each spectrum.

[0026] The rectangular mask is determined based on the relationship between the distance and the preset distance threshold;

[0027] After multiplying the frequency domain information with the rectangular mask, an inverse discrete Fourier transform is performed to obtain the high-frequency component, and the low-frequency component is obtained based on the original image and the high-frequency component.

[0028] The high-frequency component is multiplied by a preset random matrix to obtain the processed high-frequency component; the preset random matrix consists of 0 and 1.

[0029] Based on the processed high-frequency components and the low-frequency components, the processed image corresponding to the original image is obtained.

[0030] In some embodiments, the expression for the processed image corresponding to the original image is:

[0031]

[0032] Where, x i' This represents the i'th original image. x represents i' The corresponding processed image, F represents the discrete Fourier transform, F -1 The inverse discrete Fourier transform, z i' x represents i' The frequency domain information, This refers to the rectangular mask. This refers to the high-frequency components. M represents the low-frequency component. i' Let M represent the preset random matrix; where M is the random matrix. i' and Size and x i' They are the same size.

[0033] In some embodiments, the step of using the sample video as input to the initial pre-trained forgery detection model, and the multiple reference images as input to the auxiliary model, and iteratively training the initial pre-trained forgery detection model based on the gradient-weighted class activation map generated for each video frame by the initial pre-trained forgery detection model, and the gradient-weighted class activation map generated for each reference image by the auxiliary model, to obtain the pre-trained forgery detection model, includes:

[0034] During the t-th training iteration, the sample video is input into the t-th training model to generate the t-th gradient-weighted class activation map for each video frame of the sample video. Simultaneously, the t-th training samples are obtained from the multiple reference images and input into the auxiliary model to generate the gradient-weighted class activation map for each training sample in the t-th training sample, as well as the prediction result for each training sample in the t-th training sample; t is an integer greater than 0; when t is 1, the t-th training model is the initial pre-trained forgery detection model;

[0035] The total loss value for the tth time is calculated based on the gradient-weighted class activation map, the cross-spoofing attention consistency loss function, the cross-entropy loss function, the label of each training sample in the tth training sample, the gradient-weighted class activation map, and the prediction result.

[0036] Adjust the network parameters of the model to be trained in the t-th iteration based on the total loss value in the t-th iteration to obtain the model to be trained in the (t+1)-th iteration.

[0037] The training sample (t+1)th iteration is obtained from the multiple reference images. Based on the training sample (t+1th iteration), the sample video, the auxiliary model, the cross-spoofing attention consistency loss function, and the cross-entropy loss function, the training model (t+1th iteration) is trained until the pre-trained spoofing detection model is obtained.

[0038] In some embodiments, the expression for the total loss value at the t-th time is as follows:

[0039]

[0040] Where L is the total loss value for the t-th iteration, and N b N is the total number of video frames in the sample video. r f is the total number of training samples in the t-th training iteration. t (x i () represents the gradient-weighted class activation map of the t-th time in the i-th video frame of the sample video. For N r The j-th image, for The gradient-weighted class activation map, ||.||2 represents the L2 norm, y j,k for The tag, for The prediction result, H φ (.) represents the classifier head in the t-th training model, g θ Let be the network parameters of the model to be trained for the tth time.

[0041] In some embodiments, the step of iteratively training the forgery detection model to be trained using the multiple processed images and the cross-entropy loss function to obtain a basic training model includes:

[0042] During the c-th training iteration, the c-th training sample is obtained from the multiple processed images, and the c-th training sample is input into the c-th forgery detection model to obtain the prediction result of the c-th training sample; c is an integer greater than 0; when c is 1, the c-th forgery detection model is the forgery detection model to be trained.

[0043] Calculate the cross-entropy loss value for the c-th training sample based on the prediction result of the c-th training sample, the label of the c-th training sample, and the cross-entropy loss function.

[0044] Adjust the network parameters of the c-th forgery detection model based on the c-th cross-entropy loss value to obtain the (c+1)-th forgery detection model.

[0045] The c+1th training sample is obtained from the multiple processed images, and the c+1th forgery detection model is trained based on the c+1th training sample and the cross-entropy loss function until the basic training model is obtained.

[0046] In some embodiments, the multiple reference images are multiple RGB face images; the sample video is a video containing face regions.

[0047] The present invention has the following beneficial technical effects:

[0048] 1. Since the reference image used to train the model is the original image after performing a discrete Fourier transform and randomly discarding high-frequency information, this invention obtains the high-frequency and low-frequency components of the image through discrete Fourier transform, and at the same time applies the random discarding of high-frequency information to explore robust forgery clues, thereby increasing the model's generalization to invisible forgery samples and making the trained model's detection ability better.

[0049] 2. Because the cross-spoofed attention consistency loss function was used when training the model, the spoofed attention across different datasets was considered. Figure 1 The impact of consistency on the generalization of the model allows the model to focus on similar attention regions, thereby increasing the model's generalization ability on invisible forged samples.

[0050] 3. Since the sample video is the video to be detected, this invention can continue to train the initial pre-trained forgery detection model using the reference image, the video to be detected, and the cross-forgery attention consistency loss function. Then, the trained model is used to detect the video to be detected. Therefore, this invention can use the reference image and the video to be detected to fine-tune the features learned by the model. This makes the features extracted by the model more resolution and robust, thereby increasing the model's generalization ability in invisible forgery samples and making the model's detection ability better.

[0051] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0052] Figure 1 An optional flowchart of a face forgery detection method based on frequency masking and attention consistency provided in an embodiment of the present invention;

[0053] Figure 2 An exemplary test flowchart of the face forgery detection method based on frequency masking and attention consistency provided for embodiments of the present invention. Detailed Implementation

[0054] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0055] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0056] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0057] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, disclosure, and appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0058] Most existing frequency-domain forgery detection methods described in the background section use all frequency information as important forgery clues. Therefore, these methods perform well on test videos or images containing the same forgery types as the training data. However, because existing forgery datasets have obvious frequency clues, previously discovered forgery patterns often fail when detecting different types of forged faces.

[0059] The data augmentation-based methods described in the background art cannot effectively utilize the specific characteristics of forged videos from different individuals. Because different types of forgery methods produce data with significantly different distributions, these methods may perform well on some but not on all forged videos.

[0060] Existing methods do not simultaneously consider the generalization of high-frequency information and the difficulty in utilizing the specific features of different individuals to forge videos. This is something that has not been explored in previous work. Therefore, the final forgery features learned by the method in this application can be more generalized and better reflect the characteristics of different individual videos.

[0061] Figure 1 This is a flowchart of a face forgery detection method based on frequency masking and attention consistency provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes the following steps:

[0062] S101. Obtain the video to be tested.

[0063] Here, the video to be detected can be any type of video, for example, it can be a video of a face that contains a face region.

[0064] S102. Input the video to be detected into the pre-trained forgery detection model to obtain the detection result of the video to be detected; the detection result indicates whether the video to be detected is a forgery video; the pre-trained forgery detection model is obtained by iteratively training the initial pre-trained forgery detection model using multiple labeled reference images, sample videos, cross-forgery attention consistency loss function and cross-entropy loss function; the initial pre-trained forgery detection model is obtained by iteratively training the forgery detection model to be trained using multiple reference images and cross-entropy loss function; each reference image is an image obtained after performing discrete Fourier transform and randomly discarding high-frequency information from the corresponding original image; the cross-forgery attention consistency loss function is used to measure the difference between the gradient-weighted class activation map of the reference image and the gradient-weighted class activation map of the video frame in the sample video.

[0065] Here, the pre-trained forgery detection model can be a pre-trained Xception network.

[0066] Here, the multiple reference images include real images and fake images; for example, the multiple reference images can be multiple RGB face images including real face images and fake face images.

[0067] In some embodiments, the sample video can be the video to be detected as described above; in other embodiments, the sample video can be any video other than the video to be detected. When the sample video is the video to be detected as described above, the pre-trained forgery detection model will produce more accurate detection results for the video to be detected.

[0068] In some embodiments, prior to S102, the method further includes:

[0069] S201. Obtain multiple original images containing real and fake images; each original image is labeled; the label indicates whether the original image is a fake image.

[0070] Here, multiple original images constitute an original dataset. Each original image is labeled with 0 or 1, which is used to characterize whether the original image is a fake image or a real image.

[0071] S202. Perform Discrete Fourier Transform and random high-frequency information discarding on each of the multiple original images to obtain multiple processed images.

[0072] Specifically, it can be done for each of the multiple original images x i' ∈RH×W×3 Perform a discrete Fourier transform to obtain the frequency domain information z of the original image. i' For frequency domain information z i' The process yields processed frequency domain information; the low-frequency components in the processed frequency domain information are located in the central region, while the high-frequency components are located in the edge region; based on the processed frequency domain information and the original image x... i' Given the size H×W, determine the distance Dis(u,v) corresponding to the position (u,v) of each spectrum, and The rectangular mask is determined based on the relationship between the distance Dis(u,v) and the preset distance threshold d (used to divide different frequency components). Specifically, components with a distance greater than d from Dis(u,v) are categorized as high-frequency components, and components with a distance less than d from Dis(u,v) are categorized as low-frequency components; the frequency domain information z... i' and rectangular mask After multiplication, an inverse discrete Fourier transform is performed to obtain the high-frequency components. Right now And based on the original image x i' and high frequency components Obtain low-frequency components Right now For the preset random matrix M i' With high frequency components Perform a multiplication operation to obtain the processed high-frequency components. The preset random matrix consists of 0s and 1s; based on the processed high-frequency components h i' h ×M i' and low frequency components Obtain the processed image corresponding to the original image. Right now F represents the Discrete Fourier Transform, F -1 M represents the inverse discrete Fourier transform. i' and Size and x i' They are the same size.

[0073] Here, this application assumes that the center of the Fourier spectrum is the origin. Based on this assumption, the distance is defined in the frequency domain, and the boundary between the low-frequency component and the high-frequency component is set as Dis(u,v).

[0074] Here, the preset distance threshold d can be set according to actual needs.

[0075] S203. Using multiple processed images and the cross-entropy loss function, iteratively train the forgery detection model to be trained to obtain the basic training model.

[0076] Specifically, during the c-th training iteration, c-th training samples are obtained from multiple processed images and input into the c-th forgery detection model to obtain the prediction results of the c-th training samples; c is an integer greater than 0; when c is 1, the c-th forgery detection model is the forgery detection model to be trained; the c-th cross-entropy loss value is calculated based on the prediction results of the c-th training samples, the labels of the c-th training samples, and the cross-entropy loss function; the network parameters of the c-th forgery detection model are adjusted based on the c-th cross-entropy loss value to obtain the c+1-th forgery detection model; the c+1-th training samples are obtained from multiple processed images, and the c+1-th forgery detection model is trained based on the c+1-th training samples and the cross-entropy loss function until the basic training model is obtained.

[0077] Here, the cross-entropy loss function L cls The expression is: Where N is the total number of training samples in each training session. Let y be the i-th image in N. i',k for The tag, for The prediction results, H φ (.) represents the classifier head in the current training model, g θ These are the network parameters of the model to be trained in the current iteration.

[0078] Here, training can be stopped when the preset number of training iterations is reached, and the fake detection model obtained from the last training iteration can be used as the base training model. Alternatively, training can be stopped when the cross-entropy loss value obtained from several consecutive training iterations tends to stabilize, and the fake detection model obtained from the last training iteration can be used as the base training model.

[0079] S204. Use the basic training model as the initial pre-trained fake detection model and the auxiliary model respectively, use multiple processed images as multiple reference images, and obtain sample videos.

[0080] S205. The sample video is used as the input to the initial pre-trained forgery detection model, and multiple reference images are used as the input to the auxiliary model. The initial pre-trained forgery detection model is iteratively trained based on the gradient weighted class activation map generated by the initial pre-trained forgery detection model for each video frame and the gradient weighted class activation map generated by the auxiliary model for each reference image, to obtain the pre-trained forgery detection model.

[0081] Specifically, during the t-th training iteration, the sample video is input into the t-th training model to generate the t-th gradient-weighted class activation map for each video frame of the sample video. Simultaneously, t-th training samples are obtained from multiple reference images and input into the auxiliary model to generate the gradient-weighted class activation map for each training sample in the t-th training sample, as well as the prediction result for each training sample in the t-th training sample; t is an integer greater than 0; when t is 1, the t-th training model is the initial pre-trained forgery detection model; based on the t-th gradient-weighted class activation map and cross-forgery annotation... The attention consistency loss function, cross-entropy loss function, and the labels, gradient-weighted class activation maps, and prediction results of each training sample in the t-th training sample are used to calculate the total loss value for the t-th training. The network parameters of the model to be trained in the t-th training are adjusted according to the total loss value for the t-th training to obtain the model to be trained in the t+1-th training. The training samples for the t+1-th training are obtained from multiple reference images, and the model to be trained in the t+1-th training is trained according to the training samples for the t+1-th training, the sample video, the auxiliary model, the cross-spoofing attention consistency loss function, and the cross-entropy loss function until a pre-trained spoofing detection model is obtained.

[0082] For example, the expression for the total loss value in the t-th iteration is as follows:

[0083]

[0084] Where L is the total loss value in the t-th iteration, and N b N represents the total number of video frames in the sample video. r Let ft(x) be the total number of training samples in the t-th training iteration. i Let be the gradient-weighted class activation map of the t-th video frame in the sample video. For N r The j-th image, for The gradient-weighted class activation map, ||.||2 represents the L2 norm, y j,k for The tag, for The prediction results, H φ (.) represents the classifier head in the t-th training model, g θ Let be the network parameters of the model to be trained in the t-th iteration.

[0085] Here, the gradient-weighted class activation map (Grad-CAM) of the image is generated by the model during the process of generating the prediction result of the image.

[0086] Here, training can be stopped when the number of training iterations reaches a preset number, and the model to be trained obtained from the last training iteration can be used as the pre-trained fake detection model. Alternatively, training can be stopped when the cross-entropy loss value obtained from several consecutive training iterations tends to stabilize, and the model to be trained obtained from the last training iteration can be used as the pre-trained fake detection model.

[0087] To evaluate the effectiveness and generalization of the method proposed in this application, we selected several challenging face spoofing datasets for experiments, including the FaceForensics++ dataset, the DFD dataset, the Celeb-DF dataset, and the WDF dataset.

[0088] The FF++ dataset contains 4,000 fake face videos generated using four face spoofing methods (DeepFake, Face2Face, FaceSwap, and NeuralTextures) and 1,000 original video sequences.

[0089] The DFD dataset records hundreds of real videos and uses publicly available deepfake generation methods to create thousands of deepfake videos from these videos. The DFD dataset contains over 3,000 manipulated videos from 28 actors in different scenes.

[0090] The Celeb-DF dataset is a widely used dataset containing 590 original videos and 5,639 corresponding deepfake videos from various websites. Due to the use of an improved deepfake algorithm, the Celeb-DF dataset does not exhibit particularly obvious forgery clues compared to FF++.

[0091] The WDF dataset contains 3,805 real videos and 3,509 fake videos, all collected from the internet. Therefore, we cannot determine the forgery methods used in the WDF dataset. Additionally, all videos in the WDF dataset have been compressed.

[0092] Example 1:

[0093] like Figure 2 As shown, we conducted in-dataset experiments on the FF++ dataset, specifically using face images from the FF++ dataset as the original images for the MFR Training Stage. Figure 2 As shown, in the MFR Training Stage, for each original face image x i First, for x i Perform an FFT transformation to obtain z i ( Figure 2 (not shown in the image), then z i With frequency mask The masked frequency spectrum is obtained by performing a multiplication operation. Then, an IFFT transform is performed on the masked frequency spectrum to obtain the high-frequency components. (High Frequency Component), then, x i and Perform the subtraction operation to obtain the low-frequency component. (Low Frequency Component), then, the random matrix (Random Masks) M i With high frequency components Perform a multiplication operation to obtain the processed high-frequency component (Masked High Frequency Component). Then, add the low-frequency component to the processed high-frequency component to obtain x. i Image processing After that, As training sample input, Frequency Forgery Representationg θ and according to Based on the labels and prediction results, calculate L. cls Loss, and adjust g through backpropagation θ This process of iterative training continues until a well-trained Frequency Forgery Representation is obtained. θ After obtaining a well-trained Frequency ForgeryRepresentation tool... θ Subsequently, using existing technology and the trained Frequency ForgeryRepresentation obtained in this embodiment, θ Performance comparisons were performed on the FF++ dataset, and the AUC is reported as shown in Table 1. We present comparison results between image-based methods (RFM, Add-Net, F3-Net, MultiAtt, etc.) and video-based methods (STIL, HCIL, etc.). ACMF (this invention) performs competitively on the dataset, and our method's AUC exceeds 99%.

[0094] Our method is still superior to STIL (97.12%), HCIL (98.32%) and GFF (98.36%).

[0095] Xception 96.30 RFM 98.79 Add-Net 87.74 F3-Net 98.10 MultiAtt 99.29 RECCE 99.32 LTW 99.17 GFF 98.36 HCIL 98.32 STIL 97.12 ACMF 99.30

[0096] Table 1

[0097] Example 2:

[0098] As shown in Table 2, we conducted cross-dataset experiments on the Celeb-DF and WDF datasets. Specifically, during the FACRefinement Stage, we used the FrequencyForgery Representation g trained in the MFR Training Stage. θ As Frequency Forgery Representation and Frequency Forgery Representation The set of training samples from the MFR Training Stage is used as the reference image set (Training Video Frames). (y i Representing the image x i (The tags), and select one video (Individual Video Frame) from Celeb-DF and WDF respectively. Will Enter Frequency Forgery Representation Will EnterFrequency ForgeryRepresentation And according to Frequency Forgery Representation for Each image x in i Generated gradient-weighted class activation map calculate corresponding And according to the Frequency Forgery Representation for Each image x in i Generated gradient-weighted class activation map calculate corresponding according to and Calculate L ac Loss, according to for The prediction results generated for each image in the dataset, and For each image in the dataset, calculate L. cls Loss, and L ac Loss and L clsThe sum of losses is taken as the total loss, to be adjusted through backpropagation. This training process is repeated iteratively until the dataset Celeb-DF is obtained. The corresponding trained Frequency Forgery Representation and in the WDF dataset The corresponding trained Frequency Forgery Representation Then, these two trained Frequency Forgery Representations were used. Our method was compared with other methods on the Celeb-DF and WDF datasets, and the AUC was reported. As shown in Table 2, many methods performed well on their own datasets, but their performance dropped significantly when faced with data of unknown forgery types. F3-Net and RFM had an AUC of only 57% on the WDF dataset. MultiAtt had an AUC of 67.02% on the Celeb-DF dataset, but only 59.74% on the WDF dataset, while our method (ACMF) was able to adaptively adjust the attention region according to different forgery datasets and performed better on each dataset.

[0099] Method AUC(%) AUC(%) Xception 62.72 64.8 RFM 57.75 65.63 Add-Net 62.35 65.29 F3-Net 57.1 61.51 MultiAtt 59.74 67.02 RECCE 64.31 68.71 LTW 67.12 77.14 GFF 66.51 65.31 HCIL - 79.00 STIL - 75.58 ACMF 74.22 84.02

[0100] Table 2

[0101] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A face forgery detection method based on frequency masking and attention consistency, characterized in that, include: Obtain the video to be tested; The video to be detected is input into a pre-trained forgery detection model to obtain the detection result of the video to be detected; The detection result indicates whether the video to be detected is a fake video; The pre-trained forgery detection model is obtained by iteratively training an initial pre-trained forgery detection model using multiple labeled reference images, sample videos, a cross-forgery attention consistency loss function, and a cross-entropy loss function; the initial pre-trained forgery detection model is obtained by iteratively training the forgery detection model to be trained using the multiple reference images and the cross-entropy loss function. Each reference image is obtained by performing a discrete Fourier transform on the corresponding original image and randomly discarding high-frequency information; the cross-forgery attention consistency loss function is used to measure the difference between the gradient-weighted class activation map of the reference image and the gradient-weighted class activation map of the video frames in the sample video; The expression for the cross-forgery attention consistency loss function is as follows: ; in, To cross-fake attention consistency loss values, The total number of video frames in the sample video. This represents the total number of images selected from the plurality of reference images during each training iteration. The first in the sample video The gradient-weighted class activation map of each video frame The image selected from the plurality of reference images during each training session is the first one. The gradient-weighted class activation map of each image, This represents the L2 norm.

2. The face forgery detection method based on frequency masking and attention consistency according to claim 1, characterized in that, The sample video is the video to be detected or a video other than the video to be detected; the multiple reference images include real images and fake images.

3. The face forgery detection method based on frequency masking and attention consistency according to claim 1, characterized in that, Before inputting the video to be detected into the pre-trained forgery detection model to obtain the detection result of the video to be detected, the following steps are included: Acquire multiple original images, including real and fake images; each original image is labeled; the label indicates whether the original image is a fake image; Each of the multiple original images is subjected to discrete Fourier transform and high-frequency information is randomly discarded to obtain multiple processed images; The forgery detection model to be trained is iteratively trained using the multiple processed images and the cross-entropy loss function to obtain a basic training model; The basic training model is used as the initial pre-trained forgery detection model and the auxiliary model, respectively. The multiple processed images are used as the multiple reference images, and the sample video is obtained. The sample video is used as the input to the initial pre-trained forgery detection model, and the multiple reference images are used as the input to the auxiliary model. The initial pre-trained forgery detection model is iteratively trained based on the gradient-weighted class activation map generated by the initial pre-trained forgery detection model for each video frame and the gradient-weighted class activation map generated by the auxiliary model for each reference image, to obtain the pre-trained forgery detection model.

4. The face forgery detection method based on frequency masking and attention consistency according to claim 3, characterized in that, The process involves performing a Discrete Fourier Transform on each of the multiple original images and randomly discarding high-frequency information to obtain multiple processed images, including: Perform a discrete Fourier transform on each of the multiple original images to obtain the frequency domain information of the original images; The frequency domain information is processed to obtain processed frequency domain information; the low-frequency information component in the processed frequency domain information is located in the central region, and the high-frequency information is located in the edge region. Based on the processed frequency domain information and the size of the original image, determine the distance corresponding to the position of each spectrum. The rectangular mask is determined based on the relationship between the distance and the preset distance threshold; After multiplying the frequency domain information with the rectangular mask, an inverse discrete Fourier transform is performed to obtain the high-frequency component, and the low-frequency component is obtained based on the original image and the high-frequency component. The high-frequency component is multiplied by a preset random matrix to obtain the processed high-frequency component; the preset random matrix consists of 0 and 1. Based on the processed high-frequency components and the low-frequency components, the processed image corresponding to the original image is obtained.

5. The face forgery detection method based on frequency masking and attention consistency according to claim 4, characterized in that, The expression for the processed image corresponding to the original image is: ; in, Indicates the first The original image described in Zhang, express The corresponding processed image, This represents the discrete Fourier transform. This represents the inverse discrete Fourier transform. express The frequency domain information, This refers to the rectangular mask. This refers to the high-frequency components. This refers to the low-frequency component. Represents the preset random matrix; wherein, and Size and They are the same size.

6. The face forgery detection method based on frequency masking and attention consistency according to claim 3, characterized in that, The process involves using the sample video as input to the initial pre-trained forgery detection model, and the multiple reference images as input to the auxiliary model. The initial pre-trained forgery detection model is iteratively trained based on the gradient-weighted class activation map generated for each video frame by the initial pre-trained forgery detection model, and the gradient-weighted class activation map generated for each reference image by the auxiliary model, to obtain the pre-trained forgery detection model. This includes: During the t-th training iteration, the sample video is input into the t-th training model to generate the t-th gradient-weighted class activation map for each video frame of the sample video. Simultaneously, the t-th training samples are obtained from the multiple reference images and input into the auxiliary model to generate the gradient-weighted class activation map for each training sample in the t-th training sample, as well as the prediction result for each training sample in the t-th training sample; t is an integer greater than 0; when t is 1, the t-th training model is the initial pre-trained forgery detection model; The total loss value for the tth time is calculated based on the gradient-weighted class activation map, the cross-spoofing attention consistency loss function, the cross-entropy loss function, the label of each training sample in the tth training sample, the gradient-weighted class activation map, and the prediction result. Adjust the network parameters of the model to be trained in the t-th iteration based on the total loss value in the t-th iteration to obtain the model to be trained in the (t+1)-th iteration. The training sample (t+1)th iteration is obtained from the multiple reference images. Based on the training sample (t+1th iteration), the sample video, the auxiliary model, the cross-spoofing attention consistency loss function, and the cross-entropy loss function, the training model (t+1th iteration) is trained until the pre-trained spoofing detection model is obtained.

7. The face forgery detection method based on frequency masking and attention consistency according to claim 6, characterized in that, The expression for the total loss value in the t-th iteration is as follows: ; in, Let be the total loss value for the t-th time. The total number of video frames in the sample video. Let be the total number of training samples in the t-th training iteration. The first in the sample video The gradient-weighted class activation map of the t-th video frame. for The Middle One image, for The gradient-weighted class activation map, Describing the L2 norm, for The tag, for The prediction results, This refers to the classifier head in the t-th training model. Let be the network parameters of the model to be trained for the tth time.

8. The face forgery detection method based on frequency masking and attention consistency according to claim 3, characterized in that, The step of iteratively training the forgery detection model to be trained using the multiple processed images and the cross-entropy loss function to obtain a basic training model includes: During the c-th training iteration, the c-th training sample is obtained from the multiple processed images, and the c-th training sample is input into the c-th forgery detection model to obtain the prediction result of the c-th training sample; c is an integer greater than 0; when c is 1, the c-th forgery detection model is the forgery detection model to be trained. Calculate the cross-entropy loss value for the c-th training sample based on the prediction result of the c-th training sample, the label of the c-th training sample, and the cross-entropy loss function. Adjust the network parameters of the c-th forgery detection model based on the c-th cross-entropy loss value to obtain the (c+1)-th forgery detection model. The c+1th training sample is obtained from the multiple processed images, and the c+1th forgery detection model is trained based on the c+1th training sample and the cross-entropy loss function until the basic training model is obtained.

9. The face forgery detection method based on frequency masking and attention consistency according to claim 1, characterized in that, The multiple reference images are multiple RGB face images; the sample video is a video containing face regions.

Citation Information

Patent Citations

  • Anti-JPEG compression forged image detection method

    CN113255571A

  • Face image forgery detection method and related equipment

    CN115909445A