Deep Face Forgery Video Detection Method Based on rPPG Signals

Through the deep face forged video detection method based on rPPG signals, the face sequence and rPPG signals are extracted and combined with the CNN classifier for identification, the problem of difficulty in detecting multiple deep face forged videos in the existing technology is solved, and effective detection of forged videos and robustness in high-compression scenarios is achieved.

CN114882419BActive Publication Date: 2025-05-27NANJING UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202210572034.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-24
Publication Date
2025-05-27
Estimated Expiration
2042-05-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect fake face videos generated by various deep face forgery methods, and the detection performance is degraded when facing high-compression videos.

Method used

A deep face forged video detection method based on rPPG signals is adopted. By extracting face sequences from the video, selecting specific faces of interest, using a green single-channel method to extract rPPG signals, and using a CNN-based classifier for forgery identification.

Benefits of technology

It realizes effective detection of fake face videos generated by various deep face forgery methods, has certain anti-compression capabilities, and can maintain good detection performance in high-compression scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882419B_ABST
    Figure CN114882419B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting deep face forged videos based on rPPG signals, belonging to the field of artificial intelligence security. It includes obtaining a face sequence from a video; selecting a specific region of interest of the face and adopting a method based on a single green channel to extract rPPG signals; and using a CNN-based classifier to perform forgery discrimination based on the obtained rPPG signals. The present invention utilizes the characteristic that most current forgery methods are difficult to simulate the human heart rate signal, and detects forged faces by extracting rPPG signals from specific parts of the video face, and can effectively detect face forged videos generated by various forgery methods. At the same time, by adopting an rPPG signal extraction method with a certain ability to resist video compression, the detection performance of the present invention for highly compressed forged videos is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence security, and particularly relates to a method for detecting deep face forgery videos based on rPPG signals. Background Art

[0002] In recent years, with the continuous development of artificial intelligence technology, face forgery technology has also made great progress. At present, this technology has been widely used in various cultural and entertainment industries, greatly enriching daily life. However, it is worrying that while this technology brings some benefits to society, it also brings great risks to the whole society. In recent years, criminal activities such as using face forgery technology to forge pornographic videos of public figures for slander, defamation and blackmail have shown an increasing trend. These activities not only violate the portrait rights of the victims but also bring harm to the victims' lives and psychology. Even more serious is that terrorist forces can use this technology to forge videos to guide public opinion, incite violence or even provoke war. Therefore, how to effectively detect face forgery videos generated by this technology has become an urgent problem to be solved.

[0003] Currently, the existing deep face forgery detection technologies are mainly divided into three categories: artifact-based detection methods, data-driven detection methods, and information-inconsistency-based detection methods. Among them, artifact-based detection methods will generate some artifacts that are invisible to the naked eye when generating face forgery videos. Deep learning methods can be used to extract and identify these artifacts. However, such methods often have poor generalization ability and can only detect face forgery videos generated by specific face forgery methods. At the same time, with the iterative upgrade of face forgery technology, the artifacts in forged videos are becoming increasingly difficult to detect. Therefore, this type of method is likely to become ineffective in the future. Data-driven methods detect forgery by using existing image classification networks. This type of method automatically searches for the distinguishing features between real and forged through neural networks and can achieve good detection performance. However, the network structure used in this type of method is relatively complex, so it requires a large amount of time and space resources. Information-inconsistency-based detection methods are mainly divided into those based on biometric signal inconsistency, time-series inconsistency, and inconsistency with real human behavior. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for detecting forged face videos generated by multiple deep face forgery methods, and at the same time, the method should have a certain anti-compression ability, namely, a method for detecting deep face forgery videos based on rPPG signals.

[0005] Technical Solution: To solve the above technical problem, the technical solution adopted by the present invention is as follows:

[0006] A method for detecting deep face forgery videos based on rPPG signals, comprising the following steps:

[0007] Step 1: Obtain a face sequence from the video;

[0008] Step 2: Select a specific region of interest (ROI) of the face and use a green single-channel based method to extract the rPPG signal;

[0009] Step 3: Use a CNN-based classifier to perform forgery discrimination based on the obtained rPPG signal.

[0010] Further, in Step 1, the method for obtaining a face sequence from the video is as follows:

[0011] Input the video to be detected into the face extraction module. The face extraction module uses the FaceMesh method to extract faces from each frame of the video and generates a face sequence of 128 frames.

[0012] Further, in Step 2, the method for extracting the rPPG signal from a specific ROI of the face is as follows:

[0013] Step 2.1: Utilize the faces and face landmark points extracted in Step 1 to select n face landmark points in the region where the rPPG signal is relatively rich, and then construct square grids of equal size centered on the selected n landmark points respectively. The regions of the n square grids are used as the face ROI;

[0014] Step 2.2: Based on the prior knowledge that the rPPG signal is the strongest in the green channel of the three-channel pixel values, calculate the pixel mean value of the green channel within each selected square region;

[0015] Step 2.3: Process each of the 128 consecutive frames of face images using Steps 2.1 and 2.2. Finally, an n*128 signal matrix can be obtained;

[0016] Step 2.4: For the extracted pixel signal matrix, perform filtering processing on it to reduce noise signals and obtain a cleaner rPPG signal;

[0017] Step 2.5: For the rPPG signal matrix after filtering processing, save the values in the signal matrix in the form of an image as the rPPG signal map.

[0018] Further, in Step 2.3, the value of each column of the rPPG signal matrix is the pixel mean value of the n square regions selected on a single-frame face image, and the value of each row is the pixel mean value of a square region on the face over 128 consecutive frames.

[0019] Further, in step 2.4, a band-pass filter is used to further filter the rPPG signal, retaining signals with frequencies ranging from 0.6 HZ to 10 HZ.

[0020] Further, in step 3, a CNN-based classifier is used to perform forgery discrimination based on the obtained rPPG signal, and the specific method is as follows:

[0021] By using a CNN-based classifier to classify the n*128 rPPG signal map generated in step 2, scores for the signal map belonging to the real category and the forged category are obtained. If the forged score output by the classifier is higher than the real score, the face in the video is considered a forged face.

[0022] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0023] (1) The present invention proposes a deep face forgery video detection method based on rPPG signals. Compared with other detection methods, it can effectively detect face forgery videos generated by various face forgery methods and has a certain generalization ability.

[0024] (2) By using landmark points located on the forehead, nose wings, and cheeks of the face to construct the region of interest of the face, richer rPPG signals can be obtained. At the same time, by adopting the method centered on the face landmark points, it can be ensured to the greatest extent that the face parts for extracting signals between different frames are fixed, which can improve the robustness to face movement. At the same time, the position distribution of the selected face landmark points is relatively wide, which also helps to make full use of the consistency of rPPG signals at different positions on the face to identify forged videos.

[0025] (3) By using the pixel values of the green single channel to obtain the rPPG signal, a clearer rPPG signal can be obtained. At the same time, this extraction method has robustness to highly compressed videos, which enables the method to still maintain good detection performance when facing highly compressed videos. Description of the Drawings

[0026] Figure 1 is a flowchart of the deep face forgery video detection method based on rPPG signals;

[0027] Figure 2 is to obtain the face sequence from the video frames;

[0028] Figure 3 is the selection of the region of interest of the face;

[0029] Figure 4 is the generated rPPG signal map;

[0030] Figure 5It is the network structure of a CNN-based classifier. Specific implementation manner

[0031] The following further clarifies the present invention in conjunction with specific embodiments. The embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0032] As Figure 1 shown, the deep face forgery video detection method based on rPPG signals of the present application includes the following steps:

[0033] Step 1: Obtain a face sequence from the video;

[0034] Input the video to be detected into the face extraction module. The face extraction module uses the method of FaceMesh to extract faces and corresponding face landmark points from each frame of the video, generating a face sequence of 128 frames. The obtained face images are as shown in the appendix Figure 2 shown.

[0035] Step 2: Select a specific region of interest in the face and adopt a method based on the green single channel to extract rPPG signals. The method is as follows:

[0036] Step 2.1: Utilize the faces and face landmark points extracted in the first step. According to the characteristic that the rPPG signals in the forehead, nose wings, and both cheeks of the face are relatively rich, select 32 face landmark points located in the above regions, and then construct equal-sized square grids (the side length of the square grid is 6 pixels) centered on these 32 landmark points respectively. Then use the regions of these 32 square grids as the regions of interest in the face. The selection effect is as shown in the appendix Figure 3 shown. The points in the left face are the selected face landmark points, and the small square regions in the right face are the selected regions of interest in the face.

[0037] Step 2.2: Based on the prior knowledge that the rPPG signal in the green channel of the three-channel pixel values is the strongest, calculate the pixel mean value of the green channel within each selected square region. By using the method of the pixel mean value of the green channel within the region, while reducing the interference of camera quantization noise, the effective rPPG signal is maximally obtained. According to Step 2.1, for a single-frame face image, the pixel mean values of 32 regions can be obtained.

[0038] Step 2.3: Process 128 consecutive face images using the above Steps 2.1 and 2.2. Finally, a 32*128 rPPG signal matrix can be obtained, where the value of each column in the matrix is the pixel mean value of the 32 selected square regions on a single-frame face image, and the value of each row is the pixel mean value of a square region on the face in 128 consecutive frames.

[0039] Step 2.4: For the extracted pixel signal matrix, it is also necessary to perform filtering on it to reduce noise signals and obtain a cleaner rPPG signal. Here, a band-pass filter is used to further filter the rPPG signal, retaining signals with frequencies ranging from 0.6HZ to 10HZ. This is different from the interval of 0.67HZ - 4HZ selected when generally measuring heart rate problems. The interval of 0.67 - 4HZ is selected for heart rate problems because the frequency range of normal human heart rate signals is within this interval. Therefore, in this way, as much noise signal as possible can be removed while retaining the rPPG signal. However, this type of noise signal is helpful for forgery detection to a certain extent. Therefore, a larger frequency retention interval is selected for filtering.

[0040] Step 2.5: For the filtered rPPG signal matrix, the rPPG signal graph is obtained by saving the values in the signal matrix in the form of an image, as shown in the appendix Figure 4 as follows:

[0041] Step 3: Use a CNN-based classifier to perform forgery detection based on the obtained rPPG signal.

[0042] Use a CNN-based classifier to classify the 32*128 rPPG signal graph generated in Step 2, and obtain the scores of the signal graph belonging to the real category and the forged category. If the forged score output by the classifier is higher than the real score, it is considered that the face in the video is a forged face.

[0043] The structure and some training details of this CNN-based classifier are as follows: A simple four-layer convolutional neural network is adopted. Each convolutional layer is followed by a max-pooling layer, and finally 2 fully connected layers are used to output the video authenticity score. The ReLU function is used as the activation function between the first 4 convolutional layers and the last 2 fully connected layers. The last fully connected layer directly outputs the scores of the video belonging to the real video and the forged video, and no activation function is connected behind. If the output real video score is higher than the forged video score, it is considered that the video is a real video, and vice versa. At the same time, to prevent the model from overfitting, a dropout layer with a dropout probability of 0.8 is added between the last 2 fully connected layers. At the same time, no data augmentation measures are taken during the training process, and the original rPPG signal graph is directly used as the input.

[0044] The present invention belongs to the method based on the inconsistency of biological signals. By taking advantage of the characteristic that it is difficult for current deep face forgery technologies to forge the inherent heart rate signal of the human body, the rPPG signal of the face area in the video is extracted to generate an rPPG signal graph, and the generated rPPG signal graph is input into a CNN-based classifier for forgery detection.

[0045] The effectiveness and efficiency of the method of the present invention are verified through the following experiments:

[0046] Four indicators, namely acc (accuracy rate), fpr (false positive rate), fnr (false negative rate), and auc, are used as the evaluation criteria:

[0047] acc (accuracy rate): That is, the ratio of the number of samples correctly predicted by the model to the total number of predicted samples, and its calculation method is:

[0048]

[0049] fpr (false positive rate): That is, the probability that a negative sample is predicted as a positive sample by the model. Here, the positive samples in this article refer to forged videos, and the negative samples refer to real videos. Its calculation method is:

[0050]

[0051] fnr (false negative rate): That is, the probability that a positive sample is predicted as a negative sample by the model. Its calculation method is:

[0052]

[0053] auc: It is defined as the area under the ROC curve. The ROC curve is the Receiver Operating Characteristic Curve in full. Its abscissa is fpr (false positive rate), and its ordinate is tpr (true positive rate, that is, the probability that a real positive sample is predicted as a positive sample). Since the ROC curve is generally above the line y = x, the value range of auc is usually between 0.5 and 1. For a detection method, the closer its auc value is to 1, the better the performance of the method, and the closer it is to 0.5, the less practical application value the method has.

[0054] First, select the dataset. The present invention selects the FaceForensics++ dataset and the Celeb-DFV2 dataset. The FF++ dataset includes a total of 6000 videos, and the video frame rate, length, and resolution are not fixed. Its dataset consists of 1000 real videos and 5000 forged videos. Among them, the forged videos are generated based on the 1000 collected real videos using 5 forgery methods (Deepfake, Face2face, Faceshifter, Faceswap, NeuralTextures), with 1000 videos for each forgery method. Due to the large volume of the videos in the original dataset, the datasets compressed based on the c23 and c40 standards are selected respectively. Among them, the dataset compressed based on c40 is only used in subsequent experiments related to video compression. Therefore, unless otherwise specified, the FF++ dataset used in this article is default to the c23 compression standard, that is, FF++c23.

[0055] The Celeb-DFV2 dataset is currently the most widely used dataset, which includes 590 real videos collected from YouTube and 5,639 forged videos generated based on the Deepfake method. In addition, it also provides an additional 299 real videos as a supplement. The average length of the videos in this dataset is about 13 seconds, and the standard frame rate is 30 frames per second. Although the forgery method used in the Celeb-DFV2 dataset is relatively single compared to the FF++ dataset, the quality of the forged videos generated is significantly higher than that of the FF++ dataset.

[0056] 1. Overall test results on the dataset

[0057] The experimental results of the proposed method for detecting deepfake videos based on rPPG signals on the FF++ dataset and the Celeb-DFV2 dataset are as follows:

[0058] Table 1 Overall performance of the forgery detection method based on rPPG signals on the dataset

[0059]

[0060] As can be seen from Table 1, the method achieves good accuracy on both the FF++ dataset and the Celeb-DFV2 dataset, indicating that the method for detecting deepfake videos based on rPPG signals can effectively detect deep face forgery videos.

[0061] 2. Test results on different forgery methods

[0062] In addition to testing the overall performance of the method on the two datasets, we also utilized the fact that FF++ provides multiple forgery methods to separately test the detection performance of the method on different forgery methods. The detection method was trained and tested on each forgery method, and the results are shown in Table 2. As can be seen from the table, the method has good detection effects on the 5 forgery methods used in FF++. The method has extremely excellent performance for the first 4 methods, but its performance drops when faced with the NeuralTextures method. However, overall, the method still performs very well.

[0063] Table 2 Detection performance of the forgery detection method based on rPPG signals in the face of different forgery methods

[0064]

[0065] 3. Forgery detection results in high-compression scenarios

[0066] Considering that most forged videos in real-world network scenarios are compressed, the data at c23 and c40 compression levels provided by the FF++ dataset was used to conduct forgery detection tests in high-compression scenarios. The test results are shown in the following table. The method's performance decreased to some extent at the c40 compression level, but the decline was not significant, and it still maintained good detection performance.

[0067] Table 3 Detection Performance in High-Compression Scenarios

[0068]

[0069] 4.3.4 Cross-Dataset Performance Test

[0070] In addition to the above experiments, cross-dataset tests were also conducted on the detection method to detect its generalization performance. As can be seen from Table 4.4, the performance of the method in cross-dataset tests was relatively stable. Although its indicators decreased compared with the results of training and testing on a single dataset, the decline was not significant, indicating that the method has a certain generalization ability.

[0071] Table 4 Cross-Dataset Performance Test Results

[0072]

[0073] Based on the above experiments, it can be found that the deepfake video authentication method based on rPPG signals can detect face forged videos generated by various forgery methods. Experiments in high-compression scenarios show that the authentication method based on rPPG signals exhibits good robustness when facing high-compression scenarios. At the same time, the performance test on cross-datasets also reflects that the method has a certain generalization ability.

[0074] The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for detecting deep face forged videos based on rPPG signals, characterized in that, it includes the following steps: Step 1: Obtain a face sequence from the video; Input the video to be detected into the face extraction module. The face extraction module uses the FaceMesh method to extract faces from each frame of the video and generates a face sequence of 128 frames; Step 2: Select a specific region of interest in the face and use a method based on the green single channel to extract the rPPG signal; The method is as follows: Step 2.1: Utilize the faces and face landmark points extracted in Step 1 to select n face landmark points in the region where the rPPG signal is relatively rich, and then construct square grids of equal size centered on the selected n face landmark points respectively. The regions of the n square grids are used as the regions of interest in the face; Step 2.2: Based on the prior knowledge that the rPPG signal is the strongest in the green channel of the three-channel pixel values, calculate the pixel mean value of the green channel within the region of each selected square grid; Step 2.3: Process each of the 128 consecutive face images using Steps 2.1 and 2.2, and finally obtain an n*128 rPPG signal matrix; The value of each column of the rPPG signal matrix is the pixel mean value of the regions of the n square grids selected on a single-frame face image, and the value of each row is the pixel mean value of the region of a square grid on the face over 128 consecutive frames; Step 2.4: For the extracted rPPG signal matrix, perform filtering on it to reduce noise signals and obtain a cleaner rPPG signal matrix; Among them, a band-pass filter is used to further filter the rPPG signal matrix, and the signals with frequencies between 0.6HZ and 10HZ are retained; Step 2.5: For the rPPG signal matrix after filtering, save the values in the rPPG signal matrix in the form of an image as an rPPG signal map; Step 3: Use a CNN-based classifier to perform forgery identification based on the obtained rPPG signal map; Use a CNN-based classifier to classify the n*128 rPPG signal map generated in Step 2, and obtain the scores of the rPPG signal map belonging to the real category and the forged category. If the forged score output by the classifier is higher than the real score, it is considered that the face in the video is a forged face.

Citation Information

Cited By

  • Method, system and device for realizing deep counterfeit face identification, processor and computer readable storage medium thereof

    CN116012958A

  • Method, system, device, processor and computer readable storage medium thereof for realizing deep fake face identification

    CN116012958B