Deepfake video detection method and system based on double-flow graph network
By adopting a method of combining spatial and temporal features in Deepfake video detection, and introducing information entropy regularization, the problems of low identification accuracy and insufficient generalization ability in the prior art are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510103346.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-06
AI Technical Summary
The existing Deepfake video detection method is insufficient in utilizing the time and spatial characteristics of the video sequence, resulting in low recognition accuracy, insufficient generalization ability and poor robustness.
Using a detection method based on a dual-stream graph network, the feature extraction combined with space and time is used to aggregate information from different dimensions using auxiliary learning tasks, and the information entropy regularization process is used to reduce model overfitting.
It significantly improves the accuracy and generalization ability of Deepfake video detection, enhances the model's sensitivity to subtle changes in the video, and reduces the risk of overfitting.
Smart Images

Figure CN120107844A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Deepfake video tampering detection, and in particular to a Deepfake video detection method and system based on a dual-stream graph network. Background Art
[0002] With the rapid development of Deepfake technology, it is becoming easier and easier to generate realistic Deepfake videos. Therefore, our research on the identification and classification of Deepfake video tampering is particularly important. At present, most detection algorithms are based on single-frame images and ignore the information between video frames. In addition, Deepfake generation methods are diverse and widely used in a wide range of scenarios, resulting in low recognition accuracy, insufficient generalization ability, and poor robustness.
[0003] In 2020, Zhang Yixuan et al. published a face-tampering video detection method based on inter-frame differences in the Journal of Information Security. The features between video frames are also worthy of study as a basis for classification. A detection method based on local binary pattern and directional gradient histogram features was proposed, and combined with the deep learning technology of the twin network, the detection accuracy of face-tampering videos was effectively improved (Face-tampering video detection method based on inter-frame differences Zhang Yixuan, Li Gen, Cao Yun, Zhao Xianfeng, National Key Laboratory of Information Security, Institute of Information Engineering, Chinese Academy of Sciences, School of Cyberspace Security, University of Chinese Academy of Sciences).
[0004] In general, existing methods are still insufficient in utilizing the temporal and spatial features of video sequences. In cross-database tests on different databases, the detection capabilities of these methods have been significantly reduced, indicating their weak generalization capabilities. Summary of the invention
[0005] In order to overcome the defects and shortcomings of the prior art, the present invention provides a Deepfake video detection method based on a dual-stream graph network. The dual-stream graph network model of the present invention can fully extract and utilize the spatial and temporal feature information of video sequences, and utilize auxiliary learning tasks to aggregate important information of different dimensions in the input sequence, effectively capture the spatiotemporal anomalies in Deepfake videos, and through information entropy regularization processing, effectively reduce the overfitting of the model, and improve the accuracy and generalization ability of the model.
[0006] The present invention is achieved by at least one of the following technical solutions.
[0007] The Deepfake video detection method based on the dual-flow graph network includes the following steps:
[0008] Obtain the video data to be detected;
[0009] The video data to be detected is input into the trained dual-stream graph network model to detect deep forgery of the video; the training of the dual-stream graph network model includes the following steps:
[0010] S1. Preprocess the videos with faces in the dataset to obtain feature sequences;
[0011] S2. Input the feature sequence into the dual-stream graph network model to obtain the corresponding spatial attention parameters and temporal attention parameters; perform vector concatenation on the temporal attention parameters and the feature sequence to obtain the temporal feature sequence;
[0012] S3, taking the temporal concatenated feature sequence as the input to the 3-layer GRU encoder to obtain the temporal feature representation; taking the attention parameters in the spatial feature space as the input to the Dense fully connected encoder to obtain the spatial feature representation;
[0013] S4, vectorize the temporal feature representation and the spatial feature representation to obtain a classification feature sequence, input the classification feature sequence into a Dense fully connected classifier, and use the binary cross entropy as the classification loss function to calculate the loss of the output result of the Dense fully connected classifier;
[0014] S5, input the temporal feature representation into the GRU decoder, and the output result uses the mean square error to calculate the loss;
[0015] S6. Add the above two losses and add information entropy regularization as the total loss function, calculate the total loss function, perform backpropagation and update the weights in the dual-flow graph network model, and save the weight information of the best verification accuracy.
[0016] Furthermore, the preprocessing of the videos with human faces in the data set to obtain feature sequences includes the following steps:
[0017] Read the video with human face in the data set, extract the face area of the video frame by frame, locate the facial feature points in the face area, and divide the face area using the facial feature points and the Deroni triangulation method;
[0018] Calculate and fit the affine transformation matrix between the same areas in two consecutive frames in the video, calculate the single-segment feature quantity between the matrices, repeat this process for the areas obtained by face segmentation and all frames, and finally concatenate to obtain the original multi-segment feature sequence, and at the same time divide the original multi-segment feature sequence into multi-segment feature sequences.
[0019] Furthermore, the facial feature points and the Delonian triangulation method are used to divide the face into regions. The facial feature point positioning function in the insightFace model is used to identify the face, select the facial feature points, and use the Delonian triangulation method to divide the facial feature points into 7 regions, and each of the obtained regions is a triangle.
[0020] Furthermore, the affine transformation matrix between the same regions in two consecutive frames in the video is calculated and fitted, and the single-segment feature quantity between the matrices is calculated. This process is repeated for the regions obtained by face segmentation and all frames, and finally the original multi-segment feature sequence is obtained in series. At the same time, the original multi-segment feature sequence is divided into multiple segment feature sequences. The specific steps include:
[0021] The area obtained by dividing the face is used to calculate the affine transformation matrix M using the function of calculating affine transformation parameters in the OpenCV library. The affine transformation matrix M is used to calculate the average value, median value, average value, and median value of the lateral displacement, and the corresponding average value and median value are used as single-segment feature quantities. The single-segment feature quantities of all areas are merged to obtain a single-segment feature sequence, and then the remaining frames of the video are processed in the same way. The results of all frame processing are connected in series to obtain the original multi-segment feature sequence.
[0022] Furthermore, the original multi-segment feature sequences are divided into multiple-segment feature sequences, and the specific steps include: dividing the original multi-segment feature sequences into multiple fixed-size feature sequences as the multiple-segment feature sequences according to the sliding window and the sliding length.
[0023] Furthermore, the dual-stream graph network model includes a spatial feature map attention module and a temporal feature map attention module, and the feature sequence is input into the spatial feature map attention module and the temporal feature map attention module to obtain corresponding spatial attention parameters and temporal attention parameters. The specific steps include:
[0024] The feature sequence is regarded as a set of spatial feature sequences, and each element of the spatial feature sequence is modeled as a fully connected graph. The spatial feature graph attention module is used to output the spatial attention parameters.
[0025]
[0026] Among them, σ(.) represents the sigmoid function, represents the concatenation of vectors, W s ∈R 64×64 is a learnable linear transformation matrix, is the attention vector, is the graph attention coefficient, Leaky ReLU(.) is the rectified linear unit function, is the normalized result of the graph attention coefficient, Soft max j (.) is the normalized exponential function, S i , S j are two elements in the feature sequence;
[0027] The feature sequence is regarded as a set of time feature sequences, each element of the time feature sequence is modeled as a fully connected graph, and the time feature graph attention module is used to output the time attention parameters
[0028]
[0029] Among them, σ(.) represents the sigmoid function, represents the concatenation of vectors, W t ∈R 64×64 is a learnable linear transformation matrix, is the attention vector, is the graph attention coefficient, Leaky ReLU(.) is the rectified linear unit function, is the normalized result of the graph attention coefficient, Soft max j (.) is the normalized exponential function, T i 、T j are two elements in the feature sequence.
[0030] Furthermore, the hidden layer dimension of each GRU encoder of the 3-layer GRU encoder is 256;
[0031] The Dense fully connected classifier includes three fully connected layers of Dense and one layer of Sigmoid. The input dimension of the first layer of Dense is 512 and the output dimension is 1024; the input dimension of the second layer of Dense is 1024 and the output dimension is 512; the input dimension of the third layer of Dense is 512 and the output dimension is 1;
[0032] The Dense fully connected encoder includes two layers of Dense, the input dimension of the first layer of Dense is 28x64, and the output dimension is 1024; the input dimension of the second layer of Dense is 1024, and the output dimension is 256;
[0033] The GRU decoder includes a layer of gated recurrent single GRU and a layer of Dense. The hidden layer dimension of GRU is 256, and the input dimension and output dimension of Dense are 256 and 64, respectively.
[0034] Furthermore, the classification loss function formula is:
[0035]
[0036] Where D c (.) is a Dense fully connected classifier, is the attribute label of the input sequence sample, is the vector concatenation of temporal feature representation and spatial feature representation, L BCEis the binary cross entropy loss, BCE(.) is the binary cross entropy loss function;
[0037] The loss formula using mean square error is:
[0038] L MSE =MSE(G d (R t ),x);
[0039] Among them G d (.) is the GRU decoder, R t is the time feature representation, x is the input feature sequence, L MSE is the mean square error, MSE(.) is the mean square error function;
[0040] Add the above two losses and add information entropy regularization as the total loss function to calculate the total loss function:
[0041] L=L BCE +L MSE ;
[0042] Final_L=L+α*H;
[0043] Where H is the output of information entropy, α is the hyperparameter, L is the intermediate loss, and Final_L is the information entropy regularization loss.
[0044] Furthermore, it includes: a video frame-by-frame processing module, a feature sequence division module, a graph attention module, a GRU encoding and decoding module, a Dense fully connected encoding and classifier module, and a total loss calculation module;
[0045] The video frame-by-frame processing module is used to locate facial feature points in the video data in the data set, select facial features for Deroni triangulation, calculate affine transformation parameters of the triangular feature regions between frames, and concatenate the affine transformation parameters between frames to obtain a total feature sequence;
[0046] The feature sequence division module is used to divide the total feature sequence according to the set sliding window time and length to obtain multiple feature sequences;
[0047] The graph attention module is used to process multiple feature sequences in space and time to obtain corresponding spatial attention parameters and temporal attention parameters;
[0048] The GRU encoding and decoding module is used as an auxiliary learning task to reconstruct multiple feature sequences and calculate the average error as the loss. The GRU encoding and decoding module includes a GRU decoder and a fully connected layer Dense; the Dense fully connected encoding and classifier module is used as a binary classification task, taking the temporal feature representation and spatial feature representation of the vector concatenation as input and outputting the sequence forgery situation. The Dense fully connected encoding and classifier module includes a fully connected layer Dense and a Sigmoid function;
[0049] The total loss calculation module is used to sum the losses of the two tasks, perform information entropy regularization processing to calculate the total loss function, and simultaneously perform backpropagation and update the weights in the dual-stream graph network model, saving the weight information of the best verification accuracy; based on the trained dual-stream graph network model, feature extraction and prediction classification are performed, and the deepfake forgery status of each video in the output data set is output.
[0050] A computer device of the present invention comprises: a memory and a processor and a computer program stored in the memory. When the computer program is executed on the processor, the Deepfake video detection method based on the dual-flow graph network is implemented.
[0051] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0052] (1) The present invention adopts a dual-stream graph network model that combines space and time, which can capture the potential features in the video more comprehensively. At the same time, the spatial feature map attention module and the temporal feature map attention module process spatial and temporal information respectively, enhancing the model's sensitivity to subtle changes in the video.
[0053] (2) The present invention effectively reduces the risk of model overfitting by introducing information entropy regularization. The information entropy regularization strategy enhances the diversity and generalization ability of model learning, making the model perform better when processing unseen data.
[0054] (3) The present invention adopts the technical solution of facial feature point positioning and Deroni triangulation method to divide the face into regions, and extracts distinguishing features by calculating the affine transformation matrix of the region between consecutive frames. This method can effectively capture subtle changes in Deepfake videos, thereby significantly improving the accuracy of the model.
[0055] (4) The present invention introduces auxiliary learning tasks, and the model can aggregate important information of different dimensions in the input sequence, thereby improving the learning ability of the graph network model, promoting the interaction between temporal and spatial features, and improving the overall detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A flow chart of a Deepfake video detection method based on a dual-stream graph network according to an embodiment;
[0057] Figure 2 A verification flow chart of a Deepfake video detection method based on a dual-flow graph network is provided in the embodiment;
[0058] Figure 3 A schematic diagram of a video feature extraction and processing framework of an embodiment;
[0059] Figure 4 Schematic diagram of the overall network framework of a Deepfake video detection method based on a dual-stream graph network according to an embodiment;
[0060] Figure 5 Schematic diagram of facial feature area image in the embodiment. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0062] like Figure 1 , Figure 2 , Figure 3 Combination Figure 4 As shown, this embodiment provides a Deepfake video detection method based on a dual-flow graph network, comprising the following steps:
[0063] Obtain the video data to be detected; input the video data to be detected into the trained dual-stream graph network model to detect deepfake forgery in the video; the training of the dual-stream graph network model includes the following steps:
[0064] S1. Preprocess the videos with faces in the dataset to obtain feature sequences;
[0065] This embodiment uses the Deepfake video database FaceForensics++ as a training and testing data set. The FaceForensics++ database has 1,000 real video samples and 1,000 corresponding face-changing video samples of four different face-changing algorithms, and is classified into three different compression rates, namely c40 compression rate, c23 compression rate, and c0 compression rate. The real video data comes from the foreign video website YouTube, and the face-changing videos are processed through DeepFake technologies such as DeepFakes, FaceSwap, Face2Face, and NeuralTextures.
[0066] This example creates a training set, a validation set, and a test set from the c0 set in the FaceForensics++ video database in a ratio of 8:1:1. This example is based on the deep learning framework torch (2.1.0) for experiments, the graphics card used is an RTX 3080 (10GB), the CUDA version is 12.1, and the operating system used is ubuntu22.04.
[0067] The preprocessing includes: reading videos with faces in the data set, processing the videos frame by frame, extracting face areas, and locating and dividing the face areas;
[0068] In this embodiment, the open source computer vision library (OpenCV) in Python is used to extract video frames from the video, and then the face detection model of the InsightFace model (https: / / github.com / deepinsight / insightface) is used to detect the face in the video frame (Deng J, Guo J, Xue N, et al. Arcface: Additiveangular margin loss for deep face recognition [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019: 4690-4699.), and then the rectangular frame of the face position is obtained, the face area is extracted, and its two-dimensional feature point information is obtained, a total of 106 feature points, specific feature points are selected, the face feature points of the face area are located, and the face area is divided into regions using the given feature points and the De Roni triangulation method. The sequence numbers of the feature points given in this embodiment are 4, 7, 9, 14, 20, 23, 25, 30, 50, 52, 53, 61, 75, 77, 81, 83, 86, and 102; the Deroni triangulation method is used to divide the 7 triangular feature areas, such as Figure 5 shown.
[0069] S2, calculating the affine transformation parameters of the triangular feature region between frames, and concatenating the affine transformation parameters between frames to obtain the total feature sequence;
[0070] In this embodiment, the affine transformation matrix of the triangle is calculated using the function of calculating affine transformation parameters in the OpenCV library for the same triangular feature area in two consecutive frames, and the lateral displacement average value x is obtained using the affine transformation matrix. mean , the median value of lateral displacement x Median , average longitudinal displacement y mean , the middle value of longitudinal displacement y Median, and use the corresponding average and intermediate values as single-segment feature quantities; similarly, perform the same process on the remaining 6 triangular feature regions to obtain the single-segment feature quantities of each region, and combine the single-segment feature quantities of all regions to obtain a single-segment feature sequence. Perform the same process on the remaining frames of the video, and concatenate the results of all frame processing to obtain the total feature sequence of the video, that is, the original multi-segment feature sequence.
[0071] S3, dividing the total feature sequence according to the set sliding window time and length to obtain multiple feature sequences;
[0072] As an embodiment, the total feature sequence is divided into 28×64 sample sequences with 64 frames as the sliding window length and 1 second as the sliding window time. The same process is performed on other videos in the data set to obtain multiple feature sequences.
[0073] S4. Input multiple feature sequences into the dual-stream graph network model, which includes a spatial feature map attention module and a temporal feature map attention module. Input the feature sequences into the spatial feature map attention module and the temporal feature map attention module to obtain corresponding spatial attention parameters and temporal attention parameters.
[0074] In this embodiment, a multi-segment feature sequence is regarded as S, all elements in the sequence are selected as node vectors of the graph network model, and the spatial attention parameter and the temporal attention parameter are calculated; the spatial attention parameter is input as The specific calculation formula is as follows:
[0075]
[0076] Among them, σ(.) represents the sigmoid function, represents the concatenation of vectors, W s ∈R 64×64 is a learnable linear transformation matrix, is the attention vector, is the graph attention coefficient, LeakyReLU(.) is the rectified linear unit function, is the normalized result of the graph attention coefficient, Soft max j (.) is the normalized exponential function, S i , S j are two elements in the feature sequence.
[0077] Temporal Attention Parameters The specific calculation formula is as follows:
[0078]
[0079] Among them, σ(.) represents the sigmoid function, represents the concatenation of vectors, W t ∈R 64×64 is a learnable linear transformation matrix, is the attention vector, is the graph attention coefficient, LeakyReLU(.) is the rectified linear unit function, is the normalized result of the graph attention coefficient, Soft max j (.) is the normalized exponential function, T i 、T j are two elements in the feature sequence.
[0080] S5, concatenate multiple feature sequences with the time attention parameters and input them into the auxiliary learning task, reconstruct the multiple feature sequences, and calculate the average error as the loss;
[0081] In this embodiment, for the input feature sequence x and the time attention parameter h t , the input feature sequence x is concatenated with the temporal attention parameter vector to obtain a new sequence, such as Figure 4 As shown, the new sequence is input into the 3-layer GRU encoder, and the GRU encoder will output a 256-dimensional time feature representation R t , where R t The calculation formula is:
[0082]
[0083] Among them G e (.) is the GRU encoder, x is a multi-segment feature sequence, h t is the temporal attention parameter.
[0084] like Figure 4 As shown, the temporal feature representation R after being encoded by the 3-layer GRU encoder t Input to the GRU decoder to reconstruct multiple feature sequences, and then use the mean square error (MSE) as the loss function to participate in the calculation of the total loss function. The specific calculation formula is:
[0085] L MSE =MSE(G d (R t ), x);
[0086] Among them G d (.) is the GRU decoder, R t is the time feature representation, x is the input feature sequence, L MSEis the mean square error, MSE(.) is the mean square error function,
[0087] S6, input the spatial attention parameters into the fully connected encoder Dense to obtain spatial feature representation;
[0088] In this embodiment, the spatial attention parameters are flattened and input into the fully connected encoder Dense to output a 256-dimensional spatial feature representation R s :
[0089] R s =D e (flatten(h s ));
[0090] Where D e (.) represents the fully connected encoder Dense, flatten(.) represents the flattening operation of the vector, h s represents the spatial attention parameter.
[0091] S7, concatenate the temporal feature representation and the spatial feature representation and input them into Dens e Fully connected classifier, outputs forgery status and calculates loss.
[0092] In this embodiment, if Figure 4 As shown in Figure 2, the temporal feature representation and the spatial feature representation are vectorized in series to obtain a 512-dimensional input sequence, which is then input into Dens e The fully connected classifier outputs the forgery of the sequence after the Sigmoid activation function, and uses binary cross entropy as the loss function. The specific calculation formula is:
[0093]
[0094] Where D c (.) is a Dense fully connected classifier, is the attribute label of the input sequence sample, is the vector concatenation of temporal feature representation and spatial feature representation, L BCE is the binary cross entropy loss, and BCE(.) is the binary cross entropy loss function.
[0095] S8, calculate the total loss and perform information entropy regularization;
[0096] In this embodiment, the binary classification task loss and the auxiliary learning task loss are added, and information entropy regularization is added as the total loss function. The total loss function is calculated, and the specific calculation formula is:
[0097] L=L BCE +LMSE ;
[0098] Final_L=L+α*H;
[0099] Where H is the output of information entropy, α is a hyperparameter. As an embodiment, α is 0.1, L is the intermediate loss, and Final_L is the information entropy regularization loss.
[0100] S9. Use the total loss to train the dual-flow graph network model, update the weight coefficients in the entire dual-flow graph network model, and save the weight information at the best verification accuracy.
[0101] In this embodiment, Adam is used as the optimizer, the learning rate lr is set to 0.001, the neural network unit dropout rate dropout is set to 0.5, and the training data size batchsize is set to 256. The weight coefficient of the entire network model is optimized, and the network model weight information under the best verification accuracy is saved.
[0102] S10, model application: load the optimal weight information into the dual-stream graph network model, input the test set, and output the classification results of the tampered video of the test set;
[0103] In this embodiment, the best weights obtained by training video data synthesized using four different face-changing technologies are loaded into the model, and the model is set to verification mode. The test data set is input, and the classification accuracy is calculated. The accuracy is used as the performance evaluation index of the model. It represents the proportion of correctly classified samples to the total number of samples. The higher the accuracy, the better the classification effect of the model. The calculation formula of the accuracy is as follows:
[0104]
[0105] Among them, TP is the number of samples correctly predicted as positive, TN is the number of samples correctly predicted as negative, FP is the number of samples incorrectly predicted as positive, and FN is the number of samples incorrectly predicted as negative.
[0106] In this embodiment, TP is the number of samples accurately predicted as Deepfake videos, TN is the number of samples correctly predicted as real videos, FP is the number of samples incorrectly predicted as Deepfake videos, and FN is the number of samples incorrectly predicted as real videos.
[0107] The in-library test results of this embodiment trained on video data synthesized by four face-swap technologies: FaceSwap (FS), DeepFakes (DF), Face2Face (F2F), and NeuralTextures (NT) are shown in Table 1:
[0108] Table 1 FaceForensics++ database (c0) training model test results
[0109]
[0110] It can be seen from Table 1 that under various face-changing technologies, the recognition accuracy of the dual-stream graph network model of the present invention is relatively high, showing a good model classification effect.
[0111] In order to test the influence of information entropy on the overall model, this embodiment removes the information entropy related modules in the network under the same conditions. The test results are shown in Table 2:
[0112] Table 2 FaceForensics++ database (c0) training model (removing information entropy) test results
[0113]
[0114] It can be seen that the addition of information entropy has improved the accuracy of the dual-flow graph network model of the present invention to a certain extent, proving that the addition of information entropy can reduce the overfitting of the model and improve the classification ability of the model.
[0115] This embodiment also provides a system for implementing a Deepfake video detection method based on a dual-stream graph network, including: a video frame-by-frame processing module, a feature sequence division module, a graph attention module, a GRU encoding and decoding module, a Dense fully connected encoding and classifier module, and a total loss calculation module;
[0116] In this embodiment, the video frame-by-frame processing module is used to locate facial feature points in the video data in the data set, select specific feature points for De Ronig triangulation, calculate the affine transformation parameters of the triangular feature area between frames, and concatenate the affine transformation parameters between frames to obtain the total feature sequence.
[0117] In this embodiment, the feature sequence division module is used to divide the total feature sequence according to specific sliding window time and length to obtain multiple feature sequences.
[0118] In this embodiment, the graph attention module is used to process the feature sequence in space and time to obtain corresponding spatial and temporal attention parameters.
[0119] In this embodiment, the GRU encoding and decoding module is used as an auxiliary learning task to reconstruct the feature sequence and calculate the average error as the loss. The module includes GRU and Dense.
[0120] In this embodiment, the Dense fully connected encoding and classifier module is used as a binary classification task, takes the temporal and spatial feature representation of the vector concatenation as input, and outputs the sequence forgery situation. The module includes Dense and Sigmoid.
[0121] In this embodiment, the total loss calculation module is used to sum the losses of two tasks and perform information entropy regularization processing.
[0122] In this embodiment, the total loss function is calculated and back propagated at the same time, and the weights in the network are updated to save the weights of the current best situation.
[0123] In this embodiment, feature extraction and prediction classification are performed based on the trained dual-stream graph network, and the deepfake forgery status of each video in the test dataset is output.
[0124] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well.
Claims
1. A Deepfake video detection method based on a dual-flow graph network, characterized in that: The following steps are involved: Obtain the video data to be detected; Input the video data to be detected into the trained dual-stream graph network model to detect deep fakes in the video; The training of the dual-flow graph network model includes the following steps: S1. Preprocess the videos with faces in the dataset to obtain feature sequences; S2. Input the feature sequence into the dual-stream graph network model to obtain the corresponding spatial attention parameters and temporal attention parameters; perform vector concatenation on the temporal attention parameters and the feature sequence to obtain the temporal feature sequence; S3, taking the temporal concatenated feature sequence as the input to the 3-layer GRU encoder to obtain the temporal feature representation; taking the attention parameters in the spatial feature space as the input to the Dense fully connected encoder to obtain the spatial feature representation; S4, vectorize the temporal feature representation and the spatial feature representation to obtain a classification feature sequence, input the classification feature sequence into a Dense fully connected classifier, and use the binary cross entropy as the classification loss function to calculate the loss of the output result of the Dense fully connected classifier; S5, input the temporal feature representation into the GRU decoder, and the output result uses the mean square error to calculate the loss; S6. Add the above two losses and add information entropy regularization as the total loss function, calculate the total loss function, perform backpropagation and update the weights in the dual-flow graph network model, and save the weight information of the best verification accuracy.
2. The Deepfake video detection method based on dual-flow graph network according to claim 1 is characterized in that: The preprocessing of the videos with human faces in the data set to obtain a feature sequence includes the following steps: Read the video with human face in the data set, extract the face area of the video frame by frame, locate the facial feature points in the face area, and divide the face area using the facial feature points and the Deroni triangulation method; Calculate and fit the affine transformation matrix between the same areas in two consecutive frames in the video, calculate the single-segment feature quantity between the matrices, repeat this process for the areas obtained by face segmentation and all frames, and finally concatenate to obtain the original multi-segment feature sequence, and at the same time divide the original multi-segment feature sequence into multi-segment feature sequences.
3. The Deepfake video detection method based on dual-flow graph network according to claim 2 is characterized in that: Using facial feature points and the Delonian triangulation method to divide the face into regions is to use the facial feature point positioning function in the insightFace model to recognize the face, select facial feature points, and use the Delonian triangulation method to divide the facial feature points into 7 regions, and each of the obtained regions is a triangle.
4. The Deepfake video detection method based on dual-flow graph network according to claim 2 is characterized in that: Calculate and fit the affine transformation matrix between the same regions in two consecutive frames in the video, calculate the single-segment feature quantity between the matrices, repeat this process for the regions obtained by face segmentation and all frames, and finally concatenate to obtain the original multi-segment feature sequence, and at the same time divide the original multi-segment feature sequence into multi-segment feature sequences. The specific steps include: The area obtained by dividing the face is used to calculate the affine transformation matrix M using the function of calculating affine transformation parameters in the OpenCV library. The affine transformation matrix M is used to calculate the average value, median value, average value, and median value of the lateral displacement, and the corresponding average value and median value are used as single-segment feature quantities. The single-segment feature quantities of all areas are merged to obtain a single-segment feature sequence, and then the remaining frames of the video are processed in the same way. The results of all frame processing are connected in series to obtain the original multi-segment feature sequence.
5. The Deepfake video detection method based on dual-flow graph network according to claim 4 is characterized in that: The original multi-segment feature sequence is divided into multiple-segment feature sequences. The specific steps include: dividing the original multi-segment feature sequence into multiple fixed-size feature sequences as the multiple-segment feature sequences according to the sliding window and the sliding length.
6. The Deepfake video detection method based on dual-flow graph network according to claim 2 is characterized in that: The dual-stream graph network model includes a spatial feature map attention module and a temporal feature map attention module. The feature sequence is input into the spatial feature map attention module and the temporal feature map attention module to obtain the corresponding spatial attention parameters and temporal attention parameters. The specific steps include: The feature sequence is regarded as a set of spatial feature sequences, and each element of the spatial feature sequence is modeled as a fully connected graph. The spatial feature graph attention module is used to output the spatial attention parameters. Among them, σ(.) represents the sigmoid function, ⊕ represents the concatenation of vectors, and W s ∈R 64×64 is a learnable linear transformation matrix, is the attention vector, is the graph attention coefficient, Leaky ReLU(.) is the rectified linear unit function, is the normalized result of the graph attention coefficient, Soft max j (.) is the normalized exponential function, S i , S j are two elements in the feature sequence; The feature sequence is regarded as a set of time feature sequences, each element of the time feature sequence is modeled as a fully connected graph, and the time feature graph attention module is used to output the time attention parameters Among them, σ(.) represents the sigmoid function, ⊕ represents the concatenation of vectors, and W t ∈R 64×64 is a learnable linear transformation matrix, is the attention vector, is the graph attention coefficient, Leaky ReLU(.) is the rectified linear unit function, is the normalized result of the graph attention coefficient, Soft max j (.) is the normalized exponential function, T i 、T j are two elements in the feature sequence.
7. The Deepfake video detection method based on dual-flow graph network according to claim 1, characterized in that: The hidden layer dimension of each GRU encoder of the 3-layer GRU encoder is 256; The Dense fully connected classifier includes three fully connected layers of Dense and one layer of Sigmoid. The input dimension of the first layer of Dense is 512 and the output dimension is 1024; the input dimension of the second layer of Dense is 1024 and the output dimension is 512; the input dimension of the third layer of Dense is 512 and the output dimension is 1; The Dense fully connected encoder includes two layers of Dense, the input dimension of the first layer of Dense is 28x64, and the output dimension is 1024; the input dimension of the second layer of Dense is 1024, and the output dimension is 256; The GRU decoder includes a layer of gated recurrent single GRU and a layer of Dense. The hidden layer dimension of GRU is 256, and the input dimension and output dimension of Dense are 256 and 64, respectively.
8. The Deepfake video detection method based on dual-flow graph network according to claim 1, characterized in that: The classification loss function formula is: Where D c (.) is a Dense fully connected classifier, is the attribute label of the input sequence sample, R s ⊕R t is the vector concatenation of temporal feature representation and spatial feature representation, L BCE is the binary cross entropy loss, BCE(.) is the binary cross entropy loss function; The loss formula using mean square error is: L MSE =MSE(G d (R t ),x); Among them G d (.) is the GRU decoder, R t is the time feature representation, x is the input feature sequence, L MSE is the mean square error, MSE(.) is the mean square error function; Add the above two losses and add information entropy regularization as the total loss function to calculate the total loss function: L=L BCE +L MSE ; Final_L=L+α*H; Where H is the output of information entropy, α is the hyperparameter, L is the intermediate loss, and Final_L is the information entropy regularization loss.
9. A system for implementing the Deepfake video detection method based on a dual-flow graph network as described in any one of claims 1 to 8, characterized in that: include: Video frame-by-frame processing module, feature sequence division module, graph attention module, GRU encoding and decoding module, Dense fully connected encoding and classifier module, and total loss calculation module; The video frame-by-frame processing module is used to locate facial feature points in the video data in the data set, select facial features for Deroni triangulation, calculate affine transformation parameters of the triangular feature regions between frames, and concatenate the affine transformation parameters between frames to obtain a total feature sequence; The feature sequence division module is used to divide the total feature sequence according to the set sliding window time and length to obtain multiple feature sequences; The graph attention module is used to process multiple feature sequences in space and time to obtain corresponding spatial attention parameters and temporal attention parameters; The GRU encoding and decoding module is used as an auxiliary learning task to reconstruct multiple feature sequences and calculate the average error as the loss. The GRU encoding and decoding module includes a GRU decoder and a fully connected layer Dense; The Dense fully connected encoding and classifier module is used as a binary classification task, takes the temporal feature representation and spatial feature representation of the vector concatenation as input, and outputs the sequence forgery situation. The Dense fully connected encoding and classifier module includes a fully connected layer Dense and a Sigmoid function; The total loss calculation module is used to sum the losses of the two tasks, perform information entropy regularization processing to calculate the total loss function, and simultaneously perform backpropagation and update the weights in the dual-stream graph network model, saving the weight information of the best verification accuracy; based on the trained dual-stream graph network model, feature extraction and prediction classification are performed, and the deepfake forgery status of each video in the output data set is output.
10. A computer device, characterized in that: It comprises: a memory and a processor and a computer program stored in the memory. When the computer program is executed on the processor, the Deepfake video detection method based on the dual-flow graph network as described in any one of claims 1 to 8 is implemented.