Face video depth forgery detection method based on face micro expression
Through a deep forgery detection method based on face micro-expressions, the micro-action time series is extracted and combined with timing consistency learning and GRU network, the problems of high computational complexity and insufficient detection performance in the prior art are solved, and efficient and accurate forgery video detection is achieved.
Patent Information
- Application Number
- CN202510075299.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has high computational complexity and is difficult to maintain high detection performance under challenging manipulation techniques when detecting facial video depth forgery.
A deep forgery detection method based on face micro-expressions is adopted to extract the micro-action time series through the division of face interest areas, sampling of point of interest, and optical flow calculation. Combining time-sequence consistency learning and GRU network architecture, an attention mechanism is introduced to build a true and false detection model.
It improves the accuracy and generalization of forgery detection, reduces the complexity of the model, broadens the application scenarios, and can effectively capture the abnormal characteristics of the micro-action time series at a short scale.
Smart Images

Figure CN120108015A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting deep fakes in face videos, and in particular to a method for detecting deep fakes in face videos based on facial micro-expressions. Background Art
[0002] This section merely provides background information related to the present disclosure and is not necessarily prior art.
[0003] With the rapid development of computer graphics and artificial intelligence technology, face forgery technology has become more and more complex and realistic. As face forgery technology continues to evolve, forged videos generated using deep learning have become a major threat to society. Therefore, detecting and mitigating the risks posed by these forged videos has become an urgent research challenge.
[0004] Face forgery video detection algorithms can be divided into two categories: intra-frame detection-based and inter-frame detection-based. Intra-frame detection-based algorithms identify authenticity by detecting forgery traces exposed in a single image. However, due to the inconsistency in temporal sequence, they cannot achieve satisfactory accuracy and generalization. The difficulty of inter-frame detection-based methods is that compared with intra-frame detection, they need to additionally consider the forgery traces in temporal sequence. Recently, researchers have proposed some methods to check the consistency and coherence between consecutive frames by analyzing the temporal characteristics of face forgery videos. By modeling the temporal relationship between frames, these methods can identify differences and irregularities that indicate the presence of face forgery. An obvious disadvantage of video-level detection methods is the high computational complexity. One strategy to achieve a more lightweight detection is to exploit facial geometric features. These methods use geometric properties to capture temporal patterns and dependencies, thereby converting the original high-dimensional feature space into a low-dimensional representation. However, when faced with challenging manipulation techniques, these methods cannot show satisfactory detection performance.
[0005] Therefore, researching and developing effective deep fake video detection technology has important practical significance and urgent social needs.
[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0007] Purpose of the invention: The technical problem to be solved by the present invention is to provide a method for detecting deep fakes in facial videos based on facial micro-expressions in response to the shortcomings of the prior art.
[0008] In order to solve the above technical problems, the present invention discloses a method for detecting deep fakes in face videos based on facial micro-expressions, comprising the following steps:
[0009] Step 1, extracting facial micro-expression movements from the facial video to obtain a micro-movement time series;
[0010] Step 2, performing time consistency learning according to the micro-motion time series to obtain a new micro-motion time series;
[0011] Step 3: Based on the GRU network architecture, the attention mechanism is introduced to build a true-false detection model to capture the abnormal characteristics of micro-motion time series at a short scale;
[0012] Step 4, training the authenticity detection model using the new micro-motion time series described in step 2;
[0013] Step 5: Perform authenticity detection. For the unknown face video to be detected, after extracting the facial micro-motion trajectory, use the authenticity detection model trained in step 4 to detect abnormal features of the time series to obtain a classification result of whether the face video to be detected is a forged video.
[0014] Furthermore, the step of extracting facial micro-expression movements from the facial video in step 1 includes:
[0015] Step 1-1, dividing the face interest regions, that is, dividing the face in the face video into 7 interest regions and 1 irrelevant interest region;
[0016] Step 1-2, sampling of interest points in the area, that is, uniform sampling is performed in the interest area to obtain interest points;
[0017] Steps 1-3: perform optical flow calculations on adjacent frames to obtain the final micro-motion time series.
[0018] Furthermore, the temporal consistency learning described in step 2 includes:
[0019] A time series prediction network is constructed to learn the motion trajectory of the facial micro-expression movements, that is, to make predictions based on the micro-movement time series.
[0020] Furthermore, the training of the authenticity detection model in step 4 includes:
[0021] For a face video containing a known authenticity label, the original micro-motion time series is obtained according to the method described in step 1, and prediction is performed according to the method described in step 2. The original micro-motion time series is concatenated with the predicted time series and input into the authenticity detection model constructed in step 3 for authenticity classification. A loss function is constructed based on the classification results and the authenticity labels, and the model parameters are updated in an iterative manner.
[0022] Furthermore, the interest regions described in step 1-1 include: left eyebrow, right eyebrow, left eye, right eye, nose, mouth and jaw, and the remaining regions are irrelevant regions.
[0023] Furthermore, the facial interest region division described in step 1-1 includes:
[0024] Step 1-1-1, using a face detection method to capture the face area in the face video, and taking a preset number of consecutive frames as a video clip;
[0025] Step 1-1-2, using the landmarks detection method to detect a preset number of facial feature points in the first frame of the video clip, and dividing 7 facial interest regions based on the detected points.
[0026] Furthermore, the adjacent frame optical flow calculation described in steps 1-3, that is, for every two consecutive frames in the face video, the corner points in the interest-irrelevant area are extracted using a corner point extraction method, and a position change matrix is calculated based on the corner points; the trajectory of the interest point is tracked by a dense optical flow method, and the posture is corrected by the position change matrix to obtain the micro-motion offset of the interest point; and the interest point whose micro-motion offset is less than a threshold is eliminated, including:
[0027] Step 1-3-1, for the input adjacent frame image, that is, the first frame image F i and the next frame image F i+1 , in image F i The irrelevant area of interest in the image is extracted as follows:
[0028]
[0029] Among them, S corner represents the set of corner points, represents the nth corner point;
[0030] Step 1-3-2, use the Lucas-Kanade optical flow method to track the corner points in image F i+1 The position in is as follows:
[0031]
[0032] Among them, S predict Represents image F i+1 The set of corner points in , Represents image F i+1 The nth corner point in ;
[0033] Step 1-3-3, obtain the column matching point pair set S match ,as follows:
[0034]
[0035] Step 1-3-4, match the point pair set S according to the column match , calculate the affine transformation matrix H, and according to its inverse transformation matrix H -1 , for the next frame image F i+1 Perform posture correction as follows:
[0036]
[0037] in, Indicates the next frame image after posture correction;
[0038] Among them, the specific method of calculating the affine transformation matrix H is as follows:
[0039] The affine transformation is expressed as:
[0040]
[0041] Among them, (x, y) is the coordinate of the original point, (x ′ ,y ′ ) is the coordinate of the changed point, a, b, c, d are the parameters of the affine transformation matrix, tx, ty are the translation parameters; the equation is constructed by the least squares method:
[0042]
[0043] Substituting the matching point pairs into the above equations, we get the linear equation system:
[0044] A*p=b
[0045] Among them, A is the matrix containing all matching point pairs, p is the parameter vector to be solved, and b is the matrix of target points; the parameters of the affine transformation matrix are obtained by solving the equation by the least squares method:
[0046] p=(A T A) -1 A T b
[0047] Thus, the affine transformation matrix H is obtained;
[0048] Step 1-3-5, for the first frame image F i And the next frame image after posture correction Calculate its dense optical flow field Flow i ;
[0049] Step 1-3-6, for image F i The kth point of interest Estimate its coordinates in the next frame as follows:
[0050]
[0051] Wherein, M is the median filter;
[0052] Step 1-3-7, Image F i Micro-motion offset Δd i It is expressed as follows:
[0053]
[0054] Step 1-3-8, through image F i The initial value of the coordinates of the point of interest is calculated in the image F i+1 The estimated value of image F i+1 The estimated value of is updated to the initial value and the calculation is continued. According to the above method, the calculation is performed from the first frame to the last frame in the video segment to obtain the offset time sequence d of the video segment, which is expressed as follows:
[0055] d=[Δd 1 ,Δd 2 ,…,Δd T-1 ]
[0056] Wherein, T-1 represents the last frame of the video clip;
[0057] Step 1-3-9, screen each region of interest, remove the points of interest whose micro-motion offset is less than the threshold, and obtain the micro-motion time series vector S.
[0058] Furthermore, the time series prediction network described in step 2 includes:
[0059] Input layer, causal convolution layer, dilated convolution layer and output layer; among them,
[0060] Input layer, used to obtain the input micro-action time series vector S;
[0061] The causal convolution layer uses causal convolution to process the micro-action time series vector S of the input layer and obtain the embedding vector E;
[0062] The dilated convolutional layer is composed of a superposition of dilated convolutional networks. After the embedding vector E is input, the encoding vector V with extracted temporal features is obtained;
[0063] The output layer outputs the encoded vector V through the relu activation function and the softmax function to obtain the final predicted value Output of the time series.
[0064] Furthermore, the temporal consistency learning described in step 2 includes:
[0065] Step 2-1, using the micro-motion time series obtained from the real face video as a training set, training the time series prediction network in an autoregressive manner to obtain a pre-trained time series prediction network;
[0066] Step 2-2, will be from t 0 The trajectory coordinates of w time steps before the moment begins are set as the initial receptive field context, and the predicted time step is t w The trajectory coordinates at the moment;
[0067] Step 2-3, change t w The predicted value at time t is appended to the input sequence, and the current receptive field context window is moved forward by one time step, and the predicted time step is t. w+1 The trajectory coordinates at the moment;
[0068] Step 2-4: perform a preset number of iterations according to step 2-3. In each iteration, the context window moves forward one time step, keeping the length of the receptive field unchanged.
[0069] Step 2-5, given a series of trajectories of real face videos or forged face videos, use the pre-trained time series prediction network to predict the trajectories of a preset number of frames in the future using the methods of steps 2-2 to 2-4, that is, predict one frame at a time, and use the predicted value at the current moment for the prediction at the next moment;
[0070] Step 2-6, connect the original trajectory with the predicted trajectory to obtain a new micro-motion time series.
[0071] Furthermore, the authenticity detection model described in step 3 includes:
[0072] Input layer, GRU network layer, attention mechanism layer and output layer; among them,
[0073] The input layer is used to obtain the input multivariate time series;
[0074] The GRU network layer uses a gated structure, including an update gate and a reset gate, to capture long-term dependencies in the input multivariate time series;
[0075] The attention mechanism layer consists of an encoder and a decoder. The encoder generates an attention vector from the input, while the decoder generates a hidden state by taking the encoder output as input. The hidden state is divided into multiple views. The encoder uses the hidden state of the previous view decoder to assign an attention score to the hidden state of each step. The attention vector is obtained by performing a soft-max operation on the attention score.
[0076] At the output layer, the attention vector passes through the fully connected layer to complete the classification and obtain the authenticity inspection result.
[0077] Beneficial effects:
[0078] 1. The present invention adopts a video-level, end-to-end approach, taking into account the temporal inconsistency of forged video sequences, thereby improving the accuracy and generalization of forgery detection.
[0079] 2. The present invention uses the micro-motion modeling method to reduce the dimension of the video stream into a multivariate time series, overcoming the shortcomings of the inter-frame based forgery detection method, such as large input dimension and high model complexity, and broadening the practical application scenarios.
[0080] 3. The present invention designs a multivariate time series analysis network based on GRU, introduces an attention mechanism, and improves prediction accuracy by focusing on more relevant input variables. The multivariate time series based on micro-movements has a small input size in both the variable dimension and the time dimension, and there is no long-scale correlation. Therefore, the inherent motion pattern of micro-movements exists in the time domain of a short window in the time dimension, and exists between local motion units in the variable dimension. The invention uses an attention mechanism to capture this local feature well. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0082] Figure 1 This is a flowchart of the implementation of the micro-motion trajectory extraction algorithm of an embodiment of the present invention.
[0083] Figure 2 It is a system diagram of the detection method of the present invention.
[0084] Figure 3 4 is a diagram of the authenticity detection network structure in an embodiment of the present invention.
[0085] Figure 4 This is a diagram of the timing prediction network structure in an embodiment of the present invention.
[0086] Figure 5 FIG. 1 is a schematic diagram of facial feature points extracted in an embodiment.
[0087] Figure 6 The figure is a schematic diagram of multivariate time series comparison for false video generation in one embodiment.
[0088] Figure 7 The figure is a schematic diagram of multivariate time series comparison generated for real video in an embodiment. DETAILED DESCRIPTION
[0089] In order to solve the problem that the above-mentioned existing frame-based forgery detection algorithms cannot strike a good balance between computational efficiency and detection capability, the present invention proposes a forgery detection method based on facial micro-expressions, which models the global motion features of facial video streams as multivariate time series analysis, so as to achieve the purpose of reducing dimensionality and taking into account global information.
[0090] The motivation for the method proposed by the present invention comes from a statistical law: when a real face speaks, there is an inherent facial expression and movement pattern, and the fake face will destroy this inherent expression and movement pattern during the synthesis process. Therefore, this abnormal movement pattern can be detected by analyzing the changes in local facial micro-expressions. The present invention describes facial micro-expressions as the movement deviation of local muscles in the facial area caused by changes in the overall displacement of the head, such as the movement of the eyes, lips, eyebrows, etc.
[0091] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0092] A method for detecting deep fakes in facial videos based on facial micro-expressions, comprising the following steps:
[0093] The facial micro-expression extraction step extracts the dense trajectory of facial feature points by dividing the facial interest region, sampling the interest points in the region, and calculating the optical flow of adjacent frames. In order to ensure reliable and accurate extraction of facial motion trajectories, the present invention additionally compensates and optimizes the impact of the overall position change of the human head in the video on the optical flow calculation;
[0094] The temporal consistency learning step builds a temporal prediction network, which is trained only on videos with real labels to learn the intrinsic patterns of motion trajectories formed by real facial videos;
[0095] The classification model construction step is based on the classic GRU network architecture and introduces the attention mechanism to capture the abnormal characteristics of micro-motion time series at a short scale by focusing on the correlation between the time dimension and the spatial dimension;
[0096] Classification model training step: for face videos with known authenticity labels, model trajectory sequences based on facial micro-expressions. Input the extracted trajectory into the time series prediction network to obtain the predicted trajectory. Concatenate the predicted trajectory with the original trajectory and input them into the classifier for forgery detection. Use the constructed model to classify authenticity, and build a loss function based on the classification results and authenticity labels. Update the model parameters through multiple rounds of iterations to optimize the detection capability.
[0097] In the authenticity detection step, for the face video of unknown authenticity, the facial micro-expression trajectory is extracted and input into the time series prediction network to obtain the predicted trajectory. The predicted trajectory is spliced with the original trajectory, and the trained forgery detection model is used to perform time series abnormal feature detection to obtain the classification result of whether the face video to be detected is a forged video.
[0098] The facial video forgery detection method based on facial micro-expressions, wherein the facial micro-expression extraction step comprises:
[0099] Based on facial landmarks, the face is divided into 7 interest regions and 1 uninterested region. The interest regions include the left eyebrow, right eyebrow, left eye, right eye, nose, mouth, and jaw, which are 7 regions with rich micro-expression movements. The uninterested region refers to the region with less facial micro-expressions. Uniform sampling is performed in the interest region to obtain interest points;
[0100] For every two consecutive frames in the video, use the corner point extraction algorithm to extract the corner points that are irrelevant to the micro-expression movement of the face, obtain the position correspondence through corner point matching, and calculate the position change matrix that is irrelevant to the micro-expression movement;
[0101] For the extracted points of interest, the trajectory is tracked by dense optical flow method, and the micro-expression motion offset after posture correction is calculated by the position change matrix calculated in the previous step;
[0102] The points of interest with unclear micro-expression movements are removed to obtain the final multivariate time series.
[0103] In the facial video forgery detection method based on facial micro-expressions, the temporal consistency learning step comprises:
[0104] Pre-build a classic time series prediction network and micro-motion trajectories formed by real face videos only;
[0105] Use the constructed temporal prediction network to learn the intrinsic patterns of micro-motion trajectories formed by real face videos;
[0106] In the forgery detection stage, the extracted micro-motion trajectories are connected with the trajectories predicted by the temporal prediction network and input into the classification model for forgery detection;
[0107] Since the prediction model is trained on real videos, the concatenated time series of real videos are very similar to the original time series. However, for synthetic videos, since they come from different data distributions, this concatenation method can amplify the difference patterns caused by manipulation. By leveraging the spatial and temporal information captured in the facial motion trajectory, this data augmentation method can improve the accuracy and robustness of facial forgery detection.
[0108] The facial video forgery detection method based on facial micro-expressions, wherein the model building step comprises:
[0109] Construct a GRU module to extract features from the input multivariate time series in the time dimension and generate a feature vector;
[0110] Construct an attention mechanism encoder to generate an attention vector based on the input, indicating the relative importance of each input time step;
[0111] Construct an attention mechanism decoder to generate the hidden state of the current time step and use the hidden state of the decoder of the previous time step to calculate the attention score for each input time step;
[0112] By performing a softmax operation on the attention score, an attention vector is obtained. When predicting the output value, the decoder is used to focus on the input variables that are more relevant to the predicted output based on the attention vector.
[0113] The facial video forgery detection method based on facial micro-expressions, wherein the model training step comprises:
[0114] Build a dataset of forged face sample videos for training, set training rounds and optimizer, perform true and false classification, and iterate model parameters.
[0115] The method for detecting forgery of a face video based on micro-expressions of a face, wherein the authenticity detection step comprises:
[0116] The micro-motion extraction algorithm is used to extract multivariate time series, and the trained model is used for authenticity identification.
[0117] Embodiment 1:
[0118] The present invention provides a method for detecting the authenticity of a forged video based on human facial micro-expressions, comprising a human facial micro-expression extraction algorithm, a temporal consistency learning method, and a GRU-based authenticity detection model.
[0119] The facial regions of real faces come from the same data distribution, so there are inherent facial micro-expression motion modes. For synthetic faces, whether it is face-swapping, splicing or deepfakes generation, since the data from different positions of the face come from different data distributions, there will be different motion patterns between different micro-expression motion units. The algorithm models this motion pattern by calculating the displacement of facial muscle movement. The micro-expression motion extraction algorithm of the invention includes three modules: micro-expression motion interest point extraction, displacement trajectory optimization, and multivariate time series modeling. Figure 1 and Figure 2 shown.
[0120] The micro-expression action interest point extraction consists of the following steps:
[0121] For the original video, use the face detection algorithm (reference: Viola P, Jones M. Rapid object detection using a boosted cascade of simple features [C] / / Proceedings of the 2001 IEEE computer society conference on computer vision and pattern recognition. CVPR 2001. Ieee, 2001, 1: II.) to capture the face area in the video, and then convert several consecutive frames into a video clip;
[0122] For the first frame in the clip, the classic landmarks detection algorithm (reference: Kazemi V, Sullivan J. One millisecond face alignment with an ensemble of regression trees [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2014: 1867-1874.) is used to detect 68 facial feature points, and then 7 facial interest regions are divided accordingly. The principle is that the facial behavior coding system defines several motion units in the face area, and all facial expression movements can be constructed by these motion units. Based on the definition of motion units, the face can be divided into 7 interest regions, each of which represents a local micro-expression movement pattern;
[0123] In each area, a dense sampling method is used to densely and evenly sample a number of interest points for subsequent calculation of micro-expression action offsets;
[0124] For each frame, an additional area irrelevant to micro-expression movement is divided to eliminate the displacement calculation error caused by irrelevant micro-expression movements and optimize the displacement trajectory;
[0125] The displacement trajectory optimization module uses a facial alignment method to eliminate calculation errors caused by other facial movements. The previous facial alignment method solved the affine transformation matrix through the landmark position coordinates to correspond the faces of all frames to the same position. This method has certain limitations. Because the area where the landmarks are located is usually an area where the geometric features change dramatically, it contains a large amount of motion information and is greatly affected by changes in head posture and facial muscle movements. The affine transformation matrix calculated in this way cannot perform facial alignment well, and will also cause errors in the calculation of facial micro-movement displacement. In order to reduce the impact on the accuracy of facial micro-movement displacement calculation, the invention adopts the following steps to optimize the displacement trajectory:
[0126] For an input frame image Fi and the next frame Fi+1, extract the corner points in the micro-expression motion-independent area of Fi (reference: Rosten E, Drummond T. Machine learning for high-speed corner detection [C] / / Computer Vision–ECCV 2006: 9th European Conference on Computer Vision, Graz, Austria, May 7-13, 2006. Proceedings, Part I 9. Springer Berlin Heidelberg, 2006: 430-443.):
[0127]
[0128] Use the Lucas-Kanade optical flow algorithm (reference: Lucas B and Kanade T. An Iterative Image Registration Technique with an Application to Stereo Vision. Proc. Of 7th International Joint Conference on Artificial Intelligence (IJCAI), pp. 674-679.) to track the positions of these points in Fi+1:
[0129]
[0130] Get a list of matching point pairs:
[0131]
[0132] However, the disappearance of optical flow will cause tracking point offset, and abnormal conditions such as occlusion and motion blur can easily lead to feature point mismatching. These problems will affect the accuracy of facial alignment and thus affect subsequent face forgery detection. In order to improve the robustness of feature point matching, the present invention uses the RANSAC algorithm (reference: Raguram R, Chum O, Pollefeys M, et al. USAC: A universal framework for random sample consensus [J]. IEEE transactions on pattern analysis and machine intelligence, 2012, 35 (8): 2022-2038.). The RANSAC algorithm can effectively estimate model parameters in the presence of a large number of outliers, thereby improving the accuracy of matching. The affine transformation matrix H and its inverse transformation matrix H-1 are calculated through the processed matching point pairs. Finally, the next frame after facial posture correction is obtained:
[0133]
[0134] The invention uses the optical flow method to track the motion trajectory of the point of interest, thereby calculating the micro-expression motion offset between two frames. The sparse optical flow algorithm relies on the detection and tracking of image feature points, and is easily affected by factors such as image noise, occlusion, and illumination changes, resulting in feature point loss or tracking errors. In addition, the sparse optical flow algorithm can only reflect the movement of local feature points and cannot obtain the global motion information of the entire scene, which will cause some important areas to be missed. The dense optical flow can obtain the motion information of the entire image and is more robust. The invention uses a dense optical flow algorithm and combines it with a median filter to improve the accuracy of interest point tracking. Based on the dense optical flow algorithm, the invention proposes the following specific steps for multivariate time series modeling:
[0135] For the input frame Fi and the corrected next frame Fi+1, calculate its dense optical flow field Flow i (refer to: G.Two-frame motion estimation based on polynomial expansion[C] / / ImageAnalysis:13th Scandinavian Conference,SCIA 2003 Halmstad,Sweden,June 29–July2,2003 Proceedings 13.Springer Berlin Heidelberg,2003:363-370.).
[0136] For the kth interest point in the i-th frame, the following formula is used to estimate its coordinates in the next frame (M is the median filter):
[0137]
[0138] Through the initial value of the interest point coordinates of F1, its estimated value in F2 is calculated. Their difference is the micro-expression motion offset of F1. The micro-expression offset of Fi can be expressed as:
[0139]
[0140] Update the estimated value of F2 to the initial value, calculate the estimated value of F3, and so on, to obtain the offset timing sequence:
[0141] d=[Δd 1 ,Δd 2 ,...,Δd T-1 ] (Formula 7)
[0142] Each region of interest is screened to filter out some points of interest with unclear motion. For all points of interest in the jth region of interest, only the time series formed by the points of interest with significant offset are retained.
[0143] The invention also proposes a temporal consistency learning module for differentiating the different micro-motion trajectories formed by real and fake faces. The core idea is to use a temporal prediction network to learn the intrinsic patterns of the motion trajectories formed by real face videos. In this invention, a temporal prediction model based on WaveNet is used (reference: Van Den Oord A, Dieleman S, Zen H, et al. Wavenet: A generative model for raw audio [J]. arXiv preprint arXiv: 1609.03499, 2016, 12.). Figure 3 As shown, the network structure is as follows:
[0144] (1) Input layer: The input layer is the input micro-action time series vector S.
[0145] (2) Causal convolution layer. The causal convolution layer uses causal convolution (reference: Bai S, Kolter JZ, Koltun V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling [J]. arXiv preprint arXiv: 1803.01271, 2018.) to process the input layer vector to obtain the embedding vector E. Causal convolution means that the output elements can only depend on the previous input elements, but not on the future elements, ensuring that the model output does not violate the order of the data, which is crucial for time series prediction.
[0146] (3) Dilated convolutional layer. The embedding vector E is input into the dilated convolutional layer to obtain the encoding vector V that extracts the temporal features. The dilated convolutional layer is composed of a series of dilated convolutional networks (reference: Bai S, Kolter JZ, Koltun V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling [J]. arXiv preprint arXiv: 1803.01271, 2018.). By stacking multiple layers of dilated causal networks, the model can achieve a larger receptive field while avoiding high computational costs. After considering a large number of time steps within the receptive field, the model can capture the subtle and global patterns of the input data over time.
[0147] (4) Output layer: The encoded vector V is output through the relu activation function and the softmax function to obtain the final time series prediction value Output.
[0148] The network structure of the core network expansion convolution layer is as follows:
[0149] (1) The input first passes through a series of residual blocks, each of which sets an exponentially increasing dilation rate, starting from 2 in block 0. 0 To the pth block 2 p ,[2 0 ,2 1 ,2 2 ,...,2 k ]. Each consecutive p blocks are taken as a group and repeated k times to improve the model accuracy.
[0150] (2) Inside each residual block, the input first passes through a dilated convolution. Dilated convolution is a convolution operation that incorporates dilation into the convolution layer. Dilation means skipping input values with defined intervals, which are specified by the dilation rate; the output of the dilated convolution passes through a gated activation unit composed of two parallel activation functions (tanh and sigmoid) to avoid overfitting and speed up convergence; finally, the output of the gated unit is added through a one-dimensional convolution layer and a residual link to obtain the final output.
[0151] (3) The outputs from all residual blocks are stacked and then passed through a dense layer to obtain the final prediction value.
[0152] By stacking multiple layers of dilated causal convolutions, the model is able to achieve a larger receptive field while avoiding high computational cost. After considering a large number of time steps within the receptive field, the model is able to capture subtle and global patterns of the input data over time;
[0153] The algorithm flow of the temporal consistency learning module is as follows:
[0154] (1) Using the micro-motion time series obtained from real face videos, the above-mentioned timing prediction network is trained in an autoregressive manner.
[0155] (2) First, we will start from t 0 The trajectory coordinates of w time steps before the moment begins are set as the initial receptive field context, and the predicted time step is t w The trajectory coordinates at the time.
[0156] (3) Set t w The predicted value at time t is appended to the input sequence, and the receptive field context window is moved forward by one time step, and the predicted time step is t. w+1 The trajectory coordinates at the time.
[0157] (4) Perform several iterations according to the above steps. In each iteration, the context window moves forward one time step to keep the length of the receptive field unchanged.
[0158] (5) Given a series of real or fake trajectories, the previously trained time series prediction network is used to iteratively predict the trajectories of several future frames, i.e., predicting one frame at a time and using the predicted value at the previous moment for the prediction at the next moment.
[0159] (6) Connect the original trajectory and the predicted trajectory and input them into the authenticity detection model for classification.
[0160] The temporal consistency learning method creates a joint representation that captures the motion pattern by combining the original and predicted trajectories. This connection amplifies the data distribution difference between the processed real and fake trajectories, which benefits the subsequent classification task of the real and fake detection network. It is worth noting that for the real video trajectories, the predicted trajectories are well aligned with the original trajectories, while for the fake video trajectories, the difference in motion patterns leads to a clear difference between the predicted and original trajectories. This method of enhancing the data distribution difference enables the detection network to perform classification more effectively.
[0161] The invention also proposes a true-false detection model that combines the GRU network and the attention mechanism, such as Figure 4 As shown. The GRU network is composed of multiple GRU units. Each GRU unit contains an input gate and a forget gate. A gate controller controls both the input gate and the forget gate. When the gate controller is 1, the forget gate is closed and the input gate is open. When the gate controller is 0, the forget gate is open and the input gate is closed. At each time step, the memory of the previous moment will be saved, and the input of the current time step will be cleared.
[0162] The specific steps are as follows:
[0163] Define the input dimension of the GRU unit, which is the feature dimension of the time series. Stack multiple GRU layers to capture the longer-scale dependencies on the time dimension in the time series;
[0164] Construct an attention mechanism, construct an encoder to generate an attention vector from the input data. Construct a decoder, take the encoder's output as input and generate a hidden state vector. These hidden state vectors are divided into multiple groups. For each group, the encoder uses the decoder's hidden state vector of the previous group to assign an attention score to each step in the current view group. The final attention vector is generated by performing a softmax operation on these attention scores;
[0165] Finally, the attention vector is input into the fully connected layer for classification.
[0166] Embodiment 2:
[0167] Using the faceforensics++ dataset (reference: Andreas Rossler, David Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies and Matthias NieBner, "FaceForensics++: Learning to Detect Manipulated Facial Images", 2019 IEEE / CVF International Conference on Computer Vision, (2019)) as an example, the extracted facial feature points are as follows Figure 5 As shown;
[0168] For the multivariate time series generated by fake videos, the comparison between the original time series (left) and the spliced time series after the temporal consistency learning module (right) is shown in the figure. Figure 6 shown.
[0169] For the multivariate time series generated by real videos, the comparison diagram of the original time series (left) and the spliced time series (right) after the temporal consistency learning module is shown in the figure. Figure 7 shown.
[0170] It can be seen that for the real video trajectory, the alignment between the predicted trajectory and the original trajectory is good; while for the forged video trajectory, there will be obvious differences between the predicted trajectory and the original trajectory due to the difference in motion patterns.
[0171] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit, wherein the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, the invention content of a method for detecting deep fakes of a face video based on micro-expressions of a face provided by the present invention and some or all of the steps in each embodiment can be run. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0172] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of computer programs and their corresponding general hardware platforms. Based on such an understanding, the technical solutions in the embodiments of the present invention can be essentially or partly contributed to the prior art in the form of computer programs, i.e., software products, which can be stored in a storage medium and include several instructions for enabling a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, an MCU or a network device, etc.) to execute the methods described in various embodiments of the present invention or certain parts of the embodiments.
[0173] The present invention provides a method and idea for deep fake detection of face videos based on micro-expressions of human faces. There are many methods and ways to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A method for detecting deep fakes in facial videos based on facial micro-expressions, characterized in that: The following steps are involved: Step 1, extracting facial micro-expression movements from the facial video to obtain a micro-movement time series; Step 2, performing time consistency learning according to the micro-motion time series to obtain a new micro-motion time series; Step 3: Based on the GRU network architecture, the attention mechanism is introduced to build a true-false detection model to capture the abnormal characteristics of micro-motion time series at a short scale; Step 4, training the authenticity detection model using the new micro-motion time series described in step 2; Step 5: Perform authenticity detection. For the unknown face video to be detected, after extracting the facial micro-motion trajectory, use the authenticity detection model trained in step 4 to detect abnormal features of the time series to obtain a classification result of whether the face video to be detected is a forged video.
2. According to claim 1, a method for detecting deep fakes in facial videos based on facial micro-expressions is characterized in that: The step 1 of extracting micro-expression movements of a human face from the human face video includes: Step 1-1, dividing the face interest regions, that is, dividing the face in the face video into 7 interest regions and 1 irrelevant interest region; Step 1-2, sampling of interest points in the area, that is, uniform sampling is performed in the interest area to obtain interest points; Steps 1-3: perform optical flow calculations on adjacent frames to obtain the final micro-motion time series.
3. According to claim 2, a method for detecting deep fakes in facial videos based on facial micro-expressions is characterized in that: The temporal consistency learning described in step 2 includes: A time series prediction network is constructed to learn the motion trajectory of the facial micro-expression movements, that is, to make predictions based on the micro-movement time series.
4. According to claim 3, a method for detecting deep fakes in facial videos based on facial micro-expressions is characterized in that: The training of the authenticity detection model described in step 4 includes: For a face video containing a known authenticity label, the original micro-motion time series is obtained according to the method described in step 1, and prediction is performed according to the method described in step 2. The original micro-motion time series is concatenated with the predicted time series and input into the authenticity detection model constructed in step 3 for authenticity classification. A loss function is constructed based on the classification results and the authenticity labels, and the model parameters are updated in an iterative manner.
5. According to claim 4, a method for detecting deep fakes in facial videos based on facial micro-expressions is characterized in that: The regions of interest described in step 1-1 include: left eyebrow, right eyebrow, left eye, right eye, nose, mouth and jaw, and the remaining regions are irrelevant regions of interest.
6. A facial video deep fake detection method based on facial micro-expressions according to claim 5, characterized in that: The facial interest region division described in step 1-1 includes: Step 1-1-1, using a face detection method to capture the face area in the face video, and taking a preset number of consecutive frames as a video clip; Step 1-1-2, using the landmarks detection method to detect a preset number of facial feature points in the first frame of the video clip, and dividing 7 facial interest regions based on the detected points.
7. The method for detecting deep fakes in facial videos based on facial micro-expressions according to claim 6, characterized in that: The adjacent frame optical flow calculation described in step 1-3 is performed, that is, for every two consecutive frames in the face video, the corner points in the interest-irrelevant region are extracted using a corner point extraction method, and a position change matrix is calculated based on the corner points; for the interest point, the trajectory is tracked by a dense optical flow method, and the posture is corrected by the position change matrix to obtain the micro-motion offset of the interest point; Eliminating the points of interest whose micro-motion offset is less than a threshold value includes: Step 1-3-1, for the input adjacent frame image, that is, the first frame image F i and the next frame image F i+1 , in image F i The irrelevant area of interest in the image is extracted as follows: Among them, S corner represents the set of corner points, represents the nth corner point; Step 1-3-2, use the Lucas-Kanade optical flow method to track the corner points in image F i+1 The position in is as follows: Among them, S predict Represents image F i+1 The set of corner points in , Represents image F i+1 The nth corner point in ; Step 1-3-3, obtain the column matching point pair set S match ,as follows: Step 1-3-4, match the point pair set S according to the column match , calculate the affine transformation matrix H, and according to its inverse transformation matrix H -1 , for the next frame image F i+1 Perform posture correction as follows: in, Indicates the next frame image after posture correction; Among them, the specific method of calculating the affine transformation matrix H is as follows: The affine transformation is expressed as: Among them, (x, y) is the coordinate of the original point, (x′, y′) is the coordinate of the changed point, a, b, c, d are the parameters of the affine transformation matrix, tx, ty are the translation parameters; the equation is constructed by the least squares method: Substituting the matching point pairs into the above equations, we get the linear equation system: A*p=b Among them, A is the matrix containing all matching point pairs, p is the parameter vector to be solved, and b is the matrix of target points; the parameters of the affine transformation matrix are obtained by solving the equation by the least squares method: p=(A T A) -1 A T b Thus, the affine transformation matrix H is obtained; Step 1-3-5, for the first frame image F i And the next frame image after posture correction Calculate its dense optical flow field Flow i ; Step 1-3-6, for image F i The kth point of interest Estimate its coordinates in the next frame as follows: Where, M is the median filter; Step 1-3-7, Image F i Micro-motion offset Δd i It is expressed as follows: Step 1-3-8, through image F i The initial value of the coordinates of the point of interest is calculated in the image F i+1 The estimated value of image F i+1 The estimated value of is updated to the initial value and the calculation is continued. According to the above method, the calculation is performed from the first frame to the last frame in the video segment to obtain the offset time sequence d of the video segment, which is expressed as follows: d=[Δd1,Δd2,…,Δd T-1 ] Wherein, T-1 represents the last frame of the video clip; Step 1-3-9, screen each region of interest, remove the points of interest whose micro-motion offset is less than the threshold, and obtain the micro-motion time series vector S.
8. The method for detecting deep fakes in facial videos based on facial micro-expressions according to claim 7, characterized in that: The timing prediction network described in step 2 includes: Input layer, causal convolution layer, dilated convolution layer and output layer; among them, Input layer, used to obtain the input micro-action time series vector S; The causal convolution layer uses causal convolution to process the micro-action time series vector S of the input layer and obtain the embedding vector E; The dilated convolutional layer is composed of a superposition of dilated convolutional networks. After the embedding vector E is input, the encoding vector V with extracted temporal features is obtained; The output layer outputs the encoded vector V through the relu activation function and the softmax function to obtain the final predicted value Output of the time series.
9. The method for detecting deep fakes in facial videos based on facial micro-expressions according to claim 8, characterized in that: The temporal consistency learning described in step 2 includes: Step 2-1, using the micro-motion time series obtained from the real face video as a training set, training the time series prediction network in an autoregressive manner to obtain a pre-trained time series prediction network; Step 2-2, set the trajectory coordinates of w time steps before time t0 as the initial receptive field context, and predict the time step t accordingly w The trajectory coordinates at the moment; Step 2-3, change t w The predicted value at time t is appended to the input sequence, and the current receptive field context window is moved forward by one time step, and the predicted time step is t. w+1 The trajectory coordinates at the moment; Step 2-4: perform a preset number of iterations according to step 2-3. In each iteration, the context window moves forward one time step, keeping the length of the receptive field unchanged. Step 2-5, given a series of trajectories of real face videos or forged face videos, use the pre-trained time series prediction network to predict the trajectories of a preset number of frames in the future using the methods of steps 2-2 to 2-4, that is, predict one frame at a time, and use the predicted value at the current moment for the prediction at the next moment; Step 2-6, connect the original trajectory with the predicted trajectory to obtain a new micro-motion time series.
10. The method for detecting deep fakes in facial videos based on facial micro-expressions according to claim 9, characterized in that: The authenticity detection model described in step 3 includes: Input layer, GRU network layer, attention mechanism layer and output layer; among them, The input layer is used to obtain the input multivariate time series; The GRU network layer uses a gated structure, including an update gate and a reset gate, to capture long-term dependencies in the input multivariate time series; The attention mechanism layer consists of an encoder and a decoder. The encoder generates an attention vector from the input, while the decoder generates a hidden state by taking the encoder output as input. The hidden state is divided into multiple views. The encoder uses the hidden state of the previous view decoder to assign an attention score to the hidden state of each step. The attention vector is obtained by performing a soft-max operation on the attention score. The output layer passes the attention vector through the fully connected layer to complete the classification and obtain the authenticity inspection result.
Citation Information
Cited By
Micro-expression intelligent recognition system based on face image
CN121074991A
Authenticity verification method and system based on cloned virtual image
CN121096032A
A method and system for verifying authenticity based on a cloned virtual image
CN121096032B