Face forgery detection method and system based on three-dimensional gaussian, device and medium

By constructing a 3D Gaussian model and optimizing its rendering, combined with multimodal feature extraction, the problem of declining performance in existing 2D detection methods was solved, achieving efficient recognition of fake face videos.

CN122290192APending Publication Date: 2026-06-26BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2026-04-13
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing face forgery detection methods mainly focus on two-dimensional pixel and texture features. As forgery technology evolves, forgery traces in the two-dimensional domain become weaker and weaker, and the detection effect gradually declines, making it difficult to effectively identify forged videos.

Method used

A face forgery detection method based on 3D Gaussian is adopted. By obtaining the identity and pose parameter sequence of face video, an initial 3D Gaussian model is constructed by combining the FLAME model. Under preset constraints, 3D Gaussian sputtering rendering is performed to optimize the model to minimize the reconstruction residual and extract multimodal forgery features for discrimination.

Benefits of technology

It enables face forgery detection from a three-dimensional perspective, avoiding the limitations of single-dimensional judgment, improving the comprehensiveness and accuracy of detection, and effectively identifying forged videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290192A_ABST
    Figure CN122290192A_ABST
Patent Text Reader

Abstract

This application provides a method, system, device, and medium for face forgery detection based on 3D Gaussian, belonging to the field of face forgery detection technology. The method includes: acquiring a face image sequence corresponding to a face video to be detected; extracting features from the face image sequence to obtain an identity parameter sequence and a pose parameter sequence corresponding to the face image sequence; determining an initial 3D Gaussian model to characterize face features based on the identity parameter sequence and the FLAME model; under preset constraints, performing 3D Gaussian sputtering rendering based on the initial 3D Gaussian model and the pose parameter sequence, with the goal of minimizing the overall reconstruction residual between the rendered image and the face image sequence, to obtain an optimized 3D Gaussian model; extracting multimodal forgery features based on the optimized 3D Gaussian model, and determining whether the face video to be detected is a forged video based on the multimodal forgery features. This application can achieve face forgery detection in a three-dimensional dimension.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of face spoofing recognition technology, and more specifically, it relates to a face spoofing detection method, system, device, and medium based on three-dimensional Gaussian. Background Technology

[0002] With the development of technologies such as deep learning and generative adversarial networks, face spoofing technology has made great breakthroughs in the fields of images and videos.

[0003] Most existing methods for detecting face forgery primarily focus on algorithmic recognition of two-dimensional pixels, texture features, and facial expressions. These methods include, but are not limited to, binary classification based on convolutional neural networks, temporal consistency analysis, and multi-frame illumination and texture feature detection. While these methods have achieved some success, as forgery techniques evolve, fake traces in the two-dimensional domain become increasingly subtle, leading to a gradual decline in detection effectiveness. Summary of the Invention

[0004] The purpose of this application is to provide a method, system, device, and medium for detecting face forgery based on three-dimensional Gaussian, so as to achieve face forgery detection in three dimensions.

[0005] A first aspect of this application provides a method for detecting face forgery based on three-dimensional Gaussian, comprising: Obtain the face image sequence corresponding to the face video to be detected, and extract features from the face image sequence to obtain the identity parameter sequence and pose parameter sequence corresponding to the face image sequence; An initial 3D Gaussian model for characterizing facial features in face videos is determined based on the identity parameter sequence and the FLAME model. Under preset constraints, with the goal of minimizing the overall reconstruction residual between the rendered image and the face image sequence, three-dimensional Gaussian sputtering rendering is performed based on the initial three-dimensional Gaussian model and the pose parameter sequence to achieve fitting optimization of the initial three-dimensional Gaussian model and obtain the optimized three-dimensional Gaussian model. Multimodal forgery features are extracted based on the optimized 3D Gaussian model, and the detection of whether the face video is a forgery video is determined based on the multimodal forgery features.

[0006] A second aspect of this application provides a face forgery detection system based on three-dimensional Gaussian, comprising: The face image acquisition module is used to acquire the face image sequence corresponding to the face video to be detected, and to extract features from the face image sequence to obtain the identity parameter sequence and pose parameter sequence corresponding to the face image sequence. The initial Gaussian model determination module is used to determine the initial three-dimensional Gaussian model for characterizing facial features in face videos based on the identity parameter sequence and the FLAME model. The Gaussian model optimization module is used to minimize the overall reconstruction residual between the rendered image and the face image sequence under preset constraints. It performs 3D Gaussian sputtering rendering based on the initial 3D Gaussian model and the pose parameter sequence to achieve fitting optimization of the initial 3D Gaussian model and obtain the optimized 3D Gaussian model. The face forgery detection module is used to extract multimodal forgery features based on an optimized 3D Gaussian model, and to determine whether the face video to be detected is a forged video based on the multimodal forgery features.

[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described three-dimensional Gaussian-based face forgery detection method.

[0008] In a fourth aspect of this application, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described three-dimensional Gaussian-based face forgery detection method.

[0009] The beneficial effects of the face forgery detection method, system, device, and medium based on three-dimensional Gaussian provided in this application are as follows: This embodiment first extracts the identity and pose parameter sequences of the video to be detected, and constructs an initial three-dimensional Gaussian model by combining it with the FLAME model. This model is created from the perspective of the inherent geometric structure of the face, transforming the two-dimensional detection in the traditional scheme into three-dimensional detection. Secondly, under preset constraints, with the goal of minimizing the overall reconstruction residual, the model is optimized through three-dimensional Gaussian sputtering rendering. Real videos can achieve low residual fitting due to the consistency of the three-dimensional structure between frames, while fake videos produce more significant reconstruction residuals due to the contradictions in the three-dimensional structure between frames. This exposes the video to be detected as a fake video from a three-dimensional perspective. Finally, the embodiment extracts multimodal features based on the optimized three-dimensional Gaussian model. Compared with single two-dimensional features, the discrimination dimension is more comprehensive, avoiding the limitations of single-dimensional discrimination. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1A flowchart illustrating a face forgery detection method based on three-dimensional Gaussian provided in an embodiment of this application; Figure 2 This is a schematic flowchart of a method for determining a face image sequence according to an embodiment of this application; Figure 3 This is a structural block diagram of a face forgery detection system based on three-dimensional Gaussian provided in an embodiment of this application; Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0014] Please refer to Figure 1 , Figure 1 The flowchart of a face forgery detection method based on three-dimensional Gaussian provided in an embodiment of this application can be executed by an electronic device, and the method may include: S101-S104.

[0015] S101: Obtain the face image sequence corresponding to the face video to be detected, and extract features from the face image sequence to obtain the identity parameter sequence and pose parameter sequence corresponding to the face image sequence.

[0016] In this embodiment, the face image sequence refers to a set of temporally continuous and uniformly resolution RGB face images after preprocessing the face video to be detected, which can be denoted as... I= { I 1 ,I 2 ,...,I n}, where n is the number of RGB face images in the set, which is also the number of video frames in the face video to be detected. I i ∈ , This indicates the resolution of the image, and 3 indicates the number of channels.

[0017] In this embodiment, the identity parameter sequence refers to the set of parameters extracted frame by frame from the face image sequence, representing the inherent geometric features of the face, which can be denoted as... ,in, This represents the identity parameters corresponding to the face image in the nth frame.

[0018] In this embodiment, the pose parameter sequence refers to the set of parameters extracted frame by frame from the face image sequence, representing the spatial pose state of the face, denoted as . ,in, This represents the pose parameters corresponding to the nth frame of the face image, used to characterize the spatial position of the 3D face in the camera coordinate system and the viewing angle. In the subsequent construction of the 3D Gaussian model, the pose parameters can be regarded as equivalent camera pose information to characterize the changes in the viewing angle of the 3D face model relative to the imaging plane.

[0019] refer to Figure 2 In one embodiment, the sequence of face images can be determined in the following way: Let V be the video of the face to be detected. Input the video into an MTCNN or RetinaFace detector to obtain the bounding box (bbox), which is the coordinate of the face bounding box. This can be represented as [...]. x min ,y min ,x max ,y max The bounding box coordinates are used to define the rectangular region of the face in each frame. Then, based on the bounding box coordinates, the face regions are cropped and reconstructed from the original video frames to obtain the corresponding face image sequence.

[0020] S102: Determine the initial 3D Gaussian model for characterizing facial features in face videos based on the identity parameter sequence and the FLAME model.

[0021] In this embodiment, the FLAME model is a parameterized 3D face deformation model based on statistical learning, which can generate a topologically consistent 3D face triangular mesh. The initial 3D Gaussian model refers to the 3D face mesh model output by the FLAME model. More specifically, it can also refer to a 3D Gaussian point cloud model with a multi-resolution hierarchical structure constructed based on the 3D face mesh output by the FLAME model through curvature-adaptive Gaussian point sampling.

[0022] S103: Under preset constraints, with the goal of minimizing the overall reconstruction residual between the rendered image and the face image sequence, three-dimensional Gaussian sputtering rendering is performed based on the initial three-dimensional Gaussian model and the pose parameter sequence to achieve fitting optimization of the initial three-dimensional Gaussian model and obtain the optimized three-dimensional Gaussian model.

[0023] In this embodiment, the preset constraints refer to the multi-dimensional constraints pre-set to ensure that the fitting process of the 3D Gaussian model conforms to the physical laws and spatiotemporal characteristics of a real human face. These constraints may include physical constraints, geometric consistency constraints, temporal consistency constraints, and photometric consistency constraints. The rendered image refers to the face image generated by projecting it onto a 2D image plane under the camera pose corresponding to the pose parameter sequence using 3D Gaussian sputtering rendering technology, denoted as . ,in This represents the rendered image corresponding to the nth frame of the face image. The overall reconstruction residual refers to the sum of global errors between the rendered image sequence and the original face image sequence, used to characterize the degree of fit of the 3D Gaussian model to the original face image sequence. Fit optimization refers to the process of iteratively updating the initial 3D Gaussian model based on the gradient descent algorithm with the goal of minimizing the overall reconstruction residual.

[0024] In this embodiment, since the three-dimensional geometry of real face videos has natural consistency between frames, low residual fitting can be achieved through a single static three-dimensional Gaussian model; while fake videos have hidden three-dimensional structural inconsistencies, texture drift or non-physical changes between frames. Therefore, under the forced fitting in this embodiment, huge reconstruction "tension" will be generated. In the optimized three-dimensional Gaussian model, this is manifested as a significantly higher overall reconstruction residual and the presence of non-physical distortion. This embodiment achieves feature differentiation between real and fake videos at the three-dimensional level through three-dimensional modeling.

[0025] It should be noted that, in this embodiment of the application, two conditions are preset: Condition 1: A real human face video sequence can be identified by a unique three-dimensional entity. A complete explanation of projection from different perspectives: .in, For projection function, For the first Frame transformations (including rotation and translation). To observe noise, N is the total number of frames.

[0026] Condition 2: Facial features should remain constant within the same video sequence. ,in The threshold for identity change can be set based on experience.

[0027] S104: Extract multimodal forgery features based on the optimized 3D Gaussian model, and determine whether the face video to be detected is a forged video based on the multimodal forgery features.

[0028] In this embodiment, multimodal forgery features refer to a set of features extracted from the optimized 3D Gaussian model and related rendering results, which can quantify the differences between genuine and fake videos from different dimensions. In this embodiment, based on multimodal forgery features, feature information can be integrated and nonlinearly modeled through feature-level fusion and decision-level fusion to output the probability value of whether the detected face video is fake, and then a binary classification of genuine and fake videos can be achieved through a preset threshold.

[0029] As can be seen from the above, the embodiments of this application first extract the identity and pose parameter sequences of the video to be detected, and construct an initial three-dimensional Gaussian model in combination with the FLAME model. Modeling is performed from the perspective of the inherent geometric structure of the face, transforming the two-dimensional detection in the traditional scheme into three-dimensional detection. Secondly, under preset constraints, with the goal of minimizing the overall reconstruction residual, the model is optimized through three-dimensional Gaussian sputtering rendering. Real videos can achieve low residual fitting due to the consistency of the three-dimensional structure between frames, while fake videos produce more significant reconstruction residuals due to the contradiction of the three-dimensional structure between frames. This exposes the video to be detected as a fake video from the three-dimensional level. Finally, the embodiments extract multimodal features based on the optimized three-dimensional Gaussian model. Compared with single two-dimensional features, the discrimination dimension is more comprehensive, avoiding the limitations of single-dimensional discrimination.

[0030] In this embodiment, considering that the speed of facial movement changes may vary in different face videos, sampling different face videos at the same frame rate would affect the subsequent fitting results. For example, for videos with drastic facial movement changes and high motion complexity, sampling at a low frame rate would lead to the loss of key motion details such as rapid head turns, drastic expression changes, and large-angle pose deflections. This would prevent the subsequently extracted pose parameter sequence from accurately representing the actual spatial pose dynamics of the face. When performing 3D Gaussian sputtering rendering based on this incomplete pose parameter sequence, the spatiotemporal sampling density of the camera pose is insufficient, failing to fully reproduce the real motion process of the face in the video. Consequently, even if the current face video to be detected is a real video, the poor 3D Gaussian sputtering rendering result would lead to the face video being considered a fake video in the subsequent judgment process. Therefore, in one embodiment of this application, obtaining the face image sequence corresponding to the face video to be detected includes: The optical flow density, inter-frame differential energy, and head pose change rate between adjacent frames in the video of the face to be detected are calculated. A motion complexity score is obtained based on these scores. When the motion complexity score is higher than a preset threshold, a first frame rate is used for sampling; when the score is lower than the threshold, a second frame rate is used. Frame rate conversion is performed using optical flow interpolation to obtain a video frame sequence with a standardized frame rate. The first frame rate is greater than the second frame rate. Face detection is performed on the standardized frame rate video frame sequence using a multi-detector cascade, obtaining the face bounding boxes for each frame. Face alignment is performed on the detected face regions based on the face bounding boxes to obtain a face image sequence.

[0031] In this embodiment, optical flow density refers to the motion intensity characterization value, which is the arithmetic mean of the motion amplitudes of all pixels after extracting pixel-level motion vector fields from adjacent frames in the face detection video. It is used to objectively measure the overall motion degree of pixels between adjacent frames. Inter-frame difference energy refers to the average pixel value calculated by performing a difference operation on the grayscale / RGB values ​​of pixels in adjacent frames of the face detection video and taking the absolute value. It is used to quantify the degree of change in pixel information between adjacent frames. Head pose change rate refers to the first-order difference mean of the pose angle sequence calculated after extracting the head yaw angle, pitch angle, and roll angle from the sampled frames of the face detection video. The optical flow density, inter-frame difference energy, and head pose change rate are normalized and weighted to obtain a motion complexity score. The weights for the weighted calculation can be set by the user.

[0032] In this embodiment, the first frame rate and the second frame rate are sampling frame rates, which are two preset sampling frame rates to adapt to face videos with different motion complexity. The first frame rate is a high sampling frame rate, which can be set to 25~30fps, and is used to capture details such as rapid head turning and dramatic facial expression changes in high motion complexity videos. The second frame rate is a low sampling frame rate, which can be set to 10~15fps, and is used to reduce the amount of data processed for low motion complexity videos and improve overall processing efficiency.

[0033] In this embodiment, frame rate conversion is performed using optical flow interpolation technology. For videos that need to be upsampled, intermediate frames are generated using bidirectional optical flow field interpolation. For videos that need to be downsampled, weighted frame selection is performed based on motion vector time weights, thereby maintaining temporal continuity in both cases.

[0034] In one embodiment of this application, face detection employs a multi-detector cascade strategy, which calls detectors sequentially according to a preset detection accuracy order. The specific process is as follows: First, the main detector YOLO performs fast detection on the current frame, prioritizing most common scenarios due to its high inference speed. If the detection confidence score output by YOLO is lower than a preset confidence threshold (e.g., confidence score lower than 0.7), the frame is passed to the auxiliary detector RetinaFace for secondary detection. RetinaFace, through multi-scale feature extraction, has a stronger adaptability to small-sized faces and occluded scenarios, thus compensating for YOLO's shortcomings in such scenarios. For cases where, after the first two levels of detection, multiple candidate boxes or bounding box confidence scores (RetinaFace will also output corresponding confidence scores) are still lower than the preset confidence threshold, the verification detector MTCNN is further called to refine and correct the candidate regions, ultimately outputting stable face bounding box coordinates. x min ,y min ,x max ,y max The core significance of the three-level cascade lies in driving the on-demand allocation of computing resources based on detection quality requirements: the vast majority of frames only need to be detected by YOLO, while higher-precision but computationally more expensive detectors are activated only for difficult frames, thus achieving a balance between overall inference efficiency and detection robustness. This multi-detector collaborative mechanism can maintain stable detection performance under different scenarios such as changes in lighting, occlusion, and small-sized faces.

[0035] In this embodiment, after obtaining the face bounding box, adaptive face alignment and super-resolution enhancement can be performed. Specifically, the key points of the face can be detected first. In this embodiment, a 68-point key point scheme is used because the 68-point annotation system can completely cover key areas such as eyebrows, eyes, bridge of the nose, tip of the nose, lip contour, and jawline. Compared with the 5-point or 21-point scheme, it provides richer geometric constraints, which is sufficient to accurately estimate the head pose and face contour. At the same time, it has a lower computational cost than the denser 98-point or 106-point scheme, achieving a good balance between accuracy and efficiency. Therefore, 68 points is the mainstream choice in face alignment tasks. Based on the detected 68 key points, stable anchor points such as the corners of the eyes and the tip of the nose are selected to calculate the affine transformation matrix. The face is rotated and scaled to a standard frontal pose to achieve accurate alignment (the specific calculation method of the affine transformation matrix is ​​a conventional technique, and the specific method can be found by searching on Baidu). The aligned faces are then evaluated for quality, with metrics including image sharpness, brightness uniformity, and face integrity. For low-quality images with quality scores below a preset threshold, a super-resolution network is used for enhancement to restore details of blurred or low-resolution faces. Finally, all face images are standardized to 256×256 pixels, outputting a high-quality standardized face sequence, which is the face image sequence corresponding to the video of the face to be detected.

[0036] As can be seen from the above, this embodiment of the application solves the problems of subsequent fitting and judgment errors caused by fixed frame rate, single detection, and pose shift by optimizing the face image sequence acquisition process. This embodiment quantifies and scores videos with different motion complexities using optical flow density, inter-frame difference energy, and head pose change rate. High-complexity videos use a high frame rate of 25-30fps to capture key motion details, while low-complexity videos use a low frame rate of 10-15fps to improve efficiency. Optical flow interpolation is combined to ensure inter-frame continuity and avoid missing or redundant pose parameter sequences. Secondly, this embodiment allocates computing resources on demand through a multi-detector cascade strategy. YOLO quickly processes common scenes, RetinaFace adapts to small-size / occluded scenes, and MTCNN provides fine-grained correction, maintaining stable detection performance in complex scenes. 68-point alignment balances accuracy and efficiency, super-resolution enhancement repairs low-quality images, and the final output is a 256×256 normalized sequence.

[0037] In one embodiment of this application, the identity parameter sequence includes identity parameters corresponding to each frame of the face image sequence; the pose parameter sequence includes pose parameters corresponding to each frame of the face image sequence. Feature extraction is performed on the face image sequence to obtain the corresponding identity parameter sequence and pose parameter sequence, including: Perform the following feature extraction operation on each frame of the face image sequence: The three-dimensional parameter sets corresponding to the frame image are extracted using the first three-dimensional reconstruction model and the second three-dimensional reconstruction model, respectively; the consistency metric between the three-dimensional parameter sets output by the first three-dimensional reconstruction model and the second three-dimensional reconstruction model is calculated; wherein, the three-dimensional parameter sets include at least identity parameters, expression parameters, and pose parameters; the feature extraction dimension of the first three-dimensional reconstruction model is more than that of the second three-dimensional reconstruction model; When the consistency metric is higher than the preset consistency threshold, the identity parameter in the three-dimensional parameter set output by the first three-dimensional reconstruction model is used as the identity parameter corresponding to the frame image, and the pose parameter in the three-dimensional parameter set output by the first three-dimensional reconstruction model is used as the pose parameter corresponding to the frame image. When the consistency metric is lower than the preset consistency threshold, the identity parameters in the three-dimensional parameter set output by the first three-dimensional reconstruction model and the identity parameters in the three-dimensional parameter set output by the second three-dimensional reconstruction model are weighted and calculated to obtain the identity parameters corresponding to the frame image; the pose parameters in the three-dimensional parameter set output by the first three-dimensional reconstruction model and the pose parameters in the three-dimensional parameter set output by the second three-dimensional reconstruction model are weighted and calculated to obtain the pose parameters corresponding to the frame image.

[0038] In this embodiment, the first 3D reconstruction model can be an enhanced DECA (denoted as E-DECA). The original DECA model is a 3D face reconstruction model based on weakly supervised learning. Its core function is to encode a single face image into an interpretable 3D parameter representation, covering face geometry, expression, pose, lighting, and coarse-grained texture, and supporting mesh reconstruction and expression transfer based on the FLAME parameterized face model. However, the original DECA is insufficient in terms of fine-grained identity discrimination, micro-expression capture accuracy, and texture detail expression ability, making it difficult to meet the needs of high-fidelity face animation. Therefore, the following improvements are made: the identity parameter is expanded from the original 50 dimensions to 100 dimensions to capture more fine-grained individual feature differences; the expression parameter is expanded to 50 dimensions and a refined expression basis function is adopted to more accurately describe micro-expression changes; and a new 50-dimensional texture parameter is added, specifically for capturing skin texture details. With the improvements described above, Enhanced DECA can encode input face images into higher-precision 3D face parameters, thereby simultaneously reconstructing the face's geometry, expression, pose, and detailed texture, and supporting high-fidelity expression reproduction and animation-driven processing. Enhanced DECA returns six main parameters: projection (camera) parameters, identity parameters, expression parameters, pose parameters, lighting parameters, and texture parameters. Specifically, for each frame, the enhanced DECA model extracts the parameter set: Where t represents the t-th frame of the image. The projection parameters corresponding to the t-th frame image are used to describe the projection relationship from the 3D face model to the 2D image plane. They are responsible for modeling the scale and position of the face in the image to eliminate scale and depth ambiguity in single-view reconstruction. Let be the identity parameters corresponding to the t-th frame image, describing the basic individual characteristics of the face; Let be the facial expression parameters corresponding to the t-th frame image, describing the dynamic changes in facial micro-expressions; The pose parameters corresponding to the t-th frame image describe the head's movement and pose, such as head turning, elevation angle, and depression angle. Let be the lighting parameters corresponding to the t-th frame image, describing how scene lighting affects the face; Let be the texture parameters corresponding to the t-th frame image; This represents the encoder operation, which outputs six sets of parameters for each input frame.

[0039] Therefore, for a video sample, the above sets of parameters correspond to the following full sequence output: , ; ; ; , ; .in, express x One dimension, for example, the pose parameters. The standard pose parameter definition of the FLAME model is adopted, which includes a 3D axis-angle representation of global rotation and a 3D rotation representation of the neck joint, for a total of 6 dimensions. Indicates projection parameters, Indicates identity parameters, Indicates facial expression parameters. Indicates attitude parameters, Indicates lighting parameters, This represents the texture parameters.

[0040] In this embodiment, Deep3DFaceRecon is used as the second 3D reconstruction model to independently extract the second set of parameters for cross-validation. Deep3DFaceRecon is also a deep learning-based 3D face reconstruction model that has been pre-trained on a large-scale face dataset. It does not require retraining during the inference stage and directly takes standardized face images as input to output the corresponding 3D face parameters, including identity parameters, expression parameters, pose parameters, illumination parameters, and camera parameters.

[0041] In this embodiment, the parameter dimensions output by the first and second 3D reconstruction models differ because E-DECA and Deep3DFaceRecon are designed based on different parameterized face models and training strategies, each with its own independent definition of the encoding dimensions for semantic spaces such as identity and expression. To achieve cross-model parameter comparison, the two sets of parameters are first projected onto a unified common semantic space through linear mapping. Then, the Euclidean distance between the parameter vectors of corresponding dimensions is calculated, which serves as a metric for the difference in parameters between the two models. In this embodiment, the consistency threshold is not a fixed empirical value, but is determined through statistical analysis on the validation set: a large number of face samples of known quality are collected, and the parameter difference distribution of the two models on these samples is calculated. The mean of the difference distribution plus one standard deviation is used as the consistency threshold, giving the consistency threshold a data-driven rationality.

[0042] In this embodiment, when the difference between the parameters of the two models is less than the consistency threshold, it indicates that the two models have a high degree of consistency in their parameter estimations for the current frame, and the parameter extraction results are reliable. The output of the main model E-DECA is directly adopted. When the difference is greater than the consistency threshold, it indicates that there are interference factors such as occlusion, extreme pose, or abnormal lighting in the current frame, which cause the two models to diverge in their estimations. At this time, the fusion weight is dynamically calculated based on the output confidence of each model in the current frame. The model with higher confidence receives a larger weight. The two sets of parameters are weighted and fused to improve the accuracy of the final parameters.

[0043] In this embodiment, the weighted fusion is performed only on three sets of semantic parameters: identity parameters, facial expression parameters, and pose parameters. , This represents the parameter set output by the Deep3DFaceRecon model. The reason is that these three sets of parameters directly determine the geometric reconstruction of the face and the quality of facial animation, and are the input for subsequent processing. The lighting parameters and texture parameters have relatively independent effects on the rendering effect, and the two models have different expression systems for these parameters. Forcibly merging them may introduce inconsistencies. Therefore, the lighting and texture parameters are still provided separately by the main model E-DECA.

[0044] In this embodiment, the consistency metric between the three-dimensional parameter sets output by the first three-dimensional reconstruction model and the second three-dimensional reconstruction model can be calculated based on the following method: ,in Represents a consistency measure. This represents the importance weight of the p-th type of parameter in the consistency metric calculation, used to balance the contributions of the three sets of parameters—identity parameters, facial expression parameters, and posture parameters—to the overall consistency assessment. The distance function is defined as follows: and The consistency metric is calculated by projecting the Euclidean distance onto the common semantic space along the corresponding parameter dimensions. The range of this consistency metric is (0,1]. A value closer to 1 indicates a better match between the parameter estimates of the two models, while a value closer to 0 indicates a significant divergence between the two models.

[0045] In this embodiment, the obtained consistency metric can be compared with a preset consistency threshold, or with the aforementioned identity change threshold. The reason for this comparison is that identity parameters should remain highly stable within the same video sequence, and their variation range is within a small range under normal circumstances. Therefore, using the consistency level corresponding to the normal variation range of identity parameters as the threshold benchmark can effectively distinguish between reliable and unreliable parameter estimation, which has clear physical significance. The specific value is also determined through statistical analysis on the validation set, taking the lower quantile of the consistency measurement distribution of the same identity across normal frames as the threshold, rather than relying on human experience to set it.

[0046] When consistency is low ( This indicates that interference in the current frame causes significant discrepancies in the estimations of the two models, so a weighted fusion method is used: .

[0047] Among them, the fusion weight and It is dynamically calculated based on the confidence scores output by each of the two models in the current frame. Specifically, the confidence scores output by each model, after being normalized by softmax, are used as... and This allows models with higher confidence to have a larger weight in the fusion results, and satisfies... =1.

[0048] When consistency is high ( ≥ This indicates that the two models have a high degree of agreement on the parameter estimates for the current frame, and the parameter extraction results are reliable. Therefore, the output of the main model E-DECA can be directly adopted. As the final parameters of the current frame, no fusion processing is required to avoid introducing unnecessary computational overhead and potential parameter perturbations. It should be noted that the facial expression sequence ( ) is not used as direct input; its main function is to assist in interpreting dynamic changes and prevent misjudgment.

[0049] As can be seen from the above, in this embodiment, E-DECA, after dimensional expansion, can extract high-fine-grained parameters such as 100-dimensional identity and 50-dimensional facial expressions. Deep3DFaceRecon provides cross-validation, projecting the parameters onto a unified semantic space through linear mapping to calculate a consistency metric, ensuring cross-model parameter comparability. For high consistency, E-DECA parameters are directly used, balancing efficiency; for low consistency, core identity and pose parameters are dynamically weighted and fused according to confidence level, avoiding forced fusion that introduces bias. This approach leverages the fine-grained extraction advantages of E-DECA while improving parameter reliability with the help of Deep3DFaceRecon, effectively eliminating parameter bias caused by interference factors and outputting a high-precision, highly stable sequence of identity and pose parameters.

[0050] In one embodiment of this application, an initial three-dimensional Gaussian model for characterizing facial features in a face video is determined based on an identity parameter sequence and a FLAME model, including: Calculate the identity mean vector for each identity parameter in the identity parameter sequence; input the identity mean vector into the FLAME model to obtain an initial 3D face mesh composed of multiple triangular facets; sample the vertices of each triangular facet in the initial 3D face mesh based on different sampling densities to obtain an initial 3D Gaussian model used to characterize the facial features in the face video; among the different sampling densities, the largest sampling density is determined based on the mean curvature of multiple triangular facets.

[0051] In this embodiment, the identity mean vector can be represented as: The FLAME model is used to characterize the unified inherent identity geometric features of a face in a video to be detected. The initial 3D face mesh refers to the 3D face topology structure generated by the model after the identity mean vector is input into the model and decoded, which is composed of several triangular facets. The mean curvature is the quantized mean value obtained by calculating the surface curvature of a single triangular facet. It is used to characterize the geometric complexity of the facial region where the triangular facet is located. The larger the absolute value of the mean curvature, the richer the geometric details and the more dramatic the contour changes in the region (such as the eye, nose, and mouth areas); conversely, the flatter the geometric structure (such as the cheek and forehead areas).

[0052] In this embodiment, an adaptive sampling strategy is employed to set differentiated sampling densities based on the importance of the facial region. Importance is determined according to the following quantification criteria: the average curvature of each triangular facet of the FLAME mesh is calculated, and the absolute value of the curvature exceeds a preset curvature threshold. The area is identified as a high-importance area, generally corresponding to parts with rich geometric details and dramatic curvature changes, such as the eyes, mouth, and nose; the absolute value of curvature is lower than Regions are classified as low importance areas, generally corresponding to relatively flat areas such as the cheeks and forehead. High importance areas are sampled densely, while low importance areas are sampled sparsely. The ratio of sampling densities is adaptively determined based on the proportion of the mean curvature, rather than a fixed multiple, to accommodate differences in facial geometry among individuals.

[0053] Secondly, a multi-resolution Gaussian hierarchical structure is constructed, which is the initial three-dimensional Gaussian model, containing three levels: a coarse-grained layer with a high number of Gaussian points. The number of medium-grained layers is determined by uniform downsampling of FLAME mesh vertices. The number of fine-grained layers was obtained by resampling after performing a loop subdivision on the coarse-grained layer mesh. The optimization is achieved by further combining curvature adaptive refinement sampling on top of the medium-grained layer. The specific number of points in the three layers is affected by the complexity of the input mesh, typically around 500, 2000, and 8000 Gaussian points respectively. However, these values ​​are empirical references rather than fixed constraints, and the actual number of points is dynamically determined based on mesh subdivision and adaptive sampling results. A hierarchical optimization strategy is adopted, optimizing layer by layer from the coarse-grained layer to the fine-grained layer. After each layer converges, the Gaussian parameters of the current layer are interpolated and initialized for the next layer through upsampling, avoiding the instability caused by random initialization in the fine-grained layer.

[0054] As can be seen from the above, this embodiment of the application can extract the consistent inherent geometric features of faces in the video by calculating the mean vector of the identity parameter sequence, so that the initial model and the target face maintain identity matching and eliminate the modeling deviation caused by single-frame parameter perturbation. Secondly, this embodiment inputs the identity mean into the FLAME model to generate a triangular mesh, providing a geometric topological basis that conforms to the physiological structure of real faces for Gaussian point sampling. Finally, this embodiment determines the maximum sampling density based on the mean curvature of the triangular mesh, densely sampling high-detail areas such as eyes, nose, and mouth, and sparsely sampling flat areas, reducing computational redundancy while preserving key geometric details. A coarse, medium, and fine multi-resolution hierarchical structure is adopted and optimized layer by layer. The parameters of the previous layer are interpolated to initialize the next layer, avoiding optimization instability caused by random initialization.

[0055] In one embodiment of this application, 3D Gaussian sputtering rendering based on an initial 3D Gaussian model and a sequence of pose parameters refers to jointly optimizing the initial 3D Gaussian model using a static 3DGS algorithm. Specifically, the output standardized face image sequence and the corresponding camera parameter sequence are used as known inputs, and a unique time-invariant 3D Gaussian point cloud model G (the optimized 3D Gaussian model) is solved within the framework of the static 3DGS algorithm. The optimized 3D Gaussian model consists of a fixed set of 3D Gaussian primitives, each containing attributes such as its spatial location, covariance, color, and opacity, used to uniformly describe the 3D geometry and appearance structure of faces in the video.

[0056] During the optimization process, the initial 3D Gaussian model is projected and rendered back into the 2D image space under the camera pose conditions corresponding to each frame, generating a reconstructed image sequence that corresponds one-to-one with the input face frame. The optimization objective is to minimize the overall reconstruction residual between the reconstructed images and the original input images across all frames. The loss function includes both the photometric consistency loss in the pixel space and the perceptual loss in the feature space. After the two are weighted and summed, the attribute parameters of each Gaussian unit are iteratively updated through gradient descent until convergence, and finally the optimized 3D Gaussian model G is obtained.

[0057] The aforementioned fitting process can be understood as a forced fitting, the core of which is to prevent the 3D Gaussian model from introducing independent 3D structural changes in different time frames. Instead, it requires all observation frames to share the same static 3D face representation, explaining viewpoint differences only through camera pose changes. The camera pose sequence {P1,P2,...,Pn} originates from the projection parameters extracted by E-DECA, which are converted into the standard camera extrinsic parameter format required by 3DGS after coordinate system calibration. If the input video is real footage, the static 3D Gaussian model can usually explain all frames simultaneously within a reasonable error range; conversely, if the video is deepfake content, due to the potential for implicit 3D structural inconsistencies, texture drift, or non-physical changes between frames, this forced fitting process will generate significant reconstruction tension and anomalous residuals, providing a basis for subsequent extraction of forgery-sensitive features.

[0058] In this embodiment, the loss function in the above-described fitting process can include six components: Color reconstruction loss The pixel-level color difference between the rendered image and the original image is measured using the L1 norm: , where N is the total number of frames involved in the optimization.

[0059] Structural similarity loss To assess the degree of structural information preservation in an image, a complementary form of the SSIM metric is used: The reconstruction quality is comprehensively measured from three dimensions: brightness, contrast, and structure.

[0060] Deep consistency loss Ensure consistency between the reconstructed depth map and the depth map predicted by the FLAME model: ,in Depth map obtained by rendering a Gaussian model. This is a reference depth map generated by rasterizing the FLAME mesh.

[0061] Normal vector consistency loss Ensure the physical validity of the surface normal vector: ,in To render the normal vectors, Here, M is the reference normal vector for FLAME, and M is the number of sampling points.

[0062] Physical constraint loss It contains three sub-constraints. Rigid constraint. By monitoring the inter-frame invariance of the Euclidean distance between keypoint pairs, we ensure that the basic geometric structure of the face does not undergo non-physical deformation: , specifically, ∈ This represents the three-dimensional spatial coordinates of the k-th keypoint Gaussian element during the optimization process in frame t. This represents the 3D coordinates of the key point in the reference frame (the first frame of the sequence), where K is a predefined set of key points. By constraining the displacement of each keypoint from its reference position, rigidity preservation of the inter-frame geometry is achieved. Symmetry constraint. By utilizing the left-right symmetry of the human face, the corresponding Gaussian points on the left and right sides are forced to maintain a mirror-symmetric relationship: ,in, ∈ Let be the coordinate vector of the k-th keypoint Gaussian element in three-dimensional space; M Let be the mirror transformation matrix based on the sagittal plane of symmetry. ∈ To and Symmetric Gaussian points have opposite coordinates in the direction of the normal to the sagittal plane, but the same coordinates in the other two axes. By minimizing Its mirror point The Euclidean distance between them forces the 3D Gaussian point cloud to maintain the bilateral symmetry of the face, suppressing asymmetric geometric distortion. Smoothness constraint. The properties of adjacent Gaussian points should change continuously to avoid unnatural abrupt changes. Where Q is the set of adjacent Gaussian point pairs, and These are the attribute vectors corresponding to the Gaussian points. The physical constraint loss is comprehensively represented as: = + + .

[0063] Time consistency loss To ensure the stability of the Gaussian model over time, temporal smoothing constraints are applied to the rendering results of adjacent frames: .

[0064] The total loss function is obtained by weighted summation of the six loss functions: Weights The allocation logic is determined based on the contribution of each loss term to the reconstruction quality and the sensitivity to forgery detection: color loss and structure loss w1, w2 are taken as larger values ​​to ensure the basic reconstruction quality; depth and normal vector loss w3, w4 are taken as medium values ​​to maintain geometric consistency; physical constraint and temporal consistency loss w5, w6 are taken as smaller values ​​to apply soft constraints rather than forcibly restricting the model's expressive ability. The specific values ​​of each weight are determined through ablation experiments on the validation set.

[0065] An adaptive optimization strategy is employed in the above fitting process, based on the total reconstruction error of the current iteration step. The learning rate is dynamically adjusted. When the decrease in error relative to the previous iteration step is less than a preset magnitude threshold ε, it is determined that the current region is flat, and the learning rate is reduced for fine-tuning. When the decrease in error is greater than ε, it is determined that the current gradient information is sufficient, and the learning rate is maintained or appropriately increased to accelerate convergence. In the local region where Gaussian points with large gradient magnitudes are located, the density of Gaussian points in the region is adaptively increased to improve expressive power. For Gaussian points whose opacity is consistently lower than a preset opacity threshold τo, pruning operations are performed to remove redundant primitives and maintain a reasonable distribution of the Gaussian point cloud.

[0066] As can be seen from the above, the embodiments of this application employ a static 3DGS algorithm to force all frames to share the same three-dimensional Gaussian model. By explaining the differences in viewpoint through camera pose changes, real videos can achieve low residual fitting, while forged videos, due to inconsistencies in the three-dimensional structure between frames, will produce significant reconstruction tension and abnormal residuals, providing a basis for forgery detection. Secondly, the multi-dimensional loss function in this embodiment constrains the model from pixel, structural, geometric, physical, and temporal levels, avoiding non-physical deformations, ensuring the geometric rationality of the model, and further amplifying the abnormal features of forged videos. Finally, the embodiments of this application dynamically adjust the learning rate, encrypt Gaussian points in key regions, and prune redundant primitives through an adaptive optimization strategy, balancing fitting efficiency and model accuracy while reducing redundant computation.

[0067] In one embodiment of this application, the multimodal forgery features include: a reconstruction cohesion error feature characterizing the degree of physical consistency of the face video to be detected during the fitting optimization process; a geometric prior conformity feature characterizing the degree of structural similarity between the optimized 3D Gaussian model and the standard face geometric prior; and an identity parameter stability feature characterizing the degree of constancy of the identity features in the face video to be detected in the time domain; the standard face geometric prior is generated based on the FLAME model. Multimodal forgery features were extracted based on the optimized 3D Gaussian model, including: Based on the optimized 3D Gaussian model and pose parameter sequence, a reconstructed image sequence is generated, and the reconstruction cohesion error characteristics between the reconstructed image sequence and the face image sequence are calculated. The depth map sequence is obtained by rendering the optimized 3D Gaussian model. The depth map sequence is compared with the reference depth map sequence generated by the FLAME model to obtain the geometric prior conformity feature. The stability characteristics of identity parameters are obtained by calculating the change characteristics of identity parameters in the time domain based on the identity parameter sequence.

[0068] In this embodiment, since static 3DGS forces all frames to be projected onto the same 3D model, if the input is a real sequence, the optimized 3D Gaussian model can efficiently interpret each frame; otherwise, the error increases significantly. Specifically, the reconstruction cohesion error characteristics can be determined in the following way: For each frame of the input sequence I t Using the optimized 3D Gaussian model And its corresponding camera pose (pose parameter sequence) Render and generate reconstructed images The following multi-scale structural similarity (MS-SSIM) and peak signal-to-noise ratio (PSNR) are calculated at the pixel level:

[0069]

[0070] Aggregate and statistically analyze the reconstruction error characteristics of the entire video segment:

[0071] in, This means taking the average value over all frames. Take the standard deviation. Take the minimum value. The lower one. / A higher variance / minimum value usually indicates a forged image.

[0072] In this embodiment, the geometric prior conformity feature can be determined based on the following method: First, based on the optimized Gaussian model Generate depth maps in each pose. Specifically, during 3DGS rendering, a pixel-level depth representation aligned with the image space can be directly obtained by weighted accumulation of Gaussian points along each pixel direction. Since it is necessary to compare the similarity between the data-driven 3D structure obtained through forced fitting of the video and a physically reasonable prior facial geometry to determine authenticity, a standard facial geometry depth map under the same viewing perspective needs to be generated using the FLAME model. Its input is the identity parameters obtained in Phase 1. Frame-by-frame facial expression parameters and the same attitude parameters .

[0073] Based on this, for each frame, by comparison and To assess the structural differences between them, the following three geometric consistency indices are calculated: 1) Depth consistency:

[0074] To iterate through every pixel within the face region This indicates the depth map sequence obtained by rendering the optimized 3D Gaussian model. This represents the sequence of reference depth maps generated by the FLAME model.

[0075] 2) Normal vector deviation: Calculate the gradient of each pixel and its neighborhood on the depth map to obtain the surface normal vector:

[0076] in for exist The normal vector at that point, It is the inner product.

[0077] 3) Curvature Difference:

[0078] in for exist (Gaussian / mean) curvature at that point for exist The (Gaussian / mean) curvature at that point.

[0079] Based on the above features, full sequence aggregation is performed to finally construct the geometric prior conformity feature:

[0080] The larger the value, the more serious the violation of the prior knowledge of the real face by the synthesized 3D structure, and the more likely it is to be a forgery.

[0081] In this embodiment, the stability characteristics of identity parameters can be determined based on the following method: 1) Calculate the rate of change (first-order time difference) of the identity parameter sequence; if the identity vector dimension is... ,but:

[0082] It represents the amount of change in identity parameters between every two frames.

[0083] 2) Calculate the overall rate of change of identity parameters, obtained from the aggregated whole sequence:

[0084] If the changes are small, it means that the facial features of the person in the video are stable; if the changes are large, it means that the facial features of the person in the image being detected are constantly changing, and the possibility of forgery is high.

[0085] To calculate the variance contribution rate, principal component analysis (PCA) was used to obtain the variance contribution rates of the three principal components before the change in identity parameters. This reflects the main scale direction of identity drift.

[0086] Based on the above statistical characteristics, a complete identity stability vector is constructed:

[0087] If the identity parameters show high frequency and large fluctuations (such as the dispersion of high principal component contribution rates), it indicates that there is a non-physical identity switch, i.e., suspected forgery.

[0088] In one embodiment of this application, in addition to the reconstruction cohesion error feature, geometric prior conformity feature and identity parameter stability feature mentioned above, the multimodal forgery feature may also include one or more of the following: spatial domain forgery sensitive features, temporal domain forgery sensitive features and frequency domain forgery sensitive features.

[0089] In this embodiment, taking frequency domain forgery sensitive features as an example, frequency domain forgery sensitive features can be determined based on the following method: Obtain the three-dimensional parameter sequence in the foregoing embodiments as well as A multi-resolution wavelet transform was applied to each parameter sequence along the time axis, using the db4 wavelet basis and a decomposition level of 4. Low-frequency approximation coefficients and high-frequency detail coefficients were extracted, and the energy distribution within each frequency band was statistically analyzed to construct an energy distribution vector as the basic description of the frequency domain features. The spectral entropy of each parameter sequence was calculated to assess the signal complexity and randomness: in real videos, parameter changes typically exhibit a natural frequency distribution, with moderate spectral entropy, while forged videos generated by cyclic splicing or periodic editing show abnormally low spectral entropy. Based on this, anomalous frequency components were identified. The time-frequency spectrum was calculated using short-time Fourier transform to detect possible periodic forgery patterns, and discrete peak frequencies with amplitudes exceeding the background noise level in the power spectrum were recorded as anomalous frequency features.

[0090] Secondly, phase consistency analysis is performed, such as extracting the instantaneous phase from different frequency components of each parameter sequence, calculating the phase relationship across frequency bands, and evaluating the naturalness of the signal. The phase structure of real human face motion usually satisfies physical constraints, with stable phase coupling relationships between frequency components; however, in fake videos, the synthesis operation often disrupts the original phase structure, leading to disordered phase coupling relationships. Phase abrupt change points are detected, i.e., the moments when the phase difference between adjacent frames exceeds an adaptive threshold, identifying possible splicing or editing traces, and recording the location and amplitude of the abrupt change points as phase anomaly features. The temporal evolution of the phase is analyzed, and the continuity of the dynamic process is evaluated by calculating the autocorrelation function of the phase sequence.

[0091] Finally, frequency domain anomaly detection is performed. For example, spectral envelope analysis is used to fit the power spectral density curve of the parameter sequence to its envelope, identifying anomalous frequency components that deviate from the normal envelope range. Frequency points with deviations exceeding twice the standard deviation are marked as frequency domain anomalies. The spectral centroid and bandwidth are calculated. The spectral centroid is defined as the weighted average frequency of the power spectral density, and the bandwidth is defined as the root mean square width of the power spectral density. Both are used to assess the concentration of the frequency distribution. Abnormal centroid shifts or bandwidth abrupt changes can be used as forgery criteria. Spectral slope features are extracted, and the power spectral density is linearly fitted in logarithmic coordinates to obtain the high-frequency attenuation slope. The attenuation pattern of high-frequency components is analyzed: natural facial motion parameters usually exhibit a spectral slope that conforms to the 1 / f law; sequences that deviate from this law have a high degree of suspicion of forgery.

[0092] The above frequency domain analysis process ultimately outputs a frequency domain forgery sensitive feature vector. In this embodiment and subsequent embodiments, the multimodal forgery features include cohesion error features, geometric prior conformity features, identity parameter stability features, and frequency domain forgery sensitive feature vectors as examples.

[0093] As can be seen from the above, this embodiment extracts multi-dimensional forgery features such as reconstruction cohesion error, geometric prior conformity, and identity parameter stability, and expands frequency domain forgery sensitive features. This enables precise quantification of the abnormal characteristics of forged face videos from multiple dimensions, solving the problems of insufficient discrimination power of single features and difficulty in capturing deep forgery traces. Secondly, this embodiment is based on a static 3DGS single model forced fitting mechanism, which allows real videos to maintain low reconstruction errors, high conformity with FLAME geometric priors, and stable identity parameters in time. At the same time, forged videos will exhibit characteristics such as large reconstruction errors, geometric structures that violate physical priors, and abnormal fluctuations in identity parameters. Finally, this embodiment also reconstructs cohesion through MS-SSIM and PSNR statistical aggregation, constructs geometric conformity features through depth, normal vector, and curvature comparison, obtains stability features through identity change rate and PCA analysis, and captures editing traces in the frequency domain through wavelet transform, spectral entropy, and phase analysis. These multiple features complement each other, which can comprehensively explore the unnatural anomalies of forged videos in physical fitting, three-dimensional geometry, temporal changes, and frequency distribution.

[0094] In one embodiment of this application, determining whether a face video to be detected is a fake video based on multimodal forgery features includes: Multimodal forgery features are fused at the feature level to obtain a fused feature vector. The fused feature vector is then input into a pre-trained set of classifiers to obtain the forgery probability output by each classifier. Based on the dynamic characteristics of the face video to be detected and the type of each classifier, the fusion weight of each classifier is determined. Based on the forgery probability output by each classifier and the fusion weight, the final forgery probability is calculated. Based on the comparison between the final forgery probability and a preset probability threshold, it is determined whether the face video to be detected is a forged video.

[0095] In this embodiment, the multimodal forgery features can be concatenated first to obtain a concatenated vector. Secondly, the dependency relationship between each feature component in the multimodal forgery feature is calculated through a multi-head self-attention mechanism to obtain a weighted fused feature representation. Principal component analysis is performed on the weighted fused feature representation to reduce its dimensionality. The top K feature subsets with the highest correlation to the preset true and false labels are selected from the dimensionality-reduced features by the mutual information criterion to construct the fused feature vector.

[0096] In this embodiment, a feature fusion strategy driven by a multi-head self-attention mechanism can be adopted. The concatenated feature vector F is first mapped to a dimension d through a linear projection layer. model The feature space is 256, which is then input into a multi-head self-attention module with 8 heads and a feature dimension of 32 for each head. The self-attention module automatically learns by performing a dot product operation on the query matrix, key matrix, and value matrix. The dependencies between the four feature components are analyzed, and the attention weight distribution of each feature component is output. The fusion contribution of each feature component is dynamically adjusted based on the attention weights, so that the feature component with stronger discriminative power has a higher weight in the fusion result. Furthermore, a cross-attention mechanism is introduced to capture complementary information between features of different modalities, enhancing the collaborative discriminative ability across modalities.

[0097] In this embodiment, feature selection and dimensionality reduction can also be implemented. For example, principal component analysis can be applied to the feature representation after attention fusion, retaining the principal components that explain 95% of the cumulative variance, thus compressing the feature dimension to d. pca Dimension (d) pca The typical value of is 64 dimensions, which may vary slightly on different datasets, but is uniformly constrained by the 95% variance preservation criterion, and does not require manual specification of a fixed value; in actual training, d pca The feature distribution of the training set is automatically determined, typically within the range of [32, 128], requiring no manual intervention. This reduces the input dimensionality of subsequent classifiers while preserving key discriminative information. The mutual information criterion is used to evaluate the importance of the compressed feature components, calculating the mutual information value between each feature component and the true / false labels. The subset of features with the highest mutual information (Y) is selected to construct the final fused feature vector F_fused, where Y is adaptively determined based on the validation set performance. Redundant and noisy features are further removed through recursive feature elimination, eliminating the feature component with the lowest mutual information in each round, iterating until the validation set detection performance no longer improves.

[0098] In this embodiment, feature enhancement and normalization can also be performed, for example, for... The four feature components are z-score standardized to eliminate the impact of differences in feature scales on the fairness of fusion. After standardization, each component has a mean of 0 and a variance of 1. Feature enhancement techniques are used to amplify weak but important discriminative signals: the Fisher discriminant ratio of each feature component between real and fake samples in the training set is calculated, and a gain coefficient is applied to stable feature components with low discriminant ratios but small variances to enhance their effective contribution to the fused features. An adversarial training strategy is adopted, applying bounded random perturbations to the input feature vector F during training. By minimizing the KL divergence of the classification results before and after the perturbation, the robustness of the fused features to noise and interference is improved.

[0099] In this embodiment, the pre-trained classifier set refers to a model cluster composed of various structurally differentiated and functionally complementary machine learning / deep learning classification models. Each classifier is specifically optimized for different expression characteristics of multimodal features and participates in true / false probability inference together to improve the generalization and robustness of the judgment results, such as ResNet-18 (adapting to nonlinear modeling of low-dimensional features), Transformer encoder (capturing long-range dependencies), and graph neural network (modeling spatial topological relationships).

[0100] In this embodiment, a lightweight ResNet-18 can be used as the main classifier. The fused feature vector F_fused is mapped to a two-dimensional feature tensor adapted to the ResNet-18 input format. The features are nonlinearly modeled through a multi-layer residual structure, and the corresponding forgery probability P_main∈[0,1] is obtained by connecting a Sigmoid function to the output layer. A Transformer encoder is used as an auxiliary classifier. F_fused is divided into several feature segment sequences and input into a 6-layer Transformer encoder (8 attention heads per layer, 512 hidden layer dimensions). The self-attention mechanism is used to capture the long-range dependencies between feature components, and the corresponding forgery probability P_trans∈[0,1] is output at the [CLS] position. A graph neural network is used to process the structured features. The local features corresponding to each semantic region of the face are used as graph nodes, and the spatial adjacency between regions are used as graph edges. The topological relationship of each part of the face is modeled through a 3-layer graph convolutional network, and the corresponding graph-level forgery probability P_gnn∈[0,1] is output.

[0101] In this embodiment, the fusion weights of the lightweight ResNet-18, Transformer encoder, and graph neural network are dynamically adjusted according to the dynamic characteristics of the face video to be detected. The fusion weights corresponding to the lightweight ResNet-18 are denoted as w_main, the Transformer encoder as w_trans, and the graph neural network as w_gnn, satisfying w_main + w_trans + w_gnn = 1. Specifically, the temporal variation amplitude of the facial expression parameter sequence in the face video to be detected is calculated. If the variation amplitude exceeds a preset facial expression variation threshold τ_expr, the weight proportion of w_trans is increased to enhance the temporal modeling capability. The inter-frame fluctuation variance of the illumination parameters is calculated. If the variance exceeds a preset illumination variation threshold τ_light, the weight proportion of w_main is increased to strengthen the contribution of the frequency domain classifier. The above weight adjustment can be implemented through a meta-learning network. The meta-learning network takes the video statistical feature descriptor as input and outputs the optimal fusion weight vector. During the training phase, supervised optimization is performed through performance feedback on the meta-training set. The final forgery probability is: Pfake = w_main·P_main + w_trans·P_trans + w_gnn·P_gnn.

[0102] If the final probability of forgery is greater than the preset probability threshold, the face video to be detected is considered a forged video; otherwise, it is considered a real video.

[0103] In this embodiment, the confidence level of the final forgery probability can also be calibrated. For example, a temperature scaling method can be used to divide the original output logit of each classifier by the learnable temperature parameter T, adjusting the smoothness of the confidence distribution. T is optimized by minimizing the negative log-likelihood loss on the validation set. Isotonic regression is used to map the fused prediction score Pfake to the calibrated true probability, ensuring consistency between the output probability and the actual forgery ratio. Next, uncertainty estimation based on MC-Dropout is implemented. During the inference phase, Dropout is kept active and 50 forward propagation samples are performed. The mean of the sampling results is used as the final Pfake output, and the variance is used as the uncertainty measure, providing a reliability assessment reference for decision-making.

[0104] As can be seen from the above, this embodiment employs a multi-head self-attention mechanism to mine the dependencies and complementarities among multimodal forgery features. Combined with PCA dimensionality reduction, mutual information feature filtering, and standardization enhancement, it removes redundant noise while retaining core discriminative information, amplifies weak forgery feature signals, and improves the discriminative quality of fused features. Secondly, this embodiment uses ResNet-18, Transformer, and graph neural networks to construct a complementary classifier set, adapting to nonlinear modeling, long-range temporal dependencies, and spatial topological feature learning, respectively, covering multiple forgery modes. Finally, this embodiment adaptively allocates classifier fusion weights based on the dynamic characteristics of the video and achieves probability calibration through temperature scaling and isotonic regression. Combined with MC-Dropout, it completes uncertainty estimation, making the judgment results more closely resemble the real distribution, thus improving the generalization, accuracy, and decision credibility of face forgery detection.

[0105] The foregoing embodiments describe the process for determining frequency domain forgery-sensitive features. The processes for determining spatial domain forgery-sensitive features and temporal domain forgery-sensitive features can be found in the following embodiments: Sensitive features for airspace forgery can be extracted using one or more of the following three methods: (1) Extracting fine-grained reconstruction error features. The global reconstruction error directly originates from the total loss Ltotal and its sub-items output by the static Gaussian reconstruction module in the aforementioned embodiments. Based on this, further regional error analysis is performed: for example, the face is divided into multiple semantic regions such as eyes, nose, mouth, cheeks, and forehead. The local mean values ​​of pixel-level color loss Lcolor and structural similarity loss Lssim are calculated for each region to form a regional error distribution vector. At the same time, the spatial distribution pattern of the error is extracted, and the coordinates and area ratio of abnormal regions in the error set are statistically analyzed using a sliding window method as spatial domain forgery sensitive features.

[0106] (2) Geometric anomaly detection. The difference between the reconstructed depth map D_render and the standard face model reference depth map D_FLAME directly originates from the pixel-wise residual map of the depth consistency loss Ldepth. Based on this, curvature anomaly features are extracted: the second-order partial derivatives of the depth residual map are calculated to obtain the local curvature deviation distribution and identify regions that do not conform to the natural curvature distribution of the face. At the same time, continuity analysis is performed on the point-by-point cosine similarity residual of the normal vector consistency loss Lnormal to detect geometric discontinuities where the normal vector direction changes abruptly, which serve as sensitive features for spatial forgery.

[0107] (3) Extracting texture consistency features. The detailed differences between the reconstructed texture and the original texture originate from the residual of the color loss Lcolor in the high-frequency components during the stage two optimization process. Specifically, wavelet decomposition is performed on the rendered image and the original image respectively, and the difference of the high-frequency sub-band coefficients is calculated as the texture detail difference descriptor to evaluate the degree of texture mode preservation. At the same time, the gradient magnitude of the texture residual map is statistically analyzed to detect abnormal changes in sharpness at the texture boundary and identify possible synthesis traces as sensitive features for spatial forgery.

[0108] Temporal forgery-sensitive features can be extracted using one or more of the following three methods: (1) Perform inter-frame continuity analysis. The rate of change of parameters between adjacent frames is directly derived from the three-dimensional parameter sequence {pid_t, pexpr_t, ppose_t} extracted in stage one. The time derivative sequence is constructed by subtracting the parameters of adjacent frames. Based on this, the smoothness of parameter changes is analyzed: the second difference of the time derivative sequence is calculated, and possible abrupt change points are detected, that is, the frame positions where the second difference amplitude exceeds the adaptive threshold. The mean and variance of parameter changes are statistically analyzed at different time scales using the sliding window method, and multi-scale dynamic features are extracted as sensitive features for temporal forgery.

[0109] (2) Perform temporal consistency measurement. Since the temporal variation of illumination parameters originates from the illumination parameter sequence output by E-DECA, the L2 distance sequence of illumination parameters of adjacent frames can be calculated. Frames with abrupt distance changes are detected as criteria for unnatural illumination jumps and as sensitive features for temporal forgery. The temporal correlation of the pose parameter ppose_t is obtained by calculating the cosine similarity sequence of pose parameters of adjacent frames. Frame segments with a sudden drop in similarity are identified as criteria for disjointed head movements and as sensitive features for temporal forgery.

[0110] (3) Perform motion pattern analysis. The head motion trajectory consists of the projections of the rotation and translation components in the posture parameter ppose_t onto the time axis. Calculate the power spectral density of the trajectory sequence and analyze the naturalness of the motion. The temporal dynamics of facial expression changes are derived from the time derivative sequence of the facial expression parameter pexpr_t. Evaluate the rationality of facial expression transitions: calculate the KL divergence between the rate of change of facial expression parameters and the statistical prior of natural facial expression movements, using it as a feature of facial expression naturalness. Identify periodic components in the facial expression parameter sequence; if significant periodicity exists, mark it as a possible trace of cyclic forgery, serving as a sensitive feature for temporal forgery.

[0111] In one embodiment of this application, based on the judgment result of whether the face video to be detected is a forged video, corresponding detection evidence can also be generated. For example, based on the Grad-CAM method, the classification gradient can be calculated on the last convolutional feature map of the ResNet-18 main classifier to generate a heatmap showing the facial region that contributes the most to the forgery judgment. The heatmap is then superimposed on the original frame image for manual review. The curves showing the variation of each component along the time axis are labeled with the time points of abrupt changes in abnormal amplitude. If spatial and / or temporal forgery-sensitive features exist, they can be combined for verification, demonstrating the structural differences between real and forged videos in static Gaussian reconstruction and visually presenting the spatial distribution of reconstruction cohesion error. At the feature level, an attention weight matrix can also be output to characterize... The contribution of the four types of features to the final judgment is ranked. If there are spatial forgery sensitive features and / or temporal forgery sensitive features, then the spatial forgery sensitive features and / or temporal forgery sensitive features can be combined to locate abnormal facial semantic regions, mark high-confidence abnormal regions such as eyes and mouth, or determine suspicious time segments and key frames, and label the corresponding abnormal types.

[0112] Corresponding to the face forgery detection method based on 3D Gaussian in the above embodiment, Figure 3 This is a structural block diagram of a 3D Gaussian-based face forgery detection system provided in one embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 3 The face forgery detection system 20 based on three-dimensional Gaussian includes: a face image acquisition module 21, an initial Gaussian model determination module 22, a Gaussian model optimization module 23, and a face forgery detection module 24.

[0113] Among them, the face image acquisition module 21 is used to acquire the face image sequence corresponding to the face video to be detected, and to extract features from the face image sequence to obtain the identity parameter sequence and pose parameter sequence corresponding to the face image sequence. The initial Gaussian model determination module 22 is used to determine an initial three-dimensional Gaussian model for characterizing facial features in a face video based on the identity parameter sequence and the FLAME model. The Gaussian model optimization module 23 is used to perform three-dimensional Gaussian sputtering rendering based on the initial three-dimensional Gaussian model and the pose parameter sequence under preset constraints, with the goal of minimizing the overall reconstruction residual between the rendered image and the face image sequence, so as to achieve fitting optimization of the initial three-dimensional Gaussian model and obtain the optimized three-dimensional Gaussian model. The face forgery detection module 24 is used to extract multimodal forgery features based on the optimized three-dimensional Gaussian model, and to determine whether the face video to be detected is a forged video based on the multimodal forgery features.

[0114] In one embodiment of this application, the identity parameter sequence includes identity parameters corresponding to each frame of the face image sequence; the pose parameter sequence includes pose parameters corresponding to each frame of the face image sequence. The face image acquisition module 21 is specifically used to perform the following feature extraction operations for each frame of the face image sequence: The three-dimensional parameter sets corresponding to the frame image are extracted using the first three-dimensional reconstruction model and the second three-dimensional reconstruction model, respectively; the consistency metric between the three-dimensional parameter sets output by the first three-dimensional reconstruction model and the second three-dimensional reconstruction model is calculated; wherein, the three-dimensional parameter sets include at least identity parameters, expression parameters, and pose parameters; the feature extraction dimension of the first three-dimensional reconstruction model is more than that of the second three-dimensional reconstruction model; When the consistency metric is higher than the preset consistency threshold, the identity parameter in the three-dimensional parameter set output by the first three-dimensional reconstruction model is used as the identity parameter corresponding to the frame image, and the pose parameter in the three-dimensional parameter set output by the first three-dimensional reconstruction model is used as the pose parameter corresponding to the frame image. When the consistency metric is lower than the preset consistency threshold, the identity parameters in the three-dimensional parameter set output by the first three-dimensional reconstruction model and the identity parameters in the three-dimensional parameter set output by the second three-dimensional reconstruction model are weighted and calculated to obtain the identity parameters corresponding to the frame image; the pose parameters in the three-dimensional parameter set output by the first three-dimensional reconstruction model and the pose parameters in the three-dimensional parameter set output by the second three-dimensional reconstruction model are weighted and calculated to obtain the pose parameters corresponding to the frame image.

[0115] In one embodiment of this application, the initial Gaussian model determination module 22 is specifically used to calculate the identity mean vector of each identity parameter in the identity parameter sequence; Inputting the identity mean vector into the FLAME model yields an initial 3D face mesh composed of multiple triangular facets; Based on different sampling densities, the vertices of each triangular face in the initial 3D face mesh are sampled to obtain an initial 3D Gaussian model for characterizing face features in face videos. Among the different sampling densities, the largest sampling density is determined based on the mean curvature of multiple triangular facets.

[0116] In one embodiment of this application, the multimodal forgery features include: a reconstruction cohesion error feature characterizing the degree of physical consistency of the face video to be detected during the fitting optimization process; a geometric prior conformity feature characterizing the degree of structural similarity between the optimized 3D Gaussian model and the standard face geometric prior; and an identity parameter stability feature characterizing the degree of constancy of the identity features in the face video to be detected in the time domain; the standard face geometric prior is generated based on the FLAME model. The face forgery detection module 24 is specifically used to generate a reconstructed image sequence based on the optimized three-dimensional Gaussian model and pose parameter sequence, and to calculate the reconstruction cohesion error characteristics between the reconstructed image sequence and the face image sequence. The depth map sequence is obtained by rendering the optimized 3D Gaussian model. The depth map sequence is compared with the reference depth map sequence generated by the FLAME model to obtain the geometric prior conformity feature. The stability characteristics of identity parameters are obtained by calculating the change characteristics of identity parameters in the time domain based on the identity parameter sequence.

[0117] In one embodiment of this application, the face forgery detection module 24 is further configured to perform feature-level fusion of multimodal forgery features to obtain a fused feature vector; The fused feature vector is input into a pre-trained set of classifiers to obtain the forgery probability output by each classifier in the set. Based on the dynamic characteristics of the face video to be detected and the types of each classifier, the fusion weights of each classifier are determined. The final forgery probability is calculated based on the forgery probability output by each classifier and the fusion weight. Based on the comparison between the final forgery probability and the preset probability threshold, it is determined whether the face video to be detected is a forged video.

[0118] In one embodiment of this application, the face forgery detection module 24 is further configured to calculate the dependency relationship between each feature component in the multimodal forgery feature through a multi-head self-attention mechanism to obtain a weighted fused feature representation; Principal component analysis is used to reduce the dimensionality of the weighted fused feature representation. The top K feature subsets with the highest correlation to the preset true / false labels are selected from the dimensionality-reduced features using the mutual information criterion to construct the fused feature vector.

[0119] In one embodiment of this application, the face image acquisition module 21 is further used to calculate the optical flow density, inter-frame difference energy, and head pose change rate between adjacent frames in the face video to be detected. Motion complexity is scored based on optical flow density, inter-frame difference energy, and head pose change rate. When the motion complexity score is higher than the preset score threshold, the first frame rate is used for sampling; when the motion complexity score is lower than the preset score threshold, the second frame rate is used for sampling. Frame rate conversion is performed using optical flow interpolation technology to obtain a video frame sequence with a standardized frame rate; the first frame rate is greater than the second frame rate. Face detection is performed on a video frame sequence with a standardized frame rate using a multi-detector cascade to obtain the face bounding boxes for each frame. Face image sequence is obtained by aligning the detected face regions based on face bounding boxes.

[0120] See Figure 4 , Figure 4 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 4 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of each module / unit in the above system embodiments, for example... Figure 3 The functions of the face image acquisition module 21, the initial Gaussian model determination module 22, the Gaussian model optimization module 23, and the face forgery detection module 24 are shown.

[0121] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0122] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0123] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.

[0124] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation method described in the three-dimensional Gaussian-based face forgery detection method provided in the embodiments of this application, or they can execute the implementation method of the electronic device described in the embodiments of this application, which will not be repeated here.

[0125] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0126] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0127] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0128] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0129] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.

[0130] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0131] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0132] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0133] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A face forgery detection method based on three-dimensional Gaussian, characterized in that, include: Obtain the face image sequence corresponding to the face video to be detected, and extract features from the face image sequence to obtain the identity parameter sequence and pose parameter sequence corresponding to the face image sequence; Based on the identity parameter sequence and the FLAME model, an initial three-dimensional Gaussian model is determined to characterize the facial features in the face video; Under preset constraints, with the goal of minimizing the overall reconstruction residual between the rendered image and the face image sequence, three-dimensional Gaussian sputtering rendering is performed based on the initial three-dimensional Gaussian model and the pose parameter sequence to achieve fitting optimization of the initial three-dimensional Gaussian model and obtain the optimized three-dimensional Gaussian model. Multimodal forgery features are extracted based on the optimized 3D Gaussian model, and the detection of whether the face video is a forged video is determined based on the multimodal forgery features.

2. The face forgery detection method based on three-dimensional Gaussian as described in claim 1, characterized in that, The identity parameter sequence includes identity parameters corresponding to each frame of the face image sequence; the pose parameter sequence includes pose parameters corresponding to each frame of the face image sequence. The step of extracting features from the face image sequence to obtain the identity parameter sequence and pose parameter sequence corresponding to the face image sequence includes: For each frame of the face image sequence, the following feature extraction operation is performed: The three-dimensional parameter sets corresponding to the frame image are extracted using a first three-dimensional reconstruction model and a second three-dimensional reconstruction model, respectively; the consistency metric between the three-dimensional parameter sets output by the first three-dimensional reconstruction model and the second three-dimensional reconstruction model is calculated; wherein, the three-dimensional parameter sets include at least identity parameters, expression parameters, and pose parameters; the feature extraction dimension of the first three-dimensional reconstruction model is greater than that of the second three-dimensional reconstruction model; When the consistency metric is higher than the preset consistency threshold, the identity parameter in the three-dimensional parameter set output by the first three-dimensional reconstruction model is used as the identity parameter corresponding to the frame image, and the pose parameter in the three-dimensional parameter set output by the first three-dimensional reconstruction model is used as the pose parameter corresponding to the frame image. When the consistency metric is lower than the preset consistency threshold, the identity parameters in the three-dimensional parameter set output by the first three-dimensional reconstruction model and the identity parameters in the three-dimensional parameter set output by the second three-dimensional reconstruction model are weighted and calculated to obtain the identity parameters corresponding to the frame image; the pose parameters in the three-dimensional parameter set output by the first three-dimensional reconstruction model and the pose parameters in the three-dimensional parameter set output by the second three-dimensional reconstruction model are weighted and calculated to obtain the pose parameters corresponding to the frame image.

3. The face forgery detection method based on three-dimensional Gaussian as described in claim 1, characterized in that, The step of determining the initial 3D Gaussian model for characterizing facial features in the face video based on the identity parameter sequence and the FLAME model includes: Calculate the identity mean vector for each identity parameter in the identity parameter sequence; The identity mean vector is input into the FLAME model to obtain an initial three-dimensional face mesh composed of multiple triangular facets; The vertices of each triangular facet in the initial three-dimensional face mesh are sampled based on different sampling densities to obtain the initial three-dimensional Gaussian model used to characterize the facial features in the face video. Among the different sampling densities, the largest sampling density is determined based on the average curvature of the multiple triangular facets.

4. The face forgery detection method based on three-dimensional Gaussian as described in claim 1, characterized in that, The multimodal forgery features include: a reconstruction cohesion error feature characterizing the degree of physical consistency of the face video to be detected during the fitting optimization process; a geometric prior conformity feature characterizing the degree of structural similarity between the optimized 3D Gaussian model and the standard face geometric prior; and an identity parameter stability feature characterizing the degree of constancy of the identity features in the face video to be detected in the time domain; the standard face geometric prior is generated based on the FLAME model. The extraction of multimodal forgery features based on the optimized 3D Gaussian model includes: Based on the optimized 3D Gaussian model and the pose parameter sequence, a reconstructed image sequence is generated, and the reconstruction cohesion error characteristics between the reconstructed image sequence and the face image sequence are calculated. A depth map sequence is obtained by rendering based on the optimized 3D Gaussian model. The depth map sequence is then compared with a reference depth map sequence generated by the FLAME model to obtain geometric prior conformity features. Based on the identity parameter sequence, the change characteristics of the identity parameters in the time domain are calculated to obtain the stability characteristics of the identity parameters.

5. The face forgery detection method based on three-dimensional Gaussian as described in claim 1, characterized in that, The step of determining whether the face video to be detected is a fake video based on the multimodal forgery features includes: The multimodal forgery features are fused at the feature level to obtain a fused feature vector; The fused feature vector is input into a pre-trained set of classifiers to obtain the forgery probability output by each classifier in the set. Based on the dynamic characteristics of the face video to be detected and the types of each classifier, the fusion weights of each classifier are determined. The final forgery probability is calculated based on the forgery probability output by each classifier and the fusion weight. Based on the comparison between the final forgery probability and the preset probability threshold, it is determined whether the face video to be detected is a forged video.

6. The face forgery detection method based on three-dimensional Gaussian as described in claim 5, characterized in that, The step of fusing the multimodal forged features at the feature level to obtain a fused feature vector includes: The dependencies between the feature components in the multimodal forged features are calculated using a multi-head self-attention mechanism to obtain a weighted fused feature representation; Principal component analysis is performed to reduce the dimensionality of the weighted fused feature representation. The top K feature subsets with the highest correlation to the preset true / false labels are selected from the dimensionality-reduced features using the mutual information criterion to construct the fused feature vector.

7. The face forgery detection method based on three-dimensional Gaussian as described in claim 1, characterized in that, The step of obtaining the face image sequence corresponding to the face video to be detected includes: Calculate the optical flow density, inter-frame difference energy, and head pose change rate between adjacent frames in the face video to be detected; The motion complexity score is obtained based on the optical flow density, the inter-frame difference energy, and the head pose change rate. When the motion complexity score is higher than a preset score threshold, a first frame rate is used for sampling; when the motion complexity score is lower than the preset score threshold, a second frame rate is used for sampling. Frame rate conversion is performed using optical flow interpolation technology to obtain a video frame sequence with a standardized frame rate; the first frame rate is greater than the second frame rate. Face detection is performed on the video frame sequence with the standardized frame rate using a multi-detector cascade to obtain the face bounding box corresponding to each frame; The detected face regions are aligned based on the face bounding boxes to obtain the face image sequence.

8. A face forgery detection system based on three-dimensional Gaussian, characterized in that, include: The face image acquisition module is used to acquire the face image sequence corresponding to the face video to be detected, and to extract features from the face image sequence to obtain the identity parameter sequence and pose parameter sequence corresponding to the face image sequence. An initial Gaussian model determination module is used to determine an initial three-dimensional Gaussian model for characterizing facial features in the face video based on the identity parameter sequence and the FLAME model. The Gaussian model optimization module is used to perform three-dimensional Gaussian sputtering rendering based on the initial three-dimensional Gaussian model and the pose parameter sequence under preset constraints, with the goal of minimizing the overall reconstruction residual between the rendered image and the face image sequence, so as to achieve fitting optimization of the initial three-dimensional Gaussian model and obtain the optimized three-dimensional Gaussian model. The face forgery detection module is used to extract multimodal forgery features based on the optimized three-dimensional Gaussian model, and to determine whether the face video to be detected is a forged video based on the multimodal forgery features.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.