A Multimodal Digital Human Deepfake Detection Method Based on Hue Consistency Analysis

By analyzing the distribution of red hue in the HSV color space and combining audio-visual consistency, a multimodal digital human deepfake detection model is constructed, which solves the problem of insufficient robustness in existing technologies and achieves effective detection of high-quality fake videos.

CN122135263APending Publication Date: 2026-06-02UNIV OF ELECTRONICS SCI & TECH OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2026-02-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing deepfake detection methods lack robustness under supervised learning frameworks, making it difficult to effectively detect high-quality fake videos generated based on radiation fields. Furthermore, unsupervised learning methods perform poorly on fake videos with highly consistent lip synchronization. Existing datasets also lack rigorous screening, which limits detection performance.

Method used

By employing distribution analysis of red hue in the HSV color space and combining audio-visual consistency, a multimodal digital human deepfake detection model is constructed. Through speech feature extraction, visual feature extraction, red hue feature fusion, and matching detector, fake video detection is achieved under an unsupervised learning framework.

Benefits of technology

It significantly improves the accuracy and generalization ability of deepfake detection for digital humans, effectively identifies high-quality fake videos generated based on radiation fields, and enhances the robustness of the detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135263A_ABST
    Figure CN122135263A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal deepfake detection method for digital humans based on hue consistency analysis. A training dataset is obtained from real human speaking videos. A multimodal deepfake detection model based on hue consistency analysis is constructed. Speech features, visual features, and red hue features are extracted from the input video. The red hue features are fused into the visual features to obtain red hue visual features. Then, the speech features and red hue visual features are fused, and a matching score matrix is ​​generated based on the fused features. The multimodal deepfake detection model is trained using the training dataset. The trained multimodal deepfake detection model generates a matching score matrix for the video to be detected. The matching score matrices are then aggregated to obtain the overall matching score of the video to be detected, thereby achieving forgery detection. This invention can significantly improve the accuracy and generalization ability of speech and visual deepfake detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speaking digital human generation technology, and more specifically, relates to a multimodal digital human deepfake detection method based on hue consistency analysis. Background Technology

[0002] Speaker generation techniques map speech features to temporally aligned facial movements while learning interaction information from a multimodal space. As the rapid development of this technology brings increasing challenges, researchers are gradually shifting their focus from single-modal feature learning to audio-visual co-detection. Some methods directly use labels from real and fake videos to train audio-visual detection networks, while others employ audio-visual self-supervised learning for model pre-training, followed by fine-tuning using labeled data. To improve the model's generalization ability, an alternative approach is to use unsupervised learning strategies that rely solely on real data for training.

[0003] In the realm of deepfake datasets, early datasets primarily consisted of synthetic videos generated using face replacement methods. These videos often exhibited easily detectable forgery artifacts such as inconsistent boundaries. Speaker generation techniques aim to manipulate lip movements and facial expressions while maintaining the overall identity of the person, thus increasing the difficulty of deepfake detection. Most datasets consisted of forged videos generated using Generative Adversarial Network (GAN) models or videos generated using face manipulation methods. Recent datasets employ speaker generation techniques, possessing strong lip synchronization performance, posing a significant challenge to existing detection frameworks. However, these datasets lack rigorous data screening and often contain a large number of low-quality samples. Recently, researchers have introduced diffusion-based generation methods into the construction of deepfake datasets, improving the quality and diversity of synthetic heads.

[0004] Existing deepfake detection methods primarily operate within supervised learning frameworks, limiting the robustness of the detection models. Another class of unsupervised learning methods relies on the alignment of speech and video frames, or on obvious facial swapping artifacts. While these methods improve detection performance through audio-visual alignment, they still perform poorly for speaker generation models with highly consistent lip-sync. Furthermore, spoken video is mainly synthesized using Generative Adversarial Networks (GANs), diffusion models, and radiative fields. However, existing deepfake datasets primarily rely on GANs or diffusion models to construct fake data, ignoring the typical generative paradigm of radiative fields. This deficiency hinders the comprehensive evaluation of forgery detection methods. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a multimodal digital human deepfake detection method based on hue consistency analysis. By analyzing the distribution of red hue in the HSV color space, it captures the features of fake videos. At the same time, by combining audio-visual consistency, it significantly improves the accuracy and generalization ability of digital human deepfake detection.

[0006] To achieve the above-mentioned objectives, the multimodal digital human deepfake detection method based on hue consistency analysis of the present invention includes the following steps:

[0007] S1: Obtain several real videos of people speaking to form a training dataset according to actual needs;

[0008] S2: Construct a multimodal digital human deepfake detection model based on hue consistency analysis, including a speech feature extraction module, a speech mapper, a video feature extraction module, a visual mapper, a red hue extraction module, a red hue feature fusion module, a speech-visual feature fusion module, and a matching detector, wherein:

[0009] The speech feature extraction module is used to extract speech features from each audio frame of the input video. Extract speech features and send them to the speech mapper. , Indicates the number of frames;

[0010] A speech mapper is used to map speech features to a preset size to obtain speech features. And send it to the speech and visual feature fusion module. Indicates the size of a single frame feature;

[0011] The visual feature extraction module is used to extract visual features from video frames of the input video. Extract visual features and send them to the video mapper;

[0012] A visual mapper is used to map visual features to a preset size to obtain visual features. And send it to the red hue feature fusion module;

[0013] The red hue extraction module is used to extract red hue features from the input video. And send it to the red hue feature fusion module; the specific method for extracting red hue features is as follows:

[0014] Use an HSV generator to process each video frame in the input video. Convert to HSV color space to obtain HSV image For each frame of HSV image The distribution image of red hue was obtained by performing statistical analysis on the red hue. Then, the HSV features are calculated. :

[0015] ;

[0016] The red hue feature fusion module is used to fuse red hue features Fusion to visual features In the middle, the visual characteristics of red hue are obtained. And send it to the speech and visual feature fusion module;

[0017] The speech and visual feature fusion module is used to process speech features. And the visual characteristics of red hue By performing fusion, fusion characteristics are obtained. And send it to the matching detector;

[0018] The matching detector is used to determine the fusion features. Obtain the matching score matrix , of which elements Represents speech frames With video frames Match score;

[0019] S3: Use the training dataset from step S1 to train the multimodal digital human deepfake detection model to obtain the trained multimodal digital human deepfake detection model;

[0020] S4: When deepfake detection is required on a video of a person speaking, the video of the person speaking is sampled to a frame rate of [missing information]. The video is then input into the multimodal digital human deep forgery detection model trained in step S3 to obtain a matching score matrix. The overall matching score of the video to be detected is then obtained by aggregating the matching score matrix. When the overall matching score If the value exceeds a preset threshold, the video to be detected is a real video; otherwise, it is a deepfake video.

[0021] This invention presents a multimodal digital human deepfake detection method based on hue consistency analysis. It obtains a training dataset consisting of real human speaking videos; constructs a multimodal digital human deepfake detection model based on hue consistency analysis; extracts speech features, visual features, and red hue features from the input video; fuses the red hue features into the visual features to obtain red hue visual features; then fuses the speech features and red hue visual features to generate a matching score matrix based on the fused features; trains the multimodal digital human deepfake detection model using the training dataset; uses the trained multimodal digital human deepfake detection model to generate a matching score matrix for the video to be detected; and then aggregates the matching score matrices to obtain the overall matching score of the video to be detected, thereby achieving forgery detection.

[0022] The present invention has the following beneficial effects:

[0023] 1) This invention adopts an unsupervised learning framework and utilizes the red hue in the HSV color space features to achieve general deep forgery detection of speaking digital human videos;

[0024] 2) This invention captures the features of fake videos by analyzing the distribution of red hue in the HSV color space. By using the consistency analysis of visual and audio features based on red hue features, it significantly improves the accuracy and generalization ability of deep fake detection of speaking digital human videos. Attached Figure Description

[0025] Figure 1 This is a flowchart illustrating a specific implementation method of the multimodal digital human deepfake detection method based on hue consistency analysis according to the present invention.

[0026] Figure 2 This is a structural diagram of the multimodal digital human deepfake detection model in this invention;

[0027] Figure 3 This is a structural diagram of the weighting module in this embodiment;

[0028] Figure 4 This is a flowchart of the fake video generation method in this embodiment. Detailed Implementation

[0029] The specific embodiments of the present invention will now be described with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be particularly noted that in the following description, detailed descriptions of known functions and designs that might obscure the main content of the invention will be omitted here.

[0030] Example

[0031] Figure 1 This is a flowchart illustrating a specific implementation of the multimodal digital human deepfake detection method based on hue consistency analysis according to the present invention. Figure 1 As shown, the multimodal digital human deepfake detection method based on hue consistency analysis of the present invention includes the following steps:

[0032] S101: Obtain the training dataset:

[0033] To form a training dataset, several real-life videos of people speaking are collected based on actual needs.

[0034] S102: Constructing a multimodal digital human deepfake detection model:

[0035] Research has revealed that forged speaker videos often differ from genuine videos in color distribution, exhibiting blurred faces or overly perfect colors, particularly with excessively high intensity in the red channel. Therefore, the consistency between visual features based on red hue in the speaker video and the audio can be used for deepfake detection. Based on this, this invention constructs a multimodal digital human deepfake detection model based on hue consistency analysis. Figure 2 This is a structural diagram of the multimodal digital human deepfake detection model in this invention. (See diagram below.) Figure 2 As shown, the multimodal digital human deepfake detection model of this invention includes a speech feature extraction module, a speech mapper, a video feature extraction module, a visual mapper, a red hue extraction module, a red hue feature fusion module, a speech and visual feature fusion module, and a matching detector. Each module will be described in detail below.

[0036] The speech feature extraction module is used to extract speech features from each audio frame of the input video. Extract speech features and send them to the speech mapper. , Indicates the number of frames.

[0037] The specific structure of the speech feature extraction module can be set according to actual needs. In this embodiment, the pronunciation feature extraction module adopts the speech feature extractor in the AV-HuBERT (Audio-Visual Hidden Unit BERT) model.

[0038] A speech mapper is used to map speech features to a preset size to obtain speech features. And send it to the speech and visual feature fusion module. This indicates the size of a single frame feature.

[0039] The visual feature extraction module is used to extract visual features from video frames of the input video. Visual features are extracted and sent to the video mapper. Similarly, in this embodiment, the visual feature extraction module uses the visual feature extractor in the AV-HuBERT model.

[0040] A visual mapper is used to map visual features to a preset size to obtain speech features. And send it to the red hue feature fusion module.

[0041] The red hue extraction module is used to extract red hue features from the input video. And send it to the red hue feature fusion module. The specific method for extracting red hue features is as follows:

[0042] Use an HSV generator to process each video frame in the input video. Convert to HSV color space to obtain HSV image For each frame of HSV image The distribution image of red hue was obtained by performing statistical analysis on the red hue. Then, the HSV features are calculated. :

[0043] .

[0044] The red hue feature fusion module is used to fuse red hue features Fusion to visual features In the middle, the visual characteristics of red hue are obtained. And send it to the speech and visual feature fusion module.

[0045] To improve the quality of red hue visual features, this embodiment proposes a red hue feature fusion module based on adaptive weights. For example... Figure 2 As shown, in this embodiment, the red hue feature fusion module includes a red hue space feature extraction module, a visual space feature extraction module, a global fusion module, a local fusion module, an adaptive spatial weighting module, and a spatial feature fusion module, wherein:

[0046] The red hue space feature extraction module is used to extract red hue features. Extract global features of red hue separately and local features of red hue The data is then sent to the global fusion module and the local fusion module, respectively. Spatial feature extraction avoids overfitting to surface color features, ensuring that the final red hue visual features more accurately represent the video features. In this embodiment, the red hue spatial feature extraction module uses the DinoV2 model.

[0047] The visual spatial feature extraction module is used to extract visual features Extracting global visual features separately and visual local features Then, they are sent to the global fusion module and the local fusion module respectively.

[0048] The global fusion module is used to integrate global features of the red hue. and visual global features of each frame Element-wise multiplication is performed to obtain the fused global features for each frame. , Represent the Hada code product, and then fuse global features. Send to the adaptive spatial weight module and the spatial feature fusion module.

[0049] The local fusion module is used to integrate local features of the red hue. and visual local features of each frame Element-wise multiplication is performed to obtain the fused local features of each frame. Then, the local features are fused. Send to the adaptive spatial weight module and the spatial feature fusion module.

[0050] This embodiment employs a visual region attention mechanism, which uses global and local features of red hue as attention to process visual features, thereby modeling the influence of red hue features on different regions of visual features and improving the representational ability of red hue features.

[0051] The adaptive spatial weighting module is used to fuse global features. and fusion of local features Perform concatenation and generate global weights based on the concatenation features. and local weights And send it to the spatial feature fusion module.

[0052] The spatial feature fusion module is used to employ global weights and local weights For fusion of global features and fusion of local features Weighted fusion is performed to obtain the visual features of the red hue. The visual features of red hue in each frame The following formula is used for calculation:

[0053] .

[0054] The speech and visual feature fusion module is used to process speech features. And the visual characteristics of red hue By performing fusion, fusion characteristics are obtained. And send it to the matching detector.

[0055] Similarly, in this embodiment, the speech and visual feature fusion module is also implemented based on adaptive weights. For example... Figure 2 As shown, in this embodiment, the speech-visual feature fusion module includes an adaptive speech-visual weighting module and a multimodal feature fusion module, wherein:

[0056] The adaptive speech visual weighting module is used to combine speech features And the visual characteristics of red hue The data is concatenated, and speech weights are generated based on the concatenation features. and visual weight And send it to the multimodal feature fusion module.

[0057] The multimodal feature fusion module is used to employ speech weights. and visual weight speech features And the visual characteristics of red hue Weighted fusion is performed to obtain fusion features. Each frame's fused features The following formula is used for calculation:

[0058] ,

[0059] in, , Representing speech features respectively And the visual characteristics of red hue Features of each frame.

[0060] In this embodiment, the adaptive spatial weight module and the adaptive speech-visual weight module adopt the same structure. Figure 3 This is a structural diagram of the weighting module in this embodiment. For example... Figure 3 As shown, in this embodiment, the adaptive spatial weighting module and the self-speech visual weighting module include a feature concatenation layer, a linear module, a ReLU function layer, and a Softmax function layer, wherein:

[0061] The feature concatenation layer is used to concatenate two input features and send the concatenated features to the linear module.

[0062] The linear module performs a linear mapping on the concatenated features and sends the resulting features to the ReLU function layer. In this embodiment, the linear module includes four cascaded linear layers.

[0063] The ReLU function layer is used to process the received features using the ReLU activation function, and then sends the resulting features to the Softmax function layer.

[0064] The Softmax function layer is used to process the received features using the Softmax function, generating two normalized weights.

[0065] The matching detector is used to determine the fusion features. Obtain the matching score matrix , of which elements Represents speech frames With video frames The matching score is used as the matching detector in this embodiment.

[0066] As described above, when the input video of a person speaking is a real video, the visual features based on the red hue will show better consistency with the audio frame, and its matching score will be high. However, in fake videos, due to the high intensity of the red hue features, its consistency with the audio frame will be worse, and its matching score will be low.

[0067] S103: Training a multimodal digital human deepfake detection model:

[0068] The multimodal digital human deepfake detection model is trained using the training dataset from step S1 to obtain the trained multimodal digital human deepfake detection model.

[0069] In this embodiment, the loss function for training the multimodal digital human deepfake detection model is... The contrastive loss function (InfoNCE) is used, and the calculation formula is as follows:

[0070] ,

[0071] in, Indicates the first The temporal neighborhood of a frame.

[0072] By minimizing the above loss function and training the detection network, the ability to model the temporal consistency between audio and video frames can be effectively improved.

[0073] The termination condition for training the multimodal digital human deepfake detection model can be set according to actual needs. In this embodiment, a validation set is used to validate the current multimodal digital human deepfake detection model. The validation set includes real videos and deepfake videos. Then, the deepfake detection accuracy is calculated. When the deepfake detection accuracy reaches a preset threshold, the training ends.

[0074] As a crucial component in the generation of speaking digital human videos, the radiation field has not yet been included in existing benchmark datasets, thus limiting the comprehensiveness of current deepfake detection evaluations. To address this limitation and better test the performance of multimodal digital human deepfake detection models, this embodiment proposes a novel radiation field-based digital human deepfake video generation method. This method is built upon the neural radiation field method SyncTalk and the three-dimensional Gaussian sputtering method GaussianTalker, overcoming the limitation of current digital human deepfake identification datasets that rely solely on generative adversarial networks and diffusion models, ensuring that the generated deepfake videos are both cutting-edge and challenging. Figure 4 This is a flowchart of the deepfake video generation method in this embodiment. For example... Figure 4 As shown, the specific steps for generating deepfake videos in this embodiment include:

[0075] S401: Obtain real-life video of a person speaking.

[0076] Obtain according to actual needs Several real-life speaking videos of a particular individual.

[0077] S402: Obtain the driver audio signal:

[0078] Several audio tracks are randomly selected from a pre-defined dataset of fake voices as the driving audio signals for the synthesized character, and these tracks are divided into two sets of driving audio signals. and .

[0079] S403: Training a speech model:

[0080] Train using SyncTalk models respectively A character's speaking model Then, the GaussianTalker model was trained separately. A character's speaking model , .

[0081] SyncTalk is a NeRF-based speech-driven high-synchronization speaker avatar synthesis model proposed at CVPR 2024. It solves the synchronization problems of traditional methods, such as lip movement-speech mismatch, unnatural facial expressions, and head posture jitter. Through the collaborative optimization of three modules, it achieves high-resolution video generation with consistent identity, accurate lip movement, realistic facial expressions, and stable posture.

[0082] GaussianTalker, proposed in 2024, is an audio-driven real-time high-fidelity speaker avatar synthesis model based on 3D Gaussian sputtering (3DGS). Its core advantage lies in combining the fast rendering capability of 3DGS with the precise facial deformation driven by audio, achieving lip-to-speech synchronization, controllable posture, and a rendering speed of up to 120+ FPS. It solves the problems of slow rendering of traditional NeRF and loss of identity in GAN, and is suitable for scenarios such as real-time virtual humans and digital content creation.

[0083] S404: Generate deepfake videos:

[0084] Drive audio signal set The audio is used as the driving signal and input into each character's speaking model. Generate a deepfake video and add it to the verification set. Then, the driving audio signal set is... The audio is used as the driving signal and input into each character's speaking model. Generate a deepfake video and add it to the verification set.

[0085] S104: Deepfake Detection

[0086] When performing deepfake detection on a video of a person speaking, the video is sampled at a frame rate of [number missing]. The video is then input into the multimodal digital human deep forgery detection model trained in step S3 to obtain a matching score matrix. The overall matching score of the video to be detected is then obtained by aggregating the matching score matrix. When the overall matching score If the score exceeds a preset threshold, the video to be detected is a real video; otherwise, it is a deepfake video. In this embodiment, the matching score matrix aggregation is implemented using the Softmax function.

[0087] To better illustrate the technical effects of the present invention, specific examples are used to experimentally verify the present invention.

[0088] This embodiment trains the multimodal digital human deepfake detection model of this invention using real-person speaking videos, and then conducts experiments on four datasets: the publicly available AVLips dataset, the FakeAVCeleb dataset, the TalkingHeadBench dataset, and the RFAV dataset consisting of 6000 deepfake videos generated using the fake video generation method in this embodiment. For FakeAVCeleb (FKAV), a test set containing 500 real samples and 1000 carefully selected fake samples was constructed. For TalkingHeadBench (THB), the target test set was formed by merging the official test sets of all videos generated by the diffusion model. In the RFAV dataset, 2000 videos containing different identities and fake methods were used as the test set. Since THB and RFAV lack real samples, this embodiment supplements their test sets with real samples from the AVLips test set. In the cross-fake type generalization test, this embodiment divides the test set according to the fake types defined in FKAV, and then uses the area under the curve (AUC) and mean precision (AP) as evaluation metrics.

[0089] In this embodiment, five existing methods are used as comparison methods to compare and verify the present invention (denoted as RHTHFD), namely:

[0090] CViT: See the literature "Wodajo, D.; and Atnafu, S. 2021. Deepfake videodetection using convolutional vision transformer. arXiv preprint arXiv:2102.11126."

[0091] EfficientViT: See the literature "Coccomini, D. A.; Messina, N.; Gennaro, C.; and Falchi, F. 2022. Combining efficientnet and vision transformers for video deepfake detection. In International conference on image analysis and processing, 219–229. Springer."

[0092] LipFD: See the literature "Liu, W.; She, T.; Liu, J.; Li, B.; Yao, D.; and Wang, R. 2024. Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes. Advances in Neural Information Processing Systems, 37: 91131–91155."

[0093] AVAD: See the literature "Feng, C.; Chen, Z.; and Owens, A. 2023. Self-supervised video forensics by audio-visual anomaly detection. In proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 10491–10503."

[0094] AVH-Align: See the literature "Smeu, S.; Boldisor, D.-A.; Oneata, D.; and Oneata,E. 2025. Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning. In Proceedings of the Computer Vision and PatternRecognition Conference, 18815–18825."

[0095] AVBYOL: See "Grill, J.-B.; Strub, F.; Altche, F.; Tallec, C.;Richemond, P.; Buchatskaya, E.; Doersch, C.; Avila Pires, B.; Guo, Z.;Gheshlaghi Azar, M.; et al. 2020. Bootstrap your own latent a new approach to self-supervised learning. Advances in neural information processing systems,33: 21271–21284.”

[0096] VQ-GAN: See the literature "Afouras, T.; Chung, JS; Senior, A.; Vinyals, O.; and Zisserman, A. 2018. Deep audio-visual speech recognition. IEEE transactions on pattern analysis and machine intelligence, 44(12): 8717–8727."

[0097] Table 1 is a comparison table of the deepfake detection performance of the present invention and the comparative method in this embodiment.

[0098]

[0099] Table 1

[0100]

[0101] Table 2

[0102] In Tables 1 and 2, "Modality" represents the modalities involved in the method, and "Training Videos" represents the amount of training data included. As shown in Tables 1 and 2, the present invention outperforms existing methods on all test sets under different experimental settings.

[0103] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.

Claims

1. A multimodal digital human deepfake detection method based on hue consistency analysis, characterized in that, Includes the following steps: S1: Obtain several real videos of people speaking to form a training dataset according to actual needs; S2: Construct a multimodal deepfake detection model for digital humans based on hue consistency analysis, including a speech feature extraction module, a speech mapper, a video feature extraction module, a visual mapper, a red hue extraction module, a red hue feature fusion module, a speech-visual feature fusion module, and a matching detector. in: The speech feature extraction module is used to extract speech features from each audio frame of the input video. Extract speech features and send them to the speech mapper. , Indicates the number of frames; A speech mapper is used to map speech features to a preset size to obtain speech features. And send it to the speech and visual feature fusion module. Indicates the size of a single frame feature; The visual feature extraction module is used to extract visual features from video frames of the input video. Extract visual features and send them to the video mapper; A visual mapper is used to map visual features to a preset size to obtain visual features. And send it to the red hue feature fusion module; The red hue extraction module is used to extract red hue features from the input video. And send it to the red hue feature fusion module; the specific method for extracting red hue features is as follows: Use an HSV generator to process each video frame in the input video. Convert to HSV color space to obtain HSV image For each frame of HSV image The distribution image of red hue was obtained by performing statistical analysis on the red hue. Then, the HSV features are calculated. : ; The red hue feature fusion module is used to fuse red hue features Fusion to visual features In the middle, the visual characteristics of red hue are obtained. And send it to the speech and visual feature fusion module; The speech and visual feature fusion module is used to process speech features. And the visual characteristics of red hue By performing fusion, fusion characteristics are obtained. And send it to the matching detector; The matching detector is used to determine the fusion features. Obtain the matching score matrix , of which elements Represents speech frames With video frames Match score; S3: Use the training dataset from step S1 to train the multimodal digital human deepfake detection model to obtain the trained multimodal digital human deepfake detection model; S4: When deepfake detection is required on a video of a person speaking, the video of the person speaking is sampled to a frame rate of [missing information]. The video is then input into the multimodal digital human deep forgery detection model trained in step S3 to obtain a matching score matrix. The overall matching score of the video to be detected is then obtained by aggregating the matching score matrix. When the overall matching score If the value exceeds a preset threshold, the video to be detected is a real video; otherwise, it is a deepfake video.

2. The multimodal digital human deepfake detection method according to claim 1, characterized in that, The pronunciation feature extraction module uses the speech feature extractor in the AV-HuBERT model, and the visual feature extraction module uses the visual feature extractor in the AV-HuBERT model.

3. The multimodal digital human deepfake detection method according to claim 1, characterized in that, The red hue feature fusion module includes a red hue spatial feature extraction module, a visual spatial feature extraction module, a global fusion module, a local fusion module, an adaptive spatial weighting module, and a spatial feature fusion module, wherein: The red hue space feature extraction module is used to extract red hue features. Extract global features of red hue separately and local features of red hue Then, they are sent to the global fusion module and the local fusion module respectively; The visual spatial feature extraction module is used to extract visual features Extracting global visual features separately and visual local features Then, they are sent to the global fusion module and the local fusion module respectively; The global fusion module is used to integrate global features of the red hue. and visual global features of each frame Element-wise multiplication is performed to obtain the fused global features for each frame. , Represent the Hada code product, and then fuse global features. Send to the adaptive spatial weight module and the spatial feature fusion module; The local fusion module is used to integrate local features of the red hue. and visual local features of each frame Element-wise multiplication is performed to obtain the fused local features of each frame. Then, the local features are fused. Send to the adaptive spatial weight module and the spatial feature fusion module; The adaptive spatial weighting module is used to fuse global features. and fusion of local features Perform concatenation and generate global weights based on the concatenation features. and local weights And send it to the spatial feature fusion module; The spatial feature fusion module is used to employ global weights and local weights For fusion of global features and fusion of local features Weighted fusion is performed to obtain the visual features of the red hue. The visual features of red hue in each frame The following formula is used for calculation: 。 4. The multimodal digital human deepfake detection method according to claim 3, characterized in that, The red hue space feature extraction module uses the DinoV2 model.

5. The multimodal digital human deepfake detection method according to claim 3, characterized in that, The adaptive spatial weighting module includes a feature stacking layer, a linear module, a ReLU function layer, and a Softmax function layer, wherein: The feature overlay layer is used to concatenate two input features and send the concatenated features to the linear module; The linear module is used to perform a linear mapping on the concatenated features and send the resulting features to the ReLU function layer; The ReLU function layer is used to process the received features using the ReLU activation function and then send the resulting features to the Softmax function layer. The Softmax function layer is used to process the received features using the Softmax function, generating two normalized weights.

6. The multimodal digital human deepfake detection method according to claim 1, characterized in that, The speech-visual feature fusion module includes an adaptive speech-visual weighting module and a multimodal feature fusion module, wherein: The adaptive speech visual weighting module is used to combine speech features And the visual characteristics of red hue The data is concatenated, and speech weights are generated based on the concatenation features. and visual weight And send it to the multimodal feature fusion module; The multimodal feature fusion module is used to employ speech weights. and visual weight speech features And the visual characteristics of red hue Weighted fusion is performed to obtain fusion features. Each frame's fused features The following formula is used for calculation: , in, , Representing speech features respectively And the visual characteristics of red hue Features of each frame.

7. The multimodal digital human deepfake detection method according to claim 6, characterized in that, The adaptive spatial weighting module includes a feature concatenation layer, a linear module, a ReLU function layer, and an over-Softmax function layer, wherein: The feature concatenation layer is used to concatenate two input features and send the concatenated features to the linear module; The linear module is used to perform a linear mapping on the concatenated features and send the resulting features to the ReLU function layer; The ReLU function layer is used to process the received features using the ReLU activation function and then send the resulting features to the Softmax function layer. The Softmax function layer is used to process the received features using the Softmax function, generating two normalized weights.

8. The multimodal digital human deepfake detection method according to claim 1, characterized in that, The loss function used in training the multimodal digital human deepfake detection model is... The contrastive loss function is used, and the calculation formula is as follows: , in, Indicates the first The temporal neighborhood of a frame.

9. The multimodal digital human deepfake detection method according to claim 1, characterized in that, The training of the visual forgery detection model ends when the validation set, which includes real and fake videos, is used to validate the current visual forgery detection model. The validation set includes real videos and fake videos. Then, the forgery detection accuracy is calculated. When the forgery detection accuracy reaches a preset threshold, the training ends.

10. The multimodal digital human deepfake detection method according to claim 9, characterized in that, The method for generating the fake videos in the verification set is as follows: 1) Obtain according to actual needs Several real speaking videos of a single person; 2) Randomly select several audio tracks from the pre-set fake speech dataset as the driving audio signals for the synthesized character, and divide them into two driving audio signal sets. and ; 3) Train using SyncTalk models respectively A character's speaking model Then, the GaussianTalker model was trained separately. A character's speaking model , ; 4) Set the driving audio signal The audio is used as the driving signal and input into each character's speaking model. Generate a fake video and add it to the verification set; then, drive the audio signal set. The audio is used as the driving signal and input into each character's speaking model. Generate a fake video and add it to the verification set.