Deepfake Video Detection Method Based on Multimodal Contrastive Learning
By using multimodal contrast learning method in deep pseudo-video detection, the visual and audio features of video are extracted and fused, and the problems of low detection accuracy and low generalization ability in the prior art are solved, thereby achieving higher detection accuracy and anti-interference ability.
Patent Information
- Application Number
- CN202510181090.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The existing deep pseudo video detection methods fail to make full use of the multimodal information in the video, resulting in low detection accuracy, low generalization ability, and easy to overfit.
The deep pseudo-video detection method based on multimodal contrast learning is adopted to extract the visual and audio features of the video through the visual encoder and the audio encoder, and the cross-modal feature fusion module and the spatio-temporal feature extraction module are used to perform feature fusion and spatio-temporal feature extraction, and finally authenticity is judged through the classifier.
The accuracy, generalization and anti-interference ability of deep pseudo-video detection are improved, and the universality and overall performance of the detection model are enhanced.
Smart Images

Figure CN119672616B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a deepfake video detection method based on multi-modal contrast learning, which is applicable to the field of deepfake detection. Background Art
[0002] There are mainly the following types of methods for deepfake detection of videos:
[0003] Frame-based deepfake video detection method. This method first extracts the video into image frames, then uses the deepfake detection method based on images to judge authenticity, and finally adopts a certain strategy to fuse the detection results of each frame to obtain the detection result of the video. This method takes a single frame as input, ignores the temporal semantic relationship and global coherence between video frames, only focuses on low-level features such as texture details, is sensitive to interference, and performs poorly on new forgery types.
[0004] Deepfake video detection method based on temporal coherence. By adding the time dimension through RNN or LSTM, it directly detects the video. This method introduces temporal information, pays more attention to the facial movements and expressions that often show abnormalities in forged videos, and is more resistant to detection interference caused by compression or blurring.
[0005] However, both of the above two methods only use a single modality in the video, that is, the visual modality, do not consider the audio information in the video, do not make full use of the difference information between vision and audio, and the overall detection accuracy is not high.
[0006] In recent years, the joint learning of audiovisual information has begun to be studied. By training relatively independent forgery detection modules for videos and audio, and making decisions based on the correlation between the learned features. Although this method takes into account the multi-modal characteristics of the video, it does not fully integrate and interact the two modality information, and the improvement of the detection accuracy is not obvious. In addition, since the model directly trains a classifier on the deepfake video dataset, lacks prior knowledge, is prone to overfitting, achieves high accuracy on the same-distribution data, but on videos generated using new operation methods, the performance drops sharply, and the generalization ability across datasets is low. Summary of the Invention
[0007] The technical problem to be solved in the present invention is: aiming at the above existing problems, to provide a deepfake video detection method based on multi-modal contrast learning.
[0008] The technical solution adopted by the present invention is: a deepfake video detection method based on multi-modal contrast learning, characterized by including:
[0009] Input the video to be detected into a trained video forgery detection model, and output the detection result of the video to be detected;
[0010] The video forgery detection model includes:
[0011] A visual encoder, taken from a trained audio-visual contrastive learning model, is used to extract the visual features of the face sequence in the video to be detected;
[0012] An audio encoder, taken from a trained audio-visual contrastive learning model, is used to extract the audio features of the video to be detected;
[0013] A cross-modal feature fusion module is used to fuse the visual features extracted by the visual encoder and the audio features extracted by the audio encoder to obtain a feature fusion result;
[0014] A spatio-temporal feature extraction module is used to extract spatio-temporal features from the feature fusion result; the fused features not only contain visual features but also audio features, making the features more abundant;
[0015] A classifier is used to classify the video to be detected as a real video or a forged video based on the spatio-temporal features.
[0016] The training of the audio-visual contrastive learning model includes:
[0017] Using real audio-visual data as positive samples, and using the real audio-visual data with the audio shifted randomly forward or backward in time as negative samples;
[0018] Separating the audio and video in the positive and negative samples into audio and video, performing face detection and tracking on the video to extract faces, and obtaining a face sequence;
[0019] Extracting the visual features of the face sequence through the visual encoder, and extracting the audio features of the audio corresponding to the face sequence through the audio encoder;
[0020] Calculating a contrastive loss based on the visual features and audio features, reducing the distance between positive sample pairs and increasing the distance between negative sample pairs.
[0021] The cross-modal feature fusion module is used to extract features from stage3 and stage4 of the visual encoder and the audio encoder respectively for feature fusion; from stage1 to stage4, the deeper the network, the larger the receptive field and the richer the semantics of the extracted features, but the underlying detailed information may be missing; extracting features from stage3 and stage4 ensures the effective combination of high-level semantic information and low-level detailed information.
[0022] The cross-modal feature fusion module includes:
[0023] ,
[0024] ,
[0025] ,
[0026] Among them, is the fusion feature, is the audio feature output by the th stage, is the dimension of the feature, is the weight coefficient of the audio feature, is the weight coefficient of the video feature;
[0027] ,
[0028] ,
[0029] Among them, is the feature fusion result, is the weight coefficient of the fusion feature corresponding to the 3rd stage.
[0030] The visual encoder uses ResNet50 and adopts 3D convolution; the audio encoder uses ResNet18.
[0031] The training of the video forgery detection model includes:
[0032] Extract the visual encoder and the audio encoder from the trained audio-visual contrast learning model, and freeze the parameters of the visual encoder and the audio encoder.
[0033] A storage medium, on which a computer program executable by a processor is stored, and when the computer program is executed, the steps of the deepfake video detection method are implemented.
[0034] A deepfake video detection device, having a memory and a processor, and a computer program executable by the processor is stored on the memory, and when the computer program is executed, the steps of the deepfake video detection method are implemented.
[0035] The beneficial effects of the present invention are: in the present invention, the video encoder and the audio encoder are taken from the trained audio-visual contrast learning model, and audio-visual contrast learning pre-training is performed on a large number of unlabeled real speaker videos, so that the model can pre-understand the static and dynamic features of natural speaking faces, and solve the generalization problem of the detection model.
[0036] The present invention uses a cross-modal feature fusion module to cross-fuse the visual features extracted by the video encoder and the audio features extracted by the audio encoder, and finds complementary information between modalities through feature cross-fusion, so as to solve the accuracy problem of the detection model.
[0037] The present invention uses a spatio-temporal feature extraction module to extract spatio-temporal features from the feature fusion result, capture the temporal semantic relationship through the spatio-temporal features, understand the changes between frames, and solve the anti-interference problem of the detection model.
[0038] Through the visual encoder, audio encoder, cross-modal feature fusion module, spatio-temporal feature extraction module, etc. in the video forgery detection model of the present invention, the present invention has higher accuracy, generalization ability, anti-interference ability and universality, and improves the overall performance of the deepfake video detection technology. Description of the Drawings
[0039] Figure 1 It is a flow chart of the deepfake video detection method in the embodiment.
[0040] Figure 2 It is a schematic diagram of audio and video contrast self-supervised pre-training in the embodiment.
[0041] Figure 3 It is a structural block diagram of the cross-modal feature fusion module in the embodiment.
[0042] Figure 4 It is a structural block diagram of the spatio-temporal feature extraction module in the embodiment. Detailed Embodiments
[0043] Embodiment 1: As Figure 1 shown, this embodiment is a deepfake video detection method based on multi-modal contrast learning. This method uses a trained video forgery detection model to detect the video to be detected. The video forgery detection model includes a visual encoder, an audio encoder, a cross-modal feature fusion module, a spatio-temporal feature extraction module, a classifier, etc.
[0044] In this example, both the visual encoder and the audio encoder are taken from the trained audio and video contrast learning model. The visual encoder is used to extract the visual features of the face sequence in the video to be detected, and the audio encoder is used to extract the audio features of the video to be detected; the cross-modal feature fusion module is used to fuse the visual features and audio features to obtain a feature fusion result; the spatio-temporal feature extraction module is used to extract spatio-temporal features from the feature fusion result; the classifier is used to classify the video to be detected as a real video or a forged video based on the spatio-temporal features.
[0045] In this embodiment, the audio and video contrast learning model adopts audio and video contrast learning self-supervised training. As Figure 2 shown, the specific steps of this training are as follows:
[0046] 1) Construction of positive and negative sample pairs: The original audio-video pair is used as the positive sample, and the audio in the original audio-video is randomly moved forward or backward by 500 ms to 2 s as the negative sample;
[0047] 2) Audio and video preprocessing: Adjust the audio and video FPS to 25, crop it into 0.4s segments, separate the audio and video, perform face detection and tracking on the video to extract faces, obtain face sequences, with 10 frames in each sequence;
[0048] 3) Sample enhancement: Randomly mask 50% of the face sequences in the positive and negative samples, and also mask the audio at the corresponding time positions. After removing the masks, new positive and negative samples are formed
[0049] 4) Training: Preprocess the audio and video in the positive and negative samples to obtain face sequences and the audio corresponding to the time of the face sequences. Input the face sequences and audio into the visual encoder and audio encoder respectively to obtain the corresponding visual features and audio features, calculate the contrast loss between the visual features and audio features, reduce the distance between positive sample pairs, and increase the distance between negative sample pairs.
[0050] In this embodiment, the self-supervised contrast learning method is introduced into the deepfake video detection task, and some audio and video frames are randomly masked during the contrast learning. On the one hand, contrast learning does not require a dataset with manual annotations, and directly performs audio and video pre-training on a large number of unlabeled real speaker videos, mining high-level semantic information of audio and video, pre-understanding the static and dynamic features of natural faces and voices. The model is not prone to overfitting, and at the same time has higher accuracy, stronger robustness and generalization ability; on the other hand, randomly masking some audio and video frames reduces the computational amount, the model has a high capacity, stronger anti-interference ability, higher generalization ability, and overall improves the applicable conditions of the pre-trained features.
[0051] In this embodiment, the cross-modal feature fusion module is used to fuse the visual features extracted by the visual encoder and the audio features extracted by the audio encoder to obtain the feature fusion result.
[0052] In this example, the cross-modal feature fusion module extracts features from stage3 and stage4 of the visual encoder and audio encoder respectively for feature fusion, and the fusion process is as Figure 3 shown.
[0053] 1) First, multi-modal feature fusion, Figure 3 For the two fusion modules on the left in
[0054] ,
[0055] ,
[0056] where, is the audio feature output by the th stage, is the visual feature output by the th stage, is the dimension of the feature, is the weight coefficient of the audio feature, is the weight coefficient of the video feature, fusion feature As shown below:
[0057] ,
[0058] 2) Secondly, the fused multimodal features are fused again, such as Figure 3 For the rightmost fusion module, the fusion weight calculation method is as follows:
[0059] ,
[0060] ,
[0061] in, In order to obtain the final feature fusion result, is the weight coefficient of the fusion feature corresponding to the third stage.
[0062] In this embodiment, the cross-modal feature fusion module extracts features from stage 3 and stage 4 of the visual feature extractor and audio feature extractor respectively, and performs weighted fusion of features twice. The first weighted fusion is to fuse the multimodal features of stage 3 or stage 4 to mine complementary information between modalities; the second fusion is to fuse the fused multimodal features again, so that the model has both low-level texture feature extraction capabilities and high-level semantic information extraction capabilities. The fused feature information is richer, the detection model has stronger detection capabilities, stronger anti-interference, and higher generalization.
[0063] In this example, the spatiotemporal feature extraction module extracts spatiotemporal features based on the feature fusion result obtained by fusing the cross-modal feature fusion module. Figure 4 As shown on the left, the module adopts a multi-layer and multi-scale design. Each TCN-Block is composed of the right block, and the convolution kernel sizes are 3, 5, and 7 respectively.
[0064] Adversarial interference such as compression affects the robustness of spatial-level features. The forgery traces of the compressed image disappear, but the abnormal face, subtle facial movements and morphological changes in the video can still be captured from the perspective of spatiotemporal regularity and continuity. In this example, the spatiotemporal feature extraction module is based on the temporal convolutional network, and multiple temporal convolutional networks are stacked in depth and width to perform multi-dimensional and multi-scale feature extraction. The coordination of the previous and next frames of the forged video is perceived from the high-level semantic information in the spatiotemporal dimension. The detection model is more robust and less susceptible to post-processing.
[0065] In this embodiment, the spatiotemporal features extracted by the spatiotemporal feature extraction module are input into a classifier, and the classifier classifies the video to be detected as a real video or a forged video based on the spatiotemporal features.
[0066] In this embodiment, the training of the video forgery detection model includes:
[0067] 1) Construction of positive and negative sample pairs: Samples where the audio in the video is fake, the video is real, the audio is real, the video is fake, and both the audio and video are fake are used as positive samples, and samples where both the audio and video are real are used as negative samples;
[0068] 2) Preprocessing of audio and video: Adjust the audio and video FPS to 25, crop them into 0.2s segments, separate the audio and video, perform face detection and tracking on the video to extract faces, obtain a face sequence, with 5 frames in each sequence, and extract the audio corresponding to the face sequence;
[0069] 3) Construction of the network model: Extract the visual encoder and audio encoder from the trained audio and video contrastive learning model, freeze the parameters of the visual encoder and audio encoder, add a cross-modal feature fusion module, a spatio-temporal feature extraction module, and a classifier behind the visual encoder and audio encoder, and use the standard cross-entropy as the loss function;
[0070] 4) Training: The network model inputs the audio, video, and the corresponding authenticity conclusions, calculates the loss, and performs backpropagation until convergence.
[0071] The deepfake video detection method based on multi-modal contrastive learning in this embodiment specifically includes the following steps:
[0072] S100. Obtain the video to be detected;
[0073] S200. Input the video to be detected into the trained video forgery detection model, and output the detection result of the video to be detected;
[0074] S210. The video forgery detection model separates the audio and video of the video to be detected to obtain the audio and video;
[0075] S220. Perform face detection and tracking on the separated video to extract faces, obtaining a face sequence; Based on the face sequence, extract the audio corresponding to the time of the face sequence from the audio;
[0076] S230. Use the visual encoder to extract visual features from the face sequence obtained in step S220; Use the audio encoder to extract audio features from the audio extracted in step S220;
[0077] S240. Use the cross-modal feature fusion module to fuse the visual features and audio features extracted in step S230 to obtain a feature fusion result;
[0078] S250. Use the spatio-temporal feature extraction module to extract spatio-temporal features from the feature fusion result obtained in step S240;
[0079] S260. Classify the video to be detected as a real video or a forged video by using a classifier based on the spatio-temporal features extracted in step S250.
[0080] Embodiment 2: This embodiment is a storage medium on which a computer program executable by a processor is stored. When the computer program is executed, the steps of the deepfake video detection method in Embodiment 1 are implemented.
[0081] Embodiment 3: This embodiment is a deepfake video detection device having a memory and a processor. A computer program executable by the processor is stored on the memory. When the computer program is executed, the steps of the deepfake video detection method in Embodiment 1 are implemented.
Claims
1. A deep fake video detection method based on multimodal contrastive learning, characterized in that: include: Input the video to be detected into the trained video authentication model, and output the detection result of the video to be detected; The video authentication model includes: The visual encoder is taken from the trained audio and video contrast learning model and is used to extract the visual features of the face sequence in the video to be detected; The audio encoder is taken from the trained audio and video contrast learning model to extract the audio features of the video to be detected; A cross-modal feature fusion module, used to fuse the visual features extracted by the visual encoder and the audio features extracted by the audio encoder to obtain a feature fusion result; A spatiotemporal feature extraction module is used to extract spatiotemporal features from feature fusion results; A classifier is used to classify the video to be detected as a real video or a forged video based on spatiotemporal features; The training of the audio and video comparison learning model includes: The real audio and video are used as positive samples, and the audio in the real audio and video is moved forward or backward by a random time as a negative sample; Separate the audio and video in the positive and negative samples into audio and video, perform face detection and tracking on the video to extract the face, and obtain the face sequence; The visual features of the face sequence are extracted through the visual encoder, and the audio features of the audio corresponding to the face sequence are extracted through the audio encoder; Based on visual and audio features, the contrast loss is calculated to reduce the distance between positive samples and increase the distance between negative samples.
2. The deep fake video detection method based on multimodal contrastive learning according to claim 1 is characterized in that: The cross-modal feature fusion module is used to extract features from stage 3 and stage 4 of the visual encoder and audio encoder respectively for feature fusion.
3. The deep fake video detection method based on multimodal contrastive learning according to claim 2 is characterized in that: The cross-modal feature fusion module includes: , , , in, To fusion features, For the The audio features output by each stage, For the The visual features output by each stage, is the dimension of the feature, is the weight coefficient of the audio feature, is the weight coefficient of the video feature; , , in, is the feature fusion result, is the weight coefficient of the fusion feature corresponding to the third stage.
4. The deep fake video detection method based on multimodal contrastive learning according to claim 1, characterized in that: The visual encoder adopts ResNet50 and uses 3D convolution; the audio encoder adopts ResNet18.
5. The deep fake video detection method based on multimodal contrastive learning according to claim 1, characterized in that: The training of the video counterfeit detection model includes: The visual encoder and the audio encoder are extracted from the trained audio-video contrastive learning model, and the parameters of the visual encoder and the audio encoder are frozen.
6. A storage medium having stored thereon a computer program executable by a processor, characterized in that: When the computer program is executed, the steps of the deep fake video detection method according to any one of claims 1 to 5 are implemented.
7. A deep fake video detection device, comprising a memory and a processor, wherein the memory stores a computer program executable by the processor, characterized in that: When the computer program is executed, the steps of the deep fake video detection method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Face deep false detection method based on multi-model twice fusion
CN115700844A
Face deep false detection method based on multi-modal feature fusion
CN115880749A