Prompt learning-based multi-mode deep counterfeit video detection device and method
Through the multimodal deep fake video detection method based on prompt learning, combined with learnable visual and audio prompts and basic models, visual and audio deep fake features are extracted and predicted, and the multimodal features are aligned by convolutional neural networks, the problem that existing methods are difficult to learn fine-grained audio and video consistency features under limited data is solved, and high-accuracy multimodal deep fake video detection is achieved.
Patent Information
- Application Number
- CN202510014022.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The existing multimodal depth forged video detection methods are difficult to learn fine-grained audio-visual consistency characteristics when data is limited, and the coarse-grained constraints cannot effectively detect frame-level consistency.
A multimodal deep fake video detection method based on prompt learning is adopted, and visual and audio deep pseudo-features are extracted through learningable visual and audio prompts combined with basic models, and prediction is performed using multi-layer perceptron blocks. A convolutional neural network is introduced to align the time dimensions, fuse multimodal features for video authenticity prediction, and at the same time, a frame-level cross-modal feature matching loss function is introduced to learn fine-grained audio-visual consistency features.
Effectively detecting multimodal deep fake videos improves the accuracy and generalization of detection, can learn audio and video consistency characteristics at the fine-grained level, and output the authenticity of audio and visual modalities at the same time.
Smart Images

Figure CN120047865A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video forensics, and particularly relates to a multi-modal deepfake video detection device and method based on prompt learning. Background Art
[0002] A deepfake video is a realistic video forged by using video or audio generation technology based on deep learning to synthesize or tamper with visual and audio content in the video. Usually, this technology is used to replace the face or voice of a person in the video, making them appear to have performed or said behaviors and remarks that they have never done or said.
[0003] Currently, researchers have proposed a series of forgery video detection methods based on multi-modal information. Since there is a natural consistency between lip movement and pronunciation in real data, most multi-modal methods utilize audiovisual consistency for forgery detection, including detecting phoneme-viseme consistency for specific pronunciations, capturing audiovisual consistency using contrastive learning, designing a new multi-modal network architecture, introducing representation learning to learn audiovisual consistency features, etc. However, these methods need to train all parameters of the network on a limited dataset, and many methods only utilize audiovisual consistency but ignore the artifacts within the forgery modality. In addition, existing methods usually calculate the contrastive loss at the segment level, and this coarse-grained constraint cannot learn fine-grained frame-level consistency features. Summary of the Invention
[0004] To address the deficiencies of the above-mentioned existing technologies, the present invention proposes a multi-modal deepfake video detection device and method based on prompt learning. First, a learnable visual prompt is used in combination with a visual base model to model visual signals, extract visual deepfake features, and use a multi-layer perceptron block to predict the authenticity of the visual modality. Then, a learnable audio prompt is introduced in combination with a speech recognition base model to model audio signals, extract audio deepfake features, and combine with a multi-layer perceptron block to predict the authenticity of the audio modality. Finally, a convolutional neural network is used to align visual features and audio features in the time dimension, and the aligned multi-modal features are fused to predict the authenticity of the video. During the network training process, in addition to the cross-entropy loss function, a new frame-level cross-modal feature matching loss is introduced to learn fine-grained audiovisual consistency features.
[0005] The specific technical solution of the present invention is as follows:
[0006] A multi-modal deepfake video detection device based on prompt learning, characterized by comprising: a visual and audio data preprocessing module, a visual deepfake feature extraction and prediction module, an audio deepfake feature extraction and prediction module, a multi-modal feature alignment module, a cross-modal feature matching module, and a video prediction module, wherein,
[0007] The visual and audio data preprocessing module is used to split the input video data into small segments and extract the corresponding visual content and audio signals of the small segments;
[0008] The visual deepfake feature extraction and prediction module is used to extract visual deepfake features from the visual content extracted by the visual and audio content preprocessing module and predict the authenticity of the visual modality by using a visual base model and learnable visual cues;
[0009] The audio deepfake feature extraction and prediction module converts the audio sampling points extracted by the visual and audio content preprocessing module into spectrograms, and then extracts audio deepfake features and predicts the authenticity of the audio modality by using a speech recognition base model and learnable audio cues;
[0010] The multi-modal feature alignment module is used to align the visual features extracted by the visual deepfake feature extraction and prediction module and the audio features extracted by the audio deepfake feature extraction and prediction module in the time dimension;
[0011] The cross-modal feature matching module is used to perform frame-level matching on the video features and audio features aligned by the multi-modal feature alignment module to learn fine-grained audio-visual consistency features;
[0012] The video prediction module is used to fuse the audio-visual multi-modal features output by the cross-modal feature matching module and predict the authenticity of the video.
[0013] Further, the visual and audio data preprocessing module resamples the video and audio to 25fps and 16000Hz respectively, and the output small segments contain 16 consecutive video frames and 10240 audio sampling points.
[0014] Further, the visual deepfake feature extraction and prediction module, the audio deepfake feature extraction and prediction module, and the video prediction module all use a multi-layer perceptron to further predict the extracted features.
[0015] A multi-modal deepfake video detection method based on prompt learning, characterized by including the following steps:
[0016] S1: Construct a multi-modal deepfake video training dataset with audio-visual signals, having video-level annotations and annotations for each of the audio and visual modalities, where 1 represents a fake video / modality and 0 represents a real video / modality;
[0017] S2: Split the video data obtained in step S1 into small segments, and extract the corresponding T video frames and N audio sampling points. The dimensions of all video frames are unified as C×H×W; where C is the number of channels, H is the height, and W is the width.
[0018] S3: Input the multi-frame visual content in step S2 into the visual deepfake feature extraction and prediction network, and use learnable visual cues and a visual base model to obtain the visual features and the authenticity result of the visual modality. The size of the input data is T×C×H×W, and the output feature dimension is T×D v ; where D v is the dimension of the visual features.
[0019] S4: Input the audio signal in step S2 into the audio deepfake feature extraction and prediction network, and use learnable audio cues and a speech recognition base model to obtain the audio deepfake features and the authenticity result of the audio modality. The input audio signal is first converted into a spectrogram with size N m ×N t , where N m and N t are the number of channels and the number of tokens in the spectrogram respectively. The output feature dimension is where D a is the dimension of the audio features.
[0020] S5: Fuse the visual features and audio features output in step S3 and step S4 and input them into the multi-modal classification network to predict the authenticity corresponding to this segment, as the final segment-level prediction result;
[0021] S6: Repeat steps S2 - S5 until the loss function converges to complete the training. Finally, fix all the parameters in the visual feature extraction and prediction network in step S3, the audio deepfake feature extraction and prediction network in step S4, and the multi-modal classification network in step S5;
[0022] S7: Video detection;
[0023] S7-1: For the input multi-modal video, repeatedly execute step S2 to obtain multiple consecutive non-overlapping input segments;
[0024] S7-2: Use all the parameters in the visual feature extraction and prediction network finally fixed in step S6 to execute step S3 to obtain the visual features and the detection result of the visual modality, and add the visual modality result to the visual modality list;
[0025] S7-3: Use all the parameters in the audio deepfake feature extraction and prediction network finally fixed in step S6 to execute step S4 to obtain the audio features and the detection result of the audio modality, and add the audio modality result to the audio modality list;
[0026] S7-4: Using all the parameters in the multi-modal classification network finally fixed in step S6, the visual features obtained in step S7-2, and the audio features obtained in step S7-3, execute step S5 to obtain the final segment-level detection result, and add the segment-level result to the segment-level list.
[0027] S7-5: When the prediction of all small segments within the video is completed, calculate the average values of the prediction scores of the visual modality list, the audio modality list, and the segment-level list respectively, as the authenticity probability of the visual modality of the video, the authenticity probability of the audio modality of the video, and the authenticity probability of the video.
[0028] Furthermore, the loss function in step S3 is the cross-entropy loss function:
[0029]
[0030] where L v is the loss function, N is the number of samples, represents the label of the visual modality category c of the nth sample, and the value of category c is 0 or 1, represents the prediction result of the visual modality category c of the nth sample;
[0031] The specific steps are as follows:
[0032] S3-1: Divide each video frame obtained in step S2 into N v image patches of p×p, use the frozen embedding layer of the visual base model CLIP to extract visual embedding features, and splice the frozen classification tokens to obtain the initial input features The size of the spliced features is (N v +1)×D v ;
[0033] S3-2: The features output in step S3-1 are input into the frozen CLIP encoder. Before inputting into each encoding layer, insert learnable visual cues into the output features of the previous layer. The visual features output by the tth frame image in the lth encoding layer can be expressed as:
[0034]
[0035] where is the ith vector in the output features of the tth frame image in the lth layer, is the learnable sequential visual cue, is the number of encoder layers; N vpt is the number of visual cue vectors, and the subscript v represents visual.
[0036] S3-3: The vector corresponding to the first learnable prompt in the output of the last layer of step S3-2 is the feature of the corresponding video frame. By aggregating frame-level features in the time dimension, the clip feature f can be obtained, which can be expressed as:
[0037]
[0038] S3-4: Based on the clip-level visual features of step S3-3, a visual classification head with an MLP structure is introduced to predict the authenticity of the visual modality, which can be expressed as:
[0039]
[0040] is the authenticity probability of the visual modality.
[0041] Furthermore, the loss function in step S4 is the cross-entropy loss function:
[0042]
[0043] where N is the number of samples, represents the label of the audio modality category c of the nth sample, represents the prediction result of the audio modality category c of the nth sample;
[0044] The specific steps are as follows:
[0045] S4-1: Transform the audio sampling points obtained in step S2 into a spectrogram x a , whose feature size is N m ×N t , where N m and N t are the number of channels and the number of tokens in the spectrogram respectively, and then input it into the frozen embedding layer of the speech recognition basic model Whisper to obtain the initial embedding features whose size is
[0046] S4-2: The features output in step S4-1 are input into the frozen Whisper encoder, and a learnable short-time audio prompt is introduced before inputting into each encoding layer D p is the dimension of the audio prompt. This prompt is concatenated with the output of the previous layer in the feature dimension. To ensure that the feature dimension remains unchanged before and after concatenation, a fully connected layer FC l is introduced in each layer. In addition, this method uses the speech global feature to capture the global feature, and this process can be expressed as:
[0047]
[0048] Among them, cat(·,·) represents concatenation in the feature dimension, is the i-th vector in the output feature of the l-th layer;
[0049] S4-3: Since the audio is a sequence signal, the segment-level audio signal f can be directly calculated on the sequence dimension using the output of the last layer in step S4-2 a , and this process can be expressed as:
[0050]
[0051] Among them, is the number of encoder layers;
[0052] S4-4: Based on the segment-level audio features output in step S4-3, an audio classification head with an MLP structure is introduced to predict the authenticity of the audio modality, which can be expressed as:
[0053]
[0054] is the probability of the authenticity of the audio modality.
[0055] Furthermore, the loss function in step S5 is the cross-entropy loss function:
[0056]
[0057] Among them, N is the number of samples, represents the label of the n-th sample category c, represents the prediction result of the n-th sample category c;
[0058] The specific steps are as follows:
[0059] S5-1: The visual features f obtained in S3-2 and S4-2 v and the audio features f a can be expressed as:
[0060]
[0061] Two convolutional layers are introduced in this method to align the visual features and audio features in the time dimension, which can be expressed as:
[0062]
[0063] Among them, k v and s v are the parameters of the convolutional layer conv v for the visual features, k aand s a is the convolutional layer conv for audio features a parameters. Set k v = s v = 1, to align the sequence lengths of visual and audio features to T;
[0064] S5-2: Concatenate the audio-visual features f' v and f′ a output in step S5-1 on the feature dimension, and output the fused feature D f is the dimension of the fused feature.
[0065] S5-3: Aggregate the fused features output in step S5-2 on the time dimension to extract the multi-modal segment-level feature f f , which can be expressed as:
[0066]
[0067] S5-4: Input the segment-level feature output in step S5-3 into the video classification head of the MLP structure to predict the authenticity of the input segment, which can be expressed as:
[0068] y pred = softmax(MLP(f f )).
[0069] y pred is the authenticity probability of the video.
[0070] Furthermore, the loss function L in step S6 includes L in step 3 v , L in step 4 a , L in step 5 f and the frame-level cross-modal feature matching loss function L CMFM , specifically:
[0071]
[0072] L = L v + L a + L f + L CMFM
[0073] where L mp,nq is an intermediate process variable, and α, β, and γ are all adjustable hyperparameters.
[0074] f v ′ and f a ′is the output result of step S5-1, B is the batch size, T is the number of image frames of the visual content, m and n are intermediate variables representing the m-th or n-th sample in the batch, and p and q are also intermediate variables representing the indices of the p-th or q-th frame of the audio or visual content. is the visual feature corresponding to the p-th frame in the m-th sample of the training mini-batch is the audio feature corresponding to the p-th frame in the m-th sample of the training mini-batch, y m is the label corresponding to the m-th sample, D(·,·) is the cosine distance between two features, N l is the number of non-zero values in the summation process.
[0075] The beneficial effects of the present invention are as follows:
[0076] 1. The multi-task audio-visual prompt learning method of the present invention can detect multi-modal deepfake videos;
[0077] 2. The audio stream and video stream in the two-stream network designed by the present invention combine prompt learning with a visual base model and a speech recognition base model, and have good generalization ability;
[0078] 3. The cross-modal feature matching loss function designed by the present invention can learn fine-grained audio-visual consistency features;
[0079] 4. The present invention can output the authenticity of audio and visual modalities while outputting the authenticity of the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be construed as limiting the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:
[0081] Figure 1 is the structural diagram of the multi-modal deepfake video detection device based on prompt learning of the present invention;
[0082] Figure 2 is the schematic diagram of the training process of the multi-modal deepfake video detection method based on prompt learning of the present invention;
[0083] Figure 3 is the schematic diagram of the testing process of the multi-modal deepfake video detection method based on prompt learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0084] To better understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0085] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0086] As Figure 1 shown, a multi-modal deepfake video detection device based on prompt learning, characterized by comprising: a visual and audio data preprocessing module, a visual deepfake feature extraction and prediction module, an audio deepfake feature extraction and prediction module, a multi-modal feature alignment module, a cross-modal feature matching module, and a video prediction module. Among them,
[0087] The visual and audio data preprocessing module is used to split the input video data into small segments and extract the corresponding visual content and audio signals of the small segments;
[0088] The visual deepfake feature extraction and prediction module is used to extract visual deepfake features from the visual content extracted by the visual and audio content preprocessing module using a visual base model and learnable visual prompts and predict the authenticity of the visual modality;
[0089] The audio deepfake feature extraction and prediction module converts the audio sampling points extracted by the visual and audio content preprocessing module into spectrograms, and then extracts audio deepfake features using a speech recognition base model and learnable audio prompts and predicts the authenticity of the audio modality;
[0090] The multi-modal feature alignment module is used to align the visual features extracted by the visual deepfake feature extraction and prediction module and the audio features extracted by the audio deepfake feature extraction and prediction module in the time dimension;
[0091] The cross-modal feature matching module is used to perform frame-level matching on the video features and audio features aligned by the multi-modal feature alignment module to learn fine-grained audio-visual consistency features;
[0092] The video prediction module is used to fuse the audio-visual multi-modal features output by the cross-modal feature matching module and predict the authenticity of the video.
[0093] Preferably, the visual and audio data preprocessing module resamples the video and audio to 25fps and 16000Hz respectively, and the output small segments contain 16 consecutive video frames and 10240 audio sampling points.
[0094] Preferably, the visual deepfake feature extraction and prediction module, the audio deepfake feature extraction and prediction module, and the video prediction module all use a multi-layer perceptron to further predict the extracted features.
[0095] As Figure 2 shown, a multi-modal deepfake video detection method based on prompt learning includes the following steps:
[0096] S1: Construct a multi-modal deepfake video training dataset with audio-visual signals, with video-level annotations and annotations for each of the audio and visual modalities, where 1 represents a fake video / modality and 0 represents a real video / modality;
[0097] S2: Split the video data obtained in step S1 into small segments, and extract the corresponding T video frames and N audio sampling points. The dimensions of all video frames are unified as C×H×W;
[0098] S3: Input the multi-frame visual content in step S2 into the visual deepfake feature extraction and prediction network, and use learnable visual prompts and a visual base model to obtain visual features and the authenticity result of the visual modality. The size of the input data is T×C×H×W, and the output feature dimension is T×D v ;
[0099] S4: Input the audio signal in step S2 into the audio deepfake feature extraction and prediction network, and use learnable audio prompts and a speech recognition base model to obtain audio deepfake features and the authenticity result of the audio modality. The input audio signal is first converted into a spectrogram, with size N m ×N t where N m and N t are the number of channels and the number of tokens in the spectrogram respectively, and the output feature dimension is
[0100] S5: Fuse the visual features and audio features output in step S3 and step S4 and input them into the multi-modal classification network to predict the authenticity corresponding to the segment, as the final segment-level prediction result;
[0101] S6: Repeat steps S2 - S5 until the loss function converges, complete the training, and finally fix all the parameters in the visual feature extraction and prediction network in step S3, the audio deepfake feature extraction and prediction network in step S4, and the multi-modal classification network in step S5;
[0102] S7: Video detection, as Figure 3 shown;
[0103] S7-1: For the input multi-modal video, repeatedly execute step S2 to obtain multiple consecutive non-overlapping input segments;
[0104] S7-2: Use all the parameters in the finally fixed visual feature extraction and prediction network in step S6 to execute step S3 to obtain visual features and detection results of the visual modality, and add the visual modality results to the visual modality list;
[0105] S7-3: Use all the parameters in the finally fixed audio deepfake feature extraction and prediction network in step S6 to execute step S4 to obtain audio features and detection results of the audio modality, and add the audio modality results to the audio modality list;
[0106] S7-4: Use all the parameters in the finally fixed multi-modal classification network in step S6, the visual features obtained in step S7-2, and the audio features obtained in step S7-3 to execute step S5 to obtain the final segment-level detection results, and add the segment-level results to the segment-level list.
[0107] S7-5: When the prediction of all small segments inside the video is completed, calculate the average value of the prediction scores of the visual modality list, the audio modality list, and the segment-level list respectively as the authenticity probability of the visual modality of the video, the authenticity probability of the audio modality of the video, and the authenticity probability of the video.
[0108] In some embodiments, the loss function in step S3 is the cross-entropy loss function:
[0109]
[0110] where N is the number of samples, represents the label of the visual modality category c of the nth sample, represents the prediction result of the visual modality category c of the nth sample;
[0111] The specific steps are as follows:
[0112] S3-1: Divide each video frame obtained in step S2 into N v image patches of p×p, use the frozen embedding layer of the visual foundation model CLIP to extract visual embedding features, and splice the frozen classification tokens to obtain the initial input features The size of the spliced features is (N v +1)×D v ;
[0113] S3-2: Input the features output in step S3-1 into the frozen CLIP encoder. Before inputting into each encoding layer, insert learnable visual cues into the output features of the previous layer. The tth frame image at the lth encoding layer The output visual features can be expressed as:
[0114]
[0115] where is the i-th vector in the output features of the t-th frame image at the l-th layer, is a learnable sequence visual cue, is the number of encoder layers;
[0116] S3-3: The vector corresponding to the first learnable cue in the output of the last layer of step S3-2 is the feature of the corresponding video frame. By aggregating the frame-level features in the time dimension, the clip feature f can be obtained and can be expressed as:
[0117]
[0118] S3-4: Based on the clip-level visual features of step S3-3, a visual classification head with an MLP structure is introduced to predict the authenticity of the visual modality, which can be expressed as:
[0119]
[0120] In some embodiments, the loss function in step S4 is the cross-entropy loss function:
[0121]
[0122] where N is the number of samples, represents the label of the n-th sample's audio modality category c, represents the prediction result of the n-th sample's audio modality category c;
[0123] The specific steps are:
[0124] S4-1: Transform the audio sampling points obtained in step S2 into a spectrogram x a , whose feature size is N m ×N t , where N m and N t are respectively the number of channels and the number of tokens in the spectrogram, and then input it into the frozen embedding layer of the speech recognition base model Whisper to obtain the initial embedding features whose size is
[0125] S4-2: The features output in step S4-1 are input into the frozen Whisper encoder, and a learnable short-time audio cue is introduced before inputting into each encoding layer This prompt is the output of the previous layer Concatenate in the feature dimension. To ensure that the feature dimension remains unchanged before and after concatenation, a fully connected layer FC is introduced in each layer l . In addition, this method uses to capture global features. This process can be expressed as:
[0126]
[0127] where cat(·,·) represents concatenation in the feature dimension is the i-th vector in the output feature of the l-th layer
[0128] S4-3: Since audio is a sequence signal, the output of the last layer in step S4-2 can be directly used to calculate the segment-level audio signal f in the sequence dimension a , and this process can be expressed as:
[0129]
[0130] where is the number of encoder layers
[0131] S4-4: Based on the segment-level audio features output in step S4-3, an audio classification head with an MLP structure is introduced to predict the authenticity of the audio modality, which can be expressed as:
[0132]
[0133] In some embodiments, the loss function in step S5 is the cross-entropy loss function:
[0134]
[0135] where N is the number of samples represents the label of the n-th sample of class c represents the prediction result of the n-th sample of class c
[0136] The specific steps are as follows:
[0137] S5-1: The visual features f v and audio features f a obtained in S3-2 and S4-2 can be expressed as:
[0138]
[0139] This method introduces two convolutional layers to align the visual features and audio features in the time dimension, which can be expressed as:
[0140]
[0141] Among them, k v and s v are the parameters of the convolutional layer conv v for visual features, and k a and s a are the parameters of the convolutional layer conv a for audio features. Set k v = s v = 1, to align the sequence lengths of visual and audio features to T;
[0142] S5-2: Concatenate the audiovisual features f‘ v and f′ a output in step S5-1 on the feature dimension, and output the fused feature
[0143] S5-3: Aggregate the fused features output in step S5-2 on the time dimension, and extract the multimodal segment-level feature f f , which can be expressed as:
[0144]
[0145] S5-4: Input the segment-level feature output in step S5-3 into the video classification head of the MLP structure to predict the authenticity of the input segment, which can be expressed as:
[0146] y pred = softmax(MLP(f f )).
[0147] In some embodiments, the loss function L in step S6 includes L v in step 3, L a in step 4, L f in step 5, and the frame-level cross-modal feature matching loss function L CMFM , specifically:
[0148]
[0149] L = L v + L a + L f + L CMFM
[0150] Among them, f v ′ and f a ′ are the output results of step S5-1, is the visual feature corresponding to the p-th frame in the m-th sample of the training mini-batch, To train the audio features corresponding to the p-th frame in the m-th mini-batch sample, y m is the label corresponding to the m-th sample, D(·,·) is the cosine distance between two features, and N l is the number of non-zero values in the summation process.
[0151] Preferably, the number of iterations e in step S6 is set to 13.
[0152] To verify the effectiveness and practicality of the present invention, the FakeAVCeleb and KoDF datasets are used for training and evaluation. First, FakeAVCeleb is divided into a training set, a validation set, and a test set. On the training set, the model is trained according to steps S1 - S6, using Adam as the optimizer of the model, with the learning rate set to 0.001, and a total of 15 iterations are trained. At the 12th iteration, the learning rate decays to 1 / 10 of the original. Finally, the model with the best evaluation index on the validation set is saved as the final result.
[0153] The test set of FakeAVCeleb and the audio-driven fake video data in KoDF are used for model evaluation. Specifically, 100 real videos and 100 fake videos are sampled from the data corresponding to the corresponding categories in KoDF to form the test set. For these two datasets, the trained model is evaluated according to the above step S7. Among them, the accuracy of the FakeAVCeleb dataset is 99.8%, the AUC is 99.9%, the accuracy of the KoDF dataset is 92.0%, and the AUC is 94.5%, which belong to good results, indicating that the present invention is effective and feasible.
[0154] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multimodal deep fake video detection device based on prompt learning, characterized in that: include: Visual and audio data preprocessing module, visual deep fake feature extraction and prediction module, audio deep fake feature extraction and prediction module, multimodal feature alignment module, cross-modal feature matching module, video prediction module, among which, The visual and audio data preprocessing module is used to divide the input video data into small segments and extract the visual content and audio signals corresponding to the small segments; The visual deep fake feature extraction and prediction module is used to extract visual deep fake features from the visual content extracted by the visual and audio content preprocessing module and predict the authenticity of the visual modality using a visual base model and learnable visual cues; The audio deep fake feature extraction and prediction module converts the audio sampling points extracted by the visual and audio content preprocessing module into a spectrogram, and then uses the speech recognition basic model and learnable audio cues to extract audio deep fake features and predict the authenticity of the audio modality; The multimodal feature alignment module is used to align the visual features extracted by the visual deep pseudo feature extraction and prediction module and the audio features extracted by the audio deep pseudo feature extraction and prediction module in the time dimension; The cross-modal feature matching module is used to perform frame-level matching on the video features and audio features aligned by the multi-modal feature alignment module to learn fine-grained audio and video consistency features; The video prediction module is used to fuse the audio and video multimodal features output by the cross-modal feature matching module and predict the authenticity of the video.
2. According to the multimodal deep fake video detection device based on prompt learning according to claim 1, it is characterized in that The visual deep fake feature extraction and prediction module, the audio deep fake feature extraction and prediction module, and the video prediction module all use a multi-layer perceptron to further predict the extracted features.
3. A multimodal deep fake video detection device based on prompt learning according to claim 1 or 2, characterized in that: The visual and audio data preprocessing module resamples the video and audio to 25fps and 16000Hz respectively, and the output small clip contains 16 consecutive video frames and 10240 audio sampling points.
4. A multimodal deep fake video detection method based on the device according to any one of claims 1 to 3, characterized in that: The steps include: S1: Construct a multimodal deep fake video training dataset with audio and video signals, with video-level annotations and annotations of the two modalities, audio and visual, where 1 represents a fake video / modality and 0 represents a real video / modality; S2: Divide the video data obtained in step S1 into small segments, and extract the corresponding T video frames and N audio sampling points. The dimensions of all video frames are unified to C×H×W, where C is the channel, H is the height, and W is the width. S3: Input the multi-frame visual content in step S2 into the visual deep fake feature extraction and prediction network, and use the learnable visual cues and visual basic model to obtain the authenticity of the visual features and visual modalities. The size of the input data is T×C×H×W, and the output feature dimension is T×D v ; where D v is the dimension of visual features. S4: Input the audio signal in step S2 into the audio deep fake feature extraction and prediction network, and use the learnable audio prompts and speech recognition basic model to obtain the true and false results of the audio deep fake features and audio modality. The input audio signal is first converted into a spectrogram of size N m ×N t , where N m and N t are the number of channels and tokens in the spectrogram respectively, and the output feature dimension is Where D a is the dimension of the audio feature. S5: The visual features and audio features outputted from step S3 and step S4 are integrated and input into the multimodal classification network to predict the authenticity of the segment, as the final segment-level prediction result; S6: Repeat steps S2 to S5 until the loss function converges and the training is completed, and finally fix all parameters in the visual feature extraction and prediction network in step S3, the audio deep pseudo feature extraction and prediction network in step S4, and the multimodal classification network in step S5; S7: video detection; S7-1: for the input multimodal video, repeatedly perform step S2 to obtain a plurality of continuous non-overlapping input segments; S7-2: using all parameters in the visual feature extraction and prediction network finally fixed in step S6, executing step S3 to obtain the detection results of the visual features and the visual modality, and adding the visual modality results to the visual modality list; S7-3: Utilize all parameters in the audio deep pseudo feature extraction and prediction network finally fixed in step S6, execute step S4, obtain the detection results of the audio features and audio modality, and add the audio modality results to the audio modality list; S7-4: Using all the parameters in the multimodal classification network finally fixed in step S6, as well as the visual features obtained in step S7-2 and the audio features obtained in step S7-3, execute step S5 to obtain the final segment-level detection result, and add the segment-level result to the segment-level list. S7-5: After all the small segments within the video are predicted, the average values of the prediction scores of the visual modality list, audio modality list, and segment-level list are calculated as the visual modality authenticity probability of the video, the audio modality authenticity probability of the video, and the authenticity probability of the video.
5. According to claim 4, a multimodal deep fake video detection method based on prompt learning is characterized in that: The loss function in step S3 is a cross entropy loss function: Among them, L v is the loss function, N is the number of samples, Represents the label of the nth sample visual modality category c, and the value of category c is 0 or 1. Represents the prediction result of the nth sample visual modality category c; The specific steps are: S3-1: Divide each video frame obtained in step S2 into N v p×p image patches, use the embedding layer frozen by the visual base model CLIP to extract visual embedding features, and concatenate the frozen classification tokens to get the initial input features The feature size after splicing is (N v +1)×D v ; S3-2: The features output from step S3-1 are input into the frozen CLIP encoder. Before being input into each encoding layer, the output features of the previous layer are inserted with learnable visual cues. The t-th frame image is at the l-th encoding layer. Output visual features It can be expressed as: in, is the i-th vector in the output feature of the t-th frame image at the l-th layer, is a learnable sequence of visual cues, is the number of encoder layers; N vpt is the number of visual cue vectors, and the subscript v stands for visual. S3-3: The vector corresponding to the first learnable hint in the last layer output of step S3-2 is the feature of the corresponding video frame. By aggregating the frame-level features in the time dimension, the segment feature f can be obtained. v , which can be expressed as: S3-4: Based on the segment-level visual features of step S3-3, the visual classification head of the MLP structure is introduced to predict the authenticity of the visual modality, which can be expressed as: is the true or false probability of the visual modality.
6. According to claim 4, a multimodal deep fake video detection method based on prompt learning is characterized in that: The loss function in step S4 is a cross entropy loss function: Where N is the number of samples, represents the label of the nth sample audio modality category c, represents the prediction result of the audio modality category c of the nth sample; The specific steps are: S4-1: Convert the audio sampling points obtained in step S2 into a spectrogram x a , whose characteristic size is N m ×N t , where N m and N t They are the number of channels and the number of tokens in the spectrogram, respectively, and then input into the frozen embedding layer of the speech recognition basic model Whisper to obtain the initial embedding features Its size is S4-2: The features output from step S4-1 are input to the frozen Whisper encoder and then input to each encoding layer. Previously introduced learnable short audio cues D p is the dimension of the audio prompt. This prompt is consistent with the output of the previous layer. In the feature dimension splicing, in order to ensure that the feature dimension remains unchanged before and after splicing, a fully connected layer FC is introduced in each layer l In addition, this method uses the global characteristics of speech To capture global features, this process can be expressed as: Among them, cat(·,·) represents concatenation in the feature dimension, is the i-th vector in the output feature of the l-th layer; S4-3: Since audio is a sequence signal, the output of the last layer of step S4-2 can be directly used to calculate the segment-level audio signal f in the sequence dimension a , this process can be expressed as: in, is the number of encoder layers; S4-4: Based on the segment-level audio features outputted in step S4-3, an audio classification head with an MLP structure is introduced to predict the authenticity of the audio modality, which can be expressed as: is the probability of the audio mode being true or false.
7. The multimodal deep fake video detection method based on prompt learning according to claim 4 is characterized in that: The loss function in step S5 is a cross entropy loss function: Where N is the number of samples, represents the label of the nth sample category c, Represents the prediction result of the nth sample category c; The specific steps are: S5-1: Visual features f obtained by S3-2 and S4-2 v and audio feature f a It can be expressed as: This method introduces two convolutional layers to align visual features and audio features in the time dimension, which can be expressed as: Among them, k v and v It is the convolution layer conv for visual features v The parameter k a and a It is the convolution layer conv for audio features a Parameters. Set k v =s v =1, to align the sequence length of visual and audio features to T; S5-2: Concatenate the audio and video features f output from step S5-1 in the feature dimension v ‘ and f a ′ , output fusion features D f is the dimension of fusion features. S5-3: Aggregate the fusion features output by step S5-2 in the time dimension to extract multimodal segment-level features f f , which can be expressed as: S5-4: Input the segment-level features output from step S5-3 into the video classification head of the MLP structure to predict the authenticity of the input segment, which can be expressed as: and pred =softmax(MLP(f f )). y pred is the authenticity probability of the video.
8. The multimodal deep fake video detection method based on prompt learning according to claim 4 is characterized in that: The loss function L in step S6 includes the loss function L in step 3 v , L in step 4 a , L in step 5 f And the frame-level cross-modal feature matching loss function L CMFM , specifically: L=L v +L a +L f +L CMFM Among them, l mp,nq is an intermediate process variable, and α, β, and γ are all adjustable hyperparameters. f′ v and f′ a is the output result of step S5-1, B is the batch size, T is the number of image frames of visual content, m and n are intermediate variables, representing the mth or nth sample in the batch, p and q are also intermediate variables, representing the index of the pth frame or qth frame audio or visual content. To train the visual features corresponding to the p-th frame in the m-th sample of a small batch, To train the audio features corresponding to the p-th frame in the m-th sample of the mini-batch, y m is the label corresponding to the mth sample, D(·,·) is the cosine distance between two features, N l is the number of non-zero values in the summation process.
Citation Information
Patent Citations
Multi-modal fusion detection method for deeply-forged audio and video
CN116797896A
Multi-modal detection method and device for deeply-forged video and computer equipment
CN119131661A
Method and device for estimating the authenticity of audio or video content and associated computer program
WO2024003016A1
Cited By
Multi-modal forged video detection method based on multi-head addition cross attention mechanism
CN120635786A