Deep counterfeit content detection method

Through audio and video consistency detection and dual-mode forgery detection models, the problem of insufficient robustness of single-mode forgery detection in the existing technology is solved, real-time forgery content detection and user warnings are realized in mobile video calls, and detection accuracy and security are improved.

CN120340142APending Publication Date: 2025-07-18HANGZHOU ZHONGKE RUIJIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510227403.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing forgery detection technologies usually only identify fake content for a single mode, ignoring the correlation between different modes, resulting in weak robustness in multimodal forgery attacks, lacking user interaction mechanisms and real-time detection capabilities, limiting its application in mobile video call scenarios.

Method used

The audio-video consistency detection model and the dual-modal forged detection model are adopted. Consistent features are extracted through visual encoder and audio encoder, combined with feature fusion module and detector, consistency analysis and forged detection of audio-video features are realized, audio-video comparison learning pre-training is introduced, and a large amount of real data is used for model pre-training, and the model is optimized to adapt to mobile resource limitations.

Benefits of technology

It effectively improves the accuracy and real-timeness of multi-modal forgery detection, can timely identify forged content and warn users, ensure the security of mobile video calls, avoid user privacy leakage, and adapt to resource restrictions on mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340142A_ABST
    Figure CN120340142A_ABST
Patent Text Reader

Abstract

The invention relates to a deep counterfeit content detection method. The method is suitable for the field of computer artificial intelligence deep counterfeit content detection. According to the technical scheme, the method comprises the following steps: acquiring video content to be detected, separating an audio from a video, obtaining an audio sequence based on the audio, and obtaining a video face sequence based on the video; inputting the audio sequence and the video face sequence into a trained audio and video consistency detection model to obtain an audio and video consistency detection result of the to-be-detected video content; the audio and video consistency detection model comprises a visual encoder which is pre-trained through audio and video comparison learning; the audio encoder and the visual encoder are subjected to audio and video comparison learning pre-training together; the feature fusion module is used for fusing the consistency features extracted by the visual encoder and the consistency features extracted by the audio encoder to obtain fusion features; and the audio and video consistency detector is used for obtaining the audio and video consistency probability based on the fusion features of the feature fusion module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting deepfake content, which is applicable to the field of computer artificial intelligence deepfake content detection. Background Art

[0002] With the rapid development of mobile communication technology and intelligent terminal devices, video calls have become an important communication method in people's daily life and work. Mobile video calls are widely used in social, business, education and other fields due to their convenience and real-time nature.

[0003] However, in recent years, technologies based on deep learning have been continuously developing, and deepfake technology has been increasingly refined. This technology can generate highly realistic false video and audio content, making it difficult for ordinary users to identify the false content generated. The mobile terminal is one of the main battlefields for deepfake content attacks. Attackers can use the audio and video call functions of various mobile communication software to access deepfake content to carry out attack behaviors.

[0004] In the video call scenario, forgery attacks mainly manifest in two forms: video forgery and audio forgery. Video forgery replaces or tamps with the face in a video call, replacing the facial expressions, postures and other information of the target user with the attacker's content. Typical forgery methods include technologies such as face swapping, expression transfer, and deep synthesis. Audio forgery is based on speech synthesis technology, generating false voice content by imitating the voiceprint characteristics of the target user to achieve deception or misdirection. Audio forgery can be combined with video forgery to enhance authenticity, further increasing the difficulty of detection and protection.

[0005] Currently, forgery detection technologies usually only identify forged content for a single modality (such as audio or video), ignoring the correlation between different modalities. For example, a video forgery detection model can effectively identify abnormal facial expressions, but cannot determine whether the accompanying audio is forged. Although audio forgery detection can identify abnormalities in voice features, it may lack sufficient information support in a video scenario. This single-modal detection method shows weak robustness in multi-modal forgery attacks.

[0006] Existing forgery detection methods lack a user interaction mechanism and a user risk reminder mechanism combined with the actual application scenario. Most of them are that users actively upload files (pictures, videos, audios, etc.) for detection. Especially in the mobile video call scenario, users need to obtain the detection results of forged content and corresponding risk reminders in real time so as to take protective measures in time. This deficiency in user experience limits the wide application of forgery detection technologies. Summary of the Invention

[0007] The technical problem to be solved by the present invention is: in view of the above problems, to provide a method for detecting deepfake content.

[0008] The technical solution adopted by the present invention is: a method for detecting deepfake content, including: Obtain the video content to be detected, separate the audio and video, and obtain an audio sequence based on the audio and a video face sequence based on the video; Input the audio sequence and the video face sequence into a trained audio-visual consistency detection model to obtain the audio-visual consistency detection result of the video content to be detected; The audio-visual consistency detection model includes: A visual encoder, pre-trained by audio-visual contrast learning, for extracting the consistency features of face and voice from the video face sequence; An audio encoder, pre-trained together with the visual encoder by audio-visual contrast learning, for extracting the consistency features of voice and face from the audio sequence; A feature fusion module for fusing the consistency features extracted by the visual encoder and the consistency features extracted by the audio encoder to obtain fused features; An audio-visual consistency detector for obtaining the audio-visual consistency probability based on the fused features of the feature fusion module.

[0009] A method for detecting deepfake content, including: Obtain the video content to be detected, separate the audio and video, and obtain an audio sequence based on the audio and a video face sequence based on the video; Input the audio sequence and the video face sequence into a trained audio-visual bimodal forgery detection model to obtain the forgery detection result of the video content to be detected; The audio-visual bimodal forgery detection model includes: A visual encoder, pre-trained by audio-visual contrast learning, for extracting the consistency features of face and voice from the video face sequence; A video forgery detection classifier for obtaining the face forgery probability based on the consistency features of the visual encoder; An audio encoder, pre-trained together with the visual encoder by audio-visual contrast learning, for extracting the consistency features of voice and face from the audio sequence; An audio forgery detection classifier for obtaining the audio forgery probability based on the consistency features of the audio encoder.

[0010] A method for detecting deepfake content, including: Obtain the video content to be detected, separate the audio and video, and obtain an audio sequence based on the audio and a video face sequence based on the video; Input the audio sequence and the video face sequence into a multi-modal deepfake content detection model to obtain the multi-modal deepfake content detection result of the video content to be detected; The multi-modal deepfake content detection model is obtained by fusing multiple models, including a trained audio-visual consistency detection model and a trained audio-visual bimodal forgery detection model; The audio-visual consistency detection model includes: The first visual encoder, pre-trained by audio-visual contrast learning, is used to extract the consistency features of face and voice from the video face sequence; The first audio encoder, pre-trained by audio-visual contrast learning together with the first visual encoder, is used to extract the consistency features of voice and face from the audio sequence; The feature fusion module is used to fuse the consistency features extracted by the first visual encoder and the consistency features extracted by the first audio encoder to obtain the fused features; The audio-visual consistency detector is used to obtain the audio-visual consistency probability based on the fused features of the feature fusion module; The audio-visual bimodal forgery detection model includes: The second visual encoder, pre-trained by audio-visual contrast learning, is used to extract the consistency features of face and voice from the video face sequence; The video forgery detection classifier is used to obtain the face forgery probability based on the consistency features of the second visual encoder; The second audio encoder, pre-trained by audio-visual contrast learning together with the second visual encoder, is used to extract the consistency features of voice and face from the audio sequence; The audio forgery detection classifier is used to obtain the audio forgery probability based on the consistency features of the second audio encoder.

[0011] The obtaining of the audio sequence based on the audio includes: denoising the audio to eliminate background noise interference to obtain the audio sequence; The obtaining of the video face sequence based on the video includes: extracting frames from the video, performing face detection, and tracking to extract faces to obtain the video face sequence.

[0012] The audio-visual contrast learning pre-training includes: Regarding real audio-visual pairs as positive samples; regarding the audio and video in different audio-visuals, and the audio-visual pairs with the audio in the original audio-visual randomly moved forward or backward by a certain time as negative samples; Separating the audio and video from the positive and negative samples, obtaining the audio sequence based on the audio, and obtaining the video face sequence based on the video; Inputting the video face sequence and the audio sequence into the visual encoder and the audio encoder respectively to obtain the video face features and the audio features, calculating the contrast loss between the video face features and the audio features, reducing the distance of the positive sample pairs, and increasing the distance of the negative sample pairs.

[0013] The visual encoder uses Vision Transformer; the audio encoder uses MamBa.

[0014] The audio-visual consistency detector uses ResNet18.

[0015] The video forgery detection classifier uses ResNet18; the audio forgery detection classifier uses ResNet18.

[0016] A storage medium stores a computer program executable by a processor, and when the computer program is executed, the steps of the detection method are implemented.

[0017] A multi-modal deepfake content detection device has a memory and a processor. The memory stores a computer program executable by the processor, and when the computer program is executed, the steps of the detection method are implemented.

[0018] A method for deploying a model on a mobile device includes: On the PC side, the trained multi-modal deepfake content detection model is subjected to knowledge distillation, model pruning, and model quantization to obtain a lightweight deepfake detection model. On the mobile device, the lightweight deepfake detection model is adapted to the mobile inference framework to obtain a deepfake detection mobile model, completing the deployment of the model on the mobile device. The multi-modal deepfake content detection model is obtained by fusing multiple models, including a trained audio-visual consistency detection model and a trained audio-visual bimodal forgery detection model. The audio-visual consistency detection model includes: The first visual encoder, pre-trained by audio-visual contrast learning, is used to extract the consistency features of face and voice from the video face sequence. The first audio encoder, pre-trained by audio-visual contrast learning together with the first visual encoder, is used to extract the consistency features of voice and face from the audio sequence. The feature fusion module is used to fuse the consistency features extracted by the first visual encoder and the consistency features extracted by the first audio encoder to obtain fused features. The audio-visual consistency detector is used to obtain the audio-visual consistency probability based on the fused features of the feature fusion module. The audio-visual bimodal forgery detection model includes: The second visual encoder, pre-trained by audio-visual contrast learning, is used to extract the consistency features of face and voice from the video face sequence. The video forgery detection classifier is used to obtain the face forgery probability based on the consistency features of the second visual encoder. The second audio encoder, pre-trained through audio-visual contrastive learning together with the second visual encoder, is used to extract the consistency features of the voice and the face from the audio sequence; The audio forgery detection classifier is used to obtain the audio forgery probability based on the consistency features of the second audio encoder.

[0019] A mobile device, on which the deep forgery detection mobile model is deployed by using the model mobile deployment method.

[0020] The beneficial effects of the present invention are as follows: Consistency should be maintained between audio and video, which means that the visual content in the video and the voice signal in the audio need to be synchronized in time and semantics. If there is a significant inconsistency between the two, it may be caused by forgery technology. In order to detect the feature consistency of audio and video, the present invention uses a visual encoder and an audio encoder pre-trained through audio-visual contrastive learning to extract the consistency features of the face and the voice, and based on the consistency features extracted by the visual encoder and the audio encoder, identifies the inconsistency between audio and video, and timely detects and warns the user of potential forged content.

[0021] In the scenario of real-time video calls on mobile devices, forged content attacks are manifested not only through the video modality (such as face tampering, false expression generation, etc.), but also through the cooperation of the audio modality (such as voice synthesis, voiceprint forgery, etc.). In order to effectively detect such multi-modal forged content, the present invention performs forgery detection for the corresponding modalities through two parts: video and audio; introduces audio-visual contrastive learning pre-training, audio-visual feature consistency analysis, and audio-visual bimodal forgery detection, effectively improving the accuracy and real-time performance of forgery detection, and providing guarantee for the security of mobile video calls. When performing real-time detection during a video call, by periodically recording video call segments and inputting them into the model, first obtain the screen permission, and then input 3-second video call segments recorded at regular intervals each time into the model for forgery identification.

[0022] Audio-visual contrastive learning is an effective method for representation learning through a large number of unlabeled data. A large number of real video data such as speeches and interviews can be directly used for model pre-training without additional annotation of forged data. The present invention introduces the method of contrastive learning into the detection of deep forged content, directly uses a large number of real video data for audio-visual contrastive learning pre-training, learns the natural consistency features between audio and video modalities, and pre-understands the consistency features of the natural face and voice in the same video, so as to more easily discover the abnormal associations in the forged content. The model obtained by the present invention through supervised training on the basis of the pre-trained model has stronger generalization and robustness and higher accuracy than the model without pre-training.

[0023] In the present invention, model distillation, pruning, quantization and other methods are adopted for inference optimization and deployment on mobile devices, aiming to reduce the model size, computational complexity and memory occupancy to adapt to the resource limitations of mobile devices and ensure the performance and efficiency of detection on mobile devices.

[0024] In the present invention, contrastive pre-training is carried out through audio-visual data, consistency analysis of audio-visual features is performed, and forgery detection of audio-visual bimodal is carried out. Among them, the contrastive pre-training of audio-visual data learns the natural consistency features between audio-visual modalities, and pre-understands the consistency features between natural human faces and voices in the same video, so as to more easily discover abnormal associations in forged content. The consistency analysis of audio-visual features and the forgery detection of video bimodal effectively improve the accuracy of forgery detection, can timely alarm the attacks of multi-modal forged content, and provide guarantee for the security of mobile video calls. All detection models perform offline detection on mobile devices, and the data does not need to be uploaded to the cloud, avoiding the problem of user privacy leakage. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a flowchart of audio-visual consistency detection in the embodiment.

[0026] Figure 2 It is a flowchart of contrastive learning pre-training of audio-visual in the embodiment.

[0027] Figure 3 It is a flowchart of audio-visual bimodal forgery detection in the embodiment.

[0028] Figure 4 It is a flowchart of multi-modal deep fake content detection in the embodiment.

[0029] Figure 5 It is a flowchart of model mobile deployment in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] In order to better understand the technical solution of the present application, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0031] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0033] Embodiment 1: This embodiment is a method for detecting deepfake content, which specifically includes the following steps (see Figure 1 ): S1. Obtain the video content to be detected, separate the audio and video, obtain the audio sequence based on the audio, and obtain the video face sequence based on the video.

[0034] In this example, the audio and video are separated from the video content to be detected, the video is framed, face detection is performed, and the face is tracked and extracted to obtain the video face sequence; the audio is denoised to eliminate background noise interference to obtain the audio sequence.

[0035] S2. Input the audio sequence and the video face sequence into the trained audio-visual consistency detection model to obtain the audio-visual consistency detection result of the video content to be detected.

[0036] In this embodiment, the audio-visual consistency detection model includes: a visual encoder, an audio encoder, a feature fusion module, an audio-visual consistency detector, etc.

[0037] In this example, the visual encoder uses Vision Transformer and is pre-trained by audio-visual contrast learning to extract the consistency features of face and voice from the video face sequence; the audio encoder uses MamBa and is pre-trained by audio-visual contrast learning together with the visual encoder to extract the consistency features of voice and face from the audio sequence.

[0038] The audio-visual contrast learning pre-training in this embodiment is as Figure 2 shown, including: 1) Use the real audio-visual pair as the positive sample; use the audio and video in different audio-visuals, and the audio-visual pair with the audio in the original audio-visual randomly moved forward or backward by 500 ms to 3 s as the negative sample; 2) Separate the audio and video from the positive and negative samples, frame the video, perform face detection, track and extract the face to obtain the video face sequence; denoise the audio to eliminate background noise interference to obtain the audio sequence; 3) Input the video face sequence and the audio sequence into the visual encoder and the audio encoder respectively to obtain the video face features and the audio features, calculate the contrast loss between the video face features and the audio features, reduce the distance of the positive sample pair, increase the distance of the negative sample pair, and the loss function uses cross-entropy loss.

[0039] In this embodiment, the feature fusion module uses cross-attention to fuse the consistency features extracted by the visual encoder and the consistency features extracted by the audio encoder to obtain the fusion features; the audio-visual consistency detector uses ResNet18 to determine the audio-visual consistency probability based on the fusion features of the feature fusion module.

[0040] In this example, the audio - video consistency detection model is trained using a training set. Each sample in the training set has audio and video content and the audio - video consistency probability corresponding to this audio and video content. For positive samples, the probability label is 1, and for negative samples, the probability label is 0.

[0041] When training the audio - video consistency detection model in this embodiment, based on the audio and video content in the sample, a video face sequence and an audio sequence are obtained and input into a visual encoder and an audio encoder respectively to obtain corresponding features. After fusing the features, they are input into an audio consistency detector to obtain the audio consistency probability. Combining with the audio - video consistency probability in the sample, the cross - entropy loss is calculated and backpropagation is performed for training.

[0042] Embodiment 2: This embodiment is a method for detecting deep - fake content, which specifically includes the following steps (see Figure 3 ): S1. Obtain the video content to be detected, separate the audio and video, and based on the audio, obtain an audio sequence, and based on the video, obtain a video face sequence.

[0043] In this example, the audio and video are separated from the video content to be detected. The video is frame - extracted, face - detected, and the face is tracked and extracted to obtain a video face sequence; the audio is denoised to eliminate background noise interference to obtain an audio sequence.

[0044] S2. Input the audio sequence and the video face sequence into the trained audio - video bimodal forgery detection model to obtain the forgery detection result of the video content to be detected.

[0045] In this example, the audio - video bimodal forgery detection model includes: a visual encoder, a video forgery detection classifier, an audio encoder, an audio forgery detection classifier, etc.

[0046] In this example, the visual encoder uses Vision Transformer and is pre - trained by audio - video contrastive learning, which is used to extract the consistency features of face and voice from the video face sequence; the audio encoder uses MamBa and is pre - trained by audio - video contrastive learning together with the visual encoder, which is used to extract the consistency features of voice and face from the audio sequence.

[0047] The audio - video contrastive learning pre - training in this embodiment includes: 1) Using real audio - video pairs as positive samples; using the audio and video from different audio - videos, and the audio - video pairs with the audio in the original audio - video randomly moved forward or backward by 500 ms to 3 s as negative samples; 2) Separating the audio and video from the positive and negative samples, frame - extracting the video, face - detecting, and tracking and extracting the face to obtain a video face sequence; denoising the audio to eliminate background noise interference to obtain an audio sequence; 3) The video face sequence and the audio sequence are respectively input into the visual encoder and the audio encoder to obtain video face features and audio features. Calculate the contrast loss between the video face features and the audio features, reduce the distance of the positive sample pair, and increase the distance of the negative sample pair. The loss function uses the cross-entropy loss.

[0048] In this embodiment, the video forgery detection classifier uses ResNet18 to obtain the face forgery probability based on the consistency features extracted by the visual encoder; the audio forgery detection classifier uses ResNet18 to obtain the audio forgery probability based on the consistency features extracted by the audio encoder.

[0049] In this example, the audio-visual bimodal forgery detection model is trained using the training set. Each sample in the training set has audio and video content and the corresponding face forgery probability and audio forgery probability for this audio and video content.

[0050] When the audio-visual bimodal forgery detection model in this embodiment is trained, the video face sequence and the audio sequence are obtained based on the audio and video content in the sample, and are respectively input into the visual encoder and the audio encoder to obtain the corresponding features. The corresponding features are respectively input into the corresponding forgery detection classifiers to obtain the face forgery probability and the audio forgery probability. Combine the face forgery probability and the audio forgery probability in the sample, calculate the loss, and perform backpropagation for training.

[0051] Embodiment 3: This embodiment is a method for detecting deepfake content, which specifically includes the following steps (see Figure 4 ) S1. Obtain the video content to be detected, separate the audio and the video, and obtain the audio sequence based on the audio and the video face sequence based on the video.

[0052] In this example, the audio and the video are separated from the video content to be detected. The video is frame-extracted, face-detected, and the face is tracked and extracted to obtain the video face sequence; the audio is denoised to eliminate background noise interference to obtain the audio sequence.

[0053] S2. Input the audio sequence and the video face sequence into the multi-modal deepfake content detection model to obtain the multi-modal deepfake content detection result of the video content to be detected.

[0054] In this embodiment, the multi-modal deepfake content detection model is obtained by fusing a trained audio-visual consistency detection model and a trained audio-visual bimodal forgery detection model. The audio-visual consistency detection model includes a first visual encoder, a first audio encoder, a feature fusion module, and an audio-visual consistency detector; the audio-visual bimodal forgery detection model includes a second visual encoder, a video forgery detection classifier, a second audio encoder, and an audio forgery detection classifier.

[0055] In this example, the first visual encoder uses Vision Transformer and is pre-trained by audio-visual contrast learning to extract the consistency features of face and voice from the video face sequence; the first audio encoder uses MamBa and is pre-trained by audio-visual contrast learning together with the first visual encoder to extract the consistency features of voice and face from the audio sequence.

[0056] In this example, the second visual encoder uses Vision Transformer and is pre-trained by audio-visual contrast learning to extract the consistency features of face and voice from the video face sequence; the second audio encoder uses MamBa and is pre-trained by audio-visual contrast learning together with the second visual encoder to extract the consistency features of voice and face from the audio sequence.

[0057] The audio-visual contrast learning pre-training in this embodiment includes: 1) Using the real audio-visual pairs as positive samples; using the audio and video in different audio-visuals, and the audio-visual pairs with the audio in the original audio-visual randomly shifted forward or backward by 500 ms to 3 s as negative samples; 2) Separating the audio and video from the positive and negative samples, extracting frames from the video, performing face detection, and tracking to extract faces to obtain the video face sequence; denoising the audio to eliminate background noise interference to obtain the audio sequence; 3) Inputting the video face sequence and the audio sequence into the visual encoder and the audio encoder respectively to obtain the video face features and the audio features, calculating the contrast loss between the video face features and the audio features, reducing the distance of the positive sample pairs, and increasing the distance of the negative sample pairs. The loss function uses cross-entropy loss.

[0058] In this embodiment, the feature fusion module uses cross-attention to fuse the consistency features extracted by the visual encoder and the consistency features extracted by the audio encoder to obtain the fused features; the audio-visual consistency detector uses ResNet18 to determine the audio-visual consistency probability based on the fused features of the feature fusion module.

[0059] In this example, the audio-visual consistency detection model is trained using the training set. Each sample in the training set has audio-visual content and the corresponding audio-visual consistency probability. The probability label for the positive sample is 1, and the probability label for the negative sample is 0.

[0060] When training the audio-visual consistency detection model in this embodiment, the video face sequence and the audio sequence are obtained based on the audio and video content in the sample, and are respectively input into the visual encoder and the audio encoder to obtain corresponding features. After fusing the features, they are input into the audio consistency detector to obtain the audio consistency probability. Combining with the audio-visual consistency probability in the sample, the cross-entropy loss is calculated and backpropagated for training.

[0061] In this embodiment, the video forgery detection classifier uses ResNet18 to obtain the face forgery probability based on the consistency features extracted by the visual encoder; the audio forgery detection classifier uses ResNet18 to obtain the audio forgery probability based on the consistency features extracted by the audio encoder.

[0062] In this example, the audio-visual bimodal forgery detection model is trained using the training set. Each sample in the training set has audio and video content, and the corresponding face forgery probability and audio forgery probability for this audio and video content.

[0063] When training the audio-visual bimodal forgery detection model in this embodiment, the video face sequence and the audio sequence are obtained based on the audio and video content in the sample, and are respectively input into the visual encoder and the audio encoder to obtain corresponding features. The corresponding features are respectively input into the corresponding forgery detection classifiers to obtain the face forgery probability and the audio forgery probability. Combining with the face forgery probability and the audio forgery probability in the sample, the loss is calculated and backpropagated for training.

[0064] In this embodiment, the multimodal deepfake content detection model has a forgery detection result module, which is used to determine the multimodal deepfake content detection result based on the audio consistency probability output by the audio consistency detector, the face forgery probability output by the video forgery detection classifier, and the audio forgery probability output by the audio forgery detection classifier.

[0065] Embodiment 4: This embodiment is a storage medium on which a computer program executable by a processor is stored. When the computer program is executed, the steps of the deepfake content detection method in Embodiment 1 or 2 or 3 are implemented.

[0066] Embodiment 5: This embodiment is a multimodal deepfake content detection device, which has a memory and a processor. A computer program executable by the processor is stored on the memory. When the computer program is executed, the steps of the deepfake content detection method in Embodiment 1 or 2 or 3 are implemented.

[0067] Embodiment 6: As Figure 5 shown, this embodiment is a method for deploying a model on a mobile device, which specifically includes the following steps: Ⅰ. On the PC side, the multi-modal deepfake content detection model trained in Embodiment 3 is subjected to knowledge distillation, model pruning, and model quantization to obtain a lightweight deepfake detection model.

[0068] 1) Knowledge distillation. Use the well-trained large model in the two-stage training as the teacher network and the lightweight network as the student network. Under the same data, guide the student model to learn by taking the predicted value of the teacher network for the sample as the prediction target of the student network. Through the guidance of the teacher model, let the student model learn the generalization ability of the teacher model. Model distillation can reduce the volume and computational complexity of the model.

[0069] 2) Model pruning. Adopt a structured pruning algorithm to remove redundant network connections and parameters, maintain a relatively regular network structure, and improve the running efficiency and speed of the model.

[0070] 3) Model quantization. Adopt INT8 quantization. Convert floating-point parameters into fixed-point numbers or low-precision floating-point numbers to reduce the memory occupancy and computational amount of the model.

[0071] Ⅱ. On the mobile side, adapt the lightweight deepfake detection model to the mobile inference framework to obtain a mobile deepfake detection model and complete the mobile deployment of the model. In this example, a deep learning inference framework for mobile devices is used for inference, enabling the model to be adapted to a wider range of mobile hardware.

[0072] Embodiment 7: This embodiment is a mobile device on which a mobile deepfake detection model is deployed using the model mobile deployment method of Embodiment 6.

Claims

1. A method for detecting deepfake content, characterized in that, Including: Obtain the video content to be detected, separate the audio and video, obtain an audio sequence based on the audio, and obtain a video face sequence based on the video; Input the audio sequence and the video face sequence into the trained audio-visual consistency detection model to obtain the audio-visual consistency detection result of the video content to be detected; The audio-visual consistency detection model includes: A visual encoder, pre-trained by audio-visual contrast learning, for extracting the consistency features of the face and voice from the video face sequence; An audio encoder, pre-trained by audio-visual contrast learning together with the visual encoder, for extracting the consistency features of the voice and face from the audio sequence; A feature fusion module for fusing the consistency features extracted by the visual encoder and the consistency features extracted by the audio encoder to obtain fused features; An audio-visual consistency detector for obtaining the audio-visual consistency probability based on the fused features of the feature fusion module.

2. A deepfake content detection method, characterized in that Including: Obtain the video content to be detected, separate the audio and video, obtain an audio sequence based on the audio, and obtain a video face sequence based on the video; Input the audio sequence and the video face sequence into the trained audio-visual bimodal forgery detection model to obtain the forgery detection result of the video content to be detected; The audio-visual bimodal forgery detection model includes: A visual encoder, pre-trained by audio-visual contrast learning, for extracting the consistency features of the face and voice from the video face sequence; A video forgery detection classifier for obtaining the face forgery probability based on the consistency features of the visual encoder; An audio encoder, pre-trained by audio-visual contrast learning together with the visual encoder, for extracting the consistency features of the voice and face from the audio sequence; An audio forgery detection classifier for obtaining the audio forgery probability based on the consistency features of the audio encoder.

3. A method for detecting deepfake content, characterized in that, Including: Obtain the video content to be detected, separate the audio and video, obtain an audio sequence based on the audio, and obtain a video face sequence based on the video; Input the audio sequence and the video face sequence into the multi-modal deepfake content detection model to obtain the multi-modal deepfake content detection result of the video content to be detected; The multi-modal deepfake content detection model is obtained by fusing multiple models, including the trained audio-visual consistency detection model and the trained audio-visual bimodal forgery detection model; The audio-visual consistency detection model includes: The first visual encoder, pre-trained by audio-visual contrast learning, for extracting the consistency features of the face and voice from the video face sequence; The first audio encoder, pre-trained by audio-visual contrast learning together with the first visual encoder, for extracting the consistency features of the voice and face from the audio sequence; A feature fusion module for fusing the consistency features extracted by the first visual encoder and the consistency features extracted by the first audio encoder to obtain fused features; An audio-visual consistency detector for obtaining the audio-visual consistency probability based on the fused features of the feature fusion module; The audio-visual bimodal forgery detection model includes: The second visual encoder, pre-trained by audio-visual contrast learning, for extracting the consistency features of the face and voice from the video face sequence; A video forgery detection classifier for obtaining a face forgery probability based on the consistency features of the second visual encoder; A second audio encoder, pre-trained by audio-visual contrast learning together with the second visual encoder, for extracting the consistency features of the voice and the face from the audio sequence; An audio forgery detection classifier for obtaining an audio forgery probability based on the consistency features of the second audio encoder.

4. The deepfake content detection method according to claim 1 or 2 or 3, characterized in that: The obtaining of the audio sequence based on the audio includes: denoising the audio to eliminate background noise interference to obtain the audio sequence; The obtaining of the video face sequence based on the video includes: extracting frames from the video, detecting faces, and tracking and extracting faces to obtain the video face sequence.

5. The deepfake content detection method according to claim 1 or 2 or 3, characterized in that, The audio-visual contrast learning pre-training includes: Using real audio-visual pairs as positive samples; using audio and video from different audio-visuals, and audio-visual pairs with the audio in the original audio-visual randomly shifted forward or backward by a certain time as negative samples; Separating the audio and video from the positive and negative samples, obtaining an audio sequence based on the audio, and obtaining a video face sequence based on the video; Inputting the video face sequence and the audio sequence into the visual encoder and the audio encoder respectively to obtain video face features and audio features, calculating the contrast loss between the video face features and the audio features, reducing the distance of the positive sample pairs, and increasing the distance of the negative sample pairs.

6. The deepfake content detection method according to claim 1 or 2 or 3, characterized in that: The visual encoder uses Vision Transformer; the audio encoder uses MamBa.

7. The deepfake content detection method according to claim 1 or 3, characterized in that: The audio-visual consistency detector uses ResNet18.

8. The deepfake content detection method according to claim 2 or 3, characterized in that: The video forgery detection classifier uses ResNet18; the audio forgery detection classifier uses ResNet18.

9. A storage medium having stored thereon a computer program executable by a processor, characterized in that: When the computer program is executed, it implements the steps of the detection method described in any one of claims 1 to 8.

10. A multi-modal deep fake content detection device, having a memory and a processor, with a computer program stored on the memory that can be executed by the processor, characterized in that: When the computer program is executed, it implements the steps of the detection method described in any one of claims 1 to 8.

11. A method for deploying a model on a mobile device, characterized in that, Including: On the PC side, the trained multi-modal deep fake content detection model is subjected to knowledge distillation, model pruning, and model quantization to obtain a lightweight deep fake detection model; On the mobile side, the lightweight deep fake detection model is adapted to the mobile inference framework to obtain a mobile deep fake detection model, completing the mobile deployment of the model; The multi-modal deep fake content detection model is obtained by fusing multiple models, including a trained audio-visual consistency detection model and a trained audio-visual bimodal forgery detection model; The audio-visual consistency detection model includes: A first visual encoder, pre-trained by audio-visual contrast learning, for extracting the consistency features of the face and the voice from the video face sequence; A first audio encoder, pre-trained by audio-visual contrast learning together with the first visual encoder, for extracting the consistency features of the voice and the face from the audio sequence; A feature fusion module for fusing the consistency features extracted by the first visual encoder and the consistency features extracted by the first audio encoder to obtain fused features; An audio-visual consistency detector for obtaining an audio-visual consistency probability based on the fused features of the feature fusion module; The audio-visual bimodal forgery detection model includes: The second visual encoder, pre-trained by audio-visual contrast learning, is used to extract the consistency features of faces and voices from video face sequences; The video forgery detection classifier is used to obtain the face forgery probability based on the consistency features of the second visual encoder; The second audio encoder, pre-trained by audio-visual contrast learning together with the second visual encoder, is used to extract the consistency features of voices and faces from audio sequences; The audio forgery detection classifier is used to obtain the audio forgery probability based on the consistency features of the second audio encoder.

12. A mobile device, characterized in that, The deepfake detection mobile model is deployed on the mobile device by using the model mobile deployment method described in claim 11.