Multimodal deepfake detection method and apparatus, electronic device, and storage medium

By constructing a video feature extraction network using a superpixel-based forgery trace mining transformer and the Information Bottleneck (IB) strategy, the problems of large cross-modal differences and redundant information in multimodal deep forgery detection are solved, achieving high accuracy and strong generalization ability in detection.

CN121958940BActive Publication Date: 2026-07-24BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
Filing Date
2026-04-01
Publication Date
2026-07-24

Smart Images

  • Figure CN121958940B_ABST
    Figure CN121958940B_ABST
Patent Text Reader

Abstract

The application provides a multi-modal deep fake detection method and device, electronic equipment and storage medium, belonging to the technical field of deep fake detection. The method comprises: obtaining initial audio information and initial video information; inputting the information into an audio network and a video feature extraction network respectively for feature extraction to obtain audio features and video features; the video feature extraction network comprises a plurality of superpixel-based fake trace mining transformer modules connected in sequence; inputting the audio features and the video features into a trained multi-modal information refining model, removing noise and redundant information based on an information bottleneck (IB) strategy, and then performing encoding and fusion into multi-modal features, and performing fake detection based on the multi-modal features. The application can improve the accuracy of multi-modal fake detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of deepfake detection technology, and more specifically, relates to a multimodal deepfake detection method, apparatus, electronic device and storage medium. Background Technology

[0002] Deepfake technology uses deep learning methods to replace someone else's face in an image or video with another person's face. The malicious use of this technology poses a serious threat to social stability. With the simultaneous development of audio and visual forgery technologies, current unimodal methods are no longer effective in utilizing multimodal information; therefore, multimodal deepfake detection has emerged.

[0003] Existing multimodal deepfake detection methods have certain limitations. For example, some researchers have proposed detecting multimodal deepfakes by mining forgery traces within and across modalities. However, differences exist between different modalities, and reducing cross-modal differences during feature fusion is insufficient. This is because single-modal features often contain a large amount of redundant information and noise, which, being irrelevant to deepfake detection, interferes with the extraction of effective discriminative features during the fusion stage, ultimately affecting detection accuracy. Therefore, there is an urgent need to explore a multimodal deepfake detection method to improve detection accuracy. Summary of the Invention

[0004] The purpose of this application is to provide a multimodal deep forgery detection method, device, electronic device and storage medium. Based on superpixel forgery trace mining, it can effectively remove redundant information and noise in video. Furthermore, by using the forgery trace emphasis information bottleneck method, redundant information in multimodal mode is further removed to reduce cross-modal gap, and ultimately improve the accuracy of multimodal forgery detection.

[0005] A first aspect of this application provides a multimodal deepfake detection method, comprising: Obtain initial audio and initial video information; Initial audio information is input into an audio network for feature extraction to obtain audio features, and initial video information is input into a video feature extraction network for feature extraction to obtain video features; the video feature extraction network includes multiple superpixel-based forgery detection transformer modules connected in sequence; Audio and video features are input into a trained multimodal information refinement model. Noise and redundant information in the audio and video features are removed based on the Information Bottleneck (IB) strategy. After encoding, they are fused into multimodal features, and forgery detection is performed based on the multimodal features. Each transformer module performs the following operations during the feature extraction process of the initial video information: The input of the transformer module is subjected to position encoding, superpixel projection, and local feature enhancement to obtain the output of the transformer module. The input of the first transformer module is the initial video information, and the input of the remaining transformer modules is the output of the previous transformer module.

[0006] A second aspect of this application provides a multimodal deepfake detection device, comprising: The data acquisition unit is used to acquire initial audio and initial video information; The feature extraction unit is used to input the initial audio information into the audio network to extract audio features and input the initial video information into the video feature extraction network to extract video features. The video feature extraction network includes multiple superpixel-based forgery detection transformer modules connected in sequence. The data processing unit is used to input audio and video features into a trained multimodal information refinement model, remove noise and redundant information from the audio and video features based on the Information Bottleneck (IB) strategy, encode and fuse them into multimodal features, and perform forgery detection based on the multimodal features. The second processing unit is used to encode the single-modal features corresponding to audio and video respectively and then fuse them into multimodal features, and perform forgery detection based on the multimodal features; Specifically, the feature extraction unit is used in the process of extracting features from the initial video information to: The input of the transformer module is subjected to position encoding, superpixel projection, and local feature enhancement to obtain the output of the transformer module. The input of the first transformer module is the initial video information, and the input of the remaining transformer modules is the output of the previous transformer module.

[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the multimodal deepfake detection method described above.

[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the multimodal deepfake detection method described above.

[0009] The beneficial effects of the multimodal deepfake detection method, apparatus, electronic device, and storage medium provided in this application are as follows: This embodiment constructs a video feature extraction network by introducing multiple superpixel-based forgery trace mining transformer modules. This network can progressively focus on subtle yet crucial forgery traces in the image, significantly improving the perception of forged video regions and overcoming the shortcomings of traditional methods in capturing local forgery features. Simultaneously, this embodiment combines an audio network to extract audio features. Both audio and video features are input into a multimodal information refinement model based on the Information Bottleneck (IB) strategy. Before fusion, noise and redundant information in the audio and video features are removed, effectively suppressing interference information irrelevant to the detection task in single-modal features and enhancing the discriminativeness and robustness of cross-modal fusion.

[0010] The multimodal deepfake detection method in this embodiment realizes multi-level feature mining and purification fusion from local details to global semantics, which not only improves the accuracy of multimodal deepfake detection, but also enhances the model's generalization ability in complex scenarios, and can better cope with the challenges brought by current multimodal forgery technology. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a multimodal deepfake detection method provided in an embodiment of this application; Figure 2 A schematic block diagram illustrating a multimodal deepfake detection method provided in an embodiment of this application; Figure 3 A schematic diagram of a video feature extraction network provided in an embodiment of this application; Figure 4 This is a diagram illustrating the effect of the IB strategy provided in one embodiment of this application; Figure 5 This is a structural block diagram of a multimodal deepfake detection device provided in an embodiment of this application; Figure 6 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.

[0014] It is understood that in the embodiments of this application, data such as video information and audio information are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0015] In related technologies, with the increasing threat posed by synthetic images and videos, researchers have proposed various deepfake detection methods. These methods are based on single-modal deepfake detection. Currently, single-modal deepfake detection methods can be divided into two categories: single-frame-based detection methods and video-based detection methods.

[0016] Single-frame-based methods focus on utilizing spatial information within a single frame for forgery detection. For example, convolutional neural networks (CNNs) can uncover finer-grained forgery traces by monitoring neuronal behavior, utilizing frequency-domain perceptual cues, residual information, and identity-perceptual cues. However, these frame-level methods often neglect crucial temporal information necessary for detecting face forgeries in videos. Therefore, some researchers have studied temporal inconsistencies across frames in videos, hoping to effectively improve deepfake detection performance, such as deepfake detection methods based on visual Transformers and diffusion models.

[0017] While single-modal deepfake detection methods have demonstrated good performance, they often neglect audio information synchronized with the video. Fusing audio data can provide more comprehensive clues for deepfake detection. Related technologies, by combining audio-visual speech recognition with a dual-label strategy, can effectively detect deepfakes even when audio, video, or cross-modal data are tampered with, or even when one modality is missing. However, this multimodal approach fails to fully exploit intra-modal and cross-modal forgery clues, resulting in limited detection performance. Therefore, this application proposes the Audio-Visual Information Refinement (AVIR) framework, based on the concepts of superpixels and information bottlenecks, to improve the performance of multimodal deepfake detection.

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0019] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a multimodal deepfake detection method provided in an embodiment of this application. The method may include steps S101 to S103.

[0020] S101: Obtain initial audio and initial video information.

[0021] In this embodiment, initial audio information refers to the unprocessed raw audio data to be detected, and initial video information refers to the unprocessed raw video data to be detected. Initial audio and initial video information can be obtained by stripping the audio tracks from the raw audio and video files.

[0022] S102: Input the initial audio information into the audio network to extract features and obtain audio features. Input the initial video information into the video feature extraction network to extract features from the initial video information and obtain video features. The video feature extraction network includes multiple superpixel-based forgery trace mining transformer modules connected in sequence.

[0023] In this embodiment, reference Figure 2 After inputting the initial audio and video information into the AVIR framework, the first step is to extract the single-modal features of the audio and video through two feature extractors, namely the audio network and the video feature extraction network.

[0024] Specifically, this embodiment uses an audio network to obtain initial audio information. Extracting audio features The audio network can be a VGGish network. This embodiment extracts video features from the initial video information using a video feature extraction network. This video feature extraction network includes a 3D convolutional module and multiple sequentially connected Superpixels-based Defects Mining Transformer (SDMT) modules. Initial video information The size is ,in, T For video frame rate, H and W These represent the height and width of the video, respectively, and 3 represents the RGB color channels of each frame. After inputting the above initial video information into the video feature extraction network, the convolutional kernel of the 3D convolutional module will... The convolution module slides along the three dimensions of "frame-height-width," simultaneously capturing temporal dynamics between video frames and spatial texture information within a single frame, such as facial details and background features. After processing by the 3D convolution module, the processed initial video information is transmitted to multiple sequentially connected superpixel-based forgery detection transformer modules for feature extraction.

[0025] Each transformer module performs the following operations during the feature extraction process of the initial video information: The input of the transformer module is subjected to position encoding, superpixel projection, and local feature enhancement to obtain the output of the transformer module. The input of the first transformer module is the initial video information, and the input of the remaining transformer modules is the output of the previous transformer module.

[0026] In this embodiment, reference Figure 3 Each SDMT block includes a position encoder, a superpixel projection module, and a feedforward network. The number of SDMT blocks is [number missing]. N= .

[0027] Among them, positional encoding is used to add temporal and spatial positional information to the initial video information after input processing, which solves the problem that the Transformer model in the existing technology is not sensitive to feature position, so that the subsequent processing module can perceive the spatial coordinates of superpixels in the video frame and the temporal order between frames, thereby accurately capturing the positional association of forgery traces.

[0028] The superpixel projection module is used to re-divide the map tiles based on regional correlation, and aggregate and map the location-encoded video information into superpixel-level feature representations. That is, it transforms continuous pixel-level features into superpixel feature blocks with local semantics, thereby extracting global context representations more efficiently, reducing computational redundancy (reducing the number of map tiles), and at the same time preserving the local texture and structural information of forgery traces.

[0029] The feedforward network further extracts discriminative information from the superpixel features, enhancing the features' ability to express forgery patterns, and ultimately outputs superpixel-level forgery trace features that can be used for deep forgery detection. Finally, the initial video information is processed through N SDMT modules to obtain video features. .

[0030] S103: Input audio and video features into the trained multimodal information refinement model, remove noise and redundant information from the audio and video features based on the Information Bottleneck (IB) strategy, encode and fuse them into multimodal features, and perform forgery detection based on the multimodal features.

[0031] In this embodiment, the multimodal information refinement model is also known as the Artifacts-Emphasis Information Bottleneck (AEIB) model. AEIB, based on the information bottleneck strategy, learns minimal and sufficient unimodal and multimodal representations. Here, "features without redundancy and noise" constitute the minimal and sufficient representation. For unimodal features, AEIB can filter out noise irrelevant to the detection task, thereby refining the unimodal representation; for multimodal features, AEIB focuses on eliminating redundancy and reducing cross-modal gaps in the fused multimodal representation.

[0032] In this embodiment, noise refers to redundant or external information in the input that is irrelevant to deepfake detection, such as background, lighting variations, or other non-essential elements. Such irrelevant noise limits model performance. Redundancy refers to information that can be inferred from other existing information. For example, suppose... For input features, These are characteristics refined by IB, and If for tags ,satisfy Then it is called Compared to It is redundant. Here, the feature that is neither redundant nor noisy is the minimal sufficient representation, which retains only the information shared with the task and removes non-shared components. The minimal sufficient representation aims to represent data in the most concise way, retaining the necessary information while avoiding unnecessary redundancy.

[0033] In this embodiment, after inputting audio and video features into a trained multimodal information refinement model, the model can remove noise and redundant information from each single-modal feature of the audio and video features based on the Information Bottleneck (IB) strategy. It can also remove noise and redundant information from the audio-video fusion features, ensuring that the final multimodal representation retains only information relevant to multimodal forgery clues. Finally, the multimodal features (audio and video features) extracted by the trained multimodal information refinement model are input into a multilayer perceptron module to calculate the probability that the video has been tampered with.

[0034] The multilayer perceptron is an existing model that can include multiple fully connected layers, dropout layers, and an output layer. In this embodiment, multimodal features are first input into the first fully connected layer, where a linear transformation is performed and the ReLU activation function is used to capture the complex relationships between features, outputting an activated feature vector. This feature vector is then input into the dropout layer, which randomly selects some data to discard according to a set discard probability. The remaining data is then scaled to obtain a regularized feature vector. The dropout layer prevents overfitting, avoids excessive reliance on certain feature dimensions, and improves generalization ability. The regularized feature vector is then sequentially input into multiple fully connected layers to obtain the output vector of the last layer. This output vector is then input into the output layer, where a linear transformation is performed and a Sigmoid activation function is applied to obtain the video tampering probability.

[0035] As can be seen from the above, this embodiment constructs a video feature extraction network by introducing multiple superpixel-based forgery trace mining transformer modules. This network can focus layer by layer on subtle and crucial forgery traces in the image, significantly improving the perception ability of video forgery areas and overcoming the shortcomings of traditional methods in capturing local forgery features. Simultaneously, this embodiment combines an audio network to extract audio features. Both audio and video features are input into a multimodal information refinement model based on the Information Bottleneck (IB) strategy. Before fusion, noise and redundant information in the audio and video features are removed, effectively suppressing interference information irrelevant to the detection task in single-modal features and enhancing the discriminativeness and robustness of cross-modal fusion.

[0036] In summary, the multimodal deepfake detection method in this embodiment achieves multi-level feature mining and purification fusion from local details to global semantics. This not only improves the accuracy of multimodal deepfake detection but also enhances the model's generalization ability in complex scenarios, enabling it to better address the challenges posed by current multimodal forgery technologies.

[0037] In one embodiment of this application, the output of the transformer module is obtained by performing position encoding, superpixel projection, and local feature enhancement on the input of the first transformer module, including: After positional encoding of the initial video information, N tiles are obtained; A global context representation is obtained by superpixel projection of N map tiles using a k-means-based superpixel algorithm. The global context representation is enhanced with local features through a convolutional feedforward network and then fused with the global context representation to obtain the output of the first transformer module.

[0038] In this embodiment, the architecture reference of the superpixel-based forgery trace mining transform is as follows: Figure 3 The initial video information input First, the image passes through a 3D convolutional module, then sequentially through N transformer modules. Based on regional correlations, the image is re-divided into patches, thereby extracting global contextual representations more efficiently. Finally, a feedforward network enhances the local representations, ultimately yielding the video features. .

[0039] In this embodiment, it is assumed that the initial video information after 3D convolution processing still uses It means that it will After being input into the first transformer module of SDMT, the output after position encoding, superpixel projection, and local feature enhancement is as follows: , ,

[0040] Where X represents the output of the position encoding. Y represents the video position coding function, and Y represents the output of the superpixel projection. This represents the superpixel projection attention function. This represents the output result of local feature enhancement. For a feedforward network, LN indicates layer normalization.

[0041] In one embodiment, a global context representation is obtained by superpixel projection of N map tiles using a k-means-based superpixel algorithm, including: Determine the pixel sequence based on N map tiles; The k-means-based superpixel algorithm maps pixel sequences to superpixel sequences, performs self-attention operations on the superpixel sequences to obtain local features, and the data volume of the superpixel sequence is less than that of the pixel sequence. Global context representation is obtained based on local features and pixel sequences.

[0042] In this embodiment, a k-means-based superpixel algorithm is used during superpixel projection. First, the pixel sequence, i.e., the token sequence, is determined based on N map tiles. (in , C (For RGB color channels), the token sequence is mapped to a super-token sequence. , M represents the data size of the super-token, where M is less than N. This embodiment improves the expressive power of each token by aggregating semantically related tokens together using a k-means-based superpixel algorithm to form new super-tokens.

[0043] Specifically, the initial super-token is sampled by dividing the tiles into a grid and averaging the tokens within the grid. Thus, the association mapping can be obtained. :

[0044] in This represents the weight in the t-th iteration. , i Indicates the index of the original pixel tag. j Indicates the index of the superpixel marker. t Indicates the number of iterations. Indicates the first i The feature vector of each original pixel label Indicates the first j The first superpixel markert -1 iterations of the label vector. Correspondingly, the new superpixel cluster centers are updated through a weighted sum of pixel features:

[0045] In each iteration t, we base our calculations on the center of the previous superpixel. Calculate its relationship with all pixel tags The similarity is used to obtain the association mapping. After T iterations, the superpixel center no longer changes, resulting in a stable final correlation mapping. To ensure efficiency, this embodiment only calculates the association between each token and its adjacent super-tokens.

[0046] In one embodiment, since the super-token sequence aggregates visually similar regions, a self-attention mechanism helps to highlight forgery traces. Therefore, this embodiment employs a self-attention operation on the super-token S sequence:

[0047] in This is a scaling factor (to prevent gradient vanishing). For parameters, respectively The linear transformation. In this embodiment, the self-attention operation does not impose neighborhood restrictions to ensure the propagation of long-range information.

[0048] While self-attention-based super-tokens can better detect global forgery traces based on region relevance, they may lose some local information during projection. Therefore, we utilize association mapping. super-token (i.e., local features) Upsampled back to the original pixel sequence token and integrate super-token With visual pixel sequence token Obtain a global context representation to preserve both global and local information.

[0049] As can be seen from the above, this embodiment uses a superpixel-based forgery trace mining transformer to cluster forgery traces and form identifiable local regions, thereby eliminating redundant and noise information contained in the video information, which is beneficial for subsequent cross-modal detection.

[0050] In one embodiment of this application, when obtaining single-modal features and Subsequently, to further reduce redundancy and narrow the cross-modal gap, we propose an information bottleneck strategy for forgery trace enhancement, which is used to learn a minimal multimodal representation containing only the information necessary for deep forgery detection.

[0051] In this embodiment, the Information Bottleneck (IB) strategy is first briefly introduced: like Figure 4 As shown, in the original feature representation part, all audio and video information is used for prediction. In the minimum sufficient representation part, information can be divided into modality non-shared information and modality shared information. For non-shared information, each modality is refined to extract task-relevant components; for shared information, redundancy and noise are further removed. Therefore, according to... Figure 4 As can be seen from the minimal sufficient representation, the minimal sufficient representation consists of three parts: audio non-shared task-related information (i.e., audio-specific task-related information), Figure 4 (Part marked with ①) Video non-shared task information (i.e., video-specific task information). Figure 4 (Part ② of the standard number) and information related to multimodal tasks (i.e., modality sharing information, Figure 4 (Part ③ of the standard notation). The gray area represents redundant information in the multimodal task; that is, this information is useful but relatively redundant. Light yellow and light blue represent useless audio and video information, respectively. This minimal sufficient representation retains less task-related information from the input multimodal information than other sufficient representations, but it is sufficient to complete the prediction task.

[0052] In this embodiment, since the difference in the distribution of single-modal information can significantly hinder the effective mining of modal information by the AEIB model, the AEIB model first uses the IB strategy to extract single-modal features, and then uses the IB strategy to extract multimodal features after the fusion of audio and video features.

[0053] Specifically, in one embodiment, the training process of the multimodal information refinement model (AEIB model) includes: The multimodal information refinement model is trained based on historical audio and video features and corresponding label information; the overall optimization objective function of the multimodal information refinement model includes the objective function of unimodal IB and the objective function of multimodal IB. Training stops when the multimodal information refinement model reaches the required number of iterations or the overall optimization objective function meets the preset conditions, thus obtaining the trained multimodal information refinement model.

[0054] In this embodiment, some symbol definitions are first introduced. The training set used by the model during training is denoted as:

[0055] in, and Audio information With video information Modal labels, and Multimodal joint labeling: If any modality is tampered with, it is defined as forgery. For all labels, 0 represents forgery, 1 represents authenticity, and Q represents the total amount of data.

[0056] In this embodiment, the AEIB model first employs the IB strategy to extract single-modal features, aiming to achieve the ideal state of minimum noise. For the input video features... IB can learn compact representations By reducing noise and redundancy, it can be completely replaced. Used for forgery detection (i.e., maximizing) Clearly, the most informative representations often contain noise or redundancy. Therefore, it is necessary to... and Constraints are imposed on the mutual information between the audio and video modalities. The objective function of the single-modal IB is... for:

[0057] In the formula, This represents the objective function of the single-mode IB. True or false labels representing audio features. This represents the compact audio representation learned through the IB strategy. True or false labels representing video features. This represents a compact video representation learned through the IB strategy. Represents the mutual information function. The weight parameters represent those associated with the audio features. This represents the weight parameters associated with video features.

[0058] Among them, mutual information Defined as:

[0059]

[0060] in Indicates the Kullback-Leibler divergence. For joint distribution, and These are marginal distributions. In the above objective function... In the middle, the first and third terms can make the representation compact. and The goal is to maximize its predictive power for labels, while the second and fourth terms are represented by constraint encoding. Compared with the original input information , Mutual information between them forces them to reduce the retention of redundant or irrelevant information in the input.

[0061] Mutual Information The definition method is the same as mutual information This will not be elaborated upon here.

[0062] In one embodiment, the AEIB model performs unimodal processing on the audio features and video features respectively to obtain two refined unimodal representations, and then fuses the two refined unimodal representations into a multimodal feature. To further reduce redundancy and obtain the final compact multimodal features Multimodal features are processed using the IB (Introduction By Interpreter) strategy. The objective function of multimodal IB is:

[0063] In the formula, This represents the objective function of multimodal IB. This represents the label information corresponding to audio and video features. Represents multimodal features, This represents the compact multimodal representation learned through the IB strategy. Represents the mutual information function. The weight parameters represent those associated with the multimodal features. express and Mutual information functions between them.

[0064] In this embodiment, This enables the AEIB model to further extract useful information from different modalities. Considering that audio features and visual features may correspond to different labels, it is defined as follows:

[0065] Mutual Information The definition of multimodal IB encourages information sharing between modalities when their labels are identical (both real or both fake); conversely, it suppresses unnecessary modal coupling when labels are inconsistent (e.g., only one modality is tampered with), thereby improving the ability to detect some fake samples. The goal of multimodal IB is to promote... While maximizing the discriminative power of the labels, we should minimize our reliance on the original multimodal input information.

[0066] In one embodiment, the overall optimization objective function of the AEIB model for:

[0067] In the training process of a multimodal information refinement model, the final trained model is obtained by minimizing the overall optimization objective function or reaching the required number of iterations. The number of iterations can be set empirically. This is achieved by minimizing... The model encourages the encoded representation to be as discriminative as possible to the label, while "forgetting" redundant and noisy information in the input. This form allows us to achieve end-to-end backpropagation by single sampling of random encoded samples, and guarantees that the gradient is an unbiased estimate of the true desired gradient.

[0068] Finally, the multimodal features (audio and video features) extracted by the trained multimodal information refinement model are input into the multilayer perceptron module to calculate the probability that the video has been tampered with.

[0069] In one embodiment of this application, in addition to training the multimodal information refinement model, the overall multimodal deep forgery detection framework of this application embodiment can also be trained, namely... Figure 2 The framework is used for training, and the final loss function of deepfake detection is then determined. It can be written as:

[0070] in , and These are the weighting coefficients for each item. Represents the binary cross-entropy loss. This is a tag to indicate whether a video (including audio) is real or fake.

[0071] In summary, this application proposes an information bottleneck (IB) model to enhance forgery traces by introducing an information bottleneck strategy to further eliminate redundancy and bridge cross-modal gaps. This model filters out noise irrelevant to the prediction task in single-modal and multi-modal representations, thereby obtaining more discriminative refined features and making the detection rate of forgery information more accurate.

[0072] In one embodiment, we conducted multimodal deepfake detection experiments on the following datasets: DFDC, FakeAVCeleb, and DefakeAVMiT. Details of each dataset are as follows: DFDC dataset: This dataset contains over 120,000 videos, including 19,154 real videos and 100,000 deepfake videos. FakeAVCeleb dataset: This audio-video multimodal deepfake detection dataset contains both video and audio deepfake samples and achieves accurate lip-sync. It employs three deepfake methods—FaceSwap, DeepFaceLab, and FSGAN—and includes three forgery types: "real video & fake audio," "fake video & real audio," and "fake video & fake audio," with the entire dataset containing over 20,000 fake videos. DefakeAVMiT dataset contains 540 real videos and 6,480 deepfake videos.

[0073] Specifically, for visual information, each video segment is edited into 30 frames, and facial features are extracted using a face detector. When multiple faces are detected in a single frame, only the face with the largest area is retained. The extracted faces are then resized to 128×128 and a data augmentation process involving random cropping and random horizontal flipping is applied. For audio information, the audio is resampled to 16 kHz and converted into a 96×64 log-Melogram as audio input. The hyperparameter settings of the multimodal deep forgery detection framework are as follows: , , Information bottleneck hyperparameters , and The values ​​are set to 1, 1, and 2 respectively. The latent feature space dimension for both audio and video is 1000. For the learning parameters, the batch size is set to 24, the learning rate is set to 0.0001, the weight decay is set to 0.0005, and the number of iterations is set to 150.

[0074] This embodiment trains the multimodal deep forgery detection framework based on the parameters and dataset set above, resulting in a trained multimodal deep forgery detection framework.

[0075] Corresponding to the multimodal deepfake detection method in the above embodiments, Figure 5 This is a structural block diagram of a multimodal deepfake detection device provided in one embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 5 The multimodal deepfake detection device 20 includes: a data acquisition unit 21, a feature extraction unit 22, and a data processing unit 23. Among them, the data acquisition unit 21 is used to acquire initial audio information and initial video information; The feature extraction unit 22 is used to input the initial audio information into the audio network to extract features and obtain audio features, and input the initial video information into the video feature extraction network to extract features from the initial video information and obtain video features; the video feature extraction network includes multiple superpixel-based forgery trace mining transformer modules connected in sequence; The data processing unit 23 is used to input audio features and video features into the trained multimodal information refinement model, remove noise and redundant information from the audio features and video features based on the information bottleneck IB strategy, and then fuse them into multimodal features after encoding, and perform forgery detection based on the multimodal features.

[0076] Specifically, the feature extraction unit 22 is used in the process of feature extraction from the initial video information to: The input of the transformer module is subjected to position encoding, superpixel projection, and local feature enhancement to obtain the output of the transformer module. The input of the first transformer module is the initial video information, and the input of the remaining transformer modules is the output of the previous transformer module.

[0077] In one embodiment of this application, the feature extraction unit 22 is specifically used for: After positional encoding of the initial video information, N tiles are obtained; A global context representation is obtained by superpixel projection of N map tiles using a k-means-based superpixel algorithm. The global context representation is enhanced with local features through a convolutional feedforward network and then fused with the global context representation to obtain the output of the first transformer module.

[0078] In one embodiment of this application, when the feature extraction unit 22 performs superpixel projection on N patches using a k-means-based superpixel algorithm to obtain a global context representation, it is specifically used for: Determine the pixel sequence based on N map tiles; The k-means-based superpixel algorithm maps pixel sequences to superpixel sequences, performs self-attention operations on the superpixel sequences to obtain local features, and the data volume of the superpixel sequence is less than that of the pixel sequence. Global context representation is obtained based on local features and pixel sequences.

[0079] In one embodiment of this application, the training process of the multimodal information refinement model includes: The multimodal information refinement model is trained based on historical audio and video features and corresponding label information; the overall optimization objective function of the multimodal information refinement model includes the objective function of unimodal IB and the objective function of multimodal IB. Training stops when the multimodal information refinement model reaches the required number of iterations or the overall optimization objective function meets the preset conditions, and the trained multimodal information refinement model is obtained.

[0080] In one embodiment of this application, the objective function of the single-mode IB is:

[0081] In the formula, This represents the objective function for a single-mode IB. True or false labels representing audio features. This represents the compact audio representation learned through the IB strategy. True or false labels representing video features This represents a compact video representation learned through the IB strategy. Represents the mutual information function. The weight parameters represent those associated with the audio features. This represents the weight parameters associated with video features.

[0082] In one embodiment of this application, the objective function of the multimodal IB is:

[0083] In the formula, This represents the objective function of multimodal IB. This represents the label information corresponding to audio and video features. Represents multimodal features, This represents the compact multimodal representation learned through the IB strategy. Represents the mutual information function. The weight parameters represent those associated with the multimodal features. express and Mutual information functions between them.

[0084] In one embodiment of this application, The expression is:

[0085] In the formula, This represents the compact audio representation learned through the IB strategy. This represents a compact video representation learned through the IB strategy. This indicates that the labels for both audio and video modalities are the same.

[0086] See Figure 6 , Figure 6 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 6The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the units in the above-described device embodiments, for example... Figure 5 The functions of the data acquisition unit 21, feature extraction unit 22, and data processing unit 23 shown are illustrated.

[0087] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0088] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0089] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory.

[0090] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation method described in the multimodal deep forgery detection method provided in the embodiments of this application, or they can execute the implementation method of the electronic device described in the embodiments of this application, which will not be repeated here.

[0091] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0092] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0093] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0094] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0095] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces or units, or they may be electrical, mechanical, or other forms of connection.

[0096] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0097] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0098] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A multimodal deepfake detection method, characterized in that, include: Obtain initial audio and initial video information; The initial audio information is input into an audio network for feature extraction to obtain audio features, and the initial video information is input into a video feature extraction network for feature extraction to obtain video features; the video feature extraction network includes multiple superpixel-based forgery trace mining transformer modules connected in sequence; The audio features and video features are input into a trained multimodal information refinement model. Noise and redundant information in the audio and video features are removed based on the Information Bottleneck (IB) strategy. After encoding, they are fused into multimodal features. Forgery detection is performed based on the multimodal features. During the feature extraction process of the initial video information, each of the converter modules performs the following operations: After performing position encoding, superpixel projection, and local feature enhancement on the input of the transformer module, the output of the transformer module is obtained; the input of the first transformer module is the initial video information, and the input of the remaining transformer modules is the output of the previous transformer module.

2. The multimodal deepfake detection method as described in claim 1, characterized in that, After performing position encoding, superpixel projection, and local feature enhancement on the input of the first transformer module, the output of the transformer module is obtained, including: The initial video information is positionally encoded to obtain N map tiles; The global context representation is obtained by superpixel projection of the N map tiles using a k-means-based superpixel algorithm. The global context representation is enhanced with local features through a convolutional feedforward network and then fused with the global context representation to obtain the output of the first transformer module.

3. The multimodal deepfake detection method as described in claim 2, characterized in that, The step of using a k-means-based superpixel algorithm to perform superpixel projection on the N patches to obtain a global context representation includes: Determine the pixel sequence based on the N map tiles; The k-means-based superpixel algorithm maps the pixel sequence to a superpixel sequence, performs a self-attention operation on the superpixel sequence to obtain local features, and the data volume of the superpixel sequence is less than the data volume of the pixel sequence. A global context representation is obtained based on the local features and the pixel sequence.

4. The multimodal deepfake detection method as described in claim 1, characterized in that, The training process of the multimodal information refinement model includes: The multimodal information refinement model is trained based on historical audio and video features and corresponding label information; the overall optimization objective function of the multimodal information refinement model includes the objective function of unimodal IB and the objective function of multimodal IB. When the multimodal information refinement model reaches the required number of iterations or the overall optimization objective function meets the preset conditions, training stops, and the trained multimodal information refinement model is obtained.

5. The multimodal deepfake detection method as described in claim 4, characterized in that, The objective function of the single-modal IB is: In the formula, This represents the objective function for a single-mode IB. True or false labels representing audio features. This represents the compact audio representation learned through the IB strategy. True or false labels representing video features. This represents a compact video representation learned through the IB strategy. Represents the mutual information function. The weight parameters represent those associated with the audio features. This represents the weight parameters associated with video features.

6. The multimodal deepfake detection method as described in claim 4, characterized in that, The objective function of the multimodal IB is: In the formula, This represents the objective function of multimodal IB. This represents the label information corresponding to audio and video features. Represents multimodal features, This represents the compact multimodal representation learned through the IB strategy. Represents the mutual information function. The weight parameters represent those associated with the multimodal features. express and Mutual information functions between them.

7. The multimodal deepfake detection method as described in claim 6, characterized in that, The expression is: In the formula, This represents the compact audio representation learned through the IB strategy. This represents a compact video representation learned through the IB strategy. This indicates that the labels for both audio and video modalities are the same.

8. A multimodal deepfake detection device, characterized in that, include: The data acquisition unit is used to acquire initial audio and initial video information; The feature extraction unit is used to input the initial audio information into an audio network for feature extraction to obtain audio features, and to input the initial video information into a video feature extraction network for feature extraction to obtain video features; the video feature extraction network includes multiple superpixel-based forgery trace mining transformer modules connected in sequence. The data processing unit is used to input the audio features and the video features into a trained multimodal information refinement model, remove noise and redundant information from the audio features and video features based on the information bottleneck (IB) strategy, encode and fuse them into multimodal features, and perform forgery detection based on the multimodal features. Specifically, the feature extraction unit, in the process of extracting features from the initial video information, is used for: After performing position encoding, superpixel projection, and local feature enhancement on the input of the transformer module, the output of the transformer module is obtained; the input of the first transformer module is the initial video information, and the input of the remaining transformer modules is the output of the previous transformer module.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.