Multi-mode depth forgery detection model for time forgery positioning
The multi-modal deepfake detection model addresses the limitations of existing methods by enhancing fine-scale feature capture and adaptive modal fusion, achieving precise video forgery localization through ConvNeXt-based feature extraction and fusion.
Patent Information
- Application Number
- CN202510775152.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The existing depth forgery detection methods are insufficient to express subtle scale forgery features in the time forgery positioning task, ignore the artifacts of high-frequency regions, and cannot adaptively adjust the modal importance during the multimodal feature fusion process, resulting in inaccurate judgment and positioning of forgery frames.
Visual and audio features are extracted using a ConvNeXt-based feature encoder, combined with a high-frequency feature extraction module and an adaptive multimodal feature fusion module, and dynamically adjust the modal feature weights through attention maps to generate the final forged boundary prediction.
It improves the ability to capture subtle scale forged features, enhances the identification of high-frequency artifacts, ensures the retention of key forged feature information, improves the accuracy and robustness of multimodal feature fusion, and realizes the accurate positioning of forged clips in video.
Smart Images

Figure CN120318593A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of network content security, cyberspace security technology, etc. Specifically, it is a multi-modal deepfake detection model for time forgery localization. Background Art
[0002] Digital video, as an efficient information carrier, plays an important role in the information dissemination of modern society. To better display and utilize digital video, the accompanying video editing technology has also been developing rapidly. Currently, the most representative one is the Deepfake technology. Since its birth, it has spread rapidly on the network and been continuously updated and iterated due to its advantage of being able to generate high-quality videos in a short time with only low-cost and low-threshold operations.
[0003] The purpose of traditional deepfake detection is to judge the authenticity of a video by detecting artifacts in the spatial and frequency domains of the image and unnatural biometric features (such as lip movements inconsistent with audio, abnormal blinking actions, etc.). A large number of research methods have been proven to have strong capabilities to judge the authenticity of videos. However, all these traditional detection methods assume that the detection task is a binary classification task, and they only focus on identifying whether the content of a complete video has been tampered with. This limitation will seriously affect their detection effectiveness in complex scenarios. The current deepfake technology already has the ability to accurately tamper with key semantics. For example, in a public speech by the founder of a certain technology company, the statement "Autonomous driving has sufficient safety" is replaced with "hidden danger" through AI synthesis technology. This millisecond-level voice-text synchronous forgery not only retains the original voice rhythm features but also creates strongly misleading false evidence through semantic inversion. This requires that the detection method can not only detect the authenticity of the video but also be able to locate the start and end positions of the forged segments in the video. Currently, the research on time forgery localization is still in its infancy. Such methods are achieved by combining traditional deepfake binary classification detection technology with localization technologies such as Proposal Relation Block (PRB) and boundary matching loss. Although the current methods have made certain progress, there are still the following problems: 1. Existing methods tend to analyze the overall image semantic features while ignoring the forgery traces at the fine scale. In the scenario of temporal forgery localization tasks, the forged segments are very short, and the deep forgery technique of Expression Swap is used. The forged areas are usually organs such as the mouth or eyes rather than the whole face. In this case, micro-scale features such as edges, textures, and corners are more conducive to reflecting forgery traces than the overall semantics of the image. Moreover, due to the use of continuous convolutions in existing methods, the fine-scale features across layers may be reduced or even disappear, resulting in insufficient representation of forgery features by the network, thus affecting the judgment of forged frames and leading to deviations or even failures in temporal forgery localization.
[0004] 2. Existing methods only capture image tampering artifacts through the RGB color channel while ignoring the artifacts in the high-frequency region. Although the operations of deep forgery will destroy the correlation of the RGB color channel and the natural optical laws, with the continuous progress of deep forgery technology, the tampering artifacts in the RGB color channel become increasingly natural. In fact, existing deep forgery technologies still leave obvious forgery traces in the high-frequency region. Extracting such high-frequency features can reveal the statistical differences between forgery and authenticity. However, existing methods do not take this into account.
[0005] 3. Existing methods all involve multi-modalities, which means that it is necessary to fuse visual and auditory modal features. In fact, the importance of vision and audition in reflecting forgery facts at the same moment is different. However, existing methods cannot dynamically adjust the importance of each modal feature according to the specific situation of the input features during the multi-modal feature fusion process, which may lead to overemphasis or neglect of a certain modal feature during the modal fusion process, resulting in the model being unable to capture some key forgery features, thus affecting the judgment and localization of forged frames in the video. Summary of the Invention
[0006] The purpose of the present invention is to solve the problems of insufficient representation of fine-scale forgery features in existing detection methods and the inability of existing methods to effectively capture image forgery artifacts, and to provide a multi-modal deep forgery detection model for temporal forgery localization, which includes a feature extraction module, an adaptive multi-modal feature fusion module, a frame classification module, and a boundary localization module. The feature extraction module is constructed based on the ConvNeXt network, providing the model with excellent ability to capture fine forgery features from image frame sequences and audio streams. The adaptive multi-modal feature fusion module can better obtain multi-modal fusion feature representations by adjusting the fusion weights of different modal features. The frame classification module can generate frame-level authenticity predictions based on the feature representations. The boundary localization module can generate the final forgery boundary prediction through visual, audio, fusion feature representations, and frame classification predictions.
[0007] The present invention is realized through the following technical solutions: A multi-modal deepfake detection model for time forgery localization, comprising a feature extraction module, an adaptive multi-modal feature fusion module, a frame classification module, and a boundary localization module; The feature extraction module is mainly composed of a visual feature extraction module and an audio feature extraction module. The visual feature extraction module is constructed based on a ConvNeXt 3D visual encoder and is used to extract visual feature representations; the audio feature extraction module is constructed based on a ConvNeXt audio encoder and is used to extract audio feature representations; The adaptive multi-modal feature fusion module fuses the visual feature representation and the audio feature representation in the feature channel dimension by constructing an attention map to obtain a multi-modal fusion feature; The frame classification module is designed to process the visual feature representation and the audio feature representation obtained by the feature extraction module to obtain visual frame-level prediction labels and audio frame-level prediction labels; The boundary localization module is designed to process the visual feature representation and the audio feature representation obtained by the feature extraction module, the multi-modal fusion feature obtained by the adaptive multi-modal feature fusion module, and the visual frame-level prediction labels and audio frame-level prediction labels obtained by the frame classification module to obtain a final forged boundary prediction.
[0008] To better implement the multi-modal deepfake detection model for time forgery localization of the present invention, the following setting method is particularly adopted: When the visual feature extraction module extracts visual feature representations, it includes the following steps: 1.1.1) Use TorchVision to obtain an image frame sequence from the original video; 1.1.2) Preprocess the image frames: First, read the original RGB color channel information and convert it into a color channel tensor representation; then, obtain the high-frequency component information of the image through a designed high-frequency feature extraction module and convert it into a high-frequency component tensor representation; the high-frequency feature extraction module is composed of a high-pass filter based on the Laplacian of Gaussian operator; 1.1.3) Apply the ConvNeXt 3D network to extract the feature representations of the RGB color channel and the high-frequency components of the image respectively; the network structure of the ConvNeXt 3D network includes a total of four sequences, namely sequence one, sequence two, sequence three, and sequence four. Each sequence is composed of 3D ConvNeXt blocks, and the numbers are 3, 3, 9, and 3 respectively; 1.1.4) Fuse the feature representations of the RGB color channel and the high-frequency components of the image obtained in step 1.1.3) through an adaptive visual feature fusion module designed based on a weighted parameter matrix to obtain the final visual feature representation.
[0009] To better implement a multi-modal deepfake detection model for time forgery localization according to the present invention, the following setting method is specifically adopted: Apply the ConvNeXt 3D network to extract the feature representations of the RGB color channels and high-frequency components of the image, including the following specific steps: 1.1.3.1) Input the tensor representations of the color channels and high-frequency components respectively, and perform initial modeling on the image through an initial 3D convolutional layer to provide basic features for subsequent network layers; 1.1.3.2) Pass through a sequence one containing 3 3D ConvNeXt blocks, perform layer normalization, and add the input and output results of each block through residual connection; the final tensor shape remains unchanged; 1.1.3.3) Pass through a downsampling layer to reduce the feature space dimension; 1.1.3.4) Repeat steps 1.1.3.2) to 1.1.3.3) three times. Each time, sequence one is replaced by sequence two, sequence three, and sequence four respectively to further extract feature information; 1.1.3.5) After step 1.1.3.4), map the features to the specified feature representation through a fully connected layer and two convolutional layers to facilitate multi-modal fusion operation with subsequent audio features.
[0010] To better implement a multi-modal deepfake detection model for time forgery localization according to the present invention, the following setting method is specifically adopted: Step 1.1.4) includes the following specific steps: 1.1.4.1) Construct two tensor parameters with the same feature representations of the RGB color channels and high-frequency components obtained in step 1.1.3) respectively, and obtain two parameter matrices by training the two tensor parameters; 1.1.4.2) Perform dot product operations on the two tensor parameter matrices with the feature representations of the RGB color channels and high-frequency components respectively to obtain two weighted feature components; 1.1.4.3) Then add the two weighted feature components to obtain the final visual feature representation.
[0011] To better implement a multi-modal deepfake detection model for time forgery localization according to the present invention, the following setting method is specifically adopted: When the audio feature extraction module extracts audio feature representations, it includes the following steps: 1.2.1) Use TorchVision to obtain the original audio data from the original video; 1.2.2) Convert the original audio data into a mel spectrogram through short-time Fourier transform and mel filters; 1.2.3) Apply the ConvNeXt network to extract the feature representation of the audio; the network structure of the ConvNeXt network consists of three sequences: Sequence One, Sequence Two, and Sequence Three. Each sequence is composed of ConvNeXt blocks, and the numbers of blocks are 3, 3, and 3 respectively.
[0012] To better implement the multi-modal deepfake detection model for temporal forgery localization of the present invention, the following setting method is particularly adopted: Step 1.2.3) includes the following steps: 1.2.3.1) Input the Mel spectrogram (audio tensor), and perform initial modeling on the audio information through an initial convolutional layer to provide basic features for subsequent network layers; 1.2.3.2) Pass through Sequence One containing 3 ConvNeXt blocks, perform layer normalization, and add the input and output results of each block through residual connection; the final tensor shape remains unchanged; 1.2.3.3) Pass through a downsampling layer to reduce the feature space dimension; 1.2.3.4) Repeat steps 1.2.3.2) to 1.2.3.3) twice. Each time, Sequence One is replaced by Sequence Two and Sequence Three respectively to further extract feature information; 1.2.3.5) Map the features to the specified feature representation through a fully connected layer and a convolutional layer to obtain the audio feature representation, which is convenient for multi-modal fusion operation with the previously obtained visual features.
[0013] To better implement the multi-modal deepfake detection model for temporal forgery localization of the present invention, the following setting method is particularly adopted: The adaptive multi-modal feature fusion module obtains the multi-modal fusion feature through the following steps: 2.1) Concatenate the visual feature representation and the audio feature representation obtained by the feature extraction module along the channel dimension to obtain the original concatenated feature; 2.2) The original concatenated feature passes through a convolutional layer with a size of 1×1 and a batch normalization layer for normalization processing. The obtained result passes through an ReLU activation function, and then passes through a convolutional layer with a size of 3×3 and Sigmoid activation function to obtain an attention weight containing visual and audio channels; 2.3) The attention weight obtained in step 2.2) is replicated and concatenated in the channel dimension to make its size the same as the original concatenated feature obtained in step 2.1) to obtain the final attention feature map; 2.4) Perform weighted calculation on the original concatenated feature obtained in step 2.1) and the attention feature map obtained in step 2.3) to obtain the final multi-modal fusion feature of vision and audio.
[0014] To better implement a multi-modal deep fake detection model for time forgery localization according to the present invention, the following setting method is particularly adopted: The weighted calculation process in step 2.4) includes the following steps: 2.4.1) Perform a dot product operation on the original concatenated features and the attention feature map; 2.4.2) Perform an addition operation on the result obtained in step 2.4.1) with the original concatenated features; 2.4.3) Perform dimensionality reduction processing on the result obtained in step 2.4.2) through a convolutional layer with a size of 1×1 to obtain the final multi-modal fusion features of vision and audio.
[0015] To better implement a multi-modal deep fake detection model for time forgery localization according to the present invention, the following setting method is particularly adopted: The frame classification module processes the visual feature representation and the audio feature representation through the following steps to obtain the visual frame-level prediction label and the audio frame-level prediction label: 3.1) Perform dimensionality reduction processing on the visual feature representation obtained from the visual feature extraction module through a convolutional layer with a size of 1×1 to obtain a binary visual frame-level prediction label, and this label continues to participate in the subsequent boundary localization module processing; 3.2) Perform dimensionality reduction processing on the audio feature representation obtained from the audio feature extraction module through a convolutional layer with a size of 1×1 to obtain a binary audio frame-level prediction label, and this label continues to participate in the subsequent boundary localization module processing.
[0016] To better implement a multi-modal deep fake detection model for time forgery localization according to the present invention, the following setting method is particularly adopted: The boundary localization module obtains the final forged boundary prediction through the following processing steps: 4.1) Perform an addition operation on the visual feature representation and the visual frame-level prediction label, and then obtain the visual position-aware boundary map and the visual channel-aware boundary map through PRB processing; 4.2) Perform an addition operation on the audio feature representation and the audio frame-level prediction label, and then obtain the audio position-aware boundary map and the audio channel-aware boundary map through PRB processing; 4.3) Perform an addition operation on the visual frame-level prediction label, the audio frame-level prediction label, and the multi-modal fusion features, and then obtain the fusion position-aware boundary map and the fusion channel-aware boundary map through PRB processing; 4.4) Perform an addition operation on the position-aware boundary map and the channel-aware boundary map of the corresponding modality obtained in steps 4.1) to 4.3), and then aggregate them through a convolutional layer with a size of 2×2 to obtain the position-channel boundary map of the corresponding modality; 4.5) Finally, calculate the final boundary map using various boundary maps obtained in steps 4.1) to 4.4).
[0017] Compared with the prior art, the present invention has the following advantages and beneficial effects: (1) Aiming at the problem of insufficient representation of fine-scale forgery features in existing detection methods, the present invention proposes a feature encoder based on ConvNeXt. This encoder retains the locality advantage of the convolutional network through the strategies of depthwise separable convolution and group convolution, and enhances the ability to capture fine-scale features of images through more effective parameter allocation. In addition, compared with traditional convolutional networks, this encoder increases the network width by adding more convolutional kernel numbers in each layer, further improving the ability to capture fine-scale features of images. This encoder also inherits the skip connection design in the residual network and introduces a new global response normalization layer (GRN, Global Response Normalisation) and Gaussian error linear unit (GELU, Gaussian Error Linear Unit) to help maintain the stability of feature flow across layers and reduce the loss of features during transmission between layers.
[0018] (2) Aiming at the problem that existing methods cannot effectively capture image forgery artifacts, the present invention proposes a high-frequency feature extraction module to enhance the forgery feature representation of the visual modality. This module consists of a high-pass filter based on the Laplacian of Gaussian operator, which is specifically used to extract the high-frequency component information in the image and highlight its detailed features. In addition, a visual feature adaptive fusion module for image RGB color and high-frequency component features is also proposed. This module realizes pixel-level fusion of adaptive RGB color and high-frequency component features by introducing a weighted parameter matrix, and overall improves the model's ability to capture and represent visual forgery features.
[0019] (3) Aiming at the problem that the existing methods cannot adaptively adjust the importance of different modality features during the multi-modal feature fusion process, the present invention proposes an adaptive multi-modal feature fusion module, which improves the multi-modal feature fusion effect. This module first concatenates the audio features and visual features to form a preliminary fusion feature, and then generates a feature attention matrix through calculation to dynamically evaluate the importance of each modality feature. On this basis, the original features are weighted and combined with the attention matrix to obtain the final multi-modal fusion feature. This module can adaptively adjust the fusion weights of visual and audio features to ensure that the key forged feature information is fully retained, thereby significantly improving the accuracy and robustness of multi-modal feature fusion and providing a more reliable feature representation for temporal forgery localization. Brief Description of the Drawings
[0020] Figure 1 This is the overall architecture diagram of the multi-modal deepfake detection model described in the present invention.
[0021] Figure 2 This is the feature extraction module described in the present invention Figure 3 This is the filter kernel parameter diagram.
[0022] Figure 4 This is the adaptive visual feature fusion module.
[0023] Figure 5 This is the adaptive multi-modal feature fusion module.
[0024] In Figure 1 , is the average operation; is the soft non-maximum suppression.
[0025] In Figure 2 , is the sequence; is the block; is the 3D convolutional layer; is the linear layer; is the 2D convolutional layer; is the 1D convolutional layer; is the layer normalization; is the Gaussian error linear unit; is the global response normalization; is the rearrangement. Detailed Embodiments
[0026] The present invention will be further described in detail below in conjunction with embodiments, but the embodiments of the present invention are not limited thereto.
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents the selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0028] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.
[0029] Glossary: ConvNeXt 3D: The next generation of 3D convolutional networks.
[0030] ConvNeXt: The next generation of convolutional networks.
[0031] TorchVision: A computer vision toolkit.
[0032] 3D ConvNeXt: The next generation of 3D convolutional networks.
[0033] PRB: Proposal relationship block.
[0034] Example 1: The present invention designs a multi-modal deepfake detection model for temporal forgery localization, which solves the problems of insufficient representation of subtle-scale forgery features in existing detection methods and the inability of existing methods to effectively capture image forgery artifacts. It includes a feature extraction module, an adaptive multi-modal feature fusion module, a frame classification module, and a boundary localization module; The feature extraction module is mainly composed of a visual feature extraction module and an audio feature extraction module. The visual feature extraction module is constructed based on the visual encoder of ConvNeXt 3D and is used to extract visual feature representations; the audio feature extraction module is constructed based on the audio encoder of ConvNeXt and is used to extract audio feature representations; The adaptive multi-modal feature fusion module fuses the visual feature representation and the audio feature representation in the feature channel dimension by constructing an attention map to obtain a multi-modal fusion feature; The frame classification module is designed to process the visual feature representation and the audio feature representation obtained by the feature extraction module to obtain visual frame-level prediction labels and audio frame-level prediction labels; The boundary localization module is designed to process the visual feature representation and the audio feature representation obtained by the feature extraction module, the multi-modal fusion feature obtained by the adaptive multi-modal feature fusion module, and the visual frame-level prediction labels and audio frame-level prediction labels obtained by the frame classification module to obtain a final forged boundary prediction.
[0035] Embodiment 2: This embodiment is further optimized on the basis of the above embodiment. The same parts as the foregoing technical solutions will not be described herein again. To better implement a multi-modal deep fake detection model for temporal forgery localization according to the present invention, the following setting method is specifically adopted: When the visual feature extraction module extracts the visual feature representation, it includes the following steps: 1.1.1) Use TorchVision to obtain an image frame sequence from the original video; 1.1.2) Preprocess the image frames: First, read the original RGB color channel information and convert it into a color channel tensor representation; then obtain the high-frequency component information of the image through a designed high-frequency feature extraction module and convert it into a high-frequency component tensor representation; the high-frequency feature extraction module is composed of a high-pass filter based on the Laplacian of Gaussian operator; 1.1.3) Apply the ConvNeXt 3D network to extract the feature representations of the RGB color channel and the high-frequency components of the image respectively; the network structure of the ConvNeXt 3D network includes a total of four sequences, namely sequence one, sequence two, sequence three, and sequence four. Each sequence is composed of 3D ConvNeXt blocks, and the numbers are 3, 3, 9, and 3 respectively; 1.1.4) Fuse the feature representations of the RGB color channel and the high-frequency components of the image obtained in step 1.1.3) through an adaptive visual feature fusion module designed based on a weighted parameter matrix to obtain the final visual feature representation.
[0036] Embodiment 3: This embodiment is further optimized on the basis of any of the above embodiments. The same parts as the foregoing technical solutions will not be described herein again. To better implement a multi-modal deep fake detection model for temporal forgery localization according to the present invention, the following specific steps are adopted for applying the ConvNeXt 3D network to extract the feature representations of the RGB color channel and the high-frequency components of the image respectively: 1.1.3.1) Input the color channel tensor representation and the high-frequency component tensor representation respectively. After an initial 3D convolutional layer, perform initial modeling on the image to provide basic features for subsequent network layers. 1.1.3.2) Pass through Sequence One containing 3 3D ConvNeXt blocks, perform layer normalization, and add the input and output results of each block through residual connection; the final tensor shape remains unchanged. 1.1.3.3) Pass through a downsampling layer to reduce the dimensionality of the feature space. 1.1.3.4) Repeat steps 1.1.3.2) to 1.1.3.3) three times. Each time, replace Sequence One with Sequence Two, Sequence Three, and Sequence Four respectively to further extract feature information. 1.1.3.5) After step 1.1.3.4), map the features to the specified feature representation through a fully connected layer and two convolutional layers to facilitate multimodal fusion operations with subsequent audio features.
[0037] Example 4: This example is further optimized based on any of the above examples. The same parts as the previous technical solutions will not be elaborated here. To better implement a multimodal deepfake detection model for time forgery localization of the present invention, the following setting method is particularly adopted: Step 1.1.4) includes the following specific steps: 1.1.4.1) Construct two tensor parameters with the same feature representations of the RGB color channel and high-frequency components obtained in step 1.1.3) respectively, and obtain two parameter matrices through training. 1.1.4.2) Perform dot product operations on the two parameter matrices with the feature representations of the RGB color channel and high-frequency components respectively to obtain two weighted feature components. 1.1.4.3) Then add the two weighted feature components to obtain the final visual feature representation.
[0038] Example 5: This example is further optimized based on any of the above examples. The same parts as the previous technical solutions will not be elaborated here. To better implement a multimodal deepfake detection model for time forgery localization of the present invention, the following setting method is particularly adopted: When the audio feature extraction module extracts audio feature representations, it includes the following steps: 1.2.1) Use TorchVision to obtain the original audio data from the original video. 1.2.2) Convert the original audio data into a mel spectrogram through short-time Fourier transform and mel filters. 1.2.3) Apply the ConvNeXt network to extract the feature representation of the audio; the network structure of the ConvNeXt network includes a total of three sequences, namely Sequence One, Sequence Two, and Sequence Three. Each sequence is composed of ConvNeXt blocks, and the numbers are 3, 3, and 3 respectively.
[0039] Example 6: This example further optimizes on the basis of any of the above examples. The same parts as the foregoing technical solutions will not be elaborated here. Further, to better implement a multi-modal deep fake detection model for time forgery localization described in the present invention, the following setting method is specifically adopted: Step 1.2.3) includes the following steps: 1.2.3.1) Input the Mel spectrogram (audio tensor), and perform initial modeling on the audio information through an initial convolutional layer to provide basic features for subsequent network layers; 1.2.3.2) Pass through Sequence One containing 3 ConvNeXt blocks, perform layer normalization, and add the input and output results of each block by means of residual connection; the final tensor shape remains unchanged; 1.2.3.3) Pass through a downsampling layer to reduce the feature space dimension; 1.2.3.4) Repeat steps 1.2.3.2) to 1.2.3.3) twice. Each time, Sequence One is replaced by Sequence Two and Sequence Three respectively to further extract feature information; 1.2.3.5) Map the features to the specified feature representation through a fully connected layer and a convolutional layer to obtain the audio feature representation, which is convenient for multi-modal fusion operation with the previously obtained visual features.
[0040] Example 7: This example further optimizes on the basis of any of the above examples. The same parts as the foregoing technical solutions will not be elaborated here. Further, to better implement a multi-modal deep fake detection model for time forgery localization described in the present invention, the following setting method is specifically adopted: The adaptive multi-modal feature fusion module obtains the multi-modal fusion feature through the following steps: 2.1) Concatenate the visual feature representation and the audio feature representation obtained by the feature extraction module along the channel dimension to obtain the original concatenated feature; 2.2) The original concatenated feature passes through a convolutional layer with a size of 1×1 and a batch normalization layer for normalization processing. The obtained result passes through an ReLU activation function, and then passes through a convolutional layer with a size of 3×3 and an Sigmoid activation function to obtain an attention weight containing visual and audio channels; 2.3) The attention weights obtained in step 2.2) are replicated and concatenated in the channel dimension to make their size the same as the original concatenated features obtained in step 2.1), resulting in the final attention feature map; 2.4) Weighted calculation is performed on the original concatenated features obtained in step 2.1) and the attention feature map obtained in step 2.3) to obtain the final multi-modal fusion features of vision and audio.
[0041] Example 8: This embodiment is a further optimization based on any of the above embodiments. The same parts as the foregoing technical solutions will not be described herein again. To better implement a multi-modal deepfake detection model for time forgery localization described in the present invention, the following setting method is particularly adopted: The weighted calculation process in step 2.4) includes the following steps: 2.4.1) Perform a dot product operation on the original concatenated features and the attention feature map; 2.4.2) Perform an addition operation on the result obtained in step 2.4.1) with the original concatenated features again; 2.4.3) After the result obtained in step 2.4.2) is further dimension-reduced through a convolutional layer with a size of 1×1, the final multi-modal fusion features of vision and audio are obtained.
[0042] Example 9: This embodiment is a further optimization based on any of the above embodiments. The same parts as the foregoing technical solutions will not be described herein again. To better implement a multi-modal deepfake detection model for time forgery localization described in the present invention, the following setting method is particularly adopted: The frame classification module processes the visual feature representation and the audio feature representation through the following steps to obtain the visual frame-level prediction label and the audio frame-level prediction label: 3.1) Dimension-reduce the visual feature representation obtained from the visual feature extraction module through a convolutional layer with a size of 1×1 to obtain a binary-class visual frame-level prediction label, and this label continues to participate in the subsequent boundary localization module processing process; 3.2) Dimension-reduce the audio feature representation obtained from the audio feature extraction module through a convolutional layer with a size of 1×1 to obtain a binary-class audio frame-level prediction label, and this label continues to participate in the subsequent boundary localization module processing process.
[0043] Example 10: This embodiment is further optimized on the basis of any of the above embodiments. The same parts as the foregoing technical solutions will not be described herein again. To better implement a multi-modal deep fake detection model for time forgery localization of the present invention, the following setting method is specifically adopted: The boundary localization module obtains the final forged boundary prediction through the following processing steps: 4.1) Perform an addition operation on the visual feature representation and the visual frame-level prediction label, and then PRB obtain a visually position-aware boundary map and a visually channel-aware boundary map after processing; 4.2) Perform an addition operation on the audio feature representation and the audio frame-level prediction label, and then PRB obtain an audio position-aware boundary map and an audio channel-aware boundary map after processing; 4.3) Perform an addition operation on the visual frame-level prediction label, the audio frame-level prediction label, and the multi-modal fusion feature, and then PRB obtain a fusion position-aware boundary map and a fusion channel-aware boundary map after processing; 4.4) Perform an addition operation on the corresponding modality's position-aware boundary map and channel-aware boundary map obtained in steps 4.1) to 4.3), and then aggregate through a 2×2 convolutional layer to obtain the corresponding modality's position-channel boundary map; 4.5) Finally, calculate the final boundary map using various boundary maps obtained in steps 4.1) to 4.4).
[0044] Embodiment 11: A multi-modal deep fake detection model for time forgery localization, combined with Figures 1 - 5 as shown, includes a feature extraction module, an adaptive multi-modal feature fusion module, a frame classification module, and a boundary localization module.
[0045] Combined with Figure 2 as shown, the feature extraction module is mainly composed of a visual feature extraction module constructed by a ConvNeXt 3D-based visual encoder and an audio feature extraction module constructed by a ConvNeXt-based audio encoder.
[0046] The visual feature extraction module is used to extract visual feature representations , including the following steps: S1. Use TorchVision to obtain an image frame sequence from the original video .
[0047] S2. Preprocess the image frames: First, read the original RGB color channel information and convert it into a color channel tensor representation ; Then, the high-frequency component (HF) information of the image is obtained through the designed high-frequency feature extraction module and converted into a high-frequency component tensor representation. ; The high-frequency feature extraction module consists of a high-pass filter based on the Laplacian of Gaussian operator, and the parameters of the filter kernel are as Figure 3 shown.
[0048] S3. Apply the ConvNeXt 3D network to extract the feature representations of the RGB color channels of the image and the feature representation of the high-frequency component ; It includes the following specific steps: S3.1 Input the color channel tensor representation and the high-frequency component tensor representation , and perform initial modeling on the image through an initial 3D convolutional layer to provide basic features for subsequent network layers; S3.2) Pass through a sequence one containing 3 3D ConvNeXt blocks, perform layer normalization, and add the input and output results of each block through residual connection; the final tensor shape remains unchanged; S3.3) Pass through a downsampling layer to reduce the feature space dimension; S3.4) Repeat steps S3.2) to S3.3) three times. Each time, sequence one is replaced by sequence two, sequence three, and sequence four respectively to further extract feature information; S3.5) After step S3.4), map the features to the specified feature representation through a fully connected layer and two convolutional layers to obtain the feature representation of the RGB color channels of the image and the feature representation of the high-frequency component , which is convenient for multi-modal fusion operation with subsequent audio features.
[0049] Among them, the network structure of the ConvNeXt 3D network includes a total of four sequences: sequence one, sequence two, sequence three, and sequence four. Each sequence is composed of 3D ConvNeXt blocks, and the numbers are 3, 3, 9, and 3 respectively.
[0050] S4. Combine Figure 4 as shown, and fuse the feature representation of the RGB color channels of the image and the feature representation of the high-frequency component through the adaptive visual feature fusion module designed based on the weighted parameter matrix to obtain the final visual feature representation ; It includes the following specific steps: S4.1. Construct two feature representations of the RGB color channels obtained in the same step as S3 and the feature representation of the high-frequency component Two identical tensor parameters and , two tensor parameters and are used to obtain two parameter matrices (RGB color image weighting parameter matrix) and (high-frequency image weighting parameter matrix) through training. The calculation formula is as follows: ; S4.2. The two parameter matrices and are respectively dot-product operated with the feature representation of the RGB color channel and the feature representation of the high-frequency component to obtain two weighted feature components; S4.3. Then, the two weighted feature components are added together to obtain the final visual feature representation .
[0051] An audio feature extraction module for extracting an audio feature representation ; including the following steps: A1. Use TorchVision to obtain the original audio data from the original video .
[0052] A2. Convert the original audio data into a Mel spectrogram through short-time Fourier transform and Mel filter .
[0053] A3. Apply the ConvNeXt network to extract the feature representation of the audio; including the following steps: A3.1. Input the Mel spectrogram (audio tensor) , and perform initial modeling of the audio information through an initial convolutional layer to provide basic features for subsequent network layers; A3.2. Pass through a sequence one containing 3 ConvNeXt blocks, perform layer normalization processing, and add the input and output results of each block through residual connection; the final tensor shape remains unchanged; A3.3. Pass through a downsampling layer to reduce the feature space dimension; A3.4. Repeat steps A3.2 to A3.3 twice. Each time, sequence one is replaced by sequence two and sequence three respectively to further extract feature information; A3.5. Map the features to the specified feature representation through a fully connected layer and a convolutional layer to obtain the audio feature representation , which is convenient for multi-modal fusion operation with the previously obtained visual features.
[0054] Among them, the network structure of the ConvNeXt network includes a total of three sequences: Sequence One, Sequence Two, and Sequence Three. Each sequence is composed of ConvNeXt blocks, and the numbers are 3, 3, and 3 respectively.
[0055] Combined with Figure 5 As shown, the adaptive multi-modal feature fusion module fuses the visual feature representation and the audio feature representation in the feature channel dimension by constructing an attention map and the audio feature representation to obtain the multi-modal fusion feature ; when performing the fusion, the following steps are included: B1. Concatenate the visual feature representation and the audio feature representation obtained by the feature extraction module along the channel dimension to obtain the original concatenated feature .
[0056] B2. The original concatenated feature is normalized through a convolutional layer with a size of 1×1 and a batch normalization layer, and the obtained result passes through a ReLU activation function, and then through a convolutional layer with a size of 3×3 and Sigmoid the processing of the activation function to obtain an attention weight containing visual and audio channels , and the calculation formula is as follows: ; where Z is the enhanced intermediate feature representation.
[0057] ReLU (Rectified Linear Unit): A piecewise linear activation function, defined as , used to introduce non-linearity.
[0058] BN (Batch Normalization): A normalization technique that normalizes the input of each layer of the neural network. Calculate the mean and variance for a small batch of data (mini-batch), and normalize the input to a distribution with a mean of 0 and a variance of 1.
[0059] Conv 1×1 (1×1 convolution): An operation that uses a 1×1 convolution kernel to extract local spatial features, capturing local patterns (such as edges, textures) in an image or feature map.
[0060] Sigmoid: A non-linear activation function that maps the input to the interval (0,1), defined as , and the output can be interpreted as a probability, often used in binary classification problems.
[0061] Conv3×3 The operation of using a 3×3 convolutional kernel to extract local spatial features (3×3 convolution) captures local patterns (such as edges and textures) in an image or feature map.
[0062] B3. The attention weights obtained via step B2 Perform replication and splicing in the channel dimension to make its size the same as the original spliced features obtained in step B1 to obtain the final attention feature map. . The calculation formula is as follows: , where C is the number of repetitions of the attention weights for audio-visual channel splicing, is the attention weight of the audio channel, A v is the attention weight of the video channel.
[0063] B4. For the original spliced features obtained via step B1 and the attention feature map obtained in step 2.3) perform weighted calculation to obtain the final multi-modal fusion features of vision and audio .
[0064] In step B4, the weighted calculation process includes the following steps: B4.1. Perform a dot product operation on the original spliced features and the attention feature map ; B4.2. Perform an addition operation on the result obtained in step B4.1 with the original spliced features again; B4.3. After reducing the dimension of the result obtained in step B4.2 via a 1×1 convolutional layer, obtain the final multi-modal fusion features of vision and audio .
[0065] The frame classification module (including the visual frame classification module and the audio frame classification module) aims to process the visual feature representation and the audio feature representation obtained by the feature extraction module to obtain visual frame-level prediction labels and audio frame-level prediction labels; when processing, it includes the following specific steps: C1. Reduce the dimension of the visual feature representation obtained from the visual feature extraction module through a 1×1 convolutional layer to obtain a binary-class visual frame-level prediction label , and this label continues to participate in the subsequent boundary localization module processing process.
[0066] C2. Perform dimensionality reduction on the audio features obtained from the audio feature extraction module through a 1×1 convolutional layer to obtain a frame-level prediction label for a binary classification audio which continues to participate in the subsequent boundary localization module processing.
[0067] The boundary localization module (including the visual boundary matching layer, the fusion boundary matching layer, and the audio boundary matching layer) aims to process the visual feature representation and audio feature representation obtained from the feature extraction module, the multi-modal fusion feature obtained from the adaptive multi-modal feature fusion module, and the frame-level prediction label for vision and the frame-level prediction label for audio
[0068] to obtain the final forged boundary prediction. When the boundary localization module processes to obtain the final forged boundary prediction it includes the following specific steps: D1. Perform an addition operation on the visual feature representation and the frame-level prediction label for vision, and then obtain the visual position-aware boundary map PRB and the visual channel-aware boundary map after processing, and the calculation process is as follows:
[0069] D2. Perform an addition operation on the audio feature representation and the frame-level prediction label for audio, and then obtain the audio position-aware boundary map and the audio channel-aware boundary map PRB after processing, and the calculation process is as follows:
[0070] D3. Perform an addition operation on the frame-level prediction label for vision and the frame-level prediction label for audio and the multi-modal fusion feature and then obtain the fusion position-aware boundary map and the fusion channel-aware boundary map after processing by PRB, and the calculation process is as follows:
[0071] D4. Add the position-aware boundary maps and channel-aware boundary maps obtained in steps D1 to D3, and then aggregate them through a 2×2 convolutional layer to obtain the position-channel boundary map of the corresponding modality . The calculation process is as follows: ; D5. Finally, calculate the final forged boundary prediction using various boundary maps obtained in steps D1 to D4. The calculation formula is as follows:
[0072] where represents the weights of different boundary maps, which are obtained by stacking the features of different modalities and the boundary maps through a 1×1 convolutional layer and then performing an average operation. The calculation formula is as follows: . Among them, is the video boundary map; is the audio boundary map, ; For the inference stage, use soft non-maximum suppression (S-NMS) to remove duplicate prediction results and eliminate prediction redundancy. The finally obtained boundary map will be used as the inference result of the model.
[0073] As described above, it is only a preferred embodiment of the present invention, and does not impose any formal restrictions on the present invention. Any simple modification or equivalent change made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A multi-modal deepfake detection model for time forgery localization, characterized in that: It includes a feature extraction module, an adaptive multi-modal feature fusion module, a frame classification module, and a boundary localization module; The feature extraction module is mainly composed of a visual feature extraction module and an audio feature extraction module. The visual feature extraction module is constructed based on a ConvNeXt 3D visual encoder and is used to extract visual feature representations; the audio feature extraction module is constructed based on a ConvNeXt audio encoder and is used to extract audio feature representations; The adaptive multi-modal feature fusion module fuses the visual feature representation and the audio feature representation in the feature channel dimension by constructing an attention map to obtain a multi-modal fusion feature; The frame classification module processes the visual feature representation and the audio feature representation obtained by the feature extraction module to obtain visual frame-level prediction labels and audio frame-level prediction labels; The boundary localization module processes the visual feature representation and the audio feature representation obtained by the feature extraction module, the multi-modal fusion feature obtained by the adaptive multi-modal feature fusion module, and the visual frame-level prediction labels and audio frame-level prediction labels obtained by the frame classification module to obtain a final forged boundary prediction.
2. The multimodal deepfake detection model for time forgery localization according to claim 1, wherein: When the visual feature extraction module extracts visual feature representations, it includes the following steps: 1.1.1) Use TorchVision to obtain an image frame sequence from the original video; 1.1.2) Preprocess the image frames: First, read the original RGB color channel information and convert it into a color channel tensor representation; then obtain the high-frequency component information of the image through a designed high-frequency feature extraction module and convert it into a high-frequency component tensor representation; the high-frequency feature extraction module consists of a high-pass filter based on the Laplacian of Gaussian operator; 1.1.3) Apply the ConvNeXt 3D network to extract the feature representations of the RGB color channels and high-frequency components of the image respectively; the network structure of the ConvNeXt 3D network includes a total of four sequences: Sequence One, Sequence Two, Sequence Three, and Sequence Four. Each sequence is composed of 3D ConvNeXt blocks, and the numbers are 3, 3, 9, and 3 respectively; 1.1.4) Pass the result of step 1.1.3) through an adaptive visual feature fusion module designed based on a weighted parameter matrix to fuse the feature representations of the RGB color channels and high-frequency components of the image to obtain the final visual feature representation.
3. The multimodal deepfake detection model for time forgery localization according to claim 2, characterized in that: Applying the ConvNeXt 3D network to extract the feature representations of the RGB color channels and high-frequency components of the image respectively includes the following specific steps: 1.1.3.1) Input the color channel tensor representation and the high-frequency component tensor representation respectively, and perform initial modeling on the image through an initial 3D convolutional layer; 1.1.3.2) Pass through Sequence One containing 3 3D ConvNeXt blocks, perform layer normalization processing, and add the input and output results of each block through residual connection; 1.1.3.3) Pass through a downsampling layer to reduce the feature space dimension; 1.1.3.4) Repeat steps 1.1.3.2) to 1.1.3.3) three times. Each time when repeating, Sequence One is replaced by Sequence Two, Sequence Three, and Sequence Four respectively; 1.1.3.5) After step 1.1.3.4), map the features to the specified feature representation through a fully connected layer and two convolutional layers.
4. The multimodal deepfake detection model for time forgery localization according to claim 2, wherein: Step 1.1.4) includes the following specific steps: 1.1.4.1) Construct two tensor parameters with the same feature representations of the RGB color channels and high-frequency components obtained in two steps 1.1.3) respectively, and obtain two parameter matrices by training the two tensor parameters; 1.1.4.2) Perform dot product operations on the two parameter matrices and the feature representations of the RGB color channels and high-frequency components respectively to obtain two weighted feature components; 1.1.4.3) Then add the two weighted feature components to obtain the final visual feature representation.
5. A multimodal deepfake detection model for time forgery localization according to claim 1 or 2 or 3 or 4, characterized in that: When the audio feature extraction module extracts audio feature representations, it includes the following steps: 1.2.1) Use TorchVision to obtain the original audio data from the original video; 1.2.2) Convert the original audio data into a mel spectrogram through short-time Fourier transform and mel filters; 1.2.3) Apply the ConvNeXt network to extract the feature representation of the audio; the network structure of the ConvNeXt network includes a total of three sequences, namely Sequence One, Sequence Two, and Sequence Three. Each sequence is composed of ConvNeXt blocks, and the numbers are 3, 3, and 3 respectively.
6. The multimodal deepfake detection model for time forgery localization according to claim 5, characterized in that: Step 1.2.3) includes the following steps: 1.2.3.1) Input the mel spectrogram and perform initial modeling of the audio information through an initial convolutional layer; 1.2.3.2) Pass through Sequence One containing 3 ConvNeXt blocks, perform layer normalization processing, and add the input and output results of each block through the residual connection method; 1.2.3.3) Pass through a downsampling layer to reduce the feature space dimension; 1.2.3.4) Repeat steps 1.2.3.2) to 1.2.3.3) twice. Each time when repeating, Sequence One is replaced by Sequence Two and Sequence Three respectively; 1.2.3.5) Map the features to the specified feature representation through a fully connected layer and a convolutional layer to obtain the audio feature representation.
7. A multi-modal deepfake detection model for time forgery localization according to claim 1 or 2 or 3 or 4, characterized in that: The adaptive multi-modal feature fusion module obtains the multi-modal fusion feature through the following steps: 2.1) Concatenate the visual feature representation and the audio feature representation obtained by the feature extraction module along the channel dimension to obtain the original concatenated feature; 2.2) The original splicing features are normalized through a 1×1 convolutional layer and a batch normalization layer, and the obtained result passes through a ReLU activation function, and then through a 3×3 convolutional layer and Sigmoid activation function to obtain an attention weight containing visual and audio channels; 2.3) The attention weights obtained in step 2.2) are replicated and concatenated in the channel dimension to make its size consistent with the original concatenated feature obtained in step 2.1) to obtain the final attention feature map; 2.4) Perform weighted calculation on the original concatenated feature obtained in step 2.1) and the attention feature map obtained in step 2.3) to obtain the final visual and audio multi-modal fusion feature.
8. The multimodal deepfake detection model for time forgery localization according to claim 7, characterized in that: The weighted calculation process in step 2.4) includes the following steps: 2.4.1) Perform dot product operation on the original concatenated feature and the attention feature map; 2.4.2) Perform an addition operation on the result obtained in step 2.4.1) and the original stitching feature; 2.4.3) After reducing the dimension of the result obtained in step 2.4.2) through a 1×1 convolutional layer, the final visual and audio multi-modal fusion feature is obtained.
9. A multi-modal deepfake detection model for time forgery localization according to claim 1 or 2 or 3 or 4, characterized in that: The frame classification module processes the visual feature representation and the audio feature representation through the following steps to obtain the visual frame-level prediction label and the audio frame-level prediction label: 3.1) Reduce the dimension of the visual feature representation obtained from the visual feature extraction module through a 1×1 convolutional layer to obtain a binary-class visual frame-level prediction label; 3.2) Reduce the dimension of the audio feature representation obtained from the audio feature extraction module through a 1×1 convolutional layer to obtain a binary-class audio frame-level prediction label.
10. A multi-modal deepfake detection model for time forgery localization according to claim 1 or 2 or 3 or 4, characterized in that: The boundary localization module obtains the final forged boundary prediction through the following processing steps: 4.1) Perform an addition operation on the visual feature representation and the visual frame-level prediction label, and then obtain the visual position-aware boundary map and the visual channel-aware boundary map after being processed by PRB; 4.2) Perform an addition operation on the audio feature representation and the audio frame-level prediction label, and then obtain the audio position-aware boundary map and the audio channel-aware boundary map after being processed by PRB; 4.3) Perform an addition operation on the visual frame-level prediction label, the audio frame-level prediction label, and the multi-modal fusion feature, and then obtain the fusion position-aware boundary map and the fusion channel-aware boundary map after being processed by PRB; 4.4) Perform an addition operation on the corresponding modality's position-aware boundary map and channel-aware boundary map obtained in steps 4.1) to 4.3), and then aggregate them through a 2×2 convolutional layer to obtain the corresponding modality's position-channel boundary map; 4.5) Finally, calculate the final boundary map using various boundary maps obtained in steps 4.1) to 4.4).
Citation Information
Patent Citations
Training method of video scene boundary detection model and scene boundary detection method
CN116128043A
Video crowd counting method based on double-branch space-time interaction network
CN118781553A
Dynamic Memory Network
US20170024645A1
Cited By
Multi-modal forged video detection method based on multi-head addition cross attention mechanism
CN120635786A
Video time positioning method and device, and electronic equipment
CN121482679A