Deep forgery detection method and system based on audio and video time domain fusion
By using the audio-visual time domain fusion method in deep forgery detection, and using the Transformer encoder for self-attention learning and multi-scale time domain fusion, the problem that a single mode detection method is difficult to capture the discontinuity between audio-visual modes is solved, and the accuracy and robustness of the detection are improved.
Patent Information
- Application Number
- CN202510161275.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing single-modal depth forgery detection methods are difficult to capture discontinuity and forgery traces between audio and video modes, and when faced with multimodal depth forgery, the detection accuracy and robustness are affected.
Deep forgery detection method based on audio and video time domain fusion is adopted, video and audio features are extracted through residual networks and feedforward networks, and self-attention learning and multi-scale time domain fusion are used to capture the discontinuity between audio and video modes.
The accuracy and robustness of deep forgery detection are improved, and the discontinuity of deep forgery videos can be more effectively captured between audio and video modes, and the detection ability of multimodal depth forgery is enhanced.
Smart Images

Figure CN119992422A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of digital security and multimedia digital forensics technology, and in particular to a deep fake detection method and system based on audio and video time domain fusion. Background Art
[0002] The combination of artificial intelligence technology and multimedia synthesis technology has enabled deep fakes to develop rapidly in the fields of video and audio, and the visual effects of deep fake videos have become more realistic. However, this has also brought major security risks to network security and even social stability. At present, deep fake technologies such as text-to-speech (TTS), voice conversion (VC), face swapping, and face reenactment can efficiently generate high-quality deep fake media files in specific scenarios, further exacerbating the phenomenon of media file fraud in cyberspace and seriously damaging the public's trust in the authenticity of media files.
[0003] In order to deal with the potential threat of deep fake media to cyberspace and social public opinion, a large number of detection methods for single-modal media files (such as audio and video) have been proposed and achieved good performance. The methods for video single-modal deep fake detection can be divided into texture feature-based methods and semantic feature-based methods. Texture feature-based methods detect video single-modal deep fake by capturing discriminative frame-level features such as face blending boundaries and tampering traces in the frequency domain. Semantic feature-based methods are committed to capturing facial movement or discontinuity of adjacent frames in the time domain for video single-modal deep fake detection. For audio single-modal deep fake detection, most existing audio single-modal deep fake detection methods extract audio features such as mel-scale frequency cepstral coefficients (MFCC), short time Fourier transform (STFT) and constant q cepstral coefficients (CQCC) from the original waveform signal to detect the authenticity of audio media files.
[0004] However, high-quality deep fakes usually forge and tamper with both video and audio, thereby obtaining more realistic false information or fake news. Deep fake detection methods targeting a single modality are usually affected in detection accuracy and robustness when dealing with this multimodal deep fake technology. In addition, single-modal deep fake detection methods also find it difficult to capture discontinuities or traces of forgery between modalities, and it is difficult to fully utilize the tampering traces left by multimodal deep fake technology in forged videos. How to deal with this increasingly complex deep fake scenario is an urgent problem that technicians in the field of deep fake detection need to solve.
[0005] Therefore, there is an urgent need for a targeted deep fake detection method and system based on audio and video time domain fusion. Summary of the invention
[0006] The purpose of the present invention is to provide a deep fake detection method based on audio and video time domain fusion. By capturing the discontinuity of audio and video modalities in the time domain, the tampering traces of deep fake videos can be fully explored, thereby improving the accuracy and robustness of the deep fake detection method.
[0007] In a first aspect, the present application provides a deep fake detection method based on audio and video time domain fusion, the method comprising:
[0008] Preprocessing of the video to be detected: Frame division and face area extraction of the video to be detected, and audio extraction at the same time;
[0009] Video modality feature extraction: Use residual networks to learn deep fake representations of video modalities and extract video features with strong representation capabilities;
[0010] Audio modality feature extraction: Use a feedforward network to perform deep fake representation learning on the audio modality and extract audio features with strong representation capabilities;
[0011] Self-attention learning: Contains three Transformer-based self-attention branches with shared weights. Two of the branches extract deep fake features from video features and audio features respectively, and the other branch merges the outputs of the video modality feature extraction and audio modality feature extraction steps to jointly learn the audio and video modalities and capture the discontinuity of deep fake videos between multiple modalities.
[0012] Multi-scale time-domain fusion: The deep video forgery features, deep audio forgery features, and discontinuity features extracted based on Transformer self-attention are merged, and the multi-scale time-domain convolutional network is used for feature fusion. The final decision result is obtained through the time-domain pooling layer and the linear layer.
[0013] During the model training process, cross entropy loss is introduced to improve the model's detection ability by shortening the distance between the model's prediction results and the labeled labels.
[0014] In a second aspect, the present application provides a deep fake detection system based on audio and video time domain fusion, the system comprising:
[0015] Preprocessing module for the video to be detected: used to perform frame segmentation and face area extraction operations on the video to be detected, and perform audio extraction at the same time;
[0016] Video modality feature extraction module: used to use residual networks to learn deep fake representations of video modalities and extract video features with strong representation capabilities;
[0017] Audio modality feature extraction module: used to perform deep fake representation learning on the audio modality using a feedforward network and extract audio features with strong representation capabilities;
[0018] Self-attention module: It contains three Transformer-based self-attention branches with shared weights. Two of them mine deep fake features from video features and audio features respectively, and the other branch merges the outputs of the video modality feature extraction and audio modality feature extraction modules to jointly learn the audio and video modalities and capture the discontinuity of deep fake videos between multiple modalities.
[0019] Multi-scale time-domain fusion module: used to merge the deep video forgery features, deep audio forgery features, and discontinuity features between multiple modalities mined by the Transformer-based self-attention module, use a multi-scale time-domain convolutional network for feature fusion, and obtain the final decision result through the time-domain pooling layer and the linear layer;
[0020] During the model training process, cross entropy loss is introduced to improve the model's detection ability by shortening the distance between the model's prediction results and the labeled labels.
[0021] In a third aspect, the present application provides a deep fake detection system based on audio and video time domain fusion, the system comprising a processor and a memory:
[0022] The memory is used to store program code and transmit the program code to the processor;
[0023] The processor is used to execute any one of the four possible methods of the first aspect according to the instructions in the program code.
[0024] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to be executed by a processor to implement any one of the four possible methods in the first aspect.
[0025] Beneficial Effects
[0026] The present invention provides a deep fake detection method and system based on audio and video time domain fusion, which utilizes a self-supervised feature extractor to capture the discontinuity between audio and video modes, firstly uses a residual network to extract video features, uses a feedforward network to extract audio features, and then uses a Transformer encoder to fuse the merged video features and audio features; at the same time, the video features and audio features are also respectively input into the Transformer encoder for high-dimensional feature extraction; finally, the video features, audio features, and audio and video fusion features obtained by the Transformer encoder are merged together, deep feature fusion is performed through a multi-scale time domain convolutional network, and the final detection result is obtained through a time domain pooling layer and a linear layer, thereby overcoming the problem that the single-modality deep fake detection method of the prior art is difficult to capture the discontinuity or forgery traces between modes, and is difficult to fully utilize the tampering traces left in the forged video by the multi-modal deep fake technology.
[0027] The method and system of the present invention have the following advantages and effects:
[0028] The time-domain convolutional network is used to fuse deep fake traces of video and audio modalities and to explore the discontinuities of deep fake videos in video and audio modalities.
[0029] Multi-scale convolution kernels are used to capture the feature inconsistencies of video and audio modalities at different time lengths, thereby enhancing the accuracy and robustness of model detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0031] Figure 1 is a flow chart of the method of the present invention;
[0032] Figure 2 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION
[0033] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0034] This application provides a deep fake detection method based on audio and video time domain fusion, such as Figure 1As shown, the method includes:
[0035] Preprocessing of the video to be detected: Frame division and face area extraction of the video to be detected, and audio extraction at the same time;
[0036] Video modality feature extraction: Use residual networks to learn deep fake representations of video modalities and extract video features with strong representation capabilities;
[0037] Audio modality feature extraction: Use a feedforward network to perform deep fake representation learning on the audio modality and extract audio features with strong representation capabilities;
[0038] Self-attention learning: Contains three Transformer-based self-attention branches with shared weights. Two of the branches extract deep fake features from video features and audio features respectively, and the other branch merges the outputs of the video modality feature extraction and audio modality feature extraction steps to jointly learn the audio and video modalities and capture the discontinuity of deep fake videos between multiple modalities.
[0039] Multi-scale time-domain fusion: The deep video forgery features, deep audio forgery features, and discontinuity features extracted based on Transformer self-attention are merged, and the multi-scale time-domain convolutional network is used for feature fusion. The final decision result is obtained through the time-domain pooling layer and the linear layer.
[0040] During the model training process, cross entropy loss is introduced to improve the model's detection ability by shortening the distance between the model's prediction results and the labeled labels.
[0041] In some preferred embodiments, the video to be detected is preprocessed, and the specific steps are as follows:
[0042] The video to be detected is divided into frames, and the face area of each video frame is extracted. The extracted face area is used as the input of video modality feature extraction;
[0043] The audio slices corresponding to each group of video frames input for video modality feature extraction are intercepted as the input for audio modality feature extraction.
[0044] In some preferred embodiments, the video modality feature extraction comprises the following specific steps:
[0045] The video modality feature extraction includes a two-dimensional residual network to extract features from the input video frame group v, as shown in the following formula:
[0046]
[0047] in, represents a two-dimensional residual network, Fv It represents the forged features obtained by extracting the video modality features from the video frame group v, that is, the video features with strong representation ability are extracted from the face area obtained from the video to be tested.
[0048] In some preferred embodiments, the audio modality feature extraction comprises the following specific steps:
[0049] The audio modality feature extraction is composed of a feedforward neural network. The input of the audio modality feature extraction is an audio slice a corresponding to the video frame group input by the video modality feature extraction. The feedforward neural network containing multiple layers of neurons is used to extract from the audio a, as shown in the following formula:
[0050]
[0051] in, represents a feedforward neural network, F a It represents the forged features obtained by extracting audio modal features from audio a, that is, extracting audio features with strong representation ability from the audio in the video to be tested.
[0052] In some preferred embodiments, the self-attention learning comprises the following specific steps:
[0053] Self-attention learning consists of three branches based on Transformer encoders. The three branches share weights and use express;
[0054] The video modality self-attention branch forges features F from the video modality. v Further mining video forgery features with strong representation capabilities is shown in the following formula:
[0055]
[0056] The audio modality self-attention branch forges features F from the audio modality. a Further mining audio forgery features with strong representation capabilities is shown in the following formula:
[0057]
[0058] Multimodal self-attention branch forges features F from video and audio v 、F a The discontinuity caused by the deep fake operation between the two modalities is captured in the following equation:
[0059]
[0060] in, represents the three self-attention branches with shared weights, Video forgery features mined by the video modality self-attention branch, The audio forgery features mined by the audio modality self-attention branch, Discontinuities between multimodal features mined by the multimodal self-attention branch.
[0061] In some preferred embodiments, the multi-scale time domain fusion comprises the following specific steps:
[0062] First, the video forgery feature and the audio forgery feature are subtracted to obtain the difference between the two modal features, which is the audio and video modality difference;
[0063] Then, the discontinuity between the multi-modalities and the audio and video modal differences are merged to obtain the final merged result of the video forgery feature, the audio forgery feature, and the discontinuity between the multi-modalities, as shown in the following formula:
[0064]
[0065] The final merge result F fusion As input, features are fused and further screened through a multi-scale time-domain convolutional network, and the final decision result is obtained through the time-domain pooling layer and the linear layer.
[0066] Specifically, an embodiment can be introduced as follows.
[0067] S1. Preprocessing of the video to be detected.
[0068] S11. Perform frame division operation on the video at 24 frames per second.
[0069] S12. Use Dlib to draw a rectangular area to detect the face area of the video frame. Take the detected face as the center and frame an area of 224×224 pixels. The number of frames in each group of video frames is 16.
[0070] S13. Capture the audio slices corresponding to each group of video frames.
[0071] S2. Video modality feature extraction module, which extracts features from video modality.
[0072] S21. This embodiment uses ViViT as a video frame feature extractor. For an input video frame of 16×3×224×224, the output is a 768-dimensional vector F v .
[0073] S3. Audio modality feature extraction module, which extracts features from audio modality.
[0074] S31. This embodiment uses a feed-forward neural network (FFN) as an audio feature extractor. The output of FFN is a 768-dimensional vector F a .
[0075] S4. Self-attention module, further filtering the deep features of video and audio modalities, and capturing the discontinuity between audio and video modalities.
[0076] S41. This embodiment uses the Transformer encoder as the feature extractor of the self-attention module. The 768-dimensional video features obtained in S21 are input into the Transformer self-attention encoder. The output of the encoder is also a 768-dimensional video deep fake feature.
[0077] S42. Input the 768-dimensional audio features obtained in S31 into the Transformer self-attention encoder to obtain 768-dimensional audio deep fake features
[0078] S43. Combine the video features and audio features obtained in S21 and S31, and input them into the Transformer self-attention encoder to obtain 768-dimensional audio and video forgery features. Capturing inconsistencies in audio and video modalities in deepfake videos.
[0079] S5. Multi-scale temporal fusion of audio and video features.
[0080] S51. Take the video deep fake features obtained in S41 Audio deep fake features obtained with S42 The difference between The above formula means directly taking and The absolute value of the difference between two 768-dimensional vectors.
[0081] S52. With F diff Merge to obtain the input features of the multi-scale time domain fusion module In the above formula, ⊕ represents a merge operation.
[0082] S53. The multi-scale temporal convolutional network (MS-TCN) in this implementation performs the feature F fusion Perform feature fusion and screening in the time domain. The last two layers of MS-TCN are the time domain pooling layer and the linear layer, which output the final decision result.
[0083] in, Represents the video label predicted by the model (there are two types of labels: 0 and 1, 0 represents real video and 1 represents fake video).
[0084] S6. Training and testing
[0085] This example uses the FakeAVCeleb dataset for model training and prediction. The cross entropy loss between the real video label and the predicted label is calculated and the model parameters are updated through the back propagation algorithm. The cross entropy loss function is defined as follows:
[0086]
[0087] Among them, y i represents the true label of the i-th sample, represents the label of the i-th sample predicted by the model, Represents the average loss of n training samples.
[0088] Figure 2 This is an architecture diagram of a deep fake detection system based on audio and video time domain fusion provided in this application, and the system includes:
[0089] Preprocessing module for the video to be detected: used to perform frame segmentation and face area extraction operations on the video to be detected, and perform audio extraction at the same time;
[0090] Video modality feature extraction module: used to use residual networks to learn deep fake representations of video modalities and extract video features with strong representation capabilities;
[0091] Audio modality feature extraction module: used to perform deep fake representation learning on the audio modality using a feedforward network and extract audio features with strong representation capabilities;
[0092] Self-attention module: It contains three Transformer-based self-attention branches with shared weights. Two of them mine deep fake features from video features and audio features respectively, and the other branch merges the outputs of the video modality feature extraction and audio modality feature extraction modules to jointly learn the audio and video modalities and capture the discontinuity of deep fake videos between multiple modalities.
[0093] Multi-scale time-domain fusion module: used to merge the deep video forgery features, deep audio forgery features, and discontinuity features between multiple modalities mined by the Transformer-based self-attention module, use a multi-scale time-domain convolutional network for feature fusion, and obtain the final decision result through the time-domain pooling layer and the linear layer;
[0094] During the model training process, cross entropy loss is introduced to improve the model's detection ability by shortening the distance between the model's prediction results and the labeled labels.
[0095] The present application provides a deep fake detection system based on audio and video time domain fusion, the system comprising: the system comprising a processor and a memory:
[0096] The memory is used to store program code and transmit the program code to the processor;
[0097] The processor is used to execute the method described in any one of all embodiments of the first aspect according to the instructions in the program code.
[0098] The present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to be executed by a processor to implement the method described in any one of all the embodiments of the first aspect.
[0099] In a specific implementation, the present invention further provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, the program may include some or all of the steps in each embodiment of the present invention. The storage medium may be a disk, an optical disk, a read-only storage memory (abbreviated as: ROM) or a random access memory (abbreviated as: RAM), etc.
[0100] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution in the embodiments of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention or some parts of the embodiments.
[0101] The same and similar parts between the various embodiments of this specification can be referred to each other. In particular, for the embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.
[0102] The above-described embodiments of the present invention do not limit the protection scope of the present invention.
Claims
1. A deep fake detection method based on audio and video time domain fusion, characterized in that: The method comprises: The video to be detected is divided into frames, and the face area of each video frame is extracted. The extracted face area is used as the input for extracting the video modal features; At the same time, audio extraction is performed, and audio slices corresponding to each group of video frames are intercepted from the input for extracting video modal features as input for extracting audio modal features; Extracting video modality features, including using residual networks to learn deep fake representations of video modalities and extracting video features with strong representation capabilities; Extract audio modality features, including using a feedforward network to perform deep fake representation learning on the audio modality and extract audio features with strong representation capabilities; Use self-attention learning, including three Transformer-based self-attention branches with shared weights. Two of the branches extract deep fake features from video features and audio features respectively, and the other branch merges the outputs of the video modality feature extraction and audio modality feature extraction steps to jointly learn the audio and video modalities and capture the discontinuity of deep fake videos between multiple modalities; Use multi-scale time-domain fusion, including merging deep video forgery features, deep audio forgery features, and discontinuity features between multiple modalities extracted based on Transformer self-attention, using a multi-scale time-domain convolutional network for feature fusion, and obtaining the final decision result through the time-domain pooling layer and the linear layer; During the model training process, cross entropy loss is introduced to improve the detection ability of the model by shortening the distance between the model prediction results and the labeled labels.
2. The method according to claim 1, characterized in that: The video modality feature extraction described above includes the following specific steps: The video modality feature extraction includes a two-dimensional residual network to extract features from the input video frame group v, as shown in the following formula: F v =M θv (v); Among them, M θv represents a two-dimensional residual network, F v It represents the forged features obtained by extracting the video modality features from the video frame group v, that is, the video features with strong representation ability are extracted from the face area obtained from the video to be tested.
3. The method according to claim 1, characterized in that: The audio modal feature extraction described above comprises the following specific steps: The audio modality feature extraction is composed of a feedforward neural network. The input of the audio modality feature extraction is an audio slice a corresponding to the video frame group input by the video modality feature extraction. The feedforward neural network containing multiple layers of neurons is used to extract from the audio a, as shown in the following formula: F a =M θa (a); Among them, M θa represents a feedforward neural network, F a It represents the forged features obtained by extracting audio modal features from audio a, that is, extracting audio features with strong representation ability from the audio in the video to be tested.
4. The method according to claim 1, characterized in that: The self-attention learning described above has the following specific steps: Self-attention learning consists of three branches based on Transformer encoders. The three branches share weights and use M θe express; The video modality self-attention branch forges features F from the video modality. v Further mining video forgery features with strong representation capabilities is shown in the following formula: The audio modality self-attention branch forges features F from the audio modality. a Further mining audio forgery features with strong representation capabilities is shown in the following formula: Multimodal self-attention branch forges features F from video and audio v 、F a The discontinuity caused by the deep fake operation between the two modalities is captured in the following equation: in, represents the three self-attention branches with shared weights, Video forgery features mined by the video modality self-attention branch, The audio forgery features mined by the audio modality self-attention branch, Discontinuities between multimodal features mined by the multimodal self-attention branch.
5. The method according to claim 1, characterized in that: The multi-scale time domain fusion described above has the following specific steps: First, the video forgery feature and the audio forgery feature are subtracted to obtain the difference between the two modal features, which is the audio and video modality difference; Then, the discontinuity between the multi-modalities and the audio and video modal differences are merged to obtain the final merged result of the video forgery feature, the audio forgery feature, and the discontinuity between the multi-modalities, as shown in the following formula: The final merge result F fusion As input, features are fused and further screened through a multi-scale time-domain convolutional network, and the final decision result is obtained through the time-domain pooling layer and the linear layer.
6. A deep fake detection system based on audio and video time domain fusion, characterized in that: The system comprises: Preprocessing module for the video to be detected: used to perform frame segmentation and face area extraction operations on the video to be detected, and perform audio extraction at the same time; Video modality feature extraction module: used to use residual networks to learn deep fake representations of video modalities and extract video features with strong representation capabilities; Audio modality feature extraction module: used to perform deep fake representation learning on the audio modality using a feedforward network and extract audio features with strong representation capabilities; Self-attention module: It contains three Transformer-based self-attention branches with shared weights. Two of them mine deep fake features from video features and audio features respectively, and the other branch merges the outputs of the video modality feature extraction and audio modality feature extraction modules to jointly learn the audio and video modalities and capture the discontinuity of deep fake videos between multiple modalities. Multi-scale time-domain fusion module: used to merge the deep video forgery features, deep audio forgery features, and discontinuity features between multiple modalities mined by the Transformer-based self-attention module, use a multi-scale time-domain convolutional network for feature fusion, and obtain the final decision result through the time-domain pooling layer and the linear layer; During the model training process, cross entropy loss is introduced to improve the model's detection ability by shortening the distance between the model's prediction results and the labeled labels.
7. A deep fake detection system based on audio and video time domain fusion, characterized in that: The system comprises a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor; The processor is used to execute the method according to any one of claims 1 to 5 according to the instructions in the program code.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program codes, and the program codes are used to be executed by a processor to implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-operation detection method based on multi-scale feature fusion and multi-branch prediction
CN113850284A
Face counterfeit video detection model training method and apparatus, and computing device
CN116129502A
Multi-modal fusion detection method for deeply-forged audio and video
CN116797896A