Deep counterfeit video detection method based on multi-identity internal aggregation
By using the changed Efficient-B0 network and random channel single-head attention mechanism in deep forgery video detection, combined with spatial and temporal attention mechanism, the problem of deep forgery video detection in multiple identities is solved, and efficient and accurate detection results are achieved.
Patent Information
- Application Number
- CN202510177828.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
Existing deep fake video detection methods are difficult to achieve efficient detection in complex multi-identity scenarios, especially in local time period forgery and multi-identity or non-main character forgery scenarios, which are prone to misjudgment.
The changed Efficient-B0 network is used to extract shallow texture difference information and shallow frequency domain information, and a random channel single-head attention mechanism is designed for feature fusion. The interframe difference information is mined through spatial attention and temporal attention mechanisms to realize deep fake video detection of multi-identity internal aggregation.
It realizes efficient detection of deep forged videos in complex multi-identity scenarios, reduces the misjudgment rate, and reveals the dynamic evolution of forgery through attention scores, improving the interpretability and verification of detection.
Smart Images

Figure CN120125978A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of forged video detection, and specifically to a method for detecting deepfake videos with multi-identity internal aggregation. Background Art
[0002] With the rapid development of various generative models, deepfake video technology has gradually become the focus of social attention. However, malicious synthetic media content has spread widely on social platforms, triggering multiple security incidents such as political smear, blackmail, and false information dissemination, posing a serious threat to personal privacy and social trust.
[0003] For the detection of deepfake videos, existing methods can be divided into video-frame-based detection and video-frame-interrelationship-based detection methods. For video-frame-based detection methods, when dealing with videos with local-period forgery, that is, only some segments are forged for a certain identity, when counting the processing results of all video frames, due to the choice of activation function, the final result is affected by the number of true and false video frames, and it is extremely easy to cause misjudgment. For video-frame-interrelationship-based detection methods, such methods simply input a set of face images into the model to generate the final prediction result, without considering the scenarios of multi-identity or non-main character forgery, so it is impossible to achieve an efficient detection effect in complex scenarios.
[0004] Therefore, for deepfake videos in multi-identity complex scenarios, a detection method that can accurately identify them is needed. Summary of the Invention
[0005] To solve the deficiencies in the prior art, the present invention provides a method for detecting deepfake videos with multi-identity internal aggregation. By using the modified Efficient-B0 network to extract shallow texture difference information and shallow frequency domain information, designing a random-channel single-head attention mechanism for feature fusion, and using spatial attention and temporal attention to mine frame-interframe difference information, the detection of deepfake videos in multi-identity complex scenarios is solved.
[0006] To achieve the above object, the specific solution adopted by the present invention is as follows: A method for detecting deepfake videos with multi-identity internal aggregation mainly includes the following steps: Step S1, extract the face images in the video to be detected, and classify the face images according to the identities of the persons to whom they belong; Step S2, uniformly extract several face images containing each person's identity from the classified face images, and preprocess the extracted several face images to form an input sequence; Step S3: Use the modified Efficient-B0 network as the backbone network to extract shallow texture difference information and shallow frequency domain information for each face image in the input sequence. Concatenate the extracted shallow spatial domain texture difference information and shallow frequency domain information and input them into the random channel single-head attention mechanism for feature fusion to obtain a fused feature map. Then use the modified Efficient-B0 network to process the fused feature map and output a sequence of feature maps. Step S4: Linearly project the sequence of feature maps and add a CLS token, and then input them into the spatial attention mechanism and the temporal attention mechanism respectively. Execute the spatial attention mechanism on all the feature maps in the sequence of feature maps, and then execute the temporal attention mechanism on the feature maps belonging to the same person identity. Extract the temporal attention score and spatial attention score of each feature map, and add the spatial attention score and temporal attention score of each feature map as the weighted attention score of the feature map. Step S5: Separate the CLS token from the sequence of feature maps processed in Step S4, input the separated CLS token into a multi-layer perceptron for final video-level authenticity classification; normalize the weighted attention scores of all feature maps and group them according to the belonging identity information. Add up the weighted attention scores of all feature maps of each person identity to obtain the final score of each person identity, convert the final score of each person identity into a probability value, and label it on the corresponding person identity in the video to reveal the dynamic evolution of forgery.
[0007] Further, in Step S2, the method for preprocessing the face image is as follows: (1) Adjust the size of the facial area of the face image. (2) Randomly perform data augmentation on the face image, and then convert the image after data augmentation into a tensor to form an input sequence.
[0008] Further, in Step S3, the Efficient-B0 network contains 9 layers. The modified Efficient-B0 network is obtained by using average difference convolution to replace the ordinary convolution in the first layer. The random channel single-head attention mechanism is arranged between the second layer and the third layer of the modified Efficient-B0 network.
[0009] Further, use the first two layers of the modified Efficient-B0 network to extract texture difference information in the shallow spatial domain features of the image, and use the first two layers of the Efficient-B0 to extract the frequency domain information of the image transformed into the frequency domain; use the remaining seven layers to process the fused feature map.
[0010] Further, in step S4, each feature map in the feature map sequence is divided into 49 patches, linearly mapped to 512 dimensions, and then a CLS token is added as the subsequent input.
[0011] Further, in step S5, the multi-layer perceptron consists of a normalization layer and a fully connected layer.
[0012] Further, in step S5, when making the authenticity division according to the probability value, the threshold is set to 0.5 for the final authenticity prediction.
[0013] Beneficial effects:
[0014] (1) The present invention combines forgery features and inter-frame differences as the classification basis, extracts and mines forgery features from multi-identity face image sequences, and uses the temporal attention mechanism and the spatial attention mechanism to aggregate multi-identity feature information, realizing the efficient detection of deepfake videos in complex multi-identity scenarios.
[0015] (2) A random channel single-head attention mechanism is proposed. By randomly selecting some channels to perform the single-head attention mechanism, it can effectively fuse spatial and frequency domain information, reduce the computational amount, introduce the interaction between channels, and increase the diversity of the model.
[0016] (3) It is proposed to use the temporal attention mechanism and the spatial attention mechanism to achieve internal aggregation of multiple identities, realize video-level prediction in complex multi-identity scenarios, and label the prediction results to the corresponding identities in the video using attention scores to clearly reveal the dynamic evolution of forgery, improving interpretability and verifiability. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is the overall architecture diagram of the detection method of the present invention.
[0018] Figure 2 It is the architecture diagram of the random channel single-head attention mechanism in the present invention.
[0019] Figure 3 It is the schematic flow diagram of realizing internal aggregation of multiple identities through the temporal attention mechanism and the spatial attention mechanism in the present invention.
[0020] Figure 4 It is the schematic diagram of the final prediction result of the input video using the detection method. DETAILED DESCRIPTION OF THE INVENTION
[0021] The technical solution of the present invention will be clearly and completely described below in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.
[0022] The present invention provides a method for detecting deepfake videos with multi-identity internal aggregation. Please refer to Figure 1 , which mainly includes the following steps: Step S1: Extract the face images in the video to be detected and classify the face images according to the identities of the persons to whom they belong; Step S2: Uniformly extract several face images containing each person's identity from the classified face images, and preprocess the extracted several face images to form an input sequence; Step S3: Use the modified Efficient-B0 network as the backbone network to extract the shallow texture difference information and shallow frequency domain information of each face image in the input sequence, splice the extracted shallow spatial domain texture difference information and shallow frequency domain information and input them into the random channel single-head attention mechanism for feature fusion to obtain a fused feature map; then use the Efficient-B0 network to process the fused feature map and output a sequence of feature maps; Among them, the Efficient-B0 network contains 9 layers. The modified Efficient-B0 network is obtained by replacing the ordinary convolution in the first layer with average differential convolution. The first two layers of the modified Efficient-B0 network are used to extract the texture difference information in the shallow spatial domain features of the image, and the first two layers of Efficient-B0 are used to extract the frequency domain information of the image transformed into the frequency domain; the random channel single-head attention mechanism is arranged between the second layer and the third layer of the modified Efficient-B0 network. After using the random channel single-head attention mechanism for feature fusion to obtain a fused feature map, the remaining seven layers are used to process the fused feature map; Step S4: Linearly project the sequence of feature maps and add a CLS token, and then input them into the spatial attention mechanism and the temporal attention mechanism respectively. Please refer to Figure 3 , perform the spatial attention mechanism on all the feature maps in the sequence of feature maps, and then perform the temporal attention mechanism on the feature maps belonging to the same person's identity, extract the temporal attention score and spatial attention score of each feature map, and add the spatial attention score and temporal attention score of each feature map as the weighted attention score of the feature map; Step S5: Isolate the CLS token from the sequence of feature maps processed in step S4, and input the isolated CLS token into a multi-layer perceptron (MLP head) for final video-level authenticity classification; normalize the weighted attention scores of all feature maps and group them according to the affiliated identity information. Add up the weighted attention scores of all feature maps of each person's identity to obtain the final score of each person's identity, convert the final score of each person's identity into a probability value, and label it on the corresponding person's identity in the video to reveal the dynamic evolution of forgery.
[0023] It should be noted that in step S5, when performing authenticity classification according to the probability value, the threshold is set to 0.5 for final video-level prediction.
[0024] Furthermore, the multi-layer perceptron consists of a normalization layer and a fully connected layer.
[0025] Figure 4 A schematic diagram showing the output prediction result using the detection method of the present invention is given. By Figure 4 It can be seen that the present invention can effectively track the characters appearing in the multi-identity video, can dynamically capture and analyze the forgery evolution process of the characters in the video, and realize effective deepfake detection.
[0026] Specific embodiments are given below to further clearly, completely, and detailedly illustrate the technical solution of the present invention. This embodiment is the best embodiment based on the technical solution of the present invention, but the protection scope of the present invention is not limited to the following embodiments.
[0027] Embodiment 1
[0028] This embodiment provides a method for detecting deepfake videos with multi-identity internal aggregation, including the following steps: Step S1: First, extract the face images in the input video, use the face detector MTCNN to crop the face area, then use the InceptionResnetV1 model for feature extraction and similarity calculation, and then classify the extracted faces by identity, while discarding the identity classes with less than 3 extraction quantities.
[0029] Step S2: Uniformly extract 16 face images of each identity from the classified face images, and preprocess the extracted several face images to form an input sequence.
[0030] The specific method for preprocessing the face images is as follows: (1) Adjust the size of the face area of the above 16 face images, which is unified to 224*224 pixels in this embodiment; (2) Randomly perform data augmentation on the face images. In this embodiment, means including adding Gaussian noise, horizontally flipping the images, adjusting the brightness and contrast of the images, translating or rotating the images can be used to process the face images. After that, the images after the above data augmentation are converted into tensor tensors to form an input sequence.
[0031] Step S3: Use the modified Efficient-B0 network as the backbone network to extract shallow texture difference information and shallow frequency domain information for each face image in the input sequence. Concatenate the extracted shallow spatial domain texture difference information and shallow frequency domain information and input them into the random channel single-head attention mechanism for feature fusion to obtain a fused feature map. Then use the Efficient-B0 network to process the fused feature map and output a sequence of feature maps.
[0032] In this embodiment, the ordinary convolution in the first layer of the Efficient-B0 network is replaced with an average difference convolution (ADC) to obtain the modified Efficient-B0 network. The network structures of the modified Efficient-B0 network and the random channel single-head attention mechanism are shown in Table 1.
[0033] Table 1 Network structures of the modified Efficient-B0 network and the random channel single-head attention mechanism Layer Operator Resolution #Channels 1 ADC or VC, K7×7 224×224 32 2 MBConv1, K3×3 112×112 16 - RCSA 112×112 32 3 MBConv6, K3×3 112×112 24 4 MBConv6, K5×5 56×56 40 5 MBConv6, K3×3 28×28 80 6 MBConv6, K5×5 14×14 112 7 MBConv6, K5×5 14×14 192 8 MBConv6, K3×3 7×7 320 9 Conv1×1 & Pooling & FC 7×7 1280
[0034] Please refer to Figure 2 , in this embodiment, the single-head attention mechanism module randomly selects half of the channels in the input features for attention calculation, effectively fuses the channel information and reduces the computational complexity. Using this module, both the diversity of the input is retained, and the global dependence information is captured through self-attention, improving the generalization ability of the model.
[0035] Specifically, this module randomly selects half of the channels of the input feature map, generates queries (Q), keys (K), and values (V) for the selected feature map, and calculates the relationships between these features to generate a global feature representation. Concatenate the attention-weighted features calculated above with the features of the unselected channels to finally form a complete fused feature map.
[0036] Specifically, for the input feature X with the shape of [B, C, H, W], where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map respectively. Randomly select C / 2 channels, apply a convolution operation to the selected channels to generate queries (Q), keys (K), and values (V), and make these features suitable for self-attention calculation through a flattening operation. Calculate the attention weights through the dot product of the queries (Q) and keys (K), and perform normalization. Use the normalized weights to weight the values (V) to obtain the attention-weighted features, and adjust the shape to [B, C / 2, H, W]. Concatenate the weighted features with the features of the unselected channels to generate a feature map with the shape of [B, C, H, W], and generate the final fused feature map through projection.
[0037] Step S4: Divide each feature map in the feature map sequence into 49 patches, linearly map them to 512 dimensions, and then add a CLS token as the subsequent input.
[0038] Assume the feature map sequence is described as [B, F, C, H, W], where B represents the batch size, F represents the number of input sequences, which is 16 in this embodiment, C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map. After the feature map is divided, it is converted to [B, (F * H * W), C], and (F * H * W) represents the product of the number of input sequences and the number of patches. After linear mapping, C in [B, (F * H * W), C] is transformed into 512, that is, [B, (F * H * W), 512].
[0039] Furthermore, for the linearly mapped input sequence [B, (F * H * W), 512], embed a classification token (CLS token) at the beginning, and the shape becomes [B, 1 + (F * H * W), 512].
[0040] Input the processed feature map sequence into the spatial attention mechanism and the temporal attention mechanism. First, perform the spatial attention mechanism on all feature maps, and then perform the temporal attention mechanism on the feature maps belonging to the same identity. Extract the temporal attention score and spatial attention score of each feature map, and add the spatial attention score and temporal attention score of each feature map as the weighted attention score of the feature map.
[0041] Step S5: Isolate the CLS token from the sequence of feature maps processed in Step S4, and input the isolated CLS token into a multi-layer perceptron for final video-level authenticity classification; perform normalization on the weighted attention scores of all feature maps and group them according to the affiliated identity information. Add up the weighted attention scores of all feature maps for each person's identity to obtain the final score for each person's identity, convert the final scores of each person's identity into probability values, and label them on the corresponding person's identity in the video to reveal the dynamic evolution of forgery.
[0042] The CLS token generates corresponding scores through a multi-layer perceptron, and then the scores are mapped to the interval from 0 to 1 through a sigmoid function. A threshold of 0.5 is set to generate the final classification result of the input video.
[0043] The experimental results in the ForgeryNet public dataset are shown in Tables 2 and 3. The results show that the method of the present invention has achieved the optimal AUC value in six forgery methods, and it basically remains at 99%. This indicates that when using the method of the present invention for video-level prediction, the method of combining frame image forgery traces with inter-frame inconsistencies has extremely strong classification performance. In addition, in terms of ACC, the proposed method also remains at about 98%, which is not much different from the comparison method. This confirms that the method of the present invention has excellent detection effects when detecting a single forgery mode.
[0044] Table 2 Video-level prediction results for six forgery methods on ForgeryNet
[0045] Table 3 Video-level prediction results (TNR, TPR) for six forgery methods on ForgeryNet
[0046] Meanwhile, the present invention selects videos with multiple identities in the scene in the ForgeryNet dataset and conducts experiments on the model based on multiple identities. The experimental results are shown in Table 4. In the tests with two and three identities, the AUC of the method of the present invention is ahead of MINTIME, and the accuracy in multi-identity videos with three or more identities is also better than MINTIME.
[0047] Table 4 Multi-identity video-level prediction results on ForgeryNet
[0048] The deepfake video detection method proposed by the present invention extracts shallow texture difference information and shallow frequency domain information through the modified Efficient-B0 network embedded with average differential convolution. Then, a random channel single-head attention mechanism is used for feature fusion to obtain a fused feature map, and the fused feature map is processed using the modified Efficient-B0 network. Secondly, a spatial attention mechanism and a temporal attention mechanism are used to process the feature map, and the weighted attention scores of each feature map are calculated. Finally, a multi-layer perceptron is used to achieve the final deepfake video detection, and this detection method has better accuracy when detecting multi-identity videos.
[0049] As described above, it is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Any equivalent transformation or modification made according to the essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A multi-identity internal aggregation deep fake video detection method, characterized in that: The main steps are as follows: Step S1, extracting face images from the video to be detected, and classifying the face images according to the identities of the people they belong to; Step S2, evenly extracting a number of face images containing various identities from the classified face images, and pre-processing the extracted face images to form an input sequence; Step S3, using the modified Efficient-B0 network as the backbone network, extracting shallow texture difference information and shallow frequency domain information for each face image in the input sequence, splicing the extracted shallow spatial texture difference information and shallow frequency domain information and inputting them into the random channel single-head attention mechanism for feature fusion to obtain a fused feature map; then using the modified Efficient-B0 network to process the fused feature map and output a feature map sequence; Step S4: linearly project the feature map sequence and add CLS token, and then input them into the spatial attention mechanism and the temporal attention mechanism respectively. The spatial attention mechanism is executed on all feature maps in the feature map sequence, and then the temporal attention mechanism is executed on the feature maps belonging to the same person identity, and the temporal attention score and spatial attention score of each feature map are extracted. The spatial attention score and temporal attention score of each feature map are added as the weighted attention score of the feature map; Step S5, separate the CLS token from the feature map sequence processed by step S4, and input the separated CLS token into a multi-layer perceptron for final video-level authenticity classification; normalize the weighted attention scores of all feature maps and group them according to their identity information, add up the weighted attention scores of all feature maps of each character identity to obtain the final score of each character identity, convert the final score of each character identity into a probability value, and mark it on the corresponding character identity in the video to reveal the dynamic evolution of forgery.
2. According to the multi-identity internal aggregation deep fake video detection method of claim 1, it is characterized in that: In step S2, the method for preprocessing the face image is: (1) Adjust the size of the facial area of the face image; (2) Randomly perform data augmentation on the face images, and then convert the data augmented images into tensors to form an input sequence.
3. According to the multi-identity internal aggregation deep fake video detection method of claim 1, it is characterized in that: In step S3, the Efficient-B0 network includes 9 layers, and the modified Efficient-B0 network is obtained by using average difference convolution instead of ordinary convolution in the first layer. The random channel single-head attention mechanism is set between the second and third layers of the modified Efficient-B0 network.
4. According to the method of detecting deep fake videos with multi-identity internal aggregation as described in claim 3, it is characterized in that: The first two layers of the modified Efficient-B0 network are used to extract texture difference information in the shallow spatial features of the image, and the first two layers of Efficient-B0 are used to extract frequency domain information transformed into a frequency domain image; the remaining seven layers are used to process the fused feature map.
5. According to the method of multiple identity internal aggregation deep fake video detection in claim 1, it is characterized in that: In step S4, each feature map in the feature map sequence is divided into 49 patches and linearly mapped to 512 dimensions, and then the CLS token is added as the subsequent post-input.
6. The method for detecting deep fake videos with multi-identity internal aggregation according to claim 1, characterized in that: In step S5, the multilayer perceptron consists of a normalization layer and a fully connected layer.
7. The method for detecting deep fake videos with multi-identity internal aggregation according to claim 1, characterized in that: In step S5, when true and false classification is performed according to the probability value, the threshold is set to 0.5 for the final true and false prediction.
Citation Information
Cited By
Deep pseudo detection method and device and server
CN121280759A
Deep counterfeit video detection method, device, equipment and medium
CN121582858A
Deepfake video detection method, apparatus, device and medium
CN121582858B