Deepfake Video Detection Method and System Based on Intra-frame and Inter-frame Feature Differentiation
By extracting and comparing the differentiation of intra- and inter-frame features of deep-fake videos, the problem of insufficient detection accuracy of low-quality and hybrid artificial intelligence synthetic forged face videos in the prior art is solved, and higher detection accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202210718973.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-23
AI Technical Summary
The prior art still has a lot of room for improvement in the accuracy of the detection algorithm when detecting fake face videos synthesized by low-quality and mixed artificial intelligence, and the dynamic changes of single-frame images have not been fully considered.
The deep fake video detection method based on intra-inter-feature differentiation is adopted. By extracting the intra-frame features and inter-frame features of the video, and calculating their differentiated and Euclidean distances, the authenticity detection of the deep fake video is achieved.
It significantly improves the accuracy and accuracy of deep forged video detection, and can more effectively combat forged videos, capture the intra-frame characteristics and their differentiation of forged videos.
Smart Images

Figure CN115147758B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of forged video detection, and particularly relates to a deep fake video detection method and system based on the difference between intra-frame and inter-frame features. Background Art
[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] With the development of deep learning technology, more and more synthetic videos have emerged in people's daily lives; with powerful capabilities, it is possible to synthesize forged images and videos that are difficult for humans to detect, making it very difficult for people to identify these synthetic forged images with the naked eye.
[0004] In the prior art, with face synthesis technology, by clicking a few times on a mobile device, attractive services such as face swapping and facial expression processing can be provided; however, this artificial intelligence technology has significant security and privacy issues, such as threatening facial recognition systems, etc., which can easily cause great social impacts and hazards. In recent years, with the large emergence of deep fake face synthesis videos, a series of detection technologies for such videos have emerged. In addition, complex face forgery detection systems have been developed by using biometric information (such as blinking or head pose and advanced training sets) to train detectors capable of identifying deep fake videos.
[0005] As the inventor understands, although technologies based on capturing a single feature, such as convolutional network traces, bioactivity detection, and intra-frame features of video pictures, have made good progress in detection tasks, there is still much room for improvement in the accuracy of detection algorithms for low-quality and mixed artificial intelligence synthetic forged face videos. Currently, in face identity swapping, single-frame images are the starting point of forgery technology, and little consideration is given to the dynamic changes between single-frame images. Operating on a single modality of an object may lead to inconsistencies in other modalities. Summary of the Invention
[0006] To solve the above problems, the present disclosure proposes a deep fake video detection method and system based on the difference between intra-frame and inter-frame features. By the difference between the intra-frame features and inter-frame features of the forged video, it is possible to more effectively combat the forged video, capture the intra-frame features and inter-frame features of the forged video and their differences, and thus improve the detection of deep fake videos.
[0007] According to some embodiments, the first solution of the present disclosure provides a deep fake video detection method based on the difference between intra-frame and inter-frame features, adopting the following technical solution:
[0008] A deep fake video detection method based on the difference between intra-frame and inter-frame features, comprising:
[0009] Obtain the original data of the deepfake video;
[0010] Extract intra-frame features and inter-frame features based on the obtained original data;
[0011] Calculate the difference and Euclidean distance between the extracted intra-frame features and inter-frame features;
[0012] Complete the detection of the authenticity of the deepfake video according to the obtained difference and Euclidean distance.
[0013] As a further technical limitation, after obtaining the original data of the deepfake video, perform unified processing on the obtained original data, extract the original data of the obtained deepfake video into picture frames in units of frames, and suppress the interference of irrelevant video backgrounds except for the face part.
[0014] Furthermore, use bottleneck attention optimization for feature optimization of picture frames, and use the attention mechanism to extract facial region features to suppress irrelevant backgrounds.
[0015] Furthermore, during the extraction process of picture frames, use the face recognition library to locate and crop the face region in the picture frames, and save the pictures uniformly as 128×128.
[0016] As a further technical limitation, during the process of extracting intra-frame features, use RGB images, represent the inter-frame flow through dense optical flow, and focus on the regions with large facial changes based on optical flow features to complete the extraction of intra-frame features.
[0017] As a further technical limitation, during the process of extracting inter-frame features, estimate the optical flow tracking of the movement offsets of all pixel points by calculating the offset vectors of all pixel points on the front and rear two-frame images to complete the extraction of inter-frame features.
[0018] As a further technical limitation, use the cross-entropy loss function to calculate the loss functions of intra-frame features and inter-frame features respectively, calculate the weighted sum of the loss functions, and obtain the difference between intra-frame features and inter-frame features.
[0019] According to some embodiments, the second solution of the present disclosure provides a deepfake video detection system based on the difference between intra-frame and inter-frame features, and adopts the following technical solutions:
[0020] A deepfake video detection system based on the difference between intra-frame and inter-frame features, comprising:
[0021] An acquisition module configured to acquire the original data of the deepfake video;
[0022] An extraction module configured to extract intra-frame features and inter-frame features based on the acquired original data;
[0023] A calculation module configured to calculate the difference and Euclidean distance between the extracted intra-frame features and inter-frame features;
[0024] A detection module configured to complete the detection of the authenticity of the deepfake video based on the obtained difference and Euclidean distance.
[0025] According to some embodiments, the third solution of the present disclosure provides a computer-readable storage medium, adopting the following technical solution:
[0026] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the method for detecting deepfake videos based on the difference between intra-frame and inter-frame features described in the first aspect of the present disclosure.
[0027] According to some embodiments, the fourth solution of the present disclosure provides an electronic device, adopting the following technical solution:
[0028] An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for detecting deepfake videos based on the difference between intra-frame and inter-frame features described in the first aspect of the present disclosure.
[0029] Compared with the prior art, the beneficial effects of the present disclosure are as follows:
[0030] The present disclosure constructs a dual-network structure, in which two improved CNN sub-networks are used to extract intra-frame and inter-frame features from the input data respectively for training. A contrastive loss function is used to associate these two sub-networks and capture the disharmony between intra-frame and inter-frame features; the cross-entropy loss function is used after the fully connected layers of the intra-frame sub-network and the inter-frame sub-network respectively, and it is combined with the contrastive loss to form the overall loss function of the network to improve the learning effect of the network;
[0031] The present disclosure adds a bottleneck attention module to the network to optimize the input feature map based on its global feature statistical information. The bottleneck attention module is placed at the bottleneck of the model, enabling the lower-level features to benefit from the context information. A lightweight module design is adopted to make the entire program run in an efficient manner.
[0032] The present disclosure significantly improves the detection accuracy and correct rate in related work such as deepfake detection. Description of the Drawings
[0033] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the disclosure. The schematic embodiments and descriptions thereof of the disclosure are used to explain the disclosure and do not constitute an improper limitation of the disclosure.
[0034] Figure 1 is a flowchart of the deepfake video detection method based on intra-frame and inter-frame feature differentiation in the first embodiment of the present disclosure;
[0035] Figure 2 is an overall architecture diagram of the deepfake video detection method based on intra-frame and inter-frame feature differentiation in the first embodiment of the present disclosure;
[0036] Figure 3 is a structural diagram of the intra-frame sub-network in the first embodiment of the present disclosure;
[0037] Figure 4 is a structural diagram of the bottleneck attention module in the first embodiment of the present disclosure;
[0038] Figure 5 is a structural block diagram of the deepfake video detection system based on intra-frame and inter-frame feature differentiation in the second embodiment of the present disclosure. Detailed implementation manners
[0039] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0040] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0041] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0042] Without conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other.
[0043] Embodiment 1
[0044] The first embodiment of the present disclosure introduces a deepfake video detection method based on intra-frame and inter-frame feature differentiation.
[0045] As Figure 1 and Figure 2 shown, a deepfake video detection method based on intra-frame and inter-frame feature differentiation includes:
[0046] Step S01: Obtain the original data of the deepfake video;
[0047] Step S02: Extract intra-frame features and inter-frame features based on the obtained original data;
[0048] Step S03: Calculate the difference and Euclidean distance between the extracted intra-frame features and inter-frame features;
[0049] Step S04: Complete the detection of the authenticity of the deepfake video according to the obtained difference and Euclidean distance.
[0050] As one or more embodiments, in step S01, search and collect relevant deepfake datasets, unify the formats of the collected datasets, generate labels, perform face localization on pictures, crop pictures, etc. for preprocessing. Then divide the datasets.
[0051] The comparison datasets used are all publicly available datasets, including FaceForensics++, DeepfakeTIMIT, UADFV, and Celeb DF datasets. Most of these datasets are frontal face videos of people collected from Youtube or other public video websites, and then forged videos are created on this basis through different forgery means. Since the data in the collected datasets are all circulated on the Internet in the form of videos and there is no unified format and authenticity label among them, before the formal experiment starts, it is necessary to perform unified data preprocessing on each dataset involved, that is, to unify the data format; for videos with different formats and sizes in the dataset, extract all videos into pictures in units of frames.
[0052] In order to improve the effect of network training and suppress the interference of irrelevant video backgrounds other than the face part during training, for all picture frames extracted from the dataset videos, in this embodiment, the face recognition library is used to locate and crop the face area in the picture frame and uniformly adjust it to a picture size of 128×128 for saving. After the data format is unified, each picture needs to be labeled according to the authenticity of the video in the dataset and the division of the training data.
[0053] In this embodiment, datasets such as Face Forensics++ are divided into a training set, a validation set, and a test set in a ratio of 5:1:1 to evaluate the generalization ability of the trained network model on video files of unknown authenticity. It should be noted that, different from some previous studies, each dataset contains videos from different identities (real and fake). This aspect is very important for a fair evaluation and prediction of the generalization ability of the fake detection system for unknown identities. And in order to test the generalization model on datasets with different forgery methods, as opposed to training and detecting data only for a single dataset, this embodiment will adopt the dual-stream network HOR (Detecting Deep fake Videos using the Disharmony between Intra- and Inter-frame Maps) to evaluate the detection ability of training data mixed with multiple different forgery methods.
[0054] As one or more embodiments, in step S02, corresponding intra-frame features and inter-frame features are extracted from the acquired data to be used as the original input to train the subsequent dual-stream detection network model.
[0055] The overall network architecture in this embodiment is a dual-stream network composed of two Convolutional Neural Networks (CNNs) sub-networks, namely an intra-frame sub-network for extracting intra-frame spatial features and an inter-frame sub-network for extracting inter-frame temporal features.
[0056] The intra-frame sub-network extracts intra-frame features from the RGB images of the video, and the inter-frame sub-network extracts inter-frame features from the dense optical flow map of the video; for the inter-frame flow, dense optical flow is used to represent the inter-frame flow. Because the basic principle of deep fakes in video processing is to process each forged frame image and then connect the generated forged images. This will cause larger deformations in areas with significant facial changes (such as eyes, lips) due to the lack of coherence. Compared with other traditional techniques based on intra-frame image features, using optical flow features can better focus on areas with significant facial changes and improve the detection accuracy.
[0057] It is determined whether the video is real or fake by judging the difference between the intra-frame features and the inter-frame features; the difference between the intra-frame features and the inter-frame features is captured by a contrast loss function, which makes the Euclidean distance between the intra-frame and inter-frame features smaller for real videos and larger for forged videos.
[0058] Specifically, as Figure 2As shown, the intra-frame sub-network extracts intra-frame features from the RGB images of the video, and the inter-frame network extracts inter-frame features from the pre-processed dense optical flow maps of the video; the network architectures of the intra-frame and inter-frame sub-networks are both based on the ResNet network using separable convolutions, and the residual learning mechanism adopted solves problems such as slow network convergence and degraded training effects.
[0059] During the process of inter-frame feature extraction, the dense optical flow algorithm is adopted. Specifically:
[0060] Optical flow can be mainly divided into sparse optical flow and dense optical flow. The principles of the two are similar, but sparse optical flow only estimates the movement of specifically selected pixel points on the image in the front and back two frames of images. Dense optical flow, on the other hand, is an optical flow tracking estimation algorithm that calculates the offset vectors of all pixel points on the front and back two frames of images to achieve the estimation of the movement offsets of all pixel points.
[0061] The main algorithm idea of the dense optical flow algorithm is to first determine the weight of each pixel point according to the pixel values and coordinates of other pixel points in its nearby neighborhood, and then expand the coordinates of this point with a polynomial. Regarding the image as a function of a two-dimensional signal (the output image is a grayscale image), the dependent variable is as shown in formula (1). Then, a quadratic polynomial is used to approximately model the image, as shown in formula (2), where A represents a 2x2 symmetric matrix, b is a 2x1 vector matrix, and c is a scalar:
[0062] x = (x, y) T (1)
[0063] f(x) ~ x T Ax + b T x + c (2)
[0064] Then, through coefficient conversion, the right side of formula (2) can be written as formula (3):
[0065]
[0066] The two-dimensional signal space (Cartesian coordinate system) of the original image needs a six-dimensional vector as coefficients to be converted to a space with (1, x, y, x 2 , y 2 , xy) as the basis functions, and substituting the positions x, y of different pixel points to obtain the grayscale values of different pixel points. To obtain the six coefficients of each pixel point in each frame of the image, the Farneback algorithm sets a (2n + 1) × (2n + 1) neighborhood around each pixel point, and then uses the (2n + 1) 2 pixel points in the pixel point neighborhood as the sample points of the least squares method for fitting.
[0067] In a grayscale value matrix of size (2n + 1)×(2n + 1) within the neighborhood of a pixel, the matrix is split and combined into (2n + 1) 2 ×1 vectors f in column - major order. At the same time, it is known that the transformation matrix B with (1, x, y, x 2 , y 2 , xy) as the basis functions has a dimension of (2n + 1) 2 ×6 (that is, a matrix composed of 6 column vectors b i jointly). The dimension of the common coefficient vector r within the neighborhood is 6×1. Then there is formula (4):
[0068] f = B×r = (b 1 b 2 b 3 b 4 b 5 b 6 )×r (4)
[0069] When using the least - squares method to solve, the Farneback algorithm uses a two - dimensional Gaussian distribution to weight the sample error of each pixel within the neighborhood. In the (2n + 1)×(2n + 1) matrix of the two - dimensional Gaussian distribution within the neighborhood of each pixel, the matrix is split into (2n + 1) 2 ×1 vectors a in column - major order. As shown in formula (5), the original transformation matrix B of the basis functions will be transformed into:
[0070] B = (a·b 1 a·b 2 a·b 3 a·b 4 a·b 5 a·b 6 ) (5)
[0071] By transforming the basis function matrix B again in a dual way, the coefficient vector of each pixel in a single image can be obtained. Then, through parameter vector calculation and local blurring processing, the optical flow field can be obtained. After obtaining the optical flow field, in order to make the input data structure of the inter - frame flow correspond to the RGB three - layer data structure of the intra - frame flow, the two - layer optical flow data matrix is supplemented to a three - layer matrix.
[0072] The network architecture of the intra - frame sub - network consists of 36 convolutional layers. This network is based on the Xception network model, which shows strong learning ability in image vision.
[0073] The inter-frame sub-network is used to explore the spatio-temporal long-range context correlation of the key regions of faces in video frames to enhance the ability of learning representation. In this embodiment, dense optical flow is used to infer the displacement process and direction of pixel points in the image from two consecutive frames of images, and then the inter-frame sub-network captures features from the dense optical flow.
[0074] Specifically, the calcOpticalFlowFarneback function provided in OpencCV is used. This function adopts the Farneback dense optical flow algorithm based on image pyramid modeling. The algorithm builds a three-layer image pyramid based on all pixel points in two consecutive frames, and the size of each layer is half of the previous layer. In three iterations, the window size used to construct the flow is set to 15. Then, the Farneback dense optical flow algorithm is calculated for each layer of the two image pyramids from top to bottom. By establishing the image pyramid, it is easier for the optical flow to capture objects with larger movement amplitudes.
[0075] In terms of the overall structure of the inter-frame sub-network, it is generally similar to the intra-frame sub-network. However, in order to better extract inter-frame features, an LSTM (Long Short-Term Memory) network layer is added at the end of the network. This can solve the problems of gradient explosion and gradient disappearance during the training process of long sequences. Similar to the visual flow, in this embodiment, a softmax layer is added at the end of the inter-frame sub-network, and the output is incorporated into the cross-entropy loss of the inter-frame mode.
[0076] To improve the detection ability of the network, this embodiment introduces a Figure 3 bottleneck attention module as shown. The bottleneck attention module uses the attention mechanism to extract key features from the facial regions of video frames and suppress irrelevant background information.
[0077] The bottleneck attention module uses the attention mechanism to extract key features from the facial regions of video frames and suppress irrelevant background information; the input of this module is the set of frame-level feature maps at the network bottleneck.
[0078] The feature map consists of where C, H, and W represent the number of channels, height, and width of the feature map respectively. For a given input feature map the bottleneck attention module infers a 3D attention map The final output F′ of the bottleneck attention module is calculated as follows:
[0079]
[0080] The specific implementation of the 3D attention map M(F) is by calculating the channel attention and the spatial attention On two independent branches. For channel attention, the input tensor F first uses a global average pooling layer to softly encode the global information in each channel. Then, a multi-layer perceptron (MLP) with one hidden layer is used to calculate the cross-channel attention from the channel vector M C (F). To adapt to the size of the output data of the spatial branch, a Batch Normalization layer is added after the MLP in the bottleneck attention module. Briefly, the channel attention M C (F) is calculated as follows:
[0081] M C (F) = BN(MLP(AvqPool(F))) (7)
[0082] The spatial branch highlights or ignores features at different spatial positions by generating a spatial attention map . The most important thing is to have a large receptive field in the spatial dimension to effectively utilize context information. Dilated convolutions can be used to efficiently expand the receptive field. The spatial branch adopts the bottleneck structure proposed by ResNet, saving the number of parameters and computational overhead. Specifically, to integrate and compress the feature map in the channel dimension, 1×1 convolutions are used to project the features to reduce the dimension to Briefly, the formula for spatial attention is as follows:
[0083]
[0084] where c represents the convolution operation, BN represents the batch normalization operation, and the superscript represents the size of the convolution filter.
[0085] Finally, after obtaining the channel attention M C (F) and the spatial attention M S (F), the two attention branches are combined to produce the final M(F). Since the channel attention M C (F) and the spatial attention M S (F) have different dimensions, before combining M S (F) and M C (F), the attention maps are extended to After element-wise summation, the sigmoid function is used to obtain the final 3D attention map M(F) in the range of 0 to 1. The 3D attention map is multiplied by the original input feature map F and then added to the original input feature map F to obtain the output feature map of the final bottleneck attention module, as shown in formula (6).
[0086] The main advantages of using self-attention mechanisms in CNNs are efficient global context modeling and effective backpropagation (i.e., model training). The global context allows the model to better identify locally ambiguous patterns and focus on important parts. Therefore, capturing and leveraging the global context is crucial for various visual tasks. In this regard, CNN models usually stack many convolutional layers or use pooling operations to ensure that features have a large receptive field.
[0087] As one or more embodiments, in step S03, the present embodiment uses a cross-loss function for the intra-frame and inter-frame features extracted by the intra-frame sub-network and the inter-frame sub-network to calculate the difference between the two and uses the intra-frame sub-network and the inter-frame sub-network to learn distinct unimodal features through cross-entropy loss.
[0088] After the sub-networks extract features from the input data respectively, the extracted features are output to the contrast loss function and the cross-entropy loss function through fully connected layers. The contrast loss function captures the difference between the extracted intra-frame features and inter-frame features and measures it by the Euclidean distance between them.
[0089] The cross-entropy loss has the advantages of simplicity and efficiency and is a commonly used loss function in deepfake detection tasks. However, in the HOR detection network proposed in the present embodiment, the contrast loss is used as a key component of the objective function. Initially, the contrast loss function mainly acts in aspects related to dimensionality reduction, and its theoretical basis is that the similarity of data samples in the feature space is not affected by dimensionality reduction (feature extraction). Therefore, the present embodiment utilizes the characteristic that the contrast loss can effectively reflect the similarity between samples and proposes a differential detection method based on intra-frame features and inter-frame features. The contrast loss maximizes the difference score of the manipulated video while minimizing the difference score of the real video. Compared with traditional detection networks using cross-entropy loss functions, it has better detection effects.
[0090] The contrast loss maximizes the difference score of the manipulated video while minimizing the difference score of the real video. The contrast loss function is shown in formula (9), where y i is the label of video v i , and the margin is a hyperparameter. The difference score represents the Euclidean distance between the intra-frame feature f a and the inter-frame feature f e of the intra-frame sub-network and the inter-frame sub-network respectively. In addition, the present embodiment uses the cross-entropy loss of the intra-frame sub-network and the inter-frame sub-network to learn feature representations respectively. These loss functions are defined in formula (11) (intra-frame) and formula (12) (inter-frame). The total loss is the weighted sum of these three losses, as shown in formula (13):
[0091]
[0092]
[0093]
[0094]
[0095] L = L c +L a +L e (13)
[0096] As one or more embodiments, in step S04, the authenticity of the video is detected by judging according to the Euclidean distance between the output features of the intra-frame sub-network and the output features of the inter-frame sub-network. In order to mark the test video, in this embodiment, 1{d_t^i < τ} is used, where 1{.} represents the logical indicator function of comparing the Euclidean distance d_t^i with the threshold τ. τ is determined by the training set. In this embodiment, by calculating the Euclidean distances of the real and fake videos in the training set, the midpoint between the averages of the real and fake videos is used as the representative value of τ.
[0097] If a video is represented in different modalities, manipulating any one modality alone will result in some differences between the modalities. Generally speaking, the difference between the intra-frame and inter-frame features of a real video is significantly smaller than that of a fake video; and the measure of the difference is expressed by the Euclidean distance, that is, the larger the Euclidean distance, the greater the difference, and the smaller the Euclidean distance, the smaller the difference; the authenticity of the video is determined by comparing the size of the Euclidean distance with the specified threshold.
[0098] Based on the differentiation between the intra-frame and inter-frame features of deepfake videos, this embodiment proposes a new deepfake detection method with a dual-network architecture. By constructing a dual-network structure, two improved CNN sub-networks are respectively used to extract intra-frame and inter-frame features, and a contrast loss function is used to capture the differentiation between the two; the contrast loss function is used to represent the relationship between the intra-frame features and the inter-frame features. In other studies based on forged audio and video, the contrast loss function is used to detect whether the audio and the picture in the forged audio and video are consistent; considering the latest public datasets, a thorough experimental evaluation is carried out. The evaluation results verify the superiority of this embodiment in face video detection, and also confirm the hypothesis of the differentiation between RGB images and optical flow; the network composition architecture in this embodiment also includes an intra-frame sub-network and an inter-frame sub-network, and these sub-networks are designed to learn distinguishable unimodal features through the cross-entropy loss. Subsequent experiments show that the network with the additional cross-entropy loss has a higher detection accuracy than the network that only uses the contrast loss.
[0099] Embodiment 2
[0100] Embodiment 2 of the present disclosure introduces a deepfake video detection system based on the difference between intra-frame and inter-frame features.
[0101] As Figure 5 shown, a deepfake video detection system based on the difference between intra-frame and inter-frame features includes:
[0102] An acquisition module configured to acquire the original data of the deepfake video;
[0103] An extraction module configured to extract intra-frame features and inter-frame features based on the acquired original data;
[0104] A calculation module configured to calculate the difference and Euclidean distance between the extracted intra-frame features and inter-frame features;
[0105] A detection module configured to complete the detection of the authenticity of the deepfake video according to the obtained difference and Euclidean distance.
[0106] The detailed steps are the same as those of the deepfake video detection method based on the difference between intra-frame and inter-frame features provided in Embodiment 1, and will not be elaborated here.
[0107] Embodiment 3
[0108] Embodiment 3 of the present disclosure provides a computer-readable storage medium.
[0109] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the deepfake video detection method based on the difference between intra-frame and inter-frame features as described in Embodiment 1 of the present disclosure.
[0110] The detailed steps are the same as those of the deepfake video detection method based on the difference between intra-frame and inter-frame features provided in Embodiment 1, and will not be elaborated here.
[0111] Embodiment 4
[0112] Embodiment 4 of the present disclosure provides an electronic device.
[0113] An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the deepfake video detection method based on the difference between intra-frame and inter-frame features as described in Embodiment 1 of the present disclosure.
[0114] The detailed steps are the same as those of the deepfake video detection method based on the difference between intra-frame and inter-frame features provided in Embodiment 1, and will not be elaborated here.
[0115] Although the specific embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present disclosure are still within the protection scope of the present disclosure.
Claims
1. A method for detecting deepfake videos based on the difference between intra-frame and inter-frame features, characterized in that, it includes: Obtain the original data of the deepfake video; After obtaining the original data of the deepfake video, perform uniform processing on the obtained original data, extract the original data of the obtained deepfake video into picture frames in units of frames, and suppress the interference of irrelevant video backgrounds except for the face part; Extract intra-frame features and inter-frame features based on the acquired raw data; perform feature optimization of the picture frames using bottleneck attention optimization, extract facial region features using the attention mechanism, and suppress irrelevant backgrounds; the input of the bottleneck attention module is the set of frame-level feature maps at the network bottleneck; the feature map consists of where C, H, and W represent the number of channels, height, and width of the feature map respectively; for a given input feature map the bottleneck attention module infers a 3D attention map the final output F of the bottleneck attention module ′ is calculated as follows: The specific implementation of the 3D attention map M(F) is achieved by calculating the channel attention and the spatial attention on two independent branches; Channel Attention M C (F) is calculated as follows: M C (F) = BN(MLP(AvgPool(F)); The calculation formula of spatial attention is as follows: Where c represents convolution operation, BN represents batch normalization operation, and the superscript represents the size of the convolution filter; Calculate the difference and Euclidean distance between the extracted intra-frame features and inter-frame features; According to the obtained difference and Euclidean distance, complete the detection of the authenticity of the deepfake video.
2. A method for detecting deepfake videos based on the difference between intra-frame and inter-frame features as described in claim 1, characterized in that, During the extraction of picture frames, use the face recognition library to locate and crop the face area in the picture frames, and save the pictures uniformly as 128×128.
3. A method for detecting deepfake videos based on the difference between intra-frame and inter-frame features as described in claim 1, characterized in that, During the extraction of intra-frame features, adopt RGB images, represent the inter-frame flow through dense optical flow, and focus on the eyes and / or lips based on the optical flow features to complete the extraction of intra-frame features.
4. A method for detecting deepfake videos based on the difference between intra-frame and inter-frame features as described in claim 1, characterized in that, During the extraction of inter-frame features, by calculating the offset vectors of all pixel points on the front and rear two-frame images, estimate the optical flow tracking of the moving offsets of all pixel points to complete the extraction of inter-frame features.
5. A method for detecting deepfake videos based on the difference between intra-frame and inter-frame features as described in claim 1, characterized in that, Adopt the cross-entropy loss function to calculate the loss functions of intra-frame features and inter-frame features respectively, calculate the weighted sum of the loss functions, and obtain the difference between intra-frame features and inter-frame features.
6. A deepfake video detection system based on the difference between intra-frame and inter-frame features, characterized in that, it includes: An acquisition module configured to obtain the original data of the deepfake video; after obtaining the original data of the deepfake video, perform uniform processing on the obtained original data, extract the original data of the obtained deepfake video into picture frames in units of frames, and suppress the interference of irrelevant video backgrounds except for the face part; An extraction module, which is configured to extract intra-frame features and inter-frame features based on the acquired original data; perform feature optimization of image frames using bottleneck attention optimization, extract facial region features using the attention mechanism, and suppress irrelevant backgrounds; the input of the bottleneck attention module is a set of frame-level feature maps at the network bottleneck; the feature map is composed of where C, H, and W represent the number of channels, height, and width of the feature map respectively; for a given input feature map the bottleneck attention module infers a 3D attention map the final output F of the bottleneck attention module ′ is calculated as follows: The specific implementation of the 3D attention map M(F) is achieved by calculating the channel attention and the spatial attention on two independent branches; Channel attention M C (F) is calculated as follows: M C (F) = BN(MLP(AvgPool(F)); The calculation formula of spatial attention is as follows: Where c represents convolution operation, BN represents batch normalization operation, and the superscript represents the size of the convolution filter; A calculation module configured to calculate the difference and Euclidean distance between the extracted intra-frame features and inter-frame features; A detection module configured to complete the detection of the authenticity of the deepfake video according to the obtained difference and Euclidean distance.
7. A computer-readable storage medium with a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the deepfake video detection method based on intra-frame and inter-frame feature differentiation described in any one of claims 1-5.
8. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, wherein, when the processor executes the program, it implements the steps in the deepfake video detection method based on intra-frame and inter-frame feature differentiation described in any one of claims 1-5.
Citation Information
Patent Citations
Method for detecting Deepfake video based on multi-feature fusion
CN111860414A
Multi-task learning AI face change video detection method for unbalanced data
CN114494953A