A Strong Generalization Deepfake Face Detection Method Based on Local Anomaly

By adopting a strong generalization method based on local anomalies in deep forged face detection, using the adaptive airspace rich model and noise flow input Resnet18 network, the problem of poor performance in cross-data set detection is solved, and effective generalization and high-precision detection of unknown data sets are achieved.

CN114926885BActive Publication Date: 2025-06-10NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210598271.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-30
Publication Date
2025-06-10
Estimated Expiration
2042-05-30

AI Technical Summary

Technical Problem

The existing fake face detection methods have poor performance when detecting across data sets, making it difficult to achieve effective generalization of unknown data sets.

Method used

A strong generalized deep fake face detection method based on local anomalies is adopted, and the adaptive airspace rich model and noise flow input backbone network Resnet18 is input to the backbone network Resnet18, combined with the local anomaly module to calculate the abnormal score to improve detection accuracy and improve performance on unknown data sets.

Benefits of technology

It significantly improves the generalization ability of the detection model in unknown domains, can achieve top detection accuracy without auxiliary data sets, and is suitable for a variety of forgery algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926885B_ABST
    Figure CN114926885B_ABST
Patent Text Reader

Abstract

The present invention discloses a strong generalization deepfake face detection method based on local anomalies. By adopting a second-order local anomaly learning module, anomalies in local regions are mined from the depth feature map of face images to achieve the detection of genuine and fake faces. First, the module decomposes the neighborhood of local features in different directions and distances, and then uses convolution to calculate the first-order and second-order local anomaly maps. A local enhancement module is adopted to improve the discrimination between local features in real and forged regions, thereby ensuring the accuracy of calculating local anomalies. An improved adaptive spatial rich model is used, and a learnable high-pass filter helps to mine subtle noise features, forming a two-stream structure with the local anomaly branch. Without pixel-level annotations and external synthetic data, the present invention uses a simple ResNet18 as the backbone network, and the average cross-database accuracies on the four sub-datasets of FF++ reach 91.28%, 97.20%, 97.32%, and 96.82%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of forged face detection, and in particular to a strong generalization deep forged face detection method based on local anomalies. Background Art

[0002] With the rapid development of network communication technology, the efficiency and scope of information dissemination have been greatly improved, and information security issues cannot be underestimated. Forged face videos emerging due to the rapid development of deep learning mainly use deep learning methods to tamper with the expressions, attributes, and identities of real faces, as well as generate completely non-existent fake faces. Currently, a series of software such as ZAO and Fake App have reached a cumulative download volume of over 100 million. At the same time, with the support of more than 1,500 open-source tools on the current Internet, the forged products with ultra-high resolution have achieved seamless editing visually, and even dynamic tampered products have exceeded the scope that can be distinguished by the human eye.

[0003] While the deep forgery technology realizes face replacement, it can fully fit the facial expressions and feature details of the face. It can not only replace the face but also control the changes in its facial expressions. Currently, many methods have emerged for the detection of forged faces, but most of the current forgery detection methods rely on training data-driven, and perform excellently on the datasets seen by the model, but perform poorly in cross-dataset detection. Due to the diversity of forgery methods, the biggest difficulty in the current face forgery algorithm lies in its generalization. Summary of the Invention

[0004] Object of the Invention: Aiming at the above problems, the present invention proposes a strong generalization deep forged face detection method based on local anomalies. Only by the information carried by the forged image itself can local anomaly signals be exposed, the detection accuracy of face forged videos be improved, and the performance of the detection algorithm on unknown datasets be greatly improved.

[0005] Technical Solution: To achieve the object of the present invention, the technical solution adopted by the present invention is: A strong generalization deep forged face detection method based on local anomalies, comprising the following steps:

[0006] (1) Decompose the true and false face videos in the training dataset into frames, convert the video format file into a continuous sequence of image frames, for the sequence of image frames, use a face detector to detect the face positions; crop the face frames for each image frame to obtain a continuous training set of face images;

[0007] (2) Input the face image obtained in (1) into the adaptive spatial rich model. By imposing a constraint condition on the filter kernel of the spatial rich model, make its central element remain -1 and the sum of the remaining elements remain 1. When performing filtering processing with this constrained spatial rich model, adaptively update each filtering element in the high-pass filter to extract high-frequency noise features;

[0008] (3) Input the face image obtained in (1) into the noise stream, that is, input the high-frequency noise features extracted by the adaptive spatial rich model in (2) into the first 3 blocks of the backbone network Resnet18 model, and calculate the binary cross-entropy loss;

[0009] (4) Input the face image obtained in (1) into the RGB stream, that is, input the face image in (1) into the first 3 blocks of the Resnet18 model in sequence, and perform local enhancement after each block;

[0010] (5) Extract the feature map from the last pooling layer of the RGB stream backbone network in (4) and input it into the local anomaly module to calculate the anomaly score. Combine the binary cross-entropy loss in (3) to obtain the final anomaly loss. Use backpropagation to update the local anomaly module and the backbone network to obtain the trained local anomaly detection network;

[0011] (6) Obtain the dataset to be detected, crop the face image, and input it into the trained local anomaly detection network for final face authenticity classification.

[0012] Further, in the step (2), the specific operation of the adaptive spatial rich model is as follows:

[0013] (2.1) Select 3 high-pass filters with central elements being 2, 4, and 12 respectively from 30 filters of the spatial rich model;

[0014] (2.2) Expand the selected 3 filters to dimension 3 in the channel dimension to match the 3 channels of the input image;

[0015] (2.3) Quantize the 3 filters respectively with 2, 4, and 12 as quantization factors, so that the central element value is -1 and the sum of the remaining element values is 1;

[0016] (2.4) After each backpropagation in network training, reset its central element value to -1 and the sum of the remaining elements to 1.

[0017] Further, in the step (5), the operation of the second-order local anomaly model is as follows:

[0018] (5.1) The source feature map X = F(I) ∈ R extracted from the pooling layer H×W×CDivide it into n×n image patches in the direction of the channel dimension. Here, I represents the input image, F represents the backbone network, H, W, and C represent the length, width, and number of channels of the feature map respectively. The size of each image patch is h×w, where h and w represent the length and width of the image patch respectively, and h = H / / n, w = W / / n;

[0019] (5.2) Let (m, n) be the horizontal and vertical coordinate indices of the image patch. Then, for an image patch located at (m, n), it is represented by X mn Calculate the similarity between this image patch and the nearest neighbor image patches in the four directions of directly above, directly below, directly left, and directly right of it respectively;

[0020] (5.3) Incorporate the image patches that are the second nearest neighbors in the four directions into the calculation range of similarity. That is, for an image patch X mn Calculate the similarity pairwise for the eight pairs of image patches {X mn , X m+i,n+j}, where i ∈ {0, ±1, ±2}, j ∈ {0, ±1, ±2};

[0021] (5.4) Design a 1×1 convolution operation f to project the feature map into a unified new dimension space, then use a fully connected layer to output the true / false binary classification probability of the feature, and finally introduce the binary cross-entropy loss to constrain the convolutional network:

[0022]

[0023] (5.5) Under the supervision of the true and false labels in the dataset, take the average of the similarities of each pair of graphic patches in the eight pairs of image patches described in (5.3), and use the probability output by the fully connected layer in (5.4) as the result of true and false classification.

[0024] Furthermore, use MTCNN to perform face detection on the image frame sequence frame by frame. MTCNN will return 3 sets of return values:

[0025] 1) The probability that the image contains a face; 2) The position information of the face rectangle box, represented by (x, y, w, h), where x and y represent the horizontal and vertical coordinates of the upper left corner of the detected face rectangle with the upper left corner point of the image as the origin, and w and h represent the width and height of the rectangle box respectively; 3) The positions of 5 key points of the detected face;

[0026] Set a face probability threshold. When the probability of detecting a face returned by MTCNN is lower than the threshold, do not crop the image; for the detected face, first calculate the center coordinate point P of the face box according to the following formula center :

[0027]

[0028] With P center as the center coordinate, taking the longer side of w and h as the reference, expand it by α times, and the expansion formula is as follows:

[0029]

[0030] where Rect new represents the position information of the expanded face rectangle box, and its four elements also represent the horizontal and vertical coordinates of the upper left corner of the new rectangle box and its width and height respectively.

[0031] Furthermore, in step (4), local enhancement is performed after each block, and deformation is performed after the feature map is divided into blocks after each convolutional block. Specifically, the following operations are defined in the convolutional network:

[0032]

[0033] where x represents the input feature, y represents the output feature of the same size as x, both represented in vector form; i is the index coordinate of the image block whose similarity needs to be calculated, j is the index coordinate of the surrounding image block associated and calculated; C(x) is the regularization coefficient; g is a linear embedding layer, which is a unary function; f represents the calculation operation of the similarity between i and j pairwise, which is a Gaussian function, expressed as follows:

[0034]

[0035] where is the pairwise similarity calculation; the above Gaussian function is extended to calculate the similarity in the new dimensional space after projection, as follows:

[0036]

[0037] where θ(x i ) = W θ x i and φ(x j ) = W φ x j are two embedding layers;

[0038] The linear embedding layer g is expressed as: g(x j ) = W g x j , where W g is a learned weight matrix, which is a 1×1×1 convolutional operation in space; is also set as the regularization coefficient; through this operation, the receptive field is concentrated on the extraction of the required local information.

[0039] Advantageous effects: Compared with the prior art, the technical solution of the present invention has the following advantageous technical effects:

[0040] The method for detecting deepfake faces with strong generalization based on local anomalies disclosed by the present invention can greatly improve the generalization ability of the detection model in the unknown domain. Compared with models with deeper network layers such as Xception, the present invention only requires the shallow network of Resnet18 as the backbone architecture, and is not limited to the type of forgery algorithm. It can be generalized to the unknown domain without an auxiliary dataset and achieve top detection accuracy. At the same time, the present invention can also be arbitrarily linked to a deep convolutional model as a module and has an effect on various forgery algorithms. Description of the drawings

[0041] Figure 1 is the overall flowchart of the method of the present invention;

[0042] Figure 2 is the flowchart of the Adaptive Spatial Rich Model (ASRM);

[0043] Figure 3 is the main structure of the backbone network Resnet18;

[0044] Figure 4 is the schematic diagram of the local enhancement module structure;

[0045] Figure 5 is the schematic diagram of the local anomaly module structure. Detailed implementation manners

[0046] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0047] As Figure 1 shown is the overall flowchart of a method for detecting deepfake faces with strong generalization based on local anomalies of the present invention, and the steps are as follows:

[0048] (1) Decompose the frames of the true and false face videos in the training dataset, convert the video format file into a continuous sequence of image frames, and for the sequence of image frames, use the MTCNN face detector to detect the face positions. Crop the face frames for each image frame to obtain a continuous training set of face images.

[0049] Specifically, when using MTCNN to perform face detection on the sequence of image frames frame by frame, MTCNN will return three sets of return values: 1) the probability that the image contains a face; 2) the position information of the face rectangle frame, represented by (x, y, w, h), where x and y represent the horizontal and vertical coordinates of the upper left corner of the detected face rectangle with the upper left corner point of the image as the origin, and w and h represent the width and height of the rectangle frame respectively; 3) the positions of the 5 key points of the detected face.

[0050] The present invention sets the face probability threshold to 0.85, that is, when the probability of detecting a face returned by MTCNN is lower than 0.85, the image is not cropped.

[0051] For the detected face, first calculate the center coordinate point P of the face bounding box according to the following formula center :

[0052]

[0053] Taking P center as the center coordinate, and using the longer side of w and h as a reference, expand it by α times. The expansion formula is as follows:

[0054]

[0055] Among them, Rect new represents the position information of the expanded face rectangle box, and the four elements also represent the horizontal and vertical coordinates of the upper left corner of the new rectangle box and its width and height respectively.

[0056] (2) Input the face image obtained in (1) into the Adaptive Spatial Rich Model (ASRM). By imposing a constraint condition on the filter kernel of the Spatial Rich Model (SRM), make its central element remain -1, and the sum of the remaining elements remain 1; when performing filtering processing on the constrained spatial rich model, adaptively update each filtering element in the high-pass filter to extract high-frequency noise features.

[0057] The structure of the Adaptive Spatial Rich Model (ASRM) is as Figure 2 shown; select 3 high-pass filters with central elements of 2, 4, and 12 respectively from the 30 filters of the Spatial Rich Model (SRM), and expand the selected 3 filters to a dimension of 3 in the channel dimension (that is, these filters are repeated 3 times in the RGB channels) to match the 3 channels of the input image, as Figure 2 shown. Quantize the 3 filters with 2, 4, and 12 as quantization factors respectively, so that their central element values are -1 and the sum of the remaining element values is 1, so that the filters can maintain the characteristics of high-pass filters when learning through the network.

[0058] (3) Input the face image obtained in (1) into the noise stream, that is, input the high-frequency noise features extracted by ASRM in (2) into the first 3 blocks of the backbone network Resnet18 model, and calculate the Binary CrossEntropy Loss (BCE Loss). The main structure of the backbone network Resnet 18 is as Figure 3 shown.

[0059] (4) Input the face image obtained in (1) into the RGB stream, that is, input the face image in (1) into the first 3 blocks of the Resnet18 model in sequence, and perform local enhancement (LEM) after each block. As Figure 4 shown, in order to alleviate the problem of too large receptive field in the deep network and ensure the effectiveness of the local information of the extracted feature map, the feature map after each convolutional block is deformed after being divided into blocks. More specifically, the following operations are defined in the convolutional network:

[0060]

[0061] Among them, x represents the input feature, y represents the output feature of the same size as x, both represented in vector form. i is the index coordinate of the image block for which the similarity needs to be calculated, and j is the index coordinate of all possible surrounding image blocks that may be associated and calculated. C(x) is the regularization coefficient; g is a linear embedding layer, which is a unary function; f represents the calculation operation of the similarity between i and j pairwise, and it is essentially a Gaussian function, expressed as follows:

[0062]

[0063] Among them, is the pairwise similarity calculation, and this calculation method is easier to implement in the deep convolutional neural network. In order to facilitate the calculation of the similarity in the new dimensional space after projection, the above Gaussian function is simply extended as follows:

[0064]

[0065] Among them, θ(x i ) = W θ x i and φ(x j ) = W φ x j are two embedding layers. g is a linear embedding layer, which is essentially a unary function g(x j ) = W g x j , where W g is a weight matrix to be learned, which is a 1×1×1 convolutional operation in space. Also set as the regularization coefficient. Through this operation, the receptive field is made more focused on extracting the required local information.

[0066] (5) Extract the feature map from the last pooling layer of the RGB stream backbone network in (4) and input it into the local anomaly module (Secondorder anomaly) to calculate the anomaly score. Combine it with the binary cross-entropy loss in (3) to obtain the final anomaly loss (Anomaly Loss). Use backpropagation to update the local anomaly module and the backbone network to obtain the trained local anomaly detection network.

[0067] As Figure 5 shown, first, divide the source feature map X = F(I) ∈ R H×W×C extracted from the pooling layer into n×n image patches along the channel dimension direction, where I represents the input image, F represents the backbone network, H, W, and C represent the length, width, and number of channels of the feature map respectively, the size of each image patch is h×w, and h and w represent the length and width of the image patch respectively, where h = H / / n and w = W / / n.

[0068] Next, start to extract local features. Let (m, n) be the horizontal and vertical coordinate indices of the image patch. Then, for an image patch located at the position (m, n), it is represented by X mn and calculate the similarity between this image patch and its four nearest neighbor image patches directly above, directly below, directly to the left, and directly to the right.

[0069] To expand the local information of an image patch, include the four next-nearest neighbor image patches in the calculation range of similarity. That is, for an image patch X mn and its eight image patches above, below, left, right, second above, second below, second left, and second right to form eight pairs {X mn , X m+i,n+j}, i ∈ {0, ±1, ±2}, j ∈ {0, ±1, ±2}, and calculate the similarity pairwise.

[0070] In the figure X, F(m, n) represents the coordinates of an image patch in the H and W directions after a series of convolutional operations F. F(m + i, n + j) represents its surrounding image patches.

[0071] In addition, to facilitate the similarity calculation, a 1×1 convolutional operation f is designed to project the feature map into a unified new dimension space. From a spatial perspective, the pairs formed between these image patches are actually three-dimensional vectors, and their sizes are 2×1×C or 1×2×C. Then, use the fully connected layer to output the true / false binary classification probability of the features. Finally, introduce the binary cross-entropy loss to constrain the convolutional network.

[0072]

[0073] Since the pre-training of the present invention is supervised, each convolutional layer will finally output a two-dimensional grayscale image. Under the supervision of true and false labels in the dataset, the average similarity of each pair of graphic blocks in the above eight image block pairs is taken, and the probability output by the fully connected layer is used as the result of true and false classification.

[0074] (6) Obtain the dataset to be detected, crop the face image, and input it into the trained local anomaly detection network for final face true and false classification.

[0075] This embodiment is trained and tested on the large-scale forged face video dataset FF++ (including four subsets: DeepFake, Face to Face, Face Shifter, and Neural Texture). The basic information of the FF++ dataset is shown in Table 1. This embodiment tests the influence of the change of different sequence lengths N on the detection accuracy, and compares it with the well-known spatio-temporal feature extraction model CNN-LSTM. The relevant results on DFDC-P are shown in Table 2, and the results of Celeb-DF are shown in Table 3. It can be found that on both datasets, as the sequence length increases, the accuracy also increases until the number of frames reaches 15 frames, and regardless of the size of N, the accuracy of the proposed scheme of the present invention is always higher than that of the well-known CNN-LSTM model, further proving the superiority of this scheme in time-domain feature fusion.

[0076] Table 1 Basic information of the two datasets

[0077] Dataset Real Video / Fake Video Total Number of Frames (in Millions) Resolution FF++ 1000 / 4000 358.8 / 2116.8 Multi-scale DFDC-P 1131 / 4113 88.4 / 1783.3 180p - 2160p Celeb-DF 890 / 5639 358.8 / 2116.8 Multi-scale

[0078] Table 2 Influence of different numbers of frames on detection accuracy on DFDC-P

[0079] Sequence Length 3 6 9 12 15 18 This Solution 84.76 83.14 82.75 85.28 84.81 83.19 CNN-LSTM 79.08 80.50 80.28 80.78 81.91 79.75

[0080] Table 3 Intra-database cross-test accuracy of FF++

[0081] Sequence Length 3 6 9 12 15 18 This Solution 95.86 96.27 96.17 97.12 96.91 95.28 CNN-LSTM 95.22 95.06 95.13 96.53 96.38 95.28

[0082] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. A strong generalization deepfake face detection method based on local anomalies, characterized in that: The method includes the following steps: (1) Decompose the frames of the true and false face videos in the training dataset, convert the video format file into a continuous sequence of image frames, for the image frame sequence, use a face detector to detect the face positions; crop the face frames for each image frame to obtain a continuous training set of face images; (2) Input the face images obtained in (1) into the adaptive spatial domain rich model. By imposing a constraint condition on the filter kernel of the spatial domain rich model, make its central element remain -1 and the sum of the remaining elements remain 1; when the constrained spatial domain rich model performs filtering, adaptively update each filtering element in the high-pass filter to extract high-frequency noise features; (3) Input the face images obtained in (1) into the noise stream, that is, input the high-frequency noise features extracted by the adaptive spatial domain rich model in (2) into the first 3 blocks of the backbone network Resnet18 model, and calculate the binary cross-entropy loss; (4) Input the face images obtained in (1) into the RGB stream, that is, input the face images in (1) into the first 3 blocks of the Resnet18 model in sequence, and perform local enhancement after each block; (5) Extract the feature map from the last pooling layer of the RGB stream backbone network in (4) and input it into the local anomaly model to calculate the anomaly score, combine the binary cross-entropy loss in (3) to obtain the final anomaly loss, and use backpropagation to update the local anomaly model and the backbone network to obtain a trained local anomaly detection network; (6) Obtain the dataset to be detected, crop the face images, and input them into the trained local anomaly detection network for final face true / false classification.

2. The strong generalization deepfake face detection method according to claim 1, characterized in that: In the step (2), the specific operation of the adaptive spatial domain rich model is as follows: (2.1) Select 3 high-pass filters with central elements of 2, 4, and 12 respectively from 30 filters of the spatial domain rich model; (2.2) Expand the selected 3 filters to a dimension of 3 in the channel dimension to match the 3 channels of the input image; (2.3) Quantize the 3 filters respectively with 2, 4, and 12 as quantization factors, so that the central element value is -1 and the sum of the remaining element values is 1; (2.4) After each backpropagation of network training, reset its central element value to -1 and the sum of the remaining elements to 1.

3. The strong generalization deepfake face detection method according to claim 1 or 2, characterized in that: In the step (5), the local anomaly model is a second-order local anomaly model, and the operation is as follows: (5.1) The source feature map X = F(I) ∈ R extracted from the pooling layer is divided into n×n image patches along the channel dimension direction, where I represents the input image, F represents the backbone network, H, W, and C represent the length, width, and number of channels of the feature map respectively, the size of each image patch is h×w, and h and w represent the length and width of the image patch respectively, where h = H / n and w = W / n; H×W×C That is, each image patch is obtained by dividing the feature map X along the channel dimension direction into n×n non-overlapping sub-regions, and each sub-region is an image patch with a size of h×w. (5.2) Let (m, n) be the horizontal and vertical index of the image block. Then, an image block located at position (m, n) is represented by X mn Calculate the similarity between this image block and its four nearest neighbor image blocks directly above, directly below, directly to the left, and directly to the right, respectively; (5.3) Incorporate the four azimuth second-nearest neighbor image patches into the calculation scope of similarity, that is, for an image patch X mn of the eight image patch pairs {X mn , X m+i,n+j}, i ∈ {0, ±1, ±2}, j ∈ {0, ±1, ±2}, calculate the similarity pairwise; (5.4) Design a 1×1 convolution operation f to project the feature map into a unified new dimension space, then use the fully connected layer to output the true / false binary classification probability of the feature, and finally introduce the binary cross-entropy loss to constrain the convolutional network: Among them, is the cross-entropy loss, and y i represents the true value of the true / false label, represents the predicted value of the true / false label output by the fully connected layer; (5.5) Under the supervision of true and false labels in the dataset, the similarity of each pair of graphic blocks in the eight image block pairs described in (5.3) is averaged, and the probability output by the fully connected layer in (5.4) is used as the result of true and false classification.

4. The strong generalization deepfake face detection method according to claim 1 or 2, characterized in that: using MTCNN to perform face detection on the image frame sequence frame by frame, and MTCNN will return 3 sets of return values: 1) The probability that the image contains a face; 2) The position information of the face rectangle box, represented by (x, y, w, h), where x and y represent the horizontal and vertical coordinates of the upper left corner of the detected face rectangle with the upper left corner point of the image as the origin, and w and h represent the width and height of the rectangle box respectively; 3) The positions of 5 key points of the detected face; Set a face probability threshold. When the probability of detecting a face returned by MTCNN is lower than the threshold, do not crop the image. For the detected face, first calculate the center coordinate point P of the face bounding box according to the following formula center : With P center as the center coordinate, taking the longer side of w and h as the reference, expand it by α times, and the expansion formula is as follows: Among them, Rect new represents the position information of the expanded face rectangular frame, and its four elements also represent the horizontal and vertical coordinates of the upper left corner of the new rectangular frame and its width and height respectively.

5. The strong generalization deepfake face detection method according to claim 1 or 2, characterized in that: in the step (4), local enhancement is performed after each block, and deformation is performed after the feature map is divided into blocks after each convolutional block. Specifically, the following operations are defined in the convolutional network: where x represents the input feature, y represents the output feature of the same size as x, both represented in vector form; i is the index coordinate of the image block whose similarity needs to be calculated, and j is the index coordinate of the surrounding image block associated and calculated; C(x) is the regularization coefficient; g is a linear embedding layer, which is a unary function; f represents the calculation operation of the similarity between i and j pairwise, which is a Gaussian function, and is expressed as follows: Among them, is the similarity calculation between point to point; the above Gaussian function is extended to calculate the similarity in the new dimensional space after projection, as follows: where, θ(x i ) = W θ x i and φ(x j ) = W φ x j are two embedding layers; The linear embedding layer g is expressed as: g(x j ) = W g x j , where W g is a learnable weight matrix, which is a 1×1×1 convolution operation in space; is also set as the coefficient for regularization; through this operation, the receptive field is concentrated on the extraction of the required local information.

Citation Information

Patent Citations

  • Method for detecting forged face video based on deep learning

    CN112395943A

  • Deep fake face data identification method

    CN113536990A