Face forgery detection method and system based on graph convolutional neural network

By constructing a space-time composite graph based on graph convolution neural network, the problem of insufficient detection accuracy and generalization ability of face forgery in the prior art is solved, and efficient and accurate detection of face forgery is achieved.

CN120299097AInactive Publication Date: 2025-07-11NANJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510423206.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When faced with the continuous update and iterative counterfeiting technology, the existing face forgery detection methods lack detection accuracy and generalization capabilities, making it difficult to effectively identify complex and changeable counterfeiting patterns.

Method used

Using a method based on graph convolution neural network, a space-time composite graph combining spatial graphs and timing graphs is constructed. Features are generated through the encoder, and feature learning is performed by the space-time composite graph convolution module, efficient detection is performed using graph convolution technology, and authenticity prediction is performed in combination with classifiers.

Benefits of technology

It realizes efficient and accurate face forgery detection, has strong generalization capabilities, and improves detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299097A_ABST
    Figure CN120299097A_ABST
Patent Text Reader

Abstract

The invention discloses a face forgery detection method and system based on a graph convolutional neural network. The method comprises the following steps: acquiring a face video of which the authenticity is to be detected; inputting the face video of which the authenticity is to be detected into the trained detection model to obtain a judgment result of whether the video is forged or not; wherein the detection model comprises an encoder used for generating features of each frame of image according to a face video to be subjected to authenticity detection; the space-time composite diagram construction module is used for constructing a space-time composite diagram in which a space diagram and a time sequence diagram are combined according to the characteristics of each frame of image; the space-time composite graph convolution module is used for carrying out feature learning according to the space-time composite graph to obtain classification features; and the classifier is used for predicting the authenticity result of the face video according to the classification features. According to the method, the space-time composite diagram combining the space diagram and the time sequence diagram is constructed according to the space-time consistency requirement of the face video, so that efficient and accurate face counterfeiting detection with high generalization capability is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of face forgery detection, and particularly relates to a face forgery detection method and system based on a graph convolutional neural network. Background Art

[0002] With the rapid development of multimedia technology, the number of images and videos with faces as the theme is increasing day by day, and face-related technologies have also been widely applied in fields such as social security (such as face recognition) and audio-visual entertainment (such as face enhancement). However, with the progress of advanced technologies, the phenomenon of maliciously using these technologies for face forgery has become increasingly rampant. Through face deep forgery technologies (such as generative adversarial networks GAN, diffusion models, etc.), face swapping, attribute editing, and synthesis of face images and videos are realized, and then extremely realistic fake face images and videos are created. These forged contents can often pass off as genuine and spread rapidly on social media platforms, posing a serious threat to personal privacy, public safety, and social trust. For example, false celebrity videos may be used to mislead public opinion and spread false information; forged identity images may be used for fraud, harming the interests of others. It can be seen that the illegal use of face forgery technology will bring great hidden dangers to personal property and social harmony and stability.

[0003] Therefore, face forgery detection technology has emerged, aiming to distinguish the authenticity of face images and videos. Although existing face forgery detection methods can identify some forged samples to a certain extent, in the face of continuously updated forgery technologies, traditional methods gradually expose limitations. In addition, detection methods based on manual feature extraction are difficult to adapt to complex and changeable forgery patterns. With the rise of deep learning models, existing face forgery detection models based on deep learning still have deficiencies in detection accuracy and generalization ability due to the lack of effective mining of the complex structure and semantic information of face images. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a face forgery detection method and system based on a graph convolutional neural network. Starting from the spatio-temporal consistency requirements of face videos, a spatio-temporal composite graph combining a spatial graph and a temporal graph is constructed to achieve efficient, accurate, and strongly generalized face forgery detection.

[0005] The present invention provides the following technical solutions:

[0006] In the first aspect, a face forgery detection method based on a graph convolutional neural network is provided, including:

[0007] Obtain a face video to be detected for authenticity;

[0008] Input the face video to be detected for authenticity into the trained detection model to obtain the judgment result on whether the video is forged;

[0009] Among them, the detection model includes:

[0010] An encoder, which is used to generate the features of each frame of image according to the face video to be detected for authenticity;

[0011] A spatio-temporal composite graph construction module, which is used to construct a spatio-temporal composite graph combining a spatial graph and a temporal graph according to the features of each frame of image;

[0012] A spatio-temporal composite graph convolution module, which is used to perform feature learning according to the spatio-temporal composite graph to obtain classification features;

[0013] A classifier, which is used to predict the authenticity result of the face video according to the classification features.

[0014] Further, the encoder includes a plurality of convolutional units and a plurality of downsampling units, and the convolutional units and the downsampling units are connected in series alternately.

[0015] Further, the main body of the spatio-temporal composite graph is a temporal graph; the temporal graph uses the features of each frame of image as nodes and the temporal front-back relationship as the basis for node connection; each node in the temporal graph corresponds to the spatial graph of the features of each frame of image; the spatial graph uses the features of the face component regions in one frame of image as nodes and the layout characteristics of the face components as the basis for node connection.

[0016] Further, the spatio-temporal composite graph construction module includes a temporal graph construction unit and a spatial graph construction unit;

[0017] The construction of the spatio-temporal composite graph according to the features of each frame of image includes:

[0018] Input into the spatio-temporal composite graph construction module, where f represents the feature vector obtained after being processed by the encoder, T represents the number of video frames, c represents the basic number of feature channels, h represents the feature height, and w represents the feature width;

[0019] The temporal graph construction unit obtains a feature vector with the number of video frames T as the first dimension through feature recombination , and then obtains the temporal graph G t , the temporal graph G t contains T nodes with a feature size of respectively corresponding to the features of each frame;

[0020] The spatial graph construction unit obtains a feature vector with the total number of all positions in each frame of image as the first dimension , and then obtains the spatial graph G s, the spatial graph G s contains nodes with a feature size of c corresponding to the facial component region features in each frame of the image respectively.

[0021] Furthermore, the spatio-temporal composite graph convolution module includes multiple spatial graph convolution blocks and multiple spatio-temporal convolution blocks, which are alternately connected in series;

[0022] The spatial graph convolution block is used to mine spatial features of the input spatial graph, and the spatio-temporal convolution block is used to mine temporal features of the input temporal graph.

[0023] Furthermore, both the spatial graph convolution block and the temporal graph convolution block include 3 graph convolution layers, 3 graph normalization layers, and 2 activation layers, which are alternately connected in series;

[0024] The graph convolution layer is used to extract features in the graph data by aggregating the feature information of nodes and their neighbors, capture the local structure and relationship information between nodes in the graph, and provide deep feature representations. The function formula used by the graph convolution layer is as follows:

[0025] ;

[0026] where and represent the node feature matrices of the l-th and l+1-th layers, represents the adjacency matrix of the graph, represents 's degree matrix, W l represents the learnable weight of the l-th layer, represents the activation function;

[0027] The graph normalization layer is used to adjust the distribution of the data to a preset range. The function formula used by the graph normalization layer is as follows:

[0028] ;

[0029] where represents the normalized feature vector of node i, h i represents the feature vector of node i, represents the mean of all feature dimensions of node i, represents the variance of all feature dimensions of node i, represents a constant used to avoid a zero denominator;

[0030] The activation layer is used to enhance the expressive power of the model. The activation function used by the activation layer includes ReLU.

[0031] Furthermore, the classifier includes a fully connected layer, the input size of the fully connected layer is the size of the classification feature, and the output size of the fully connected layer is 2, respectively representing the probabilities of predicting the face video as real and fake.

[0032] Furthermore, the training method of the detection model includes:

[0033] Obtain a training set, the training set includes real face video data and face video data generated by forgery techniques, as well as corresponding labels, where the label values are 0 and 1, 0 represents real data, and 1 represents forged false data;

[0034] Input the real face video data and the face video data generated by forgery techniques into a pre-constructed detection model to obtain a model prediction result;

[0035] Calculate the cross-entropy loss function according to the model prediction result and the corresponding label;

[0036] Based on the Adam optimizer, for each training batch, first use forward propagation to calculate the cross-entropy loss function, then use backward propagation to calculate the partial derivatives of each weight parameter, and finally use the gradient descent method to update the weight parameters. Repeat this step until the cross-entropy loss function reaches the minimum to obtain a trained detection model.

[0037] Furthermore, the obtaining of the training set includes downloading a public dataset and constructing a private dataset;

[0038] The public datasets include FaceForensics++, DeepFake Detection Challenge;

[0039] The construction of the private dataset includes: obtaining face videos from video websites or collecting face videos; detecting faces using the Dlib tool; aligning faces using affine transformation; setting a fixed margin for the face area and then cropping; scaling the cropped face images to a fixed pixel size.

[0040] Furthermore, the expression of the cross-entropy loss function is as follows:

[0041] ;

[0042] where y represents the corresponding label, represents the model prediction result.

[0043] In a second aspect, a face forgery detection system based on a graph convolutional neural network is provided, including:

[0044] A data acquisition module for acquiring a face video to be detected for authenticity.

[0045] A detection module for inputting the face video to be detected for authenticity into a trained detection model to obtain a judgment result on whether the video is forged.

[0046] Among them, the detection model includes:

[0047] An encoder for generating features of each frame image according to the face video to be detected for authenticity.

[0048] A spatio-temporal composite graph construction module for constructing a spatio-temporal composite graph according to the features of each frame image.

[0049] A spatio-temporal composite graph convolution module for performing feature learning according to the spatio-temporal composite graph to obtain classification features.

[0050] A classifier for predicting the authenticity result of the face video according to the classification features.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] The face forgery detection method and system based on a graph convolutional neural network provided by the present invention, first, based on the encoder of the detection model, generate features of each frame image according to the face video to be detected for authenticity; then, based on the spatio-temporal composite graph construction module, starting from the spatio-temporal consistency requirement of the face video, utilize the control of forgery details by spatial features and the requirement of front-back coherence by temporal features to construct a spatio-temporal composite graph combining a spatial graph and a temporal graph; secondly, based on the spatio-temporal composite graph convolution module, perform feature learning according to the spatio-temporal composite graph to obtain classification features, fully pay attention to the interaction between the spatial graph and the temporal graph, and utilize graph convolution technology to achieve efficient face forgery detection. While ensuring the light weight of the model, it can achieve high detection efficiency, detection accuracy and strong generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic structural diagram of the face forgery detection method based on a graph convolutional neural network in an embodiment of the present invention;

[0054] Figure 2 is a schematic structural diagram of the detection model in an embodiment of the present invention;

[0055] Figure 3 is a schematic structural diagram of the spatio-temporal composite graph construction module in an embodiment of the present invention;

[0056] Figure 4 is a schematic structural diagram of the spatio-temporal composite graph convolution module in an embodiment of the present invention;

[0057] Figure 5It is a schematic structural diagram of the spatial graph convolution block and the spatio-temporal convolution block in the embodiments of the present invention. Detailed implementation manners

[0058] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0059] The term "and / or" only describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.

[0060] Embodiment 1

[0061] As Figure 1 shown, this embodiment provides a face forgery detection method based on a graph convolutional neural network, and the steps are as follows:

[0062] Step 1: Obtain a face video to be detected for authenticity.

[0063] Step 2: Input the face video to be detected for authenticity into the trained detection model to obtain a judgment result on whether the video is forged.

[0064] As Figure 2 shown, the detection model includes:

[0065] An encoder for generating features of each frame of image according to the face video to be detected for authenticity.

[0066] A spatio-temporal composite graph construction module for constructing a spatio-temporal composite graph combining a spatial graph and a temporal graph according to the features of each frame of image.

[0067] A spatio-temporal composite graph convolution module for performing feature learning according to the spatio-temporal composite graph to obtain classification features.

[0068] A classifier for predicting the authenticity result of the face video according to the classification features.

[0069] The function of the encoder can be implemented by using an existing mainstream convolutional neural network or its optimized version, such as ResNet, SCNet, Xception, etc. Its architecture is composed of multiple convolutional units and multiple downsampling units, and the convolutional units and the downsampling units are alternately connected in series.

[0070] The main body of the spatio-temporal composite graph is a time-series graph; the time-series graph uses the features of each frame of image as nodes and the time-series relationship before and after as the basis for node connection; each node in the time-series graph corresponds to the spatial graph of the features of each frame of image; the spatial graph uses the features of the face component regions in a frame of image as nodes and the layout characteristics of the face components as the basis for node connection.

[0071] As Figure 3 shown, in this embodiment, the spatio-temporal composite graph construction module includes 1 time-series graph construction unit and 1 spatial graph construction unit; constructing the spatio-temporal composite graph according to the features of each frame of image includes:

[0072] Input into the spatio-temporal composite graph construction module, where f represents the feature vector obtained after being processed by the encoder, T represents the number of video frames, c represents the basic number of feature channels, h represents the feature height, and w represents the feature width;

[0073] The time-series graph construction unit obtains a feature vector with the number of video frames T as the first dimension through feature recombination , and then obtains the time-series graph G t , the time-series graph G t contains T nodes with a feature size of respectively corresponding to the features of each frame;

[0074] The spatial graph construction unit obtains a feature vector with the total number of all positions in each frame of image as the first dimension through feature recombination , and then obtains the spatial graph G s , the spatial graph G s contains nodes with a feature size of c respectively corresponding to the face component region features in each frame of image.

[0075] As Figure 4 shown, the spatio-temporal composite graph convolution module includes multiple spatial graph convolution blocks and multiple spatio-temporal convolution blocks, and the spatial graph convolution blocks and the spatio-temporal convolution blocks are connected in series alternately; the spatial graph convolution block is used to mine spatial features of the input spatial graph G s , and the spatio-temporal convolution block is used to mine temporal features of the input time-series graph G t . In this embodiment, the spatio-temporal composite graph convolution module includes 3 spatial graph convolution blocks and 3 spatio-temporal convolution blocks.

[0076] As Figure 5 shown, both the spatial graph convolution block and the time-series graph convolution block include 3 graph convolution layers, 3 graph normalization layers, and 2 activation layers, and the graph convolution layers, graph normalization layers, and activation layers are connected in series alternately.

[0077] The graph convolutional layer is used to extract features from graph data by aggregating the feature information of nodes and their neighbors, capture the local structure and relationship information between nodes in the graph, and provide deep feature representations. The functional formula adopted by the graph convolutional layer is as follows:

[0078] ;

[0079] Among them, and represent the node feature matrices of the l-th and l+1-th layers, represents the adjacency matrix of the graph, represents 's degree matrix, W l represents the learnable weight of the l-th layer, represents the activation function.

[0080] The graph normalization layer is used to adjust the distribution of data to a preset range, making the training of the model more stable and accelerating the convergence of the training process. The functional formula adopted by the graph normalization layer is as follows:

[0081] ;

[0082] Among them, represents the normalized feature vector of node i, h i represents the feature vector of node i, represents the mean of all feature dimensions of node i, represents the variance of all feature dimensions of node i, represents a constant used to avoid division by zero.

[0083] The activation layer is used to enhance the expressive power of the model, and common activation functions such as ReLU are adopted by the activation layer.

[0084] The classifier includes 1 fully connected layer. The input size of the fully connected layer is the size of the classification features, and the output size of the fully connected layer is 2, representing the probabilities of predicting the face video as real and forged respectively.

[0085] The training method of the detection model includes the following steps:

[0086] Step A: Obtain a training set, where the training set includes real face video data V r and face video data V f generated by forgery techniques, as well as the corresponding labels y. Among them, the label values of V r and V f are 0 and 1 respectively, that is, 0 represents real data and 1 represents forged false data.

[0087] Specifically, the obtaining of the training set includes downloading public datasets and constructing private datasets;

[0088] The public datasets include FaceForensics++ (FF++), DeepFake Detection Challenge (DFDC), etc.;

[0089] The construction of the private dataset includes: obtaining face videos from video websites or collecting face videos; detecting faces using the Dlib tool; aligning faces using affine transformation; setting a fixed margin value (such as 1.3) for the face area and then cropping; scaling the cropped face images to a fixed pixel size. In this embodiment, both the length and width are set to 256 pixel values.

[0090] Step B: Input the real face video data V r and the face video data V f generated by forgery technology into the pre-constructed detection model to obtain the model prediction result .

[0091] Step C: Calculate the cross-entropy loss function according to the model prediction result and the corresponding label; the expression of the cross-entropy loss function is as follows:

[0092] ;

[0093] where y represents the corresponding label, represents the model prediction result.

[0094] Step D: Based on the Adam optimizer, for each training batch, first use forward propagation to obtain the cross-entropy loss function, then use backward propagation to obtain the partial derivatives of each weight parameter, and finally use the gradient descent method (such as the stochastic gradient descent method) to update the weight parameters. Repeat this step until the cross-entropy loss function reaches the minimum to obtain the trained detection model.

[0095] Embodiment 2

[0096] Using the method in Embodiment 1, the index result of the area AUC under the ROC curve obtained by testing the trained detection model on the public test dataset Celeb-DF-v2 is 86.16%, which significantly exceeds that of existing mainstream models, such as 75.58% of STIL, 84.71% of TALL++-EffNetB4, and 83.58% of MRL.

[0097] The metric result of the area AUC enclosed by the ROC curve obtained from testing on the publicly available test dataset DFDC is 73.12%, significantly exceeding that of existing mainstream models, such as 67.88% of STIL, 68.37% of TALL++-EffNetB4, and 71.53% of MRL.

[0098] Example 3

[0099] Based on the same inventive concept as Example 1, this example provides a face forgery detection system based on a graph convolutional neural network, including:

[0100] A data acquisition module for acquiring a face video whose authenticity needs to be detected.

[0101] A detection module for inputting the face video whose authenticity needs to be detected into a trained detection model to obtain a judgment result on whether the video is forged.

[0102] Among them, the detection model includes:

[0103] An encoder for generating features of each frame image according to the face video whose authenticity needs to be detected.

[0104] A spatio-temporal composite graph construction module for constructing a spatio-temporal composite graph according to the features of each frame image.

[0105] A spatio-temporal composite graph convolution module for performing feature learning according to the spatio-temporal composite graph to obtain classification features.

[0106] A classifier for predicting the authenticity result of the face video according to the classification features.

[0107] For the specific function implementation of the above modules, refer to the relevant content in the method of Example 1 and will not be elaborated here.

[0108] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0109] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0110] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0112] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A face forgery detection method based on a graph convolutional neural network, characterized in that Including: Obtain a face video to be detected for authenticity; Input the face video to be detected for authenticity into a trained detection model to obtain a judgment result on whether the video is forged; Wherein, the detection model includes: An encoder for generating features of each frame of image according to the face video to be detected for authenticity; A spatio-temporal composite graph construction module for constructing a spatio-temporal composite graph combining a spatial graph and a temporal graph according to the features of each frame of image; A spatio-temporal composite graph convolution module for performing feature learning according to the spatio-temporal composite graph to obtain classification features; A classifier for predicting the authenticity result of the face video according to the classification features.

2. The face forgery detection method based on a graph convolutional neural network according to claim 1, wherein The main body of the spatio-temporal composite graph is a temporal graph; the temporal graph uses the features of each frame of image as nodes and the temporal front-back relationship as the basis for node connection; each node in the temporal graph corresponds to the spatial graph of the features of each frame of image; the spatial graph uses the features of the face component regions in a frame of image as nodes and the layout characteristics of the face components as the basis for node connection.

3. The face forgery detection method based on a graph convolutional neural network according to claim 1, wherein The spatio-temporal composite graph construction module includes a temporal graph construction unit and a spatial graph construction unit; Constructing the spatio-temporal composite graph according to the features of each frame of image includes: Input the spatio-temporal composite graph construction module, where f represents the feature vector obtained after being processed by the encoder, T represents the number of video frames, c represents the basic number of feature channels, h represents the feature height, and w represents the feature width; ​ The time series graph construction unit obtains a feature vector with the video frame number T as the first dimension through feature recombination, and then obtains the time series graph G t The time series graph G t contains T nodes with a feature size of , respectively corresponding to each frame of features; The spatial graph construction unit obtains, through feature recombination, a feature vector with the total number of all positions in each frame of image as the first-dimension , and further obtains the spatial graph G s . The spatial graph G s contains nodes with a feature size of c , respectively corresponding to the facial component region features in each frame of image.

4. The face forgery detection method based on a graph convolutional neural network according to claim 1, characterized in that, The spatio-temporal composite graph convolution module includes a plurality of spatial graph convolution blocks and a plurality of spatio-temporal convolution blocks, and the spatial graph convolution blocks and the spatio-temporal convolution blocks are alternately connected in series; The spatial graph convolution block is used for mining spatial features of the input spatial graph, and the spatio-temporal convolution block is used for mining temporal features of the input temporal graph.

5. The face forgery detection method based on graph convolutional neural network according to claim 4, characterized in that, Both the spatial graph convolution block and the temporal graph convolution block include 3 graph convolution layers, 3 graph normalization layers and 2 activation layers, and the graph convolution layers, graph normalization layers and activation layers are alternately connected in series; The graph convolution layer is used for extracting features in the graph data by aggregating the feature information of nodes and their neighbors, capturing the local structure and relationship information between nodes in the graph, and providing a deep feature representation. The function formula adopted by the graph convolution layer is as follows: ; Among them, and represent the node feature matrices of the l-th and (l + 1)-th layers, represents the adjacency matrix of the graph, represents the degree matrix of, W l represents the learnable weight of the l-th layer, represents the activation function; The graph normalization layer is used for adjusting the distribution of data to a preset range. The function formula adopted by the graph normalization layer is as follows: ; Among them, represents the feature vector of node i after normalization, h i represents the feature vector of node i, represents the mean of all feature dimensions of node i, represents the variance of all feature dimensions of node i, represents a constant used to avoid a zero denominator; The activation layer is used for enhancing the expression ability of the model, and the activation function adopted by the activation layer includes ReLU.

6. The face forgery detection method based on graph convolutional neural network according to claim 1, wherein, The classifier includes a fully connected layer. The input size of the fully connected layer is the size of the classification features, and the output size of the fully connected layer is 2, respectively representing the probabilities of predicting the face video as real and forged.

7. The face forgery detection method based on a graph convolutional neural network according to claim 1, characterized in that, The training method of the detection model includes: Obtain a training set, the training set includes real face video data, face video data generated by forgery technology, and corresponding labels. Among them, the label values are 0 and 1, 0 represents real data, and 1 represents forged false data; Input the real face video data and the face video data generated by forgery technology into a pre-constructed detection model to obtain a model prediction result; Calculate the cross-entropy loss function according to the model prediction result and the corresponding label; Based on the Adam optimizer, for each training batch, first use forward propagation to calculate the cross-entropy loss function, then use backward propagation to calculate the partial derivatives of each weight parameter, and finally use the gradient descent method to update the weight parameters. Repeat this step until the cross-entropy loss function reaches the minimum to obtain the trained detection model.

8. The face forgery detection method based on graph convolutional neural network according to claim 7, characterized in that The obtaining of the training set includes downloading public datasets and constructing private datasets; The public datasets include FaceForensics++ and DeepFake Detection Challenge; The construction of the private dataset includes: obtaining face videos from video websites or collecting face videos; using the Dlib tool to detect faces; using affine transformation to align faces; setting a fixed margin for the face area and then cropping; scaling the cropped face images to a fixed pixel size.

9. The face forgery detection method based on graph convolutional neural network according to claim 7, characterized in that, The expression of the cross-entropy loss function is as follows: ; Among them, y represents the corresponding label, indicating the model prediction result.

10. A face forgery detection system based on a graph convolutional neural network, characterized in that, including: A data acquisition module for acquiring face videos to be detected for authenticity; A detection module for inputting the face videos to be detected for authenticity into the trained detection model to obtain a judgment result on whether the video is forged; Among them, the detection model includes: An encoder for generating features of each frame of image according to the face video to be detected for authenticity; A spatio-temporal composite graph construction module for constructing a spatio-temporal composite graph according to the features of each frame of image; A spatio-temporal composite graph convolution module for performing feature learning according to the spatio-temporal composite graph to obtain classification features; A classifier for predicting the authenticity result of the face video according to the classification features.

Citation Information

Patent Citations

  • Skeleton behavior recognition method based on 3D space-time diagram convolution

    CN111814719A