Video detection method, device, equipment, medium and program product

By extracting facial key points position and heart rate timing data in video detection, generating spatiotemporal features, and using graph convolution and spatiotemporal graph convolution networks, the problem of poor robustness of DeepFake video detection in the low-quality noise environment in the prior art is solved, and the detection accuracy is improved.

CN120279455APending Publication Date: 2025-07-08IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146023.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing DeepFake video detection technology has poor robustness, high error detection rate and low accuracy when facing low-quality and noisy videos.

Method used

By extracting the position information of facial key points in the video frame, generating spatiotemporal features, and combining heart rate timing data, video detection is performed using graph convolution neural network and spatiotemporal graph convolution network to improve detection accuracy.

Benefits of technology

Improves the accuracy of deep fake video detection, especially in low quality and noisy environments to identify DeepFake videos more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279455A_ABST
    Figure CN120279455A_ABST
Patent Text Reader

Abstract

The invention provides a video detection method and device, equipment, a medium and a program product, and the method comprises the steps: extracting each to-be-detected video frame in a to-be-detected video, and determining the position information of a face key point in each to-be-detected video frame; generating a first spatial-temporal feature based on the connection relationship between the face key points and the position information of the face key points in each to-be-detected video frame; determining heart rate time sequence data corresponding to each to-be-detected video frame; generating a detection result of the to-be-detected video based on the first spatial-temporal feature and the heart rate time sequence data corresponding to each to-be-detected video frame; wherein the detection result is used for predicting whether the to-be-detected video is a deep forged video. According to the invention, the detection accuracy of the deeply-forged video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video processing, and particularly to a video detection method, device, equipment, medium, and program product. Background Art

[0002] Existing DeepFake video detection technologies mainly rely on features of facial images, such as facial movements, skin textures, etc. However, in practical applications, due to reasons such as video quality degradation or adversarial attacks, the facial images may have noise or be missing, resulting in the failure of existing detection methods.

[0003] Existing DeepFake video detection solutions have poor robustness and a greatly increased false detection rate when facing face sequences in low-quality and noisy videos. The accuracy of deepfake video detection is relatively low. Summary of the Invention

[0004] This application provides a video detection method, device, equipment, medium, and program product to improve the accuracy of deepfake video detection.

[0005] According to the first aspect of the embodiments of this application, a video detection method is provided, including:

[0006] Extract each video frame to be detected in the video to be detected, and determine the position information of facial key points in each video frame to be detected;

[0007] Generate a first spatio-temporal feature based on the connection relationship between the facial key points and the position information of the facial key points in each video frame to be detected;

[0008] Determine the heart rate time series data corresponding to each video frame to be detected;

[0009] Generate a detection result of the video to be detected based on the first spatio-temporal feature and the heart rate time series data corresponding to each video frame to be detected; wherein, the detection result is used to predict whether the video to be detected is a deepfake video.

[0010] Optionally, the generating a first spatio-temporal feature based on the connection relationship between the facial key points and the position information of the facial key points in each video frame to be detected includes:

[0011] Generate the structural information of the graph based on the connection relationship between the facial key points;

[0012] Determine the first feature information of each node in the graph based on the position information of the facial key points in each video frame to be detected; wherein, one node in the graph corresponds to one facial key point;

[0013] Input the structural information of the graph and the first feature information of each node in the graph into a graph convolutional neural network to obtain first spatio-temporal features; wherein, the first spatio-temporal features are used to characterize the spatio-temporal relationship between each of the facial key points in the video to be detected.

[0014] Optionally, the heart rate time series data corresponding to each video frame to be detected includes the heart rate time series data corresponding to each of the facial key points in each video frame to be detected;

[0015] Generating the detection result of the video to be detected based on the first spatio-temporal features and the heart rate time series data corresponding to each video frame to be detected includes:

[0016] Based on the connection relationship between each of the facial key points and the connection relationship of the same facial key point between each video frame to be detected, construct the structural information of the spatio-temporal graph;

[0017] Based on the first spatio-temporal features and the heart rate time series data corresponding to each of the facial key points in each video frame to be detected, determine the second feature information of each node in the spatio-temporal graph; wherein, one node in the spatio-temporal graph corresponds to one facial key point in one video frame to be detected;

[0018] Input the structural information of the spatio-temporal graph and the second feature information of each node in the spatio-temporal graph into a spatio-temporal graph convolutional network to obtain second spatio-temporal features;

[0019] Generate the detection result of the video to be detected based on the second spatio-temporal features.

[0020] Optionally, generating the detection result of the video to be detected based on the second spatio-temporal features includes:

[0021] Perform statistical analysis on the heart rate time series data corresponding to each video frame to be detected to obtain the physiological index corresponding to the video to be detected;

[0022] Generate the detection result of the video to be detected based on the second spatio-temporal features and the physiological index corresponding to the video to be detected.

[0023] Optionally, the training process of the spatio-temporal graph convolutional network includes:

[0024] Extract each sample video frame in the sample video and determine the position information of the facial key points in each sample video frame; wherein, the sample video includes videos corresponding to various forgery categories;

[0025] Generate first sample spatio-temporal features based on the connection relationship between the facial key points and the position information of the facial key points in each sample video frame;

[0026] Determine the heart rate time series data corresponding to each of the facial key points in each sample video frame;

[0027] Based on the connection relationships between the facial key points and the connection relationships of the same facial key points between the sample video frames, construct the structural information of the sample spatio-temporal graph;

[0028] Based on the first sample spatio-temporal feature and the heart rate time series data corresponding to each of the facial key points in each sample video frame, determine the third feature information of each node in the sample spatio-temporal graph; wherein, a node in the sample spatio-temporal graph corresponds to a facial key point in a sample video frame;

[0029] Input the structural information of the sample spatio-temporal graph and the third feature information of each node in the sample spatio-temporal graph into a spatio-temporal graph convolutional network to obtain a second sample spatio-temporal feature;

[0030] Based on the second sample spatio-temporal feature, generate a sample detection result for the sample video; wherein, the sample detection result is used to predict whether the sample video is a deepfake video;

[0031] Based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, and the true detection result of the sample video, train the spatio-temporal graph convolutional network.

[0032] Optionally, the method further includes:

[0033] Generate an adversarial sample corresponding to the original video through a generative adversarial network;

[0034] Based on the original video and the adversarial sample corresponding to the original video, determine the sample video.

[0035] Optionally, the method further includes:

[0036] Identify the sample forgery category corresponding to the sample video;

[0037] The training of the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, and the true detection result of the sample video includes:

[0038] Based on the sample forgery category, determine a first facial key point and a second facial key point in the sample spatio-temporal graph; wherein, the first facial key point is located in an area with obvious forgery traces in the sample video, and the second facial key point is located in an area with unobvious forgery traces in the sample video;

[0039] Determine a first weight corresponding to the third feature information of the first facial key point in the sample spatio-temporal graph, and a second weight corresponding to the third feature information of the second facial key point in the sample spatio-temporal graph; wherein, the first weight is greater than the second weight;

[0040] Train the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the first weight corresponding to the third feature information of the first facial key point, the second weight corresponding to the third feature information of the second facial key point, the sample detection result, and the true detection result of the sample video.

[0041] Optionally, the training of the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, and the true detection result of the sample video includes:

[0042] Determine the forgery intensity corresponding to each of the facial key points in the sample spatio-temporal graph;

[0043] Based on the forgery intensity, adjust the loss function of the spatio-temporal graph convolutional network to obtain an adjusted loss function;

[0044] Train the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, the true detection result of the sample video, and the adjusted loss function.

[0045] According to the second aspect of the embodiments of the present application, a video detection device is provided, including:

[0046] An extraction unit, configured to extract each frame of the video to be detected in the video to be detected, and determine the position information of the facial key points in each frame of the video to be detected;

[0047] A generation unit, configured to generate a first spatio-temporal feature based on the connection relationship between the facial key points and the position information of the facial key points in each frame of the video to be detected;

[0048] A first processing unit, configured to determine the heart rate time series data corresponding to each frame of the video to be detected;

[0049] A second processing unit, configured to generate a detection result of the video to be detected based on the first spatio-temporal feature and the heart rate time series data corresponding to each frame of the video to be detected; wherein, the detection result is used to predict whether the video to be detected is a deep fake video.

[0050] According to a third aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor;

[0051] The memory is connected to the processor and is used for storing programs;

[0052] The processor is configured to implement the video detection method as described in the first aspect by running the programs in the memory.

[0053] According to a fourth aspect of the embodiments of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is run by a processor, the video detection method as described in the first aspect is implemented.

[0054] According to a fifth aspect of the embodiments of the present application, a computer program product is provided, including computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the video detection method as described in the first aspect.

[0055] In the present application, each video frame to be detected in the video to be detected is extracted, and the position information of the facial key points in each video frame to be detected is determined. Based on the connection relationship between the facial key points and the position information of the facial key points in each video frame to be detected, a first spatio-temporal feature is generated. In the process of generating the first spatio-temporal feature, both the connection relationship between the facial key points and the position information of the facial key points are considered, that is, the spatial relationship between the facial key points, and each video frame to be detected is also considered, that is, the temporal relationship between the facial key points. Therefore, the first spatio-temporal feature can represent the spatio-temporal relationship between the facial key points in the video to be detected. The heart rate time series data corresponding to each video frame to be detected is determined, and based on the first spatio-temporal feature and the heart rate time series data corresponding to each video frame to be detected, a detection result of the video to be detected is generated, where the detection result is used to predict whether the video to be detected is a deepfake video. Since the first spatio-temporal feature can represent the spatio-temporal relationship between the facial key points in the video to be detected, generating the detection result of the video to be detected based on the first spatio-temporal feature can improve the accuracy of deepfake video detection. Moreover, using the heart rate time series data corresponding to each video frame to be detected as additional information for assisting in deepfake video detection can further improve the accuracy of deepfake video detection. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0057] Figure 1 It is a schematic flowchart of a video detection method provided in an embodiment of the present application;

[0058] Figure 2 It is a schematic flowchart of step 102 provided in an embodiment of the present application;

[0059] Figure 3 It is a schematic flowchart of step 104 provided in an embodiment of the present application;

[0060] Figure 4 It is a schematic flowchart of step 304 provided in an embodiment of the present application;

[0061] Figure 5 It is a schematic flowchart of the training process of a spatio-temporal graph convolutional network provided in an embodiment of the present application;

[0062] Figure 6 It is a schematic flowchart of step 508 provided in an embodiment of the present application;

[0063] Figure 7 It is a schematic flowchart of step 508 provided in an embodiment of the present application;

[0064] Figure 8 It is a schematic flowchart of a video detection method provided in an embodiment of the present application;

[0065] Figure 9 It is a schematic structural diagram of a video detection device provided in an embodiment of the present application;

[0066] Figure 10 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Detailed implementation manners

[0067] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0068] Exemplary implementation environment

[0069] The video detection method according to the embodiments of the present application can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user device, a mobile device, a computing device, a wearable device, etc., and the server can be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server capable of performing cloud computing. This method can be implemented by a processor calling computer-readable program instructions stored in a memory.

[0070] Exemplary method

[0071] Please refer to Figure 1 , in an exemplary embodiment, a video detection method is provided. As Figure 1 shown, the process of the video detection method mainly includes:

[0072] Step 101, extract each video frame to be detected in the video to be detected, and determine the position information of the facial key points in each video frame to be detected.

[0073] In an exemplary embodiment, extracting each video frame to be detected in the video to be detected may be preprocessing and frame splitting the video to be detected to obtain each video frame to be detected in the video to be detected. Ensure that each frame in the video to be detected can be accurately extracted.

[0074] In an exemplary embodiment, the facial key points may refer to distinguishable facial feature points, for example, eyes, mouth, nose, etc.

[0075] In some embodiments, determining the position information of the facial key points in each video frame to be detected may include: determining the face detection region in each video frame to be detected; determining the position information of the facial key points from the face detection region in each video frame to be detected.

[0076] In an exemplary embodiment, determining the face detection region in each video frame to be detected may include: inputting each video frame to be detected into a CNN (Convolutional Neural Networks), and obtaining the face detection region in each video frame to be detected output by the CNN.

[0077] For example, in a video, the face region may not occupy the entire video frame size, and it is necessary to detect the face region first.

[0078] For example, the CNN can be a multi-scale CNN.

[0079] Extracting the face detection region in each video frame to be detected through the CNN can extract a finer-grained facial region and improve the detection accuracy.

[0080] In an exemplary embodiment, determining the position information of facial key points from the face detection regions in each video frame to be detected may include: inputting the face detection regions in each video frame to be detected into a DNN (Deep Neural Network), and obtaining the position information of the facial key points in the face detection regions of each video frame output by the DNN.

[0081] Step 102, generating a first spatio-temporal feature based on the connection relationship between facial key points and the position information of facial key points in each video frame to be detected.

[0082] In some embodiments, as Figure 2 shown, step 102 includes:

[0083] Step 201, generating the structural information of a graph based on the connection relationship between facial key points.

[0084] In an exemplary embodiment, the structural information of the graph may include the adjacency matrix of the graph. The adjacency matrix is used to describe the connection relationship between each node in the graph. A node in the graph corresponds to a facial key point. Therefore, based on the connection relationship between facial key points, the structural information of the graph can be generated.

[0085] In an exemplary embodiment, the connection relationship between facial key points is predefined. For example, the eyes are connected to the nose, and the nose is connected to the mouth.

[0086] Step 202, determining the first feature information of each node in the graph based on the position information of facial key points in each video frame to be detected.

[0087] Wherein, a node in the graph corresponds to a facial key point.

[0088] In an exemplary embodiment, the position information corresponding to each facial key point in each video frame to be detected, that is, the position information sequence corresponding to any facial key point, can be determined as the first feature information of any facial key point in the graph.

[0089] Step 203, inputting the structural information of the graph and the first feature information of each node in the graph into a graph convolutional neural network to obtain a first spatio-temporal feature.

[0090] Wherein, the first spatio-temporal feature is used to characterize the spatio-temporal relationship between each facial key point in the video to be detected.

[0091] In an exemplary embodiment, the graph convolutional neural network refers to a GCN (Graph Convolutional Networks).

[0092] In an exemplary embodiment, the first spatio-temporal feature may refer to a feature map.

[0093] There are clear spatial relationships between facial key points. For example, the movements of the eyes and mouth are interrelated, and the movement of the face is a continuous process. Therefore, when performing DeepFake video detection, not only the spatial features (such as position, size, shape) of these facial key points are concerned, but also the dynamic changes (such as movement patterns, relative movements, etc.) of these facial key points on the time axis need to be considered. Such spatio-temporal features need to be modeled in a suitable way.

[0094] In this application, based on the connection relationships between facial key points, the structural information of the graph is generated. Based on the position information of facial key points in each video frame to be detected, the first feature information of each node in the graph is determined. Here, one node in the graph corresponds to one facial key point, and the first feature information of any facial key point in the graph includes the position information corresponding to each of the facial key points in each video frame to be detected. Therefore, the first feature information of each node in the graph can reflect both the spatial relationship and the temporal relationship between the facial key points in the video to be detected. The structural information of the graph and the first feature information of each node in the graph are input into a graph convolutional neural network to obtain the first spatio-temporal feature. In the process of generating the first spatio-temporal feature, both the connection relationships between facial key points and the spatio-temporal relationships between facial key points are considered. Therefore, the first spatio-temporal feature can represent the spatio-temporal relationship between the facial key points in the video to be detected. Moreover, the graph convolutional neural network has significant advantages in processing graph-structured data, can learn the deep connections and features between nodes, and can obtain more accurate first spatio-temporal features. Generating the detection result of the video to be detected based on the first spatio-temporal feature can improve the accuracy of deepfake video detection.

[0095] In some other embodiments, step 102 may include: constructing the structural information of the spatio-temporal graph based on the connection relationships between facial key points and the connection relationships of the same facial key points between each video frame to be detected; determining the fourth feature information of each node in the spatio-temporal graph based on the position information of facial key points in each video frame to be detected; where one node in the spatio-temporal graph corresponds to one facial key point in one video frame to be detected; inputting the structural information of the spatio-temporal graph and the fourth feature information of each node in the spatio-temporal graph into a spatio-temporal graph convolutional network to obtain the first spatio-temporal feature.

[0096] Step 103, determining the heart rate time series data corresponding to each video frame to be detected.

[0097] In some embodiments, step 103 may include: obtaining the physiological monitoring signals corresponding to each video frame to be detected; determining the heart rate time series data corresponding to each video frame to be detected based on the physiological monitoring signals corresponding to each video frame to be detected.

[0098] In an exemplary embodiment, the physiological monitoring signal may refer to an rPPG (Remote Photoplethysmography) signal. Among them, rPPG is a non-contact physiological signal monitoring technology. The rPPG signal refers to the information of the periodic change of skin color caused by the cardiac cycle.

[0099] In an exemplary embodiment, obtaining the rPPG signal corresponding to each video frame to be detected may include: obtaining the rPPG signal from the face detection region in each video frame to be detected.

[0100] In an exemplary embodiment, step 103 may include: extracting the heart rate time series data from the face detection region in each video frame to be detected by using a deep learning-based rPPG algorithm. Automatically optimize blood flow detection through a convolutional network, improve the accuracy and robustness of the rPPG signal, and be able to better separate effective physiological features from the video.

[0101] In an exemplary embodiment, adaptive noise filtering and signal enhancement may be performed on the face detection region in each video frame to be detected to obtain the optimized face detection region in each video frame to be detected; obtain the rPPG signal from the optimized face detection region in each video frame to be detected. It can reduce the interference of factors such as ambient light and facial movement during the process of extracting the rPPG signal.

[0102] In an exemplary embodiment, the time domain signal in the face detection region in each video frame to be detected may be converted into a frequency domain signal, and then adaptive noise filtering and signal enhancement are performed on the frequency domain signal. Further improve the stability and accuracy of the signal.

[0103] In an exemplary embodiment, methods such as Fourier transform or wavelet transform may be used to convert the time domain signal into a frequency domain signal.

[0104] Step 104, generate a detection result of the video to be detected based on the first spatio-temporal feature and the heart rate time series data corresponding to each video frame to be detected.

[0105] Among them, the detection result is used to predict whether the video to be detected is a deepfake video.

[0106] In an exemplary embodiment, a deepfake video refers to a DeepFake video.

[0107] In some embodiments, the heart rate time series data corresponding to each video frame to be detected includes the heart rate time series data corresponding to each facial key point in each video frame to be detected.

[0108] In an exemplary embodiment, based on the position information of the facial key points in each video frame to be detected, the rPPG signals corresponding to the facial key points in each video frame to be detected can be extracted; based on the rPPG signals corresponding to the facial key points in each video frame to be detected, the heart rate time series data corresponding to each facial key point in each video frame to be detected can be determined.

[0109] In some embodiments, as Figure 3 shown, step 104 includes:

[0110] Step 301, based on the connection relationships between the facial key points and the connection relationships of the same facial key points between the video frames to be detected, construct the structural information of the spatio-temporal graph.

[0111] Step 302, based on the first spatio-temporal feature and the heart rate time series data corresponding to each facial key point in each video frame to be detected, determine the second feature information of each node in the spatio-temporal graph.

[0112] Wherein, a node in the spatio-temporal graph corresponds to a facial key point in a video frame to be detected.

[0113] In an exemplary embodiment, the first spatio-temporal feature refers to a feature map, and the feature map includes the features corresponding to each facial key point updated by the graph convolutional network.

[0114] In an exemplary embodiment, step 302 may include: determining the weight corresponding to the first spatio-temporal feature; determining the weight corresponding to the heart rate time series data; based on the first spatio-temporal feature, the weight corresponding to the first spatio-temporal feature, the heart rate time series data, and the weight corresponding to the heart rate time series data, perform weighted fusion to obtain the second feature information of each node in the spatio-temporal graph.

[0115] In an exemplary embodiment, the weight corresponding to the first spatio-temporal feature and the weight corresponding to the heart rate time series data can be dynamically adjusted. For example, an adaptive weighted fusion method (a learnable weighting mechanism using an attention mechanism) can be used to perform weighted fusion based on the first spatio-temporal feature, the weight corresponding to the first spatio-temporal feature, the heart rate time series data, and the weight corresponding to the heart rate time series data to obtain the second feature information of each node in the spatio-temporal graph.

[0116] Step 303, input the structural information of the spatio-temporal graph and the second feature information of each node in the spatio-temporal graph into the spatio-temporal graph convolutional network to obtain the second spatio-temporal feature.

[0117] In an exemplary embodiment, the spatio-temporal graph convolutional network refers to ST-GCN (Spatial-Temporal Graph Convolutional Network), which combines spatio-temporal convolution and graph convolution.

[0118] In some embodiments, to further improve robustness, graph regularization is performed on the spatio-temporal graph convolutional network. Not only is spatio-temporal feature entanglement performed during feature extraction, but a regularization term related to the spatio-temporal graph structure is also added to the loss function of the spatio-temporal graph convolutional network to optimize the structural information of the spatio-temporal graph. This regularization helps remove noise nodes in the spatio-temporal graph and strengthen the influence of key features such as facial expressions and movements in the spatio-temporal graph, thereby avoiding common forgery techniques in DeepFake videos.

[0119] In an exemplary embodiment, graph regularization refers to the graph Laplacian smoothing prior method. In addition to the conventional role of preventing overfitting, it can also constrain the smoothness between nodes, thereby avoiding misidentifications caused by noise nodes.

[0120] Step 304: Generate a detection result for the video to be detected based on the second spatio-temporal feature.

[0121] Based on the first spatio-temporal feature and the heart rate time series data corresponding to each facial key point in each video frame to be detected, the second feature information of each node in the spatio-temporal graph is determined. The first spatio-temporal feature can represent the spatio-temporal relationship between each facial key point in the video to be detected. In the process of determining the second feature information of each node in the spatio-temporal graph, both the spatio-temporal relationship between each facial key point in the video to be detected and the heart rate time series data corresponding to each video frame to be detected are considered, and the two are fused, which can further improve the accuracy of deep fake video detection. Moreover, the spatio-temporal graph convolutional network combines the advantages of the graph convolutional network and the spatio-temporal convolutional network, can efficiently process data with spatio-temporal characteristics, can extract features from both the time and space dimensions, and consider the causal relationship between time series. When processing complex spatio-temporal data, it can capture more information, improve the prediction accuracy and stability of the model, can obtain more accurate second spatio-temporal features, and generate a detection result for the video to be detected based on the second spatio-temporal feature, which can improve the accuracy of deep fake video detection.

[0122] In some embodiments, as Figure 4 shown, step 304 includes:

[0123] Step 401: Perform statistical analysis on the heart rate time series data corresponding to each video frame to be detected to obtain a physiological index corresponding to the video to be detected.

[0124] In an exemplary embodiment, the physiological indicators corresponding to the video to be detected may include physiological indicators related to the cardiac cycle, such as heart rate changes, rhythm irregularities, etc.

[0125] Step 402: Generate a detection result for the video to be detected based on the second spatio-temporal feature and the physiological indicators corresponding to the video to be detected.

[0126] In an exemplary embodiment, step 402 may include: inputting the second spatio-temporal feature and the physiological indicators corresponding to the video to be detected into a classifier to obtain a detection result for the video to be detected.

[0127] In an exemplary embodiment, for the spatio-temporal graph convolutional network, the heart rate time series data is a streaming feature input; while the physiological indicators corresponding to the video to be detected are a statistical indicator combined with prior knowledge, which is a result obtained by statistically analyzing the time scale of the entire video. Therefore, the heart rate time series data and the physiological indicators corresponding to the video to be detected are data of different dimensions. Based on the second spatio-temporal feature generated from the heart rate time series data corresponding to each video frame to be detected, generating a detection result for the video to be detected based on the second spatio-temporal feature and the physiological indicators corresponding to the video to be detected can use the heart rate time series data corresponding to each video frame to be detected and the physiological indicators corresponding to the video to be detected as additional information to assist in the detection of deepfake videos, and can further improve the accuracy of deepfake video detection.

[0128] In some other embodiments, step 304 includes: inputting the second spatio-temporal feature into a classifier to obtain a detection result for the video to be detected.

[0129] In some other embodiments, step 104 includes: performing statistical analysis on the heart rate time series data corresponding to each video frame to be detected to obtain the physiological indicators corresponding to the video to be detected; generating a detection result for the video to be detected based on the first spatio-temporal feature and the physiological indicators corresponding to the video to be detected.

[0130] In some embodiments, as Figure 5 shown, the training process of the spatio-temporal graph convolutional network includes:

[0131] Step 501: Extract each sample video frame in the sample video and determine the position information of the facial key points in each sample video frame.

[0132] Among them, the sample video includes videos corresponding to various forgery categories.

[0133] In an exemplary embodiment, the forgery categories may include abnormal facial expressions, unnatural eye movements, or mismatches between the face and voice, etc., and may also include other forgery categories, which are not limited in this application.

[0134] Step 502: Generate the first sample spatio-temporal feature based on the connection relationships between the facial key points and the position information of the facial key points in each sample video frame.

[0135] Step 503: Determine the respective heart rate time series data corresponding to the facial key points in each sample video frame.

[0136] Step 504: Construct the structural information of the sample spatio-temporal graph based on the connection relationships between the facial key points and the connection relationships of the same facial key points between each sample video frame.

[0137] Step 505: Determine the third feature information of each node in the sample spatio-temporal graph based on the first sample spatio-temporal feature and the respective heart rate time series data corresponding to the facial key points in each sample video frame.

[0138] Among them, a node in the sample spatio-temporal graph corresponds to a facial key point in a sample video frame.

[0139] Step 506: Input the structural information of the sample spatio-temporal graph and the third feature information of each node in the sample spatio-temporal graph into the spatio-temporal graph convolutional network to obtain the second sample spatio-temporal feature.

[0140] Step 507: Generate the sample detection result of the sample video based on the second sample spatio-temporal feature.

[0141] Among them, the sample detection result is used to predict whether the sample video is a deepfake video.

[0142] Step 508: Train the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, and the true detection result of the sample video.

[0143] In an exemplary embodiment, the true detection result of the sample video may refer to the true label obtained by pre-annotating the sample video, which can accurately reflect whether the sample video is a deepfake video.

[0144] The sample video includes videos corresponding to various forgery categories, and can train the spatio-temporal graph convolutional network for videos corresponding to various forgery categories, so that the trained spatio-temporal graph convolutional network can effectively identify various deepfake videos.

[0145] In some embodiments, the video detection method further includes: identifying the sample forgery category corresponding to the sample video.

[0146] In an exemplary embodiment, a trained network model (such as a GAN (Generative Adversarial Networks) or a discriminator trained adversarially) can be used to identify the sample forgery category corresponding to the sample video, so as to subsequently dynamically adjust the training strategy of the spatio-temporal graph convolutional network according to the identified sample forgery category, and pay more attention to the third feature information of the facial key points corresponding to the sample forgery category.

[0147] In some embodiments, as Figure 6 shown, step 508 includes:

[0148] Step 601, based on the sample forgery category, determine the first facial key point and the second facial key point in the sample spatio-temporal graph.

[0149] Among them, the first facial key point is located in the area with obvious forgery traces in the sample video, and the second facial key point is located in the area with unobvious forgery traces in the sample video.

[0150] For example, if the sample forgery category is unnatural eye movement, the area with obvious forgery traces in the sample video can refer to the eyes, and the area with unobvious forgery traces in the sample video can refer to the area other than the eyes in the face detection area.

[0151] Step 602, determine the first weight corresponding to the third feature information of the first facial key point in the sample spatio-temporal graph, and the second weight corresponding to the third feature information of the second facial key point in the sample spatio-temporal graph.

[0152] Among them, the first weight is greater than the second weight.

[0153] In an exemplary embodiment, an adaptive weighting mechanism can be used to dynamically adjust the first weight and the second weight, and adjust the importance of the third feature information of the first facial key point and the third feature information of the second facial key point in the spatio-temporal graph convolutional network.

[0154] Step 603, based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the first weight corresponding to the third feature information of the first facial key point, the second weight corresponding to the third feature information of the second facial key point, the sample detection result, and the true detection result of the sample video, train the spatio-temporal graph convolutional network.

[0155] The first weight corresponding to the third feature information of the first facial key point is greater than the second weight corresponding to the third feature information of the second facial key point. The first facial key point is located in the area with obvious forgery traces in the sample video, and the second facial key point is located in the area with unobvious forgery traces in the sample video. By assigning a higher first weight to the first facial key point in the area with obvious forgery traces in the sample video and a lower second weight to the second facial key point in the area with unobvious forgery traces in the sample video, the learning of the third feature information of the first facial key point in the area with obvious forgery traces in the sample video is made more profound, the influence of the third feature information of the first facial key point in the area with obvious forgery traces in the sample video can be amplified, and at the same time, the interference of the third feature information of the second facial key point in the irrelevant area (i.e., the area with unobvious forgery traces in the sample video) on predicting whether the sample video is a deepfake video is reduced. The accuracy of detecting deepfake videos by the trained spatio-temporal graph convolutional network can be improved. Identifying and weighting forgery features adaptively during the training process, especially for those subtle and imperceptible forgery traces (such as tiny facial expression changes or details after DeepFake), can further improve the robustness and generalization ability of the spatio-temporal graph convolutional network.

[0156] In some embodiments, as Figure 7 shown, step 508 includes:

[0157] Step 701, determining the forgery intensity corresponding to each facial key point in the sample spatio-temporal graph.

[0158] In an exemplary embodiment, the forgery intensity can be dynamically adjusted and learned by the spatio-temporal graph convolutional network, and it is a learnable parameter in the spatio-temporal graph convolutional network and will be fixed after the spatio-temporal graph convolutional network is trained.

[0159] Step 702, adjusting the loss function of the spatio-temporal graph convolutional network based on the forgery intensity to obtain an adjusted loss function.

[0160] Step 703, training the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, the true detection result of the sample video, and the adjusted loss function.

[0161] Adjusting the loss function of the spatio-temporal graph convolutional network based on the forgery intensity can make the spatio-temporal graph convolutional network generate a stronger optimization direction when the forgery features are significant, and for areas with weak or imperceptible forgery features, reduce the dependence on the features of this part and avoid overfitting.

[0162] In some embodiments, as Figure 8 shown, the video detection method further includes:

[0163] Step 801: Generate adversarial samples corresponding to the original video through a generative adversarial network.

[0164] In an exemplary embodiment, the generative adversarial network refers to GAN (Generative Adversarial Networks).

[0165] In an exemplary embodiment, there is a high similarity between the adversarial samples corresponding to the original video and the original video, but there are differences in certain features or details. These differences may be very subtle, even difficult to detect with the naked eye, but are sufficient to make it difficult for the discriminator to distinguish between authenticity and falsity. These differences may include pixel-level changes, slight adjustments in color or brightness, and subtle changes in texture, etc.

[0166] Step 802: Determine the sample video based on the original video and the adversarial samples corresponding to the original video.

[0167] Generate adversarial samples corresponding to the original video through a generative adversarial network to simulate different types of forged video samples, so as to enhance the robustness of the spatio-temporal graph convolutional network. These forged video samples can cover various forgery means to ensure that the spatio-temporal graph convolutional network can still make accurate judgments in the face of various forgery techniques. In this way, the spatio-temporal graph convolutional network can not only detect forgery traces in real videos but also effectively identify them in various generated forged videos.

[0168] In summary, in this application, each frame of the video to be detected is extracted, and the position information of the facial key points in each frame to be detected is determined. Based on the connection relationship between the facial key points and the position information of the facial key points in each frame to be detected, the first spatio-temporal feature is generated. In the process of generating the first spatio-temporal feature, both the connection relationship between the facial key points and the position information of the facial key points, that is, the spatial relationship between the facial key points, are considered, and each frame to be detected, that is, the temporal relationship between the facial key points, is also considered. Therefore, the first spatio-temporal feature can represent the spatio-temporal relationship between the facial key points in the video to be detected. Determine the heart rate time series data corresponding to each frame to be detected, and generate the detection result of the video to be detected based on the first spatio-temporal feature and the heart rate time series data corresponding to each frame to be detected. Among them, the detection result is used to predict whether the video to be detected is a deepfake video. Since the first spatio-temporal feature can represent the spatio-temporal relationship between the facial key points in the video to be detected, generating the detection result of the video to be detected based on the first spatio-temporal feature can improve the accuracy of deepfake video detection. Moreover, using the heart rate time series data corresponding to each frame to be detected as additional information to assist in the detection of deepfake videos can further improve the accuracy of deepfake video detection.

[0169] Exemplary device

[0170] Correspondingly, an embodiment of the present application further provides a video detection device, as Figure 9 shown, the video detection device includes:

[0171] An extraction unit 901, configured to extract each video frame to be detected in the video to be detected, and determine the position information of facial key points in each video frame to be detected;

[0172] A generation unit 902, configured to generate first spatio-temporal features based on the connection relationship between the facial key points and the position information of the facial key points in each video frame to be detected;

[0173] A first processing unit 903, configured to determine the heart rate time series data corresponding to each video frame to be detected;

[0174] A second processing unit 904, configured to generate a detection result of the video to be detected based on the first spatio-temporal features and the heart rate time series data corresponding to each video frame to be detected; wherein, the detection result is used to predict whether the video to be detected is a deepfake video.

[0175] Optionally, the generation unit 902 is specifically configured to:

[0176] Generate the structural information of the graph based on the connection relationship between the facial key points;

[0177] Determine the first feature information of each node in the graph based on the position information of the facial key points in each video frame to be detected; wherein, one node in the graph corresponds to one facial key point;

[0178] Input the structural information of the graph and the first feature information of each node in the graph into a graph convolutional neural network to obtain first spatio-temporal features; wherein, the first spatio-temporal features are used to characterize the spatio-temporal relationship between the facial key points in the video to be detected.

[0179] Optionally, the heart rate time series data corresponding to each video frame to be detected includes the heart rate time series data corresponding to each facial key point in each video frame to be detected;

[0180] The second processing unit 904 includes:

[0181] A construction subunit, configured to construct the structural information of the spatio-temporal graph based on the connection relationship between the facial key points and the connection relationship of the same facial key point between each video frame to be detected;

[0182] A first processing subunit, configured to determine second feature information of each node in the spatio-temporal graph based on the first spatio-temporal feature and the heart rate time series data corresponding to each facial key point in each video frame to be detected; wherein, one node in the spatio-temporal graph corresponds to one facial key point in one video frame to be detected;

[0183] A second processing subunit, configured to input the structure information of the spatio-temporal graph and the second feature information of each node in the spatio-temporal graph into a spatio-temporal graph convolutional network to obtain a second spatio-temporal feature;

[0184] A generating subunit, configured to generate a detection result of the video to be detected based on the second spatio-temporal feature.

[0185] Optionally, the generating subunit is specifically configured to:

[0186] Perform statistical analysis on the heart rate time series data corresponding to each video frame to be detected to obtain a physiological index corresponding to the video to be detected;

[0187] Generate a detection result of the video to be detected based on the second spatio-temporal feature and the physiological index corresponding to the video to be detected.

[0188] Optionally, the video detection device further includes:

[0189] A third processing unit, configured to extract each sample video frame in the sample video and determine the position information of the facial key points in each sample video frame; wherein, the sample video includes videos corresponding to various forgery categories;

[0190] A fourth processing unit, configured to generate a first sample spatio-temporal feature based on the connection relationship between the facial key points and the position information of the facial key points in each sample video frame;

[0191] A fifth processing unit, configured to determine the heart rate time series data corresponding to each facial key point in each sample video frame;

[0192] A constructing unit, configured to construct the structure information of the sample spatio-temporal graph based on the connection relationship between each facial key point and the connection relationship of the same facial key point between each sample video frame;

[0193] A sixth processing unit, configured to determine third feature information of each node in the sample spatio-temporal graph based on the first sample spatio-temporal feature and the heart rate time series data corresponding to each facial key point in each sample video frame; wherein, one node in the sample spatio-temporal graph corresponds to one facial key point in one sample video frame;

[0194] A seventh processing unit, configured to input the structural information of the sample spatio-temporal graph and the third feature information of each node in the sample spatio-temporal graph into a spatio-temporal graph convolutional network to obtain second sample spatio-temporal features;

[0195] An eighth processing unit, configured to generate a sample detection result of the sample video based on the second sample spatio-temporal features; wherein the sample detection result is used to predict whether the sample video is a deepfake video;

[0196] A training unit, configured to train the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, and the true detection result of the sample video.

[0197] Optionally, the video detection device further includes:

[0198] An adversarial sample generation unit, configured to generate an adversarial sample corresponding to the original video through a generative adversarial network;

[0199] A ninth processing unit, configured to determine the sample video based on the original video and the adversarial sample corresponding to the original video.

[0200] Optionally, the video detection device further includes:

[0201] An identification unit, configured to identify the sample forgery category corresponding to the sample video;

[0202] The training unit is specifically configured to:

[0203] Determine a first facial key point and a second facial key point in the sample spatio-temporal graph based on the sample forgery category; wherein the first facial key point is located in an area with obvious forgery traces in the sample video, and the second facial key point is located in an area with inconspicuous forgery traces in the sample video;

[0204] Determine a first weight corresponding to the third feature information of the first facial key point in the sample spatio-temporal graph, and a second weight corresponding to the third feature information of the second facial key point in the sample spatio-temporal graph; wherein the first weight is greater than the second weight;

[0205] Train the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the first weight corresponding to the third feature information of the first facial key point, the second weight corresponding to the third feature information of the second facial key point, the sample detection result, and the true detection result of the sample video.

[0206] Optionally, the training unit is specifically configured to:

[0207] Determine the forgery intensity corresponding to each of the facial key points in the sample spatio-temporal map;

[0208] Based on the forgery intensity, adjust the loss function of the spatio-temporal map convolutional network to obtain an adjusted loss function;

[0209] Train the spatio-temporal map convolutional network based on the structural information of the sample spatio-temporal map, the third feature information of each node in the sample spatio-temporal map, the sample detection result, the true detection result of the sample video, and the adjusted loss function.

[0210] The video detection device provided in this embodiment belongs to the same inventive concept as the video detection method provided in the above embodiments of the present application, and can execute the video detection method provided in any of the above embodiments of the present application, and has corresponding functional modules and beneficial effects for executing the video detection method. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the video detection method provided in the above embodiments of the present application, which will not be elaborated here.

[0211] The functions implemented by the above extraction unit 901, generation unit 902, first processing unit 903, and second processing unit 904 can be implemented by the same or different processors respectively, and the embodiments of the present application do not make any limitations.

[0212] It should be understood that the units in the above device can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory inside the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of a hardware circuit. By designing the hardware circuit, some or all of the functions of the units can be implemented. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and by designing the logical relationship between the components in the circuit, some or all of the functions of the above units are implemented; again, in another implementation, the hardware circuit can be implemented by a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured through a configuration file to implement some or all of the functions of the above units. All units of the above device can be all implemented in the form of a processor calling software, or all implemented in the form of a hardware circuit, or some implemented in the form of a processor calling software, and the remaining part implemented in the form of a hardware circuit.

[0213] In an embodiment of the present application, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and running capabilities, such as a CPU, microprocessor, GPU, or DSP, etc.; in another implementation, the processor can implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of this hardware circuit is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or PLD, such as an FPGA, etc. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a type of ASIC, such as an NPU, TPU, DPU, etc.

[0214] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method. For example: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0215] In addition, each unit in the above device can be integrated in whole or in part, or can be independently implemented. In one implementation, these units are integrated together and implemented in the form of an SOC. The SOC can include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The types of the at least one processor can be different. For example, it includes a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0216] Exemplary electronic device

[0217] An embodiment of the present application proposes an electronic device. Refer to Figure 10 As shown, the device includes:

[0218] A memory 200 and a processor 210;

[0219] Wherein, the memory 200 is connected to the processor 210 and is used for storing programs;

[0220] The processor 210 is used for implementing the video detection method disclosed in any of the above embodiments by running the programs stored in the memory 200.

[0221] Specifically, the above electronic device may further include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0222] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are interconnected through the bus. Among them:

[0223] The bus may include a path for transmitting information between various components of the computer system.

[0224] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0225] The processor 210 may include a main processor, and may also include a baseband chip, a modem, etc.

[0226] The memory 200 stores the program for implementing the technical solution of the present invention, and may also store an operating system and other critical services. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, etc.

[0227] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.

[0228] The output device 240 may include a device for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.

[0229] The communication interface 220 may include a device of any transceiver type for communicating with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0230] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of any one of the video detection methods provided in the above embodiments of the present application.

[0231] Exemplary computer program product and storage medium

[0232] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the video detection method according to various embodiments of the present application described in any of the above embodiments of this specification.

[0233] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0234] In addition, an embodiment of the present application may also be a storage medium, on which a computer program is stored. The computer program is executed by a processor to perform the steps in the video detection method according to various embodiments of the present application described in any of the above embodiments of this specification. Specifically, the following steps may be implemented:

[0235] Step 101: Extract each video frame to be detected in the video to be detected, and determine the position information of the facial key points in each video frame to be detected.

[0236] Step 102: Generate a first spatio-temporal feature based on the connection relationship between the facial key points and the position information of the facial key points in each video frame to be detected.

[0237] Step 103: Determine the heart rate time series data corresponding to each video frame to be detected.

[0238] Step 104: Generate a detection result of the video to be detected based on the first spatio-temporal feature and the heart rate time series data corresponding to each video frame to be detected.

[0239] Among them, the detection result is used to predict whether the video to be detected is a deepfake video.

[0240] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0241] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the similarities and common parts among the embodiments, reference can be made to each other. For device embodiments, since they are basically similar to method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description in the method embodiments.

[0242] The steps in the methods of the embodiments of this application can be adjusted, combined, and deleted according to actual needs. The technical features recorded in each embodiment can be replaced or combined.

[0243] In the embodiments of this application, the modules and sub-modules in the devices and terminals can be combined, divided, and deleted according to actual needs.

[0244] In the several embodiments provided by this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are only illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.

[0245] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0246] In addition, in each embodiment of this application, the functional modules or sub-modules can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware, or in the form of software functional modules or sub-modules.

[0247] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0248] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0249] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0250] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video detection method, characterized in that, Including: Extracting each video frame to be detected in the video to be detected, and determining the position information of facial key points in each video frame to be detected; Generating a first spatio-temporal feature based on the connection relationship between the facial key points and the position information of the facial key points in each video frame to be detected; Determining the heart rate time series data corresponding to each video frame to be detected; Generating a detection result of the video to be detected based on the first spatio-temporal feature and the heart rate time series data corresponding to each video frame to be detected; wherein, the detection result is used to predict whether the video to be detected is a deepfake video.

2. The video detection method according to claim 1, wherein The generating a first spatio-temporal feature based on the connection relationship between the facial key points and the position information of the facial key points in each video frame to be detected includes: Generating the structural information of a graph based on the connection relationship between the facial key points; Determining the first feature information of each node in the graph based on the position information of the facial key points in each video frame to be detected; wherein, one node in the graph corresponds to one facial key point; Inputting the structural information of the graph and the first feature information of each node in the graph into a graph convolutional neural network to obtain a first spatio-temporal feature; wherein, the first spatio-temporal feature is used to characterize the spatio-temporal relationship between the facial key points in the video to be detected.

3. The video detection method according to claim 1, characterized in that The heart rate time series data corresponding to each video frame to be detected includes the heart rate time series data corresponding to the facial key points in each video frame to be detected; The generating the detection result of the video to be detected based on the first spatio-temporal feature and the heart rate time series data corresponding to each video frame to be detected includes: Constructing the structural information of a spatio-temporal graph based on the connection relationship between the facial key points and the connection relationship of the same facial key points between each video frame to be detected; Determining the second feature information of each node in the spatio-temporal graph based on the first spatio-temporal feature and the heart rate time series data corresponding to the facial key points in each video frame to be detected; wherein, one node in the spatio-temporal graph corresponds to one facial key point in a video frame to be detected; Inputting the structural information of the spatio-temporal graph and the second feature information of each node in the spatio-temporal graph into a spatio-temporal graph convolutional network to obtain a second spatio-temporal feature; Generating the detection result of the video to be detected based on the second spatio-temporal feature.

4. The video detection method according to claim 3, wherein The generating the detection result of the video to be detected based on the second spatio-temporal feature includes: Performing statistical analysis on the heart rate time series data corresponding to each video frame to be detected to obtain the physiological index corresponding to the video to be detected; Generating the detection result of the video to be detected based on the second spatio-temporal feature and the physiological index corresponding to the video to be detected.

5. The video detection method according to claim 3, characterized in that, The training process of the spatio-temporal graph convolutional network includes: Extracting each sample video frame in the sample video, and determining the position information of facial key points in each sample video frame; wherein, the sample video includes videos corresponding to various forgery categories; Generating a first sample spatio-temporal feature based on the connection relationship between the facial key points and the position information of the facial key points in each sample video frame; Determine the heart rate time series data corresponding to each of the facial key points in each sample video frame; Based on the connection relationships between the facial key points and the connection relationships of the same facial key points between each sample video frame, construct the structural information of the sample spatio-temporal graph; Based on the first sample spatio-temporal feature and the heart rate time series data corresponding to each of the facial key points in each sample video frame, determine the third feature information of each node in the sample spatio-temporal graph; wherein, a node in the sample spatio-temporal graph corresponds to a facial key point in a sample video frame; Input the structural information of the sample spatio-temporal graph and the third feature information of each node in the sample spatio-temporal graph into a spatio-temporal graph convolutional network to obtain a second sample spatio-temporal feature; Based on the second sample spatio-temporal feature, generate a sample detection result of the sample video; wherein, the sample detection result is used to predict whether the sample video is a deepfake video; Based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, and the true detection result of the sample video, train the spatio-temporal graph convolutional network.

6. The video detection method according to claim 5, wherein The method further includes: Generate adversarial samples corresponding to the original video through a generative adversarial network; Based on the original video and the adversarial samples corresponding to the original video, determine the sample video.

7. The video detection method according to claim 5, wherein The method further includes: Identify the sample forgery category corresponding to the sample video; The training of the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, and the true detection result of the sample video includes: Based on the sample forgery category, determine a first facial key point and a second facial key point in the sample spatio-temporal graph; wherein, the first facial key point is located in an area with obvious forgery traces in the sample video, and the second facial key point is located in an area with inconspicuous forgery traces in the sample video; Determine a first weight corresponding to the third feature information of the first facial key point in the sample spatio-temporal graph, and a second weight corresponding to the third feature information of the second facial key point in the sample spatio-temporal graph; wherein, the first weight is greater than the second weight; Based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the first weight corresponding to the third feature information of the first facial key point, the second weight corresponding to the third feature information of the second facial key point, the sample detection result, and the true detection result of the sample video, train the spatio-temporal graph convolutional network.

8. The video detection method according to claim 5, characterized in that, The training of the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, and the true detection result of the sample video includes: Determine the forgery intensity corresponding to each of the facial key points in the sample spatio-temporal graph; Based on the forgery intensity, adjust the loss function of the spatio-temporal graph convolutional network to obtain an adjusted loss function; Train the spatio-temporal graph convolutional network based on the structural information of the sample spatio-temporal graph, the third feature information of each node in the sample spatio-temporal graph, the sample detection result, the true detection result of the sample video, and the adjusted loss function.

9. A video detection device, characterized in that, Comprising: An extraction unit, configured to extract each video frame to be detected in the video to be detected, and determine the position information of the facial key points in each video frame to be detected; A generation unit, configured to generate a first spatio-temporal feature based on the connection relationship between the facial key points and the position information of the facial key points in each video frame to be detected; A first processing unit, configured to determine the heart rate time series data corresponding to each video frame to be detected; A second processing unit, configured to generate a detection result of the video to be detected based on the first spatio-temporal feature and the heart rate time series data corresponding to each video frame to be detected; wherein, the detection result is used to predict whether the video to be detected is a deepfake video.

10. An electronic device, characterized in that, Comprising a memory and a processor; The memory is connected to the processor and is configured to store programs; The processor is configured to implement the video detection method according to any one of claims 1 to 8 by running the programs in the memory.

11. A storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is run by the processor, the video detection method according to any one of claims 1 to 8 is implemented.

12. A computer program product, characterized in that, Comprising computer program instructions, and when the computer program instructions are run by the processor, the processor is caused to execute the video detection method according to any one of claims 1 to 8.