Deep forged video detection method and system based on space-time inconsistency characteristics
By extracting the combination of time-domain features and airspace features between grayscale changes between video frames, using deep learning models forged video detection, the problems of insufficient feature utilization and misjudgment in the prior art are solved, and the accuracy and efficiency of detection are improved.
Patent Information
- Application Number
- CN202510160716.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, when detecting fake videos, the feature utilization rate is insufficient and it is easily misjudged by subtle movements and expressions in the video.
Time domain features are extracted by calculating the grayscale changes between video frames, and combined with the airspace features extracted by convolutional neural networks, and detection is used using deep learning models to improve feature utilization.
The feature utilization rate of the image to be detected is improved, the possibility of misjudgment of subtle motion expressions is reduced, and the ability to detect forged videos is enhanced.
Smart Images

Figure CN119992421A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network security technology, and in particular to a method and system for detecting deep fake videos based on spatiotemporal inconsistency features. Background Art
[0002] Fake videos are often accompanied by malicious false information and spread in large quantities, which has always been an important factor in endangering the network. Attackers use advanced artificial intelligence technology to generate more realistic fake videos, bringing new challenges to network security. Most of the current models use a single video frame as the object to be detected, extract content features from it and use them to build a detection model, which has the problem of insufficient feature utilization. At the same time, the facial expressions in the video are characterized by subtle movements and incoherent time, which can easily be ignored or cause misjudgment.
[0003] Therefore, there is an urgent need for a targeted deep fake video detection method and system based on spatiotemporal inconsistency features. Summary of the invention
[0004] The purpose of the present invention is to provide a deep fake video detection method and system based on time-space inconsistency features, in which the feature changes between frames can reflect more fine-grained correlation information, and continuous picture frames in the video are used as input data objects. The grayscale changes between two frames are calculated to represent the changes in facial expressions, and the extracted time domain features are obtained; the spatial domain features are extracted using a convolutional neural network, and the time domain features are re-extracted. Combined with the previously extracted spatial domain features, a deep learning model is used to detect the video, thereby improving the feature utilization rate of the image to be detected.
[0005] In a first aspect, the present application provides a deep fake video detection method based on spatiotemporal inconsistency features, the method comprising:
[0006] Preprocessing the video to be detected includes: extracting picture frames from the video to be detected, capturing faces on the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set includes face pictures in continuous time;
[0007] Extracting spatial features of the face dataset using a convolutional neural network;
[0008] After obtaining continuous face images in the face data set, the grayscale difference of each pixel between two continuous face images is calculated to obtain the temporal variation trend of the grayscale, and the motion vector of the pixel on different coordinate axes is obtained according to the variation trend, so as to extract the time domain features of the face image according to the motion vector;
[0009] The convolutional neural network is called again for further learning, taking the extracted spatial features and the temporal features of the face image as input at the same time, and outputting deeper temporal features;
[0010] After the deeper time domain features and spatial domain features have undergone multiple feature transformations, a fully connected neural network is used to classify the combination of the deeper time domain features and spatial domain features to obtain a detection result.
[0011] In a second aspect, the present application provides a deep fake video detection system based on spatiotemporal inconsistency features, the system comprising: a preprocessing module, a spatial feature extraction module, a temporal feature extraction module, a deep temporal feature extraction module and a classification detection module;
[0012] The preprocessing module is used to preprocess the video to be detected, including: extracting picture frames from the video to be detected, intercepting faces from the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set contains face pictures of continuous time;
[0013] The spatial feature extraction module is used to extract the spatial features of the face dataset using a convolutional neural network;
[0014] The time domain feature extraction module is used to calculate the grayscale difference of each pixel between two consecutive face pictures after obtaining the consecutive face pictures in the face data set, obtain the grayscale change trend over time, obtain the motion vector of the pixel on different coordinate axes according to the change trend, and extract the time domain features of the face picture according to the motion vector;
[0015] The deep time domain feature extraction module is used to call the convolutional neural network again for further learning, taking the extracted spatial domain features and the time domain features of the face image as input at the same time, and outputting deeper time domain features;
[0016] The classification detection module is used to classify the combination of the deeper time domain features and the deeper space domain features using a fully connected neural network after the deeper time domain features and the deeper space domain features have undergone multiple feature transformations to obtain a detection result.
[0017] In a third aspect, the present application provides a deep fake video detection system based on spatiotemporal inconsistency features, the system comprising a processor and a memory:
[0018] The memory is used to store program code and transmit the program code to the processor;
[0019] The processor is used to execute any one of the four possible methods of the first aspect according to the instructions in the program code.
[0020] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to be executed by a processor to implement any one of the four possible methods in the first aspect.
[0021] Beneficial Effects
[0022] The present invention provides a deep fake video detection method and system based on time-space inconsistency features. The feature changes between frames can reflect more fine-grained correlation information, and continuous picture frames in the video are used as input data objects. The grayscale changes between two frames are calculated to characterize the changes in facial expressions, so as to obtain the extracted time domain features. A convolutional neural network is used to extract spatial domain features, and the time domain features are re-extracted. In combination with the previously extracted spatial domain features, a deep learning model is used to detect the video, thereby improving the feature utilization rate of the image to be detected, and overcoming the problems of the prior art that a single video frame is used as the object to be detected, the feature utilization rate is insufficient, and it is easy to be misjudged by subtle movements and expressions.
[0023] The method and system of the present invention have the following advantages and effects:
[0024] Taking the continuous picture frames in the video as the input data object, the grayscale change between two frames is calculated to represent the change of facial expression, and the extracted time domain features are obtained; the spatial domain features and the time domain features are re-extracted and fused, which can improve the feature utilization rate of the image to be detected;
[0025] Also, it will not misjudge the detection due to subtle movements or expressions. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 is a flow chart of the method of the present invention;
[0028] Figure 2 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION
[0029] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0030] This application provides a deep fake video detection method based on spatiotemporal inconsistency features, such as Figure 1 As shown, the method includes:
[0031] Preprocessing the video to be detected includes: extracting picture frames from the video to be detected, capturing faces on the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set includes face pictures in continuous time;
[0032] Extracting spatial features of the face dataset using a convolutional neural network;
[0033] The differences in facial expressions over time are often subtle movements with incoherence and inconsistency, so it is important to extract time domain features.
[0034] The facial expressions in deep fake videos are usually unnatural. This application needs to further detect abnormal facial expressions and movements based on the existing subtle movement detection.
[0035] After obtaining continuous face images in the face data set, the grayscale difference of each pixel between two continuous face images is calculated to obtain the temporal variation trend of the grayscale, and the motion vector of the pixel on different coordinate axes is obtained according to the variation trend, so as to extract the time domain features of the face image according to the motion vector;
[0036] The convolutional neural network is called again for further learning, taking the extracted spatial features and the temporal features of the face image as input at the same time, and outputting deeper temporal features;
[0037] After the deeper time domain features and spatial domain features have undergone multiple feature transformations, a fully connected neural network is used to classify the combination of the deeper time domain features and spatial domain features to obtain a detection result.
[0038] The use of the fully connected neural network here can also be: after obtaining deeper temporal and spatial features, the features are input into the deep separable convolutional network and the convolutional recurrent gating network to extract the intra-frame and inter-frame information in depth, and finally input into the classifier to obtain the true and false classification results. This further improved classification detection can achieve better results.
[0039] In some preferred embodiments, the face capture of the picture frame includes: based on a face recognition model, detecting and obtaining face position information, cropping the image, numbering the video frames in sequence, and sequentially saving the face pictures.
[0040] In some preferred embodiments, the convolutional neural network adds a BN layer between the convolutional layer and the pooling layer, and uses a Relu nonlinear activation function, so that the convolutional neural network can converge quickly.
[0041] In some preferred embodiments, the number of iterations of the convolutional neural network is selected according to the characteristics of the spatial domain features and the temporal domain features, and the loss function is set to a mean square error loss function.
[0042] Figure 2 The architecture diagram of the deep fake video detection system based on spatiotemporal inconsistency features provided by the present application, the system comprising: a preprocessing module, a spatial feature extraction module, a temporal feature extraction module, a deep temporal feature extraction module and a classification detection module;
[0043] The preprocessing module is used to preprocess the video to be detected, including: extracting picture frames from the video to be detected, intercepting faces from the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set contains face pictures of continuous time;
[0044] The spatial feature extraction module is used to extract the spatial features of the face dataset using a convolutional neural network;
[0045] The time domain feature extraction module is used to calculate the grayscale difference of each pixel between two consecutive face pictures after obtaining the consecutive face pictures in the face data set, obtain the grayscale change trend over time, obtain the motion vector of the pixel on different coordinate axes according to the change trend, and extract the time domain features of the face picture according to the motion vector;
[0046] The deep time domain feature extraction module is used to call the convolutional neural network again for further learning, taking the extracted spatial domain features and the time domain features of the face image as input at the same time, and outputting deeper time domain features;
[0047] The classification detection module is used to classify the combination of the deeper time domain features and the deeper space domain features using a fully connected neural network after the deeper time domain features and the deeper space domain features have undergone multiple feature transformations to obtain a detection result.
[0048] The present application provides a deep fake video detection system based on spatiotemporal inconsistency features, the system comprising: the system comprising a processor and a memory:
[0049] The memory is used to store program code and transmit the program code to the processor;
[0050] The processor is used to execute the method described in any one of all embodiments of the first aspect according to the instructions in the program code.
[0051] The present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to be executed by a processor to implement the method described in any one of all the embodiments of the first aspect.
[0052] In a specific implementation, the present invention further provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, the program may include some or all of the steps in each embodiment of the present invention. The storage medium may be a disk, an optical disk, a read-only storage memory (abbreviated as: ROM) or a random access memory (abbreviated as: RAM), etc.
[0053] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution in the embodiments of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention or some parts of the embodiments.
[0054] The same and similar parts between the various embodiments of this specification can be referred to each other. In particular, for the embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.
[0055] The above-described embodiments of the present invention do not limit the protection scope of the present invention.
Claims
1. A deep fake video detection method based on spatiotemporal inconsistency features, characterized in that: The method comprises: Preprocessing the video to be detected includes: extracting picture frames from the video to be detected, capturing faces on the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set includes face pictures in continuous time; Extracting spatial features of the face dataset using a convolutional neural network; After obtaining continuous face images in the face data set, the grayscale difference of each pixel between two continuous face images is calculated to obtain the temporal variation trend of the grayscale, and the motion vector of the pixel on different coordinate axes is obtained according to the variation trend, so as to extract the time domain features of the face image according to the motion vector; The convolutional neural network is called again for further learning, taking the extracted spatial features and the temporal features of the face image as input at the same time, and outputting deeper temporal features; After the deeper time domain features and spatial domain features have undergone multiple feature transformations, a fully connected neural network is used to classify the combination of the deeper time domain features and spatial domain features to obtain a detection result.
2. The method according to claim 1, characterized in that: The face capture of the picture frame includes: based on the face recognition model, detecting and obtaining the face position information, cropping the image, numbering the video frames in sequence, and sequentially saving the face pictures.
3. The method according to claim 1, characterized in that: The convolutional neural network adds a BN layer between the convolutional layer and the pooling layer, and uses a Relu nonlinear activation function, so that the convolutional neural network can converge quickly.
4. The method according to any one of claims 2 or 3, characterized in that: According to the characteristics of spatial domain features and time domain features, the number of iterations of the convolutional neural network is selected, and the loss function is set to a mean square error loss function.
5. A deep fake video detection system based on spatiotemporal inconsistency features, characterized in that: The system comprises: a preprocessing module, a spatial domain feature extraction module, a temporal domain feature extraction module, a deep temporal domain feature extraction module and a classification detection module; The preprocessing module is used to preprocess the video to be detected, including: extracting picture frames from the video to be detected, intercepting faces from the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set contains face pictures of continuous time; The spatial feature extraction module is used to extract the spatial features of the face dataset using a convolutional neural network; The time domain feature extraction module is used to calculate the grayscale difference of each pixel between two consecutive face pictures after obtaining the consecutive face pictures in the face data set, obtain the grayscale change trend over time, obtain the motion vector of the pixel on different coordinate axes according to the change trend, and extract the time domain features of the face picture according to the motion vector; The deep time domain feature extraction module is used to call the convolutional neural network again for further learning, taking the extracted spatial domain features and the time domain features of the face image as input at the same time, and outputting deeper time domain features; The classification detection module is used to classify the combination of the deeper time domain features and the deeper space domain features using a fully connected neural network after the deeper time domain features and the deeper space domain features have undergone multiple feature transformations to obtain a detection result.
6. A deep fake video detection system based on spatiotemporal inconsistency features, characterized in that: The system comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the method according to any one of claims 1 to 4 according to the instructions in the program code.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program codes, and the program codes are used to be executed by a processor to implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Space-time combination detection method, device and equipment for deeply-forged video
CN116994175A