Deep forged video detection method and system based on space-time inconsistency characteristics

By extracting the combination of time-domain features and airspace features between grayscale changes between video frames, using deep learning models forged video detection, the problems of insufficient feature utilization and misjudgment in the prior art are solved, and the accuracy and efficiency of detection are improved.

CN119992421AInactive Publication Date: 2025-05-13TIANJIN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510160716.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, when detecting fake videos, the feature utilization rate is insufficient and it is easily misjudged by subtle movements and expressions in the video.

Method used

Time domain features are extracted by calculating the grayscale changes between video frames, and combined with the airspace features extracted by convolutional neural networks, and detection is used using deep learning models to improve feature utilization.

Benefits of technology

The feature utilization rate of the image to be detected is improved, the possibility of misjudgment of subtle motion expressions is reduced, and the ability to detect forged videos is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992421A_ABST
    Figure CN119992421A_ABST
Patent Text Reader

Abstract

The invention provides a deep forged video detection method and system based on space-time inconsistency characteristics, which can reflect correlation information with finer granularity through characteristic changes between frames, take continuous picture frames in a video as input data objects, and represent changes of facial expressions by calculating gray level changes between two frames. The extracted time domain features are obtained; a convolutional neural network is used for extracting spatial domain features, time domain features are extracted again, and the video is detected by using a deep learning model in combination with the previously extracted spatial domain features, so that the feature utilization rate of a to-be-detected image is improved, and the defect that the feature utilization rate is insufficient due to the fact that a single video frame is taken as a to-be-detected object in the prior art is overcome. And misjudgment by subtle action expressions is easy to occur.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to a method and system for detecting deep fake videos based on spatiotemporal inconsistency features. Background Art

[0002] Fake videos are often accompanied by malicious false information and spread in large quantities, which has always been an important factor in endangering the network. Attackers use advanced artificial intelligence technology to generate more realistic fake videos, bringing new challenges to network security. Most of the current models use a single video frame as the object to be detected, extract content features from it and use them to build a detection model, which has the problem of insufficient feature utilization. At the same time, the facial expressions in the video are characterized by subtle movements and incoherent time, which can easily be ignored or cause misjudgment.

[0003] Therefore, there is an urgent need for a targeted deep fake video detection method and system based on spatiotemporal inconsistency features. Summary of the invention

[0004] The purpose of the present invention is to provide a deep fake video detection method and system based on time-space inconsistency features, in which the feature changes between frames can reflect more fine-grained correlation information, and continuous picture frames in the video are used as input data objects. The grayscale changes between two frames are calculated to represent the changes in facial expressions, and the extracted time domain features are obtained; the spatial domain features are extracted using a convolutional neural network, and the time domain features are re-extracted. Combined with the previously extracted spatial domain features, a deep learning model is used to detect the video, thereby improving the feature utilization rate of the image to be detected.

[0005] In a first aspect, the present application provides a deep fake video detection method based on spatiotemporal inconsistency features, the method comprising:

[0006] Preprocessing the video to be detected includes: extracting picture frames from the video to be detected, capturing faces on the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set includes face pictures in continuous time;

[0007] Extracting spatial features of the face dataset using a convolutional neural network;

[0008] After obtaining continuous face images in the face data set, the grayscale difference of each pixel between two continuous face images is calculated to obtain the temporal variation trend of the grayscale, and the motion vector of the pixel on different coordinate axes is obtained according to the variation trend, so as to extract the time domain features of the face image according to the motion vector;

[0009] The convolutional neural network is called again for further learning, taking the extracted spatial features and the temporal features of the face image as input at the same time, and outputting deeper temporal features;

[0010] After the deeper time domain features and spatial domain features have undergone multiple feature transformations, a fully connected neural network is used to classify the combination of the deeper time domain features and spatial domain features to obtain a detection result.

[0011] In a second aspect, the present application provides a deep fake video detection system based on spatiotemporal inconsistency features, the system comprising: a preprocessing module, a spatial feature extraction module, a temporal feature extraction module, a deep temporal feature extraction module and a classification detection module;

[0012] The preprocessing module is used to preprocess the video to be detected, including: extracting picture frames from the video to be detected, intercepting faces from the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set contains face pictures of continuous time;

[0013] The spatial feature extraction module is used to extract the spatial features of the face dataset using a convolutional neural network;

[0014] The time domain feature extraction module is used to calculate the grayscale difference of each pixel between two consecutive face pictures after obtaining the consecutive face pictures in the face data set, obtain the grayscale change trend over time, obtain the motion vector of the pixel on different coordinate axes according to the change trend, and extract the time domain features of the face picture according to the motion vector;

[0015] The deep time domain feature extraction module is used to call the convolutional neural network again for further learning, taking the extracted spatial domain features and the time domain features of the face image as input at the same time, and outputting deeper time domain features;

[0016] The classification detection module is used to classify the combination of the deeper time domain features and the deeper space domain features using a fully connected neural network after the deeper time domain features and the deeper space domain features have undergone multiple feature transformations to obtain a detection result.

[0017] In a third aspect, the present application provides a deep fake video detection system based on spatiotemporal inconsistency features, the system comprising a processor and a memory:

[0018] The memory is used to store program code and transmit the program code to the processor;

[0019] The processor is used to execute any one of the four possible methods of the first aspect according to the instructions in the program code.

[0020] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to be executed by a processor to implement any one of the four possible methods in the first aspect.

[0021] Beneficial Effects

[0022] The present invention provides a deep fake video detection method and system based on time-space inconsistency features. The feature changes between frames can reflect more fine-grained correlation information, and continuous picture frames in the video are used as input data objects. The grayscale changes between two frames are calculated to characterize the changes in facial expressions, so as to obtain the extracted time domain features. A convolutional neural network is used to extract spatial domain features, and the time domain features are re-extracted. In combination with the previously extracted spatial domain features, a deep learning model is used to detect the video, thereby improving the feature utilization rate of the image to be detected, and overcoming the problems of the prior art that a single video frame is used as the object to be detected, the feature utilization rate is insufficient, and it is easy to be misjudged by subtle movements and expressions.

[0023] The method and system of the present invention have the following advantages and effects:

[0024] Taking the continuous picture frames in the video as the input data object, the grayscale change between two frames is calculated to represent the change of facial expression, and the extracted time domain features are obtained; the spatial domain features and the time domain features are re-extracted and fused, which can improve the feature utilization rate of the image to be detected;

[0025] Also, it will not misjudge the detection due to subtle movements or expressions. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0027] Figure 1 is a flow chart of the method of the present invention;

[0028] Figure 2 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION

[0029] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0030] This application provides a deep fake video detection method based on spatiotemporal inconsistency features, such as Figure 1 As shown, the method includes:

[0031] Preprocessing the video to be detected includes: extracting picture frames from the video to be detected, capturing faces on the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set includes face pictures in continuous time;

[0032] Extracting spatial features of the face dataset using a convolutional neural network;

[0033] The differences in facial expressions over time are often subtle movements with incoherence and inconsistency, so it is important to extract time domain features.

[0034] The facial expressions in deep fake videos are usually unnatural. This application needs to further detect abnormal facial expressions and movements based on the existing subtle movement detection.

[0035] After obtaining continuous face images in the face data set, the grayscale difference of each pixel between two continuous face images is calculated to obtain the temporal variation trend of the grayscale, and the motion vector of the pixel on different coordinate axes is obtained according to the variation trend, so as to extract the time domain features of the face image according to the motion vector;

[0036] The convolutional neural network is called again for further learning, taking the extracted spatial features and the temporal features of the face image as input at the same time, and outputting deeper temporal features;

[0037] After the deeper time domain features and spatial domain features have undergone multiple feature transformations, a fully connected neural network is used to classify the combination of the deeper time domain features and spatial domain features to obtain a detection result.

[0038] The use of the fully connected neural network here can also be: after obtaining deeper temporal and spatial features, the features are input into the deep separable convolutional network and the convolutional recurrent gating network to extract the intra-frame and inter-frame information in depth, and finally input into the classifier to obtain the true and false classification results. This further improved classification detection can achieve better results.

[0039] In some preferred embodiments, the face capture of the picture frame includes: based on a face recognition model, detecting and obtaining face position information, cropping the image, numbering the video frames in sequence, and sequentially saving the face pictures.

[0040] In some preferred embodiments, the convolutional neural network adds a BN layer between the convolutional layer and the pooling layer, and uses a Relu nonlinear activation function, so that the convolutional neural network can converge quickly.

[0041] In some preferred embodiments, the number of iterations of the convolutional neural network is selected according to the characteristics of the spatial domain features and the temporal domain features, and the loss function is set to a mean square error loss function.

[0042] Figure 2 The architecture diagram of the deep fake video detection system based on spatiotemporal inconsistency features provided by the present application, the system comprising: a preprocessing module, a spatial feature extraction module, a temporal feature extraction module, a deep temporal feature extraction module and a classification detection module;

[0043] The preprocessing module is used to preprocess the video to be detected, including: extracting picture frames from the video to be detected, intercepting faces from the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set contains face pictures of continuous time;

[0044] The spatial feature extraction module is used to extract the spatial features of the face dataset using a convolutional neural network;

[0045] The time domain feature extraction module is used to calculate the grayscale difference of each pixel between two consecutive face pictures after obtaining the consecutive face pictures in the face data set, obtain the grayscale change trend over time, obtain the motion vector of the pixel on different coordinate axes according to the change trend, and extract the time domain features of the face picture according to the motion vector;

[0046] The deep time domain feature extraction module is used to call the convolutional neural network again for further learning, taking the extracted spatial domain features and the time domain features of the face image as input at the same time, and outputting deeper time domain features;

[0047] The classification detection module is used to classify the combination of the deeper time domain features and the deeper space domain features using a fully connected neural network after the deeper time domain features and the deeper space domain features have undergone multiple feature transformations to obtain a detection result.

[0048] The present application provides a deep fake video detection system based on spatiotemporal inconsistency features, the system comprising: the system comprising a processor and a memory:

[0049] The memory is used to store program code and transmit the program code to the processor;

[0050] The processor is used to execute the method described in any one of all embodiments of the first aspect according to the instructions in the program code.

[0051] The present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to be executed by a processor to implement the method described in any one of all the embodiments of the first aspect.

[0052] In a specific implementation, the present invention further provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, the program may include some or all of the steps in each embodiment of the present invention. The storage medium may be a disk, an optical disk, a read-only storage memory (abbreviated as: ROM) or a random access memory (abbreviated as: RAM), etc.

[0053] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution in the embodiments of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention or some parts of the embodiments.

[0054] The same and similar parts between the various embodiments of this specification can be referred to each other. In particular, for the embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.

[0055] The above-described embodiments of the present invention do not limit the protection scope of the present invention.

Claims

1. A deep fake video detection method based on spatiotemporal inconsistency features, characterized in that: The method comprises: Preprocessing the video to be detected includes: extracting picture frames from the video to be detected, capturing faces on the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set includes face pictures in continuous time; Extracting spatial features of the face dataset using a convolutional neural network; After obtaining continuous face images in the face data set, the grayscale difference of each pixel between two continuous face images is calculated to obtain the temporal variation trend of the grayscale, and the motion vector of the pixel on different coordinate axes is obtained according to the variation trend, so as to extract the time domain features of the face image according to the motion vector; The convolutional neural network is called again for further learning, taking the extracted spatial features and the temporal features of the face image as input at the same time, and outputting deeper temporal features; After the deeper time domain features and spatial domain features have undergone multiple feature transformations, a fully connected neural network is used to classify the combination of the deeper time domain features and spatial domain features to obtain a detection result.

2. The method according to claim 1, characterized in that: The face capture of the picture frame includes: based on the face recognition model, detecting and obtaining the face position information, cropping the image, numbering the video frames in sequence, and sequentially saving the face pictures.

3. The method according to claim 1, characterized in that: The convolutional neural network adds a BN layer between the convolutional layer and the pooling layer, and uses a Relu nonlinear activation function, so that the convolutional neural network can converge quickly.

4. The method according to any one of claims 2 or 3, characterized in that: According to the characteristics of spatial domain features and time domain features, the number of iterations of the convolutional neural network is selected, and the loss function is set to a mean square error loss function.

5. A deep fake video detection system based on spatiotemporal inconsistency features, characterized in that: The system comprises: a preprocessing module, a spatial domain feature extraction module, a temporal domain feature extraction module, a deep temporal domain feature extraction module and a classification detection module; The preprocessing module is used to preprocess the video to be detected, including: extracting picture frames from the video to be detected, intercepting faces from the picture frames, and storing them in sequence to obtain a face data set, wherein the face data set contains face pictures of continuous time; The spatial feature extraction module is used to extract the spatial features of the face dataset using a convolutional neural network; The time domain feature extraction module is used to calculate the grayscale difference of each pixel between two consecutive face pictures after obtaining the consecutive face pictures in the face data set, obtain the grayscale change trend over time, obtain the motion vector of the pixel on different coordinate axes according to the change trend, and extract the time domain features of the face picture according to the motion vector; The deep time domain feature extraction module is used to call the convolutional neural network again for further learning, taking the extracted spatial domain features and the time domain features of the face image as input at the same time, and outputting deeper time domain features; The classification detection module is used to classify the combination of the deeper time domain features and the deeper space domain features using a fully connected neural network after the deeper time domain features and the deeper space domain features have undergone multiple feature transformations to obtain a detection result.

6. A deep fake video detection system based on spatiotemporal inconsistency features, characterized in that: The system comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the method according to any one of claims 1 to 4 according to the instructions in the program code.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program codes, and the program codes are used to be executed by a processor to implement the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Space-time combination detection method, device and equipment for deeply-forged video

    CN116994175A