A method, system and storage medium for detecting forged videos

By predefined key points and constructing sparsely associated optical flow feature maps, the problems of high complexity of existing video authenticity detection technology and high computing resource consumption are solved, and the effects of fast detection speed and high accuracy are achieved, which are suitable for embedded platforms.

CN114120198BActive Publication Date: 2025-06-27WUHAN MARITIME COMMUNICATION RESEARCH INSTITUTE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111431151.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-06-27
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

The existing video authenticity detection technology is complex, consumes a lot of computing resources, has a long training time, and requires high hardware resources when deployed on embedded platforms, and has slow detection speed.

Method used

By predefined multiple key points, divided them into multiple regions, calculate the optical flow value and associated optical flow value of each key point, build a sparsely associated optical flow feature map, and input it into a detection model based on a convolutional neural network to output the detection results.

Benefits of technology

It achieves fast detection speed, high accuracy, and is suitable for embedded platforms, occupies less memory and has low hardware requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120198B_ABST
    Figure CN114120198B_ABST
Patent Text Reader

Abstract

The present invention discloses a forged video detection method, system and storage medium. The method includes the steps of: predefined a plurality of key points, and dividing the plurality of key points into a plurality of regions; extracting images from the video to be detected, and detecting key points on each extracted image; calculating the optical flow value of each key point according to the coordinate displacement of the same key point between two adjacent frames, and then calculating the associated optical flow value of each key point according to the optical flow value of the key point and the optical flow values of other key points in the region to which the key point belongs, and constructing a sparse associated optical flow feature map according to the associated optical flow value; inputting the sparse associated optical flow feature map into a trained detection model to output a detection result. The present invention has the advantages of fast detection speed and high accuracy at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and deep learning, and more specifically, relates to a forged video detection method, system and storage medium. Background Art

[0002] The video forgery method refers to a technology that uses a deep neural network to replace real human faces in original images or videos with synthetic human faces to create false information, and has been misused for creating fake news, pranks and financial fraud, causing serious impacts on society. The purpose of the video authenticity detection method is to detect such forged videos and maintain information security.

[0003] Video authenticity detection methods can be divided into two categories: single-frame-based methods and sequence-based methods. With the continuous development of deep learning technology, the models used in both methods are becoming more and more complex, and the demand for computing resources and time is also increasing, resulting in high training costs and slow detection speeds.

[0004] The XceptionNet algorithm is a milestone in single-frame detection methods and is usually used as a baseline for algorithm performance comparison. The XceptionNet algorithm has tens of millions of parameters and consists of 14 modules composed of 36 convolutional layers. To improve performance and generalization ability, subsequent single-frame-based methods are more complex than a single CNN model. The new method called Face X-ray is a method for determining the authenticity of a face image by predicting a grayscale image and identifying its location when there is a mixed boundary. To generate a grayscale image of the same size as the input image, the HRNet model used by the Face X-ray method requires a large amount of computing resources. Another new multi-attention detection network consists of an attention module, a texture enhancement block, and a bilinear attention pool. The architecture of this network is very complex and consumes a lot of computing resources.

[0005] Most state-of-the-art video authenticity detection techniques only analyze the spatial information of a single frame and rarely explore the temporal information between frames. However, the temporal information between consecutive frames is crucial for detecting the authenticity of videos and helps detect unnatural artifacts that exist between video frames. 3D convolutional neural networks (3DCNNs) are a sequence-based approach, in which the R3D model outperforms models such as C3D and I3D and is the best-performing 3DCNN model for detection. However, 3D convolutions in 3DCNNs have more parameters and a larger computational cost than 2D convolutions in general CNNs. In addition, recurrent neural networks are powerful tools for fully exploiting temporal information and are therefore also used to extract temporal information. A residual network algorithm based on LSTM consists of ConvLSTM units and residual paths connected. Another is a method based on convolutional neural networks (CNNs) and recurrent neural networks (RNNs) with an automatic weighting mechanism to emphasize the most reliable regions of the detected face when determining sequence-level predictions. Although the EfficientNet-b5 model used provides a good trade-off between network parameters and classification accuracy, the automatic face weighting block and GRU model add additional overhead.

[0006] The complexity of the above methods increases sharply while the accuracy improves slightly. These models impose a huge pressure on computing resources and require a long training time. Especially when dealing with increasingly large datasets, the algorithms usually require joint training of high-performance multi-block GPUs, and the training time is measured in days. In addition, when these models are deployed to embedded platforms, they have high requirements for hardware resources and a slow detection speed. Summary of the Invention

[0007] In view of at least one defect or improvement requirement of the prior art, the present invention provides a forged video detection method, system and storage medium, which has the advantages of fast detection speed and high accuracy.

[0008] To achieve the above object, according to the first aspect of the present invention, there is provided a forged video detection method, including the steps of:

[0009] Pre-define a plurality of key points and divide the plurality of key points into a plurality of regions;

[0010] Extract images from the video to be detected and detect key points on the extracted images;

[0011] Calculate the optical flow value of each key point according to the coordinate displacement of the same key point between two adjacent frames, and then calculate the associated optical flow value of each key point according to the optical flow value of the key point and the optical flow values of other key points in the region to which the key point belongs. Construct a sparse associated optical flow feature map according to the associated optical flow values of each key point on multiple frames of images;

[0012] Input the sparse correlation optical flow feature map into the trained detection model to output the detection result.

[0013] Further, the extracted image is a face image, and multiple predefined key points are all facial key points. The facial key points are divided into twelve regions: left eye, right eye, left eyelid, right eyelid, left eyebrow, right eyebrow, left cheek, right cheek, upper lip, lower lip, nose, and head.

[0014] Further, detecting key points on the extracted image includes the steps of:

[0015] It is predetermined that each key point has a unique index number. Denote the coordinate values of the detected key points as (x i,j , y i,j ), where i represents the index number of the key point and j represents the frame number. Generate a key point detection file to record the coordinate values of all key points in each frame. If no key points are detected, no key point detection file is generated.

[0016] Further, if some key points cannot be detected in the extracted image, define the undetected key points as missing points. Then, use predefined special values to represent the coordinate values of the missing points in the key point detection file.

[0017] Further, denote the number of key points as N. If key points are detected in N + 1 consecutive frames of images, construct a sparse correlation optical flow feature map based on the correlation optical flow values of the key points in the N + 1 frames of images.

[0018] Further, calculate the optical flow value of a key point and the weighted sum of the optical flow values of other key points in the region to which the key point belongs as the correlation optical flow value of the key point.

[0019] Further, the sparse correlation optical flow feature map contains features of inconsistent facial expressions.

[0020] Further, the detection model is implemented based on a convolutional neural network and includes 6 convolutional layers, 4 max pooling layers, and 3 fully connected layers. For each convolutional layer, the size of the convolutional kernel is 3×3, and the stride of the last convolutional layer is 2.

[0021] According to the second aspect of the present invention, there is provided a forged video detection system, including:

[0022] A key point definition module for predefined multiple key points and dividing the multiple key points into multiple regions;

[0023] A detection module for extracting images from the video to be detected and detecting key points on each frame of the extracted image;

[0024] A feature extraction module, which is used to calculate the optical flow value of each key point according to the coordinate displacement of the same key point between two adjacent frames, and then calculate the associated optical flow value of each key point according to the optical flow value of the key point and the optical flow values of other key points in the area where the key point belongs, and construct a sparse associated optical flow feature map according to the associated optical flow values of each key point on multiple frames of images;

[0025] An authenticity verification module, which is used to input the sparse associated optical flow feature map into the trained detection model and output the detection result.

[0026] According to the third aspect of the present invention, a storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0027] Generally speaking, compared with the prior art, the present invention has the following beneficial effects:

[0028] (1) By calculating the associated optical flow value of each key point according to the optical flow value of the key point and the optical flow values of other key points in the area where the key point belongs, the present invention takes into account the spatio-temporal information and the motion correlation between key points. On the one hand, it can greatly compress the size and dimension of the input features, reduce the number of training parameters and training time of the detection model, is suitable for deployment on an embedded platform, occupies less memory, has low hardware requirements, and has a fast detection speed. On the other hand, it also has a high detection accuracy.

[0029] (2) Further, the extracted image is a face image, the predefined multiple key points are all facial key points, and the extracted sparse associated optical flow feature map contains the motion information of facial muscle groups and the spatio-temporal information of facial expression changes, and can make full use of the characteristics such as the stiffness, disharmony of facial muscles and the inconsistency of expression changes in the forged face video for forged video detection. Description of the Drawings

[0030] Figure 1 is a flowchart of a forged video detection method according to an embodiment of the present invention;

[0031] Figure 2 is a schematic diagram of key point area division according to an embodiment of the present invention;

[0032] Figure 3 is a schematic diagram of a sparse associated optical flow feature map according to an embodiment of the present invention. Detailed Embodiment

[0033] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0034] As Figure 1 shown, a method for detecting forged videos according to an embodiment of the present invention includes the steps of:

[0035] S1, predefined a plurality of key points, and divide the plurality of key points into a plurality of regions.

[0036] The predefined key points and regions are for subsequent data processing.

[0037] Further, the predefined plurality of key points are all facial key points, and the plurality of facial key points are divided into twelve regions: left eye, right eye, left eyelid, right eyelid, left eyebrow, right eyebrow, left cheek, right cheek, upper lip, lower lip, nose and head.

[0038] In one embodiment, as Figure 2 shown, 68 facial key points are predefined, each key point has a unique index number, which are respectively from 1 to 68, and the corresponding relationship between the region (i.e., the facial motion region) and the 68 facial key points is: left eye (key points 43-48), right eye (key points 37-42), left eyelid (key points 23-27 and 43-46), right eyelid (key points 18-22 and 37-40), left eyebrow (key points 23-27), right eyebrow (key points 18-22), left cheek (key points 12-16, 28-31, 36, 53-55), right cheek (key points 2-7, 28-32 and 49-51), upper lip (key points 49-55 and 61-65), lower lip (key points 56-60 and 66-68), nose (key points 28-36) and head (key points 1-31, 34, 52, 63, 67, 68) these twelve motion regions.

[0039] There is a situation where a certain key point belongs to two or more regions. For example, key points 23 - 27 belong to both the left eyebrow and the left eyelid at the same time. The reason is that the movement of the left eyebrow (such as frowning, raising the eyebrows, etc.) requires the joint action of points 23 - 27, and the movement of the left eyelid (such as blinking, squinting, rolling the eyes, etc.) requires key points 23 - 27 and 43 - 46 to complete together. When calculating the associated optical flow value, it is calculated according to the region, that is, the associated optical flow values of this key point in the two regions are calculated separately, and the results are added together, indicating that the importance and contribution rate of this key point in the two regions are cumulative processes. In this way, key points that appear in multiple regions contain more information and are more important than key points that only appear in one region.

[0040] S2. Extract images from the video to be detected, and detect key points on the extracted images.

[0041] Extracting images from the video to be detected can be extracting each consecutive frame image. This is because the sparse associated optical flow feature map extracted by the present invention can solve the problem of excessive calculation amount. In this way, extracting consecutive frames does not have the problem of excessive calculation amount, and consecutive frames have very good correlation in time, which can capture the entire process of facial expression changes. In addition, a micro-expression can flash by in 1 / 25 seconds, and the frame rate of ordinary videos is 30 FPS, which can roughly record the change process of expressions. If interval sampling is used, important information about expression changes may be lost.

[0042] Detecting key points can be implemented by any method in the prior art.

[0043] Further, before extracting key points, first intercept the portrait from the extracted image, so that extracting images from the video to be detected can realize the segmentation of the portrait area and the background area, achieving the purpose of eliminating the interference caused by shooting scene changes, shooting equipment jitter, etc.

[0044] Further, detecting key points includes the steps: predetermining that each key point has a unique index number, and recording the coordinate values of the detected key points as (x i,j , y i,j ), where i represents the index number of the key point, j represents the frame number, and generating a key point detection file to record the coordinate values of all key points in each frame. If no key points are detected, no key point detection file is generated.

[0045] In one embodiment, the VideoCapture function in the OpenCV open-source library is then used to obtain a video handle, and each frame of the picture is obtained in sequence. Then, a folder is created for each video, and the MTCNN face detection algorithm is used to intercept the face images in each frame of the picture and save the.png pictures named according to the frame number in the corresponding folder. The face key-point detection algorithm in the Dlib library is used to detect the coordinate values (x i,j , y i,j ) of 68 face key points on the face of the same person in each frame of the picture, where i represents the index of the key point and j represents the frame number. The coordinate values of 68 points are saved in an.npy file named according to the frame number and stored in the folder corresponding to each video.

[0046] Furthermore, if some key points cannot be detected in the extracted image, the undetected key points are defined as missing points, and the coordinate values of the missing points are represented by predefined special values in the key-point detection file. For example, when the person's head rotates during shooting, such as a side face or a lowered head, the face information will be missing. At this time, all 68 face key points cannot be detected, and the undetected key points are defined as missing points and represented by special values. These missing points contain the motion information of the head and are of great significance for fake video detection.

[0047] In one embodiment, the special value is taken as 50, representing 50 px. Because the displacement of key points between two adjacent frames is usually only a few pixel points. The range of the length / width of the face image area is between 40 px and 200 px, and the special value of 50 px is greater than the normal displacement value and the displacement value caused by the forgery algorithm, so the missing points can be specially marked.

[0048] S3. Calculate the optical flow value of each key point according to the coordinate displacement of the same key point between two adjacent frames, then calculate the associated optical flow value of each key point according to the optical flow value of the key point and the optical flow values of other key points in the area to which the key point belongs, and construct a sparse associated optical flow feature map according to the associated optical flow values of each key point on multiple frames of images.

[0049] The associated optical flow value describes the motion correlation between multiple key points. The ordinary optical flow value can only reflect the individual displacement of each point. However, the change of facial expression is not caused by the independent change of each key point, but by the overall movement of the facial muscle group. According to the movement law of the muscle group, the present invention divides the key points into different movement areas, and reflects the overall movement trend of the muscle by calculating the associated optical flow value of the key points in each movement area. This associated optical flow value contains more spatial information and movement information than the ordinary optical flow value, and is more helpful for discovering the stiffness and inconsistency of facial expressions in forged videos.

[0050] The sparse associated optical flow feature map describes the motion correlation between multiple key points in multiple frames of images. The multi-frame sparse associated optical flow feature map extracts the spatial and temporal information in multiple frames of face images, depicting the motion process of facial muscle groups within 2 seconds. This is conducive to discovering phenomena such as sudden expression changes, expression retention, stiffness, etc. in forged videos.

[0051] Other key points in the area where the key point belongs are obtained according to the definition in step S1.

[0052] By calculating the associated optical flow value of each key point based on the optical flow value of the key point and the optical flow values of other key points in the area where the key point belongs, on the one hand, spatio-temporal information is considered, and on the other hand, the motion correlation between key points is also considered.

[0053] Specifically for facial key points, the extracted sparse associated optical flow feature map contains the motion information of facial muscle groups and the spatio-temporal information of facial expression transformation, including features such as facial muscle stiffness, disharmony, and inconsistent expression changes in forged face videos. Using these features can effectively detect forged videos.

[0054] Furthermore, denote the number of key points as N. If key points are detected in N + 1 consecutive frames of images, a sparse associated optical flow feature map is constructed according to the associated optical flow values of the key points in the N + 1 frames of images.

[0055] In one embodiment, first read 69 consecutive.npy files, and form the face coordinate matrices X and Y (X ∈ R 69×68 , Y ∈ R 69×68) from 69 groups of coordinate values. As shown in Equation 1, the matrices X and Y respectively represent the abscissa values and ordinate values of 68 facial key points in 69 consecutive frames.

[0056]

[0057] Then, calculate the displacements u and v in the horizontal and vertical directions between two adjacent frames for the same point. u i,j is the optical flow value of key point i in the horizontal direction in the j-th frame, and v i,j is the optical flow value of key point i in the vertical direction in the j-th frame, as shown in the following formula:

[0058] u i,j = f(x i,j ) = x i,j+1 - x i,j v i,j = f(y i,j ) = y i,j+1 - y i,j

[0059] Two matrices U and V (U ∈ R68×68, V ∈ R68×68) corresponding to matrices X and Y are obtained. Matrices U and V respectively represent the optical flow values of 68 key points in the horizontal and vertical directions in 69 consecutive frames, as shown in the following formula:

[0060]

[0061] Then, the associated optical flow values of 68 key points in 69 adjacent face images are fused into a sparse associated optical flow feature map.

[0062] The calculation method of the associated optical flow value (p i,j , q i,j ) between two adjacent frames for each key point is to add the optical flow value of each key point to the weighted sum of the optical flow values of other key points within the same facial motion area as the associated optical flow value of the key point.

[0063] Furthermore, for the case where a certain key point belongs to two or more regions, the associated optical flow value corresponding to each region to which the key point belongs is calculated based on the optical flow value of the key point and the optical flow values of other key points in each region to which the key point belongs, and then the final associated optical flow value is calculated based on the associated optical flow values corresponding to the regions to which the key point belongs. For example, for the case where a key point belongs to two regions, when calculating the associated optical flow value, it is calculated according to the region, that is, the associated optical flow value 1 of the region 1 to which the key point belongs is calculated based on the optical flow value of the key point and the optical flow values of other key points in the region 1 to which the key point belongs, and the associated optical flow value 2 of the region 2 to which the key point belongs is calculated based on the optical flow value of the key point and the optical flow values of other key points in the region 2 to which the key point belongs, and then the associated optical flow value 1 is added to the associated optical flow value 2 to obtain the final associated optical flow value of the key point, indicating that the importance and contribution rate of the key point in the two regions are cumulative processes. In this way, the key points that appear in multiple regions contain more information and are more important than the key points that only appear in one region.

[0064] In one embodiment, the calculation formula of the associated optical flow value (p i,j , q i,j ) is:

[0065]

[0066] Where D i represents the set of other key points within the region where the i-th key point is located, and k is the key point in the set D i . |u i,j -u k,j | is the absolute value of the horizontal distance difference between the i-th key point and the j-th key point; is the weight of the key point optical flow value. It is inversely proportional to the distance between two key points, indicating that the closer the key points are to the i-th key point, the greater the motion correlation and the higher the contribution rate. The meaning of the weight in the vertical direction is the same.

[0067] The associated optical flow values (p i,j , q i,j ) form the face sparse associated optical flow feature map F (F ∈ R68×68×2), which represents the associated optical flow values of 68 key points on the face of the same person in 69 consecutive frames of the video.

[0068] S4. Input the sparse associated optical flow feature map into the trained detection model to output the detection result.

[0069] The detection model is implemented based on a convolutional neural network. The processing of the detection model includes the following steps:

[0070] (1) Download the open-source dataset, which is divided into a training set, a test set, and a validation set;

[0071] Pre-download the FF++ dataset. In one embodiment, it mainly targets low-quality (video quality is C40) video data. This dataset consists of 1000 original video sequences and 1000 videos forged by each of the four forgery methods, for a total of 5000 videos. The four forgery methods are Deepfakes, Face2Face, FaceSwap, and Neural Textures methods. Among the 1000 videos of each method, 720 specific videos are selected for training, 140 videos for validation, and 140 videos for testing.

[0072] (2) Perform the same processing as in steps S2 and S3 on the videos in the training set, test set, and validation set. Extract the sparse associated optical flow feature map from the videos. Additionally, label the videos, train the convolutional neural network of the detection model using the training set and save the parameters, test the detection model using the test set, and validate the detection model using the validation set.

[0073] In one example, the specific process of the design and training of the convolutional neural network is as follows:

[0074] Construct a lightweight convolutional neural network. As shown in Table 1, the model of the convolutional neural network consists of 6 convolutional layers, 4 max-pooling layers, and 3 fully connected layers. For each convolutional layer, the size of the convolutional kernel is 3×3, which means convolving the motion trajectories of 3 adjacent face key points in 4 adjacent frames. The stride of the last convolutional layer is 2, which convolves the feature size from 7×7 to 4×4. Through the 4 max-pooling layers and the last convolutional layer of the entire network, the size of the feature map is reduced to 1 / 16 of its original size.

[0075] Table 1. Convolutional Neural Network Structure Table

[0076]

[0077]

[0078] Train the convolutional neural network. Use the sparse correlation optical flow feature map to train the convolutional neural network. Input the feature map obtained by preprocessing the training set into the convolutional neural network, predict the classification result, and calculate the cross-entropy between it and the corresponding true and false classification labels as the loss function. Backpropagate according to the loss function to train the network parameters.

[0079]

[0080] The loss function is shown as above, where m represents the number of samples, x is the input feature, y is the true classification label, h θ (x (i) ) is the predicted classification result, and θ is the parameter of model training.

[0081] After each round of training iteration, input the feature map of the validation set into the network for prediction to obtain the accuracy of the validation set. When the accuracy of the training set converges or is close to convergence, and the accuracy of the validation set is the highest, end the training of the model and save the parameters of the model at this time as the optimal parameters of the model.

[0082] Test the effect of the detection model.

[0083] First, load the trained model parameters, and then predict 140 videos in the test set in turn. For each video, first use the MTCNN face detection algorithm to intercept the face image, and then use the Dlib face key point detection algorithm to detect 68 key points of the same person's face in 69 consecutive frames before performing subsequent feature extraction and classification prediction operations; otherwise, abandon this round and start detecting 69 consecutive frames from the next frame of the picture.

[0084] Then, form coordinate matrices X and Y with 68 face key points in 69 consecutive frames, and calculate the corresponding optical flow matrices U and V, and combine them into a face sparse correlation optical flow feature map F.

[0085] Input the feature map F into the convolutional neural network for prediction, and output the probability values of true and false and the predicted label values. Save the predicted label values of this round and the corresponding true labels.

[0086] Statistically calculate the overall accuracy of the multiple judgment results of all videos in the entire test set.

[0087] Statistical analysis is performed on all predicted label values and true label values of all samples in the test set to calculate the overall accuracy of the test set and the accuracy corresponding to each method. The results are shown in the following table:

[0088] Table 2. Test Results of This Experiment

[0089]

[0090] Table 3 shows the comparison of the method proposed in the present invention and the existing optimal algorithm in terms of performance, including the computer resources used and the training time.

[0091] Table 3. Performance Comparison of Different Algorithms

[0092]

[0093] The overall accuracy of the present invention is slightly lower than that of the XceptionNet algorithm and the R3D algorithm. However, the method of the present invention uses the fewest parameters, has the lowest GPU occupancy, and the fastest training speed. As shown in Table 3, the parameters of the XceptionNet algorithm and the R3D algorithm are dozens of times, or even more than a hundred times, that of the method of the present invention. In addition, the training time of the present invention is only 8 minutes, while other algorithms require several days. During the training process, the GPU memory used by the present invention is one-tenth of that of other methods. This fully demonstrates the advantages of the present invention in terms of performance: extremely short training time, few parameters, and low GPU resource occupancy. Finally, when deployed on an embedded platform, the advantage of the few parameters of the present invention will be further reflected, the model has low hardware requirements, and the detection speed is fast.

[0094] A forged video detection system according to an embodiment of the present invention includes:

[0095] A key point definition module, configured to pre-define a plurality of key points and divide the plurality of key points into a plurality of regions;

[0096] A detection module, configured to extract images from a video to be detected and detect key points on each extracted image;

[0097] A feature extraction module, configured to calculate the optical flow value of each key point according to the coordinate displacement of the same key point between two adjacent frames, and then calculate the associated optical flow value of each key point according to the optical flow value of the key point and the optical flow values of other key points in the region to which the key point belongs, and construct a sparse associated optical flow feature map according to the associated optical flow values of each key point on multiple frames of images;

[0098] A forgery identification module, configured to input the sparse associated optical flow feature map into a trained detection model and output a detection result.

[0099] The implementation principle and technical effect of the system are similar to those of the above method, and will not be elaborated here.

[0100] An embodiment of the present invention further provides a storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the technical solution of any one of the above-described embodiments of the forged video detection method. Its implementation principle and technical effect are similar to those of the above method, and will not be elaborated here.

[0101] It must be noted that in any of the above embodiments, the methods do not necessarily need to be executed in the order of the serial numbers. As long as it cannot be inferred from the execution logic that they must be executed in a certain order, it means that they can be executed in any other possible order.

[0102] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for detecting forged videos, characterized in that, Including the steps: Pre - define multiple key points and divide the multiple key points into multiple regions; Extract images from the video to be detected and detect key points on the extracted images; Calculate the optical flow value of each key point according to the coordinate displacement of the same key point between two adjacent frames, and then calculate the associated optical flow value of each key point according to the optical flow value of the key point and the optical flow values of other key points in the region to which the key point belongs, including calculating the weighted sum of the optical flow value of the key point and the optical flow values of other key points in the region to which the key point belongs as the associated optical flow value of the key point; Construct a sparse associated optical flow feature map according to the associated optical flow values of each key point on multiple frames of images, including, recording the number of key points as N, if key points are detected in N + 1 consecutive frames of images, construct a sparse associated optical flow feature map according to the associated optical flow values of the key points in the N + 1 frames of images; Input the sparse associated optical flow feature map into the trained detection model and output the detection result.

2. The method for detecting forged videos according to claim 1, wherein, The extracted image is a face image, and the multiple pre - defined key points are all facial key points, which are divided into twelve regions: left eye, right eye, left eyelid, right eyelid, left eyebrow, right eyebrow, left cheek, right cheek, upper lip, lower lip, nose and head.

3. The method for detecting forged videos according to claim 1, wherein The detecting key points on the extracted image includes the steps: Each key point is predetermined to have a unique index number, and the coordinate value of the detected key point is recorded as (x i,j ,y i,j ), where i represents the index number of the key point and j represents the sequence number of the frame. A key point detection file is generated to record the coordinate values ​​of all key points in each frame. If no key point is detected, no key point detection file is generated.

4. The method for detecting forged videos according to claim 3, characterized in that, If some key points cannot be detected in the extracted image, define the undetected key points as missing points, and then represent the coordinate values of the missing points with pre - defined special values in the key point detection file.

5. The method for detecting forged videos according to claim 2, wherein The sparse associated optical flow feature map contains features of inconsistent facial expressions.

6. The method for detecting forged videos according to claim 1, characterized in that, The detection model is implemented based on a convolutional neural network, including 6 convolutional layers, 4 max - pooling layers and 3 fully - connected layers. For each convolutional layer, the size of the convolutional kernel is 3×3, and the stride of the last convolutional layer is 2.

7. A forged video detection system, characterized in that, Including: A key point definition module for pre - defining multiple key points and dividing the multiple key points into multiple regions; A detection module for extracting images from the video to be detected and detecting key points on each frame of the extracted images; A feature extraction module for calculating the optical flow value of each key point according to the coordinate displacement of the same key point between two adjacent frames, and then calculating the associated optical flow value of each key point according to the optical flow value of the key point and the optical flow values of other key points in the region to which the key point belongs, where including calculating the weighted sum of the optical flow value of the key point and the optical flow values of other key points in the region to which the key point belongs as the associated optical flow value of the key point; Constructing a sparse associated optical flow feature map according to the associated optical flow values of each key point on multiple frames of images, including, recording the number of key points as N, if key points are detected in N + 1 consecutive frames of images, constructing a sparse associated optical flow feature map according to the associated optical flow values of the key points in the N + 1 frames of images; A forgery detection module for inputting the sparse associated optical flow feature map into the trained detection model and outputting the detection result.

8. A storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Detection method for video target based on image significance and characteristic prior model

    CN108122247A

  • Face counterfeit video detection method based on spatial local binary pattern and optical flow gradient

    CN111797702A