False video detection method and false video detection device using same
By extracting video segments on the computing device, performing feature extraction and time-focusing model processing, and generating enhanced feature maps to judge the authenticity of the video, the problem of difficult identification of false videos in the prior art is solved, and efficient and accurate false video detection is achieved.
Patent Information
- Application Number
- CN202311716108.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively identify false videos in deep forgery technology detection, resulting in the widespread dissemination of malicious or misinformation.
The video segments in the target video are extracted by the computing device, a feature extraction program is executed, a feature map is obtained, and the time focus model is input to process it. The enhanced feature map is generated through multiplication and addition operations, and finally the full connection layer is input to judge the authenticity of the video.
It realizes efficient and accurate identification of fake videos, avoids being deceived by fake videos, and improves the accuracy and generality of detection.
Smart Images

Figure CN120107841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image detection method, and in particular to a false video detection method and a false video detection device using the method. Background Art
[0002] Deep fake detection plays an extremely critical role in the digital age. Its industrial value not only involves public safety and social stability, but also personal privacy protection and data credibility. Therefore, improving the accuracy of deep fake detection is crucial to curb the widespread dissemination of malicious or false information. Summary of the invention
[0003] An embodiment of the present invention provides a false video detection method, which is performed via a computing device, wherein the computing device includes a processor. The method includes: extracting N video segments from a target video via the processor, wherein each video segment includes M image frames; performing a feature extraction procedure on the NxM image frames to obtain a feature map corresponding to the NxM image frames; inputting the feature map into a temporal attention model to obtain a concentrated feature; multiplying the concentrated feature with the feature map to obtain a first enhanced feature map; adding the first enhanced feature map to the feature map to obtain a second enhanced feature map; and inputting the second enhanced feature map into a fully connected layer to obtain a judgment result output by the fully connected layer, wherein the judgment result indicates whether the target video is true or false.
[0004] In an embodiment of the present invention, the step of executing the feature extraction procedure on the NxM image frames to obtain the feature maps corresponding to the NxM image frames includes: performing a data preprocessing operation on the NxM image frames to obtain a first image sequence; performing a first reshaping operation on the first image sequence to obtain a second image sequence; inputting the second image sequence into a pre-trained neural network model to obtain a first feature sequence; performing a second reshaping operation on the first feature sequence to obtain a second feature sequence; and inputting the second feature sequence into an average pooling layer to obtain the feature maps corresponding to the NxM image frames.
[0005] In an embodiment of the present invention, the data preprocessing operation includes: performing a normalization operation on each of the NxM image frames so that the value of each pixel of each image frame is between 0 and 1; and performing a data expansion operation on each normalized image frame in the normalized NxM image frames to obtain the first image sequence.
[0006] In an embodiment of the present invention, the first reshaping operation includes performing a dimensionality reduction operation on the first image sequence to obtain the second image sequence.
[0007] In an embodiment of the present invention, the pre-trained neural network model includes one of the following: ResNet model, EffNet model, VGG16 model, InceptionV3 model, Xception model, wherein the pre-trained neural network model performs feature extraction on the input second image sequence to obtain the first feature sequence.
[0008] In an embodiment of the present invention, the temporal attention model includes: a first branch, including a first average pooling layer, a first convolutional layer, a first function, and another first convolutional layer; a second branch, including a second maximum pooling layer, a second convolutional layer, a second function, and another second convolutional layer; and a Sigmoid function, wherein the feature map is input to the first average pooling layer and the second maximum pooling layer of the first branch and the second branch, respectively, and the output results of the first branch and the second branch are added and input to the Sigmoid function to output the concentrated features.
[0009] In an embodiment of the present invention, the first function and the second function include one of the following: a Leaky ReLU function, a PReLU function, a RReLU function, and a SELU function.
[0010] In an embodiment of the present invention, the judgment result determined to be true is used to indicate that the target video is a real video that does not contain any fake content, and the judgment result determined to be false is used to indicate that the target video is a fake video that contains fake content, wherein the fake video is a real video that only has native content before the fake content is added.
[0011] Another embodiment of the present invention provides a false video detection device, comprising a processor. The processor is used to: extract N video segments from a target video, wherein each video segment includes M image frames; perform a feature extraction procedure on the NxM image frames to obtain a feature map corresponding to the NxM image frames; input the feature map into a temporal attention model to obtain a concentrated feature; multiply the concentrated feature with the feature map to obtain a first enhanced feature map; add the first enhanced feature map to the feature map to obtain a second enhanced feature map; and input the second enhanced feature map into a fully connected layer to obtain a judgment result output by the fully connected layer, wherein the judgment result indicates whether the target video is true or false.
[0012] Based on the above, the present invention provides a false video detection method and false video detection device, which can extract the video segment of the target video part to perform feature extraction on the video segment, obtain the corresponding feature map, input the feature map into the time attention model to obtain the enhanced feature map, and obtain the judgment result according to the enhanced feature map to determine whether the target video is a real video or a false video. In this way, the time attention model is used to find out whether the target video has discontinuous forged content, thereby efficiently and accurately detecting possible false videos, thereby avoiding being deceived by false videos. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a block diagram of a computing device for executing a false video detection method according to an embodiment of the present invention.
[0014] Figure 2 is an operation flow chart of a false video detection method according to an embodiment of the present invention.
[0015] Figure 3 FIG. 4 is a schematic diagram of extracting N video segments from a target video according to an embodiment of the present invention.
[0016] Figure 4 is a schematic diagram of obtaining a feature map through a feature extraction procedure according to an embodiment of the present invention.
[0017] Figure 5 is a schematic diagram of a false video detection method according to an embodiment of the present invention.
[0018] Figure 6 FIG. 4 is a schematic diagram of the operation of a time attention module according to an embodiment of the present invention.
[0019] Figure 7 is a schematic diagram illustrating the accuracy of different detection methods according to an embodiment of the present invention.
[0020] Figure 8 FIG. 4 is a schematic diagram illustrating the accuracy of different detection methods corresponding to different image interferences according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0022] Reference will now be made in detail to exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.
[0023] Please refer to Figure 1In this embodiment, the computing device 100 (also referred to as a false video detection device) includes a processor 110 , a communication circuit unit 120 , a storage circuit unit 130 , an input / output unit, and a memory 150 .
[0024] The processor 110 is, for example, a microprogrammed control unit (MCU), a central processing unit (CPU), a graphics processing unit (GPU), a programmable microprocessor (Microprocessor), an application specific integrated circuit (ASIC), a programmable logic device (PLD) or other similar devices.
[0025] The communication circuit unit 120 is used to transmit or receive data through wired or wireless communication. In this embodiment, the communication circuit unit may have a wireless communication circuit module (not shown) and support one or a combination of the Global System for Mobile Communication (GSM) system, Wireless Fidelity (WiFi) system, and Bluetooth communication technology, but is not limited thereto. The processor 110 may receive the target video from the Internet or other electronic devices via the communication circuit unit 120.
[0026] The input / output unit 130 includes an input device and an output device. The input device is, for example, a microphone, a touch pad, a touch panel, a keyboard, a mouse, etc., which is used to allow the user to input data or control the functions that the user wants to operate. The output device is, for example, a display (which is used to receive data of the display screen to display an image), a speaker (which can be used to receive audio data to emit sound effects), etc., but the present case is not limited to this. In an embodiment, the input / output unit 130 may include a touch screen. The touch screen is used to display various information and control interfaces of the projector. The input / output unit 130 can display the target video or the judgment result.
[0027] The storage circuit unit 140 can store data via instructions from the processor 110. The storage circuit unit includes any type of hard disk drive (HDD) or non-volatile memory storage device (such as SSD or flash memory). The storage circuit unit can store the target video and various algorithm models provided in this embodiment.
[0028] The memory 150 is used to temporarily store instructions or data executed by the processor 110, such as a dynamic random access memory (DRAM), a static random access memory (SRAM), etc. In this embodiment, the processor 110 can temporarily store various executed models, deduction data, feature data, video data, and data sequences in the memory 150.
[0029] Please refer to Figure 2 In step S210, N video segments in the target video are extracted, wherein each video segment includes M image frames.
[0030] For example, see Figure 3 , the processor 110 may randomly select N video segments VS(1) to VS(N) in the target video TV. Each of these video segments has the same number of image frames, which is M. That is, the processor 110 extracts a total of NxM image frames VS. N is, for example, 8 or other positive integers, and M is, for example, 16 or other positive integers greater than 1.
[0031] Please go back Figure 2 In step S220, the processor 110 executes a feature extraction procedure on the NxM image frames to obtain feature maps corresponding to the NxM image frames.
[0032] In detail, the feature extraction procedure includes: performing a data preprocessing operation on NxM image frames to obtain a first image sequence; performing a first reshaping operation on the first image sequence to obtain a second image sequence; inputting the second image sequence into a pre-trained neural network model to obtain a first feature sequence; performing a second reshaping operation on the first feature sequence to obtain a second feature sequence; and inputting the second feature sequence into an average pooling layer to obtain a feature map corresponding to the NxM image frames.
[0033] Please refer to Figure 4 The processor 110 performs a data preprocessing operation (A401) on the NxM image frames VS to obtain a first image sequence PS1. The first image sequence PS1 can be represented by a vector, such as a vector dimension size of [N, M, C, H, W]. Wherein N is the number of video segments, M is the number of image frames of each video segment, C is the number of color channels (e.g., pixel information of three primary colors, the number of color channels is 3), H is the height of each image frame (the number of pixels in the first direction), and W is the width of each image frame (the number of pixels in the second direction).
[0034] In this embodiment, the data preprocessing operation includes: performing a normalization operation on each of the NxM image frames so that the value of each pixel of each image frame is between 0 and 1; and performing a data expansion operation on each of the normalized NxM image frames to obtain a first image sequence.
[0035] For example, the target video is extracted in a manner of 8 video segments (batch=8) and a sequence length of 16 (each video segment has 16 image frames). Each of the NxM image frames of the 300×300 three-channel color image (300×300×3) obtained from the target video is normalized so that the value of each pixel is between 0 and 1, and a data expansion operation (including Gaussian blur, vertical flip, random brightness contrast, translation, scaling, rotation, and filling) is performed. Finally, the NxM image frames VS, for example, can form a first image sequence PS1[8, 16, 3, 300, 300]. Among them, N is 8, M is 16, C is 3, H is 300, and W is 300.
[0036] Next, the processor 110 performs a first reshaping operation (A402) on the first image sequence PS1 to obtain a second image sequence PS2. The first reshaping operation includes performing a dimensionality reduction operation on the first image sequence to obtain a second image sequence.
[0037] The second image sequence PS2 can be represented by a vector, for example, the vector dimension size is [N*M, C, H, W]. It should be noted that the second image sequence PS2 is reduced by one dimension compared to the first image sequence PS1 by the first reshaping operation. The vector dimension size of the second image sequence PS2 is, for example, [8*16, 3, 300, 300].
[0038] Next, the processor 110 inputs the second image sequence PS2 into the pre-trained neural network model NM1 (A403) to obtain a first feature sequence FS1 (A404). The first feature sequence FS1 can be represented by a vector, such as a vector dimension of [N*M, D, H', W']. In this embodiment, D is, for example, the number of channels of a feature map; H' is, for example, the height of a feature map; and W' is, for example, the width of a feature map.
[0039] In this embodiment, the pre-trained neural network model NM1 includes one of the following: ResNet model, EffNet model, VGG16 model, InceptionV3 model, Xception model, wherein the pre-trained neural network model performs feature extraction on the input second image sequence to obtain a first feature sequence FS1. Assume that the pre-trained neural network model is EffNet-b7. The second image sequence PS2 [8*16, 3, 300, 300] is input to the pre-trained neural network model EffNet-b7, and finally outputs the first feature sequence FS1. The vector dimension size of the first feature sequence FS1 is, for example, [8*16, 2560, 10, 10]. Wherein, N is 8, M is 16, D is 2560, H' is 10, and W' is 10. Wherein, 2560 represents the number of channels of the feature map, and (10, 10) represents the height and width of the feature map. In other words, after being processed by the EffNet-b7 model, each input image is converted into a feature map of size 2560x10x10. Such a feature map usually contains high-level features extracted from the original image.
[0040] Next, the processor 110 performs a second reshaping operation (A405) on the first feature sequence FS1 to obtain a second feature sequence FS2. The second feature sequence FS2 can be represented by a vector, such as a vector dimension size of [N, M, E, F]. The vector dimension size of the second feature sequence FS2 is, for example, [8, 16, 320, 400]. In this embodiment, E is, for example, the height of the feature map; F is, for example, the width of the feature map.
[0041] Next, the processor 110 inputs the second feature sequence FS2 into the average pooling layer (Average Pooling 2D) AVP (A406) to obtain a feature map S (A407) corresponding to the NxM image frames VS. The parameters of the average pooling layer AVP are, for example: Kernel size = (20, 24); Stride = None. The second feature sequence FS2 can be further reduced in dimension through the average pooling layer AVP to reduce the amount of calculation, thereby reducing the sensitivity of the subsequent convolutional layer to position information. The feature map S can be represented by a vector, such as a vector dimension size of [N, M, G1, G2]. In this embodiment, G1 is, for example, the height of the feature map; G2 is, for example, the width of the feature map. In this example, the vector dimension size of the feature map S is, for example, [8, 16, 16, 16]. After the obtained feature map S passes through the temporal attention module, its result is element-wise multiplied with the feature map S (element-by-element multiplication), and finally added to the feature map S to obtain the final enhanced feature for determining whether the target video is a fake video.
[0042] Please come back Figure 2 In step S230, the processor 110 inputs the feature map S into the temporal attention model to obtain a concentrated feature. In step S240, the processor 110 multiplies the concentrated feature with the feature map S to obtain a first enhanced feature map ES1. In step S250, the processor 110 adds the first enhanced feature map ES1 to the feature map S to obtain a second enhanced feature map ES2. Finally, in step S260, the processor 110 inputs the second enhanced feature map ES2 into the fully connected layer to obtain a judgment result output by the fully connected layer, wherein the judgment result indicates whether the target video is true or false.
[0043] The judgment result determined as true is used to indicate that the target video is a real video without any fake content. The judgment result determined as false is used to indicate that the target video is a fake video containing fake content, wherein the fake video is a real video with only original content before the fake content is added. In other words, a real video refers to a video produced with actual scenes, real performances or real-life photography, presenting real scenes, characters and events, rather than videos processed with other special effects, such as deepfake technology.
[0044] The following use Figure 5 and Figure 6 The relevant details of steps S230 to S260 are described below.
[0045] Please refer to Figure 5 First, as described above, the processor 110 performs a feature extraction procedure on the extracted verification image frames VS (ie, NxM image frames) of the target video to obtain a feature map S (A500).
[0046] The processor 110 inputs the feature map S into the temporal attention model TA (A501 or S230) to obtain a condensed feature CF (also expressed as T(S)). The condensed feature CF can be expressed using the following formula (1):
[0047] T(S)=Sigmoid(Conv1×1(Avg.Pooling(S))+Conv1×1(Max.Pooling(S))) (1)
[0048] Next, the processor 110 multiplies the condensed feature CF with the feature map S (A502 or S240) to obtain ES1. The first enhanced feature map ES1 is, for example, T(S)xS. When performing multiplication here, T(S) will automatically copy the third and fourth dimensions. For example, the function torch.tile can be used to implement it, that is, torch.tile(T(S), [1, 1, 16, 16]). In more detail, in PyTorch, the torch.tile function can copy a tensor along a specified direction. This function is often used to expand the size of a tensor or copy its contents. Simply put, this function is used to repeat the contents of a tensor in a specific direction.
[0049] Next, the processor 110 performs an addition operation on the first enhanced feature map ES1 and the feature map S (A503 or S250) to obtain a second enhanced feature map ES2. The second enhanced feature map ES2 is, for example, (T(S)xS+S).
[0050] Finally, in step S260, the processor 110 inputs the second enhanced feature map ES2 to the flattening layer FTL (A504), and then inputs the output result to the fully connected layer FCL (A505) to obtain the judgment result (A506) output by the fully connected layer FCL. That is to say, steps A505 to A506 are: using a flattening layer, the input second enhanced feature map ES2 is expanded into [8*16*16*16, 1], and the fully connected layer (1 neuron) is connected in series. Among them, the time attention model TA and the fully connected layer FCL are trained by using the cross entropy (which is used for binary classification) function as the loss function (loSS function). In more detail, the output of the time attention model TA passes through the fully connected layer FCL, and a linear combination of weights is performed. The output of the fully connected layer FCL usually includes an activation function, such as a Sigmoid function or a Softmax function, to generate the final classification result. The loss function is used to measure the difference between the model prediction and the actual label. It is the objective function in the training process. For binary classification, the binary cross-entropy loss function is usually used. This loss function is common when measuring the distance between two probability distributions, and is particularly suitable for binary classification problems. The training process of the model includes forward propagation to calculate the prediction, and then using the cross entropy loss function to calculate the difference between the prediction and the actual label. Next, the weights of the model are updated through backpropagation and optimization algorithms to minimize the loss. This process is repeated until the model converges.
[0051] In another embodiment, the processor 110 inputs the feature map S into the enhanced temporal attention model to obtain an enhanced condensed feature or inputs it into the temporal attention model twice to obtain an enhanced condensed feature. The enhanced temporal attention model is twice the enhanced temporal attention model TA. For example, the enhanced condensed feature is T(S)2. In this embodiment, the first enhanced feature map ES1 obtained subsequently is, for example, T(S)2xS, and the second enhanced feature map ES2 is, for example, (T(S)2xS+S).
[0052] For more details, please refer to Figure 6, the temporal attention model TA includes: a first branch BH1, a second branch BH2 and a Sigmoid function SM. The first branch BH1 includes a first average pooling layer AVP1, a first convolutional layer CNN11, a first function, and another first convolutional layer CNN12. The second branch BH2 includes a second maximum pooling layer MXP2, a second convolutional layer CNN21, a second function, and another second convolutional layer CNN22.
[0053] Via the processor 110 , the feature map S is input to the first average pooling layer AVP1 and the second maximum pooling layer MXP2 ( A601 ) of the first branch BH1 and the second branch BH2 , respectively.
[0054] Next, the output results of the first average pooling layer AVP1 and the second maximum pooling layer MXP2 are input to the first convolutional layer CNN11 and the second convolutional layer CNN21 (A602).
[0055] For example, the first average pooling layer AVP1 (whose parameters are, for example, kernel size = 1×1, strides = (16, 16)) downsamples the feature map S, reduces parameters, removes redundant information, and retains the overall data features, and then inputs it to the first convolutional layer CNN11 (kernel size = 1×1, strides = 1). After being processed by the first function, the processed result is input to another first convolutional layer CNN12 (kernel size = 1×1, strides = 1) (A603) to obtain the output result of the first branch BH1, for example, [8, 16, l, 1].
[0056] For another example, the second maximum pooling layer MXP2 (pooling size = 1×1, strides = (16, 16)) downsamples the feature map S, reduces parameters, removes redundant information, and retains important data features, and then inputs it to the second convolutional layer CNN21 (kernel size = 1×1, strides = 1). After being processed by the second function, the processed result is input to another second convolutional layer CNN22 (kernel size = 1×1, strides = 1) (A603) to obtain the output result of the first branch BH2, whose vector dimension size is, for example, [8, 16, 1, 1].
[0057] Finally, the output results of the first branch BH1 and the second branch BH2 are added (A604), and the added result is input to the Sigmoid function SM to output the condensed feature CF (A605). The vector dimension size of the condensed feature CF is, for example, [8, 16, 1, 1].
[0058] Simply put, the temporal attention model performs three steps. Among them, the first step: use the average pooling layer and the maximum pooling layer to reduce the height and width to 1, perform dimensionality reduction processing on the features, and extract key features. The second step: use two 1×1 convolutions and adopt the first and second functions. The third step: after the final addition, pass the sigmoid function SM to obtain the concentrated feature CF. That is, by passing the feature map S through the maximum pooling layer, the average pooling layer, and the convolution layer respectively, the outputs of the two branches are added, and then the sigmoid function is passed to calculate the attention map based on the time dimension (concentrated feature CF), the purpose of which is to further capture the temporal discontinuity features in deep fake technology videos.
[0059] In addition, using two branches corresponding to the average pooling layer and the maximum pooling layer, key features can be extracted and the number of model parameters can be greatly reduced. This is especially important for large models, because the more parameters a model has, the more computing resources and memory space it requires. By reducing the parameters of the model, the detection model can be processed and trained more efficiently. On the other hand, the purpose of using two 1×1 convolutional layers in each branch is that the first convolutional layer is used to reduce the number of dimensions, and the second convolutional layer is used to increase the number of dimensions. In this way, the complexity of the entire model can be effectively controlled.
[0060] It should be noted that, in this embodiment, the first function and the second function include one of the following: Leaky ReLU function, PReLU function, RReLU function and SELU function. The first function and the second function are used to solve the "dead neuron" (DeadReLUProblem) problem. The traditional ReLU function sets the gradient to 0 in the negative part, which may cause the neuron to be unable to be updated in the subsequent training process, which is called the "dead neuron" problem. The functions listed above can solve the "dead neuron" problem.
[0061] It is worth mentioning that, compared with using only one fully connected (FC) layer, the temporal attention model provided by the embodiment of the present invention can reduce the number of model parameters used to one eighth of the original. In addition, the temporal attention model calculates concentrated features based on the time dimension, the purpose of which is to further capture the temporal discontinuity in the forged fake video, reduce the number of model parameters, and improve the effectiveness and training efficiency of the detection model.
[0062] The following use Figure 7 , Figure 8 To illustrate the accuracy of the false video detection method provided by the present invention.
[0063] Please refer to Figure 7As shown in Table T71, the traditional detection method uses EffNet-b7 as a pre-trained model to obtain feature maps, and uses the LSTM algorithm to judge the feature maps for fake videos; the detection method provided in this embodiment uses EffNet-b7 as a pre-trained model to obtain feature maps, and uses the twice-time attention model TA2 to judge the feature maps for fake videos. If the training data set is used for testing, the judgment accuracy of the traditional method and the present method are 97% and 98.4% respectively. The model is trained using the FaceForensics++ (FF++) training data set consisting of 1,000 real videos and 4,000 fake videos. The fake videos in the training data set are generated by four different deepfake technologies: Deepfake (DF) 1, FaceSwap (FS) 2, Face2Face (F2F) and NeuralTexture (NT).
[0064] If another test data set Celeb-DF (Y. Li, X. Yang, P. Sun, H. Qi and S. Lyu, "Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics," 2020 IEEE / CVFConference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020, pp. 3204-3213, doi: 10.1109 / CVPR42600.2020.00327) is used for testing, the judgment accuracy of the traditional method and this method is 79.6% and 84.1% respectively; if the test data set DFDC (DeepFake Detection Challenge Dataset: Dolhansky, Brian, et al. "The deepfake detection challenge (dfdc) dataset." arXiv preprint arXiv:2006.07397(2020)) was used for testing. The accuracy of the traditional method and the present method were 73% and 81.2% respectively. The above test results show that the accuracy of the traditional method is high for trained data, but the accuracy for the untrained test data set will drop significantly. In other words, the detection method provided by this embodiment through the use of the temporal attention model is more versatile (better generalization ability).
[0065] On the other hand, as shown in Table T72, the accuracy recorded in Table T72 can reflect that when the same pre-trained model is used, the number of times the temporal attention model is used will affect the accuracy. It can be seen that the accuracy of the detection method (EffNet-b7+TA2) using the temporal attention model twice is greater than the accuracy of the detection method (EffNet-b7+TA) using the temporal attention model only once, and it also greatly exceeds the accuracy of the detection method (EffNet-b7) not using the temporal attention model.
[0066] Please refer to Figure 8 As shown in Table T81, the detection method provided in this embodiment is more tolerant to interference than the traditional detection method (Multiple-attention [Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu, "Multiattentional deepfake detection," in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2021, pp. 2185-2194], Face X-ray [Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, DongChen, Fang Wen, and Baining Guo, "Face x-ray for more general face forgery detection," in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2020, pp. 5001-5010]). That is, when the target video is interfered with in different ways, the average accuracy rate for judging whether the interfered target video is a fake video is the highest.
[0067] Based on the above, the present invention provides a false video detection method and false video detection device, which can extract the video segment of the target video part to perform feature extraction on the video segment, obtain the corresponding feature map, input the feature map into the time attention model to obtain the enhanced feature map, and obtain the judgment result according to the enhanced feature map to determine whether the target video is a real video or a false video. In this way, the time attention model is used to find out whether the target video has discontinuous forged content, thereby efficiently and accurately detecting possible false videos, thereby avoiding being deceived by false videos.
[0068] Although the present invention has been disclosed as above by way of embodiments, it is not intended to limit the present invention. Any person having ordinary knowledge in the technical field may make some changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by the scope of the attached patent application.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A false video detection method, performed by a computing device, the computing device comprising a processor, It is characterized in that The method comprises: via the processor to: Extracting N video segments from the target video, where each video segment contains M image frames; Execute a feature extraction procedure on the NxM image frames to obtain feature maps corresponding to the NxM image frames; Inputting the feature map into a temporal attention model to obtain condensed features; Multiplying the concentrated feature with the feature map to obtain a first enhanced feature map; Adding the first enhanced feature map to the feature map to obtain a second enhanced feature map; and The second enhanced feature map is input into a fully connected layer to obtain a judgment result output by the fully connected layer, wherein the judgment result indicates whether the target video is true or false.
2. The false video detection method according to claim 1, It is characterized in that The step of executing the feature extraction procedure on the NxM image frames to obtain the feature maps corresponding to the NxM image frames comprises: Performing a data preprocessing operation on the NxM image frames to obtain a first image sequence; performing a first reshaping operation on the first image sequence to obtain a second image sequence; Inputting the second image sequence into a pre-trained neural network model to obtain a first feature sequence; performing a second reshaping operation on the first feature sequence to obtain a second feature sequence; and The second feature sequence is input into the average pooling layer to obtain the feature maps corresponding to the NxM image frames.
3. The false video detection method according to claim 2, It is characterized in that The data preprocessing operation includes: Performing a normalization operation on each of the NxM image frames so that a value of each pixel of each image frame is between 0 and 1; and A data expansion operation is performed on each of the normalized NxM image frames to obtain the first image sequence.
4. The false video detection method according to claim 2, It is characterized in that The first reshaping operation includes performing a dimensionality reduction operation on the first image sequence to obtain the second image sequence.
5. The false video detection method according to claim 2, It is characterized in that The pre-trained neural network model includes one of the following: ResNet model, EffNet model, VGG16 model, InceptionV3 model, Xception model, wherein the pre-trained neural network model performs feature extraction on the input second image sequence to obtain the first feature sequence.
6. The false video detection method according to claim 1, It is characterized in that The time attention model includes: A first branch includes a first average pooling layer, a first convolutional layer, a first function, and another first convolutional layer; A second branch includes a second maximum pooling layer, a second convolutional layer, a second function, and another second convolutional layer; and Sigmoid function, wherein the feature map is input to the first average pooling layer and the second maximum pooling layer of the first branch and the second branch respectively, and the output results of the first branch and the second branch are added and input to the Sigmoid function to output the concentrated features.
7. The false video detection method according to claim 6, It is characterized in that The first function and the second function include one of the following: a Leaky ReLU function, a PReLU function, a RReLU function, and a SELU function.
8. The false video detection method according to claim 1, It is characterized in that in The judgment result determined to be true is used to indicate that the target video is a real video that does not contain any fake content, and The determination result determined to be false is used to indicate that the target video is a false video containing fake content, wherein the false video is the real video having only native content before the fake content is added.
9. A false video detection device, It is characterized in that include: A processor, wherein the processor is configured to: Extracting N video segments from the target video, where each video segment contains M image frames; Execute a feature extraction procedure on the NxM image frames to obtain feature maps corresponding to the NxM image frames; Inputting the feature map into a temporal attention model to obtain condensed features; Multiplying the concentrated feature with the feature map to obtain a first enhanced feature map; Adding the first enhanced feature map to the feature map to obtain a second enhanced feature map; as well as The second enhanced feature map is input into a fully connected layer to obtain a judgment result output by the fully connected layer, wherein the judgment result indicates whether the target video is true or false.
10. The false video detection device according to claim 9, It is characterized in that The step of executing the feature extraction procedure on the NxM image frames to obtain the feature maps corresponding to the NxM image frames comprises: Performing a data preprocessing operation on the NxM image frames to obtain a first image sequence; performing a first reshaping operation on the first image sequence to obtain a second image sequence; Inputting the second image sequence into a pre-trained neural network model to obtain a first feature sequence; performing a second reshaping operation on the first feature sequence to obtain a second feature sequence; and The second feature sequence is input into the average pooling layer to obtain the feature maps corresponding to the NxM image frames.
11. The false video detection device according to claim 10, It is characterized in that The data preprocessing operation includes: Performing a normalization operation on each of the NxM image frames so that a value of each pixel of each image frame is between 0 and 1; and A data expansion operation is performed on each of the normalized NxM image frames to obtain the first image sequence.
12. The false video detection device according to claim 10, It is characterized in that The first reshaping operation includes performing a dimensionality reduction operation on the first image sequence to obtain the second image sequence.
13. The false video detection device according to claim 10, It is characterized in that The pre-trained neural network model includes one of the following: ResNet model, EffNet model, VGG16 model, InceptionV3 model, Xception model, wherein the pre-trained neural network model performs feature extraction on the input second image sequence to obtain the first feature sequence.
14. The false video detection device according to claim 9, It is characterized in that The time attention model includes: A first branch includes a first average pooling layer, a first convolutional layer, a first function, and another first convolutional layer; A second branch includes a second maximum pooling layer, a second convolutional layer, a second function, and another second convolutional layer; and Sigmoid function, wherein the feature map is input to the first average pooling layer and the second maximum pooling layer of the first branch and the second branch respectively, and the output results of the first branch and the second branch are added and input to the Sigmoid function to output the concentrated features.
15. The false video detection device according to claim 14, It is characterized in that The first function and the second function include one of the following: a LeakyReLU function, a PReLU function, a RReLU function, and a SELU function.
16. The false video detection device according to claim 9, It is characterized in that in The judgment result determined to be true is used to indicate that the target video is a real video that does not contain any fake content, and The determination result determined to be false is used to indicate that the target video is a false video containing fake content, wherein the false video is the real video having only native content before the fake content is added.