Video face forgery detection method and system based on time-space domain collaborative information bottleneck
By adopting a detection method based on the time-space domain collaborative information bottleneck in video face forgery detection, the problem of insufficient detection accuracy and generalization in the prior art is solved, and effective detection of multiple forgery modes is achieved, and detection accuracy and generalization are improved.
Patent Information
- Application Number
- CN202510216898.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-06
AI Technical Summary
Existing video face forgery detection technology is difficult to effectively deal with multiple random combinations of forgery modes, timing forgery, and high-quality and low-quality video forgery, resulting in insufficient detection accuracy and generalization.
The video face forgery detection method based on the time-space-coordinated information bottleneck is adopted. By training the video face forgery detection network model based on the time-space-coordinated information bottleneck, the detection accuracy and generalization are improved by using the time-space-coordinated information bottleneck technology.
It improves the accuracy and generalization of video face forgery detection, and can effectively deal with various forgery modes such as identity modification, expression replay, and local content tampering, reduces the rate of missed and false detection, and improves the accuracy of high-quality and low-quality video forgery detection.
Smart Images

Figure CN120108049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision, pattern recognition, machine learning, and the like, and in particular to a video face forgery detection method and system based on spatiotemporal domain collaborative information bottleneck. Background Art
[0002] Video forgery detection technology needs to detect many unknown and complex image forgery patterns. People are the main content of most videos, and there are forgeries such as face swapping and expression replay. Traditional video forgery detection technology mainly targets face images in specific scenarios and specific forgery methods.
[0003] The main difficulties of video face forgery detection technology are: 1. Forged face images may contain multiple random combinations of forgery patterns such as identity modification, expression replay, and local content tampering. 2. Forged videos may be forged in multiple time sequences, and the forgery methods for each time sequence are different. 3. The forgery effects of forgery methods are becoming more and more realistic, which increases the difficulty of detecting forged faces in high-quality videos. 4. In terms of low-quality video forgery, because of the interference of multiple degradation modes of the image, the feature differences between unforged and forged videos are more difficult to be detected by machines, which limits the generalization and accuracy of video face forgery detection models.
[0004] Information bottleneck has achieved remarkable results in many fields of computer vision, especially in task-related feature compression technology. There is no method and system for video face forgery detection using information bottleneck in the prior art. Summary of the invention
[0005] In order to solve the above technical problems, the present invention proposes a video face forgery detection method based on the information bottleneck compression technology to optimize the spatiotemporal domain features and improve the accuracy and generalization of video forgery detection. The method is implemented in accordance with the following steps: A video face forgery detection method based on the spatiotemporal domain collaborative information bottleneck comprises the steps of:
[0006] Step S1, collect real videos and fake videos, and divide the data into a training set and a test set;
[0007] Step S2, training a video face forgery detection network model based on spatiotemporal collaborative information bottleneck to determine whether an input video is forged;
[0008] Step S3: Use the trained video face forgery detection network model based on spatiotemporal collaborative information bottleneck to predict the forgery identification result from the video face image to be identified.
[0009] The present invention also proposes a video face forgery detection system based on spatiotemporal collaborative information bottleneck, comprising:
[0010] The collection module is used to collect real videos and fake videos and divide the data into training sets and test sets;
[0011] A training module, used for training a video face forgery detection network model based on a spatiotemporal collaborative information bottleneck that can determine whether an input video is forged;
[0012] The forgery result identification module predicts the forgery identification result from the video face image to be identified based on the video face forgery detection network model.
[0013] The present invention also proposes an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the above method is performed.
[0014] The present invention also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the above method.
[0015] Beneficial effects of the present invention:
[0016] The above method of the present invention can determine whether a real-life video is forged by a machine through a video face forgery detection network based on a collaborative information bottleneck in the spatiotemporal domain.
[0017] 1. Through information optimization based on information bottlenecks, the generalization of the model is improved, and it can cope with a variety of random combination forgery modes such as identity modification, expression replay, and local content tampering;
[0018] 2. Improve the accuracy of time series forgery detection and reduce missed detection and false detection rates by solving information bottlenecks in the time and space domains;
[0019] 3. Deeply mine facial image features through spatial and temporal collaborative information bottlenecks to effectively handle high-quality facial forgery detection;
[0020] 4. Use information bottleneck technology to effectively filter out noise interference in low-quality images and improve the accuracy of face forgery detection in low-quality videos. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flow chart of the video face forgery detection method based on spatiotemporal domain collaborative information bottleneck in the present invention. DETAILED DESCRIPTION
[0022] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0023] In order to solve the problem of video forgery detection, the present invention proposes a video face forgery detection method and system based on spatiotemporal collaborative information bottleneck, which improves the forgery detection accuracy and generalization of the video forgery detection method in dealing with unknown forgery methods in different scenarios.
[0024] Figure 1 The video face forgery detection method and system flow chart based on spatiotemporal collaborative information bottleneck proposed by the present invention are shown as follows: Figure 1 The method shown includes the following steps:
[0025] Step S1: Collect real videos and fake videos, divide the data into training sets and test sets; prepare a neural network model for feature compression , and the operator for mutual information calculation The neural network model M contains a convolutional layer and a sigmoid activation layer. Operator is the KL divergence (relative entropy) operator. The neural network model M includes a neural network model for time domain feature compression in step S2. Neural network model for spatial feature compression . Mutual information calculation operator In the loss function and used in. Specifically:
[0026] Step S11: extract frames from the captured video.
[0027] Step S12: Perform face detection, cropping and alignment on the image in step S11.
[0028] Step S2: Using the sample images in the training data set, train a video face forgery detection network model based on the spatiotemporal collaborative information bottleneck to determine whether the video is forged. Input the continuous face image frames into the 3D convolution model (including three feature dimensions: horizontal space, vertical space and time series) to extract time domain features and obtain time domain correlation features , the corresponding parameter update algorithm is ,in is the learning rate, is the image data, is the corresponding label, is the time domain correlation feature The network parameters, Spatial correlation features The network parameters are calculated; the time domain correlation features are compressed using the time domain information bottleneck to filter out the time domain noise information that interferes with the video face forgery detection, that is, a neural network consisting of a convolutional layer and a sigmoid activation layer is constructed. , the sigmoid activation layer is used to obtain the feature attention map ,then , The larger the median value, the more local features are retained, and the smaller the value, the more local features are discarded; the compressed time domain correlation features are input into the three-dimensional convolution model Extract spatial features in the spatial domain and obtain spatial correlation features , the corresponding parameter update algorithm is ,in is the learning rate, is the image data, is the corresponding label, is the time domain correlation feature The network parameters, Spatial correlation features The network parameters are calculated; the spatial correlation features are compressed using the spatial information bottleneck to obtain compressed spatial correlation features, thereby filtering out the spatial noise information that interferes with video face forgery detection, that is, constructing a neural network consisting of a convolutional layer and a sigmoid activation layer. , the sigmoid activation layer is used to obtain the feature attention map ,then , The larger the median value, the more local features are retained, and the smaller the value, the more local features are discarded. Finally, through the classification network (including an average pooling layer and an FC (fully connected) layer), the video face forgery detection result is obtained, and the classification loss is the BCE (binary cross entropy) loss. The model training calculates the loss through the loss function described below.
[0029] Specifically, the objective function of the video face forgery detection network based on the time-space domain collaborative information bottleneck includes three parts, namely: the first part is the loss function of compressed perception between compressed time-domain correlation features and time-domain correlation features; the second part is the loss function of compressed perception between compressed spatial domain correlation features and spatial domain correlation features; the third part is the binary cross entropy loss function between the real video face classification label and the predicted video face classification label. In the present invention, the objective function shown below can be constructed:
[0030] ,
[0031] in, Label real video faces and predicted video face classification labels The classification loss function between is the time domain correlation feature Correlation characteristics with compressed time domain The loss function of compressed sensing between is the weight of the loss. Spatial correlation features Correlation features with compressed airspace The loss function of compressed sensing between is the weight of the loss. In the loss function, and The role of is to reduce the mutual information between the temporal and spatial domain correlation features and the compressed temporal and spatial domain correlation features. The role of is to increase the mutual information between the compressed spatiotemporal correlation features and the forgery detection classification results. This loss function enables the model to retain important information related to forgery identification while reducing data interference. and They are information bottleneck networks for information compression in time domain and space domain respectively. is the mutual information calculation operator of the feature (i.e., the KL divergence operator, which is used to measure the distance between two data or feature distributions), is the data batch size.
[0032] Step S3, using the trained video face forgery detection network model based on spatiotemporal collaborative information bottleneck, predicting the forgery identification result from the video face image to be identified.
[0033] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, applications, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A video face forgery detection method based on spatiotemporal collaborative information bottleneck, characterized in that: Includes steps: Step S1, collect real videos and fake videos, and divide the data into a training set and a test set; Step S2, training a video face forgery detection network model based on spatiotemporal collaborative information bottleneck to determine whether an input video is forged; Step S3: Use the trained video face forgery detection network model based on spatiotemporal collaborative information bottleneck to predict the forgery identification result from the video face image to be identified.
2. The method according to claim 1, characterized in that The step S1 comprises: Step S11: extract frames from the captured video; Step S12: Perform face detection, cropping and alignment on the image in step S11.
3. The method according to claim 1, characterized in that The step S2 includes: selecting continuous face image frames from the training data set, inputting them into a three-dimensional convolution model for time domain feature extraction, and obtaining time domain correlation features; compressing the time domain correlation features using a time domain information bottleneck to filter out time domain noise information that interferes with video face forgery detection; inputting the compressed time domain correlation features into the three-dimensional convolution model for spatial domain feature extraction to obtain spatial domain correlation features; compressing the spatial domain correlation features using a spatial domain information bottleneck to obtain compressed spatial domain correlation features, which are used to filter out spatial domain noise information that interferes with video face forgery detection; finally, obtaining a video face forgery detection result through a classification network; using the difference between the real video face classification label and the predicted video face classification label, adjusting the parameters in the video face forgery detection network model based on the time-space domain collaborative information bottleneck through gradient back propagation, iterating the above training process until convergence, and obtaining a trained video face forgery detection network model based on the time-space domain collaborative information bottleneck.
4. The method according to claim 3, characterized in that: 3D Convolutional Model It contains three feature dimensions: horizontal space, vertical space and time series. Time domain feature extraction obtains time domain correlation features. , the corresponding parameter update algorithm is ,in is the learning rate, is the image data, is the corresponding label, is the time domain correlation feature The network parameters, Spatial correlation features ; Build a neural network with a convolutional layer and a sigmoid activation layer , the sigmoid activation layer is used to obtain the feature attention map ,then .
5. The method according to claim 4, characterized in that Input the compressed time-domain correlation features into the 3D convolutional model Extract spatial features in the spatial domain and obtain spatial correlation features , the corresponding parameter update algorithm is ,in is the learning rate, is the image data, is the corresponding label, is the time domain correlation feature The network parameters, Spatial correlation features ; Build a neural network with a convolutional layer and a sigmoid activation layer , the sigmoid activation layer is used to obtain the feature attention map ,then .
6. The method according to claim 5, characterized in that The loss function used for gradient back propagation in step S2 includes three parts: the first part is the loss function of compressed perception between compressed time domain correlation features and time domain correlation features; the second part is the loss function of compressed perception between compressed spatial domain correlation features and spatial domain correlation features; the third part is the binary cross entropy loss function between the real video face classification label and the predicted video face classification label.
7. The method according to claim 6, characterized in that The objective function used to calculate the loss function value in step S2 is: , in, Label real video faces and predicted video face classification labels The classification loss function between is the time domain correlation feature Correlation characteristics with compressed time domain The loss function of compressed sensing between is the weight of the loss, Spatial correlation features Correlation features with compressed airspace The loss function of compressed sensing between is the weight of the loss, is the operator for calculating mutual information, operator is the KL divergence operator.
8. A video face forgery detection system based on spatiotemporal collaborative information bottleneck, characterized in that: include: The collection module is used to collect real videos and fake videos and divide the data into training sets and test sets; A training module, used for training a video face forgery detection network model based on a spatiotemporal collaborative information bottleneck that can determine whether an input video is forged; The forgery result identification module predicts the forgery identification result from the video face image to be identified based on the video face forgery detection network model.
9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Executable instructions are stored thereon, and when the instructions are executed by a processor, the processor implements the method according to any one of claims 1 to 7.