Method and system for detecting fake face video based on regionalized space-time Transform

By using the regionalized spatiotemporal Transformer method, the falsified face video detection is performed using the difference in spatiotemporal characteristics of different areas of the face, which solves the limitations of the global analysis strategy in the existing technology, and achieves more accurate and robust fake face video detection.

CN120472296APending Publication Date: 2025-08-12WUHAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510507761.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Most of the existing fake face video detection technologies adopt global feature analysis strategies, ignoring the spatial and temporal feature differences caused by the forgery process in different areas of the face, resulting in insufficient detection accuracy and robustness.

Method used

The regionalized space-time Transformer method is adopted to accurately detect fake face videos through face analysis and area division, area adaptive feature extraction, regionalized space-time Transformer module and feature fusion, and use the forgery process to generate spatial-time feature differences in different areas of the face.

Benefits of technology

It improves the accuracy and robustness of fake face video detection, overcomes the limitations of global analysis, and can detect fake face videos more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472296A_ABST
    Figure CN120472296A_ABST
Patent Text Reader

Abstract

The invention discloses a counterfeited face video detection method based on a regionalized space-time Transform. The method comprises the following steps: constructing a counterfeited face video data set containing real and counterfeited samples; the method comprises the following steps: constructing a fake face video detection network comprising a face analysis and region division module, a region adaptive feature extraction module, a regionalization space-time Transform module, an intra-region feature fusion module and an inter-region feature fusion module; training the detection network using a data set; and utilizing the trained network model to predict true and false labels of the input video sequence. According to the method, spatial-temporal feature differences generated in different areas of a face in a counterfeiting process are utilized, and key links such as feature extraction (adopting area adaptive convolution), spatial-temporal modeling (adopting spatial-temporal Transform with intra-area / inter-area attention) and feature fusion (adopting an intra-area / inter-area fusion mechanism) are penetrated through a regionalized processing strategy; the method aims to overcome the limitation of global analysis, and can detect the fake face video more accurately and robustly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of security technology, and in particular relates to a method for detecting forged face videos, and specifically to a method and system for detecting forged face videos based on a regionalized spatiotemporal Transformer. Background Art

[0002] Fake face video detection aims to determine the authenticity of facial videos. It is a binary classification task that is crucial for maintaining information security. Existing methods, particularly Transformer-based models, show potential in capturing spatiotemporal features, but most adopt a global perspective. However, the traces produced by the forgery process in different regions of the face (such as the T-zone and edges) vary significantly in spatiotemporal dimensions. Existing methods generally ignore this regional feature variability and fail to implement targeted modeling. This limits their detection accuracy and robustness against complex forgeries, making them difficult to meet the needs of practical applications.

[0003] With the development of deep learning in the field of computer vision, research on forged face video detection methods has made significant progress and achieved remarkable results. In 2023, Zhao et al. proposed a method for detecting forged face videos based on a factorized spatiotemporal self-attention mechanism in the IEEE Transactions on Information Forensics and Security. In 2023, Xu et al. proposed a method for effectively encoding the spatiotemporal characteristics of video clips and retaining them in a static thumbnail at the IEEE International Conference on Computer Vision, using the Swin-Transformer to detect forged face videos. In 2024, Ba et al. proposed a method for detecting forged face videos (UMF) that combines multi-region feature extraction with information-theoretic constraints at the AAAI Conference on Artificial Intelligence.

[0004] Existing forged face video detection techniques, particularly Transformer-based methods, have made some progress in capturing spatiotemporal traces of deepfakes. However, most employ global feature analysis strategies, ignoring the spatiotemporal differences in facial features across different regions during the forgery process. Exploiting and leveraging these regional differences is crucial for optimizing existing detection methods. Therefore, developing detection methods that can effectively identify and leverage regional spatiotemporal specificity is crucial for improving the reliability and defense capabilities of deepfake detection. Summary of the Invention

[0005] In order to solve the problem that most existing forged face video detection technologies adopt global feature analysis strategies and ignore the differences in spatiotemporal features generated by the forgery process in different regions of the face, the present invention provides a forged face video detection method based on a regionalized spatiotemporal Transformer. It utilizes the differences in spatiotemporal features generated by the forgery process in different regions of the face, and uses a regionalized processing strategy throughout key links such as feature extraction, spatiotemporal modeling and feature fusion, aiming to overcome the limitations of global analysis and detect forged face videos more accurately and robustly.

[0006] According to one aspect of the present invention, a method for detecting fake face videos based on a regionalized spatiotemporal Transformer is provided, comprising: Obtain forged face video data to be detected; Inputting forged face video data to be detected into a trained forged face video detection network and outputting a detection result; wherein the training of the forged face video detection network includes: Build a fake face video dataset; A fake face video detection network was constructed, including: a face parsing and region division module, which is used to obtain the regional mask matrix corresponding to the three regions of the face image; a regional adaptive feature extraction module, which is used to generate the regional guidance matrix based on the regional mask matrix and use the regional guidance matrix to independently extract the features of the three regions; a regionalized spatiotemporal Transformer module, which is used to further extract the relationship between the three regions; an intra-regional feature fusion module, which integrates the internal spatiotemporal features of each region; and an inter-regional feature fusion module, which integrates the features of different regions to obtain the final decision vector. Based on the designed loss function, the network is trained on the constructed fake face video dataset, and the trained fake face video detection network is output.

[0007] As a further technical solution, when the face analysis and region division module obtains the region mask matrix corresponding to the three regions of the face image, it also includes: Input the video frame into BiSeNet to obtain the initial face mask; Use facial key point detection method to locate the T zone and generate a T zone mask; Perform dilation and erosion operations on the initial face mask to obtain the edge area mask; Combining the initial face mask, T-zone mask and edge region mask to determine other facial region masks and background region masks; Combining these masks generates a region mask matrix, where each element value represents the region class of the corresponding pixel.

[0008] As a further technical solution, the regional adaptive feature extraction module generates the regional guidance matrix in the following manner: For each layer of the feature extraction network, the mode of the pixel or feature region category in the receptive field of the convolution kernel of the current layer is calculated based on the region attribution information of the previous layer, the region attribution of each position on the feature map of the current layer is determined, and the region guidance matrix of the layer is generated.

[0009] As a further technical solution, the region adaptive feature extraction module performs the region adaptive convolution operation as follows: An independent convolution kernel group is set for each predefined face area. During the convolution calculation, based on the regional guidance matrix of the current layer, only the convolution kernel group corresponding to the specific area on the feature map is used for feature calculation.

[0010] As a further technical solution, the regionalized spatiotemporal Transformer module further extracts the three-region relationship, including: The input feature sequence is divided into subsequences of different face regions according to the regional guidance matrix, and the background region features are ignored; In the spatial dimension, we apply the intra-region attention mechanism to the subsequences and the inter-region attention mechanism to the sequences and combine the results to obtain spatial features. The spatial features are pooled, and the intra-region and inter-region attention mechanisms are applied to the pooled features in the temporal dimension. The results are combined to obtain the temporal features.

[0011] As a further technical solution, the intra-region feature fusion module includes the following steps when fusing the spatiotemporal features within each region: The spatial features and temporal features are weightedly fused using learnable weights to obtain fused spatiotemporal features.

[0012] As a further technical solution, the inter-region feature fusion module includes the following steps when fusing features of different regions: The fused spatiotemporal features are globally pooled to obtain a global context vector. The fused spatiotemporal features are pooled by region to obtain a regional decision vector. All regional decision vectors are concatenated and the fusion weight of each region is calculated. The weighted sum of the regional decision vectors is performed according to the fusion weight to obtain the fused regional decision vector. The fusion region decision vector is concatenated with the global context vector and the final decision vector is calculated.

[0013] As a further technical solution, the designed loss function is: , in is the cross entropy loss function of the facial T-zone, is the cross entropy loss function in the marginal area, is the cross entropy loss function for other facial regions, is the cross entropy loss function of the final prediction result, 、 、 Represents the loss function weight parameters of the three regions respectively. It is the weight parameter of the final prediction result loss function, which controls the balance between regional loss and final prediction loss.

[0014] According to one aspect of the present invention, a forged face video detection system based on a regionalized spatiotemporal Transformer is provided, comprising: A data acquisition module is used to acquire forged face video data to be detected; The data detection module is configured to input the forged face video data to be detected into the trained forged face video detection network and output a detection result. The training of the forged face video detection network includes: Build a fake face video dataset; A fake face video detection network was constructed, including: a face parsing and region division module, which is used to obtain the regional mask matrix corresponding to the three regions of the face image; a regional adaptive feature extraction module, which is used to generate the regional guidance matrix based on the regional mask matrix and use the regional guidance matrix to independently extract the features of the three regions; a regionalized spatiotemporal Transformer module, which is used to further extract the relationship between the three regions; an intra-regional feature fusion module, which integrates the internal spatiotemporal features of each region; and an inter-regional feature fusion module, which integrates the features of different regions to obtain the final decision vector. Based on the designed loss function, the network is trained on the constructed fake face video dataset, and the trained fake face video detection network is output.

[0015] According to one aspect of the present invention, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the steps of the method for detecting fake face videos based on a regionalized spatiotemporal Transformer.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This paper utilizes the differences in spatiotemporal features generated in different areas of the face during the forgery process. Through a regionalized processing strategy, it runs through key links such as feature extraction (using region-adaptive convolution), spatiotemporal modeling (using spatiotemporal Transformer with intra / inter-regional attention) and feature fusion (using intra / inter-regional fusion mechanism), aiming to overcome the limitations of global analysis and detect forged face videos more accurately and robustly. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, a brief introduction will be given below to the drawings used in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is a network diagram of a face parsing and region division module according to an embodiment of the present invention; Figure 2 is a network diagram of the regional attention module in the regionalized spatiotemporal Transformer module of an embodiment of the present invention; Figure 3 is a network diagram of the inter-region attention module in the regionalized spatiotemporal Transformer module of an embodiment of the present invention; Figure 4 This is a network structure diagram of the intra-region feature fusion module constructed in an embodiment of the present invention; Figure 5 This is a network diagram of an inter-region feature fusion module constructed in an embodiment of the present invention; Figure 6 This is a flowchart of the overall framework of the forged face video detection method based on the regionalized spatiotemporal Transformer constructed in an embodiment of the present invention. DETAILED DESCRIPTION

[0019] The terms "including" and "having" and any variations thereof in the description and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions, for example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to the steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.

[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In addition, the technical features in the various embodiments or single embodiments provided by the present invention are arbitrarily combined with each other to form a new technical solution. This combination is not restricted by the sequence of steps and / or structural composition mode, but must be based on the ability of ordinary technicians in this field to implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that this combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0021] Please see Figure 6 The embodiment of the present invention provides a method for detecting fake face videos based on a regionalized spatiotemporal Transformer, comprising the following steps: Step 1: Build a dataset.

[0022] Build a dataset of fake face videos, including both real and fake videos. Build the dataset required for training and testing fake face detection models.

[0023] To build the dataset required for the forged face video detection model, we first collected both real and forged videos. All videos in the training and test sets were processed identically, with four randomly sampled segments from each video, each containing eight consecutive frames of facial images. We then used a face detection tool to detect faces in each frame and cropped the detected facial regions to a resolution of 300×300. Finally, we applied an affine transformation to the cropped facial images to align key points and ensure input consistency.

[0024] Step 2: Build a fake face video detection network.

[0025] The fake face video detection network is trained using the training data set obtained in step 1 to obtain a model that can distinguish the authenticity of videos. Specifically, it includes a face parsing and region division module, a regional adaptive feature extraction module, a regionalized spatiotemporal Transformer module, an intra-regional feature fusion module, and an inter-regional feature fusion module. First, the face parsing module obtains the regional mask matrix corresponding to the three regions of the face image; secondly, the feature extraction module generates a regional guidance matrix and uses the guidance matrix to independently extract the features of the three regions; next, the regionalized spatiotemporal Transformer module further extracts the relationship between the three regions; then, the intra-regional feature fusion module fuses the internal spatiotemporal features of each region; finally, the inter-regional feature fusion module fuses the features of different regions; and finally, the final decision vector is obtained and input into the classifier to obtain the true or false label.

[0026] The face parsing and region segmentation module inputs the video frame into BiSeNet to obtain the initial face mask; uses the face key point detection method Dlib to locate the T area and generate the T area mask; performs dilation and erosion operations on the initial face mask to calculate the edge area mask; combines the initial face mask, T area mask and edge area mask to determine other facial area masks and background area masks; finally, combines these masks to generate the region mask matrix, where each element value represents the region category of the corresponding pixel, see Figure 1 .

[0027] The regional adaptive feature extraction module contains 5 convolution blocks and 4 pooling layers. Each convolution block contains a convolution layer, batch normalization (BN) and RELU activation function. It aims to convert the input image into a feature sequence containing regional information as the input token of the regionalized spatiotemporal Transformer module.

[0028] The regionalized spatiotemporal Transformer module divides the input feature sequence into subsequences of different facial regions (facial T-zone, edge region, and other facial regions) according to the regional guidance matrix, and ignores the background region features; in the spatial dimension, it applies the intra-region attention mechanism to the subsequences and the inter-region attention mechanism to the sequence, and combines the results to obtain spatial features. It performs a pooling operation on the spatial features; in the temporal dimension, it applies the intra-region and inter-region attention mechanisms to the pooled features and combines the results to obtain temporal features. For details, see Figure 2 and Figure 3 .

[0029] The feature fusion module in the region uses learnable weights to perform weighted fusion of spatial features and temporal features to obtain fused spatiotemporal features. Figure 4 .

[0030] The inter-region feature fusion module performs global pooling on the fused spatiotemporal features to obtain a global context vector, pools the fused spatiotemporal features by region to obtain a regional decision vector, concatenates all regional decision vectors and calculates the fusion weight of each region, and performs weighted summation of each regional decision vector based on the fusion weight to obtain a fused regional decision vector. The fused regional decision vector is concatenated with the global context vector and calculated to obtain the final decision vector. For details, see Figure 5 .

[0031] The construction process specifically includes the following sub-steps: Step 2.1: The video sequence to be detected in the data obtained in step 1 Input the face parsing and region division module to generate the corresponding region mask. Considering the relative stability of the face structure in a short time, all frames in the video sequence can be processed based on the first frame. regional division.

[0032] First, use the BiSeNet network to train the original face image Perform segmentation and generate initial face mask BiSeNet can efficiently segment the face area and provide a basis for subsequent area division. The process can be expressed as:

[0033] In order to accurately locate the facial features, the facial key point extraction tool is used to extract the key point coordinates of the face image. Based on these key points, a polygonal area of the T zone is constructed and filled to obtain the T zone mask. Next, in order to extract the face edge information, the initial face mask Perform dilation and erosion operations. First, use the dilation operation to expand range, get , and then use the erosion operation to reduce range, get Finally, the difference between the dilation and erosion results is calculated to obtain the edge mask The above process is as follows:

[0034] in, represents the expansion operation, Represents the corrosion operation, and kernel represents the convolution kernel used in the expansion corrosion operation. Indicates that the pixel (x, y) belongs to the edge area, Indicates that the pixel (x, y) does not belong to the edge area. In order to obtain the remaining face area except the T area and the edge, the initial face mask is used Subtract T-zone mask and edge mask , get the remaining face area mask . Mask the initial face Invert to get the background area mask , .

[0035] Finally, according to the four masks generated 、 、 、 , construct a region mask matrix , each element in the matrix represents the category to which the corresponding pixel in the original image belongs.

[0036]

[0037] Step 2.2: Compare the region mask matrix obtained in step 2.1 with the sequence to be detected Input region adaptive feature extraction module. This module firstly extracts the feature according to the region mask matrix At each level of feature extraction, a corresponding regional guidance matrix G is generated. This matrix is then used to guide the region-adaptive convolution operation: for each location on the feature map, the corresponding convolution kernel group is dynamically selected for calculation based on the region indicated by the regional guidance matrix G. In this way, features are extracted independently for different regions, ensuring that spurious features in local regions can be effectively modeled. Ultimately, this module outputs a feature sequence F containing rich regional information as the input to the regionalized spatiotemporal Transformer module.

[0038] Step 2.3: Input the feature sequence F obtained in step 2.2 into the regionalized spatiotemporal Transformer module. This module performs processing in the spatial and temporal dimensions in turn. In the spatial dimension, the regional guidance matrix G is used to divide the features in F into corresponding face region subsequences. (Ignore the background area.) Apply intra-region spatial attention (independent calculation within each region) and inter-region spatial attention (only allowing interaction between different regions) to these subsequences, and add the results of the two attentions to generate a spatial feature map. Next, the spatial feature map Perform a pooling operation to pool each region feature of each frame into one token. Then, in the temporal dimension, apply similar intra-regional temporal attention and inter-regional temporal attention mechanisms to the pooled feature sequence, and combine the results to capture the temporal dynamics across frames and the temporal correlation between regions to generate a temporal feature map. The whole process can be expressed as:

[0039] In this embodiment, the calculation method of intra-region attention and inter-region attention is as follows (taking the spatial dimension as an example, the method of the time dimension is similar): 1) For the calculation of attention within the region, first perform the face region subsequence , design independent query, key and value projection matrices for them respectively, and use different weights for different regions. Applicable area First, we linearly transform the features of each region to obtain the query, key, and value:

[0040] Then perform self-attention calculation on the token in each area:

[0041] After calculating the spatial attention features of each region, the attention results of all regions are spliced together to obtain .

[0042] 2) In order to calculate the attention between regions, it is necessary to perform the attention calculation globally, not just in each region. Therefore, we first convert the region-independent projection matrix obtained in the intra-region attention stage into The global Q, K, and V projection matrices are concatenated. Unlike intra-region attention, which only focuses on the interaction of patches within the same region, inter-region attention only allows patches between different regions to interact.

[0043] In the standard self-attention calculation, each patch will be calculated with all other patches, including patches belonging to the same region. However, in the intra-region attention stage, the feature interactions within the same region have been calculated separately, so an inter-region attention mask matrix is constructed based on the regional guidance matrix. . It is used to shield the attention calculation within the same area, ensuring that attention only flows between different areas. When calculating the attention between regions, Add to Then perform the softmax operation.

[0044]

[0045] Finally, the attention in the area is output and inter-region attention output Add together to get the spatial feature map .

[0046] Step 2.4: The spatial feature map obtained in step 2.3 and time feature maps Input region feature fusion module. This module is used for each predefined face region. Introducing an independent learnable fusion weight The weight is used to adaptively weight the spatial and temporal features corresponding to each region to obtain the fused spatiotemporal features. The process can be expressed as .

[0047] Step 2.5: Fusion spatiotemporal features obtained in step 2.4 Input inter-region feature fusion module, first divide the Pooling is performed by region to extract regional decision vectors At the same time, the fusion features in the region Pooling is performed to obtain the global context vector , used to supplement the overall information. The above process is as follows:

[0048] Next, the decision vectors of the three regions are concatenated, and the inter-region fusion weight is calculated through a two-layer fully connected (FC) network and an activation function. This weight is used to measure the contribution of each region to the fusion region decision vector. According to the calculated fusion weight, the decision vectors of each region are weighted summed to obtain the fusion region decision vector. The process is as follows:

[0049] Finally, the fusion region decision vector is concatenated with the global context vector, and the final decision vector is calculated through the fully connected layer. to predict the final result.

[0050]

[0051] Step 3: Use the fake face video dataset obtained in step 1 to train the fake face video detection network of the regionalized spatiotemporal Transformer, calculate the sum of various loss functions, and backpropagate to update the fake face video detection network parameters until the sum of various loss functions converges; During training, the network updates the overall network parameters according to the loss function. The overall loss function is expressed as: , in is the cross entropy loss function of the facial T-zone, is the cross entropy loss function in the marginal area, is the cross entropy loss function for other facial regions, is the cross entropy loss function of the final prediction result, 、 、 Represents the loss function weight parameters of the three regions respectively. is the weight parameter of the final prediction result loss function, which controls the balance between regional loss and final prediction loss. The specific loss functions are as follows:

[0052] in, Represent the prediction results of the three regions respectively. is the final prediction result after fusion, is the true label.

[0053] Step 4: Use the trained network model and input the test dataset obtained in step 1 to predict the labels of each fake face video sequence and real video sequence.

[0054] The implementation of each embodiment of the present invention is based on programmed processing performed by a device with processor functionality. Therefore, in practical engineering, the technical solutions and functionalities of each embodiment of the present invention are encapsulated into various modules. Based on this reality, and in addition to the aforementioned embodiments, an embodiment of the present invention provides a forged face video detection system based on a regionalized spatiotemporal Transformer. This system is used to implement the forged face video detection method based on a regionalized spatiotemporal Transformer described in the aforementioned method embodiments.

[0055] The system includes: a data acquisition module for acquiring forged face video data to be detected; a data detection module for inputting the forged face video data to be detected into a trained forged face video detection network and outputting a detection result; wherein, the training of the forged face video detection network includes: constructing a forged face video data set; constructing a forged face video detection network, including: a face parsing and region division module for obtaining a region mask matrix corresponding to three regions of a face image; a region adaptive feature extraction module for generating a region guide matrix according to the region mask matrix and independently extracting features of the three regions using the region guide matrix; a regionalized spatiotemporal Transformer module for further extracting the relationship between the three regions; an intra-regional feature fusion module for fusing the internal spatiotemporal features of each region; an inter-regional feature fusion module for fusing features of different regions to obtain a final decision vector; based on a designed loss function, network training is performed on the constructed forged face video data set, and the trained forged face video detection network is output.

[0056] The embodiment of the present invention provides a fake face video detection system based on a regionalized spatiotemporal Transformer. In view of the fact that most existing fake face video detection technologies adopt a global feature analysis strategy and ignore the differences in spatiotemporal features generated by the forgery process in different regions of the face, the system adopts the aforementioned modules, utilizes the differences in spatiotemporal features generated by the forgery process in different regions of the face, and uses a regionalized processing strategy throughout key links such as feature extraction, spatiotemporal modeling, and feature fusion, aiming to overcome the limitations of global analysis and detect fake face videos more accurately and robustly.

[0057] It should be noted that the system embodiments provided by the present invention are not only used to implement the methods in the above-mentioned method embodiments, but also used to implement the methods in other method embodiments provided by the present invention. The only difference lies in the setting of corresponding functional modules, and the principles thereof are basically the same as the principles of the above-mentioned system embodiments provided by the present invention. As long as those skilled in the art refer to the specific technical solutions in other method embodiments on the basis of the above-mentioned system embodiments, obtain corresponding technical means and technical solutions composed of these technical means by combining technical features, and on the premise of ensuring the practicality of the technical solutions, improve the modules in the above-mentioned system embodiments to obtain corresponding system class embodiments for implementing the methods in other method class embodiments.

[0058] Based on the same inventive concept as the aforementioned embodiment, an embodiment of the present invention further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions cause the computer to execute the steps of the method for detecting forged face videos based on a regionalized spatiotemporal Transformer, including: Obtain forged face video data to be detected; Inputting forged face video data to be detected into a trained forged face video detection network and outputting a detection result; wherein the training of the forged face video detection network includes: Build a fake face video dataset; A fake face video detection network was constructed, including: a face parsing and region division module, which is used to obtain the regional mask matrix corresponding to the three regions of the face image; a regional adaptive feature extraction module, which is used to generate the regional guidance matrix based on the regional mask matrix and use the regional guidance matrix to independently extract the features of the three regions; a regionalized spatiotemporal Transformer module, which is used to further extract the relationship between the three regions; an intra-regional feature fusion module, which integrates the internal spatiotemporal features of each region; and an inter-regional feature fusion module, which integrates the features of different regions to obtain the final decision vector. Based on the designed loss function, the network is trained on the constructed fake face video dataset, and the trained fake face video detection network is output.

[0059] In summary, the present invention discloses a method for detecting forged face videos based on a regionalized spatiotemporal Transformer. The method primarily includes the following steps: constructing a forged face video dataset containing both real and forged samples; constructing a forged face video detection network comprising a face parsing and region division module, a region-adaptive feature extraction module, a regionalized spatiotemporal Transformer module, an intra-regional feature fusion module, and an inter-regional feature fusion module; training the detection network using the dataset; and using the trained network model to predict the authenticity label of an input video sequence. The present invention utilizes the differences in spatiotemporal features generated by the forgery process in different regions of the face. Through a regionalized processing strategy, the method encompasses key steps such as feature extraction (using region-adaptive convolution), spatiotemporal modeling (using a spatiotemporal Transformer with intra-regional and inter-regional attention), and feature fusion (using an intra-regional and inter-regional fusion mechanism). The method aims to overcome the limitations of global analysis and enable more accurate and robust detection of forged face videos.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting fake face videos based on regionalized spatiotemporal Transformer, characterized by: include: Obtain forged face video data to be detected; Inputting forged face video data to be detected into a trained forged face video detection network and outputting a detection result; wherein the training of the forged face video detection network includes: Build a fake face video dataset; A fake face video detection network was constructed, including: a face parsing and region division module, which is used to obtain the regional mask matrix corresponding to the three regions of the face image; a regional adaptive feature extraction module, which is used to generate the regional guidance matrix based on the regional mask matrix and use the regional guidance matrix to independently extract the features of the three regions; a regionalized spatiotemporal Transformer module, which is used to further extract the relationship between the three regions; an intra-regional feature fusion module, which integrates the internal spatiotemporal features of each region; and an inter-regional feature fusion module, which integrates the features of different regions to obtain the final decision vector. Based on the designed loss function, the network is trained on the constructed fake face video dataset, and the trained fake face video detection network is output.

2. The method for detecting fake face videos based on regionalized spatiotemporal Transformer according to claim 1, characterized in that: When the face analysis and region division module obtains the region mask matrix corresponding to the three regions of the face image, it also includes: Input the video frame into BiSeNet to obtain the initial face mask; Use facial key point detection method to locate the T zone and generate a T zone mask; Perform dilation and erosion operations on the initial face mask to obtain the edge area mask; Combining the initial face mask, T-zone mask and edge region mask to determine other facial region masks and background region masks; Combining these masks generates a region mask matrix, where each element value represents the region class of the corresponding pixel.

3. The method for detecting fake face videos based on regionalized spatiotemporal Transformer according to claim 1, characterized in that: The regional adaptive feature extraction module generates the regional guidance matrix in the following way: For each layer of the feature extraction network, the mode of the pixel or feature region category in the receptive field of the convolution kernel of the current layer is calculated based on the region attribution information of the previous layer, the region attribution of each position on the feature map of the current layer is determined, and the region guidance matrix of the layer is generated.

4. The method for detecting fake face videos based on regionalized spatiotemporal Transformer according to claim 3, characterized in that: The region adaptive feature extraction module performs the region adaptive convolution operation as follows: An independent convolution kernel group is set for each predefined face area. During the convolution calculation, based on the regional guidance matrix of the current layer, only the convolution kernel group corresponding to the specific area on the feature map is used for feature calculation.

5. The method for detecting fake face videos based on regionalized spatiotemporal Transformer according to claim 1, characterized in that: The regionalized spatiotemporal Transformer module further extracts three-region relationships, including: The input feature sequence is divided into subsequences of different face regions according to the regional guidance matrix, and the background region features are ignored; In the spatial dimension, we apply the intra-region attention mechanism to the subsequences and the inter-region attention mechanism to the sequences and combine the results to obtain spatial features. The spatial features are pooled, and the intra-region and inter-region attention mechanisms are applied to the pooled features in the temporal dimension. The results are combined to obtain the temporal features.

6. The method for detecting fake face videos based on regionalized spatiotemporal Transformer according to claim 1, characterized in that: When fusing the spatiotemporal features within each region, the intra-region feature fusion module includes: The spatial features and temporal features are weightedly fused using learnable weights to obtain fused spatiotemporal features.

7. The method for detecting fake face videos based on regionalized spatiotemporal Transformer according to claim 1, characterized in that: When fusing features of different regions, the inter-region feature fusion module includes: The fused spatiotemporal features are globally pooled to obtain a global context vector. The fused spatiotemporal features are pooled by region to obtain a regional decision vector. All regional decision vectors are concatenated and the fusion weight of each region is calculated. The weighted sum of the regional decision vectors is performed according to the fusion weight to obtain the fused regional decision vector. The fusion region decision vector is concatenated with the global context vector and the final decision vector is calculated.

8. The method for detecting fake face videos based on regionalized spatiotemporal Transformer according to claim 1, characterized in that: The designed loss function is: , in is the cross entropy loss function of the facial T-zone, is the cross entropy loss function in the marginal area, is the cross entropy loss function for other facial regions, is the cross entropy loss function of the final prediction result, 、 、 Represents the loss function weight parameters of the three regions respectively. It is the weight parameter of the final prediction result loss function, which controls the balance between regional loss and final prediction loss.

9. A fake face video detection system based on regionalized spatiotemporal Transformer, characterized by: include: A data acquisition module is used to acquire forged face video data to be detected; The data detection module is configured to input the forged face video data to be detected into the trained forged face video detection network and output a detection result. The training of the forged face video detection network includes: Build a fake face video dataset; A fake face video detection network was constructed, including: a face parsing and region division module, which is used to obtain the regional mask matrix corresponding to the three regions of the face image; a regional adaptive feature extraction module, which is used to generate the regional guidance matrix based on the regional mask matrix and use the regional guidance matrix to independently extract the features of the three regions; a regionalized spatiotemporal Transformer module, which is used to further extract the relationship between the three regions; an intra-regional feature fusion module, which integrates the internal spatiotemporal features of each region; and an inter-regional feature fusion module, which integrates the features of different regions to obtain the final decision vector. Based on the designed loss function, the network is trained on the constructed fake face video dataset, and the trained fake face video detection network is output.

10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, which enable the computer to execute the steps of the method for detecting fake face videos based on regionalized spatiotemporal Transformer as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Face anti-counterfeiting detection method and device based on Transform domain information fusion

    CN121214527A