Face forgery detection method based on double-flow complementary architecture

By introducing a dual-stream complementary architecture, VicIA module, DevER module and DSR-KD framework into the facial forgery detection model, the problems of high parameter volume and high computation volume of the existing models are solved, and a lightweight facial forgery detection model is realized, which is suitable for multimedia devices with resource constrained.

CN119992619APending Publication Date: 2025-05-13ANHUI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510046204.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Due to the high parameter volume and high computing volume, existing facial forgery detection models are difficult to deploy to multimedia devices with limited storage and computing resources, resulting in limited promotion and application.

Method used

A facial forgery detection method based on a dual-flow complementary architecture is adopted, and feature extraction and enhancement is introduced by introducing VicIA module and DevER module, and knowledge distillation is used to design a lightweight detection model.

Benefits of technology

It realizes the complexity of the model while ensuring detection performance, and is suitable for resource-constrained multimedia devices, improving the lightweightness of the facial forgery detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992619A_ABST
    Figure CN119992619A_ABST
Patent Text Reader

Abstract

The invention discloses a face forgery detection method based on a double-flow complementary architecture, and belongs to the technical field of computer vision, and the method comprises the following steps: S1, constructing a network architecture; s2, knowledge distillation training; and S3, detecting face forgery. According to the ResNet backbone network reconstructed based on embedded space-frequency interactive convolution, an internal single-flow transmission mode is converted into a double-flow transmission mode; a pair of complementary lightweight extraction modules is designed based on an information complementation concept, and heuristic information and non-heuristic information of the image information are extracted respectively; in addition, an inverse residual attention module is utilized to align an intermediate feature map between the student model and the teacher model, in addition, a feature difference calculation module based on multi-level self-adaption is designed, the intermediate feature difference can be evaluated from the perspective of different feature map sizes, and therefore the learning progress of students can be described more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a facial forgery detection method based on a dual-stream complementary architecture. Background Art

[0002] Facial forgery technology has become one of the most commonly used tools for criminals to fabricate false news and spread rumors. Although this technology has shown great commercial value in multimedia applications (such as live streaming and film production), its rampant abuse has not only seriously undermined the authenticity and authority of digital resources, but also greatly lowered the threshold for cyber forgery crimes. Unfortunately, although the performance of facial forgery detection models has been continuously improved, the high number of parameters and high computational complexity have become the biggest factors hindering its promotion to multimedia devices with limited storage and computing resources (such as mobile terminals and monitoring systems). In order to meet this challenge, exploring a facial forgery detection model that can be deployed in resource-constrained scenarios has become a research hotspot in the field of multimedia security and forensics.

[0003] As the performance of facial forgery detection models has steadily improved, the need for large parameters and high computational load has become the biggest obstacle to curbing forgery crimes. For example, the overall parameter volume of the MAT detection model is as high as 417.65M, and its computational load has reached 99.8G; although the F2Trans detection model is much lighter than MAT, its parameter volume and computational load have reached 117.52M and 25.86G respectively; even the GHiEF version of the HiEF model with smaller parameter volume and computational load has 24.03M and 5.24G respectively. Therefore, it is increasingly important to design a more lightweight facial forgery model while ensuring competitive performance. From these perspectives, exploring highly lightweight solutions is becoming increasingly important for deep fake detection.

[0004] How to more comprehensively and lighten the facial forgery detection model is a problem to be solved urgently. To this end, the present invention proposes a facial forgery detection method based on a dual-stream complementary architecture. Summary of the invention

[0005] The technical problem to be solved by the present invention is: how to more comprehensively lightweight the facial forgery detection model, thereby solving the problem of incomplete model compression in the prior art, and providing a facial forgery detection method based on a dual-stream complementary architecture.

[0006] like Figure 6 As shown, the present invention solves the above technical problems through the following technical solutions, and the present invention comprises the following steps:

[0007] S1: Network architecture construction

[0008] A facial forgery detection network architecture is constructed, including a downsampling layer, an extraction layer, a fusion and complementary layer, and a classification layer. The VicIA module and the DevER module are introduced in the extraction layer. The VicIA module aggregates local context information without introducing additional parameters, and enhances edge features through the DevER module. Its parameters are optimized to half of the original size. While aggregating neighboring information and enhancing abnormal edges, the single-stream feature map obtained by the downsampling layer is processed by the VicIA module and the DevER module and then divided into a dual-stream feature map. The dual-stream feature map is sequentially passed to multiple fusion and complementary layers for repeated fusion and complementation, and finally the final feature map is output by element-by-element addition. The final feature map is flattened after adaptive average pooling through the classification layer and mapped to the classification category by the fully connected layer.

[0009] S2: Knowledge Distillation Training

[0010] The teacher model and the student model structures are designed based on the facial forgery detection network architecture constructed in step S1; in the teacher model, there are two extraction layers and four fusion complementary layers, and the two extraction layers are respectively located before the first fusion complementary layer and the third fusion complementary layer; in the student model, there is one extraction layer and five fusion complementary layers, and the extraction layer is located before the first fusion complementary layer; the rest of the structures of the student model of the teacher model are the same; the DSR-KD framework is used to perform knowledge distillation training to obtain the student model after knowledge distillation training, that is, the facial forgery detection model;

[0011] S3: Facial Forgery Detection

[0012] The facial image to be detected is input into the facial forgery detection model to obtain the facial forgery detection result.

[0013] Furthermore, in step S1, the specific processing process of the VicIA module is as follows:

[0014] S101: In space, the non-heuristic self-information graph I(f i ) is mapped to between 0 and 1, and then element-by-element multiplication is performed to obtain the spatial attention map A. s ; In the channel, average the global channel I(f i ), then linearly transformed Convert to channel attention map A c ; The specific calculation formula is as follows:

[0015]

[0016] S102: A s With A c Appended to the feature map f obtained by the downsampling layer, the formula is as follows:

[0017]

[0018] in, Represents element-wise multiplication;

[0019] Finally, the non-heuristic information flow feature graph f output by the VicIA module is obtained v .

[0020] Furthermore, in step S101, the non-heuristic self-information graph I (f i ) is calculated as follows:

[0021] I(f i (x,y))=-logP s (f i (x,y))-logP c (f i (x,y)

[0022] in, f i (x, y) represents the pixel value of the feature map f of the i-th channel at (x, y), σ represents the scale parameter of the radial basis function, and m and n represent the spatial dimension R s The index of the current neighboring pixel on the graph, and i represents the channel dimension R c The index of the current neighboring channel.

[0023] Furthermore, in step S1, the DevER module uses a learnable constrained separable convolution LCSConv to enhance the abnormal facial edge information. The specific processing process is as follows:

[0024] S111: Use LCSConv to extract and enhance abnormal edge information;

[0025] S112: Enhance the feature map f obtained by the downsampling layer to a heuristic information flow feature map f d :

[0026] f d =LCSConv(ω i (f))

[0027] Among them, ω i represents the convolution kernel weight of the i-th channel LCSConv, i∈c, c is the number of channels of the input feature map f.

[0028] Furthermore, in step S111, the convolution kernel size of LCSConv is fixed to 5×5, and the constraint rules of the convolution kernel of LCSConv are as follows:

[0029]

[0030] Among them, (w / 2,h / 2) represents the center position of the current convolution kernel.

[0031] Furthermore, in step S2, in the DSR-KD framework, the RevRA module is used to align the knowledge between the layers of the model, and then the HieDC module is used to obtain the difference in knowledge between the aligned student model and the teacher model, thereby evaluating the learning progress of the student model.

[0032] Furthermore, in the RevRA module, the specific processing process is as follows:

[0033] S201: For one of the two-stream data, first pass the inverse residual attention map of the previous layer The size of the feature map of this layer is changed to Same attention map

[0034] S202: Then splice and Then it is handed over to the two-dimensional 1×1 convolution C 1×1 Output two attention maps and

[0035] S203: and Append to and Then, after element-by-element addition, an adaptive intermediate attention map is generated.

[0036] S204: After the RevRA module processes the dual-stream data, the student model obtains a series of intermediate feature maps that can be compared with the teacher model and Aligned Inverse Residual Attention Map and

[0037] Furthermore, in step S201, the inverse residual attention map calculation formula of the student model is as follows:

[0038]

[0039] Among them, k∈{K|1,2,3,4,5} represents the last five layers of the face forgery detection network, Represents the feature map of the current layer.

[0040] Furthermore, in the HieDC module, the specific processing process is as follows:

[0041] S211: HieDC module based on the inverse residual attention map of the student model and the intermediate feature map f of the teacher model t k The size S dynamically generates a new size sequence H, and the generation rules are as follows:

[0042]

[0043] S212: Inverse residual attention map of the student model and the intermediate feature map f of the teacher model t k Generate a series of difference maps using adaptive average pooling based on H;

[0044] S213: Obtain the difference value of the current level by calculating the mean square error of the paired student and teacher model difference maps pixel by pixel. The calculation formula is as follows:

[0045]

[0046] Where h is the sequence index of H;

[0047] The difference values ​​of multiple levels are integrated into the multi-level MSE loss, and the calculation formula is as follows:

[0048]

[0049] S214: Calculate the multi-level MSE loss for all aligned layers And finally after balancing, it is synthesized into the level difference loss

[0050]

[0051] Among them, K is the number of intermediate layers of the network, and hierarchy represents the number of elements in the sequence H;

[0052] S215: After the HieDC module processes the dual-stream data, the hierarchical difference loss of the dual stream is obtained and

[0053] Furthermore, in step S2, during the knowledge distillation training process, the total loss is defined as follows:

[0054]

[0055] in, is the cross entropy loss, used as the classification loss of the student model, It is the two-stream knowledge distillation loss, which is used as the distillation loss of the DSR-KD framework;

[0056] Cross Entropy Loss The definition is as follows:

[0057]

[0058] Where N is the number of images, is the predicted value of the i-th image, y i is the label value of the image;

[0059] Two-stream knowledge distillation loss The definition is as follows:

[0060]

[0061] Among them, α is a hyperparameter that balances the two-stream hierarchical difference loss.

[0062] Compared with the prior art, the present invention has the following advantages:

[0063] 1. The DSC-Light of the present invention is a ResNet backbone network reconstructed based on embedded space-frequency interactive convolution, which retains the hierarchical design structure of ResNet, but converts its internal single-stream transmission mode into dual-stream transmission. In addition, in order to take into account both lightweight and high efficiency, a pair of complementary lightweight extraction modules are designed based on the concept of information complementarity to extract heuristic information and non-heuristic information of image information respectively. In order to differentiate the design of the teacher model and the student model, the extraction layer is placed at different positions, and the number of channels of the student model is halved to compress the complexity of the model as much as possible.

[0064] 2. The present invention also proposes a new customized feature knowledge distillation mechanism, which adapts to the unique internal dual-stream structure of the above detector and uses an inverse residual attention module that can mimic the human learning process to align the intermediate feature maps between the student model and the teacher model. In addition, in order to evaluate the difference between the two models in a more fine-grained manner, a feature difference calculation module based on multi-level adaptation is designed, which can evaluate the intermediate feature difference from the perspective of different feature map sizes, thereby more accurately describing the student's learning progress. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is the overall architecture diagram of DSC-Light in the embodiment of the present invention. The face images in the figure are images in the FaceForensics++ dataset, which is an open source dataset;

[0066] Figure 2 is a schematic diagram of the structure of two lightweight modules in an embodiment of the present invention, wherein a is a VicIA module and b is a DevER module;

[0067] Figure 3 is a schematic diagram of the structure of two customized modules of the DSR-KD framework in an embodiment of the present invention, wherein a is the RevRA module and b is the HieDC module;

[0068] Figure 4 is a complexity comparison diagram of different models in the embodiments of the present invention;

[0069] Figure 5 is a classification decision heat map of the facial forgery detection model in an embodiment of the present invention, wherein the first row is the original image of the forged image, the second row is the CAM class activation response image, and the third row is the classification decision heat map obtained by using Grad-CAM; the face images in the figure are images in the FaceForensics++ dataset and the CelebDF dataset, both of which are open source datasets;

[0070] Figure 6 It is a flow chart of the facial forgery detection method based on the dual-stream complementary architecture of the present invention. DETAILED DESCRIPTION

[0071] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given, but the protection scope of the present invention is not limited to the following embodiment.

[0072] In order to cope with the challenge of the surge in the overall scale of current facial forgery detection models, we proposed a new solution for lightweight facial forgery detection models from the perspectives of reconstructing the backbone network and compressing the network model while taking into account both detection performance and model scale. As a result, we obtained a dual-stream complementary lightweight deepfake detector (Dual-Stream Complementary Lightweight Deepfake Detector, DSC-Light), which effectively utilizes the dual-stream residual knowledge distillation mechanism (Dual-Stream Residual Knowledge Distillation Mechanism, DSR-KD) to enhance its extraction of dual-stream complementary information and further compress its complex internal structure. Its overall architecture is as follows: Figure 1As shown. The method of the present invention cleverly exploits the complementarity of non-heuristic (requires iterative training to obtain) and heuristic (can be extracted without training) information, and combines them to improve detection performance while maintaining the compactness of the model. The architecture design of DSC-Light has two novel lightweight extraction modules: the Vicinity Information Aggregation (VicIA) module and the Deviant Edge Refinemen (DevER) module. The VicIA module aggregates local contextual information without introducing additional parameters, which helps capture the fine-grained details necessary for detecting subtle traces of tampering. On the other hand, the DevER module improves the edge features, and its parameters are optimized to half of the original size, thereby reducing complexity while retaining basic features. By combining these modules, DSC-Light achieves an effective trade-off between feature representation and feature expression. In order to further improve the performance of DSC-Light and reduce its complexity, we introduce the DSR-KD mechanism with a mirror structure. It simulates human learning through the Reverse Residual Attention (RevRA) module, enabling the student model to extract knowledge from the teacher model through different feature streams. In addition, DSR-KD utilizes the Hierarchical Differences Calculation (HieDC) module to evaluate the differences between the teacher and student models, optimize the distillation process, and improve the transfer efficiency and accuracy.

[0073] Traditional facial forgery detection methods often stack extraction layers or extraction modules before the backbone network, which makes the detection network more cumbersome. In order to make the detection network more lightweight, we use a single-stream structure based on the improved embedded two-stream convolution to reconstruct the internal structure of the ResNet backbone network in DSC-Light. Looking back at its overall structure, DSC-Light consists of four layers: downsampling layer, extraction layer, fusion complementary layer and classification layer. When given a 3-channel RGB image h and w represent the height and width of image x respectively. DSC-Light first passes it to the downsampling layer. The downsampling layer uses a 7×7 two-dimensional convolution kernel to obtain global wide features under a large receptive field. And directly reduce the scale of the information flow input to the subsequent network layer, where c represents the number of output channels, and h′ and w′ represent the height and width after reduction. While aggregating neighboring information and enhancing abnormal edges, the single stream f of DSC-Light is divided into complementary dual streams f by the extraction layer. v and f d .f v and fd have the same scale, that is, the number of channels, height and width of the feature maps of the two streams are consistent. Then, f v and f d It is sequentially passed into multiple fusion and complementary layers for repeated fusion and complementation, and finally output f by element-by-element addition. m Finally, f m After being flattened by the classification layer for adaptive average pooling, it is then mapped to the classification category by the fully connected layer.

[0074] Since DSC-Light has a two-stream structure, the traditional single-stream feature knowledge distillation method cannot be applied to this network structure. To this end, we customized the DSR-KD framework that can be split for offline knowledge distillation. We use the inverse residual link network to distill the intermediate feature maps of each layer of the two-stream of the student model (S-DSC-Light) and The attention link is performed, where the number of layers k∈{1,2,3,4,5}. Next, the inverse residual attention map of the student model with attention is calculated and Intermediate feature maps with the pre-trained teacher model (T-DSC-Light) and The difference between the various levels between them, and then the two-stream knowledge distillation loss is obtained This loss is used to evaluate the knowledge gap between the teacher model and the student model, and forces the teacher model to better guide the student. The DSR-KD framework complements the facial forgery detection capability of S-DSC-Light, achieving a more lightweight network while taking into account detection performance.

[0075] In order to highly lightweight the detector, we design the number and position of extraction layers in the teacher model and the student model differently. The powerful teacher model contains multiple extraction layers in its backbone network, for example, the extraction layer is placed before the first and third layers of the backbone network; the streamlined student model only uses the extraction layer before the first layer of the backbone network. The extraction layer is a crucial layer in DSC-Light, which not only extracts complementary non-heuristic and heuristic information, but also acts as a diversion layer to divide the single-stream information within the network into two-stream information. The detailed structure is as follows: Figure 2 shown.

[0076] Neighborhood information aggregation module: Neighborhood information is a kind of non-heuristic image information, and the acquisition of non-heuristic information does not require training network weights. Neighborhood information only calculates the correlation between the current pixel and the surrounding pixels. This calculation does not generate network parameters, but only increases the amount of calculation, thus ensuring that DSC-Light will not weaken its lightweight due to the use of this module. The essence of neighbor information is a kind of self-information calculated based on the difference between pixels. The concept of self-information comes from information theory, which is used to measure the uncertainty of information or the amount of information. Its calculation formula is shown in formula (1):

[0077] I(θ)=-log2P(θ) (1)

[0078] Among them, I(θ) represents the self-information of event θ, and P(θ) represents the probability of event θ. The self-information I(θ) is inversely proportional to the probability of event θ P(θ). The higher P(θ), the smaller I(θ). Vice versa, the lower P(θ), the larger I(θ).

[0079] We generalize the calculation of self-information to the calculation of pixel differences. In the spatial dimension and channel dimension, radial basis functions are used to approximate the joint probability distribution P of the current pixel and its neighboring pixels. s (θ) and P c (θ), and regard them as self-information to calculate P(θ) in formula (1). The representation of the joint probability distribution is shown in formula (2):

[0080]

[0081] Among them, f i (x, y) represents the pixel value of the feature map f of the i-th channel at (x, y). σ represents the scale parameter of the radial basis function, which controls its influence range in the current dimension. m and n represent the spatial dimension R s The index of the current neighboring pixel on the graph, and i represents the channel dimension R c The index of the current adjacent channel on the . The complete self-information calculation is shown in formula (3).

[0082] I(f i (x,y))=-log P s (f i (x,y))-log P c (f i (x,y)) (3)

[0083] We convert the calculated self-information map into an attention map. In space, we transform the non-inspired self-information map I(f i ) is mapped to between 0 and 1, and then element-by-element multiplication is performed to obtain the spatial attention map A.s In the channel, we average the global channel I(f i ), then linearly transformed Convert to channel attention map A c . Each I(f i ) is shown in formula (4):

[0084]

[0085] Then, we will A s With A c Append to the original feature f, as shown in formula (5):

[0086]

[0087] in, Represents element-wise multiplication.

[0088] Finally, the non-heuristic information flow f output by the VicIA module is obtained v .

[0089] Abnormal edge enhancement module: As a supplement to non-heuristic information, heuristic image information (such as mixed boundaries, texture anomalies, and high-frequency noise) is essential for facial forgery detection. Different from the parameter-free characteristics of the non-heuristic information acquisition process, heuristic information requires continuous iterative learning to optimize network parameters. Therefore, in order to maintain the lightweight of DSC-Light, we designed a DevER module with fewer parameters to enhance the heuristic information.

[0090] Heuristic forged information often exists in the edge areas of image content (such as the sides of the cheeks, around the nose, and under the lips). These edge areas usually contain both abnormal forged traces and normal image content. In order to enhance the heuristic information in a directionally manner, we use constrained convolution to replace the original convolution. It is well known that constrained convolution is an efficient tool for enhancing specific local forged information. However, the weights of traditional constrained convolutions are often set in advance manually. For example, SFIConv uses multi-channel separable constrained convolutions with fixed convolution kernel weights to extract forged information. Constrained convolutions with fixed convolution kernel weights often rely too much on the initialization of weights, resulting in the inability to adaptively enhance various edge information. Therefore, we propose to use learnable constrained separable convolution (LCSConv) to enhance facial abnormal edge information.

[0091] The convolution kernel weight of this convolution is variable, and it is optimized through repeated iterations. The center value of the convolution kernel is artificially set to -1, and the surrounding pixel values ​​are regularized to ensure that the sum of all pixel values ​​except the center value is 1. The constraint rule of the convolution kernel of LCSConv can be expressed as the following formula (6):

[0092]

[0093] Among them, ω i represents the convolution kernel weight corresponding to the i-th channel, and (w / 2,h / 2) represents the center position of the current convolution kernel. Since the center value of the convolution kernel is fixed to -1, when abnormal edge information is extracted, the abnormal area will be given a large negative weight, thereby suppressing the expression of the surrounding normal information and enhancing the abnormal edge information.

[0094] f d =LCSConv(ω i (f)) (7)

[0095] Considering the need to obtain more comprehensive local information under a larger receptive field, the convolution kernel size of LCSConv is fixed to 5×5. Considering that traditional constrained convolution often faces the problem of large number of parameters and computation, we imitate the depthwise separable convolution and equip each channel with a separate learnable constrained separable convolution. This design makes the DevER module much lighter than the traditional enhancement module. By using the DevER module, we enhance the output f of the downsampling layer to f d , the formula is shown in the above formula (7), where i∈c, c is the number of channels of the input feature map f.

[0096] The two extraction modules (VicIA module and DevER module) of the above extraction layer aggregate non-heuristic information and enhance heuristic information without increasing the number of parameters and calculations. Next, these complementary dual-stream information will be interactively fused and propagated in the fusion interaction layer. The transformation module uses 1×1 convolution to align the channel dimensions of the dual-stream features to facilitate the element-by-element fusion of the dual streams. Finally, the fusion interaction layer is sent to the classification layer for mapping and classification after the dual streams are reconstructed into a single stream.

[0097] In order to adapt to the unique internal two-stream structure of DSC-Light and achieve a high degree of lightweight for S-DSC-Light, we customized a new two-stream residual knowledge distillation (DSR-KD) framework. The DSR-KD framework can transfer non-heuristic information and heuristic information from T-DSC-Light to S-DSC-Light. How to align the intermediate feature maps of the teacher model and the student model is a crucial issue in feature knowledge distillation. In the DSR-KD framework, we designed the RevRA module to align the knowledge between the layers of the model. Then the HieDC module is used to obtain the difference in knowledge between the aligned student model and the teacher model, and then evaluate the learning progress of the student model.

[0098] Reverse residual attention module: The traditional feature knowledge distillation alignment method is often to align only the feature maps of the same level between the teacher and student models, or to align the feature maps of each layer of the teacher model with all levels of the student model. These methods always make it impossible for the student model to fully learn the knowledge of the teacher model. We customized the RevRA module for the structure of DSC-Light to summarize the intermediate features of the student model to coordinate the knowledge learned by the student model.

[0099] The RevRA module uses an inverse residual link mechanism with attention to generate the intermediate attention map of the student model. The positive residual link network is from shallow to deep, so that the deep network can obtain the original input information; while the inverse residual link network is from deep to shallow, so that the shallow network can reversely obtain the deep output information. It is not constructive to apply this reverse residual link mechanism in ordinary network training, but because the inverse residual link mechanism is closer to the human learning process, it has a subtle effect in the feature alignment of the teacher model and the student model in feature knowledge distillation. With growth, people will have a deeper understanding and firm memory of what they have learned in the past; similarly, with training, the student model will learn more general and comprehensive knowledge of the teacher model in iterations, rather than just learning the features of the corresponding layer. The formula of this inverse residual link mechanism is shown in the following formula (8):

[0100]

[0101] Among them, k∈{K1,2,3,4,5} represents the last five layers of DSC-Light. Represents the student inverse residual attention map after being processed by the RevRA module. Represents the feature map of the current layer. We also use mirror distillation to transfer knowledge to the non-inspired information flow and the heuristic information flow of DSC-Light synchronously, and transfer the m s Divided into m sv With msd , the teacher model’s f t is divided into f tv and f td .

[0102] Directly learning or memorizing complex knowledge mechanically is not conducive to understanding the deep meaning behind it, but requires an adaptive learning process. The same is true in the feature distillation process. The shallow network of the student model is not conducive to its understanding if it is exposed to unprocessed deep features too early. Therefore, in order to imitate this learning process in feature knowledge distillation, we use the attention mechanism to reversely learn the deep knowledge of the student model. The mechanism diagram is as follows: Figure 3 shown.

[0103] First, the RevRA module first converts the inverse residual attention map passed by the previous layer The size of the feature map of this layer is changed to Same attention map Then splice and Then it is handed over to the two-dimensional 1×1 convolution C 1×1 Output two attention maps and The specific formula is shown in formula (9).

[0104]

[0105] at last, and Append to and Then, after element-by-element addition, an adaptive intermediate attention map is generated. The formula is shown in the following formula (10).

[0106]

[0107] The above statement generally refers to the processing process of each stream in DSR-KD. Since both streams have similar structures, the overall distillation process is essentially the same as the above description. After being processed by the two-stream residual knowledge distillation framework, the student model can obtain a series of intermediate feature maps that can be used with the teacher model. and Aligned Inverse Residual Attention Map and

[0108] Hierarchical difference calculation module: In the knowledge transfer process of feature knowledge distillation, evaluating the difference between the intermediate feature maps of the teacher model and the student model is an indispensable step. Accurate evaluation of the difference often enables the student model to better obtain the guidance of the teacher model and ultimately achieve performance comparable to that of the teacher model. The DSR-KD framework designs a HieDC module to calculate the difference in knowledge, such as Figure 3 shown.

[0109] The HieDC module can generate dynamic multi-level difference maps when calculating the difference at each layer. First, the HieDC module generates dynamic multi-level difference maps based on the student inverse residual attention map. and the teacher's intermediate feature map f t k The size S (i.e., height or width) is used to dynamically generate a new size sequence H. The generation rule is shown in the following formula (11):

[0110]

[0111] Then, students and teachers t k Adaptive average pooling is used to generate a series of difference maps based on H. Difference maps of different sizes contain different focus points of DSC-Light. Next, the difference value of the current level is obtained by calculating the mean squared error (MSE) of the paired student and teacher difference maps pixel by pixel. The calculation formula of MSE is shown in the following formula (12):

[0112]

[0113] Where h is the sequence index of H. The difference values ​​of multiple levels are integrated into the multi-level MSE loss, which is calculated as formula (13):

[0114]

[0115] To avoid the impact of the continuously reduced difference map on the loss, we add the same weight as the reduced scale to balance the MSE. Finally, we calculate the MSE between all aligned layers. And finally after balancing, it is synthesized into the level difference loss The calculation formula is shown in formula (14):

[0116]

[0117] The above hierarchical difference loss is the calculation result of a single stream in the DSR-KD framework, and the other stream is calculated in the same way. Finally, we can obtain the hierarchical difference loss of the two streams through the HieDC module. and

[0118] Overall process: In the training of T-DSC-Light and S-DSC-Light, we use cross entropy loss as the classification loss of the network, which is defined as:

[0119]

[0120] Where N is the number of images, is the predicted value of the i-th image, y i is the label value of the image.

[0121] In the DSR-KD framework, we use a two-stream knowledge distillation loss The distillation loss of this framework is defined as:

[0122]

[0123] Among them, α is a hyperparameter to balance the two-stream loss. Finally, the total loss can be defined as:

[0124]

[0125] Among them, λ is the hyperparameter that balances S-DSC-Light and DSR-KD frameworks during the distillation process. The complete training process of DSC-Light can be divided into two steps. The first step is to pre-train T-DSC-Light to obtain powerful teacher model parameters; the second step is to use the DSR-KD framework to perform offline knowledge transfer on S-DSC-Light, enhance the performance of the student model and achieve lightweight. The detailed process is shown in Table 1 below:

[0126] Table 1 Complete training process

[0127]

[0128]

[0129] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A facial forgery detection method based on a two-stream complementary architecture, characterized in that: The following steps are involved: S1: Network architecture construction A facial forgery detection network architecture is constructed, including a downsampling layer, an extraction layer, a fusion and complementary layer, and a classification layer. The VicIA module and the DevER module are introduced in the extraction layer. The VicIA module aggregates local context information without introducing additional parameters, and enhances edge features through the DevER module. Its parameters are optimized to half of the original size. While aggregating neighboring information and enhancing abnormal edges, the single-stream feature map obtained by the downsampling layer is processed by the VicIA module and the DevER module and then divided into a dual-stream feature map. The dual-stream feature map is sequentially passed to multiple fusion and complementary layers for repeated fusion and complementation, and finally the final feature map is output by element-by-element addition. The final feature map is flattened after adaptive average pooling through the classification layer and mapped to the classification category by the fully connected layer. S2: Knowledge Distillation Training The teacher model and the student model structures are designed based on the facial forgery detection network architecture constructed in step S1; in the teacher model, there are two extraction layers and four fusion complementary layers, and the two extraction layers are respectively located before the first fusion complementary layer and the third fusion complementary layer; in the student model, there is one extraction layer and five fusion complementary layers, and the extraction layer is located before the first fusion complementary layer; the rest of the structures of the student model of the teacher model are the same; the DSR-KD framework is used to perform knowledge distillation training to obtain the student model after knowledge distillation training, that is, the facial forgery detection model; S3: Facial Forgery Detection The facial image to be detected is input into the facial forgery detection model to obtain the facial forgery detection result.

2. The facial forgery detection method based on a dual-stream complementary architecture according to claim 1, characterized in that: In step S1, the specific processing process of the VicIA module is as follows: S101: In space, the non-heuristic self-information graph I(f i ) is mapped to between 0 and 1, and then element-by-element multiplication is performed to obtain the spatial attention map A. s ; In the channel, average the global channel I(f i ), then linearly transformed Convert to channel attention map A c ; The specific calculation formula is as follows: S102: A s With A c Appended to the feature map f obtained by the downsampling layer, the formula is as follows: in, Represents element-wise multiplication; Finally, the non-heuristic information flow feature graph f output by the VicIA module is obtained v .

3. The facial forgery detection method based on a dual-stream complementary architecture according to claim 2 is characterized in that: In step S101, the non-heuristic self-information graph I(f i ) is calculated as follows: I(f i (x,y))=-logP s (f i (x,y))-logP c (f i (x,y)) in, f i (x, y) represents the pixel value of the feature map f of the i-th channel at (x, y), σ represents the scale parameter of the radial basis function, and m and n represent the spatial dimension R s The index of the current neighboring pixel on the graph, and i represents the channel dimension R c The index of the current neighboring channel.

4. The facial forgery detection method based on a dual-stream complementary architecture according to claim 2, characterized in that: In step S1, the DevER module uses a learnable constrained separable convolution LCSConv to enhance the abnormal facial edge information. The specific processing process is as follows: S111: Use LCSConv to extract and enhance abnormal edge information; S112: Enhance the feature map f obtained by the downsampling layer to a heuristic information flow feature map f d : f d =LCSConv(ω i (f)) Among them, ω i represents the convolution kernel weight of the i-th channel LCSConv, i∈c, c is the number of channels of the input feature map f.

5. The facial forgery detection method based on a dual-stream complementary architecture according to claim 4 is characterized in that: In step S111, the convolution kernel size of LCSConv is fixed to 5×5, and the constraint rules of the convolution kernel of LCSConv are as follows: Among them, (w / 2,h / 2) represents the center position of the current convolution kernel.

6. The facial forgery detection method based on a dual-stream complementary architecture according to claim 4, characterized in that: In step S2, in the DSR-KD framework, the RevRA module is used to align the knowledge between the layers of the model, and then the HieDC module is used to obtain the difference in knowledge between the aligned student model and the teacher model, thereby evaluating the learning progress of the student model.

7. The facial forgery detection method based on dual-stream complementary architecture according to claim 6, characterized in that: In the RevRA module, the specific processing process is as follows: S201: For one of the two-stream data, first pass the inverse residual attention map of the previous layer The size of the feature map of this layer is changed to Same attention map S202: Then splice and Then it is handed over to the two-dimensional 1×1 convolution C 1×1 Output two attention maps and S203: and Append to and Then, after element-by-element addition, an adaptive intermediate attention map is generated. S204: After the RevRA module processes the dual-stream data, the student model obtains a series of intermediate feature maps that can be compared with the teacher model and Aligned Inverse Residual Attention Map and 8. The facial forgery detection method based on dual-stream complementary architecture according to claim 7, characterized in that: In step S201, the inverse residual attention map calculation formula of the student model is as follows: Among them, k∈{K|1,2,3,4,5} represents the last five layers of the face forgery detection network, Represents the feature map of the current layer.

9. The facial forgery detection method based on dual-stream complementary architecture according to claim 7, characterized in that: In the HieDC module, the specific processing process is as follows: S211: HieDC module based on the inverse residual attention map of the student model and the intermediate feature map f of the teacher model t k The size S dynamically generates a new size sequence H, and the generation rules are as follows: S212: Inverse residual attention map of the student model and the intermediate feature map f of the teacher model t k Generate a series of difference maps using adaptive average pooling based on H; S213: Obtain the difference value of the current level by calculating the mean square error of the paired student and teacher model difference maps pixel by pixel. The calculation formula is as follows: Where h is the sequence index of H; The difference values ​​of multiple levels are integrated into the multi-level MSE loss, and the calculation formula is as follows: S214: Calculate the multi-level MSE loss for all aligned layers And finally after balancing, it is synthesized into the level difference loss Among them, K is the number of intermediate layers of the network, and hierarchy represents the number of elements in the sequence H; S215: After the HieDC module processes the dual-stream data, the hierarchical difference loss of the dual stream is obtained and 10. The facial forgery detection method based on dual-stream complementary architecture according to claim 9, characterized in that: In step S2, during the knowledge distillation training process, the total loss is defined as follows: in, is the cross entropy loss, used as the classification loss of the student model, It is the two-stream knowledge distillation loss, which is used as the distillation loss of the DSR-KD framework; Cross Entropy Loss The definition is as follows: Where N is the number of images, is the predicted value of the i-th image, y i is the label value of the image; Two-stream knowledge distillation loss The definition is as follows: Among them, α is a hyperparameter that balances the two-stream hierarchical difference loss.