A Face Forgery Detection Method Based on Feature-Enhanced Mamba Framework
The feature-enhanced Mamba framework addresses the limitations of existing face forgery detection by integrating local and global feature modules with Hilbert scanning, improving accuracy and efficiency in face forgery detection.
Patent Information
- Application Number
- CN202510585552.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Existing face forgery detection methods face challenges in modeling long-range dependencies in convolutional neural networks, leading to incomplete utilization of global association information, and Transformer frameworks suffer from high computational complexity, causing inefficiencies and resource constraints, especially in large-scale data processing.
A feature-enhanced Mamba framework is introduced, incorporating local and global feature enhancement modules, utilizing Hilbert scanning for encoding and decoding, and combining with residual and semi-feature pyramid networks to extract and fuse features, guided by a weighted loss function for improved detection accuracy.
The framework effectively captures global and local features, enhancing detection accuracy and reducing computational overhead, ensuring efficient and reliable face forgery identification.
Smart Images

Figure CN120088838B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically, to a face forgery detection method based on a feature-enhanced Mamba framework. Background Art
[0002] The development of deep learning has made face forgery technology increasingly mature. Through technologies such as generative adversarial networks, realistic face images can be generated. The forged content is not only used to create pornographic videos and fake news, but may also be used for political attacks and financial fraud, etc., bringing great threats to fields such as politics, security, transportation, and finance. Therefore, how to identify forged face images and maintain social security and public order has become an urgent problem to be solved.
[0003] Although existing face forgery detection methods have achieved good results, when extracting face features, convolutional neural networks still have the problem of being difficult to model long-range dependence relationships. The inability to model long-range dependence relationships will cause the detection model to be unable to fully utilize global correlation information during the training process, resulting in incorrect detection results. The Transformer framework has the problem that the computational complexity is high, leading to excessive time consumption for forgery detection. The high computational complexity not only increases the hardware cost, but may also cause the system to be unable to operate normally under limited resources. At the same time, it becomes more difficult to process large-scale data, which may cause the system to run slowly or even crash. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a face forgery detection method based on a feature-enhanced Mamba framework, apply the Mamba framework to the field of deep forgery detection, and innovatively transform the original framework. On this basis, a feature-enhanced Mamba framework is proposed. This framework can accurately capture the global and local features of images, achieving good results in complex and changeable face forgery detection and being able to effectively identify forged face images.
[0005] The present invention adopts the following technical solutions to achieve the invention purpose:
[0006] A face forgery detection method based on a feature-enhanced Mamba framework, characterized by including the following steps:
[0007] S1: Preparation of the data set;
[0008] S2: Construct a feature-enhanced Mamba framework;
[0009] This framework introduces a locality-enhanced state space module on the basis of the original Mamba framework. This module includes a global feature enhancement module and a local feature enhancement module;
[0010] S3: Training of the model;
[0011] The specific image feature extraction process is as follows:
[0012] S31: Feature extraction of the global feature enhancement module;
[0013] S32: Feature extraction of the local feature enhancement module;
[0014] S33: Obtaining of the fused features;
[0015] S4: Construction of the loss function.
[0016] As a further limitation of this technical solution, the specific steps of S31 are:
[0017] After the picture is processed by the residual connection network and the semi-feature pyramid network, global information is extracted It is processed through the global feature enhancement module to obtain the enhanced global feature information ;
[0018] Global enhanced feature The extraction formula is as follows:
[0019] (1);
[0020] Where: n ∈ N represents the number of global feature enhancement modules;
[0021] LN() represents layer normalization;
[0022] SSMs() represents the state space model;
[0023] represents the global feature enhancement module encoder;
[0024] represents the global feature enhancement module decoder;
[0025] represents the activation function;
[0026] represents the depthwise separable convolution with a convolution kernel size of 3x3;
[0027] represents the linear layer;
[0028] represents the global feature enhancement module;
[0029] In the global feature enhancement module, both the global feature enhancement module encoder and the global feature enhancement module decoder use the Hilbert scanning method for scanning;
[0030] The generation formula of the Hilbert curve is as follows:
[0031] (2);
[0032] Where: represents the nth-order Hilbert matrix. When n > 1, the Hilbert matrix is constructed by recursively combining four nth-order Hilbert matrices , , and ;
[0033] In the formula, is a matrix of all 1s, which is used to maintain the structure of the matrix during the combination process;
[0034] When n = 1, is constructed as follows:
[0035] (3).
[0036] As a further limitation of this technical solution, the global feature enhancement module includes layer normalization, a linear layer, depth convolution, a Sigmoid linear unit activation function, a global feature enhancement module encoder, a state space model, a global feature enhancement module decoder, and a residual connection.
[0037] As a further limitation of this technical solution, the specific steps of S32 are as follows:
[0038] Input the local face features obtained after being processed by the residual connection network and the semi-feature pyramid network into two parallel convolution blocks for processing to enhance the local features The extraction formula is as follows:
[0039] (4);
[0040] Where: represents a 1x1 convolution;
[0041] represents a depthwise separable convolution with a convolution kernel size of kxk.
[0042] As a further limitation of this technical solution, the specific steps of S33 are as follows:
[0043] By concatenating along the channel dimension, fuse the enhanced global features and local features together to form fused features, and the final output of this module Obtained through 1×1 2D convolution to restore the channel count to match that of the input and the residual connection, with the formula as follows:
[0044] (5);
[0045] Where: Is the input to the locality-enhanced state space module;
[0046] Concat(G o ,L k5 ,L k7 ) is the fused feature obtained after processing by the locality-enhanced state space module;
[0047] Conv2D 1x1 () is a 1x1 2D convolution.
[0048] As a further limitation of this technical solution, the specific steps of S4 are:
[0049] Input the feature map extracted by the feature-enhanced Mamba framework Into the fully connected layer FC, and classify through the classifier to determine whether the face image is real or forged. In this process, through the cross-entropy loss function L CE , L L1 Loss function and L L2 Loss function guide the model to train;
[0050] The formula of the cross-entropy loss function Is:
[0051] (6);
[0052] Where: N Is the number of samples;
[0053] Is the true label of the i th sample;
[0054] Is the i th sample's predicted probability;
[0055] The formula for the loss function is:
[0056] (7);
[0057] Where: Is the predicted value of the i th sample;
[0058] The calculation formula of the loss function is:
[0059] (8);
[0060] Combining these three loss functions and forming a total loss function through weighted summation L total , during the training process, making the value of the total loss function L total decrease continuously, thereby optimizing the performance of the model and improving the accuracy of classification. The formula of the total loss function L total is:
[0061] (9);
[0062] Where: α , β and γ are the weights of the three loss functions.
[0063] Compared with the prior art, the advantages and positive effects of the present invention are:
[0064] This study proposes an innovative solution based on the Mamba architecture with feature enhancement, and creatively proposes a locality-enhanced state space module on the basis of the original Mamba framework. This module includes a global feature enhancement module and a local feature enhancement module. In the global feature enhancement module, the decoder and encoder adopt the Hilbert scanning method, and scan in different directions to effectively extract the global features of the face image. In the local feature enhancement module, through the processing of the convolution module and the depth convolution module, the local features of the face image are effectively extracted. Finally, the extracted local enhanced features and global enhanced features are fused to fully obtain the feature information of the face image, improving the accuracy of face image forgery detection, effectively maintaining network space security, and maintaining social security and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a schematic diagram of the global feature enhancement module of the present invention.
[0066] Figure 2 It is the Hilbert scanning method of the present invention.
[0067] Figure 3 It is the structural diagram of the feature-enhanced Mamba framework of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0068] The following will describe in detail a specific embodiment of the present invention in conjunction with the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific embodiment.
[0069] The present invention includes the following steps:
[0070] S1: Preparation of the dataset.
[0071] The dataset selected for the present invention is CelebA, where 162,079 face images are used as the training set and 40,520 face images are used as the test set, and all images are adjusted to a size of 256×256;
[0072] S2: Construct a Mamba framework based on feature enhancement.
[0073] This framework introduces a locality-enhanced state space module on the basis of the original Mamba framework. This module includes a global feature enhancement module and a local feature enhancement module; through the processing of the locality-enhanced state space module, the pre-extracted features can be processed to obtain the fused vector of the enhanced global features and local features, further enhancing the feature representation ability, and being able to effectively capture the feature information of face images, thereby improving the accuracy of complex and variable face forgery detection.
[0074] S3: Training of the model.
[0075] First, we input the pictures in the dataset into the residual connection network. The residual connection network extracts feature information of different scales through its multiple convolutional layers and downsampling layers, and then inputs the extracted feature information of different scales into the semi-feature pyramid network to fuse the feature maps of different scales to generate a richer multi-scale feature representation. Then, the output feature map is input into the Mamba framework based on feature enhancement. This framework introduces a locality-enhanced state space module. This module processes the input image through the global feature enhancement module and the local feature enhancement module, further enhancing the feature representation ability.
[0076] The specific image feature extraction process is as follows:
[0077] S31: Feature extraction of the global feature enhancement module.
[0078] The specific steps of S31 are as follows:
[0079] The picture extracts global information after being processed by the residual connection network and the semi-feature pyramid network Through the processing of the global feature enhancement module, enhanced global feature information is obtained ;
[0080] Global information Processed by N global feature enhancement modules to obtain globally enhanced features ; Globally enhanced features The extraction formula is as follows:
[0081] (1);
[0082] Where: n ∈ N represents the number of global feature enhancement modules, and N is usually taken as 2 or 3;
[0083] LN() represents layer normalization;
[0084] SSMs() represents the state space model;
[0085] represents the global feature enhancement module encoder;
[0086] represents the global feature enhancement module decoder;
[0087] represents the activation function;
[0088] represents the depthwise separable convolution with a convolution kernel size of 3x3;
[0089] represents the linear layer;
[0090] represents the global feature enhancement module;
[0091] In the global feature enhancement module, both the global feature enhancement module encoder and the global feature enhancement module decoder use the Hilbert scanning method for scanning; the global feature enhancement module encoder uses the Hilbert scanning method to encode the input feature map into a sequence suitable for processing by the state space model (SSM). The global feature enhancement module decoder then uses the Hilbert scanning method to decode the feature sequence output by the state space model back to the spatial dimension of the original input feature map. This can ensure that the decoded feature map can accurately recover the structure and information of the input feature map while retaining the global modeling ability.
[0092] Hilbert scanning is a space-filling curve that can convert a two-dimensional feature map into a one-dimensional sequence while retaining the correlation of local and global information. The generation formula of the Hilbert curve is as follows:
[0093] (2);
[0094] Where: represents the n - order Hilbert matrix. When n > 1, the Hilbert matrix is constructed by recursively combining four n - order Hilbert matrices , ( transpose), ( left - right flip) and ( up - down flip);
[0095] in the formula, is an all - ones matrix used to maintain the structure of the matrix during the combination process;
[0096] parameters in the formula such as 4n, 4n + 1, etc. are used to adjust the indices and positions of the matrix during the recursive process to ensure that the generated Hilbert curve can correctly fill the space.
[0097] When n = 1, the construction is as follows:
[0098] (3).
[0099] The Hilbert scanning method is as Figure 2 shown.
[0100] Figure 2 shows the Hilbert scanning in 8 directions, Figure 2 from (a) to (h) are forward scanning, reverse scanning, width - height forward scanning, width - height reverse scanning, 90° rotation forward scanning, 90° rotation reverse scanning, width - height 90° rotation forward scanning, width - height 90° rotation reverse scanning respectively.
[0101] The global feature enhancement module includes layer normalization (LN), linear layer, depth convolution, Sigmoid Linear Unit activation function, global feature enhancement module encoder, state space model (SSM), global feature enhancement module decoder and residual connection. The internal structure schematic diagram of the global feature enhancement module is as Figure 1 shown.
[0102] S32: Feature extraction of the local feature enhancement module
[0103] The specific steps of the said S32 are as follows:
[0104] The local face features obtained after being processed by the residual connection network and the semi - feature pyramid network It is input to two parallel convolutional blocks for processing. Each block includes a 1×1 convolutional block (Conv 1x1), a k×k depthwise convolutional block (DWConv kxk), and another 1×1 convolutional block (Conv 1x1), where k is the size of the depthwise convolution kernel. Each convolutional block includes a 2D convolutional layer, an instance normalization 2D layer, and a SiLU activation function. Enhance local features The extraction formula is as follows:
[0105] (4);
[0106] Where: represents a 1x1 convolution;
[0107] represents a depthwise separable convolution with a convolution kernel size of kxk.
[0108] S33: Obtaining the fused features.
[0109] The specific steps of the said S33 are as follows:
[0110] The enhanced global features and local features are fused together through concatenation along the channel dimension to form fused features, and the final output of this module is obtained through a 1×1 2D convolution to restore the channel count to match the channel count of the input and the residual connection. The formula is as follows:
[0111] (5);
[0112] Where: is the input of the locality-enhanced state space module;
[0113] Concat(G o ,L k5 ,L k7 ) is the fused feature obtained after being processed by the locality-enhanced state space module;
[0114] Conv2D 1x1 () is a 1x1 2D convolution.
[0115] S4: Constructing the loss function.
[0116] The specific steps of the said S4 are as follows:
[0117] The feature map extracted by the feature enhancement Mamba framework is input into the fully connected layer FC and classified by a classifier to determine whether the face image is real or forged. In this process, through the cross-entropy loss function L CE 、 LL1 The loss function and L L2 the loss function guides the model training;
[0118] The cross-entropy loss function is used to measure the difference between the predicted class of the model and the true label distribution. For the problem of face forgery detection, the cross-entropy loss function has the formula:
[0119] (6);
[0120] where: N is the number of samples;
[0121] is the true label (0 or 1) of the i th sample;
[0122] is the predicted probability (between 0 and 1) of the i th sample;
[0123] The loss function measures the absolute difference between the predicted classification value and the true value. The calculation formula is:
[0124] (7);
[0125] where: is the predicted value of the i th sample;
[0126] The loss function measures the squared difference between the predicted classification value and the true value. The calculation formula is:
[0127] (8);
[0128] Combining these three loss functions and forming a total loss function through weighted summation L total to continuously reduce the value of the total loss function L total during the training process, thereby optimizing the performance of the model and improving the classification accuracy. The formula for the total loss function L total is:
[0129] (9);
[0130] where: α , β and γ are the weights of the three loss functions.
[0131] The Adam optimizer is adopted as the core algorithm for network training, and the total loss function is minimized through the Adam optimizer L total , enabling the model to gradually learn the discriminative features between real and forged faces during the training process, improving the accuracy of discrimination, and selecting AUC (Area Under the Curve) as the core evaluation index of the model performance to help us determine the optimal model parameter configuration.
[0132] The specific embodiments of the present invention disclosed above are only for illustration, but the present invention is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A face forgery detection method based on the feature-enhanced Mamba framework, characterized in that, It includes the following steps: S1: Preparation of the dataset; S2: Construct a Mamba framework based on feature enhancement; On the basis of the original Mamba framework, this framework introduces a locality-enhanced state space module, which includes a global feature enhancement module and a local feature enhancement module; S3: Training of the model; The specific image feature extraction process is as follows: S31: Feature extraction of the global feature enhancement module; S32: Feature extraction of the local feature enhancement module; S33: Obtaining of the fused features; S4: Construction of the loss function; The specific steps of S31 are as follows: The global information is extracted after processing the picture through the residual connection network and the semi-feature pyramid network It is processed through the global feature enhancement module to obtain the enhanced global feature information ; Global enhanced features The extraction formula is as follows: (1); Where: n ∈ N represents the number of global feature enhancement modules; LN() represents layer normalization; SSMs() represents the state space model; It represents the global feature enhancement module encoder; It represents the decoder of the global feature enhancement module; represents an activation function; It represents a depthwise separable convolution with a convolution kernel size of 3x3; It represents a linear layer; represents the global feature enhancement module; In the global feature enhancement module, both the global feature enhancement module encoder and the global feature enhancement module decoder use the Hilbert scan method for scanning; The generation formula of the Hilbert curve is as follows: (2); Wherein: represents an n - order Hilbert matrix. When n > 1, the Hilbert matrix is constructed by recursively combining four n - order Hilbert matrices , , and ; In the formula is a matrix of all ones, used to maintain the structure of the matrix during the combination process; When n = 1, The structure is as follows: (3)。 2. The face forgery detection method based on the feature-enhanced Mamba framework according to claim 1, wherein: The global feature enhancement module includes layer normalization, linear layer, depth convolution, Sigmoid linear unit activation function, global feature enhancement module encoder, state space model, global feature enhancement module decoder, and residual connection.
3. The face forgery detection method based on the feature-enhanced Mamba framework according to claim 2, wherein: The specific steps of S32 are as follows: The local face features obtained after being processed by the residual connection network and the semi-feature pyramid network are input into two parallel convolutional blocks for processing to enhance the local features The extraction formula is as follows: (4); Wherein: represents a 1x1 convolution; It represents a depthwise separable convolution with a convolution kernel size of kxk.
4. The face forgery detection method based on the feature-enhanced Mamba framework according to claim 3, wherein: The specific steps of S33 are as follows: The enhanced global features and local features are fused together through concatenation along the channel dimension to form fused features, which are the final output of this module. This is obtained through a 1×1 2D convolution to restore the channel count to match that of the input and the residual connection, as shown in the following formula: (5); Wherein: is the input of the local enhancement state space module; Concat(G o ,L k5 ,L k7 ) is the fused feature obtained after being processed by the locality-enhanced state space module; Conv2D 1x1 () is a 2D convolution with a kernel size of 1x1.
5. The face forgery detection method based on the feature-enhanced Mamba framework according to claim 4, wherein: The specific steps of S4 are as follows: The feature map extracted by the feature-enhanced Mamba framework is input into the fully connected layer FC and classified by the classifier to determine whether the face image is real or forged. During this process, the cross-entropy loss function L CE and L L1 the loss function L L2 guide the model to train; Cross-entropy loss function The formula is as follows: (6); Wherein: N is the number of samples; is the i true label of the is the i predicted probability of the nth sample; The calculation formula of the loss function is as follows: (7); Wherein: is the predicted value of the i th sample; The calculation formula of the loss function is as follows: (8); Combining these three loss functions, a total loss function is formed by weighted summation L total , during the training process, making the value of the total loss function L total constantly decrease, thereby optimizing the performance of the model and improving the accuracy of classification. The formula for the total loss function L total is as follows: (9); Wherein: α , β and γ are the weights of three loss functions.
Citation Information
Patent Citations
Face forgery detection method based on multi-feature fusion network
CN118015714A
Underwater image enhancement method based on Mama multi-feature enhancement fusion
CN119809951A