Face forgery detection method based on feature enhancement Mama framework
By introducing the local enhancement state space module of the feature enhancement Mamba framework in face forgery detection, and using Hilbert scanning and convolution modules to extract features, the problem of difficult to model long-range dependencies and high computational complexity in face feature extraction in the prior art is solved, and face forgery detection with high accuracy and low resource consumption is achieved.
Patent Information
- Application Number
- CN202510585552.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing face forgery detection methods are difficult to model long-range dependencies when extracting face features, resulting in errors in detection results; while the high computational complexity of the Transformer framework leads to excessive time-consuming forgery detection, which increases hardware costs and may lead to system crashes.
A face forgery detection method based on feature enhancement Mamba framework is proposed. By introducing a local enhancement state space module, including a global feature enhancement module and a local feature enhancement module, the Hilbert scanning method is used to extract global features, and local features are extracted through the convolution module, and finally fuse them to improve detection accuracy.
This method can accurately capture the global and local features of the image, improve the accuracy of face forgery detection, effectively identify forged face images, reduce system resource consumption, and ensure stable operation of the system.
Smart Images

Figure CN120088838A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically, to a face forgery detection method based on a feature-enhanced Mamba framework. Background Art
[0002] The development of deep learning has made face forgery technology increasingly mature. Through technologies such as generative adversarial networks, realistic face images can be generated. The forged content is not only used to create pornographic videos and false news, but may also be used for political attacks and financial fraud, etc., posing a huge threat to fields such as politics, security, transportation, and finance. Therefore, how to identify forged face images and maintain social security and public order has become an urgent problem to be solved.
[0003] Although existing face forgery detection methods have achieved good results, when extracting face features, convolutional neural networks still have the problem of being difficult to model long-range dependence relationships. The inability to model long-range dependence relationships will cause the detection model to be unable to fully utilize global correlation information during the training process, resulting in incorrect detection results. The Transformer framework has the problem that the high computational complexity leads to excessive time consumption for forgery detection. The high computational complexity not only increases the hardware cost, but may also cause the system to malfunction under limited resources. At the same time, it becomes more difficult to process large-scale data, which may lead to slow system operation or even crashes. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a face forgery detection method based on a feature-enhanced Mamba framework, apply the Mamba framework to the field of deepfake detection, and make an innovative transformation of the original framework. On this basis, a feature-enhanced Mamba framework is proposed. This framework can accurately capture the global and local features of images, achieving good results in complex and variable face forgery detection and being able to effectively identify forged face images.
[0005] The present invention adopts the following technical solutions to achieve the invention purpose: A face forgery detection method based on a feature-enhanced Mamba framework, characterized by including the following steps: S1: Preparation of the data set; S2: Construct a feature-enhanced Mamba framework; This framework introduces a locality-enhanced state space module on the basis of the original Mamba framework. This module includes a global feature enhancement module and a local feature enhancement module; S3: Training of the model; The specific image feature extraction process is as follows: S31: Feature extraction of the global feature enhancement module; S32: Feature extraction of the local feature enhancement module; S33: Obtaining the fused features; S4: Constructing the loss function.
[0006] As a further limitation of this technical solution, the specific steps of S31 are as follows: After processing the picture through the residual connection network and the semi-feature pyramid network, global information is extracted Processed by the global feature enhancement module to obtain the enhanced global feature information ; Global enhanced features The extraction formula is as follows: (1); Where: n ∈ N represents the number of global feature enhancement modules; LN() represents layer normalization; SSMs() represents the state space model; represents the global feature enhancement module encoder; represents the global feature enhancement module decoder; represents the activation function; represents the depthwise separable convolution with a convolution kernel size of 3x3; represents the linear layer; represents the global feature enhancement module; In the global feature enhancement module, both the global feature enhancement module encoder and the global feature enhancement module decoder use the Hilbert scanning method for scanning; The generation formula of the Hilbert curve is as follows: (2); Where: represents the nth-order Hilbert matrix. When n > 1, the Hilbert matrix is constructed by recursively combining four nth-order Hilbert matrices , , and ; In the formula, is a matrix of all 1s, used to maintain the structure of the matrix during the combination process; When n = 1, is constructed as follows: (3).
[0007] As a further limitation of this technical solution, the global feature enhancement module includes layer normalization, a linear layer, depth convolution, a Sigmoid linear unit activation function, a global feature enhancement module encoder, a state space model, a global feature enhancement module decoder, and a residual connection.
[0008] As a further limitation of this technical solution, the specific steps of S32 are as follows: The local face features obtained after being processed by the residual connection network and the semi-feature pyramid network are input into two parallel convolutional blocks for processing to enhance the local features The extraction formula is as follows: (4); Where: represents a 1x1 convolution; represents a depthwise separable convolution with a convolution kernel size of kxk.
[0009] As a further limitation of this technical solution, the specific steps of S33 are as follows: The enhanced global features and local features are fused together through concatenation along the channel dimension to form fused features, and the final output of this module is obtained through a 1×1 2D convolution to restore the channel count to match the channel counts of the input and the residual connection. The formula is as follows: (5); Where: is the input of the locality-enhanced state space module; Concat(G o ,L k5 ,L k7 ) is the fused feature obtained after being processed by the locality-enhanced state space module; Conv2D 1x1 () is a 1x1 2D convolution.
[0010] As a further limitation of this technical solution, the specific steps of S4 are as follows: The feature map extracted by the feature enhancement Mamba framework is input into the fully connected layer FC and classified by a classifier to determine whether the face image is real or forged. During this process, through the cross-entropy loss function L CE , L L1 loss function and LL2 The loss function guides the training of the model; Cross-entropy loss function The formula is: (6); Where: N is the number of samples; is the true label of the i th sample; is the i th sample's predicted probability; The calculation formula of the loss function is: (7); Where: is the predicted value of the i th sample; The calculation formula of the loss function is: (8); Combining these three loss functions and forming a total loss function through weighted summation L total , during the training process, making the value of the total loss function L total decrease continuously, thereby optimizing the performance of the model and improving the classification accuracy. The formula of the total loss function L total is: (9); Where: α , β and γ are the weights of the three loss functions.
[0011] Compared with the prior art, the advantages and positive effects of the present invention are: This study presents an innovative solution based on the Mamba architecture with feature enhancement. On the basis of the original Mamba framework, a locality-enhanced state space module is creatively proposed. This module includes a global feature enhancement module and a local feature enhancement module. In the global feature enhancement module, the decoder and encoder adopt the Hilbert scanning method, scanning in different directions to effectively extract the global features of the face image. In the local feature enhancement module, through the processing of the convolutional module and the depth convolutional module, the local features of the face image are effectively extracted. Finally, the extracted local enhanced features and global enhanced features are fused to fully obtain the feature information of the face image, improving the accuracy of face image forgery detection, effectively maintaining network space security, and safeguarding social security and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 FIG. is a schematic diagram of the global feature enhancement module of the present invention.
[0013] Figure 2 FIG. is the Hilbert scanning method of the present invention.
[0014] Figure 3 FIG. is a structural diagram of the feature-enhanced Mamba framework of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The following combines the drawings to describe in detail a specific embodiment of the present invention, but it should be understood that the protection scope of the present invention is not limited by the specific embodiment.
[0016] The present invention includes the following steps: S1: Preparation of the data set.
[0017] The data set selected for the present invention is CelebA, where 162,079 face images are used as the training set and 40,520 face images are used as the test set, and all images are adjusted to a size of 256×256; S2: Construct a Mamba framework based on feature enhancement.
[0018] This framework introduces a locality-enhanced state space module on the basis of the original Mamba framework. This module includes a global feature enhancement module and a local feature enhancement module; through the processing of the locality-enhanced state space module, the pre-extracted features are processed to obtain a fused vector of enhanced global features and local features, further enhancing the feature representation ability, and being able to effectively capture the feature information of the face image, thereby improving the accuracy of complex and variable face forgery detection.
[0019] S3: Training of the model.
[0020] First, we input the images in the dataset into the residual connection network. Through its multiple convolutional layers and downsampling layers, the residual connection network extracts feature information at different scales. Then, the feature information at different scales extracted is input into the semi-feature pyramid network to fuse the feature maps at different scales to generate a richer multi-scale feature representation. Next, the output feature map is input into the Mamba framework based on feature enhancement, which introduces a locality-enhanced state space module. This module processes the input image through a global feature enhancement module and a local feature enhancement module, further enhancing the feature representation ability.
[0021] The specific image feature extraction process is as follows: S31: Feature extraction of the global feature enhancement module.
[0022] The specific steps of S31 are as follows: After the image is processed by the residual connection network and the semi-feature pyramid network, global information is extracted It is processed through the global feature enhancement module to obtain enhanced global feature information ; Global information Is processed by N global feature enhancement modules to obtain global enhanced features ; Global enhanced features The extraction formula is as follows: (1); Where: n ∈ N represents the number of global feature enhancement modules, and N is usually taken as 2 or 3; LN() represents layer normalization; SSMs() represents the state space model; Represents the global feature enhancement module encoder; Represents the global feature enhancement module decoder; Represents the activation function; Represents the depthwise separable convolution with a convolution kernel size of 3x3; Represents the linear layer; Represents the global feature enhancement module; In the global feature enhancement module, both the global feature enhancement module encoder and the global feature enhancement module decoder use the Hilbert scanning method for scanning; the global feature enhancement module encoder uses the Hilbert scanning method to encode the input feature map into a sequence suitable for processing by the state space model (SSM). The global feature enhancement module decoder then uses the Hilbert scanning method to decode the feature sequence output by the state space model back to the spatial dimension of the original input feature map. This can ensure that the decoded feature map can accurately recover the structure and information of the input feature map while retaining the global modeling ability.
[0023] Hilbert scanning is a space-filling curve that can convert a two-dimensional feature map into a one-dimensional sequence while preserving the correlation of local and global information. The generation formula of the Hilbert curve is as follows: (2); Where: represents the nth-order Hilbert matrix. When n > 1, the Hilbert matrix is constructed by recursively combining four nth-order Hilbert matrices , ( transposed), ( flipped left and right), and ( flipped up and down); in the formula, is a matrix of all 1s, which is used to maintain the structure of the matrix during the combination process; parameters in the formula such as 4n, 4n + 1, etc. are used to adjust the index and position of the matrix during the recursive process to ensure that the generated Hilbert curve can correctly fill the space.
[0024] When n = 1, is constructed as follows: (3).
[0025] The Hilbert scanning method is as Figure 2 shown.
[0026] Figure 2 shows the Hilbert scanning in 8 directions, Figure 2 from (a) to (h) are forward scanning, reverse scanning, width-height forward scanning, width-height reverse scanning, 90° clockwise forward scanning, 90° clockwise reverse scanning, width-height 90° clockwise forward scanning, and width-height 90° clockwise reverse scanning respectively.
[0027] The global feature enhancement module includes layer normalization (LN), a linear layer, depth convolution, a Sigmoid Linear Unit activation function, a global feature enhancement module encoder, a state space model (SSM), a global feature enhancement module decoder, and a residual connection. The schematic diagram of the internal structure of the global feature enhancement module is as shown in Figure 1 shown.
[0028] S32: Feature extraction of the local feature enhancement module
[0029] The specific steps of S32 are as follows: The local face features obtained after being processed by the residual connection network and the semi-feature pyramid network are input into two parallel convolutional blocks for processing. Each block includes a 1×1 convolutional block (Conv 1x1), a k×k depth convolutional block (DWConv kxk), and another 1×1 convolutional block (Conv 1x1), where k is the size of the depth convolution kernel. Each convolutional block includes a 2D convolutional layer, an instance norm 2D normalization layer, and a SiLU activation function. Enhance the local features The extraction formula is as follows: (4); Where: represents a 1x1 convolution; represents a depthwise separable convolution with a convolution kernel size of kxk.
[0030] S33: Obtaining the fused features.
[0031] The specific steps of S33 are as follows: The enhanced global features and local features are fused together by concatenation along the channel dimension to form fused features. The final output of this module is obtained through a 1×1 2D convolution to restore the channel count to match the channel counts of the input and the residual connection. The formula is as follows: (5); Where: is the input of the locality-enhanced state space module; Concat(G o ,L k5 ,L k7 ) is the fused feature obtained after being processed by the locality-enhanced state space module; Conv2D 1x1 () is a 1x1 2D convolution.
[0032] S4: Construction of the loss function.
[0033] The specific steps of S4 are as follows: Input the feature map extracted by the feature-enhanced Mamba framework into the fully connected layer FC, and classify it through a classifier to determine whether the face image is real or forged. In this process, through the cross-entropy loss function L CE and L L1 the loss function and L L2 the loss function guide the model to train; The cross-entropy loss function is used to measure the difference between the category predicted by the model and the true label distribution. For the problem of face forgery detection, the cross-entropy loss function has the formula: (6); Where: N is the number of samples; is the true label (0 or 1) of the i th sample; is the predicted probability (between 0 and 1) of the i th sample; The loss function measures the absolute difference between the predicted classification value and the true value. The calculation formula is: (7); Where: is the predicted value of the i th sample; The loss function measures the squared difference between the predicted classification value and the true value. The calculation formula is: (8); Combine these three loss functions and form a total loss function through weighted summation L total , and make the value of the total loss function L total decrease continuously during the training process, so as to optimize the performance of the model and improve the classification accuracy. The formula of the total loss function L total is: (9); Where: α , β and γ are the weights of the three loss functions.
[0034] The Adam optimizer is adopted as the core algorithm for network training, and the total loss function is minimized through the Adam optimizer L total , enabling the model to gradually learn the discriminative features of real and forged faces during training, improving the accuracy of discrimination, and selecting AUC (Area Under the Curve) as the core evaluation index of the model performance to help us determine the optimal model parameter configuration.
[0035] The specific embodiments of the present invention disclosed above are only for illustration, but the present invention is not limited thereto. Any changes that can be conceived by those skilled in the art shall fall within the protection scope of the present invention.
Claims
1. A face forgery detection method based on feature-enhanced Mamba framework, characterized in that: The following steps are involved: S1: Dataset preparation; S2: Constructing the Mamba framework based on feature enhancement; The framework introduces a local enhancement state space module based on the original Mamba framework, which includes a global feature enhancement module and a local feature enhancement module. S3: Model training; The specific image feature extraction process is as follows: S31: Feature extraction of global feature enhancement module; S32: Feature extraction of local feature enhancement module; S33: Acquisition of fusion features; S4: Loss function construction.
2. The method for detecting face forgery based on feature-enhanced Mamba framework according to claim 1, characterized in that: The specific steps of S31 are: The image is processed by the residual connection network and the semi-feature pyramid network to extract the global information Processed by the global feature enhancement module, the enhanced global feature information is obtained ; Global Enhancement Features The extraction formula is as follows: (1); Where: n∈N represents the number of global feature enhancement modules; LN() stands for layer normalization; SSMs() represents state space models; It represents the global feature enhancement module encoder; It represents the decoder of the global feature enhancement module; represents the activation function; It represents a depth-separable convolution with a kernel size of 3x3; represents the linear layer; It represents the global feature enhancement module; In the global feature enhancement module, both the global feature enhancement module encoder and the global feature enhancement module decoder are scanned using the Hilbert scanning method; The formula for generating the Hilbert curve is as follows: (2); in: It represents the n-order Hilbert matrix. When n>1, the Hilbert matrix is obtained by recursively combining four n-order Hilbert matrices , , and to construct; In the formula is an all-1 matrix used to maintain the structure of the matrix during the combination process; When n=1, The construction is as follows: (3)。 3. The face forgery detection method based on feature-enhanced Mamba framework according to claim 2 is characterized in that: The global feature enhancement module includes layer normalization, linear layer, depth convolution, Sigmoid linear unit activation function, global feature enhancement module encoder, state space model, global feature enhancement module decoder and residual connection.
4. The method for detecting face forgery based on feature-enhanced Mamba framework according to claim 2, characterized in that: The specific steps of S32 are: The local features of the face obtained after processing by the residual connection network and the semi-feature pyramid network Input to two parallel convolution blocks for processing to enhance local features The extraction formula is as follows: (4); in: It represents a 1x1 convolution; It represents a depth-wise separable convolution with a kernel size of kxk.
5. The method for detecting face forgery based on feature-enhanced Mamba framework according to claim 4, characterized in that: The specific steps of S33 are: By concatenating along the channel dimension, the enhanced global features and local features are fused together to form fused features. The final output of this module is Obtained by a 1×1 2D convolution to restore the channel count to match the channel count of the input and residual connection, the formula is as follows: (5); in: Enhance the input of the state-space module for locality; Concat(G o ,L k5 ,L k7 ) is the fusion feature obtained after being processed by the local enhancement state space module; Conv2D 1x1 () is a 1x1 2D convolution.
6. The method for detecting face forgery based on feature-enhanced Mamba framework according to claim 4, characterized in that: The specific steps of S4 are: Feature enhancement Mamba framework extracts the feature map The input is sent to the fully connected layer FC and classified by the classifier to determine whether the face image is real or fake. In this process, the cross entropy loss function is used L CE , L L1 The loss function and L L2 The loss function guides the model to train; Cross Entropy Loss Function The formula is: (6); in: N is the sample size; It is i The true labels of samples; It is i The predicted probability of samples; The loss function calculation formula is: (7); in: It is i The predicted value of samples; The loss function calculation formula is: (8); These three loss functions are combined to form a total loss function by weighted summation L total , so that the total loss function during training L total The value of is continuously reduced, thereby optimizing the performance of the model and improving the accuracy of classification. The total loss function L total The formula is: (9); in: α , β and γ are the weights of the three loss functions.
Citation Information
Patent Citations
Face-changing video detection method and system based on overall counterfeit trace and local detail information extraction
CN117935381A
Face forgery detection method based on multi-feature fusion network
CN118015714A
Face forgery detection method based on local forgery area detection
CN118135641A
Crowd counting method based on adaptive global perception and multi-scale feature fusion
CN119048993A
Micro-expression recognition method and system based on state space model and double-flow fusion
CN119649427A
Cited By
Deep forgery detection method and device for HSIC feature alignment
CN121305654A
Deepfake detection method and apparatus using hsic feature alignment
CN121305654B
Face forgery detection method based on watermark feature assistance and cross-task distillation
CN121747207A
A Face Spoofing Detection Method Based on Watermark Feature Assistance and Cross-Task Distillation
CN121747207B
Steel bar detection model, method and system
CN122073010A