3D mask face presentation attack detection method based on micro expression
By constructing a 3D mask face presentation attack detection method based on micro-expression, using video data preprocessing, texture and micro-expression feature extraction network, combined with multi-scale differential convolution and feature fusion module, the accuracy and robustness of 3D mask face presentation attack detection is solved, and the accuracy and robustness of detection are improved.
Patent Information
- Application Number
- CN202510451593.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively detect 3D mask face presentation attacks, especially in terms of texture details and depth information, and is poorly robust, making it difficult to deal with deceptive attacks under different lighting and perspectives.
A 3D mask face presentation attack detection method based on micro-expression is constructed, and a network is extracted through video data preprocessing, texture features and micro-expression feature, combined with a multi-scale differential convolutional network and feature fusion module, and a multi-branch collaborative optimization loss function is used for training to improve detection accuracy and robustness.
It improves the accuracy and robustness of 3D mask face presentation attack detection, ensuring the security of the face recognition system, especially in complex lighting and material conditions.
Smart Images

Figure CN120340143A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method for detecting 3D mask face presentation attacks based on microexpressions. Background Art
[0002] Face presentation attack is a technology that presents a face in front of a camera through media such as photos and videos to deceive a face recognition system. In recent years, with the rapid development of deepfake technology, the attack means have become increasingly realistic, posing a serious threat to the security of face recognition systems. Among them, 3D mask attacks have become a more challenging attack method due to their real three-dimensional structure and more realistic deception effects. Existing face presentation attack detection methods mostly rely on the deep supervised features extracted by convolutional networks, but these methods often ignore fine-grained information and the mutual connection between depth information and texture information.
[0003] There are significant differences between real faces and 3D mask attacks in texture details and depth structures. For example, 3D mask attacks usually lack the dynamic microexpression features of real faces and show specific deception traces in material reflection, facial details (such as wrinkles around the eyes and micro-textures of the nose wings), and three-dimensional structures. However, current deep learning networks mostly focus on extracting high-order semantic information and pay insufficient attention to these key fine-grained features, resulting in limited accuracy in detecting 3D mask attacks. In addition, the deception mode of 3D mask attacks is highly robust, and their performance under different lighting conditions and viewing angles is more deceptive, further increasing the detection difficulty. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention proposes a method for detecting 3D mask face presentation attacks based on microexpressions, aiming to improve the accuracy and robustness of 3D mask face presentation attack detection, thereby ensuring the security of face recognition systems.
[0005] The present invention adopts the following technical solutions to solve the technical problems:
[0006] A method for detecting 3D mask face presentation attacks based on microexpressions according to the present invention is characterized by including the following steps:
[0007] Step 1, obtain a video data set and perform preprocessing to obtain the k-th dual-channel input matrix of any video sample E ;
[0008] Step 2, construct a texture feature extraction sub-network, including: a shallow face feature extraction unit and a texture feature map generation unit, and process to obtain the k-th texture feature map ;
[0009] Step 3. Construct a micro-expression feature extraction network, including: a spatial attention module, a multi-level attention module, a feature aggregation module, and an adaptive attention feature fusion module, and process to obtain the k-th micro-expression feature map ;
[0010] Step 3.1. The spatial attention module processes to obtain the k-th spatial feature map ;
[0011] Step 3.2. The multi-level attention module processes to obtain the k-th multi-level fusion feature map ;
[0012] Step 3.3. The feature aggregation module processes and to obtain the k-th aggregated feature map ;
[0013] Step 3.4. The adaptive attention feature fusion module processes , and to obtain the k-th micro-expression feature map ;
[0014] Step 4. Construct a joint feature extraction sub-network, including: a multi-scale differential convolution network and a feature fusion module, and process and to obtain the k-th joint feature map , thereby obtaining the joint feature map set of the video sample E ; is the number of feature matrices;
[0015] Step 5. Construct a total loss function , which is used to train the detection network composed of the texture feature extraction sub-network, the micro-expression feature extraction network and the joint feature extraction sub-network, and obtain a 3D mask face presentation attack detection model with optimal parameters to realize 3D mask face presentation attack detection.
[0016] The feature of the 3D mask face presentation attack detection method based on micro-expression according to the present invention also lies in that the step 1 includes the following steps:
[0017] Step 1.1. Obtain a video data set, and set the true class label of any video sample E in the video data set as , represents the video sample It is a 3D mask face presentation attack video, indicating a video sample is a real face video;
[0018] Step 1.2: Generate a feature matrix from the video sample E , where represents the th feature matrix in is the number of rows of the feature matrix, is the number of columns of the feature matrix; is the number of feature matrices;
[0019] Step 1.3: Decompose the th feature matrix into the th actual feature matrix and the th device noise matrix , thereby constructing the th objective function using Equation (1) and optimizing it using the alternating direction method of multipliers to obtain the th optimal actual feature matrix :
[0020] (1)
[0021] In Equation (1), is the rank of; represents 0 - norm of; is the th regularization term coefficient;
[0022] Step 1.4: Convert from the RGB color space to the CHROM color space and generate the th first chrominance component and the th second chrominance component , thereby forming the th two - channel input matrix .
[0023] Furthermore, the said Step 2 includes the following steps:
[0024] Step 2.1: The shallow face feature extraction unit performs a convolution operation on to obtain the k - th shallow face feature matrix ;
[0025] Step 2.2: The texture feature map generation unit includes S CDC convolution blocks;
[0026] Input it into the s-th CDC convolution block for processing to obtain the s-th face feature matrix ; thus, S face feature matrices are obtained, and after splicing the S face feature matrices along the channels, the k-th texture feature map is formed . .
[0027] Furthermore, the s-th CDC convolution block in step 2.2 obtains the s-th face feature matrix according to the following process :[[]]
[0028] Step 2.2.1: Perform a convolution operation on to obtain the s-th ordinary convolution feature ;
[0029] Step 2.2.2: Perform a central difference convolution operation on to obtain the s-th central difference convolution feature ;
[0030] Step 2.2.3: Add and element-wise to generate the s-th face feature matrix .
[0031] Furthermore, step 3.1 includes:
[0032] Step 3.1.1: Perform a linear transformation on , , through three learnable linear transformation matrices to respectively generate the k-th query matrix , the k-th key matrix and the k-th value matrix ;
[0033] Step 3.1.2: Reshape and along the height and width dimensions respectively, such that each channel of and is flattened horizontally, and then extract the k-th horizontal query matrix and the k-th horizontal key matrix from the two flattened matrices respectively;
[0034] Step 3.1.3: Obtain the k-th horizontal attention weight matrix using Equation (2):
[0035] (2)
[0036] In formula (2), T represents transpose;
[0037] Step 3.1.4: Obtain the k-th horizontal attention score using formula (2) :
[0038] (3)
[0039] In formula (3), Softmax represents the activation function; represents the scaling factor;
[0040] Step 3.1.5: Obtain the k-th horizontal direction attention weight using formula (4) :
[0041] (4)
[0042] In formula (4), represents element-wise multiplication, is the k-th horizontal distance matrix;
[0043] Step 3.1.7: According to the process of Step 3.1.2 - Step 3.1.5, extract the k-th vertical query and from and the k-th vertical key , thereby calculating the k-th vertical attention weight matrix , the k-th vertical attention score , the k-th vertical distance matrix , and the k-th vertical direction attention weight ;
[0044] Step 3.1.8: After performing matrix multiplication on and , obtain the k-th intermediate result ; After performing matrix multiplication on and , obtain the k-th spatial feature .
[0045] Furthermore, the said Step 3.2 includes:
[0046] Step 3.2.1: Perform convolution operation on to obtain the k-th convolution feature map ;
[0047] Step 3.2.2: Evenly divide along the channel axis into n sub-feature maps, where the i-th sub-feature map is denoted as , , each sub - feature map has the same spatial dimension, and the number of channels of each sub - feature map is reduced to ;
[0048] Step 3.2.3, perform a convolution operation on to compress the number of channels and output the \(i\) - th compressed feature map ;
[0049] Step 3.2.4, after splicing \(n\) compressed feature maps along the channel axis, generate the \(k\) - th multi - level fusion feature map .
[0050] Furthermore, the said Step 3.3 includes:
[0051] Step 3.3.1, obtain the \(k\) - th fusion feature map using Equation (5):
[0052] (5)
[0053] In Equation (5), and are two weight coefficients to be learned, and + = 1, ;
[0054] Step 3.3.2, perform a convolution operation on to compress the number of channels and obtain the \(k\) - th feature map ; apply a convolution with a kernel of 1×1 to for per - channel spatial feature extraction to obtain the \(k\) - th spatial depth feature map ; then perform a convolution operation on to compress the number of channels to the target dimension and obtain the \(k\) - th aggregated feature .
[0055] Furthermore, the said Step 3.4 includes:
[0056] Step 3.4.1, after connecting and along the channel dimension, form the \(k\) - th channel - fused feature map , and then perform a convolution operation on to compress it to the target dimension and obtain the \(k\) - th target feature map ;
[0057] Step 3.4.2, after connecting and along the channel dimension, obtain the \(k\) - th new feature map ;
[0058] Step 3.4.3: Perform average pooling and max pooling operations on respectively, and obtain the k-th average pooling feature and the k-th max pooling feature ; Pass respectively through the Sigmoid function and for normalization to obtain the attention weight of the k-th average pooling feature and the attention weight of the k-th max pooling feature ;
[0059] Step 3.4.4: Obtain the k-th weighted feature map using Equation (6):
[0060] (6)
[0061] In Equation (6), denotes element-wise multiplication;
[0062] Step 3.4.5: Perform a convolution operation on to compress its number of channels and obtain the k-th micro-expression feature map .
[0063] Furthermore, the said Step 4 includes:
[0064] Step 4.1: Concatenate and along the channel dimension to form the k-th preliminary joint feature map ;
[0065] Step 4.2: The multi-scale differential convolution network processes using multiple CDC convolution blocks, connects the obtained features at different levels together, and thus forms the k-th multi-scale differential feature map ;
[0066] Step 4.3: The feature fusion module performs a convolution operation on to obtain the k-th convolutionally compressed feature map ;
[0067] Step 4.4: Calculate the k-th spatial position weight using Equation (7):
[0068] (7)
[0069] In Equation (7), denotes the global average pooling operation, denotes the max pooling operation, Expressed as the Sigmoid function, is a convolution operation with a convolution kernel of 1×1;
[0070] Step 4.5, Multiply element-wise with to obtain the feature at the k-th spatial position ;
[0071] Step 4.6, Use several convolution kernels of different scales to perform convolution operations on , output feature maps corresponding to different scales and fuse them crosswise to generate the k-th joint feature map .
[0072] Furthermore, the said Step 5 includes:
[0073] Step 5.1, Concatenate in the channel dimension to form the global feature map ;
[0074] Step 5.2, Input into the fully connected layer for processing, and output the predicted class probability value of E , where represents the probability that E is a 3D mask attack, represents the probability that E is a real human face, and satisfies ;
[0075] Step 5.3, Use Equation (8) to construct the cross-entropy loss :
[0076] (8)
[0077] In Equation (8), represents the predicted class probability value of the i-th type, represents the true class label of the i-th type;
[0078] Step 5.4, Use Equation (9) to construct the center loss :
[0079] (9)
[0080] In Equation (9), represents the class center matrix of Y and is initialized to the mean value of; is the Frobenius norm;
[0081] Step 5.5, Use Equation (10) to construct the total loss function :
[0082] (10)
[0083] In formula (10), is the dynamic balance coefficient;
[0084] Step 5.6: Process the detection network using the backpropagation algorithm and continuously perform forward propagation to calculate the total loss , and update the parameters and class centers of the detection network through backpropagation until the total loss converges or reaches the preset maximum number of iterations to obtain a 3D mask face presentation attack detection model with optimal parameters.
[0085] Compared with the existing technology, the beneficial effects of the present invention are as follows:
[0086] 1. The present invention constructs a video dataset preprocessing module. Through innovative data preprocessing strategies, the data quality is significantly improved. In step 1.3, the feature matrix is decomposed and optimized using the alternating direction method of multipliers to effectively remove device noise and provide pure data for model training; in step 1.4, the feature matrix is converted from the RGB color space to the CHROM color space to generate a two-channel input matrix, reducing light interference and enhancing the expression of human skin color information. These operations specifically solve the problems that traditional technologies are vulnerable to noise and light at the data level and lay a solid foundation for subsequent analysis.
[0087] 2. The present invention constructs an innovative multi-modal feature extraction and fusion mechanism: the CDC convolutional block in the texture feature extraction sub-network combines ordinary convolution and central difference convolution to more finely extract texture details; the micro-expression feature extraction network uses a multi-module collaborative mechanism to deeply mine micro-expression information. On this basis, the joint feature extraction sub-network effectively fuses texture and micro-expression features, giving play to the complementary advantages of both and providing rich information for the model to distinguish 3D masks and real faces, which is difficult to achieve by traditional technologies.
[0088] 3. The present invention constructs a loss function optimized by multi-branch collaboration: by combining cross-entropy loss and center loss to construct a total loss function, the cross-entropy loss ensures the classification accuracy of the model, and the center loss enhances the intra-class feature compactness. The two cooperate with each other through the dynamic balance coefficient, effectively improving the classification performance and generalization ability of the model, breaking the limitations of traditional single-loss function training, and improving the detection accuracy and robustness of the model against 3D mask attacks, especially in complex lighting and material conditions. Description of the Drawings
[0089] Figure 1 is the overall flowchart of the present invention;
[0090] Figure 2 is the network framework diagram of the present invention. Detailed Embodiments
[0091] In this embodiment, a method for detecting 3D mask face presentation attacks based on microexpressions first obtains samples from a multi-modal video dataset and generates a feature matrix; then constructs a texture feature extraction network to extract texture features, constructs a microexpression feature network to extract microexpression features, and constructs a joint feature extraction module to fuse texture features and microexpression features; constructs a loss function with multi-branch collaborative optimization, and uses the Adam optimizer to train and optimize model parameters; finally, inputs the video to be tested to test the model to ensure that the model can effectively distinguish between real faces and presented attack faces; specifically, referring to Figure 1 , this method is carried out according to the following steps:
[0092] Step 1. Obtain a video dataset and perform preprocessing to obtain the th two-channel input matrix of any video sample E , which specifically includes the following steps:
[0093] Step 1.1. Obtain a video dataset, select the publicly available CASIA-SURF 3DMask dataset, which contains 50 subjects, covering two types of attack types: 3D silicone masks and resin masks, with a total of 1,200 video clips. Each sample contains three-modal data of visible light, depth map, and near-infrared, unified the resolution to 256×256, and the frame rate is 30fps; divide the dataset according to the subject ID: 35 people in the training set (840 videos), 5 people in the validation set (120 videos), and 10 people in the test set (240 videos); set the true class label of any video sample E in the video dataset as , indicating that the video sample is a 3D mask face presentation attack video, indicating that the video sample is a real face video. The reason for selecting this dataset is that it contains various attack types and multi-modal data, which can fully train the model's recognition ability for different situations. Dividing the dataset is for training, validating, and testing the model respectively to ensure the generalization performance of the model. In practical applications, if there are other similar publicly available datasets or high-quality datasets collected by oneself, they can also be divided and used in a similar manner.
[0094] Step 1.2. Extract the face region from video sample E every 10 frames, detect 68-point landmarks for alignment and cropping through OpenFace 2.0, with the output size of 256×256, and generate a feature matrix , where represents the th feature matrix in is the number of rows of the feature matrix, is the number of columns of the feature matrix; is the number of feature matrices; taking the extraction of the face region every 10 frames as an example, it comprehensively considers the computing resources and the frequency of micro-expression changes, which can not only capture the micro-expression information but also not generate too much redundant data. In actual operation, the extraction frequency can be appropriately adjusted according to the video content and hardware conditions.
[0095] Step 1.3, decompose the th feature matrix into the th actual feature matrix and the th device noise matrix , so as to construct the th objective function using Equation (1), and use the alternating direction multiplier method for optimization to obtain the th optimal actual feature matrix :
[0096] (1)
[0097] In Equation (1), is the rank of , which is used to reduce the irrelevant noise components in L; represents the 0-norm of, which is used to ensure the sparsity of N; is the th regularization term coefficient. In this embodiment, is set to 0.1. In actual calculation, by continuously iterating the alternating direction multiplier method, the objective function is gradually optimized to remove noise, making the obtained actual feature matrix purer, which is beneficial to subsequent feature extraction and analysis. If the value of is changed, it will affect the degree of noise removal and the retention of the feature matrix. For example, increasing may remove more noise, but it may also lose some useful features, and it needs to be adjusted according to the experimental results.
[0098] Step 1.4, convert from the RGB color space to the CHROM color space, and generate the th first chrominance component and the th second chrominance component , so as to form the th two-channel input matrix . The RGB color space is greatly affected by light. Converting to the CHROM color space can highlight the chrominance information of the human face and reduce light interference. In actual applications, this conversion can make the model more stable in obtaining face features under different lighting conditions and improve the detection accuracy.
[0099] Step 2: Refer to Figure 2 , construct a texture feature extraction sub-network, including: a shallow face feature extraction unit and a texture feature map generation unit, and process to obtain the k-th texture feature map , the specific steps are as follows:
[0100] Step 2.1: The shallow face feature extraction unit performs a convolution operation on with a convolution kernel of 7×7, a stride of 2, a padding of 3, and a ReLU activation function to obtain the k-th shallow face feature matrix ;
[0101] Selecting a convolution kernel of 7×7, a stride of 2, and a padding of 3 is to cover a larger receptive field when extracting features, obtain richer local information, and control the amount of computation and the size of the feature map. The ReLU activation function can increase the non-linear expression ability of the network, enabling the model to learn more complex feature relationships. In practical applications, changing the convolution kernel size, stride, and padding value will affect the feature extraction effect and computational efficiency, and need to be optimized through experiments.
[0102] Step 2.2: The texture feature map generation unit includes S CDC convolution blocks, including the following steps:
[0103] Input into the s-th CDC convolution block for processing to obtain the s-th face feature matrix ; thus obtaining S face feature matrices, and through the method of concatenating along the channel, the S face feature matrices are concatenated to form the k-th texture feature map , the specific process is as follows:
[0104] Step 2.2.1: Perform a convolution operation on to obtain the s-th ordinary convolution feature . Ordinary convolution can extract basic texture features of the image, such as edges, contours, etc.
[0105] Step 2.2.2: Perform a central difference convolution operation on to obtain the s-th central difference convolution feature . Central difference convolution can highlight the local changes of pixels in the image and capture subtle texture features more sensitively.
[0106] Step 2.2.3: Add and element-wise to generate the s-th face feature matrix ;
[0107] This fusion method combines the advantages of ordinary convolution and central difference convolution, enabling the obtained face feature matrix to contain richer texture information. After obtaining S face feature matrices, through channel concatenation, the S face feature matrices are concatenated to form the k-th texture feature map. Channel concatenation can retain the texture features extracted by different convolutional blocks and enrich the expression of the texture feature map.
[0108] Step 3: Construct a micro-expression feature extraction network, including: a spatial attention module, a multi-level attention module, a feature aggregation module, and an adaptive attention feature fusion module, and process to obtain the k-th micro-expression feature map , and the specific steps are as follows:
[0109] Step 3.1: The spatial attention module processes to obtain the k-th spatial feature map , and the specific process is as follows:
[0110] Step 3.1.1: Through three learnable linear transformation matrices , , perform a linear transformation on to generate the k-th query matrix , the k-th key matrix and the k-th value matrix ; The linear transformation can map the input features to different spaces, preparing for subsequent calculation of attention weights and enabling the model to focus on features in different regions.
[0111] Step 3.1.2: Reshape and along the height and width dimensions respectively. The reshaping operation changes the shape of the matrix, facilitating the extraction of horizontal feature information, helping the model capture key features in the horizontal direction, and flattening each channel of and in the horizontal direction. Thus, the k-th horizontal query matrix and the k-th horizontal key matrix are extracted from the two flattened matrices respectively, including the following steps:
[0112] Step 3.1.3: Use Equation (2) to obtain the k-th attention weight matrix :
[0113] (2)
[0114] In Equation (2), T represents the transpose;
[0115] Step 3.1.4: Obtain the k-th horizontal attention score using Equation (2). :
[0116] (3)
[0117] In Equation (3), Softmax represents the activation function; represents the scaling factor;
[0118] Step 3.1.5: Obtain the k-th horizontal direction attention weight using Equation (4). :
[0119] (4)
[0120] In Equation (4), represents element-wise multiplication; the calculation of the attention weight matrix and score can measure the correlation degree between features at different positions, highlight important horizontal features, and enable the model to pay more attention to key information in the horizontal direction.
[0121] Step 3.1.7: According to the process of Step 3.1.2 - Step 3.1.5, extract the k-th vertical query and from and the k-th vertical key , so as to calculate the k-th vertical attention weight matrix , the k-th vertical attention score , the k-th vertical distance matrix , and the k-th vertical direction attention weight
[0122] Step 3.1.8: After performing matrix multiplication on and , obtain the k-th intermediate result ; after performing matrix multiplication on and , obtain the k-th spatial feature ; by calculating the attention weight in the vertical direction and matrix multiplication operations, comprehensively capture spatial features and obtain a spatial feature map containing important spatial information.
[0123] Step 3.2: The multi-level attention module processes to obtain the k-th multi-level fusion feature map , and the specific process is as follows:
[0124] Step 3.2.1: Perform a convolution operation on to obtain the k-th convolution feature map ; the convolution operation can further extract features and enrich the feature expression.
[0125] Step 3.2.2: Evenly divide along the channel axis into n sub - feature maps. Among them, the i - th sub - feature map is denoted as , . Each sub - feature map has the same spatial dimension, and the number of channels of each sub - feature map is reduced to along the channel axis into n sub - feature maps, where the i - th sub - feature map is denoted as , of the original; Channel splitting can extract features from different channel scales, enabling the model to learn information at different levels. ; Channel splitting can extract features from different channel scales, enabling the model to learn information at different levels.
[0126] Step 3.2.3: Perform a convolution operation on with a 3×3 convolution kernel to compress the number of channels and output the i - th compressed feature map ; Compressing the number of channels can reduce the computational load while retaining key features. Perform a convolution operation on with a 3×3 convolution kernel to compress the number of channels and output the i - th compressed feature map ; Compressing the number of channels can reduce the computational load while retaining key features. ; Compressing the number of channels can reduce the computational load while retaining key features.
[0127] Step 3.2.4: After splicing the n compressed feature maps along the channel axis, generate the k - th multi - level fusion feature map ; The splicing operation fuses features of different scales to obtain a multi - level fusion feature map containing rich hierarchical information. ; The splicing operation fuses features of different scales to obtain a multi - level fusion feature map containing rich hierarchical information.
[0128] Step 3.3: The feature aggregation module processes and and performs adaptive weighted fusion to obtain the k - th aggregated feature map . The specific process is as follows: and to perform adaptive weighted fusion to obtain the k - th aggregated feature map . The specific process is as follows: The specific process is as follows:
[0129] Step 3.3.1: Obtain the k - th fusion feature map using Equation (5): :
[0130] (5)
[0131] In Equation (5), and are two weights to be learned, and + = 1, , ensuring the stability of the fusion. The weighted fusion combines the information of the spatial feature map and the multi - level fusion feature map, making the fusion result more representative.
[0132] Step 3.3.2: Perform a convolution operation on with a 1×1 convolution kernel to compress the number of channels and obtain the k - th feature map ; Apply a 1×1 convolution kernel to perform per - channel spatial feature extraction on to obtain the k - th spatial depth feature map ; Then perform a convolution operation on with a 1×1 convolution kernel to compress the number of channels to the target dimension and obtain the k - th aggregated feature Perform a convolution operation on with a 1×1 convolution kernel to compress the number of channels and obtain the k - th feature map ; Apply a 1×1 convolution kernel to perform per - channel spatial feature extraction on to obtain the k - th spatial depth feature map ; Then perform a convolution operation on with a 1×1 convolution kernel to compress the number of channels to the target dimension and obtain the k - th aggregated feature ; Apply a 1×1 convolution kernel to perform per - channel spatial feature extraction on to obtain the k - th spatial depth feature map ; Then perform a convolution operation on with a 1×1 convolution kernel to compress the number of channels to the target dimension and obtain the k - th aggregated feature to perform per - channel spatial feature extraction on to obtain the k - th spatial depth feature map ; Then perform a convolution operation on with a 1×1 convolution kernel to compress the number of channels to the target dimension and obtain the k - th aggregated feature ; Then perform a convolution operation on with a 1×1 convolution kernel to compress the number of channels to the target dimension and obtain the k - th aggregated feature to compress the number of channels to the target dimension and obtain the k - th aggregated feature The multiple convolution operations further compress the feature dimension and extract more compact and effective features.
[0133] Step 3.4: The adaptive attention feature fusion module processes , and to obtain the k-th micro-expression feature map . The specific process is as follows:
[0134] Step 3.4.1: After connecting and along the channel dimension, the k-th channel fusion feature map is formed. Then, a convolution operation is performed on with a convolution kernel of 1×1 to compress it to the target dimension, obtaining the k-th target feature map . Channel connection fuses features from different sources, and the convolution operation compresses the dimension to make the features more compact.
[0135] Step 3.4.2: After connecting and along the channel dimension, the k-th new feature map is obtained;
[0136] Step 3.4.3: Average pooling and max pooling operations are respectively performed on to obtain the k-th average pooling feature and the k-th max pooling feature ; The sigmoid function is respectively used to normalize and to obtain the attention weight of the k-th average pooling feature and the attention weight of the k-th max pooling feature;
[0137] Step 3.4.4: The k-th weighted feature map is obtained using Equation (6):
[0138] (6)
[0139] In Equation (6), denotes element-wise multiplication;
[0140] Step 3.4.5: A convolution operation is performed on with a convolution kernel of 1×1 to compress its number of channels, obtaining the k-th micro-expression feature map . Through weighted fusion and convolution operations, a micro-expression feature map containing rich micro-expression information is obtained.
[0141] Step 4: Construct a joint feature extraction sub-network, including: a multi-scale differential convolution network and a feature fusion module, and process and to obtain the k-th joint feature map , so as to obtain the set of joint feature maps of video sample E ; is the number of feature matrices, including the following steps:
[0142] Step 4.1: Concatenate and on the channel dimension to form the k-th preliminary joint feature map ; The concatenation operation fuses texture features and micro-expression features, providing more comprehensive information for subsequent processing.
[0143] Step 4.2: The multi-scale differential convolution network uses multiple CDC convolution blocks to process . There are 3 CDC convolution blocks. The first one: a 3×3 central differential convolution with a dilation rate of 1 and an output channel number of 512; the second one: a 3×3 central differential convolution with a dilation rate of 2 and an output channel of 256; the third one: a 3×3 central differential convolution with a dilation rate of 3 and an output channel of 128; Central differential convolutions with different dilation rates can capture feature information at different scales. When the dilation rate is 1, basic local features can be obtained. After the dilation rate increases, the feature changes in a wider area can be noticed. This can enrich the feature expression. After obtaining features at different levels, they are connected together to form the k-th multi-scale differential feature map ; In practical applications, if the feature differences in the dataset are large, the number of CDC convolution blocks can be appropriately increased or decreased, and the dilation rate and output channel number can be adjusted. The optimal parameters can be determined through experiments.
[0144] Step 4.3: The feature fusion module performs a convolution operation on with a 1×1 convolution kernel to obtain the k-th convolution-compressed feature map ; The 1×1 convolution kernel can compress the number of channels without changing the spatial dimension of the feature map, reducing the computational amount and fusing the information between channels. In some scenarios with extremely high requirements for computational speed, further optimizing the convolution kernel parameters or adopting a more efficient convolution method can be considered.
[0145] Step 4.4: Calculate the k-th spatial position weight using Equation (7):
[0146] (7)
[0147] In Equation (7), represents the global average pooling operation, represents the max pooling operation, Expressed as the Sigmoid function, is a convolution operation with a 1×1 convolution kernel; by combining global average pooling and max pooling, the global information of features can be obtained from different perspectives, and then through 1×1 convolution and Sigmoid function processing, the spatial position weights are obtained, highlighting the important spatial positions in the feature map. In other similar image feature processing tasks, this way of calculating spatial position weights can also be used to enhance the features of key regions.
[0148] Step 4.5, Multiply element-wise with to obtain the k-th spatial position feature ;
[0149] Step 4.6, Use convolution kernels of several different scales to perform convolution operations on respectively. Here, convolution kernels of sizes 7×7, 5×5, and 3×3 are used respectively. After outputting the feature maps of corresponding different scales and cross-connecting and fusing them, the k-th combined feature map is generated. Convolution kernels of different sizes can extract features of different scales. Large convolution kernels focus on features in larger regions, and small convolution kernels focus on local details. Cross-connecting and fusing can fully integrate these features, making the combined feature map more representative. In practical applications, more combinations of convolution kernels of different scales can be tried according to the model effect.
[0150] Step 5, Construct the total loss function , which is used to train the detection network composed of the texture feature extraction sub-network, the micro-expression feature extraction network, and the combined feature extraction sub-network to obtain the 3D mask face presentation attack detection model with optimal parameters to achieve 3D mask face presentation attack detection, including the following steps:
[0151] Step 5.1, After splicing in the channel dimension, a global feature map is formed; the splicing operation integrates all the combined feature map information of the samples, providing comprehensive data for subsequent classification.
[0152] Step 5.2, Input into the fully connected layer for processing, and output the predicted class probability value of E, where represents the probability that E is a 3D mask attack, represents the probability that E is a real face, and satisfies ; the fully connected layer can map the high-dimensional feature vector to the classification space to obtain the predicted probability value. In practical applications, the number of layers and the number of neurons of the fully connected layer can be adjusted according to the complexity of the model and the size of the dataset.
[0153] Step 5.3. Construct the cross-entropy loss using Equation (8) :
[0154] (8)
[0155] In Equation (8), represents the probability value of the i-th predicted class, represents the i-th true class label; the cross-entropy loss is used to measure the difference between the predicted probability and the true label, guiding the model to train in the direction of correct classification. During the training process, the smaller the value of the cross-entropy loss, the closer the predicted result of the model is to the true label.
[0156] Step 5.4. Construct the center loss using Equation (9) :
[0157] (9)
[0158] In Equation (9), represents the class center matrix of Y and is initialized as the mean of; is the Frobenius norm; the center loss makes the features of the same class more compact and the features of different classes more separated, improving the generalization ability of the model. At the beginning of training, the class center matrix may not be very accurate, but it will be gradually optimized as the training progresses.
[0159] Step 5.5. Construct the total loss function using Equation (10) :
[0160] (10)
[0161] In Equation (10), is the dynamic balance coefficient; by adjusting the value of , the proportion of the cross-entropy loss and the center loss in the total loss can be balanced. For example, when is close to 1, the model pays more attention to the cross-entropy loss, that is, it pays more attention to the accuracy of classification; when is close to 0, the model pays more attention to the center loss, emphasizing the compactness of intra-class features. In actual training, the optimal value of can be determined through multiple experiments.
[0162] Step 5.6. Use the Adam optimizer with the parameter settings: learning rate η = 0.001, decaying by 0.5 times every 10 epochs down to 0.0001 at the lowest; momentum parameters β1 = 0.9, β2 = 0.999. Apply the backpropagation algorithm to process the detection network and continuously perform forward propagation to calculate the total loss and update the parameters of the detection network and the class center through backpropagation until the total loss Converge or reach the preset maximum number of iterations to obtain an optimal parameter 3D mask face presentation attack detection model. The Adam optimizer combines the advantages of momentum and adaptive learning rate, and can update model parameters quickly and stably. The learning rate decay strategy can accelerate the model convergence speed in the initial stage of training and avoid overfitting in the later stage. During the actual training process, the loss curve and the performance of the model on the validation set can be observed to further adjust the optimizer parameters.
[0163] Step 6, Model testing and deployment, verify the ability of the model to distinguish real faces from 3D mask attacks through a multi-scenario independent test set:
[0164] Step 6.1, Introduce new types of forged attack videos outside the training set (such as 3D masks of different materials, deep fakes, etc.) for testing, and focus on evaluating the model's recognition ability for unseen presentation attack types to ensure that the system has the robustness to continuously protect against unknown threats. In practical applications, with the development of technology, new forged attack methods emerge continuously. Regularly testing the model with new attack videos can timely detect the deficiencies of the model and make improvements. For example, if it is found that the model has a low recognition rate for 3D masks of a certain new material, more relevant data can be collected specifically and the model can be retrained.
[0165] Step 6.2, Select natural face videos collected in real scenarios (including variables such as different lighting, angles, expressions, etc.) for testing, and verify the false alarm risk of the model by calculating the false alarm rate, requiring the index to be lower than the industry benchmark value. The complex factors in real scenarios will affect the model performance. Testing the false alarm rate can evaluate the reliability of the model in actual use. If the false alarm rate is too high, it can be checked whether there are problems in the feature extraction or classification decision-making links of the model, and the model structure or parameters can be optimized.
[0166] Step 6.3, If the attack recognition rate ≥ 99% and the false alarm rate ≤ the threshold, it is determined that the model passes the test and can be deployed online; if any index does not meet the standard, it is necessary to trace back the training data, adjust the network structure (or optimize the loss function) specifically, and re-enter the training and testing loop. During actual deployment, factors such as the running efficiency of the model and the hardware resource requirements also need to be considered to ensure that the model can run stably in the actual application environment.
[0167] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0168] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium. When the computer program is run by a processor, it executes the steps of the above method.
Claims
1. A 3D mask face presentation attack detection method based on micro-expression, characterized in that, Including the following steps: Step 1: Obtain a video dataset and perform preprocessing to obtain the th dual-channel input matrix of any video sample E ; Step 2: Construct a texture feature extraction sub-network, including: a shallow face feature extraction unit and a texture feature map generation unit, and process to obtain the k-th texture feature map ; Step 3: Construct a micro-expression feature extraction network, including: a spatial attention module, a multi-level attention module, a feature aggregation module, and an adaptive attention feature fusion module, and process to obtain the k-th micro-expression feature map ; Step 3.
1. The spatial attention module processes to obtain the k-th spatial feature map ; Step 3.2, the multi-level attention module processes to obtain the k-th multi-level fusion feature map ; Step 3.3, the feature aggregation module processes and to obtain the k-th aggregated feature map ; Step 3.
4. The adaptive attention feature fusion module processes , and to obtain the k-th micro-expression feature map ; Step 4. Construct a joint feature extraction sub-network, including: a multi-scale differential convolution network and a feature fusion module, and process and to obtain the k-th joint feature map , thereby obtaining the set of joint feature maps of the video sample E; is the number of feature matrices; Step 5: Construct the total loss function , which is used to train the detection network composed of the texture feature extraction sub-network, the micro-expression feature extraction network, and the joint feature extraction sub-network, to obtain a 3D mask face presentation attack detection model with optimal parameters, so as to realize 3D mask face presentation attack detection.
2. The 3D mask face presentation attack detection method based on micro-expression according to claim 1, wherein The said step 1 includes the following steps: Step 1.1: Obtain a video dataset and set the true class label of any video sample E in the video dataset as , indicating that the video sample is a 3D mask face presentation attack video, indicating that the video sample is a real face video; Step 1.2: Generate a feature matrix from video sample E , where represents the th feature matrix in , and is the number of rows of the feature matrix, is the number of columns of the feature matrix; is the number of feature matrices; Step 1.
3. Decompose the th feature matrix into the th actual feature matrix and the th device noise matrix , so as to construct the th objective function using Equation (1), and optimize it using the alternating direction method of multipliers to obtain the th optimal actual feature matrix : (1) In formula (1), is the rank of; denotes the 0-norm of; is the th regularization term coefficient; Step 1.4, convert from the RGB color space to the CHROM color space, and generate the first chrominance component and the second chrominance component , thereby forming the dual-channel input matrix .
3. The method for detecting 3D mask face presentation attacks based on micro-expressions according to claim 2, characterized in that, The said step 2 includes the following steps: Step 2.1: The shallow face feature extraction unit performs a convolution operation on to obtain the k-th shallow face feature matrix ; Step 2.2, the texture feature map generation unit includes S CDC convolutional blocks; Input into the s-th CDC convolutional block for processing to obtain the s-th facial feature matrix ; thus, S facial feature matrices are obtained. After concatenating the S facial feature matrices along the channels, the k-th texture feature map is formed .
4. The 3D mask face presentation attack detection method based on micro-expression according to claim 3, characterized in that, The s-th CDC convolutional block in step 2.2 obtains the s-th facial feature matrix through the following process : Step 2.2.
1. Convolve to obtain the s-th ordinary convolution feature . Step 2.2.2, perform central difference convolution operation on to obtain the s-th central difference convolution feature ; Step 2.2.3, after performing element-by-element addition on and , generate the s-th facial feature matrix .
5. The method for detecting 3D mask face presentation attacks based on microexpressions according to claim 4, wherein The said step 3.1 includes: Step 3.1.1: Perform linear transformations on through three learnable linear transformation matrices , , respectively to generate the k-th query matrix , the k-th key matrix and the k-th value matrix ; Step 3.1.2, Reshape and along the height and width dimensions respectively, such that each channel of and is flattened horizontally, and then extract the k-th horizontal query matrix and the k-th horizontal key matrix from the two flattened matrices respectively; Step 3.1.3: Obtain the k-th horizontal attention weight matrix using Equation (2). : (2) In formula (2), T represents transpose; Step 3.1.4: Obtain the k-th horizontal attention score using Equation (2). :[[]]END]] (3) In Equation (3), Softmax represents the activation function; represents the scaling factor; Step 3.1.5: Obtain the k-th horizontal direction attention weight using Equation (4). : (4) In Equation (4), represents element-by-element multiplication, is the k-th horizontal distance matrix; Step 3.1.
7. According to the process of Step 3.1.2 - Step 3.1.5, extract the k-th vertical query and from and the k-th vertical key , so as to calculate the k-th vertical attention weight matrix , the k-th vertical attention score , the k-th vertical distance matrix , and the k-th vertical direction attention weight ; Step 3.1.8, after performing matrix multiplication on and , the k-th intermediate result is obtained; after performing matrix multiplication on and , the k-th spatial feature is obtained.
6. The method for detecting 3D mask face presentation attacks based on microexpressions according to claim 5, characterized in that, The said step 3.2 includes: Step 3.2.
1. Convolve to obtain the k-th convolutional feature map . Step 3.2.2, divide uniformly along the channel axis into n sub-feature maps, where the i-th sub-feature map is denoted as , , each sub-feature map has the same spatial dimension, and the number of channels of each sub-feature map is reduced to of the original; Step 3.2.3, perform a convolution operation on to compress the number of channels and output the i-th compressed feature map ; Step 3.2.4: After splicing n compressed feature maps along the channel axis, generate the k-th multi-level fusion feature map .
7. The method for detecting 3D mask face presentation attacks based on micro-expressions according to claim 6, wherein The said step 3.3 includes: Step 3.3.1: Obtain the k-th fused feature map using Equation (5) :[[-END]] (5) In formula (5), and are two weight coefficients to be learned, and + = 1, ; Step 3.3.2, perform a convolution operation on to compress the number of channels and obtain the k-th feature map ; apply a convolution with a 1×1 convolution kernel to perform per-channel spatial feature extraction and obtain the k-th spatial depth feature map ; then perform a convolution operation on to compress the number of channels to the target dimension and obtain the k-th aggregated feature .
8. The method for detecting 3D mask face presentation attacks based on micro-expressions according to claim 7, wherein The said step 3.4 includes: Step 3.4.1: After connecting and along the channel dimension, the k-th channel fusion feature map is formed. Then, convolution operation is performed on to compress it to the target dimension, and the k-th target feature map is obtained; Step 3.4.2, connect and along the channel dimension to obtain the k-th new feature map ; Step 3.4.3: For perform average pooling and max pooling operations respectively, and correspondingly obtain the k-th average pooling feature and the k-th max pooling feature ; respectively normalize through the Sigmoid function and to obtain the attention weight of the k-th average pooling feature and the attention weight of the k-th max pooling feature; Step 3.4.
4. Obtain the k-th weighted feature map using Equation (6). : (6) In formula (6), represents element-wise multiplication; Step 3.4.5, perform a convolution operation on to compress its number of channels and obtain the k-th micro-expression feature map .
9. The 3D mask face presentation attack detection method based on micro-expression according to claim 8, wherein The said step 4 includes: Step 4.1: Concatenate and on the channel dimension to form the k-th preliminary joint feature map ; Step 4.2: The multi-scale differential convolution network uses multiple CDC convolution blocks to process it, obtain features at different levels, and then connect them together to form the k-th multi-scale differential feature map ; Step 4.3, the feature fusion module performs a convolution operation to obtain the k-th convolutionally compressed feature map ; Step 4.4: Calculate the k-th spatial position weight using Equation (7) :[[]]END]] (7) In formula (7), represents the global average pooling operation, represents the max pooling operation, represents the Sigmoid function, is a convolution operation with a 1×1 convolution kernel; Step 4.5, multiply with element by element to obtain the k-th spatial position feature ; Step 4.
6. Use several convolution kernels of different scales to perform convolution operations. After outputting feature maps corresponding to different scales and performing cross-connection fusion, generate the k-th combined feature map .
10. The method for detecting 3D mask face presentation attacks based on micro-expressions according to claim 9, characterized in that, The said step 5 includes: Step 5.
1. After splicing in the channel dimension, a global feature map is formed . Step 5.2, input into the fully connected layer for processing, and output the predicted class probability value of E , where represents the probability that E is a 3D mask attack, represents the probability that E is a real human face, and satisfies ; Step 5.
3. Construct the cross-entropy loss using Equation (8). :[[]]END]] (8) In formula (8), represents the probability value of the i-th predicted category, represents the label of the i-th true category; Step 5.
4. Construct the center loss using Equation (9). :[[]]END]] (9) In formula (9), represents the class center matrix of Y and is initialized to the mean value of; is the Frobenius norm; Step 5.
5. Construct the total loss function using Equation (10). :[[]]END]] (10) In formula (10), is the dynamic balance coefficient; Step 5.6: Process the detection network using the backpropagation algorithm, and continuously perform forward propagation to calculate the total loss , update the parameters and class centers of the detection network through backpropagation until the total loss converges or reaches the preset maximum number of iterations, and obtain a 3D mask face presentation attack detection model with optimal parameters.