Cross-domain face anti-fraud system and method based on Resnet convolution modulation
By introducing a convolutional modulation module and a feature fusion strategy into the ResNet convolutional network, the problems of insufficient generalization performance and high computational complexity of cross-domain face anti-fraud methods in cross-domain scenarios are solved, achieving lightweight and efficient cross-domain face anti-fraud recognition.
Patent Information
- Application Number
- CN202510931407.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-11-21
AI Technical Summary
Existing cross-domain face fraud prevention methods have insufficient generalization performance in cross-domain scenarios, high computational complexity, and are difficult to deploy on mobile devices. Furthermore, traditional models such as ResNet fail to fully exploit local texture and subtle artifact information.
We introduce a Convolutional Modulation Module (ConvMod) and a multi-level feature fusion strategy, combined with the ResNet convolutional network. The ConvMod generates dynamic attention weights to focus on key regions and suppress noise. Feature fusion enhances discriminative power, and supervised contrastive learning and gradient projection domain alignment strategies improve the model's generalization ability.
While maintaining the model's lightweight nature, it enhances the feature extraction and generalization performance of the cross-domain face anti-fraud model, improves the ability to identify fake faces, reduces computational complexity, and is suitable for mobile deployment.
Smart Images

Figure CN120997888A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of computer vision, and relates to biometric recognition, in particular to a cross-domain face anti-fraud system and method based on Resnet convolution modulation. BACKGROUND
[0002] Face recognition technology is widely used in mobile payment, access control system and other scenarios, but its security is vulnerable to counterfeit attacks. For example, print attacks, video replay attacks, 3D mask attacks, etc. Existing face anti-fraud methods mainly extract discriminative features of live samples and fake samples through deep learning models. However, in practical applications, there are significant domain differences between training data and test data, such as resolution, sensor type, lighting conditions, etc., which leads to a significant decline in the generalization performance of the model in unknown domains.
[0003] Currently, deep learning methods to solve cross-domain face anti-fraud problems mainly fall into the following categories: Transformer-based methods, meta-learning-based methods, generation-based methods, and alignment-based methods. Among them, VisionTransformer uses self-attention mechanism to capture long-range dependencies in face images, enhancing the model's ability to distinguish real faces from fake faces. Meta-learning learns a model that can quickly adapt to new domains by simulating domain changes during the training phase, thereby enhancing the generalization ability of the face anti-fraud model. Generation-based methods use generative adversarial networks (GAN) or other methods to generate diverse live samples, thereby enhancing the model's ability to recognize various counterfeit attacks. Alignment-based methods reduce the differences between domains by aligning the domains during training, allowing the model to perform well in specific domains while maintaining its ability to generalize to unseen domains.
[0004] However, Vision Transformer and other structures improve feature expression through global attention, but their computational complexity is high, making it difficult to deploy on mobile devices. Existing alignment methods mostly use traditional Resnet and other models to extract features, without explicitly designing modules to enhance the discriminative features of deep convolution outputs, resulting in insufficient mining of subtle artifacts such as local texture and reflections in cross-domain scenarios. In summary, although existing face anti-fraud methods improve the performance and generalization ability of the model to some extent, there is still room for improvement in terms of cross-domain generalization, computational efficiency, and feature extraction ability. SUMMARY
[0005] In view of the deficiencies of the prior art, the present application proposes a cross-domain face anti-fraud system and method based on Resnet convolution modulation, which adds a convolution modulation module (Convolutional Modulation Block, ConvMod) in a deep convolutional network and combines a multi-level feature fusion strategy to maintain the lightweight of the model while improving the ability to obtain deep features, aiming to improve the performance and generalization ability of the cross-domain face anti-fraud model.
[0006] A cross-domain face anti-fraud system based on Resnet convolution modulation includes a visible light camera, an embedded device, a display module, and an alarm module.
[0007] The visible light camera is used to capture real-time RGB picture information.
[0008] The embedded device is deployed with a trained face detection network and a cross-domain face anti-fraud model. The face detection network is used to detect and crop the RGB picture information captured by the visible light camera, and then the cross-domain face anti-fraud model is used to complete real-time inference of face feature extraction and liveness judgment.
[0009] The display module displays the face capture frame and the liveness judgment result in real time according to the detection result of the embedded device, and pops up a prompt when a fake face appears.
[0010] The alarm module includes a buzzer alarm and a rotating warning light, which will sound an alarm when the liveness judgment result is fake face for three consecutive times.
[0011] The cross-domain face anti-fraud method based on Resnet convolution modulation specifically includes the following steps: Step 1, convert the video data into image frames, then use the face detection network to detect and crop the images, and adjust them to a uniform size. Then perform two random image enhancement operations on the cropped images, convert the enhanced images to Tensor tensors and perform standardization processing as training samples . Label the training samples and domain sources . Where i represents the sample index.
[0012] Step 2, for the training samples processed in step 1 , based on Resnet convolution modulation, perform multi-scale feature encoding: Step 2-1, feature preprocessing Use an initial convolution with a kernel size of 7x7 to preliminarily downsample the initial features X of the training samples , and then perform normalization, activation function, and pooling processing to obtain the preprocessed features .
[0013] Step 2-2, residual block group feature extraction Based on the Resnet residual learning framework, a basic feature extraction network is constructed by stacking L residual blocks. For the feature map processed by the initial convolution layer Feature extraction is performed, where H and W represent the height and width of the feature map, respectively, and C0 represents the number of channels of the feature map. The first Output of the L-th residual block is: where, =1,2,…L, is a dimension adaptation function used to adjust the shape of to be consistent with . represents a convolution kernel with a size of 3x3. is a normalization operation. represents a ReLU activation function.
[0014] Step 2-3, convolution modulation module For high-order features output by the residual block group, convolution modulation is performed. First, a depth separable convolution (DWConv) is used to generate a spatial weight matrix A: where, is a linear transformation matrix used to generate intermediate features Q, represents a depth convolution operation with a kernel size of k x k.
[0015] Then, the high-order features are linearly transformed to generate a value feature matrix V, and are element-wise fused (Hadamard product) with the spatial weight matrix A, outputting the modulated features Z: where, is a linear transformation matrix used to generate the feature matrix V, represents a Hadamard product.
[0016] Step 3, in view of the problems of overfitting and loss of some important global information when extracting detailed features in the convolution modulation module, based on feature complementarity, the features of the basic feature extraction network and the convolution modulation module are fused by linear addition: Use global average pooling to enhance the feature map Dimensionality reduction is performed, compressing the spatial dimension into a single value to obtain a global feature representation. The pooled feature map is then flattened into a one-dimensional vector, input into a fully connected layer, and output as a sample. This represents the probability of a live sample.
[0017] Step 4, Loss Calculation: Step 4-1: Calculate the contrast loss Supervised contrastive learning (SupCon) is used for feature separation. It leverages label information to cluster similar samples and separate dissimilar samples in the feature space, improving the model's discriminative ability. For a batch of training samples, a contrastive loss is defined. for: Where 2b is the batch size and τ is the temperature parameter. This indicates that the batch contains samples A set of sample indexes that have the same label and belong to the same domain. Represents a set Size. Table Sample Enhanced feature map The feature vector representation after global average pooling and normalization.
[0018] Step 4-2: Calculate the classification loss In multi-domain scenarios, for each domain Train a classifier separately The goal is to minimize the binary cross-entropy loss over this domain. The final total classification loss... Classify losses by field The average value is used for sharing feature extractors. ϕ Domain-specific classification head βe Learning: Where E represents the number of training time domains. A set of fields .
[0019] Step 4-3: Calculation of Total Loss Taking into account the comparative loss and classification loss Set the total loss for: in It is a comparative loss The coefficient.
[0020] Step 5, Backpropagation and domain alignment Step 5-1, Backpropagation At the beginning of training, the initialization feature extractor based on the pre-trained model of ImageNet is used, and the SGD with momentum and weight decay strategy is adopted to prevent overfitting. Set StepLR to decay the learning rate according to the fixed period. Based on the total loss Backpropagation is performed.
[0021] Step 5-2, Domain alignment In order to further reduce the distribution difference between different domains and improve the generalization ability of the model, after the specified epoch in the training process, the gradient projection type domain alignment strategy is used to update the direction of each domain classifier , and the geometric alignment of the hyperplane between different domains is realized. The domain alignment loss is defined as: wherein, denotes the feature extractor fixed at the th domain, and the set of classifier parameters allowed to be used is effective. denotes the alignment direction set formed by linear interpolation starting from the current domain classifier and towards the farthest domain classifier , which is used to constrain the consistency of the classifier direction.
[0022] The classifier is updated in the geometric space to its projection form on the set : wherein is a balance factor, 2 represents the index corresponding to the domain classifier farthest from in the direction.
[0023] Through this projection update operation performed after backpropagation, the intra-domain classifier is guided to a more consistent direction in the geometric space while maintaining its discriminative ability, thereby effectively improving the cross-domain robustness.
[0024] After enabling the projection update operation of the domain alignment strategy, the total loss is modified as: Step 6, input the face image into the model trained in step 5, output the probability of the face image as a live sample, and realize cross-domain face anti-fraud recognition.
[0025] Compared with the prior art, the present application has the following beneficial results: 1. A new convolution modulation module ConvMod is introduced in the traditional Resnet feature extraction network, which generates spatially adaptive dynamic attention weights through large kernel depth convolution, focuses on key areas, and dynamically weights the dynamic weights and basic features through Hadamard product, highlights important features, suppresses noise, and alleviates the problem that the traditional residual block has a limited receptive field and cannot model long-range dependencies, so that the extracted features are more suitable for cross-domain face anti-fraud and other fine-grained discrimination tasks.
[0026] 2. The low-level features x1 extracted by the Resnet four-layer residual block and the high-level features extracted by the ConvMod are fused to form multi-scale discrimination ability. After introducing feature fusion, the discrimination ability can be incrementally enhanced based on the x1 features, avoiding excessive modification of the distribution features by the ConvMod, improving the accuracy of feature extraction, and ultimately improving the generalization ability of the model. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1 The overall framework diagram for face anti-fraud model training; Figure 2 The architecture diagram of the feature extraction module; Figure 3 The architecture diagram of the convolution modulation module ConvMod; Figure 4 The loss function calculation flowchart. DETAILED DESCRIPTION
[0028] The present application will be further explained in conjunction with the accompanying drawings; A cross-domain face anti-fraud system based on Resnet convolution modulation includes a visible light camera, an embedded device, a display module, and an alarm module.
[0029] The visible light camera is installed at a position 1.6 m above the top of the access control gate, with a pitch angle ≤15° and an effective detection distance of 0.5 m~1.2 m, used to capture RGB picture information and connect the embedded device and the display module through the power supply line.
[0030] The embedded device is deployed with a trained face detection network and a cross-domain face anti-fraud model, which performs face detection and cropping on the RGB picture information captured by the visible light camera through the face detection network, and then uses the cross-domain face anti-fraud model to complete real-time inference of face feature extraction and liveness judgment.
[0031] The display module displays the face capture frame and the liveness judgment result in real time according to the detection results of the embedded device.
[0032] The alarm module includes a buzzer alarm and a rotating warning light.
[0033] The visible light camera acquires a video stream at a frame rate of 25 fps, and after hardware H.264 decoding, the video stream is input into an embedded device. The embedded device transmits the living body judgment result to a display module. When p≥0.95, a green "trusted authentication" prompt is displayed, and the channel permission is opened. When 0.8≤p<0.95, a yellow "secondary verification" interface is displayed, and a voice prompt is started. When p<0.8, a 3D anti-fake verification mode is triggered, and the user is required to perform blinking and shaking actions for further verification. When p<0.8 for three consecutive times, the alarm module is started.
[0034] The cross-domain face anti-fraud method based on Resnet convolution modulation specifically includes the following steps: Step 1, the embodiment takes four public data sets OULU-NPU (O), Idiap Replay-Attack (I), CASIA-FASD (C), and MSU-MFSD (M) as the original data sources. Among them, OULU-NPU has a total of 4950 videos, including 990 real face videos and 3960 fake face videos, Idiap Replay-Attack has a total of 1200 videos, including 200 real face videos and 1000 fake face videos, CASIA-FASD has a total of 600 videos, including 150 real face videos and 450 fake face videos, and MSU-MFSD has a total of 280 videos, including 70 real face videos and 210 fake face videos.
[0035] Each data set is regarded as a separate domain. First, an image is extracted from the original video stream every 15 frames, and then MTCNN is used for face detection and cropping, and the size is adjusted to 255x255 pixels. The leave-one-out test protocol is used, that is, in each training, three data sets are used as training data, and the remaining one data set is used as test data. For example, OCI→M means that OULU-NPU, CASIA-FASD, and Idiap Replay-Attack are regarded as three different training data domains, and MSU-MFSD is regarded as cross-domain test data.
[0036] As shown in Figure 1 , for the images in the training set, two random image enhancement operations are performed, including flipping or cropping, and then the enhanced images are converted into Tensor tensors and standardized as training samples . The labels and domain sources of the training samples are labeled . Wherein i represents the sample index.
[0037] Step 2, for the training samples processed in step 1 , multi-scale feature encoding is performed based on Resnet convolution modulation: Step 2-1, feature preprocessing As shown in Figure 2 , for the training sample First, a 7x7 convolution kernel is used to quickly capture low-level features such as edges and textures, and a step size of 2 is used to quickly downsample the input size. Batch normalization and ReLU activation are used to alleviate the problem of gradient explosion and enhance the model's expression ability. Finally, the features are downsampled by max pooling, further compressing the spatial features to obtain the feature map , where H and W represent the height and width of the feature map, and C0 represents the number of channels of the feature map.
[0038] Step 2-2, residual primitive feature extraction Based on the Resnet residual learning framework, a basic feature extraction network is constructed by stacking 4 residual blocks. For the feature map processed by the initial convolution layer, feature extraction is performed, and the output of the first residual block is : where =1,2,…4, denotes a 3x3 convolution kernel. is a normalization operation. is a dimension adaptation function. denotes a ReLU activation function.
[0039] Step 2-3, convolution modulation As shown in Figure 3 , the features extracted by the basic feature extraction network contain more global and structured information, but for the face anti-fraud task, local detail features are also needed. At the same time, the network structure of Resnet is fixed and cannot be dynamically adjusted to adapt to the needs of different tasks. To cope with unknown fields, for the high-order features output by the basic feature extraction network, a dynamic spatial modulation mechanism is introduced to enhance the long-range dependency modeling capability. First, a depthwise separable convolution DWConv is used to generate a spatial weight matrix A: where is a linear transformation matrix used to generate intermediate features Q, denotes a depth convolution operation with a kernel size of k x k.
[0040] Then the value matrix is generated by linear projection and fused with the spatial weight matrix A at the element level, and the modulated feature Z is output: wherein, is a linear transformation matrix for generating the feature matrix V, represents the Hadamard product.
[0041] Step 3, fuse the features of the base feature extraction network and the convolution modulation module by linear addition, and output the enhanced feature map : By direct addition, it can be ensured that the original features will not be completely covered or ignored. Even if the convolution modulation module fails to effectively extract some key information, the original features can still play a role in the final feature representation, reducing the risk of important information loss caused by the introduction of the ConvMod module.
[0042] Use global average pooling to reduce the dimension of the enhanced feature map , compress the spatial dimension to a single value, and obtain the global feature representation. Then flatten the pooled feature map into a one-dimensional vector, which reduces the number of parameters while enhancing the generalization ability of the model. Then flatten the pooled feature map into a one-dimensional vector, which facilitates subsequent processing by the fully connected layer. The corresponding pseudo code is shown in Table 1: Table 1 Step 4, loss calculation As shown in Figure 4 , different UUIDs are assigned to samples from different training sets, and each UUID corresponds to a . For sample features from different domains, input three fully connected layers according to the UUIDs, predict the probability of belonging to a live sample, and calculate the contrast loss and classification loss according to the separability and alignment of cross-domain face anti-fraud.
[0043] Step 4-1, calculate the contrast loss Use supervised contrast learning SupCon to achieve the separability of the model, and distinguish whether the sample is a fake sample or a live sample. The goal of SupCon is to pull samples of the same class closer together and repel samples of different classes in the embedding space, forming compact and separable feature clusters, and improving the generalization ability of the face anti-fraud model.
[0044] The pseudo code of SupCon calculation is shown in Table 2: Table 2 Step 4-2, classification loss In the multi-domain scenario, for each domain Train a classifier , the goal is to minimize the binary cross-entropy loss on this domain. The final total classification loss is the average of the individual environment classification losses, used for learning the shared feature extractor ϕ and the domain-specific classification head βe .
[0045] Classification loss The pseudo code of the calculation process is shown in Table 3: Table 3 Step 4-3, total loss calculation Considering separability and alignment, taking into account the contrast loss and the domain alignment-based classification loss , the total loss is set as: Where is the coefficient of the contrast loss . Since directly corresponds to the core performance indicator of face anti-fraud, its weight is fixed at 1 to avoid the sensitivity of the model to classification errors due to weighting.
[0046] Step 5, backpropagation based on total loss After backpropagation, to further reduce the distribution difference between different domains, a gradient projection-based domain alignment strategy is introduced after a specified number of training rounds, which updates the direction of each domain classifier to achieve geometric alignment of hyperplanes between different domains. The defined domain alignment loss at this time can be expressed as: After enabling the projection update operation of the domain alignment strategy, the total loss is modified as: Step 6, model training process The steps 2 and 3 are cyclically trained using the weights based on ImageNet pre-training as the initial weights of the basic feature extraction network, and the best one model and the model of the last round of training are finally saved, and the best index and the average index of the last 10 rounds are output. The SGD optimizer is used, and the initial learning rate is set to 5e-3, the learning rate is set to half at the 40th and 80th rounds, and the training round number is 100 in most protocols, 300 in the training round number of the most protocols. The batch size is set to 96 or 126, 0.995, , and the alignment strategy starts to be enabled at the 20th round.
[0047] To prove the effectiveness of the method, the experimental results are compared with the common methods in the prior art, the evaluation indexes of the experimental results are selected as best HTER, best AUC, mean HTER and mean AUC, the best performance results are as shown in Table 4, and the average performance of the last 10 rounds is as shown in Table 5: Table 4 Table 5 Wherein HTER (Half Total Error Rate) is a half total error rate, a classification performance of a model, AUC (Area Under the Curve) is an area under a ROC curve (Receiver Operating Characteristic Curve), a comprehensive performance of a model. The best index represents the best result that can be achieved by a model, and the mean index reflects the stability of a model in actual deployment. It can be known from the data in the table that the best and average HTER and AUC of the method are improved in most protocols, which indicates that the method has better generalization ability in the face anti-fraud task in the cross-domain scene.
Claims
1. A cross-domain face anti-fraud method based on ResNet convolutional modulation, characterized by: Specifically, the following steps are included: Step 1: Collect training samples for face liveness detection Label HeYu Source Construct a training set; where i represents the sample index; Step 2: Establish a cross-domain face anti-fraud model for training samples. First, preprocessing is performed, and then the data is input into the ResNet network for feature extraction. High-order features of ResNet network output A spatial weight matrix A is generated using depthwise separable convolution, and a value feature matrix V is generated through linear transformation. The spatial weight matrix A and the value feature matrix V are then fused element-wise to obtain the modulation feature Z. The modulation feature Z is then combined with higher-order features... Add them together to obtain the enhanced feature map. ; Enhanced feature map After dimensionality reduction and expansion, the samples are input into the classifier to obtain the data. The predicted probability of a live sample; Step 3: Use supervised contrastive learning (SupCon) for feature separation and calculate the defined contrastive loss. ; Calculate the classification loss for each domain The average value is used as the classification loss. Total loss Set as contrast loss and classification loss The weighted sum is used to train the cross-domain face anti-fraud model in step 2; Step 4: Use the cross-domain face anti-fraud model trained in Step 3 to detect the input image and output the liveness probability to achieve cross-domain face anti-fraud recognition.
2. The cross-domain face anti-fraud method based on ResNet convolutional modulation as described in claim 1, characterized in that: A face detection network is used to detect faces in the input image, and the face regions in the image are cropped out, adjusted to a uniform size, converted into Tensor tensors and standardized, and then input into a cross-domain face anti-fraud model for liveness detection.
3. The cross-domain face anti-fraud method based on ResNet convolutional modulation as described in claim 2, characterized in that: The face detection network is MTCNN.
4. The cross-domain face anti-fraud method based on ResNet convolutional modulation as described in claim 1, characterized in that: Use an initial convolution with a kernel size of 7×7 on the training samples. The initial features X are initially downsampled, and then normalized, activated, and pooled to output the preprocessed results.
5. The cross-domain face anti-fraud method based on ResNet convolutional modulation as described in claim 1, characterized in that: The ResNet network consists of four cascaded residual blocks.
6. The cross-domain face anti-fraud method based on ResNet convolutional modulation as described in claim 1, characterized in that: The spatial weight matrix A and the value feature matrix V are fused element-wise through the Hadamard product.
7. The cross-domain face anti-fraud method based on ResNet convolutional modulation as described in claim 1, characterized in that: In the initial training phase of the cross-domain face anti-fraud model, an initial feature extractor based on an ImageNet pre-trained model is used. Momentum-driven SGD and weight decay strategies are employed to prevent overfitting. The learning rate decays at a fixed period, based on the total loss. Perform backpropagation; after reaching the specified training epoch, use a gradient projection-based neighborhood alignment strategy to adjust the total loss. Revised to: in, It is a comparative loss coefficient, Indicates the domain alignment loss: in, Represents the set of fields. Feature extractor Fixed at the first The set of classifier parameters that are valid and allowed to be used within a domain; Indicates the current domain classifier Starting from the farthest domain classifier Alignment direction set formed by linear interpolation; classifier In geometric space, it is updated to its position in the set. Projection format on: in As a balance factor, 2 indicates and The domain classifier furthest in direction The corresponding index; After enabling the domain alignment strategy for projection update operations, the total loss will be... Revised to: 。 8. A cross-domain face anti-fraud system based on ResNet convolutional modulation, characterized in that: The device is used to implement the method as described in any one of claims 1 to 7, and includes a visible light camera, an embedded device, a display module, and an alarm module. The visible light camera is used to capture RGB image information in real time; The embedded device is equipped with a trained face detection network and a cross-domain face anti-fraud model. The face detection network performs face detection and cropping on the RGB image information captured by the visible light camera, and then the cross-domain face anti-fraud model is used to complete real-time reasoning for face feature extraction and liveness detection. The display module displays the face capture frame and liveness detection result in real time based on the detection results of the embedded device, and pops up a prompt when a fake face is detected. The alarm module is used to issue an alarm when the liveness detection results are three consecutive times that the face is fake.
9. A cross-domain face anti-fraud system based on ResNet convolutional modulation as described in claim 8, characterized in that: The alarm module includes a buzzer alarm and a rotating warning light.
10. A cross-domain face anti-fraud system based on ResNet convolutional modulation as described in claim 8, characterized in that: When the probability of liveness detection p ≥ 0.95, the display module shows a green "Trusted Authentication" prompt, granting access to the channel; When the probability of liveness detection is 0.8≤p<0.95, the display module shows a yellow "Secondary Verification" interface and initiates voice prompts; When the probability of a liveness detection p < 0.8, the 3D anti-counterfeiting verification mode is triggered, and the display module prompts the user to blink and shake their head for further verification. The alarm module is activated when the probability of the judgment is p<0.8 for three consecutive times.