Living body detection method based on fraudulent face enhancement and style feature self-supervision

By generating screens and printing attack face image data enhancement and self-supervision of style features, the problem of insufficient generalization ability of the face fraud detection model under a limited data set is solved, and higher recognition accuracy and adaptability are achieved, and suitable for embedded terminal devices.

CN120356268APending Publication Date: 2025-07-22SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510327470.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing face fraud detection model is difficult to effectively identify unknown attacks under training in limited data sets, and its generalization capabilities are insufficient, especially when facing changes in different devices, lighting and attack methods.

Method used

Data enhancement is performed by generating screen attack faces and printing attack face images, combining the style feature self-supervision module and the recombination contrast loss module, the model's learning ability of fraud clues is enhanced, and the domain discriminator loss is constructed using adversarial learning and gradient inversion to narrow the style feature distance between different views of the same image, and the distance between different attack methods and real and fake faces.

Benefits of technology

It improves the generalization performance of the model when facing unknown attacks, can more effectively identify fraudulent faces, reduce the demand for the number of data sets, and is suitable for mobile phones and embedded terminal devices, improving the generalization ability and recognition accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356268A_ABST
    Figure CN120356268A_ABST
Patent Text Reader

Abstract

The invention discloses a living body detection method based on fraudulent face enhancement and style feature self-supervision, and the method comprises the following steps: dividing a data set, extracting a face, and constructing a face and a corresponding authenticity label and an attack mode label; constructing a data field based on a screen attack face, a printing attack face and a data enhanced real face; extracting content features and style features of the face; constructing domain discriminator loss through antagonism learning and gradient inversion; reducing the distance of the same face between different views, learning attack information between style features, and constructing style feature self-supervision loss; recombining the content features and the style features, and constructing recombining comparison loss; performing classification prediction based on a classifier, and constructing classification loss; and training based on the total loss to obtain an optimal network model, and outputting a prediction result based on the optimal network model. According to the method, the generalization ability of true and false face detection can be improved, and the true and false face detection performance is excellent in limited face data set training and unknown attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face liveness detection, and particularly relates to a liveness detection method based on fraud face enhancement and style feature self-supervision. Background Art

[0002] Compared with fingerprint and iris recognition technologies, the acquisition of facial features is relatively more convenient and is widely used in fields related to identity verification and security management. In recent years, with the development of deep learning technology, face recognition technology has been significantly improved, but the security of face recognition has also exposed hidden dangers, especially the use of forged faces (such as using photos or video playback) to deceive face recognition systems. To address this challenge, people have proposed face anti-spoofing technology to ensure the security of face recognition systems. As a pre-security measure to prevent forged attacks, face anti-spoofing technology usually regards the authenticity of a face as a binary classification problem or a multi-classification problem to distinguish whether it is a forged face.

[0003] Currently, the mainstream face fraud detection algorithms are mainly divided into two categories: traditional machine learning methods and deep learning-based methods. Traditional machine learning usually relies on manually designed features, such as using features like Local Binary Pattern (LBP), Difference of Gaussians (DoG), etc. However, manually designed features cannot comprehensively represent fraud clues, resulting in poor generalization ability of the model and difficulty in distinguishing the authenticity of faces when facing unseen attacks. Deep learning-based methods mainly adopt methods such as domain adaptation, domain generalization, anomaly detection, and multi-modal fusion strategies, with a certain improvement in generalization ability. However, due to differences in image capture devices, attack methods, lighting conditions, and data resolution in different datasets, the recognition performance of the model is low when the data distribution differences are large.

[0004] Therefore, the most important challenge currently faced by face fraud detection technology is how to have excellent true and false face detection performance when the model trained with a limited face dataset faces unknown attacks. Summary of the Invention

[0005] To overcome the defects and deficiencies of the existing technologies, the present invention provides a live detection method based on fraud face enhancement and style feature self-supervision. The present invention uses a screen attack face generator, a printed attack face generator, and a real face enhancer for data augmentation, enabling the model to learn more fraud methods on a limited face dataset and increasing the diversity of styles. In addition, a style feature self-supervision module is introduced to supervise the style features of different views, enabling the model to focus on the fraud clues in the image rather than being disturbed by environmental factors such as lighting. A recombination contrast loss module is introduced to recombine the content features and style features, and perform comparison and classification according to screen attack faces, printed attack faces, and real faces, making it easier for the model to learn the differences between different fraud methods and improving the generalization ability of the model through finer-grained division.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention provides a live detection method based on fraud face enhancement and style feature self-supervision, including the following steps:

[0008] Divide the dataset, extract frames from the face video and then intercept face images, and construct face images and corresponding true / false labels and attack mode labels;

[0009] Generate screen attack face images and printed attack face images, and perform data augmentation on real face images;

[0010] Construct the original dataset's face images, screen attack face images, printed attack face images, and enhanced real face images into a data domain;

[0011] Extract the content features and style features of the face images in the data domain;

[0012] Through adversarial learning and gradient reversal, make the content features extracted from different domains similar, and construct a domain discriminator loss;

[0013] Narrow the distance between the same face image in different views, learn the attack information between style features, and construct a style feature self-supervision loss;

[0014] Recombine the content features and style features, narrow the distance between face images of the same attack mode, and increase the distance between face images of different attack modes and between true / false face images, and construct a recombination contrast loss;

[0015] Based on a classifier, perform classification prediction on real face images, screen attack face images, and printed attack face images, and construct a classification loss;

[0016] Construct a total loss based on the domain discriminator loss, style feature self-supervision loss, recombination contrast loss, and classification loss;

[0017] Train the optimal network model based on the total loss and output the prediction result based on the optimal network model.

[0018] As a preferred technical solution, generate screen attack face images and printed attack face images, and perform data augmentation on real face images, specifically including:

[0019] Generate screen attack faces by adding moiré patterns, screen reflections, pixelation, light spots, motion blur artifacts, and halos;

[0020] Generate printed attack faces by adding ink traces, SFC - halftone, BN - halftone, and color distortion;

[0021] Convert the real face image in the RGB domain of different shooting devices, and simulate real faces taken by different devices by changing the color gamut.

[0022] As a preferred technical solution, extract the content features of face images in the data domain, specifically including:

[0023] Input the face image into a network composed of convolution and max - pooling for preliminary feature extraction to generate preliminary features, then input it into a network composed of three groups of residual blocks for intermediate feature extraction to generate intermediate features, and finally input it into a network composed of convolution, BN layer, and ReLU function for final feature extraction to generate content features;

[0024] The processing of the BN layer is expressed as:

[0025]

[0026]

[0027] where μ batch represents the mean of all samples in the batch, σ batch represents the standard deviation of all samples in the batch, ∈ is a small constant to prevent division - by - zero errors during calculation, γ and β are scaling factors and offsets to be learned, and H, W, C represent the height, width, and number of channels of the image.

[0028] As a preferred technical solution, extract the style features of face images in the data domain, specifically including:

[0029] The face image is input into a network composed of convolution and max pooling for preliminary feature extraction to generate preliminary features, and then input into a multi-scale network composed of three groups of residual blocks and a feature pyramid structure for intermediate feature extraction to generate three feature maps. After the three feature maps are subjected to feature fusion operations of convolution and IN layers, multi-scale features are generated. Finally, the multi-scale features are input into a network composed of convolution, IN layer, max pooling layer and fully connected layer for final feature extraction to generate style features;

[0030] The processing of the IN layer is expressed as:

[0031]

[0032] Among them, μ inst represents the mean of the sample feature map, σ inst represents the standard deviation of the sample feature map, ∈ is a small constant to prevent division by zero error during calculation, γ and β are scaling factor and offset to be learned, and H, W represent the height and width of the image.

[0033] As a preferred technical solution, through adversarial learning and gradient reversal, the content features extracted from different domains are made similar, and the domain discriminator loss is constructed, specifically including:

[0034] Through adversarial learning, the model learns general content features. The content features are passed through an adaptive average pooling layer and a Reshape layer, and output features are obtained through a double-layer fully connected layer, a ReLU function activation layer, and a Dropout layer;

[0035] The domain discriminator loss is expressed as:

[0036]

[0037] Among them, L adv (G,D) represents the supervised loss, G represents the content feature extractor, D represents the domain discriminator, x represents the input image, y represents the region label of the image, M represents the number of source domains, represents the indicator function.

[0038] As a preferred technical solution, a dynamic coefficient λ is introduced in the initial stage of adversarial learning, expressed as:

[0039]

[0040] Among them, cur_iter represents the number of current training iterations, and sum_iter represents the total number of iterations.

[0041] As a preferred technical solution, the distance between the same face image in different views is reduced, the attack information between style features is learned, and the style feature self-supervised loss is constructed, specifically including:

[0042] The face image is generated into two sub-views view1 and view2 through data enhancement, and the corresponding style features are extracted and style features The projector is used to map the two style features into a new feature space. The function calculates the similarity of two style features through L self-supervised The loss function brings the distance between different views of the same image closer, which is specifically expressed as:

[0043]

[0044] Among them, G style represents the style feature extractor, Proj represents the projector, and L self-supervised The network is optimized for the self-supervised loss, and τ is the temperature coefficient.

[0045] As an optimal technical solution, the content features and style features are reorganized to reduce the distance between face images with the same attack mode, increase the distance between face images with different attack modes and between real and fake face images, and construct a reorganization contrast loss, specifically including:

[0046] The style features are pooled globally to obtain a global description of each channel, and two parameters γ and β are generated through a multi-layer perceptron.

[0047] After the content feature is convolved, the AdaIN layer is applied to normalize it, and it is adjusted by γ and β. The new feature z is obtained by activating the layer through the ReLU function. Finally, the feature z is convolved and the AdaIN layer, γ and β are adjusted again, and it is added to the original content feature to obtain the final style assembly feature. The operation of the AdaIN layer is expressed as:

[0048]

[0049] Where μ(x) and σ(x) represent the mean and standard deviation of the channel dimension, and γ and β represent the affine parameters generated from the style input;

[0050] The reconstruction contrast loss is expressed as:

[0051]

[0052] Among them, y i =1 indicates samples of the same category, y i =0 indicates samples of different categories, m is the interval coefficient, T anchor represents the anchor point, i.e., the self-assembly feature, T positive represents the reassembled features with the same style features as the anchor point, T negative Reassembled features that represent different style features from the anchor.

[0053] As a preferred technical solution, a classification loss is constructed and expressed as:

[0054]

[0055] where N represents the total number of samples, C represents the number of categories, y i,c represents the category label, and p i,c represents the probability that the model predicts that sample i belongs to category c.

[0056] The present invention also provides a live detection system based on fraud face enhancement and style feature self-supervision, including: a data preprocessing module, a screen attack face generator, a printed attack face generator, a real face enhancer, a data domain construction module, a content feature extractor, a style feature extractor, a domain discriminator module, a style feature self-supervision module, a reconstruction contrast loss construction module, a classifier, a network model training module, and a prediction result output module;

[0057] The data preprocessing module is used to divide the data set, extract frames from the face video and then intercept face images, and construct face images and corresponding true / false labels and attack mode labels;

[0058] The screen attack face generator is used to generate screen attack face images;

[0059] The printed attack face generator is used to generate printed attack face images;

[0060] The real face enhancer is used to perform data enhancement on real face images;

[0061] The data domain construction module is used to construct a data domain from the face images of the original data set, as well as screen attack face images, printed attack face images, and enhanced real face images;

[0062] The content feature extractor is used to extract the content features of the face images in the data domain;

[0063] The style feature extractor is used to extract the style features of the face images in the data domain;

[0064] The domain discriminator module is used to make the content features extracted from different domains similar through adversarial learning and gradient reversal, and construct a domain discriminator loss;

[0065] The style feature self-supervision module is used to narrow the distance between the same face image in different views, learn the attack information between style features, and construct a style feature self-supervision loss;

[0066] The recombinant contrast loss construction module is used to recombine the content features and style features, reduce the distance between face images of the same attack method, increase the distance between face images of different attack methods and between real and fake face images, and construct the recombinant contrast loss;

[0067] The classifier classifies and predicts real face images, screen attack face images, and printed attack face images to construct a classification loss;

[0068] The network model training module is used to train based on the total loss to obtain an optimal network model, and the total loss is constructed based on the domain discriminator loss, style feature self-supervised loss, recombinant contrast loss, and classification loss;

[0069] The prediction result output module is used to output the prediction result based on the optimal network model.

[0070] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0071] (1) Based on the screen attack face generator, printed attack face generator, and real face enhancer, the present invention performs data augmentation on face images in a limited dataset to simulate screen and printed attacks in the invisible domain; at the same time, through augmentations of different intensities, the model can more easily learn common fraud clues, such as moiré information that appears in screen attacks and paper texture information that appears in printed attacks, etc., which can greatly enhance the trainable face images from the perspective of the data domain, thereby improving the generalization ability of the model.

[0072] (2) Based on the style feature self-supervised module, the present invention projects the view style features of the same image after different data augmentations and narrows the distance between them, which can effectively inhibit the learning of environmental features such as illumination and chromaticity by the style feature extractor, enabling the model to focus more on the common style features of different views, such as unchanged moiré information and unchanged paper texture information, etc. By optimizing the style feature extractor through the style feature self-supervised module, it can more effectively extract style features related to face fraud clues, thereby improving the generalization ability of the model.

[0073] (3) Based on the recombinant contrast loss module, the present invention recombines the content features and style features, and converts the traditional real and fake face binary classification problem into a ternary classification problem of screen attack face, printed attack face, and real face. By narrowing the distance between faces of the same attack method and widening the distance between faces of different attack methods and real faces, the three different categories are separated in the spatial domain, facilitating the discrimination of the subsequent classifier. Compared with the traditional method, the present invention can force the model to learn the differences between printed attack faces and screen attack faces, enhance the model's understanding ability of fraud clues, and thereby improve the generalization ability of the model.

[0074] (4) The network of the present invention has good generalization performance when facing invisible attacks, and belongs to a lightweight network. The data augmentation module greatly reduces the demand for the number of face datasets of the model, and is suitable for deployment on devices such as mobile phones and embedded terminals, with high practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 It is a schematic flow diagram of the live detection method based on fraud face enhancement and style feature self-supervision of the present invention;

[0076] Figure 2 It is a schematic diagram of the overall network architecture of the live detection system based on fraud face enhancement and style feature self-supervision of the present invention;

[0077] Figure 3 It is a schematic diagram of the data augmentation effect of the screen attack face generator of the present invention;

[0078] Figure 4 It is a schematic diagram of the data augmentation effect of the printed attack face generator of the present invention;

[0079] Figure 5 It is a schematic diagram of the data augmentation effect of the real face enhancer of the present invention;

[0080] Figure 6(a) is a schematic diagram of the network structure of the content feature extractor of the present invention;

[0081] Figure 6(b) is a schematic diagram of the network structure of the residual block of the present invention;

[0082] Figure 7 It is a schematic diagram of the network structure of the style feature extractor of the present invention;

[0083] Figure 8 It is a schematic diagram of the network structure of the style feature self-supervision module of the present invention;

[0084] Figure 9 It is a schematic diagram of the operation of the recombination contrast loss module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0085] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0086] Embodiment 1

[0087] Taking the training and testing of this embodiment using four publicly available live detection video datasets, namely Replay-Attack, CASIA-FASD, MSU_MFSD, and OULU, as an example, the implementation process of this embodiment will be introduced in detail. The Replay-Attack database contains 1200 videos. If the file path contains "real", it is a real face; if the file path contains "print", it is a printed attack face; if the file path contains "video", it is a screen attack face. The CASIA-FASD dataset contains 600 videos. If the folder ends with 1, 2, or HR_1, it is a real face; if the folder ends with 3, 4, 5, 6, HR_2, or HR_3, it is a printed attack face; if the folder ends with 7, 8, or HR_4, it is a screen attack face. The MSU-MFSD dataset contains 280 videos. If the file path contains "scene01", it is a real face; if the file path contains "printed", it is a printed attack face; if the file path contains "video", it is a screen attack face. The OULU dataset contains 4950 videos. If the folder ends with 1, it is a real face; if the folder ends with 2 or 3, it is a printed attack face; if the folder ends with 4 or 5, it is a screen attack face.

[0088] This embodiment runs on the Ubuntu system of Linux and is implemented using the Pytorch deep learning framework. The graphics card model is NVIDIA GeForce RTX 4090 GPU, the Python version is 3.10.6, and the Pytorch version is 2.5.1+cu124.

[0089] As Figure 1 shown, this embodiment provides a live detection method based on fraud face enhancement and style feature self-supervision, including the following steps:

[0090] S1. Construct a data preprocessing module;

[0091] In this embodiment, the ffmpeg program is used to frame the video dataset for subsequent image-level processing. The MTCNN algorithm is used to detect the target face area, generate a position box and crop it. The Lanczos interpolation algorithm is used to adjust the image resolution to H×W×C (256×256×3 in this embodiment), where H, W, and C represent the height, width, and number of channels of the image;

[0092] In this embodiment, the publicly available dataset is divided. After extracting frames from the face videos, face images are intercepted, and text files containing face images, corresponding true / false labels, and specific attack methods are constructed according to different attack methods. The attack methods include screen attacks, printed attacks, etc.;

[0093] S2. Construct a screen attack face generator;

[0094] As Figure 3 shown, input the real face image into the screen attack face generator. Generate the screen attack face by adding moiré patterns, screen reflections, pixelation, light spots, trailing artifacts, and halos. The moiré patterns and screen reflections are weighted and superimposed on the real face using the constructed moiré pattern database and shooting background database; Pixelation generates large pixels to simulate the low-resolution scenario of the display screen by shrinking and then enlarging the image; Light spots simulate the light reflection on the display screen through random distribution and blurring effects; Trailing artifacts simulate the effects caused by shooting and movement of the display screen through horizontal stripes and motion blur. The specific process is as follows:

[0095] As Figure 3 in (a), it is the original real face image. As Figure 3 in (b), for adding moiré patterns, generate an image with moiré patterns by simulating pixel misalignment between the digital display and the camera, and superimpose and fuse the real face image and the moiré pattern image according to different weight ratios to generate a face with moiré patterns.

[0096] As Figure 3 in (c), for adding screen reflections, collect the image during the shooting of the camera and blur the face information on the image to obtain the screen reflection background image, and superimpose and fuse the real face image and the screen reflection background image according to different weight ratios to generate a face with screen reflections.

[0097] As Figure 3 in (d), for pixelation, first shrink the image to a certain ratio, and then use nearest neighbor interpolation to re-enlarge the image to the original size to generate a face with pixelation effect.

[0098] As Figure 3 in (e), for light spots, simulate the light spot reflection on the smooth display screen through random distribution and blurring effects to generate a face with light spots.

[0099] As Figure 3 in (f), for trailing artifacts, add horizontal stripes to the face image to simulate the scan line defect of the screen display, disrupt the face details by controlling the transparency, and simulate the motion blur caused by the slow screen response time by increasing the directional blur of the operator to generate a face with trailing artifacts.

[0100] As Figure 3 in (g), for halos, create a halo layer with a central gradient, use Gaussian blur to spread it, simulate the light diffusion around the highlighted area in the screen display, and superimpose it on the face image to generate a face with halo effect.

[0101] S3. Build a printed attack face generator;

[0102] As Figure 4 shown, input the real face image into the printed attack face generator, and generate a printed attack face by adding ink traces, SFC - halftone, BN - halftone, and color distortion. The ink traces spread in a circular area by selecting the center point and use Gaussian blur to smooth the circular edge; SFC - halftone maps through the halftone matrix after adjusting the gray level of each pixel channel to obtain the SFC - halftone pattern, and then fuses the SFC - halftone pattern with the face to simulate printing distortion; BN - halftone fuses the BN - halftone pattern and the face with different weights; color distortion converts RGB from the original color domain to the CMYK color domain of different printers and then converts it back to the original RGB color domain to simulate color distortion during printing. The specific process is as follows:

[0103] As Figure 4 in (a), it is the original real face image. As Figure 4 in (b), for the ink traces, several diffusion points are randomly generated on the face image to simulate the effect of inkjet diffusion. The diffusion range simulates edge gradient through Gaussian blur and then is superimposed on the original face to generate a face with an ink smudging effect.

[0104] As Figure 4 in (c), for SFC - halftone, the face image is reduced, the RGB channels are separated, and corresponding halftone patterns are generated for each channel. After restoring the pattern to the original size, they are superimposed to generate an SFC - halftone face with printing characteristics.

[0105] As Figure 4 in (d), for BN - halftone, an image from the texture background (blue noise) database is randomly selected, cropped and adjusted in size to match the face image, and then superimposed with different transparencies to generate a BN - halftone face with printing characteristics.

[0106] As Figure 4 in (e), for color distortion, load the color profile of the face, map the image from the sRGB domain to the CMYK profile through ImageCms.profileToProfile, and then remap it back to the target RGB profile to simulate the resampling during photographing the printed attack face image;

[0107] S4. Build a real - face enhancer to convert the real face image in the RGB domains of different shooting devices to simulate the differences of different shooting devices;

[0108] AsFigure 5 As shown in Figure 5 , the real face image is input into the real face enhancer, and the diversity of the face image is enhanced by transforming the color gamut. For example, the color is changed by increasing the mainstream RGB domain of the image to simulate real faces captured by different devices. The specific process is as follows:

[0109] As Figure 5 shown in (a) of Figure 5 , for the original real face image, read the face image and load all ICC color profiles under the RGB configuration file, and map the image from one color profile to another (avoid mapping to itself) through the color_deversity function. As Figure 5 shown in (b)-(e) of Figure 5 , multiple real face images are generated;

[0110] S5. Construct a content feature extractor;

[0111] In this embodiment, the content feature extractor is used to extract the content features of the face image. The input face image refers to the face image enhanced by different attacks and the face image in the original database. The size of the face image and the images in the original database are both 256×256×3. All images in the same dataset and the images after data augmentation are regarded as a data domain. For example, the CASIA-FASD dataset can be regarded as data domain 1, and the OULU dataset can be regarded as data domain 2;

[0112] The specific process is as follows: As shown in Fig. 6(a), the content feature extractor is divided into three parts. First, the 256×256×3 image is input into a network composed of 7×7 convolution and max pooling for preliminary feature extraction to generate preliminary features of 64×64×64. Then, it is input into a network composed of three groups of residual blocks for intermediate feature extraction to generate intermediate features of 16×16×256. The network structure of the residual block is shown in Fig. 6(b). Finally, it is input into a network composed of 3×3 convolution, BN layer, and ReLU function for final feature extraction to generate content features of 8×8×256. During the content feature extraction process, by normalizing each sample batch, the BN layer can alleviate the problem of gradient vanishing or explosion, promote the stability of network training, and accelerate the convergence of the model. The formula is as follows:

[0113]

[0114]

[0115] where μ batch represents the mean of all samples in the batch, σ batch represents the standard deviation of all samples in the batch, ∈ is a small constant to prevent division by zero error during calculation, and γ and β are scaling factors and offsets to be learned to ensure the network's expression ability for content features;

[0116] S6. Construct a style feature extractor;

[0117] In this embodiment, the style feature extractor is used to extract the style features of face images. The input face images refer to the face images enhanced by different face attacks and the face images in the original database. The size of the face images and the images in the original database is 256×256×3. All the images in the same dataset and the images after data augmentation are regarded as a data domain. For example, the CASIA-FASD dataset can be regarded as data domain 1, and the OULU dataset can be regarded as data domain 2;

[0118] Such as Figure 7 As shown, the style feature extractor is divided into three parts. First, the 256×256×3 image is input into a network composed of 7×7 convolution and max pooling for preliminary feature extraction, generating preliminary features of 64×64×64. Then, it is input into a multi-scale network composed of three groups of residual blocks and a feature pyramid structure for intermediate feature extraction, generating three feature maps of 64×64×64, 32×32×128, and 16×16×256. After the three feature maps are subjected to feature fusion operations of 3×3 convolution and IN layer, multi-scale features of 16×16×256 are generated. Finally, it is input into a network composed of 3×3 convolution, IN layer, max pooling layer, and fully connected layer for final feature extraction, generating style features of 256. Among them, the feature pyramid will aggregate the shallow features output by the three groups of residual blocks, further improve the feature extraction ability of the image through multi-scale feature extraction, and obtain the final style features through the IN layer. The formula is as follows:

[0119]

[0120]

[0121] Among them, μ inst represents the mean of the sample feature map, σ inst represents the standard deviation of the sample feature map, ∈ is a small constant to prevent division by zero error during calculation, and γ and β are scaling factors and offsets to be learned to ensure the network's expression ability for style features;

[0122] S7. Construct a domain discriminator module;

[0123] In this embodiment, the domain discriminator module makes the content features extracted by the content feature extractor inseparable in different domains, and makes the content features extracted in different domains similar through adversarial learning and gradient reversal;

[0124] The specific process is as follows. The domain discriminator module enables the model to learn general content features through adversarial learning. It is constructed by the domain discriminator and the gradient reversal layer. The content features of 8×8×256 are reduced to features of 256 through the adaptive average pooling layer and the Reshape layer, and the output features of size M are obtained through the double-layer fully connected layer, the ReLU function activation layer, and the Dropout layer. The domain discriminator module composed of the discriminator and the gradient reversal layer realizes the effect of adversarial learning of the model to ensure that the content features extracted by the content feature extractor are indistinguishable in different domains. Different domain annotations (i.e., Figure 2 data domain 1 and data domain 2) of different databases are used to supervise the output of the domain discriminator module. Since adversarial learning is unstable in the initial stage and the model performance is easily interfered by factors such as noise, a dynamic coefficient λ is introduced, which can effectively suppress the influence of the initial noise signal in adversarial learning on the training stability of the model. The formula is as follows:

[0125]

[0126] Among them, L adv (G, D) represents the supervision loss, G represents the content feature extractor, D represents the domain discriminator, x represents the input image, y represents the region label of the image, M represents the number of source domains, represents the indicator function, λ represents the dynamic coefficient, cur_iter represents the number of current training iterations, and sum_iter represents the total number of iterations;

[0127] S8. Construct a style feature self-supervised module to narrow the distance between the view style features of the same image after different data augmentations, enable the model to learn the attack information between style features, and suppress the learning of environmental information such as illumination and hue;

[0128] As Figure 8 shown, the style feature self-supervised module narrows the distance between different views of the same image to force the style feature extractor to extract fraud clue features. Specifically, the face image image after screen enhancement, print enhancement, or real enhancement is generated into two sub-views view1 and view2 through different basic data augmentations such as cropping, rotation, and color transformation, and the corresponding style features style are obtained through the style feature extractor G and Subsequently, they are mapped into a new feature space by the projector Proj for subsequent feature operations. The similarity of the two style features is calculated through the function, and the L self-supervisedThe loss function reduces the distance between different views of the same image, suppressing environmental interferences in the style features learned by the model, such as light sensitivity and chromaticity changes, enabling the style feature extractor to learn style information that does not change due to data augmentation, such as moiré patterns and paper textures, thereby enhancing the model's ability to express fraud clues. The specific formula is as follows:

[0129]

[0130] Among them, view1 and view2 represent different data augmentations of the same face image, G style represents the style feature extractor, Proj represents the projector, L self-supervised is the self-supervised loss optimization network, and τ is the temperature coefficient;

[0131] S9. Construct a recombined contrastive loss module to recombine the content features and style features, reduce the distance between faces of the same attack method, and increase the distance between faces of different attack methods, real and fake faces;

[0132] As Figure 9 shown, the recombined contrastive loss module recombines different content features and style features, and performs subsequent reduction and increase operations on different recombined features in the feature space. Specifically, the recombined contrastive loss module uses the AdaIN layer and convolutional operations, as well as residual mapping for combination to generate recombined features, and reduces the distance between self-assembled and recombined features of the same attack category, and increases the distance between self-assembled and recombined features of different attack categories. First, the style features are globally averaged to obtain a global description of each channel, and then two parameters γ and β are generated through a multi-layer perceptron. The content features are normalized by applying AdaIN after 3x3 convolution and adjusted by γ and β, and a new feature z is obtained through the ReLU function activation layer. Finally, the feature z is convolved by 3×3 and the adjustment of AdaIN, γ, and β is applied again, and added to the original content features to obtain the final style assembled feature. According to the above operations, the feature composed of the original content features and style features is called the self-assembled feature, and the one composed of the original content features and other style features is called the recombined feature. By reducing the distance between recombined features of the same category and increasing the distance between recombined features of different categories for subsequent classification. For example, (real face content feature, real face style feature) is the self-assembled feature, and the distance between the recombined features of (real face content feature, screen attack face style feature) and (real face content feature, printed attack face style feature) is increased. The specific formula is as follows:

[0133]

[0134] Among them, μ(x) and σ(x) represent the mean and standard deviation of the channel dimension, and γ and β represent the affine parameters generated from the style input. y iWhen it is equal to 1, it represents samples of the same category, and the goal is to make the features of the same category closer. y i When it is equal to 0, it represents samples of different categories, and the goal is to make the features of different categories farther apart. m is the margin coefficient to ensure that there is at least a certain distance between the features of different categories. T anchor represents the anchor point, that is, the self-assembled feature. T positive represents the re-assembled feature of the same style as the anchor point. T negative represents the re-assembled feature of a different style from the anchor point;

[0135] S10. Build a classifier, input the self-assembled features to obtain the predicted values of real faces, printed attack faces, and screen attack faces;

[0136] In this embodiment, the classifier classifies and predicts real faces, screen attack faces, and printed attack faces. The specific formula is as follows:

[0137]

[0138] Among them, N represents the total number of samples, C represents the number of categories (here C = 3, representing real faces, printed attack faces, and screen attack faces), y i,c represents the category label, p i,c represents the probability that the model predicts that sample i belongs to category c;

[0139] S11. Build a network model training module to supervise the four output results of the domain discriminator module, the style feature self-supervised module, the re-assembled contrast loss module, and the classifier, obtain the domain discriminator loss, the style feature self-supervised loss, the re-assembled contrast loss, and the classification loss, and then obtain the total loss function. The model is trained using the total loss function;

[0140] In this embodiment, the network model training module inputs the face images into the neural network for training, constrains the model optimization by calculating the total loss function, and saves the optimal network model and corresponding weights. Specifically, first, the RGB face images are input into the network model for training. The training process is optimized using the Adam optimizer, and the initial learning rate is set to 1×10 -4 , the exponential decay rate β1 of the first moment estimate is 0.9, and the exponential decay rate β2 of the second moment estimate is 0.999. By minimizing the total loss function L total constrain the network and update the weight parameters of the network model. After training, save the best network model and corresponding weights according to the results of the benchmark test;

[0141] The specific formula of the total loss function L total involved in this process is as follows:

[0142] L total = L adv(G, D) + αL rcl + βL ce + γL self-supervised

[0143] where L adv (G, D) represents the domain discriminator loss, L rcl represents the reconstruction contrast loss, L ce represents the classifier loss, L self-supervised represents the style feature self-supervised loss, and α, β, and γ are the weight coefficients of the three loss functions. In this embodiment, α, β, and γ are set to 0.5, 1, and 0.1;

[0144] S12. Use the saved optimal model and weights for testing;

[0145] In this embodiment, load the optimal model and weights obtained from training to construct a test network, input the images of the test data set for network model testing, and measure the performance of the network model according to the benchmark metrics;

[0146] The specific process is as follows: load the trained model and weights, and input the cropped test set faces into the model, and verify the model performance through various benchmark metrics. The benchmark metrics used include False Positive Rate (FPR), False Negative Rate (FNR), True Positive Rate (TPR), Half Total Error Rate (HTER), and Area Under Curve (AUC);

[0147] TP, FP, FN, and TN are the four basic metrics for calculating the false positive rate and false negative rate; TP represents when the label is true and the prediction is true, FP represents when the label is false and the prediction is true, FN represents when the label is true and the prediction is false, and TN represents when the label is false and the prediction is false;

[0148] The false positive rate represents the proportion of false faces misjudged as real faces among all false faces. The specific calculation formula is as follows:

[0149]

[0150] The false negative rate represents the proportion of real faces misjudged as false faces among all real faces. The specific calculation formula is as follows:

[0151]

[0152] The true positive rate represents the proportion of real faces that are also judged as real faces among all real faces. The specific calculation formula is as follows:

[0153]

[0154] The semi-total error rate represents the average of the false positive rate and the false negative rate. The smaller the semi-total error rate, the smaller the false alarm rate and the miss rate, and the higher the model accuracy. The specific calculation formula is as follows:

[0155]

[0156] The receiver operating characteristic curve plots the false positive rate on the abscissa and the true positive rate on the ordinate. The larger the area under the curve, the better the performance of the model;

[0157] In this embodiment, cross-database experiments were carried out on four publicly available video databases, namely CASIA-MFSD, Replay-Attack, MSU-MFSD, and OULU, and tested according to the commonly used protocols in the field of face fraud detection. Specifically, the four groups of experiments were set as the OCI library crossing the M library, the OMI library crossing the C library, the OCM library crossing the I library, and the ICM library crossing the O library. The optimal experimental results are shown in the following table;

[0158] Table 1 Cross-database test HTER

[0159]

[0160] Table 2 Cross-database test AUC

[0161]

[0162] This embodiment is compared with the PatchNet and IADG methods. The PatchNet algorithm was published in CVPR in 2022, which proposed a simple algorithm for face anti-counterfeiting detection through refined patch recognition; the IADG algorithm was published in CVPR in 2023, which proposed a method to enhance the cross-domain generalization ability of the face fraud model by removing domain-specific style information and using style adaptation Liao Zheng. These two papers published in top conferences are both representative in the field of face fraud detection.

[0163] Through the test results in Table 1 and Table 2, the proposed invention achieved the best results in the experiments of OCI library crossing M library and OCM library crossing I library, with the semi-total error rates being 5.30% and 6.92% respectively. At the same time, it also achieved quite competitive results in the experiments of OMI library crossing C library and ICM library crossing O library, with the semi-total error rates being 8.84% and 10.32% respectively. Correspondingly, the area under the receiver operating characteristic curve achieved the best results in the experiments of OMI library crossing C library and OCM library crossing I library, being 97.05% and 98.01% respectively. It also achieved quite competitive results in the experiments of OCI library crossing M library and ICM library crossing O library, with the area under the receiver operating characteristic curve being 98.32% and 95.97% respectively. The experiments prove that the method of this embodiment has excellent generalization performance when facing invisible attacks.

[0164] As Figure 2 shown, this embodiment provides a live detection system based on fraud face enhancement and style feature self-supervision, including: a data preprocessing module, a screen attack face generator, a printed attack face generator, a real face enhancer, a data domain construction module, a content feature extractor, a style feature extractor, a domain discriminator module, a style feature self-supervision module, a recombination contrast loss construction module, a classifier, a network model training module, and a prediction result output module;

[0165] In this embodiment, the data preprocessing module is used to divide the data set, extract frames from the face video and then intercept face images, and construct face images and corresponding authenticity labels and attack mode labels; the screen attack face generator is used to generate screen attack face images; the printed attack face generator is used to generate printed attack face images; the real face enhancer is used to perform data enhancement on real face images; the data domain construction module is used to construct the face images in the original data set, as well as screen attack face images, printed attack face images, and enhanced real face images into a data domain; the content feature extractor is used to extract the content features of the face images in the data domain; the style feature extractor is used to extract the style features of the face images in the data domain; the domain discriminator module is used to make the content features extracted from different domains similar through adversarial learning and gradient reversal, and construct a domain discriminator loss; the style feature self-supervised module is used to narrow the distance between the same face image in different views, learn the attack information between style features, and construct a style feature self-supervised loss; the recombination contrast loss construction module is used to recombine the content features and style features, narrow the distance between face images of the same attack mode, and increase the distance between face images of different attack modes, real and fake face images, and construct a recombination contrast loss; the classifier classifies and predicts real face images, screen attack face images, and printed attack face images, and constructs a classification loss; the network model training module is used to train based on the total loss to obtain an optimal network model, and the total loss is constructed based on the domain discriminator loss, style feature self-supervised loss, recombination contrast loss, and classification loss; the prediction result output module is used to output the prediction result based on the optimal network model.

[0166] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A live detection method based on fraud face enhancement and self-supervised style features, characterized in that, It includes the following steps: Divide the dataset, extract frames from the face video and then intercept face images, and construct face images and corresponding true / false labels and attack mode labels; Generate screen attack face images and printed attack face images, and perform data augmentation on real face images; Construct a data domain from the face images in the original dataset, as well as screen attack face images, printed attack face images, and augmented real face images; Extract the content features and style features of the face images in the data domain; Through adversarial learning and gradient reversal, make the content features extracted from different domains similar, and construct a domain discriminator loss; Reduce the distance between the same face image in different views, learn the attack information between style features, and construct a style feature self-supervised loss; Recombine the content features and style features, reduce the distance between face images with the same attack mode, and increase the distance between face images with different attack modes, true / false face images, and construct a recombined contrast loss; Based on the classifier, classify and predict real face images, screen attack face images, and printed attack face images, and construct a classification loss; Construct a total loss based on the domain discriminator loss, style feature self-supervised loss, recombined contrast loss, and classification loss; Train based on the total loss to obtain an optimal network model, and output a prediction result based on the optimal network model.

2. The live detection method based on fraud face enhancement and style feature self-supervision according to claim 1, wherein Generate screen attack face images and printed attack face images, and perform data augmentation on real face images, specifically including: Generate screen attack faces by adding moiré patterns, screen reflections, pixelation, light spots, motion blur artifacts, and halos; Generate printed attack faces by adding ink traces, SFC - halftone, BN - halftone, and color distortion; Convert real face images in the RGB domain of different shooting devices, and simulate real faces taken by different devices by changing the color gamut.

3. The live detection method based on fraud face enhancement and self-supervised style features according to claim 1, wherein, Extract the content features of the face images in the data domain, specifically including: Input the face image into a network composed of convolution and max pooling for preliminary feature extraction to generate preliminary features, then input it into a network composed of three groups of residual blocks for intermediate feature extraction to generate intermediate features, and finally input it into a network composed of convolution, BN layer, and ReLU function for final feature extraction to generate content features; The processing of the BN layer is expressed as: Among them, μ batch represents the mean of all samples in the batch, σ batch represents the standard deviation of all samples in the batch, ∈ is a small constant to prevent division-by-zero errors during calculation, γ and β are scaling factors and offsets to be learned, and H, W, and C represent the height, width, and number of channels of the image.

4. The live detection method based on fraud face enhancement and self-supervised style features according to claim 1, wherein Extract the style features of the face images in the data domain, specifically including: Input the face image into a network composed of convolution and max pooling for preliminary feature extraction to generate preliminary features, then input it into a multi-scale network composed of three groups of residual blocks and a feature pyramid structure for intermediate feature extraction to generate three feature maps. After the three feature maps are fused through convolution and IN layer feature fusion operations, multi-scale features are generated, and finally input into a network composed of convolution, IN layer, max pooling layer, and fully connected layer for final feature extraction to generate style features; The processing of the IN layer is expressed as: Among them, μ inst represents the mean of the sample feature map, σ inst represents the standard deviation of the sample feature map, ∈ is a small constant to prevent division by zero errors during calculation, γ and β are scaling factors and offsets to be learned, and H and W represent the height and width of the image.

5. The live detection method based on fraud face enhancement and self-supervised style features according to claim 1, wherein Through adversarial learning and gradient reversal, make the content features extracted from different domains similar, and construct a domain discriminator loss, specifically including: The model learns general content features through adversarial learning. The content features pass through an adaptive average pooling layer and a Reshape layer, and output features are obtained through a two-layer fully connected layer, a ReLU function activation layer, and a Dropout layer. The domain discriminator loss is expressed as: Among them, L adv (G, D) represents the supervised loss, G represents the content feature extractor, D represents the domain discriminator, x represents the input image, y represents the region label of the image, M represents the number of source domains, represents the indicator function.

6. The live detection method based on fraud face enhancement and self-supervised style features according to claim 1, wherein In the initial stage of adversarial learning, a dynamic coefficient λ is introduced, which is expressed as: where cur_iter represents the number of current training iterations, and sum_iter represents the total number of iterations.

7. The live detection method based on fraud face enhancement and self-supervised style features according to claim 1, wherein Reduce the distance between the same face image in different views, learn the attack information between style features, and construct a style feature self-supervised loss, specifically including: The face image is generated into two sub-views view1 and view2 through data enhancement, and the corresponding style features are extracted and style features The projector is used to map the two style features into a new feature space. The function calculates the similarity of two style features through L self-supervised The loss function brings the distance between different views of the same image closer, which is specifically expressed as: Among them, G style represents the style feature extractor, Proj represents the projector, and L self-supervised is the self-supervised loss optimization network, and τ is the temperature coefficient.

8. The live detection method based on fraud face enhancement and self-supervised style features according to claim 1, wherein Recombine the content features and style features, reduce the distance between face images with the same attack method, increase the distance between face images with different attack methods, real and fake face images, and construct a recombined contrast loss, specifically including: The global description of each channel is obtained by global average pooling of the style features, and two parameters γ and β are generated through a multi-layer perceptron. The content features are normalized by applying the AdaIN layer after convolution, and adjusted by γ and β. The new feature z is obtained through the ReLU function activation layer. Finally, the feature z is convolved and the adjustment of the AdaIN layer, γ and β is applied again, and added to the original content features to obtain the final style assembly feature. The operation of the AdaIN layer is expressed as: where μ(x) and σ(x) represent the mean and standard deviation of the channel dimension, and γ and β represent the affine parameters generated from the style input. The recombined contrast loss is expressed as: Among them, y i = 1 indicates samples of the same category, and y i = 0 indicates samples of different categories. m is the interval coefficient, and T anchor represents the anchor point, that is, the self-assembly feature, and T positive represents the re-assembly feature with the same style feature as the anchor point, and T negative represents the re-assembly feature with a different style feature from the anchor point.

9. The live detection method based on fraud face enhancement and self-supervised style features according to claim 1, wherein Construct a classification loss, which is expressed as: Among them, N represents the total number of samples, C represents the number of categories, and y i,c represents the category label, and p i,c represents the probability that the model predicts that sample i belongs to category c.

10. A live detection system based on fraud face enhancement and self-supervised style features, characterized in that, Including: A data preprocessing module, a screen attack face generator, a print attack face generator, a real face enhancer, a data domain construction module, a content feature extractor, a style feature extractor, a domain discriminator module, a style feature self-supervised module, a recombined contrast loss construction module, a classifier, a network model training module, and a prediction result output module; The data preprocessing module is used to divide the dataset, extract face images after extracting frames from face videos, and construct face images and corresponding true / false labels and attack method labels. The screen attack face generator is used to generate screen attack face images. The print attack face generator is used to generate print attack face images. The real face enhancer is used to perform data augmentation on real face images. The data domain construction module is used to construct a data domain from the face images of the original dataset, as well as screen attack face images, print attack face images, and enhanced real face images. The content feature extractor is used to extract the content features of the face images in the data domain. The style feature extractor is used to extract the style features of the face images in the data domain. The domain discriminator module is used to make the content features extracted from different domains similar through adversarial learning and gradient reversal, and construct a domain discriminator loss. The style feature self-supervised module is used to reduce the distance between the same face image in different views, learn the attack information between style features, and construct a style feature self-supervised loss. The reconstructed contrast loss construction module is used to reconstruct the content features and style features, reduce the distance between face images of the same attack method, increase the distance between face images of different attack methods and between real and fake face images, and construct the reconstructed contrast loss; The classifier classifies and predicts real face images, screen attack face images, and printed attack face images to construct a classification loss; The network model training module is used to train based on the total loss to obtain an optimal network model, and the total loss is constructed based on the domain discriminator loss, style feature self-supervised loss, reconstructed contrast loss, and classification loss; The prediction result output module is used to output the prediction result based on the optimal network model.