General face living body detection method based on abnormality prompt enhanced guidance
By constructing a dual-input embedding fusion module and a hybrid expert model, and using a latent diffusion model to generate anomaly prompts, the problem of insufficient generalization ability of the liveness detection model across datasets is solved, achieving faster training convergence and more robust liveness detection.
Patent Information
- Application Number
- CN202511177867.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing liveness detection models lack generalization ability across datasets and fail to fully utilize the potential value of anomaly alerts, resulting in poor performance when faced with unseen datasets.
A dual-input embedding fusion module and an anomaly feature dynamic extraction module based on a hybrid expert model are constructed. Anomaly prompts are generated through a latent diffusion model, and features are dynamically extracted using a gating mechanism and an expert network, thereby enhancing the model's attention to and feature representation of anomaly regions.
It improves the model's cross-dataset generalization ability and training convergence speed, achieving more robust liveness detection results.
Smart Images

Figure CN120726703B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of cross-domain face liveness detection, and particularly relates to a general face liveness detection method based on abnormality prompt enhanced guidance. BACKGROUND
[0002] Due to the convenience and high accuracy of face recognition technology, it has been applied to various interactive intelligent applications, such as check-in and mobile payment, etc. However, the existing face recognition system is vulnerable to various presentation attacks, such as using someone's video or photo to deceive the system, thus posing a major security threat. Therefore, both academia and industry attach great importance to developing liveness detection technology to ensure the security of face recognition systems.
[0003] Previous research has achieved excellent results in the evaluation within the same dataset. However, most existing liveness detection datasets are small in size and lack diversity. No single dataset can cover all types of presentation attacks and various environmental changes, which leads to domain bias in model training. Therefore, models trained on a single dataset usually perform poorly in generalization when facing unseen datasets.
[0004] To address this challenge, many researches have focused on domain generalization techniques to enhance the model's generalization ability to unseen domains. Typically, the model is trained on multiple datasets simultaneously to find discriminative domain-invariant features for liveness detection. Some researches align the liveness detection features from different domains to make them indistinguishable in the domain. To further optimize feature alignment, other researches focus on decoupling domain-invariant features from domain-dependent features. In addition, meta-learning is also applied to help the model obtain better domain-invariant features by simulating domain migration scenarios.
[0005] Although domain generalization techniques effectively alleviate the problem of insufficient data diversity by training on multiple datasets, the performance of these methods is still constrained by the limited types of presentation attacks in the training dataset. To overcome this limitation, the AG-FAS method trains a pseudo face generator using existing large-scale real face data and uses the residual between the generated real face and the original input as an abnormality prompt to assist subsequent liveness detection. However, this method only applies the abnormality prompt directly to the cross-attention module of the ViT encoder layer, ignoring the preference of different stages of the encoder for semantic information and the attention level of different datasets to abnormal regions, thus limiting the model's generalization ability across datasets. In addition, this method also does not fully explore and utilize the potential value of the abnormality prompt. SUMMARY
[0006] The application aims to provide a general face liveness detection method based on abnormal prompt enhancement guidance, which realizes efficient use of abnormal prompts and further improves the cross-dataset generalization ability and training convergence speed of the model by constructing a double-input embedding fusion module and an abnormal feature dynamic extraction module based on a hybrid expert model.
[0007] To achieve the above-mentioned purpose, the technical scheme of the application is as follows: a general face liveness detection method based on abnormal prompt enhancement guidance, comprising the following steps:
[0008] S1, constructing an abnormal prompt generation module, a double-input embedding fusion layer, an abnormal feature dynamic extraction module based on a hybrid expert model, and a liveness detection module based on abnormal prompt enhancement guidance; the abnormal prompt generation module takes a latent diffusion model (LDM) as a backbone network and introduces an identity branch as a constraint condition for the reconstruction process; the double-input embedding fusion layer includes two symmetrical image block embedding layers; the abnormal feature dynamic extraction module includes an abnormal prompt encoder, a gating layer, and an expert network; the liveness detection module includes a ViT encoder and a classifier;
[0009] S2, inputting a face image to be detected;
[0010] The abnormal prompt generation module takes a face image as input, generates an abnormal prompt, and outputs the abnormal prompt;
[0011] The double-input embedding fusion layer takes a face image and an abnormal prompt as input, converts them into two vector sequences through two symmetrical image block embedding layers, and fuses the two vector sequences as input of the liveness detection module;
[0012] The abnormal feature dynamic extraction module takes an abnormal prompt as input, obtains deep semantic features through an abnormal prompt encoder as input features of the gating layer and the expert network, uses a gating mechanism to customize a routing scheme through the gating layer, performs a nonlinear transformation on the deep semantic features through the expert network, weights and sums the expert outputs according to the routing scheme, obtains a routing scheme output, and takes the routing scheme output as an additional input of the liveness detection module;
[0013] The liveness detection module extracts liveness detection features from the input features through the ViT encoder and sends them to the classifier, which distinguishes real faces from fake faces and outputs a detection result.
[0014] Preferably, the abnormal prompt generation module specifically includes an encoder in the LDM, a U-Net network, a decoder, and an identity feature extractor in the identity branch, and performs the following operations:
[0015] In the identity branch: the identity feature extractor extracts features from the input face image and obtains deep identity features ;
[0016] In the LDM backbone network: the encoder encodes the input face image into a clean latent code , then obtains a noisy latent code through a preset number of diffusion steps , and inputs the noisy latent code into a U-Net network, which takes the input deep identity feature as a conditional input to gradually reconstruct an optimized latent code from the noisy latent code , and the decoder converts the optimized latent code into a high-fidelity real face ; ; ;
[0017] Calculate the input face image , and the difference between the real face and each pixel point is filtered out by a threshold value to remove the pixel points with a difference less than the threshold value, and an abnormality prompt is obtained .
[0018] Preferably, the abnormality prompt generation module is trained using an open-source real face dataset, and the definition of the target function is as follows:
[0019]
[0020] Wherein, is an encoder responsible for embedding the input face image into a clean latent code ; represents the Gaussian noise added in the forward process, denotes a Gaussian distribution; is a U-Net network trained to predict noise ; denotes the time step; is an identity feature extractor; denotes the expected value, i.e. the average loss for all possible input images, noise and time steps.
[0021] Preferably, the double-input embedding fusion layer specifically performs the following operations:
[0022] The two symmetrical image block embedding layers respectively segment the input face image and the abnormality prompt into a plurality of fixed-size image blocks and flatten them, and then project them to a vector sequence of the same dimension;
[0023] The two vector sequences are fused by element-wise addition, and the fusion result is output as the input of the living body detection module ViT encoder.
[0024] Preferably, the abnormal feature dynamic extraction module specifically performs the following operations:
[0025] The abnormal prompt is sent into an abnormal prompt encoder and deep semantic features are obtained .
[0026] Through a gate , a routing scheme is obtained , so that the deep semantic features can adaptively respond to the needs of each layer of the ViT encoder for different data sets and different feature information:
[0027]
[0028]
[0029] Wherein, the gate represents a weight vector for assigning each expert the probability of being selected; represents a global average pooling operation; is a learnable weight matrix of the gate; is used to calculate the normalized probability distribution; is the total number of ViT encoder layers;
[0030] According to the routing scheme , the outputs of the experts are weighted and summed to obtain the specific input of each ViT encoder layer:
[0031]
[0032]
[0033] Wherein, represents the total number of experts, represents the output of the nth expert, and the superscript T represents transposition, represents the operation of all experts, represents a batch normalization operation, represents the output of the router scheme corresponding to the ith ViT encoder layer, represents the operation of the top selected experts, and the weights of the remaining experts are set to 0, is a preset value.
[0034] Preferably, the abnormality prompt encoder extracts features with ResNet-18 as the backbone network.
[0035] Preferably, each expert in the expert network is designed as an independent convolutional subnetwork, responsible for processing the input deep semantic features performing a nonlinear transformation, the convolutional subnetwork performs the following operations:
[0036]
[0037] wherein, represents the output of the expert, represents the operation combination, , respectively represent the linear rectifier function and convolution operation, represents the continuous application of the two compound operations.
[0038] Preferably, the ViT encoder of the living body detection module comprises a stack of ViT encoder layers, each ViT encoder layer comprising a self-attention module, a cross-attention module and a feedforward neural network in sequence.
[0039] Preferably, the living body detection module specifically performs the following operations:
[0040] The output of the dual-input embedding fusion layer is taken as the input of the first ViT encoder layer, the output of each ViT encoder layer is taken as the input of the next ViT encoder layer, and the routing scheme output of the abnormal feature dynamic extraction module is taken as an additional input of the cross-attention module in each ViT encoder layer;
[0041] The input features of each ViT encoder layer enter the self-attention module and output semantic-enhanced feature representations; the cross-attention module takes the enhanced feature representations as queries, introduces the routing scheme output as keys and values, fuses the abnormality prompt features into the current features, dynamically guides the model to focus on different abnormal areas; the fused features are input into the feedforward neural network, the feature expression ability is improved through nonlinear mapping, and the output is output; through a stack of ViT encoder layers to obtain living body detection features;
[0042] The extracted living body detection features are sent to the classifier for distinguishing real faces and fake faces.
[0043] Preferably, the classifier is a linear model realized by a fully connected layer.
[0044] Compared with the prior art, the present application has the following beneficial effects:
[0045] (1) Abnormal prompt provides additional prior: the input face and the generated abnormal prompt are simultaneously passed through a symmetrical block embedding layer for information fusion, guiding the model to focus on the global abnormal area, thereby accelerating the training convergence speed.
[0046] (2) Dynamic expert feature processing: combined with dynamic expert selection, sparse selection strategy and multi-gate mechanism, efficient calculation and robust feature learning are realized, and the semantic information preferences of different data sets and different stages of the encoder are explored.
[0047] (3) Abnormal prompt enhanced guidance: a customized routing scheme is used to provide specific abnormal prompt features for different stages of the encoder, and cross-attention is used to dynamically guide the model to focus on different abnormal areas, thereby extracting more robust live detection features. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 The flow chart of the face live detection method based on abnormal prompt enhanced guidance of the present application. DETAILED DESCRIPTION
[0049] The technical solutions of the present application will be specifically described below in combination with the drawings.
[0050] The present application provides a general face live detection method based on abnormal prompt enhanced guidance, as shown in the following steps: Figure 1
[0051] Step S1, abnormal prompt generation module construction, that is, constructing a pseudo face generator; it uses a latent diffusion model (LDM) as the backbone network, and by introducing an identity branch as a constraint condition in the reconstruction process, the consistency between the input and the generated face is enhanced.
[0052] The training objective function of the generator is defined as follows:
[0053]
[0054] wherein, is the input face image; is an encoder responsible for embedding the input face image into a clean latent code ; is the noisy latent code obtained through the forward process; represents the Gaussian noise added in the forward process, denotes a Gaussian distribution; is a U-Net network trained to predict the noise ; denotes the time step; It is an identity feature extractor that uses a pre-trained ResNet-18 to extract facial identity features (a common implementation in this field, which will not be elaborated here). This represents the expected value, which is the average loss over all possible input images, noise, and time steps.
[0055] The aforementioned pseudo-face generator is trained using only a large number of open-source real-world face datasets. The aim is to focus on learning the essential features of real faces, thereby enabling it to generate a corresponding "realistic version" of any given input face. Specifically, the generator first takes the input face image... Encoding as a clean latent code And through a preset number of diffusion steps Obtaining noisy latent coding Subsequently, U-Net used deep identity features from facial images. As a conditional input, a cross-attention module is introduced before each stage of downsampling and upsampling to fuse the features with the noisy latent code. Guided by the identity features, the model predicts the current noise and iteratively optimizes through stepwise denoising operations, ultimately reconstructing a latent code that highly matches the input face identity features. Finally, the decoder in the LDM is used to... Converted to a high-fidelity "real" face .
[0056] use Obtain the input face image and the "real" face. The degree of difference between individual pixels is determined. Then, threshold filtering is used to remove pixels with very small differences, reducing noise and ultimately yielding an anomaly alert closely related to subsequent steps. .
[0057] Step S2: Construction of the dual-input embedding fusion layer.
[0058] Error message As an additional input to the backbone network (ViT) in step S4, the purpose is to provide some prior knowledge to guide the model to focus on abnormal regions of the input image, thereby accelerating the training convergence speed.
[0059] Construct an image patch embedding layer identical to the backbone network ViT. These two symmetrical embedding layers are used to process the input face images. and error messages Specifically, the input is first divided into several fixed-size image patches and flattened, and then projected onto a vector sequence of the same dimension.
[0060] The two vector sequences are added element by element and fused together, and the final result is used as the input to the ViT encoder.
[0061] Step S3: Construct an anomaly feature dynamic extraction module based on a hybrid expert model, including an anomaly prompt encoder, a gating layer, and an expert network.
[0062] The exception message generated in step S1 Send error message encoder To obtain deep semantic features The encoder uses ResNet-18 as its backbone network.
[0063] In order to It can adaptively respond to the needs of each layer of the ViT encoder in step S4 for different datasets and different feature information, and uses a gating mechanism to customize a special routing scheme. Gate controller This represents a weight vector used to assign the probability of each expert being selected:
[0064]
[0065]
[0066] in, This indicates a global average pooling operation. It is the learnable weight matrix of the gate. Used to calculate the normalized probability distribution This represents the total number of ViT encoder layers. Each router has its own stage-specific preferences for customizing the appropriate expert mix.
[0067] The expert outputs are weighted and summed to obtain the specific input for each ViT encoder layer. In this method, experts are designed as independent convolutional subnetworks responsible for performing specialized nonlinear transformations on the input features, enabling the model to adapt to diverse tasks and heterogeneous input distributions. To maintain computational efficiency without sacrificing expressive power, lightweight convolution operations are employed.
[0068]
[0069] in, This represents the expert's output. Indicates the combination of operations. , Represent the linear rectified function and Convolution operation, This indicates that two consecutive compound operations are applied. Afterwards, we obtain the output for each routing scheme:
[0070]
[0071]
[0072] in, This represents the total number of experts. This represents the output of the nth expert, with the superscript T indicating transpose. This indicates that all expert operations were performed. This indicates a batch normalization operation. This represents the output of the router scheme corresponding to the i-th ViT encoder layer. This indicates that the top-ranked items are selected. ( For expert operations with a probability of 2, the weight of experts with a probability lower than this is set to 0. This sparse selection strategy ensures that only a subset of experts are activated, thereby reducing computational overhead.
[0073] Step S4: Construct a liveness detection module guided by anomaly alerts, using ViT as the backbone network, including a ViT encoder and a classifier; the ViT encoder includes... The ViT encoder layers are stacked, and each ViT encoder layer includes a self-attention module, a cross-attention module, and a feedforward neural network in sequence.
[0074] The output of step S2 is used as the input of the first ViT encoder layer, and the output of each ViT encoder layer is used as the input of the next ViT encoder layer. Each ViT encoder layer... Composed of a self-attention module, a cross-attention module, and a feedforward neural network, this system progressively models and optimizes the input features. First, the input features enter the self-attention module, which captures contextual information between positions in the vector sequence by constructing global dependencies, outputting a semantically enhanced feature representation. Next, the cross-attention module uses this feature as a query and incorporates the routing scheme generated in step S3 into its output. As keys and values, specific anomaly alerting features are further integrated into the current features, dynamically guiding the model to focus on different anomaly regions. Subsequently, the fused features are input into the feedforward neural network, where nonlinear mapping enhances feature representation capabilities, and the output is sent to the next encoder layer. Through the layer-by-layer collaboration of these modules, self-attention is responsible for modeling the global structure, cross-attention introduces external guidance information, and the feedforward network further processes and strengthens the features, thereby more effectively mining robust liveness detection features.
[0075] The extracted liveness detection features are fed into a classifier to distinguish between real and fake faces. This classifier is a linear model implemented by fully connected layers.
[0076] This module can be used in conjunction with most existing domain-generalization-based techniques to further enhance the model's cross-dataset liveness detection capabilities.
[0077] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.
Claims
1. A general face liveness detection method based on abnormity prompt enhanced guidance, characterized in that, The method comprises the following steps: S1, constructing an anomaly prompt generation module, a double-input embedding fusion layer, an abnormal feature dynamic extraction module based on a hybrid expert model, and a live detection module based on an anomaly prompt enhancement guide; the anomaly prompt generation module takes a latent diffusion model (LDM) as a backbone network and introduces an identity branch as a constraint condition for a reconstruction process; the double-input embedding fusion layer comprises two symmetrical image block embedding layers; the abnormal feature dynamic extraction module comprises an anomaly prompt encoder, a gating layer, and an expert network; the live detection module comprises a ViT encoder and a classifier; S2, inputting a face image to be detected; The anomaly prompt generation module takes a face image as input, generates an anomaly prompt, and outputs the anomaly prompt; The double-input embedding fusion layer takes a face image and an anomaly prompt as input, converts them into two vector sequences through two symmetrical image block embedding layers, and fuses the two vector sequences as input of the live detection module; The abnormal feature dynamic extraction module takes an anomaly prompt as input, obtains deep semantic features through an anomaly prompt encoder as input features of a gating layer and an expert network, uses a gating mechanism to customize a routing scheme through the gating layer, performs a nonlinear transformation on the deep semantic features through the expert network, weights and sums the expert outputs according to the routing scheme, obtains a routing scheme output, and takes the routing scheme output as an additional input of the live detection module; The live detection module extracts live detection features from the input features through a ViT encoder and sends the live detection features to a classifier, which distinguishes real faces from fake faces and outputs a detection result; The abnormal feature dynamic extraction module specifically performs the following operations: presenting an abnormality prompt sending the abnormality prompt to an encoder and obtaining deep semantic features ; By a gate obtaining a routing scheme to enable the deep semantic features to adaptively respond to the needs of each layer of the ViT encoder for different data sets and different feature information: where the gate represents a weight vector for assigning the selected probability of each expert; denotes the global average pooling operation; is the learnable weight matrix of the gate; is used to calculate the normalized probability distribution; is the total number of ViT encoder layers; According to the routing scheme The outputs of the experts are weighted summed to obtain the specific input for each ViT encoder layer: wherein, denotes the total number of experts, denotes the output of the nth expert, the superscript T denotes the transpose, denotes all expert operations, denotes the batch normalization operation, denotes the router scheme output corresponding to the ith ViT encoder layer, denotes the top expert operations are selected, the rest of the expert weights are set to 0, is a preset value.
2. The general face liveness detection method based on the abnormal prompt enhanced guidance according to claim 1, characterized in that, The anomaly prompt generation module specifically comprises an encoder in the LDM, a U-Net network, a decoder, and an identity feature extractor in the identity branch, and performs the following operations: In the identity branch: the identity feature extractor on the input face image perform feature extraction and obtain deep identity features ; In the LDM backbone network: the encoder encodes the input face image into a clean latent code After that, through a preset diffusion step number obtain a noisy latent code and input into the U-Net network, which takes the input deep identity feature as a conditional input, according to the noisy latent code Step by step, the optimized latent code is reconstructed , and the decoder converts into a high-fidelity real face ; Calculate Obtain the face image of the input person The difference between the real face After obtaining the difference degree of each pixel point between the real face and the face image of the input person, the pixel points with a difference less than a threshold are removed through threshold filtering to obtain an abnormality prompt .
3. The general face liveness detection method based on abnormal prompt enhanced guidance according to claim 2, characterized in that, The abnormal prompt generation module is trained using an open-source real face data set, and a target function is defined as follows: where, is the encoder, responsible for embedding the input face image into a clean latent code ; represents the Gaussian noise added in the forward process, denotes a Gaussian distribution; is a U-Net network trained to predict the noise ; denotes the time step; is the identity feature extractor; denotes the expected value, i.e. the average loss over all possible input images, noise and time steps.
4. The general face liveness detection method based on abnormal prompt enhanced guidance according to claim 1, characterized in that, The double-input embedding fusion layer specifically performs the following operations: The two symmetrical image block embedding layers respectively segment the input face image and the anomaly prompt into a plurality of image blocks of a fixed size, flatten the image blocks, and project the image blocks into vector sequences of the same dimension; The two vector sequences are added element by element to obtain a fusion result, which is output as input of a ViT encoder of the live detection module.
5. The general face liveness detection method based on the abnormal prompt enhanced guidance according to claim 1, characterized in that, The anomaly prompt encoder takes a ResNet-18 as a backbone network to extract features.
6. The general face liveness detection method based on abnormal prompt enhanced guidance according to claim 1, characterized in that, Each expert in the expert network is designed as an independent convolutional subnetwork, responsible for processing the input deep semantic features performing a non-linear transformation, the convolutional subnetwork performing the following operations: wherein, represents the output of an expert, represents an operation composition, , represent a linear rectification function and a convolution operation, respectively, represents the successive application of two compound operations.
7. The general face liveness detection method based on the abnormity prompt enhanced guidance according to claim 1, characterized in that, The ViT encoder of the live detection module comprises a stack of ViT encoder layers, each ViT encoder layer in turn comprising a self-attention module, a cross-attention module and a feed-forward neural network.
8. The general face liveness detection method based on the abnormal prompt enhanced guidance according to claim 7, characterized in that, The live detection module specifically performs the following operations: an output of a dual-input embedded fusion layer as an input of a first ViT encoder layer, an output of each ViT encoder layer as an input of a next ViT encoder layer, and an output of a routing scheme of the anomaly feature dynamic extraction module as an additional input of a cross-attention module in each ViT encoder layer The input features of each ViT encoder layer enter a self-attention module and output a semantic enhanced feature representation; a cross-attention module takes the enhanced feature representation as a query and introduces a routing scheme to output As keys and values, the abnormality prompt features are fused into the current features to dynamically guide the model to focus on different abnormal areas; the fused features are input into a feedforward neural network, the feature expression capability is improved through nonlinear mapping, and output is obtained; through The ViT encoder layers in the layer stack obtain the live detection features; The extracted live detection features are sent to the classifier to distinguish real faces from fake faces. 9.The general face liveness detection method based on the abnormal prompt enhanced guidance according to claim 1, wherein, The classifier is a linear model implemented by a fully connected layer.
Citation Information
Patent Citations
Face recognition model training method, face recognition method and device
CN117437522A
Face living body detection method and device based on use of RGB and depth map anomaly detection
CN117935378A