A generalized face anti-spoofing method based on style hybrid reconstruction

By separating features into content and style features, and adopting adversarial learning and contrast learning methods, the problem of performance degradation of traditional models in cross-domain scenarios is solved, and better generalization and robustness are achieved, and it is suitable for face anti-spoofing tasks.

CN120412113BActive Publication Date: 2025-08-29QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510913280.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-29
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

The existing face anti-spoofing domain generalization method has insufficient cross-domain adaptability. Traditional models rely too much on global statistical information, resulting in a significant decline in performance during target domain testing, and lack of methods to effectively capture local features and image statistics.

Method used

The mixed style recombination method is used to separate the features into content features and style features. Through adversarial learning, the content features are indistinguishable in different domains. Through comparative learning, the style features related to activity are emphasized, the style features related to specific domains are suppressed, and the style recombination is used to use adaptive instance normalization technology to improve the generalization ability of the model.

Benefits of technology

It improves the generalization ability and robustness of the model in the unseen domain, and enhances the accuracy and adaptability of the face anti-spoofing task in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412113B_ABST
    Figure CN120412113B_ABST
Patent Text Reader

Abstract

The present invention discloses a domain generalization method for face anti-spoofing based on style hybrid recombination, characterized in that it includes the following steps: S1: constructing a data processing model; S11: constructing an adversarial learning module; S12: constructing a contrastive learning module; S13: constructing a style assembly layer; S14: constructing a loss function; S2: training process. The present invention relates to the field of image processing technology, and in particular, to a domain generalization method for face anti-spoofing based on style hybrid recombination. The technical problem to be solved by the present invention is to provide a domain generalization method for face anti-spoofing based on style hybrid recombination, and proposes to separate complete features into content features and style features, and process them separately. The content features are made indistinguishable between different domains through adversarial learning, while the style features are emphasized with respect to activity-related information through contrastive learning, while suppressing domain-specific information. This method effectively improves the generalization ability of the model in unseen domains.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a face anti-spoofing domain generalization method based on style hybrid reconstruction. Background Art

[0002] With the rapid development of deep learning technology, domain generalization (DG) technology has gradually matured and has gradually become a core research direction for solving the problem of cross-domain model adaptability. In face anti-spoofing tasks, domain generalization technology aims to enable the model to learn domain-invariant features from multi-source domain data during the training phase, thereby maintaining robust performance in unknown target domains (such as different acquisition devices, attack types, or environmental conditions). However, existing methods face severe challenges in achieving cross-domain generalization: traditional models often over-rely on global statistical information, resulting in them implicitly learning features that are strongly related to a specific domain during source domain training, and experiencing significant performance degradation due to domain shift when tested in the target domain.

[0003] Although existing FAS methods achieve good performance in same-domain scenarios, the performance may degrade significantly in cross-domain settings. The main reasons are the limitations of training data and network capabilities, which cause the model to fall into dataset bias and thus have poor generalization ability in the new domain. To solve this problem, domain adaptation techniques are used to alleviate the differences between the source and target domains by using unlabeled target data. However, in most practical FAS scenarios, collecting sufficient unlabeled target data for training is inefficient. Therefore, domain generalization methods are proposed to achieve good generalization on unseen target domains. Domain generalization methods almost all achieve domain generalization on the complete feature representation, ignoring the subtle characteristics of global and local image statistics in FAS.

[0004] Currently, there is a lack of a domain generalization method that is more generalizable and can capture the unique properties of local features and image statistics. Summary of the Invention

[0005] The technical problem addressed by this invention is to provide a domain-generalization method for face anti-spoofing based on style hybrid reconstruction. This method proposes separating complete features into content features and style features, processing them separately. Content features are made indistinguishable across domains through adversarial learning, while style features are emphasized through contrastive learning, emphasizing activity-related information while suppressing domain-specific information. This approach effectively improves the model's generalization ability in unseen domains.

[0006] The present invention adopts the following technical solutions to achieve the invention objectives:

[0007] A face anti-spoofing domain generalization method based on style hybrid reconstruction, characterized by comprising the following steps:

[0008] S1: Construct data processing model;

[0009] S11: Construct adversarial learning module;

[0010] S12: Constructing a contrastive learning module;

[0011] S13: tectonic style assembly layer;

[0012] S14: construct loss function;

[0013] S2: training process;

[0014] S21: Input RGB images from different domains into the feature generator, and separate their complete feature representation into content features and style features. The style features are further divided into activity-related style features and domain-specific style features.

[0015] S22: The content features are fed into the Wasserstein generative adversarial network through the anti-learning module, making the content features of different domains indistinguishable;

[0016] S23: A contrastive learning strategy is used to emphasize style features related to activity and suppress features related to specific domains;

[0017] S24: Inputting style features into the Kolmogorov-Arnold network to obtain a more complete representation and the affine parameters required for adaptive instance normalization in the style reconstruction layer and ;

[0018] S25: Reorganize the correct style features and content features to improve the generalization and robustness of the data processing model.

[0019] As a further limitation of this technical solution, the adversarial learning module is specifically:

[0020] The parameters of the content feature generator are optimized by maximizing the adversarial loss function, while the parameters of the domain-specific feature generator are optimized in the opposite direction. The process is expressed as follows:

[0021] (1);

[0022] in: is the discriminator target;

[0023] is the generator target;

[0024] is the adversarial loss function;

[0025] Indicates that from the data distribution ( ) in the sampled input image , and the corresponding domain label expectations;

[0026] is the input image set , is the indicator function;

[0027] is a set of domain labels;

[0028] is the number of different data fields;

[0029] is a content feature generator;

[0030] is the domain feature generator;

[0031] To simultaneously optimize both the content feature generator and the domain feature generator, the Wasserstein generative adversarial network is used to replace the JS divergence with the Wasserstein distance to measure the difference between the generated distribution and the true distribution, providing a smoother gradient and ensuring that content features remain consistent across different domains.

[0032] For style information aggregation, due to the different scales of style features, a pyramid network is used to collect multi-layer features along the hierarchical structure.

[0033] As a further limitation of this technical solution, the comparative learning module is specifically:

[0034] Combining content features and style features to obtain self-assembly features and shuffle-assemble features ;

[0035] The given length is The input sequence, Represented as the input samples, of which , represents the randomly selected sample index, It is a random sample;

[0036] For self-assembly features , which is fed into the classifier and used with a loss function The true value of the binary ground truth signal is used for supervision;

[0037] For the shuffle-assemble feature , by using cosine similarity to measure their self-assembly features The cosine similarity calculation formula is as follows:

[0038] (2);

[0039] in: represent paradigm;

[0040] and Indicates two features to be compared;

[0041] Self-assembly features Set as the anchor point of the stylized feature space, perform gradient stopping on it to fix its position in the feature space, and then guide the shuffle-assembly feature according to the activity information To its corresponding anchor point Close or far away, so the contrast loss function It is expressed as follows:

[0042] (3);

[0043] in: Representative self-assembly characteristics ;

[0044] Represents the shuffle-assemble feature ;

[0045] Stopgrad() is an operation in the deep learning framework;

[0046] Indicates the length of a given sample sequence;

[0047] Evaluate and The active label consistency, the simulation process is as follows:

[0048] (4);

[0049] Where: label represents the binary classification label of the input sample, which is used to distinguish between living and attack samples.

[0050] As a further limitation of this technical solution, the style assembly layer realizes style transfer based on adaptive instance normalization technology. Given style features , content features , where the adaptive instance normalization technique formula is as follows:

[0051] (5);

[0052] in: and represent the channel mean and standard deviation respectively;

[0053] and The Kolmogorov-Arnold network is composed of the style input features in formula (6) Generated affine parameters;

[0054] To combine the content features and style characteristics Combined with the adaptive instance normalization technique and the convolution operator with residual mapping, a style assembly layer is established. The specific process is as follows:

[0055] (6);

[0056] Among them: KAN is a neural network model;

[0057] GAP is global average pooling;

[0058] and It is a 3×3 convolution kernel;

[0059] represents the convolution operation;

[0060] is an intermediate variable;

[0061] ReLU() is a piecewise linear activation function whose formula is ;

[0062] SAL() is the style assembly layer;

[0063] A shuffled style assembly method is proposed to form auxiliary stylized features for domain generalization;

[0064] Given a mini-batch of length The input sequence, Represents the input sample, and the content feature is represented as , and the style feature is represented as , therefore, the corresponding combined feature space The formula is as follows:

[0065] (7);

[0066] This formula means that given an input image , the corresponding content features in the corresponding style assembly layer and .

[0067] As a further limitation of this technical solution, the loss function of the model is as follows:

[0068] (8);

[0069] and There are two hyperparameters

[0070] Compared with the prior art, the advantages and positive effects of the present invention are:

[0071] By separating complete features into style features and content features, this method emphasizes activity-related style information and suppresses domain-specific style information. Furthermore, thanks to the multi-scale style features, a variety of discriminative features are collected. A contrastive learning module and an adversarial learning module are proposed. The contrastive learning module emphasizes activity-related style features while suppressing domain-specific features, while the adversarial learning module makes content features indistinguishable across domains, thereby improving the model's generalization capabilities.

[0072] This invention aims to tackle the face anti-spoofing task in deep learning. First, images from different domains are input into the feature extractor, and then their complete feature representation is separated into content features and style features. The content features are subjected to adversarial learning and input into the WGAN network, making the content features of different domains indistinguishable. The style features are further divided into style features related to activity and style features related to specific domains. At the same time, a contrastive learning strategy is adopted to emphasize style features related to activity and suppress features related to specific domains. Then, we input the style features into the KAN network to obtain a more complete representation and the affine parameters required for AdaIN adaptive instance normalization in the style reconstruction layer. γ and β , and then recombines the correct style and content features, thereby improving the generalization and robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 It is a framework diagram of the present invention.

[0074] Figure 2 Schematic diagram of the comparative learning module of the present invention. DETAILED DESCRIPTION

[0075] A specific embodiment of the present invention is described in detail below with reference to the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific embodiment.

[0076] The face anti-spoofing (FAS) task has achieved significant breakthroughs in its ongoing development. However, traditional FAS often relies on a complete image representation, neglecting the unique statistical information required for this task. To address this issue, a novel approach is proposed that separates the complete image representation into style and content information. A style hybrid assembly network is then proposed to extract and reconstruct distinct content and style features in the stylized feature space. A contrastive learning strategy is also employed to emphasize authenticity-related style information while suppressing domain-specific information. Furthermore, a real face discriminator is proposed to detect real or fake faces.

[0077] The present invention comprises the following steps:

[0078] S1: Construct data processing model.

[0079] This paper relies on the OULU-NPU, CASIA-MFSD, Replay-Attack and MSU-MFSD datasets, which contain a variety of real-world changes and aim to simulate various challenges in actual application scenarios. They provide a rich variety of attack types, including printing attacks, video replay attacks and 3D mask attacks, to evaluate and improve the performance of face anti-counterfeiting algorithms. The diversity and complexity of these datasets make them important resources for studying face anti-counterfeiting technology.

[0080] S11: Construct adversarial learning module.

[0081] In the field of face authentication (FAS), content information is typically represented by common factors, primarily semantic features and physical attributes. In FAS tasks, style information can be further divided into two components: domain-specific style information and liveness-related style information. Therefore, in our network architecture, content features and style features are extracted and captured separately through a two-stream approach.

[0082] Specifically, the feature generator acts as a shallow embedding network that is responsible for capturing low-level information at multiple scales. Subsequently, the content feature extractor and style feature extractor further process these features by using specific normalization layers such as batch normalization (BN) and instance normalization (IN).

[0083] For the aggregation of content information, we assume that the distribution differences between different domains are small. This assumption is based on two facts: first, samples from different domains all contain facial regions, and therefore they share a common semantic feature space; second, physical attributes (such as shape and size) are generally similar for both real and offensive face representations. Based on these observations, we employ an adversarial learning approach to ensure that the generated content features are indistinguishable between different domains. This strategy helps improve the model's generalization ability, enabling it to better adapt to different face anti-forgery scenarios.

[0084] In this way, our network is able to effectively separate and utilize content features and style features, thereby achieving higher accuracy and robustness in face anti-forgery tasks. We adopt an adversarial learning module to ensure that content features are indistinguishable and can therefore be used in different domains.

[0085] The adversarial learning module is specifically:

[0086] The parameters of the content feature generator are optimized by maximizing the adversarial loss function, while the parameters of the domain-specific feature generator are optimized in the opposite direction. The process is expressed as follows:

[0087] (1);

[0088] in: The discriminator's goal is to correctly classify the source domain of the content features. That is, through adversarial learning, the content features are constrained to be domain-independent representations, reducing the model's reliance on domain-specific statistical features and improving cross-domain generalization performance.

[0089] As the generator goal, by maximizing the discriminator loss, the generated content features Unable to correctly classify the discriminator into the original domain;

[0090] is the adversarial loss function, which is a comprehensive indicator of the performance of the generator and the discriminator;

[0091] Indicates that the data distribution Midsample input image , and the corresponding domain label expectations;

[0092] is the input image set , is the indicator function, when the domain predicted by the discriminator With the real domain label If the ground truth label is 1, otherwise it is 0. , then only hour , the rest are 0;

[0093] is a set of domain labels;

[0094] is the number of different data fields;

[0095] is a content feature generator;

[0096] is the domain feature generator;

[0097] To optimize both the content feature generator and the domain feature generator, the Wasserstein generative adversarial network (GAN) uses the Wasserstein distance instead of the Jensen-Shannon Divergence (JS divergence) to measure the difference between the generated distribution and the true distribution. This provides a smoother gradient and ensures that content features remain consistent across domains.

[0098] For style information aggregation, due to the different scales of style features, pyramid networks are used to collect multi-layer features along the hierarchical structure.

[0099] JS divergence is a similarity measure between probability distributions. It is a symmetric version of the Kullback-Leibler divergence (KL divergence) and has some better properties. For example, it is always non-negative and bounded. It measures the difference between the generated distribution and the true distribution, provides a smoother gradient, and ensures that content features remain consistent across different domains.

[0100] Pyramid Networks are a structural design that enhances model perception by fusing multi-scale features. They are widely used in computer vision tasks such as object detection and image segmentation. Their core concept is to mimic the human visual system's processing of multi-scale information. By extracting and fusing features at different levels, they improve the model's adaptability to complex scenes. For example, while rendering scene brightness primarily involves large-scale features, rendering material texture typically focuses on local-scale regions.

[0101] S12: Construct a contrastive learning module.

[0102] From the perspective of style features, a major obstacle is that domain-specific style features may mask life-related style features in cross-domain scenes, leading to misjudgment. To this end, a contrastive learning module is proposed to emphasize style features related to vividness and suppress domain-specific style features.

[0103] The contrastive learning module is specifically:

[0104] Combining content features and style features to obtain self-assembly features and shuffle-assemble features ;

[0105] The given length is The input sequence, Represented as the input samples, of which , represents the randomly selected sample index, It is a random sample;

[0106] The contrastive learning module drives the model to learn domain-invariant and highly discriminative feature representations by leveraging the differences between self-assembled features and shuffled assembled features.

[0107] For self-assembly features , which is fed into the classifier and used with a loss function The ground truth binary signal is the real and accurate data label or target value used to train, verify, and test the model. It represents the objective state of the data in the real world or the correct answer and is used for supervision.

[0108] For the shuffle-assemble feature , by using cosine similarity to measure their self-assembly features The cosine similarity calculation formula is as follows:

[0109] (2);

[0110] in: represent paradigm;

[0111] and represents the characteristics of two comparisons, which is equivalent to Mean square error of the normalized vector;

[0112] Self-assembly features Set as the anchor point of the stylized feature space, perform the gradient stop (stopgrad) operation on it to fix its position in the feature space, and then guide the shuffle-assembly feature according to the activity information To its corresponding anchor point , self-assembly characteristics Does not participate in gradient updates, only serves as a reference point, and uses back propagation to assemble features only by shuffling , forcing the model to adjust the style fusion strategy to adapt to the label constraint, close or far away, so the contrast loss function It is expressed as follows:

[0113] (3);

[0114] in: Representative self-assembly characteristics ;

[0115] Represents the shuffle-assemble feature ;

[0116] Stopgrad() is an operation in deep learning frameworks (such as PyTorch) that is used to stop the gradient from flowing through a certain computational node during training, thereby preventing the parameters corresponding to that node from being updated. Its core idea is to calculate the value normally during forward propagation, but set the gradient of that node to 0 during backpropagation.

[0117] Indicates the length of a given sample sequence;

[0118] Evaluate and The active label consistency, the simulation process is as follows:

[0119] (4);

[0120] Where: label represents the binary classification label of the input sample, which is used to distinguish between living and attack samples.

[0121] Through simple label consistency judgment, the similarity optimization direction of sample pairs is dynamically adjusted in contrastive learning.

[0122] The contrastive learning module emphasizes style features related to activity features and suppresses style features related to specific domains. This helps the model better learn activity-related features in cross-domain scenarios and helps improve the robustness and generalization ability of the model. Figure 2 shown.

[0123] S13: Shuffled Style Assembly.

[0124] The Shuffled Style Assembly layer implements style transfer based on Adaptive Instance Normalization (AdaIN). Specifically, given the style features , content features , where the adaptive instance normalization technique formula is as follows:

[0125] (5);

[0126] in: and represent the channel mean and standard deviation respectively;

[0127] and The Kolmogorov-Arnold network is composed of the style input features in formula (6) Generated affine parameters;

[0128] To combine the content features and style characteristics Combined with the adaptive instance normalization technique and the convolution operator with residual mapping, a style assembly layer (SAL) is established. The specific process is as follows:

[0129] (6);

[0130] KAN is a neural network model. The Kolmogorov–Arnold Network (KAN) is a neural network model based on the Kolmogorov–Arnold Representation Theorem (KART). It aims to more efficiently approximate complex multivariate functions using this mathematical theorem. Compared to traditional neural networks (such as MLPs and multilayer perceptrons), KAN can approximate complex functions with fewer parameters, reducing dependence on network depth and width, and facilitating analysis of the impact of input variables.

[0131] GAP stands for Global Average Pooling (GAP). Global Average Pooling (GAP) is an operation that reduces the dimensionality of feature maps. It averages all pixel values ​​of the feature map of each channel and outputs a single value. For example, the size of the input feature map is (Number of channels × height × width), after GAP, the output size is , that is, each channel corresponds to an average value;

[0132] and It is a 3×3 convolution kernel;

[0133] represents the convolution operation;

[0134] is an intermediate variable;

[0135] ReLU() is a piecewise linear activation function whose formula is , which is to perform threshold processing on the input value, setting negative values ​​to 0 and positive values ​​to remain unchanged. The introduction of nonlinear activation function enables the neural network to learn complex nonlinear mapping relationships and enhance the expressive power of the model;

[0136] SAL() is the Style Assembly Layer, which is built through the AdaIN layer and the convolution operator;

[0137] Style characteristics It not only contains information related to activity, but also contains domain-specific information, which may lead to domain bias during network optimization. To solve this problem, a shuffled style assembly method is proposed to form auxiliary stylized features for domain generalization;

[0138] Given a mini-batch of length The input sequence, represents the input sample (input image), where Content features are represented as , and the style feature is represented as , therefore, the corresponding combined feature space The formula is as follows:

[0139] (7);

[0140] This formula means that given an input image , the corresponding content features in the corresponding style assembly layer and .

[0141] Generating AdaIN parameters through KAN enables more flexible and adaptive feature normalization and adjustment of input data compared to traditional MLP (Multilayer Perceptron) models. Unlike AdaIN, which directly calculates statistics based on specific style features, KAN, through knowledge encoding and parameter generation mechanisms, learns a more comprehensive and abstract mapping from features to normalization parameters. This enables the network to better adapt to data of different types or scenarios, particularly when dealing with diverse or unknown style variations, potentially enhancing the model's generalization capabilities.

[0142] S14: Construct loss function.

[0143] The loss function of the model is as follows:

[0144] (8);

[0145] and are two hyperparameters used to balance the proportions of different loss functions.

[0146] S2: training process;

[0147] S21: Input RGB images from different domains into the feature generator, and separate their complete feature representation into content features and style features. The style features are further divided into activity-related style features and domain-specific style features.

[0148] S22: Content features are fed into a Wasserstein generative adversarial network (WGAN) through an anti-learning module, making content features from different domains indistinguishable.

[0149] S23: A contrastive learning strategy is used to emphasize style features related to activity and suppress features related to specific domains;

[0150] S24: Inputting style features into Kolmogorov-Arnold Networks (KAN) to obtain a more complete representation and the affine parameters required for adaptive instance normalization in the style reconstruction layer and ;

[0151] S25: Reorganize the correct style features and content features to improve the generalization and robustness of the data processing model.

[0152] The above disclosure is only a specific embodiment of the present invention, but the present invention is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present invention.

Claims

1. A face anti-spoofing domain generalization method based on style hybrid reconstruction, characterized by: The following steps are involved: S1: Construct data processing model; S11: Construct adversarial learning module; S12: Constructing a contrastive learning module; S13: tectonic style assembly layer; S14: construct loss function; S2: training process; S21: Input RGB images from different domains into the feature generator, and separate their complete feature representation into content features and style features. The style features are further divided into activity-related style features and domain-specific style features. S22: The content features are fed into the Wasserstein generative adversarial network through the anti-learning module, making the content features of different domains indistinguishable; S23: A contrastive learning strategy is used to emphasize style features related to activity and suppress features related to specific domains; S24: Inputting style features into the Kolmogorov-Arnold network to obtain a more complete representation and the affine parameters required for adaptive instance normalization in the style reconstruction layer and ; S25: Recombining correct style features and content features to improve the generalization and robustness of the data processing model; The adversarial learning module is specifically: The parameters of the content feature generator are optimized by maximizing the adversarial loss function, while the parameters of the domain-specific feature generator are optimized in the opposite direction. The process is expressed as follows: (1); in: is the discriminator target; is the generator target; is the adversarial loss function; Indicates that from the data distribution ( ) in the sampled input image , and the corresponding domain label expectations; is the input image set , is the indicator function; is a set of domain labels; is the number of different data fields; is a content feature generator; is the domain feature generator; To simultaneously optimize the content feature generator and the domain feature generator, the Wasserstein generative adversarial network is used to measure the difference between the generated distribution and the true distribution by using the Wasserstein distance, providing a smoother gradient so that the content features remain consistent across different domains. For style information aggregation, due to the different scales of style features, a pyramid network is used to collect multi-layer features along the hierarchical structure; The contrastive learning module is specifically: Combining content features and style features to obtain self-assembly features and shuffle-assemble features ; The given length is The input sequence, Represented as the input samples, of which , represents the randomly selected sample index, It is a random sample; For self-assembly features , which is fed into the classifier and used with a loss function The true value of the binary ground truth signal is used for supervision; For the shuffle-assemble feature , by using cosine similarity to measure their self-assembly features The cosine similarity calculation formula is as follows: (2); in: represent paradigm; and Indicates two features to be compared; Self-assembly features Set as the anchor point of the stylized feature space, perform gradient stopping on it to fix its position in the feature space, and then guide the shuffle-assembly feature according to the activity information To its corresponding anchor point Close or far away, so the contrast loss function It is expressed as follows: (3); in: Representative self-assembly characteristics ; Represents the shuffle-assemble feature ; Stopgrad() is an operation in the deep learning framework; Indicates the length of a given sample sequence; Evaluate and The active label consistency, the simulation process is as follows: (4); Where: label represents the binary classification label of the input sample, which is used to distinguish between living and attack samples.

2. The face anti-spoofing domain generalization method based on style hybrid reconstruction according to claim 1 is characterized by: The style assembly layer realizes style transfer based on adaptive instance normalization technology. , content features , where the adaptive instance normalization technique formula is as follows: (5); in: and represent the channel mean and standard deviation respectively; and The Kolmogorov-Arnold network is composed of the style input features in formula (6) Generated affine parameters; To combine the content features and style characteristics Combined with the adaptive instance normalization technique and the convolution operator with residual mapping, a style assembly layer is established. The specific process is as follows: (6); Among them: KAN is a neural network model; GAP is global average pooling; and It is a 3×3 convolution kernel; represents the convolution operation; is an intermediate variable; ReLU() is a piecewise linear activation function whose formula is ; SAL() is the style assembly layer; A shuffled style assembly method is proposed to form auxiliary stylized features for domain generalization; Given a mini-batch of length The input sequence, Represents the input sample, and the content feature is represented as , and the style feature is represented as , therefore, the corresponding combined feature space The formula is as follows: (7); This formula means that given an input image , the corresponding content features in the corresponding style assembly layer and .

3. The face anti-spoofing domain generalization method based on style hybrid reconstruction according to claim 2 is characterized by: The loss function of the model is as follows: (8); and are two hyperparameters.

Citation Information

Patent Citations

  • Face deception detection method based on domain adaptive learning and domain generalization

    CN110309798A

  • Face in-vivo detection method based on conditional adversarial domain generalization and network model architecture

    CN114078276A