Face living body detection method based on image style mixing

By using image-style mixing module to perform feature-level data enhancement in facial viscera detection, the problem of insufficient accuracy and robustness in the prior art is solved, efficient and accurate viscera detection is achieved, and hardware requirements are reduced.

CN120108047APending Publication Date: 2025-06-06YANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510013956.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When faced with complex and changing attack forms, existing facial live detection technology lacks accuracy and robustness, and some technologies have high requirements for hardware equipment, which increases deployment costs.

Method used

Using a face live detection method based on image style mixing, the image style mixing module is embedded in the first two residual blocks of the ResNet-18 neural network to perform feature-level data enhancement, improving the generalization ability and robustness of the model.

Benefits of technology

It improves the accuracy and efficiency of the facial live detection model, enhances the robustness to different scenarios and lighting conditions, can effectively resist unknown attacks, and reduces the cost of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108047A_ABST
    Figure CN120108047A_ABST
Patent Text Reader

Abstract

The invention discloses a face living body detection method based on image style mixing, and the method comprises the following steps: inputting an image, extracting primary features through a convolution layer, and generating an image feature map; deep features of a feature map are learned through a ResNet-18 neural network, an image style mixing module is integrated in the first two residual blocks to perform feature-level data enhancement, style information of different samples is mixed in a feature level, data diversity in the training process is increased, and generalized feature representation in the feature map is learned; using an average pooling layer to reduce the dimension of the feature space, and using a full connection layer to classify the authenticity of the image; and performing judgment through an output layer to obtain a face image authenticity judgment result. According to the method, feature-level data enhancement is carried out on the samples by using an image style mixing idea, the problem of data shortage is solved, hardware requirements are reduced by adopting a ResNet-18 lightweight model, and the accuracy and efficiency of a living body detection model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biometric identification technology, and in particular to a human face liveness detection method based on image style mixing. Background Art

[0002] In recent years, with the rapid progress of artificial intelligence and deep learning technology, face recognition technology has been widely used in various fields. From payment systems to public security, from smart homes to smart transportation, "face scanning" has gradually become an important part of people's daily lives. Face recognition has ushered in a blowout outbreak and is widely used in face payment, attendance, security and other aspects, bringing great convenience to life.

[0003] While face recognition technology is developing rapidly and penetrating into society, it also brings many security challenges. Under normal application conditions, face recognition algorithms can achieve very high recognition accuracy. However, they cannot resist some live attacks, such as high-definition printed photos, video replays, or 3D masks. These attacks not only threaten personal privacy, but may also cause property losses and even endanger public safety. These problems expose the limitations of traditional face recognition systems, which are unable to cope with complex and changing forms of attacks, and seriously affect the security and reliability of the system.

[0004] In order to solve this problem, face liveness detection technology is introduced to distinguish real living faces from various forms of attacks, which is crucial to ensuring the security and reliability of face recognition systems.

[0005] Although the current research on face liveness detection technology has achieved certain results, due to the development of counterfeiting technology, the attack methods have become more complex and covert, and the accuracy and robustness of liveness detection technology are facing great challenges. Secondly, the existing liveness detection model shows certain limitations when dealing with different lighting conditions, environmental changes, camera quality and other issues, resulting in insufficient generalization ability of the model. In addition, some liveness detection technologies have high requirements for hardware equipment, which increases the deployment cost and is not conducive to large-scale application. Therefore, how to improve the accuracy, generalization and deployment efficiency of liveness detection models has become the focus of current research. Summary of the invention

[0006] The problem to be solved by the present invention is to provide a method for face liveness detection based on image style mixing, so as to improve the accuracy and efficiency of a liveness detection model.

[0007] The present invention adopts the following technical solution: a method for face liveness detection based on image style mixing, comprising the following steps:

[0008] Step 1: Input face image data, construct the initial image data set and perform preprocessing;

[0009] Step 2: Input the initial image into the convolution layer, extract the primary features of the image, and generate an image feature map;

[0010] Step 3: Input the image feature map into the ResNet-18 neural network for deep feature learning. The first two residual blocks of the ResNet-18 neural network are integrated with an image style mixing module to perform feature-level data enhancement.

[0011] Step 4: Input the data-enhanced feature map into the average pooling layer and the fully connected layer for processing. The average pooling layer converts the feature map into a low-dimensional vector and sends it to the fully connected layer for classification.

[0012] Step 5: Send the classified feature map to the output layer for judgment and output the predicted value to obtain the authenticity judgment result of the face image.

[0013] Preferably, in step 1, preprocessing includes face detection and image normalization, and the preprocessed initial image data set is divided into several batches for input, each batch containing multiple face images.

[0014] Preferably, in step 2, the input initial image is regarded as a two-dimensional feature map of three channels, and the initial image is convolved through the convolution layer to extract primary features of the image, including texture features, color features, local shape features, etc., to generate an image feature map with a shape of [16, 64, 56, 56].

[0015] Preferably, in step 3, the basic building block of the ResNet-18 neural network is a residual block, and the residual block includes a main path and a residual connection;

[0016] The main path is used to learn the feature representation of the input data and extract the features of the input data. The main path includes 1 7×7 convolution layer, 4 residual blocks (each residual block consists of 2 3×3 convolution layers, a batch normalization layer, an activation function and a maximum pooling layer), 1 global average pooling layer, and 1 fully connected layer;

[0017] The residual connection directly adds the input to the output of the main path through the skip connection. The output of the residual block is the residual sum between the features learned by the main path and the input data.

[0018] Furthermore, an image style mixing module is embedded in the first two residual blocks of the ResNet-18 neural network. The image style mixing module synthesizes a new domain by mixing sample statistics of different domains at the feature level.

[0019] Preferably, in step 3, the face liveness detection problem is described as a binary classification task, and the probability of the model output and the actual label are supervised and trained using cross entropy loss.

[0020] Preferably, in step 4, the average pooling layer slides the window on the entire feature map according to the size and step size of the pooling window, calculates the average value of the area covered by each window, and forms a new output feature map of smaller size with the average value; after dimensionality reduction by the average pooling layer, the feature map is converted into a low-dimensional vector and sent to the fully connected layer for authenticity classification of the face image.

[0021] The technical solution of the present invention also provides: an electronic device, comprising:

[0022] one or more processors;

[0023] a storage device having one or more programs stored thereon;

[0024] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the above-mentioned methods for live face detection based on image style mixing.

[0025] The technical solution of the present invention also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in any of the above-mentioned face liveness detection methods based on image style mixing are implemented.

[0026] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:

[0027] Compared with the existing methods, the face liveness detection method of the present invention has better generalization performance and better results in performance evaluation indicators. Due to the mixed styles of face images, the method of the present invention improves the robustness of the model in different scenes and lighting conditions through diversified data enhancement while maintaining high accuracy, and can effectively resist unknown attacks. At the same time, due to the characteristics of ResNet-18, the cost of computing resources is reduced, and efficient liveness detection is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a flowchart of the face liveness detection method based on image style mixing of the present invention. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the application is further elaborated in detail below in conjunction with the accompanying drawings. The described embodiments are only a part of the embodiments involved in the present invention. All non-innovative embodiments of other researchers in the field on this embodiment belong to the protection scope of the present invention. At the same time, for the step numbering in the embodiment of the present invention, it is only set for the convenience of explanation, and the order between the steps is not limited in any way. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0030] In one embodiment of the present invention, a method for detecting live faces based on image style mixing is provided. Figure 1 As shown in the figure, first, the input image is passed through the convolution layer to extract primary features; then the features are learned through the ResNet-18 neural network. ResNet-18 is a lightweight model, which is easy to deploy and has a very high readiness rate in mainstream image recognition and classification; then, in view of the limited number of existing data sets, in order to maximize the use of the existing small amount of data, the image style mixing module is integrated in the ResNet Block to perform feature-level data enhancement, make up for the lack of data, and improve the generalization ability of the model; the image style mixing module increases the data diversity in the model training process by mixing the style information of different samples at the feature level, which helps the model learn a more generalized feature representation; however, although the introduction of the image style mixing module can improve the generalization ability of the model, it also increases the computational overhead; in order to reduce the amount of calculation and maintain the important information of the features, the average pooling layer is then used to reduce the spatial dimension of the features, the fully connected layer classifies the authenticity of the image, and finally the output layer gives the judgment result.

[0031] The face liveness detection method of this embodiment is based on image style mixing, and the details are as follows:

[0032] Step 1: Input face image data and build an initial image dataset.

[0033] After preprocessing of face detection and image normalization, the initial image dataset is divided into several small batches for input, each batch containing multiple images.

[0034] Step 2: The convolutional layer initially extracts image features.

[0035] The convolution layer performs convolution operations on the two-dimensional feature maps of multiple channels to generate further feature maps with a certain number of channels.

[0036] In this embodiment, the initial input image can be regarded as a two-dimensional feature map of three channels. The core operation of the convolution layer is convolution, which uses a smaller two-dimensional convolution kernel to perform a convolution operation on the part of the feature map of the input image with the same size to obtain the feature value of this part.

[0037] The size of the convolution kernel is called the receptive field, which refers to the size of the features that a convolution layer can perceive. The weight parameters in the convolution kernel are obtained through learning. In the convolution layer, after the convolution kernel completes a convolution on the feature map, it moves a certain step length and continues the convolution of the next area until all areas are completed. If the remaining edge area of ​​the feature map is not enough for convolution, the edge area can be filled based on certain rules to finally obtain a complete feature map of a channel.

[0038] In this embodiment, the initial image is convolved through the convolution layer to extract primary features of the image, including texture features, color features, local shape features, etc., to generate an image feature map with a shape of [16, 64, 56, 56].

[0039] Step 3: Perform deep feature learning through the ResNet-18 neural network, where the first two ResNet Blocks (residual blocks) integrate image style mixing modules.

[0040] In this embodiment, deep learning is performed through the ResNet-18 neural network, and residual connections are introduced to solve the gradient vanishing and gradient exploding problems in the deep neural network.

[0041] The basic building block of the ResNet-18 neural network is the ResNet Block, which consists of two main parts: First, the main path. The main path is responsible for learning the feature representation of the input data. It contains a series of convolutional layers, batch normalization layers, and activation functions that work together to further extract the features of the input data. Second, the residual connection, which directly adds the input to the output of the main path through skip connection.

[0042] In this embodiment, the main path includes 1 7×7 convolution layer, 4 residual blocks (each residual block consists of 2 3×3 convolution layers, batch normalization layer, activation function and maximum pooling layer), 1 global average pooling layer, and 1 fully connected layer. The output of the residual block is the sum of the residuals (differences) between the features learned by the main path and the input data. This design helps the gradient to propagate back to the beginning of the network more easily, preventing the problem of gradient vanishing or gradient exploding in deep networks.

[0043] In particular, since the number of face data sets in this embodiment is limited, the data sets can be enriched by embedding image style mixing modules in the first two ResNet Blocks.

[0044] The image style mixing module is used for cross-domain generalization of neural networks. It synthesizes new domains by mixing sample statistics from different domains at the feature level to enhance the generalization ability of the model.

[0045] Specifically, in this embodiment, given a batch of samples x, the image style mixing module first generates a reference sample batch based on x

[0046] If there is a domain label, then sample one instance from each domain to form x = [x i ,x j ] and swap their order to obtain Among them, shuffle(x i ) operation is to disrupt x i If there is no domain label, the order of x is directly shuffled to obtain x i .

[0047] Then, the image style mixing module calculates the mean μ(x) and standard deviation σ(x) of x, and The mean and standard deviation Next, we sample the instance-wise weight λ from the β distribution and use it to calculate the mean and standard deviation of the mixture:

[0048]

[0049] Among them, Y mix and β mix represents the mean and standard deviation of the mixture.

[0050] Finally, these mixed statistics are used to regularize x to obtain new synthetic domain samples. The regularization formula is:

[0051]

[0052] In this embodiment, new synthetic domain samples are used to input into the main path based on the convolutional neural network. After passing through multi-layer feature extraction modules (including convolutional layers, batch normalization layers, and activation functions, etc.), deep representation features that can simultaneously capture style domain information and semantic features are learned, effectively representing the content and style characteristics of the input face image, and providing more robust feature support for subsequent liveness detection tasks.

[0053] Furthermore, the method of this embodiment describes the face liveness detection problem as a binary classification task, and thus adopts the binary classification problem loss function, namely, cross entropy loss, for supervised training, and the expression is as follows:

[0054]

[0055] Among them, y i represents the true label of sample i, p i It represents the probability that sample i is a true sample, and N is the total number of samples.

[0056] Through supervised training with cross entropy loss, the error between the model's predicted probability and the true label is minimized, guiding the model to learn the discriminative features of the input data while reducing the misjudgment rate, effectively optimizing the model's classification ability, enabling it to accurately distinguish between real samples and attack samples, and providing reliable model support for face liveness detection tasks.

[0057] Step 4: Average pooling layer and fully connected layer processing.

[0058] According to the size and step size of the pooling window, the window is slid across the entire feature map, the average value of the area covered by each window is calculated, and these average values ​​are used to form a new smaller output feature map. After the dimensionality reduction of the average pooling layer, the feature map is converted into a low-dimensional vector and sent to the fully connected layer for classification.

[0059] Step 5: The output layer gives a judgment.

[0060] The output layer outputs the predicted value and gives a judgment on the authenticity of the face image.

[0061] Furthermore, this embodiment performs specific data simulation and verification of the face liveness detection method:

[0062] 1. ResNet-18 and AENet were trained on the dataset CelebA_Spoof (which is an extension of CelebA and is a large-scale face anti-counterfeiting dataset that contains a wealth of prosthetic attack scenarios and a variety of real face samples and is widely used in face liveness detection research). S And the image style mixture ResNet model, with epoch set to 80 and batch_size set to 128.

[0063] The results of training on the CelebA_Spoof dataset are shown in Table 1:

[0064] Table 1. Results of training on the CelebA_Spoof dataset

[0065] AUC (%) HETR(%) APCER (%) BPCER(%) <![CDATA[AENet S ]]> 99.81 1.1 4.62 1.35 ResNet-18 99.96 0.94 0.94 0.94 Image Style Mixture ResNet 99.98 0.55 0.55 0.55

[0066] It can be seen that the AUC of the face liveness detection method in this embodiment is higher than that of ResNet-18 and AENet. S The model shows that the method of this embodiment has better discrimination ability. At the same time, HETR of this embodiment is better than ResNet-18 and AENet S The APCER of the method in this embodiment is also lower than that of ResNet-18 and AENet. SThis shows that the method in this embodiment has better resistance to attacks. At the same time, the BPCER of the method in this embodiment is also higher than that of ResNet-18 and AENet. S This indicates that the method of this embodiment has a strong ability to recognize real faces.

[0067] 2. In addition to testing on the CelebA_Spoof test data, we also tested on the CASIA-FASD dataset (a large-scale dataset for studying face liveness detection, which contains multiple attack types, including distorted photo attacks, cut photo attacks, and video playback attacks, and has high diversity and challenges). The results of the test on the CASIA-FASD dataset are shown in Table 2 below:

[0068] Table 2. Test results on the CASIA-FASD dataset

[0069] AUC (%) HETR(%) APCER (%) BPCER(%) ResNet-18 83.75% 24.88 24.87 24.88 Image Style Mixture ResNet 86.01% 23.96 23.97 23.95

[0070] It can be seen that the AUC of the face liveness detection method of this embodiment is higher than that of the ResNet-18 model, indicating that the method of this embodiment has better discrimination ability. At the same time, the HETR of the method of this embodiment is lower than that of ResNet-18, indicating that the error rate of the method of this embodiment is lower. The APCER of the method of this embodiment is also lower than that of ResNet-18, indicating that the method of this embodiment has better resistance to attacks. The BPCER of the method of this embodiment is also lower than that of ResNet-18, indicating that the method of this embodiment has a strong ability to recognize real faces.

[0071] In an embodiment of the present invention, an electronic device is also provided, including: one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by the one or more processors, the one or more processors implement the face liveness detection method based on image style mixing described in any of the above embodiments.

[0072] In an embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods for face liveness detection based on image style mixing in the above embodiments are implemented.

[0073] In summary, the present invention utilizes the concept of image style mixing to perform feature-level data enhancement on samples to make up for the problem of data scarcity. At the same time, by adopting the ResNet-18 lightweight model, the hardware requirements are reduced, thereby effectively improving the accuracy and efficiency of the liveness detection model.

[0074] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for face liveness detection based on image style mixing, characterized in that: The steps include: Step 1: Input face image data, construct the initial image data set and perform preprocessing; Step 2: Input the initial image into the convolution layer, extract the primary features of the image, and generate an image feature map; Step 3: Input the image feature map into the ResNet-18 neural network for deep feature learning. The first two residual blocks of the ResNet-18 neural network are integrated with an image style mixing module to perform feature-level data enhancement. Step 4: Input the data-enhanced feature map into the average pooling layer and the fully connected layer for processing. The average pooling layer converts the feature map into a low-dimensional vector and sends it to the fully connected layer for classification. Step 5: Send the classified feature map to the output layer for judgment and output the predicted value to obtain the authenticity judgment result of the face image.

2. The method for face liveness detection based on image style mixing according to claim 1, characterized in that: The preprocessing includes face detection and image normalization, and the preprocessed initial image data set is divided into several batches for input, each batch containing multiple face images.

3. The method for face liveness detection based on image style mixing according to claim 2, characterized in that: In step 2, the input initial image is regarded as a two-dimensional feature map of three channels, and the initial image is convolved through the convolution layer to extract primary features of the image, including texture features, color features, and local shape features; Generates an image feature map of shape [16, 64, 56, 56].

4. The method for face liveness detection based on image style mixing according to claim 1, characterized in that: In step 3, the basic building block of the ResNet-18 neural network is a residual block, and the residual block includes a main path and a residual connection; The main path is used to learn the feature representation of the input data and extract the features of the input data. The main path includes a 7×7 convolution layer, 4 residual blocks, a global average pooling layer, and a fully connected layer; each residual block consists of two 3×3 convolution layers, a batch normalization layer, an activation function, and a maximum pooling layer; The residual connection directly adds the input to the output of the main path through the skip connection, and the output of the residual block is the residual sum between the features learned by the main path and the input data.

5. The method for face liveness detection based on image style mixing according to claim 4, characterized in that: In step 3, an image style mixing module is embedded in the first two residual blocks of the ResNet-18 neural network. The image style mixing module synthesizes a new domain by mixing sample statistics of different domains at the feature level. The method is as follows: Step 3.1: Given a sample x, generate a reference sample batch based on x If x has a domain label, sample an instance x from each of the two domains according to the domain label of x. i 、x j , forming x=[x i ,x j ], exchange the order to obtain the reference sample batch Among them, shuffle(x i ) operation is to shuffle x i The order of the elements in ; If x has no domain label, directly shuffle the order of x to get x i ; Step 3.2, calculate the mean μ(x) and standard deviation σ(x) of x, and The mean and standard deviation Step 3.3, sample instance-wise weight λ from the β distribution and calculate the mean and standard deviation of the mixture: Among them, Y mix and β mix represents the mean and standard deviation after mixing; Step 3.4: Use the mixed statistic Y mix and β mix Regularize x to obtain new synthetic domain samples. The regularization formula is: The new synthetic domain samples are used to be input into the main path based on the convolutional neural network, and pass through the multi-layer feature extraction module to learn deep representation features that simultaneously capture style domain information and semantic features, so as to effectively represent the content and style characteristics of the input face image.

6. The method for face liveness detection based on image style mixing according to claim 5, characterized in that: In step 4, the face liveness detection problem is described as a binary classification task, and the cross entropy loss is used to supervise the probability and actual label output by the model. The expression is as follows: Among them, y i represents the true label of sample i, p i It represents the probability that sample i is a true sample, and N is the total number of samples.

7. The method for face liveness detection based on image style mixing according to claim 6, characterized in that: In step 4, the average pooling layer slides the window on the entire feature map according to the size and step size of the pooling window, calculates the average value of the area covered by each window, and forms a new output feature map of smaller size with the average value; after dimensionality reduction by the average pooling layer, the feature map is converted into a low-dimensional vector and sent to the fully connected layer for authenticity classification of the face image.

8. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method for face liveness detection based on image style mixing as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the steps in the method for face liveness detection based on image style mixing described in any one of claims 1 to 7 are implemented.