Multi-level fusion iris living body detection method and device, equipment and storage medium

By using a multi-level fusion method and attention module to enhance the CNN network, extracting and fusing iris and periophthalmic features in iris live detection, the problem of degradation of detection accuracy in the prior art under complex conditions is solved, and high accuracy and low complexity iris live detection is achieved.

CN119992669AActive Publication Date: 2025-05-13BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510102807.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The existing iris live detection methods have significantly reduced detection accuracy when dealing with complex lighting conditions, sensor changes, and cross-database scenarios, and it is difficult to extract comprehensive and accurate features to detect multiple attack types, with high computational complexity and processing time.

Method used

A multi-level fusion iris live detection method is adopted, and a pre-trained dual-stream CNN network is enhanced by an attention module driven by an iris mask, and periophthalmic features and iris features are extracted, and self-attention features are fusion, combining CLAHE and HOG edge features to form a fusion three-channel image to improve the accuracy and robustness of the detection.

Benefits of technology

It effectively improves the accuracy of iris live detection, can accurately detect multiple prosthetic types under different light and sensor environments, reduces the computational complexity and processing time, and is suitable for real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992669A_ABST
    Figure CN119992669A_ABST
Patent Text Reader

Abstract

The invention provides a multi-level fusion iris living body detection method and device, equipment and a storage medium, and belongs to the technical field of iris living body detection, and the method comprises the steps: obtaining an iris mask and a fusion three-channel image based on an iris sample image; and inputting the fused three-channel image into an iris living body detection model to obtain a periorbital feature vector, an iris feature vector and a fused feature vector. And obtaining a periorbital feature vector based on the first feature map. And obtaining a second feature map based on the attention module, the iris mask and the first feature map. And obtaining an iris feature vector based on the second feature map. And obtaining a fusion feature vector based on the periorbital feature vector and the iris feature vector. And obtaining an iris living body detection result based on the periorbital feature vector, the iris feature vector and the fusion feature vector. The invention further provides a multi-level fusion iris living body detection device. The multi-level fusion iris living body detection device comprises an image layer fusion module, a feature layer fusion module and a score layer fusion module. According to the invention, the accuracy of iris living body detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of iris liveness detection, and more specifically, to a multi-level fusion iris liveness detection method, device, equipment, and storage medium. Background Art

[0002] Iris recognition has become an important technology in the field of biometrics due to its high stability, uniqueness, and non-contact characteristics. However, during the iris image acquisition phase, attackers often deceive the iris recognition system by forging or covering the texture information of the real iris, thereby achieving the purpose of impersonating others or hiding personal identity. Common prosthetic attack methods include printed irises, cosmetic contact lenses, screen-displayed irises, artificial eyes, synthetic irises, etc. Iris liveness detection is designed to detect such prosthetic attacks, thereby ensuring the security of the iris recognition system.

[0003] Existing iris liveness detection methods usually use a closed set evaluation protocol, that is, the training set and the test set share similar acquisition environments and attack types. Although the detection method performs well in this scenario, its detection accuracy is significantly reduced and its generalization ability is insufficient when dealing with unknown acquisition conditions such as complex lighting conditions, sensor changes, and cross-database scenarios. Moreover, many models that only rely on raw periocular images or normalized iris images as input find it difficult to extract comprehensive and accurate features to detect as many attack types as possible. If both periocular images and normalized iris images are used as input to extract corresponding features, the computational complexity and processing time of the entire detection model will be significantly increased, which is not conducive to the real-time operation of the model.

[0004] In summary, the existing technology still has the problem of low accuracy in iris liveness detection. Summary of the invention

[0005] The purpose of the present disclosure is to provide a multi-level fusion iris liveness detection method, device, equipment, and storage medium to improve the accuracy of iris liveness detection.

[0006] A first aspect of the embodiments of the present disclosure provides a multi-level fusion iris liveness detection method, comprising: Obtain an iris mask and a fused three-channel image based on the iris sample image; Inputting the fused three-channel image into an iris liveness detection model to obtain an eye periphery feature vector, an iris feature vector and a fused feature vector; the iris liveness detection model includes an attention module; The iris liveness detection model is used to: obtain a first feature map based on the fused three-channel image; obtain an eye-peripheral feature vector based on the first feature map; obtain a second feature map based on the attention module, the iris mask and the first feature map; obtain an iris feature vector based on the second feature map; obtain a fused feature vector based on the eye-peripheral feature vector and the iris feature vector; the second feature map is the first feature map after feature enhancement; A target classification output is obtained based on the eye periphery feature vector, the iris feature vector and the fusion feature vector, and an iris liveness detection result is obtained based on the target classification output.

[0007] A second aspect of the embodiments of the present disclosure provides a multi-level fusion iris liveness detection device, comprising: An image layer fusion module, used for obtaining an iris mask and fusing a three-channel image based on an iris sample image; A feature layer fusion module, used for inputting the fused three-channel image into an iris liveness detection model to obtain an eye periphery feature vector, an iris feature vector and a fusion feature vector; the iris liveness detection model includes an attention module; The iris liveness detection model is used to: obtain a first feature map based on the fused three-channel image; obtain an eye-peripheral feature vector based on the first feature map; obtain a second feature map based on the attention module, the iris mask and the first feature map; obtain an iris feature vector based on the second feature map; obtain a fused feature vector based on the eye-peripheral feature vector and the iris feature vector; the second feature map is the first feature map after feature enhancement; A fractional layer fusion module is used to obtain a target classification output based on the periocular feature vector, the iris feature vector and the fusion feature vector, and obtain an iris liveness detection result based on the target classification output. A third aspect of the embodiments of the present disclosure provides an electronic device, including a memory, a processor and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the above-mentioned multi-level fusion iris liveness detection method when executing the computer program.

[0008] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned multi-level fusion iris liveness detection method are implemented.

[0009] The multi-level fusion iris liveness detection method, device, equipment, and storage medium provided by the disclosed embodiments have the beneficial effects of: addressing the two challenges faced by iris liveness detection: joint detection of global attacks (such as printed irises) and local attacks (such as cosmetic contact lenses); and performance degradation caused by cross-domain (such as different collection environments and devices). This embodiment provides an accurate and effective response method, which improves the accuracy of iris liveness detection.

[0010] On the one hand, this embodiment uses an iris mask-driven attention module to enhance the pre-trained two-stream CNN network, extracts periocular features for global attacks and iris features for local attacks, and performs self-attention feature fusion based on periocular features and iris features to enhance the expression capability of fused features. Periocular features, iris features, and fused features are classified as true or false, and then fused at the fractional level, which improves the accuracy of iris liveness detection.

[0011] On the other hand, in order to reduce the impact of cross-domain, the CLAHE enhancement map and HOG edge enhancement feature map are extracted from the input image respectively to highlight the texture and edge information of the iris, and three-channel fusion is performed with the original grayscale image to form a fused input image, which is helpful for subsequent iris liveness detection to extract more discriminative features.

[0012] On the other hand, this embodiment constructs a lightweight iris liveness detection model based on iris segmentation mask-driven attention-enhanced multi-level fusion, which mines multi-level complementary information from images, features, and classification scores for iris liveness detection, and uses shared parameters in the feature extraction part of the model to reduce the model size. Specifically, the present application adopts an iris segmentation mask-driven attention-enhanced pre-trained two-stream CNN network. Through a dual-branch architecture, the model can reduce environmental background interference while enhancing the sensitivity of prosthetic features in the iris area, thereby improving robustness under different lighting and sensor environments, and enhancing the ability to accurately detect a variety of prosthetic types. Therefore, compared with the method of using periocular images and normalized iris images as inputs to extract corresponding features, the model is lighter, with less computational complexity and parameters, which is more conducive to the application of the model in practice. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0014] Figure 1A schematic diagram of the process of a multi-level fusion iris liveness detection method provided in an embodiment of the present disclosure; Figure 2 A segmentation mask and a segmentation effect diagram provided by an embodiment of the present disclosure; Figure 3 An iris training sample image provided by an embodiment of the present disclosure; Figure 4 A CLAHE enhanced image, a HOG edge feature image, and a fused three-channel image provided in an embodiment of the present disclosure; Figure 5 A schematic diagram of the structure of an iris liveness detection model provided in an embodiment of the present disclosure; Figure 6 This is an image processing effect diagram of the attention module provided in one embodiment of the present disclosure; Figure 7 A schematic diagram of the workflow of a fractional layer fusion module provided in one embodiment of the present disclosure; Figure 8 An approximately binary mask image provided by an embodiment of the present disclosure; Fig. 9 A schematic diagram of a pixel-by-pixel supervision and self-distillation weighted joint loss training process for an iris liveness detection model provided in an embodiment of the present disclosure; Fig.10 A schematic diagram of the workflow of an attention module provided in one embodiment of the present disclosure; Fig.11 A schematic diagram of the workflow of a feature layer fusion module provided in one embodiment of the present disclosure; Fig.12 A structural block diagram of a multi-level fusion iris liveness detection device provided in an embodiment of the present disclosure; Fig.13 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present disclosure. However, it should be clear to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present disclosure with unnecessary details.

[0016] In order to make the purpose, technical solutions and advantages of the present disclosure more clear, specific embodiments will be described below in conjunction with the accompanying drawings.

[0017] Please refer to Figures 1 to 11 , Figure 1This is a flow chart of a multi-level fusion iris liveness detection method provided in an embodiment of the present disclosure. The method may include S101 to S103.

[0018] S101: Obtain an iris mask and a fused three-channel image based on an iris sample image.

[0019] Figure 2 A segmentation mask and a segmentation effect diagram provided by an embodiment of the present disclosure. Figure 2 In this embodiment, the iris sample image is an iris image to be detected. Obtaining an iris mask based on the iris sample image includes: The iris region is annotated on the iris sample image, and the iris is segmented on the annotated iris sample image to obtain an iris mask.

[0020] For example, an iris sample image to be detected can be obtained by using an iris scanner or other acquisition equipment for live / prosthesis discrimination. The iris sample image is input into an iris segmentation model for iris segmentation to obtain an iris mask.

[0021] Figure 3 An iris training sample image provided by an embodiment of the present disclosure. Figure 3 During the training process, multiple live and prosthetic iris images can be collected as training sample images, the iris area can be manually marked, and a non-cooperative iris segmentation dataset including live and prosthetic irises can be constructed. A highly generalized iris segmentation model can be trained using the non-cooperative iris segmentation dataset to perform iris segmentation on iris sample images. The iris segmentation model can be constructed using a convolutional neural network, or an IrisSegNet model based on deep learning can be used.

[0022] Figure 4 The CLAHE enhancement image, HOG edge feature image and fused three-channel image provided by an embodiment of the present disclosure. Figure 4 In this embodiment, obtaining a fused three-channel image based on the iris sample image includes: Based on the iris sample image, the CLAHE enhancement image and HOG edge feature map are obtained.

[0023] The iris sample image, CLAHE enhanced image and HOG edge feature map are fused to obtain a fused three-channel image.

[0024] In this embodiment, the CLAHE enhanced image is an image obtained by applying the CLAHE technology to the iris sample image. CLAHE is Contrast-Limited Adaptive Histogram Equalization. The CLAHE enhanced image obtained after CLAHE processing has enhanced local contrast, making the details of the iris clearer and facilitating subsequent analysis and processing. HOG is the Histogram of Oriented Gradients. The HOG edge feature map highlights the edge information of the iris sample image.

[0025] In digital image processing, images are usually composed of one or more channels, each of which represents a specific attribute of the image. The fused three-channel image is to combine the iris sample image, CLAHE enhanced image and HOG edge feature map as a channel to form a new image with three channels. This three-channel fused image can simultaneously contain the information of the original image, the enhanced contrast information and the edge feature information, providing richer features for subsequent classification and recognition.

[0026] Exemplarily, the CLAHE enhancement image and HOG edge enhancement feature map are extracted from the input image to highlight the texture and edge information of the iris. The CLAHE enhancement image, the HOG edge enhancement feature map and the original grayscale image are three-channel fused to form a fused input image. The fusion of the three-channel image helps the subsequent iris liveness detection model to extract more discriminative features.

[0027] S102: Input the fused three-channel image into an iris liveness detection model to obtain an eye periphery feature vector, an iris feature vector, and a fused feature vector; the iris liveness detection model includes an attention module; The iris liveness detection model is used to: obtain a first feature map based on a fused three-channel image; obtain an eye-peripheral feature vector based on the first feature map; obtain a second feature map based on an attention module, an iris mask and the first feature map; obtain an iris feature vector based on the second feature map; obtain a fused feature vector based on the eye-peripheral feature vector and the iris feature vector; the second feature map is the first feature map after feature enhancement.

[0028] Figure 5 This is a schematic diagram of the structure of an iris liveness detection model provided by an embodiment of the present disclosure. Figure 5 In this embodiment, the first feature map is obtained based on the fusion of the three-channel image, including: Based on the common convolution layer and the maximum pooling layer in the iris liveness detection model, features are extracted from the fused three-channel image to obtain the first feature map.

[0029] In this embodiment, the iris liveness detection model is an attention-enhanced pre-trained two-stream CNN network driven by an iris segmentation mask. The iris liveness detection model includes a common convolution layer and a maximum pooling layer.

[0030] Exemplarily, the input fused three-channel image is first passed through a 7×7 common convolution layer and a 3×3 maximum pooling layer to extract low-level features to obtain a first feature map.

[0031] Figure 6 This is an image processing effect diagram of the attention module provided by an embodiment of the present disclosure. Figure 5 and Figure 6 In this embodiment, the iris liveness detection model further includes ResNet Block1, ResNet Block2, a first attention module, ResNet Block3, ResNet Block4, a second attention module, and a convolutional layer. The first attention module is an attention module added after ResNet Block2 and before ResNet Block3, and the second attention module is an attention module added after ResNetBlock4.

[0032] In this embodiment, obtaining the periocular feature vector based on the first feature map includes: Based on ResNet Block1, ResNet Block2, ResNet Block3, ResNet Block4 and convolutional layers, feature extraction is performed on the first feature map to obtain the periorbital intermediate feature map.

[0033] The periocular intermediate feature map is flattened and normalized to obtain the periocular feature vector.

[0034] In this embodiment, obtaining a second feature map based on the attention module, the iris mask and the first feature map, and obtaining an iris feature vector based on the second feature map includes: Based on ResNet Block1, ResNet Block2, the first attention module, ResNet Block3, ResNetBlock4, the second attention module and the convolution layer, feature enhancement and feature extraction are performed on the second feature map to obtain the iris intermediate feature map.

[0035] The iris intermediate feature map is flattened and normalized to obtain the iris feature vector.

[0036] Exemplarily, the pre-trained two-stream CNN network selects the residual network ResNet18. ResNet18 has four ResNetBlock modules (ResNet Block1, ResNet Block2, ResNet Block3, ResNet Block4), and each ResNetBlock module contains 4 3×3 convolutional layers. The pre-trained two-stream CNN network includes two branches: the periocular branch and the iris branch. The periocular branch can extract the global features of the input image and process the input image in the same way as the traditional pre-trained CNN network, aiming to ensure that the model remains robust when the mask information is inaccurate. The iris branch is used to extract the local feature map of the input image, and adds an iris segmentation mask-driven attention module (attention layer) after the two intermediate layers ResNet Block 2 and ResNet Block 4, namely the first attention module and the second attention module. These two iris segmentation mask-driven attention modules enable the CNN network to pay more attention to the iris area of ​​the input image, thereby enhancing the local feature extraction capability of the CNN network. The feature extraction process of the above two branches based on ResNet Block shares parameters.

[0037] The periocular branch and the iris branch undergo feature refinement through a convolution layer and a RELU layer respectively, obtaining a periocular specific intermediate feature map (i.e., periocular intermediate feature map) and an iris specific intermediate feature map (i.e., iris intermediate feature map). The RELU (Rectified Linear Unit) layer is an important activation function layer in the neural network.

[0038] The model flattens the obtained periorbital intermediate feature map and iris intermediate feature map and performs L2 norm normalization to obtain the periorbital feature vector and iris feature vector.

[0039] S103: Obtaining a target classification output based on the eye periphery feature vector, the iris feature vector and the fusion feature vector, and obtaining an iris liveness detection result based on the target classification output.

[0040] Figure 7 This is a schematic diagram of the workflow of the fractional layer fusion module provided in one embodiment of the present disclosure. Figure 5 and Figure 7 ,In this embodiment, the iris liveness detection model also includes a classifier.

[0041] The target classification output includes periocular classification output, iris classification output and fusion classification output.

[0042] The target classification output is obtained based on the periocular feature vector, iris feature vector and fusion feature vector, including: The periocular feature vector is input into the classifier to obtain the periocular classification output corresponding to the periocular feature vector.

[0043] The iris feature vector is input into the classifier to obtain the iris classification output corresponding to the iris feature vector.

[0044] The fused feature vector is input into the classifier to obtain the fused classification output corresponding to the fused feature vector.

[0045] In this embodiment, the iris liveness detection result is obtained based on the target classification output, including: Substitute the periocular classification output, iris classification output and fusion classification output into the classification score weighted fusion calculation formula to obtain the iris classification score.

[0046] If the iris classification score is greater than or equal to the classification score threshold, the iris liveness detection result is true.

[0047] If the iris classification score is less than the classification score threshold, the iris liveness detection result is false.

[0048] The calculation formula for weighted fusion of classification scores is:

[0049] in, Indicates the output of eye-peripheral classification, represents the iris classification output, represents the fusion classification output, , and is the weight coefficient, which can be , , . Represents the classification output score.

[0050] In this embodiment, the classification score threshold is a preset critical value for determining whether an iris is real or fake.

[0051] From the above, we can conclude that there are two challenges facing iris liveness detection: joint detection of global attacks (such as printed irises) and local attacks (such as cosmetic contact lenses); and performance degradation caused by cross-domain (such as different collection environments and devices). This embodiment provides an accurate and effective response method to improve the accuracy of iris liveness detection.

[0052] On the one hand, this embodiment uses an iris mask-driven attention module to enhance the pre-trained two-stream CNN network, extracts periocular features for global attacks and iris features for local attacks, and performs self-attention feature fusion based on periocular features and iris features to enhance the expression capability of fused features. Periocular features, iris features, and fused features are classified as true or false, and then fused at the fractional level, which improves the accuracy of iris liveness detection.

[0053] On the other hand, in order to reduce the impact of cross-domain, the CLAHE enhancement map and HOG edge enhancement feature map are extracted from the input image respectively to highlight the texture and edge information of the iris, and three-channel fusion is performed with the original grayscale image to form a fused input image, which is helpful for subsequent iris liveness detection to extract more discriminative features.

[0054] On the other hand, this embodiment constructs a lightweight iris liveness detection model based on iris segmentation mask-driven attention-enhanced multi-level fusion, which mines multi-level complementary information from images, features, and classification scores for iris liveness detection, and uses shared parameters in the feature extraction part of the model to reduce the model size. Specifically, the present application adopts an iris segmentation mask-driven attention-enhanced pre-trained two-stream CNN network. Through a dual-branch architecture, the model can reduce environmental background interference while enhancing the sensitivity of prosthetic features in the iris area, thereby improving robustness under different lighting and sensor environments, and enhancing the ability to accurately detect a variety of prosthetic types. Therefore, compared with the method of using periocular images and normalized iris images as inputs to extract corresponding features, the model is lighter, with less computational complexity and parameters, which is more conducive to the application of the model in practice.

[0055] Figure 8 An approximately binary mask image provided by an embodiment of the present disclosure. Fig. 9 A schematic diagram of the pixel-by-pixel supervision and self-distillation weighted joint loss training process of the iris liveness detection model provided in an embodiment of the present disclosure is provided. Figure 8 and Fig. 9 In one embodiment of the present disclosure, the multi-level fusion iris liveness detection method further includes: Input the iris sample training set into the initial iris liveness detection model to obtain the periorbital intermediate feature map, the iris intermediate feature map, the periorbital feature vector, the iris feature vector, the fusion feature vector and the target classification output; The self-distillation loss is obtained based on the periocular feature vector, the iris feature vector, the fusion feature vector, the target classification output and the self-distillation loss function; The pixel-by-pixel supervision loss is obtained based on the periorbital intermediate feature map, the iris intermediate feature map and the pixel-by-pixel supervision loss function; The weighted joint loss is obtained based on the self-distillation loss, pixel-by-pixel supervision loss and the weighted joint loss function; The initial iris liveness detection model is updated based on the weighted joint loss to obtain the iris liveness detection model.

[0056] In this embodiment, the self-distillation loss function is:

[0057] in, represents the self-distillation loss, label represents the true label of the iris sample, label=1 represents the real iris, label=0 represents the prosthetic iris, Represents the target classification output, which includes the periorbital classification output, iris classification output, and fusion classification output. When i=1, Represents the periocular classification output; when i=2, Represents the iris classification output; when i=3, represents the fusion classification output; () represents the loss calculation function between the target classification output and the true label; the fusion feature is used as the teacher model, and the periorbital feature and iris feature are used as the student model. Represents the classification output of the teacher model in the self-distillation process, that is, the fusion classification output; () represents the KL divergence loss function, () represents the fully connected layer. When i=1, represents the periocular feature vector; when i=2, represents the iris feature vector; when i=3, represents the fused feature vector; represents the eigenvector of the teacher model, i.e., the fusion eigenvector, Denotes that the self-distillation loss dynamically adjusts the weights.

[0058] In this embodiment, Indicates calculation first The sum of the squares of each element of this vector, and then taking the square root, we get and The distance between these two vectors in terms of L2 norm. This distance measures the student model feature vector after dimension alignment. and the teacher model feature vector The degree of difference between them is used to optimize the student model by minimizing this distance during the feature distillation process, so that the student model can learn the characteristics of the teacher model and improve the model performance.

[0059] The pixel-by-pixel supervision loss function is:

[0060]

[0061] in, represents the pixel-by-pixel supervision loss. When j=1, represents the periocular intermediate feature map, which is used to represent the visual information around the eye. When j=2, Represents the iris intermediate feature map, which is used to represent the visual information of the iris area. represents the approximate binary mask image generated based on the iris mask, R represents the pixel coordinate set of the iris mask area, () indicates based on The loss function is Denotes pixel-wise supervision loss dynamically adjusts weights.

[0062] In this embodiment, the pixel-by-pixel supervised loss function is used to measure the difference between the periocular intermediate feature map and the iris intermediate feature map and the expected result at the pixel level. By generating an approximate binary mask map based on the iris mask, the difference between the pixels in the iris mask area (R) is focused on. For the periocular intermediate feature map and the iris intermediate feature map, the SmoothL1 loss function is used to calculate the difference between them. The pixel-by-pixel supervised loss function calculates the loss of each pixel to ensure that the model can capture and learn the features of the target area (such as the iris area) more accurately, making the model more robust in processing details and local features, dynamically adjusting the weights as the training progresses, and gradually optimizing the intermediate feature maps.

[0063] The weighted joint loss function is:

[0064] in, represents the weighted joint loss, , is the balance coefficient, .

[0065] For example, It is a self-distillation loss with fusion features as the teacher model and periorbital features and iris features as the student model. The self-distillation loss is calculated by the Focal loss function, KL divergence loss function and L2 loss. The Focal loss function is used to calculate the true label (label) of the iris sample and the three classification outputs respectively, and the knowledge hidden in the training data set can be directly introduced from the label to all classifiers.

[0066] By introducing the KL divergence loss, the periorbital and iris classification outputs (student model) can learn knowledge from the fusion feature classification output (teacher model), thereby promoting the student model to gradually approach the output distribution of the teacher model, allowing the model to learn more comprehensive features.

[0067] Yes and Fully connected layers for dimension alignment. Through feature distillation, the fused features of the teacher model are passed to the student model and optimized through L2 loss. The detailed and high-level information is implicitly passed through the feature map, so that the student model can learn the knowledge of the teacher model more accurately.

[0068] Indicates that the self-distillation loss dynamically adjusts the weight, The initial value of can be 0.03, and as the training progresses, it increases by epoch*0.005.

[0069] In this embodiment, the self-distillation loss function is a loss function for knowledge transfer and optimization, with fusion features as the teacher model and periorbital features and iris features as the student model. The self-distillation loss function uses the Focal loss function to introduce the knowledge of the true label of the sample into the classifier, uses the KL divergence loss to let the student model learn the output distribution from the teacher model, combines the fully connected layer to align the feature vector dimensions, optimizes feature distillation through L2 loss, and can also dynamically adjust weights to encourage the model to learn more comprehensive features and improve overall performance.

[0070] For example, is the pixel-by-pixel supervision loss, and the training error is calculated by the Smooth L1 loss function. and Negative feedback is given to the initial iris liveness detection model for the next step of training to optimize and update the iris liveness detection model. Use Smooth L1 loss to optimize the intermediate feature map and The difference between them can ensure that the model can capture and learn the characteristics of the target area more robustly. Indicates that pixel-by-pixel supervision loss dynamically adjusts weights, The initial value of can be 1, and as the training progresses, it can be reduced by epoch*0.001.

[0071] In this embodiment, the acquisition of the iris sample training set may include: constructing a training library after data cleaning based on an existing public iris liveness detection data set, inputting the images in the training library into the initial iris liveness detection model for training in batches, and saving the optimal training model in the training project for liveness detection of samples to be tested.

[0072] After obtaining the iris sample training set, the data set can also be expanded, including: obtaining a single iris training sample image from each batch of training sample images, performing iris segmentation masking and three-channel fusion image joint data expansion operations on the iris training sample image, such as randomly cropping and flipping the image to expand the data set.

[0073] The self-distillation loss in this embodiment combines multiple loss functions to introduce training data knowledge into the classifier, allowing the student model to learn knowledge from the teacher model and learn comprehensive features. The pixel-by-pixel supervision loss uses Smooth L1 loss to ensure that the model robustly captures and learns target area features. The weighted joint loss balances the two and continuously optimizes during training by dynamically adjusting the weights. This embodiment allows the initial model to more accurately determine whether the iris is real after being updated, improves detection accuracy and robustness, effectively copes with complex iris samples, and improves the overall performance of the iris liveness detection model.

[0074] Fig.10 A schematic diagram of the workflow of the attention module provided in one embodiment of the present disclosure. Fig.10 In one embodiment of the present disclosure, obtaining a second feature map based on an attention module, an iris mask and a first feature map includes: Performing feature enhancement on the first feature map based on the convolutional layer in the attention module to obtain an enhanced feature map; Scaling the iris mask based on the nearest neighbor interpolation method to obtain a target iris mask with the same size as the first feature map; Obtain an attention feature map of the iris area based on the enhanced feature map and the target iris mask; A second feature map is obtained based on the attention feature map and the first feature map.

[0075] In this embodiment, the attention feature map of the iris region is obtained based on the enhanced feature map and the target iris mask, including: The enhanced feature map and the target iris mask are aligned and multiplied to obtain the attention feature map of the iris area.

[0076] In this embodiment, the attention module includes a first attention module and a second attention module. The first attention module and the second attention module have the same structure and workflow, but they are in different processing stages, so the input feature maps are also different. The first feature map can represent the feature map output by the middle layer ResNet Block 2 or the feature map output by ResNetBlock 4.

[0077] Exemplarily, for the first attention module: the input feature map (i.e., the first feature map) is the feature map output by the middle layer ResNetBlock 2. The model takes the first feature map as input, combines the iris mask corresponding to the first feature map, and obtains an enhanced feature map after processing by the convolution layer in the first attention module. The nearest neighbor interpolation method is used to scale the iris mask to the same size as the first feature map (i.e., the height H and width W are the same). The enhanced enhanced feature map and the scaled target iris mask are aligned and multiplied to obtain the attention feature map of the iris area. The attention feature map is added to the first feature map to obtain the second feature map after iris enhancement.

[0078] For the second attention module: the input feature map (i.e., the first feature map) is the feature map output by the intermediate layer ResNet Block 4.

[0079] This embodiment can effectively focus on the iris area, enhance feature representation, and increase the attention of the iris liveness detection model to iris features, which helps to improve the accuracy of iris liveness detection.

[0080] Fig.11 A schematic diagram of the workflow of a feature layer fusion module provided in an embodiment of the present disclosure. Fig.11 In one embodiment of the present disclosure, obtaining a fused feature vector based on the periocular feature vector and the iris feature vector includes: Perform vector concatenation on the periocular feature vector and the iris feature vector, and determine the attention weight corresponding to the concatenated feature vector.

[0081] The concatenated feature vectors are fused based on the attention weights to obtain an intermediate fusion vector.

[0082] The intermediate fusion vector is input into the fully connected layer to obtain the fused feature vector.

[0083] In this embodiment, feature fusion is performed on the concatenated feature vectors based on the attention weights to obtain an intermediate fusion vector, including: The concatenated feature vectors are fused based on the attention weights to obtain a fused feature vector.

[0084] The fused feature vector is average pooled and layer normalized to obtain the intermediate fused vector.

[0085] In this embodiment, the periocular feature vector is a feature representation extracted from the area around the eye. The periocular feature vector may include various feature information around the eye, such as the texture, wrinkles, shape, color and other features of the periocular skin. The iris feature vector is a feature representation extracted from the iris area. The iris feature vector may include iris-specific features, such as the texture, color, pattern, blood vessel distribution and other information of the iris. The intermediate fusion vector is an intermediate result obtained in the feature fusion process.

[0086] Attention weights are used to emphasize or focus on different parts of the concatenated feature vector. Different attention weights can be assigned according to the importance of different tasks or features. For example, in an iris recognition task, the part corresponding to the iris feature vector can be assigned a higher attention weight than the periocular feature vector. Attention weights can be learned or set based on prior knowledge.

[0087] Exemplarily, dynamic feature fusion is performed on the periocular feature vector and the iris feature vector to improve the model's ability to express different features. The periocular feature vector and the iris feature vector are concatenated, and a multi-head self-attention mechanism is used to dynamically assign different attention weights to each concatenated feature, thereby achieving more effective feature fusion. The fused features are subjected to sequence dimension average pooling and layer normalization operations to further enhance the expression ability of the fused features. The processed fused features are integrated through a fully connected layer to obtain a fused feature vector containing periocular and iris information.

[0088] This embodiment uses a multi-head self-attention mechanism to enable the iris liveness detection model to flexibly focus on important features according to task requirements, thereby improving the feature expression capability. The average pooling and layer normalization operations in the sequence dimension further enhance the fusion feature expression. The integration of the fully connected layer and the final output fusion feature vector combine the periorbital and iris information, providing a more comprehensive and effective feature representation for subsequent tasks such as iris recognition, which helps to improve the accuracy of iris liveness detection.

[0089] Corresponding to the multi-level fusion iris liveness detection method in the above embodiment, Fig.12 This is a structural block diagram of a multi-level fusion iris liveness detection device provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Fig.12 The multi-level fusion iris liveness detection device 20 includes: an image layer fusion module 21, a feature layer fusion module 22 and a score layer fusion module 23.

[0090] The image layer fusion module 21 is used to obtain an iris mask and fuse the three-channel image based on the iris sample image.

[0091] The feature layer fusion module 22 is used to input the fused three-channel image into the iris liveness detection model to obtain the eye periphery feature vector, the iris feature vector and the fused feature vector. The iris liveness detection model includes an attention module.

[0092] The iris liveness detection model is used to: obtain a first feature map based on the fused three-channel image. Obtain an eye periocular feature vector based on the first feature map. Obtain a second feature map based on the attention module, the iris mask and the first feature map. Obtain an iris feature vector based on the second feature map. Obtain a fused feature vector based on the eye periocular feature vector and the iris feature vector. The second feature map is the first feature map after feature enhancement.

[0093] The score layer fusion module 23 is used to obtain a target classification output based on the periocular feature vector, the iris feature vector and the fusion feature vector, and obtain an iris liveness detection result based on the target classification output.

[0094] In one embodiment of the present disclosure, the multi-level fusion iris liveness detection device 20 also includes: a loss training module, which is used to input the iris sample training set into the initial iris liveness detection model to obtain the periorbital intermediate feature map, the iris intermediate feature map, the periorbital feature vector, the iris feature vector, the fusion feature vector and the target classification output.

[0095] The self-distillation loss is obtained based on the periocular feature vector, iris feature vector, fusion feature vector, target classification output and self-distillation loss function.

[0096] The pixel-by-pixel supervision loss is obtained based on the periorbital intermediate feature map, the iris intermediate feature map and the pixel-by-pixel supervision loss function.

[0097] The weighted joint loss is obtained based on the self-distillation loss, pixel-by-pixel supervision loss and the weighted joint loss function.

[0098] The initial iris liveness detection model is updated based on the weighted joint loss to obtain the iris liveness detection model.

[0099] In one embodiment of the present disclosure, the loss training module is specifically used for the self-distillation loss function:

[0100] in, represents the self-distillation loss, label represents the true label of the iris sample, label=1 represents the real iris, label=0 represents the prosthetic iris, Represents the target classification output, which includes the periorbital classification output, iris classification output, and fusion classification output. When i=1, Represents the periocular classification output; when i=2, Represents the iris classification output; when i=3, represents the fusion classification output; () represents the loss calculation function between the target classification output and the true label; the fusion feature is used as the teacher model, and the periorbital feature and iris feature are used as the student model. Represents the classification output of the teacher model in the self-distillation process, that is, the fusion classification output; () represents the KL divergence loss function, () represents the fully connected layer. When i=1, represents the periocular feature vector; when i=2, represents the iris feature vector; when i=3, represents the fused feature vector; represents the eigenvector of the teacher model, i.e., the fusion eigenvector, Denotes that the self-distillation loss dynamically adjusts the weights.

[0101] The pixel-by-pixel supervision loss function is:

[0102]

[0103] in, represents the pixel-by-pixel supervision loss. When j=1, represents the periocular intermediate feature map, which is used to represent the visual information around the eye; when j=2, Represents the iris intermediate feature map, which is used to represent the visual information of the iris area; represents the approximate binary mask image generated based on the iris mask, R represents the pixel coordinate set of the iris mask area, () indicates based on The loss function is Denotes pixel-wise supervision loss dynamically adjusts weights.

[0104] The weighted joint loss function is:

[0105] in, represents the weighted joint loss, , is the balance coefficient, .

[0106] In one embodiment of the present disclosure, the feature layer fusion module 22 is specifically used to perform feature enhancement on the first feature map based on the convolutional layer in the attention module to obtain an enhanced feature map.

[0107] The iris mask is scaled based on the nearest neighbor interpolation method to obtain a target iris mask with the same size as the first feature map.

[0108] The attention feature map of the iris region is obtained based on the enhanced feature map and the target iris mask.

[0109] A second feature map is obtained based on the attention feature map and the first feature map.

[0110] In an embodiment of the present disclosure, the image layer fusion module 21 is specifically configured to obtain a CLAHE enhancement image and a HOG edge feature image based on the iris sample image.

[0111] The iris sample image, CLAHE enhanced image and HOG edge feature map are fused to obtain a fused three-channel image.

[0112] In one embodiment of the present disclosure, the feature layer fusion module 22 is further configured to extract features from the fused three-channel image based on the common convolution layer and the maximum pooling layer in the iris liveness detection model to obtain a first feature map.

[0113] In one embodiment of the present disclosure, the feature layer fusion module 22 is specifically used to perform vector splicing on the periocular feature vector and the iris feature vector, and determine the attention weight corresponding to the spliced ​​feature vector.

[0114] The concatenated feature vectors are fused based on the attention weights to obtain an intermediate fusion vector.

[0115] The intermediate fusion vector is input into the fully connected layer to obtain the fused feature vector.

[0116] See also Fig.13 , Fig.13 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Fig.13 The electronic device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303 and one or more memories 304. The processors 301, input devices 302, output devices 303 and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of the modules in the above-mentioned device embodiments, such as Fig.12 The functions of modules 21 to 23 are shown.

[0117] It should be understood that in the embodiment of the present disclosure, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0118] The input device 302 may include a touch panel, a fingerprint collection sensor (for collecting the user's fingerprint information and fingerprint direction information), a microphone, etc., and the output device 303 may include a display (LCD, etc.), a speaker, etc.

[0119] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0120] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present disclosure can execute the implementation methods described in the first and second embodiments of the multi-level fusion iris liveness detection method provided in the embodiments of the present disclosure, and can also execute the implementation methods of the electronic device 300 described in the embodiments of the present disclosure, which will not be repeated here.

[0121] In another embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by the processor, all or part of the processes in the above-mentioned embodiment method are implemented, and the computer program can also be completed by instructing the relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0122] The computer-readable storage medium may be an internal storage unit of the electronic device of any of the aforementioned embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0123] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.

[0124] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic devices and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0125] In the several embodiments provided in the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or it can be an electrical, mechanical or other form of connection.

[0126] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present disclosure.

[0127] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0128] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present disclosure, and these modifications or replacements should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.

Claims

1. A multi-level fusion iris liveness detection method, characterized in that: include: Obtain an iris mask and a fused three-channel image based on the iris sample image; Inputting the fused three-channel image into an iris liveness detection model to obtain an eye periphery feature vector, an iris feature vector and a fused feature vector; The iris liveness detection model includes an attention module; The iris liveness detection model is used to: obtain a first feature map based on the fused three-channel image; obtain an eye-peripheral feature vector based on the first feature map; obtain a second feature map based on the attention module, the iris mask and the first feature map; obtain an iris feature vector based on the second feature map; obtain a fused feature vector based on the eye-peripheral feature vector and the iris feature vector; the second feature map is the first feature map after feature enhancement; A target classification output is obtained based on the eye periphery feature vector, the iris feature vector and the fusion feature vector, and an iris liveness detection result is obtained based on the target classification output.

2. The multi-level fusion iris liveness detection method according to claim 1, characterized in that: Also includes: Input the iris sample training set into the initial iris liveness detection model to obtain the periorbital intermediate feature map, the iris intermediate feature map, the periorbital feature vector, the iris feature vector, the fusion feature vector and the target classification output; The self-distillation loss is obtained based on the periocular feature vector, the iris feature vector, the fusion feature vector, the target classification output and the self-distillation loss function; The pixel-by-pixel supervision loss is obtained based on the periorbital intermediate feature map, the iris intermediate feature map and the pixel-by-pixel supervision loss function; Obtaining a weighted joint loss based on the self-distillation loss, the pixel-by-pixel supervision loss and a weighted joint loss function; The initial iris liveness detection model is updated based on the weighted joint loss to obtain an iris liveness detection model.

3. The multi-level fusion iris liveness detection method as claimed in claim 2, characterized in that: The self-distillation loss function is: in, represents the self-distillation loss, label represents the true label of the iris sample, label=1 represents the real iris, label=0 represents the prosthetic iris, Represents the target classification output, which includes the periorbital classification output, iris classification output, and fusion classification output. When i=1, represents the periocular classification output; when i=2, Represents the iris classification output; when i=3, represents the fusion classification output; () represents the loss calculation function between the target classification output and the true label; the fusion feature is used as the teacher model, and the periorbital feature and iris feature are used as the student model. Represents the classification output of the teacher model in the self-distillation process, that is, the fusion classification output; () represents the KL divergence loss function, () represents the fully connected layer. When i=1, represents the periocular feature vector; when i=2, represents the iris feature vector; when i=3, represents the fused feature vector; represents the eigenvector of the teacher model, i.e., the fusion eigenvector, Indicates that the self-distillation loss dynamically adjusts the weight; The pixel-by-pixel supervision loss function is: in, represents the pixel-by-pixel supervision loss. When j=1, represents the periocular intermediate feature map, which is used to represent the visual information around the eye; when j=2, Represents the iris intermediate feature map, which is used to represent the visual information of the iris area; represents the approximate binary mask image generated based on the iris mask, R represents the pixel coordinate set of the iris mask area, () indicates based on The loss function is Indicates that pixel-by-pixel supervision loss dynamically adjusts weights; The weighted joint loss function is: in, represents the weighted joint loss, , is the balance coefficient, .

4. The multi-level fusion iris liveness detection method according to claim 1, characterized in that: The obtaining a second feature map based on the attention module, the iris mask and the first feature map comprises: Performing feature enhancement on the first feature map based on the convolutional layer in the attention module to obtain an enhanced feature map; Scaling the iris mask based on a nearest neighbor interpolation method to obtain a target iris mask having the same size as the first feature map; Obtaining an attention feature map of the iris region based on the enhanced feature map and the target iris mask; A second feature map is obtained based on the attention feature map and the first feature map.

5. The multi-level fusion iris liveness detection method according to claim 1, characterized in that: The step of obtaining a fused three-channel image based on the iris sample image comprises: Obtaining a CLAHE enhancement image and a HOG edge feature image based on the iris sample image; The iris sample image, the CLAHE enhancement image and the HOG edge feature image are fused to obtain a fused three-channel image.

6. The multi-level fusion iris liveness detection method according to claim 1, characterized in that: The obtaining a first feature map based on the fused three-channel image includes: Based on the common convolution layer and the maximum pooling layer in the iris liveness detection model, feature extraction is performed on the fused three-channel image to obtain a first feature map.

7. The multi-level fusion iris liveness detection method according to claim 1, characterized in that: The step of obtaining a fused feature vector based on the periorbital feature vector and the iris feature vector comprises: Performing vector concatenation on the periocular feature vector and the iris feature vector, and determining an attention weight corresponding to the concatenated feature vector; Perform feature fusion on the concatenated feature vectors based on the attention weights to obtain an intermediate fusion vector; The intermediate fusion vector is input into the fully connected layer to obtain a fusion feature vector.

8. A multi-level fusion iris liveness detection device, characterized in that: include: An image layer fusion module, used for obtaining an iris mask and fusing a three-channel image based on an iris sample image; A feature layer fusion module, used for inputting the fused three-channel image into an iris liveness detection model to obtain an eye periphery feature vector, an iris feature vector and a fusion feature vector; the iris liveness detection model includes an attention module; The iris liveness detection model is used to: obtain a first feature map based on the fused three-channel image; obtain an eye-peripheral feature vector based on the first feature map; obtain a second feature map based on the attention module, the iris mask and the first feature map; obtain an iris feature vector based on the second feature map; obtain a fused feature vector based on the eye-peripheral feature vector and the iris feature vector; the second feature map is the first feature map after feature enhancement; The fractional layer fusion module is used to obtain a target classification output based on the periocular feature vector, the iris feature vector and the fusion feature vector, and obtain an iris liveness detection result based on the target classification output.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Image-based living body target detection method, device and system

    CN109871811A

  • Iris image segmentation, positioning and normalization method based on multi-task neural network

    CN112287872A

  • Iris recognition model training method, iris recognition method and device

    CN115083006A

  • Living body detection method and system

    CN116311551A

  • Iris living body detection method and device based on multi-source information fusion

    CN117351579A