A real-time iris localization and segmentation system and method

By proposing a multi-task real-time processing network RISNet in the iris recognition system, the problem of poor iris segmentation and positioning effect in visible light environment is solved, and the iris positioning and segmentation effect with high accuracy, robustness and real-time performance is achieved.

CN115798027BActive Publication Date: 2025-06-20CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211685270.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-06-20
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

Due to the influence of noise in visible light environments, the existing iris recognition system has poor iris segmentation and positioning effects, insufficient robustness, and low real-time performance of the model.

Method used

A multi-task real-time processing network RISNet is proposed. By processing iris positioning and segmentation tasks in parallel, it adopts dual-center point confidence prediction and dual-branch regression to perform internal and external circular positioning, and a lightweight feature extraction network is designed to improve real-time performance.

Benefits of technology

In infrared and visible light environments, the accuracy and robustness of iris positioning and segmentation are significantly improved, and at least 3 times the speed improvement is achieved to meet real-time application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115798027B_ABST
    Figure CN115798027B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of biometric recognition technology, and particularly to a real-time iris localization and segmentation system and method. The system includes: a backbone network for extracting image features from an original image to obtain an image feature map; a semantic segmentation subnet for performing iris segmentation on the image feature map; a center point localization subnet for locating the inner and outer circle positions of the iris; and a size regression subnet for predicting the radii and offsets of the inner and outer circles of the iris. The RISNet of the present invention performs inner and outer circle localization in a manner of double center point confidence prediction and double-branch regression, shares the same feature map with the segmentation network, outputs the localization result and the iris mask simultaneously, and does not require post-processing; optimizes the loss function and proposes PairLoss; the experimental results show that the method of the present invention has strong robustness, high accuracy, improved localization accuracy and segmentation accuracy, and realizes speed improvement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biometric identification, and particularly to a real-time iris localization and segmentation system and method. Background Art

[0002] The iris is an annular region between the pupil and the sclera. Among numerous biological forms (such as fingerprints, faces, palm prints, gaits), iris recognition is considered the most reliable biometric identification technology due to the uniqueness and stability of iris textures. Currently, iris recognition is widely applied in fields such as intelligent unlocking, border control, and information forensics.

[0003] A complete iris recognition system generally consists of four sub-processes: image acquisition, iris segmentation and localization, feature extraction, and matching. Iris segmentation aims to generate an iris mask to distinguish iris pixels from non-iris pixels, and iris localization aims to detect the inner and outer boundaries of the iris region. Iris normalization aims to flatten the annular iris regions of different sizes along the inner circle into a fixed size (64×512). Iris segmentation and iris localization, as the preprocessing of iris recognition, jointly define the pixel regions for feature extraction and matching, and directly affect the overall recognition performance of the iris. As Figure 1 shows a complete iris recognition preprocessing (iris segmentation and iris localization) process.

[0004] Previously, the vast majority of research focused on infrared images taken under cooperative conditions (such as fixed acquisition distance, gaze, user cooperation, near-infrared illumination). Due to the fixed shooting distance, fixed light source angle, and cooperative shooting, the captured iris images have the characteristics of low environmental noise and clear contours. Currently, with the improvement of the shooting pixels of smartphones, visible light iris images taken based on mobile devices have become a research hotspot. Since the imaging conditions are unrestricted and no user cooperation is required for shooting, the iris images obtained in this environment are prone to being affected by various noises, such as incomplete iris exposure caused by user attention deviation, image blurring caused by device jitter, specular reflection, eyelid occlusion, and other problems. This makes iris segmentation and localization more challenging. Figure 2 shows the iris images obtained under different conditions and the degradation types.

[0005] Iris segmentation, as the most fundamental step in iris recognition systems, has always attracted the attention of researchers. The basic segmentation methods have evolved from the initial traditional algorithms based on manually designed features such as thresholds, regions, and edges to semantic segmentation networks based on deep learning; the research content generally improves the network performance in three directions according to the difficulties of image segmentation: increasing the segmentation fineness, enhancing the network's generalization ability for multi-scale, and learning the context spatial correlation. Liu et al. proposed the MFCN (Nianfeng, Haiqing Li, Man Zhang, Jing Liu, Zhenan Sun, and Tieniu Tan. "Accurate iris segmentation in non-cooperative environments using fully convolutional networks." In *2016 International Conference on Biometrics (ICB)*, pp. 1-8. IEEE, 2016.) method, which uses a fully convolutional network to segment the iris, applying the semantic segmentation method based on deep learning to the iris segmentation task and demonstrating the superiority of the deep learning-based segmentation method over traditional iris segmentation methods. Wang et al. proposed a baseline for iris segmentation (Wang Caiyong, Sun Zhemnan. Evaluation benchmark for iris segmentation algorithms [J]. Journal of Computer Research and Development, 2020, 57(2): 395-412), applying general semantic segmentation methods to the iris segmentation task. Methods such as FCEDN (Ehsaneddin Jalilian and Andreas Uhl, “Iris segmentation using fully convolutional encoder–decoder networks,” in Deep Learning for Biometrics, pp. 133–155. Springer, 2017.) and FCDNN (Bazrafkan, Shabab, Shejin Thavalengal, and Peter Corcoran. "An end to end deep neural network for iris segmentation in unconstrained scenarios." Neural Networks 106(2018): 79-95) use a fully convolutional encoder-decoder structure to train images of different scales to obtain multi-scale features;

[0006] Muhammad Arsalan et al. proposed the IrisDenseNet method (Muhammad Arsalan, Rizwan Ali Naqvi and Kang Ryoung Park, “Irisdensenet: Robust iris segmentation using densely connected fully convolutional networks in the images by visible light and near-infrared light camera sensors,” Sensors, vol. 18, no. 5, pp. 1501, 2018), which uses dense connections to obtain deeper fusion features; Arsalan et al. proposed the FRED-Net method (Arsalan, Muhammad, Dong Seop Kim, Min Beom Lee, Muhammad Owais, and Kang Ryoung Park. "FRED-Net: Fully residual encoder–decoder network for accurate iris segmentation." Expert Systems with Applications 122 (2019): 217-241), which uses residual blocks to obtain deeper semantic features; Wang et al. proposed the IrisParseNet method (Caiyong Wang, Jawad Muhammad, Yunlong Wang, Zhaofeng He, and Zhenan Sun, “Towards complete and accurate iris segmentation using deep multi-task attention network for non-cooperative iris recognition,” IEEE Transactions on Information Forensics and Security, vol. 15, pp. 2944–2959, 2020), which utilizes the characteristics of dilated convolution with a larger receptive field to obtain more spatial information and detailed information, and predicts more accurate iris segmentation results, etc. Although these deep learning-based methods have improved the segmentation accuracy, most of them focus on segmentation accuracy while ignoring the network size and computational speed, resulting in low real-time performance of the model.Fang et al. used the lightweight segmentation method CCNet (Fang, Zhaoyuan, and Adam Czajka. "Open Source Iris Recognition Hardware and Software with Presentation Attack Detection." 2020 IEEE International Joint Conference on Biometrics (IJCB)) to process the iris segmentation task. The model takes grayscale images as the network input. Although the processing speed of the model is improved, on visible light iris images, since some information is lost when the compressed image is converted to grayscale, the segmentation accuracy on visible light iris images is generally average.

[0007] In iris localization, most previous studies have focused on infrared images captured under cooperative conditions (fixed distance between the face and the imaging device). Two widely used localization methods are to use Dangman's integral-differential operator or circular Hough-transform to fit the inner and outer circle parameters on the mask image obtained by iris segmentation. The above methods have achieved certain results on infrared iris images. However, in the visible light environment, since the segmented mask is irregular, the localization effect is easily affected by the boundary of the iris mask, and the robustness of the method is not strong. Currently, the research on iris localization has begun to transition to the segmentation of the inner and outer circle contours and target localization. Fen et al. proposed the Iris-RCNN (Xin Feng, Wenxing Liu, et al. "Iris R-CNN: Accurate iris segmentation and localization in non-cooperative environment with visible illumination." Pattern Recognition Letters (2021)) model based on Mask-RCNN (Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. "Mask r-cnn." In Proceedings of the IEEE international conference on computer vision, pp. 2961-2969. 2017) to output feature vectors of a fixed dimension for recognition. Since it is based on Mask-RCNN and the backbone network is ResNet50 with FPN (Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. "Deep residual learning for image recognition." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778. 2016.), the number of model parameters is large (the model size is 252.1M). Although the robustness of localization is improved, it is difficult to meet the speed requirements in practical applications. The IrisParseNet method converts the problem of iris inner and outer circle localization into a contour segmentation problem and integrates it into iris segmentation, which is a multi-task network.While segmenting the iris mask, the model additionally predicts the contour masks of the inner and outer circles, and then performs post-processing on the contour masks to obtain the boundaries of the inner and outer circles of the iris. Although this method improves the positioning accuracy, on visible light images with high noise, due to the overly blurred predicted boundary contours, the positioning robustness of the model is not strong. Moreover, on the inner and outer circle contour masks, fine post-processing is required to obtain the parameters (center point coordinates and radius) for inner and outer circle positioning, which increases the overall processing time of the model.

[0008] To address the above problems, the present invention proposes a multi-task real-time processing network (Real-Time Iris localization and Segmentation multi-task Network, RISNet) that integrates localization and segmentation. The network processes the iris localization and iris segmentation tasks in parallel, while improving real-time performance, the model also maintains good localization and segmentation effects. Summary of the Invention

[0009] The object of the present invention is to provide a real-time iris localization and segmentation method for quickly and accurately performing iris localization and segmentation. The proposed method is a multi-task real-time processing network that integrates localization and segmentation. The network processes the iris localization and iris segmentation tasks in parallel. While improving real-time performance, the model also maintains good localization and segmentation effects, and at the same time solves the problems in the background technology, such as the existing localization methods based on iris masks or contour masks having weak positioning effects, low robustness, and excessive dependence on fine post-processing on visible light images.

[0010] To achieve the above object, the present invention adopts the following technical solutions:

[0011] The present invention provides a real-time iris localization and segmentation system, the system includes:

[0012] A backbone network for extracting image features from the original image to obtain an image feature map;

[0013] A semantic segmentation subnet for segmenting the iris from the image feature map;

[0014] A center point localization subnet for locating the center point positions of the inner and outer circles of the iris;

[0015] A size regression subnet for predicting the radii and offsets of the inner and outer circles of the iris.

[0016] Further, the center point localization subnet includes:

[0017] 3 convolutional modules, a 1×1 convolutional layer, and a sigmoid function;

[0018] The convolution module includes: a 3×3 convolution layer, a BatchNorm normalization layer, and a ReLU activation function.

[0019] Furthermore, the size regression subnet includes:

[0020] Three convolution modules, a convolution layer, and a sigmoid function;

[0021] The convolution module consists of: a 3×3 convolution layer, a BatchNorm normalization layer, and a ReLU activation function;

[0022] The input and output of the convolution module have the same size;

[0023] The sigmoid function normalizes the predicted value to [0, 1].

[0024] Furthermore, the size regression subnet includes:

[0025] An outer circle size regression branch and an inner circle size regression branch;

[0026] Both the outer circle size regression branch and the inner circle size regression branch include:

[0027] A radius prediction subnet and an offset prediction subnet;

[0028] The radius prediction subnet includes: a 3×3 convolution layer, a BatchNorm normalization layer, a ReLU activation function, and a 1×1 convolution layer; the input and output of the convolution layer have the same size, and the channel dimension of the input of the radius prediction subnet is 32;

[0029] The offset prediction subnet includes: a 3×3 convolution layer, a BatchNorm normalization layer, a ReLU activation function, a 1×1 convolution layer, and a sigmoid function.

[0030] Furthermore, the backbone network includes:

[0031] An Encoder module and a Decoder module;

[0032] The Encoder module includes two ConvX layers and two Conv2X layers;

[0033] The Decoder module includes four ConvX layers;

[0034] One Conv4X layer is also arranged between the Encoder module and the Decoder module;

[0035] Each layer structure of the Encoder module reduces the width and height of the image by half while doubling the depth.

[0036] Each layer structure of the Decoder module doubles the width and height of the image while reducing the depth by half.

[0037] Furthermore, the loss calculation of the center point localization subnet is as follows:

[0038] Let the loss of the center point localization subnet be Pair Loss;

[0039] Given a circular bounding box, a ground truth heatmap G is generated using a Gaussian kernel H*W*1 , and the non-zero region in the heatmap is called the Gaussian region;

[0040] Since the iris pupil region must be located within the iris region, the non-zero Gaussian region of the inner circle generated is always located within the non-zero Gaussian region of the outer circle. The overlapping region of the two is defined as A m , and the points in this region are defined as p;

[0041] Since the inner and outer circles of the iris always appear in pairs and there is an obvious constraint relationship between the radii, there are the following optimization problem and constraint conditions:

[0042]

[0043] Among them, Q represents the optimization objective, represents the model prediction value, and ∈ is the true difference between the inner circle radius and the outer circle radius;

[0044] The original optimization problem is transformed into a Lagrangian function:

[0045] L(x, λ p , λ q ) = f(x) + λ p L radius + λ q L constraint

[0046] Among them, L radius The formula is expressed as follows:

[0047]

[0048] Among them, λ kis a hyperparameter, which is set to 1 in the experiment. (i, j, k) represents the position and the category of the point. || is the L1 norm. G(i, j) is the Gaussian value corresponding to the point. Only the samples falling within the Gaussian region are regarded as positive samples, and they are all responsible for predicting the target radius. The samples outside the Gaussian region are regarded as negative samples and they do not make radius predictions;

[0049] L constraint The formula is expressed as follows:

[0050]

[0051] Among them, p i,j,inner is the coordinate with the predicted category of the inner circle, is the predicted radius of the inner circle corresponding to the coordinate; p i,j,outer is the central point coordinate with the predicted category of the outer circle, is the predicted radius of the corresponding outer circle; ∈ is the distance constraint term, which is the true distance between the inner circle radius and the outer circle radius;

[0052] Pair Loss = λ p L radius + λ q L cons

[0053] λ q and λ q are two constant terms.

[0054] Furthermore, the total loss function of the system is:

[0055] L = w loc * L loc + w reg * L reg + w mask * L mask

[0056] Among them, L reg is Pair Loss, L loc is Modified Focal Loss, L mask is the cross - entropy loss; w loc , w reg , w mask These are three hyperparameters, which are set to 1, 1, and 10 respectively.

[0057] The present invention also provides a real - time iris localization and segmentation method, including the following steps:

[0058] Adopt a backbone network to extract image features from the original image to obtain an image feature map;

[0059] Input the image feature map into the semantic segmentation subnet for iris segmentation and predict the iris mask;

[0060] Input the image feature map into the center point localization subnet and the size regression subnet, and perform inner and outer circle localization in the way of double center point confidence prediction and double-branch regression to achieve iris localization.

[0061] Further, the center point localization subnet realizes iris localization in the way of double center point confidence prediction, including:

[0062] For a heatmap, the value distribution satisfies the Gaussian distribution, and the coordinates where the maximum Gaussian value is located correspond to the center point of the circle; just sort the predicted values of the heatmap by size and only take the coordinates where the maximum value is located to directly obtain the unique center point of the circle.

[0063] The predicted value of the heatmap is equivalent to the confidence level that the coordinate point where the predicted value is located becomes the center point of the circle. The higher the confidence level, the greater the possibility that the coordinate point becomes the center point of the circle; therefore, the problem of circle center point localization is transformed into the problem of point confidence prediction.

[0064] Since the inner circle and the outer circle appear in pairs, the network always predicts two heatmaps corresponding to the inner circle and the outer circle. The problem of localizing the center points of the iris inner circle and outer circle is transformed into the problem of confidence prediction of a pair of points, that is, sort the two output heatmaps by value respectively, and take the coordinate points where the maximum values are located as the center points of the inner circle and the outer circle.

[0065] Assume that the predicted results are two heatmaps Then the value of the heatmap is the confidence level, and the coordinates corresponding to the maximum value of the heatmap are the final localization results of the inner circle center point and the outer circle center point;

[0066] The heatmap is output by the center point localization subnet and is in the form of two tensors. The width and height of the tensor are H and W, and each position has a predicted value. The range of the predicted value is 0-1, which satisfies the Gaussian distribution. The position where the maximum value is located corresponds to the center point of the circle.

[0067] Further, the size regression subnet realizes iris localization in the way of double-branch regression, including:

[0068] Let p k =(x k , y k ) represent the center point of the circular bounding box, where k is the inner circle category or the outer circle category;

[0069] Use the heat-map obtained by the center point localization subnet to determine the center point position p k ;

[0070] Based on the center point position, predicting the target radius through a radius prediction subnet, the steps include:

[0071] Define that the radius prediction subnet outputs a tensor whose width and height are the same as the heatmap output by the center point positioning subnet; let the predicted value in the tensor be r x,y , where (x, y) are coordinates, then the meaning of the predicted value r x,y is: when (x, y) is the center point of the circle, the radius value corresponding to this center point;

[0072] Based on the center point coordinates p k =(x k , y k ) of the circle output by the center point positioning subnet, according to this coordinate, in the tensor predicted by the radius prediction subnet , find the value r x,y corresponding to this coordinate, and this value is the radius value of the circle corresponding to the center point output by the center point positioning subnet.

[0073] Performing offset prediction through an offset prediction subnet, the steps include:

[0074] Since the coordinate values of the real circle may be floating-point numbers, while the center point coordinates output by the radius prediction subnet are integers, in order to make up for the coordinate accuracy loss caused by rounding of floating-point numbers, the prediction tensor of the offset prediction subnet The predicted value in the tensor is the offset value of the predicted coordinate, ranging from 0 to 1. Based on the center point coordinates output by the center point positioning subnet, in the tensor , find the offset value corresponding to this center point.

[0075] The present invention has at least the following beneficial effects:

[0076] The present invention proposes an efficient and fast end-to-end multi-task iris localization and segmentation network RISNet to simultaneously process iris localization and segmentation. First, RISNet performs inner and outer circle localization in the way of double center point confidence prediction and double-branch regression. Moreover, the localization network and the segmentation network share the same feature map, and the network head outputs the inner and outer circle localization results and the iris mask at the same time. The model output is the final result without post-processing. Second, aiming at the particularity of iris double-circle localization; finally, the loss function is optimized, and PairLoss is proposed. To this end, in order to meet the real-time requirement, a lightweight feature extraction network is designed. Experimental results on 4 public datasets: CASIA-Distance.v4, CASIA-M, NICEII, I-Social-DB (including 2 infrared environment datasets and 2 visible light environment datasets) show that the proposed method has strong localization robustness and high accuracy. Compared with the current optimal method IrisParseNet, both the localization accuracy and the segmentation accuracy are improved, and at least a 3-fold speed increase is achieved. On the CASIA-Distance dataset, a localization accuracy of 93.13% and a segmentation accuracy of 86.24% are achieved, and the localization accuracy is improved by 3.17%. On the CASIA-M dataset, the localization and segmentation accuracies are improved by 4.34% and 0.64% respectively. On the NICEII dataset, a localization accuracy of 82.85% and a segmentation accuracy of 88.82% are achieved, with improvements of 2.07% and 1.21% respectively. On the I-Social-DB dataset, a localization accuracy of 86.09% and a segmentation accuracy of 86.96% are achieved, with improvements of 1.83% and 0.86% respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0078] Figure 1 It is the flowchart of iris recognition preprocessing;

[0079] Figure 2 It is an example diagram of iris images: (a) and (b) are infrared iris images, and (c-f) are visible light iris images of different degradation types: (c) specular occlusion, (d) blur, (e) eyelid occlusion, (f) gaze deviation;

[0080] Figure 3 It is the structural diagram of the RISNet model;

[0081] Figure 4 It is the structural diagram of the backbone network;

[0082] Figure 5 It is an example diagram of a dataset;

[0083] Figure 6 It is the IoU of a circular bounding box;

[0084] Figure 7 It is a diagram of the loss change. (*) indicates the use of Pair Loss, and without (*) indicates the use of smooth-l1-loss;

[0085] Figure 8 It is a comparison of the localization and segmentation effects on different dataset test cases. (a) CASIA-D4 dataset; (b) CASIA-M dataset; (c) NICEII dataset; (d) I-Social-DB dataset. Detailed implementation manners

[0086] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.

[0087] The main objective of the present invention is to provide a real-time iris localization and segmentation method, specifically as follows:

[0088] 1. Method

[0089] RISNet is a concise, unified, and real-time end-to-end multi-task network, mainly composed of a backbone network and three parallel network heads. As Figure 3 shown, iris localization is mainly responsible for by the center point localization subnet and the size regression subnet, and iris segmentation is mainly responsible for by the semantic segmentation subnet. The three parallel subnets share the same feature map.

[0090] In the present invention, we first introduce the design idea of the RISNet network framework, including dual center point localization and dual regression design, semantic segmentation subnet, backbone network design. Then, we introduce Pair loss to optimize the radius prediction.

[0091] 1.1 Network architecture design

[0092] Dual center point confidence prediction

[0093] The iris region is defined by an inner circle and an outer circle. Regarding the inner circle and the outer circle of the iris as different target categories, then iris localization becomes an object detection task with an overlapping region.

[0094] Current center - based object detection methods such as CenterNet (Duan, Kaiwen, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian. "Centernet: Keypoint triplets for object detection." In *Proceedings of the IEEE / CVF international conference on computer vision*, pp. 6569 - 6578. 2019.), FCOS (Tian, Zhi, Chunhua Shen, Hao Chen, and Tong He. "Fcos: Fully convolutional one - stage object detection." In *Proceedings of the IEEE / CVF international conference on computer vision*, pp. 9627 - 9636. 2019) etc. use multi - level feature pyramids (Multi - level FPN) to solve the problem of detecting overlapping objects with different scales but the same center point; in addition, there are also methods using dense detection (one center predicts multiple object scales) to handle overlapping object detection. However, the above methods are not suitable for the iris detection task. For a detailed sample analysis, see Table 1 in the ablation experiment. The main reasons are as follows: (1) It is difficult to find appropriate hyperparameters to divide the inner - circle size and the outer - circle size, so multi - level FPN cannot be used; (2) Not all inner - circles and outer - circles have the same center point, so using dense prediction will introduce a large number of redundant detections.

[0095] Considering the paired characteristics of the inner - circle and outer - circle of the iris, the present invention converts the problem of locating the center points of the inner - circle / outer - circle into the problem of predicting the confidence of a pair of center points: the center - point localization subnet predicts 2 heatmaps (corresponding to the inner - circle category and the outer - circle category respectively). The value of the heatmap is the confidence, and the coordinates corresponding to the maximum value are the final localization results (the inner - circle center point and the outer - circle center point). The detailed process is as follows:

[0096] 1) Based on the center - point localization subnet of the deep neural network, according to the input image, the network automatically infers and outputs a heatmap. For this heatmap, the value distribution should ideally satisfy the Gaussian distribution in theory. The coordinates where the maximum Gaussian value is located correspond to the center point of the circle. Therefore, by simply sorting the predicted values of the heatmap according to their magnitudes and only taking the coordinates where the maximum value is located, the unique center point of the circle can be directly obtained.

[0097] 2) Based on the above representation, the predicted value of the heatmap is equivalent to the confidence level that the coordinate point where the predicted value is located becomes the center point of the circle. The greater this confidence level, the greater the possibility that the coordinate point becomes the center point of the circle. Thus, the problem of locating the center point of the circle is transformed into the problem of predicting the confidence level of the point.

[0098] 3) Since the inner circle and the outer circle appear in pairs, the network always predicts two heatmaps (corresponding to the inner circle and the outer circle respectively). Thus, the problem of locating the center points of the inner and outer circles of the iris is transformed into the problem of predicting the confidence levels of a pair of points. That is, sort the two output heatmaps by value respectively, and take the coordinate points where the maximum values are located (a total of 2), which are the center points of the inner circle and the outer circle.

[0099] Therefore, the localization subnet only needs to predict the confidence level without having to concern about the classification problem of the points. Moreover, after only one sorting, the localization subnet can output the final pair of center points (the coordinate of the center point of the inner circle and the coordinate of the center point of the outer circle).

[0100] Through the dual-center-point confidence level prediction method, the problem of center point contention for overlapping targets can be effectively avoided, and redundant detection can also be avoided, thus saving the model localization time.

[0101] Definition of positive and negative samples

[0102] In the definition of positive and negative samples around the center point, only the target center point is regarded as a positive sample, and all other positions are regarded as negative samples. The generation method of the Ground-truth heatmap is the same as that in the literature (Duan, Kaiwen, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian. "Centernet: Keypoint triplets for object detection." In *Proceedings of the IEEE / CVF international conference on computer vision*, pp. 6569-6578. 2019.). The Gaussian value corresponding to the coordinate is calculated using a Gaussian kernel. The closer to the center point, the higher the corresponding value. All generated Gaussian values are in the range of [0, 1]. During the training process, it is optimized using the Modified Focal Loss (Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar, “Focal loss for dense object detection,” ′in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2980–2988).

[0103] Center point localization subnet

[0104] The center point localization subnet consists of 3 convolutional modules, 1 convolutional layer, and a sigmoid function. The convolutional block consists of a 3×3 convolutional layer (stride of 1, padding of 1), a normalization layer, and an activation function ReLU. The input and output of the convolutional block have the same size. The output part of the subnet uses a sigmoid function to normalize the predicted value to [0, 1].

[0105] Dual regression branch

[0106] To correspond to the dual center point localization module, two parallel branches are used in the radius regression part based on the center point. Each branch consists of a radius prediction subnet and an offset prediction subnet. The radius prediction subnet consists of a 3×3 convolutional layer (pad of 1, stride of 1), a normalization layer BatchNorm, and an activation function ReLU. The input and output have the same size, and the channel dimension of its input is 32.

[0107] Let p k =(x k , yk ) represents the center point of the circular bounding box, where k is the inner circle category or the outer circle category, and the heat-map predicted by the center point positioning module is used to determine the center point position p k , and then based on this center point, the predicted value of the radius prediction subnet is the predicted target radius. The algorithm is described as follows:

[0108] 1) The radius prediction subnet outputs a tensor (with the same width and height as the heatmap output by the center point positioning subnet), and this tensor is automatically predicted by the deep neural network. Let the predicted value in the tensor be r x,y , where (x,y) are the coordinates. Then the predicted value r x,y means: when (x,y) is the center point of the circle, the radius value corresponding to this center point.

[0109] 2) Based on the center point coordinates p k =(x k , y k ) of the circle output by the center point positioning subnet, according to this coordinate, in the tensor predicted by the radius prediction subnet, find the value r x,y corresponding to this coordinate, and this value is the radius value of the circle corresponding to the center point output by the center point positioning subnet.

[0110] Since in the center point positioning module, a floor operation is used when mapping the target center point coordinates to the ground-truth heatmap, in order to recover the discretization error caused by integer operations, the model additionally predicts the center point offset. The offset prediction subnet consists of a 3×3 convolution (pad is 1, stride is 1), a normalization layer BatchNorm, an activation function ReLU, and a sigmoid layer. Similar to the literature (Duan, Kaiwen, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian. "Centernet: Keypoint triplets for object detection." In Proceedings of the IEEE / CVF international conference on computer vision, pp. 6569-6578. 2019), only the pixel points located at the target center will predict the offset values. The L1 loss is used for optimization in the experiment.

[0111] Semantic segmentation subnet

[0112] The semantic segmentation network predicts the iris mask, which consists of three 3×3 convolutional blocks, a 1×1 convolutional layer, and a sigmoid function in the channel dimension. Each 3×3 convolutional block consists of a 3×3 convolution with a stride of 1 and a pad of 1, a batch normalization layer, and a ReLU activation function. The cross-entropy loss function commonly used in segmentation methods is used for optimization.

[0113] Backbone network

[0114] Based on the Encoder-Decoder type U-shaped segmentation network and the fully convolutional type segmentation network, good results have been achieved in semantic segmentation. For example, SegNet (Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE transactions on pattern analysis and machine intelligence, 39(12):2481–2495, 2017) uses an encoder-decoder structure to recover high-resolution feature maps; PSPNet (Zhao, Hengshuang, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. "Pyramid scene parsing network." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2881-2890. 2017) designs a pyramid pool to capture global context information. However, due to high-resolution inputs and complex network connections, most methods require a large amount of computational cost, which does not meet the real-time requirements of iris segmentation and localization.Other segmentation methods such as BiseNet (Changqian Yu, Changxin Gao, Jingbo Wang, Gang Yu, and Nong Sang. Bisenet v2: Bilateral network with guided aggregation for real-time semantic segmentation. arXiv preprint arXiv:2004.02147, 2020) use a two-stream network that combines a lightweight backbone ResNet18 (Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. "Deep residual learning for image recognition." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778. 2016) and a spatial path as the encoder, but the structure is too redundant; Stdc-seg (Fan, Mingyuan, Shenqi Lai, Junshi Huang, Xiaoming Wei, Zhenhua Chai, Junfeng Luo, and Xiaolin Wei. "Rethinking bisenet for real-time semantic segmentation." In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2021) uses a single-stream network based on the STDC module as the encoder part, and then uses a deconvolution layer to expand the feature size to the same size as the input size. Although the speed is improved, a lot of detailed information is lost in multiple downsamplings.

[0115] To achieve real-time and accurate iris segmentation, considering the above deficiencies, the present invention designs an efficient backbone network that achieves a good balance between speed and accuracy. The backbone network has an Encoder-Decoder structure. The Encoder part extracts image features, and the Decoder part restores the size of the feature map. We only use 1×1 convolution and 3×3 convolution layers to construct the basic modules. The backbone network is built using basic modules with dense connections, and a total of 3 different basic modules are used: the ConvX module, which consists of a 3×3 convolution layer with a stride of 1 and a padding of 1, a batch normalization layer (BatchNorm), and a ReLU activation layer; the Conv2X module, which consists of a 1×1 convolution layer and 2 ConvX modules, and the structure is as shown in Figure 4 (b); the Conv4X module, which consists of a 1×1 convolution layer and 3 ConvX modules, and the structure is as shown in Figure 4 (c).

[0116] In the Encoder part, after each layer, the width and height of the feature map are halved, and the depth is doubled; in the Decoder part, after each layer, the width and height of the feature map are doubled, and the depth is halved. The change table of the channel dimensions of each layer: [3, 32, 64, 128, 256, 512, 256, 128, 64, 32].

[0117] The overall structure of the backbone network is as shown in Figure 4 (a). The depths of both the decoder and the encoder are 4. Between each layer of the encoder and the decoder, a shortcut connection is used for connection. The number of channels of the feature map output by the backbone network is 32, and the width and height remain unchanged.

[0118] 1.2 Center point to bounding box coordinates

[0119] In the inference stage, the center point localization module outputs 2 heatmaps representing the confidence of the inner circle center point and the confidence of the outer circle center point respectively. Sort the confidences, and the coordinate position corresponding to the maximum value is the predicted target center point (x k , y k ), k ∈ [inner, outer]. Based on this center point, the corresponding predicted radius and offset can be directly determined.

[0120] For the center point position (x, y) and the predicted radius and offset (δ x , δ k ), the calculation formula for the corresponding bounding box coordinates is as follows:

[0121]

[0122] where s is the scaling ratio of the feature map output by the backbone network compared to the size of the input image. Since the predicted radius is the size of the target's original pixels, therefore is the positioning result of the final circular bounding box.

[0123] 1.3 Double-circle Pair Loss

[0124] As mentioned above, the iris localization task is converted into a pair of center point confidence prediction and regression problems. In the field of object detection and segmentation, in most cases, smooth-l1 loss (Malik. Girshick, Ross, et al. "Rich feature hierarchies for accurate object detection and semantic segmentation." Proceedings of the IEEE conference on computer vision and pattern recognition. 2014.) and IoU loss (Yu, Jiahui, Yuning Jiang, Zhangyang Wang, Zhimin Cao, and Thomas Huang. "Unitbox: An advanced object detection network." In Proceedings of the 24th ACM international conference on Multimedia, pp. 516 - 520. 2016) are effective methods for supervised regression problems. However, these optimization methods only consider independent objects and ignore the correlation between objects. Considering the paired appearance characteristics of the inner and outer circles of the iris and the correlation between object sizes, an effective Double-circle Pair Loss (hereinafter referred to as Pair Loss) is proposed.

[0125] First, it is the problem definition. Given a circular bounding box, use a Gaussian kernel to generate the ground truth heatmap G H*W*1 , the non-zero region in the heatmap is called the Gaussian region. Due to the physiological characteristics of the iris annular region: the pupil region must be located within the iris region. Therefore, the non-zero Gaussian region of the inner circle generated is always located within the non-zero Gaussian region of the outer circle. Define the overlapping region of the two as A m, the points within this region are defined as p. Since the inner and outer circles of the iris always appear in pairs and there is an obvious constraint relationship between their radii, there are the following optimization problems and constraint conditions:

[0126]

[0127] Among them, Q represents the optimization objective, represents the model prediction value, ∈ is the true difference between the inner circle radius and the outer circle radius. Then the original optimization problem is converted into a Lagrangian function:

[0128] L(x, λ p , λ q ) = f(x) + λ p L radius + λ q L constraint (3)

[0129] Among them, the formula of L radius is expressed as follows:

[0130]

[0131] Among them, λ k is a hyperparameter, which is set to 1 in the experiment. (i, j, k) represents the position and the category of this point, || is the L1 norm, G(i, j) is the Gaussian value corresponding to this point. Only the samples falling within the Gaussian region are regarded as positive samples, and they are all responsible for predicting the target radius. The samples outside the Gaussian region are regarded as negative samples, and they do not make radius predictions.

[0132] The formula of L constraint is expressed as follows:

[0133]

[0134] Among them, p i,j,inner is the coordinate of the inner circle with the predicted category, is the predicted radius of the inner circle corresponding to the coordinate; p i,j,outer is the center point coordinate of the outer circle with the predicted category, is the predicted radius of the corresponding outer circle; ∈ is the distance constraint term, which is the true distance between the inner circle radius and the outer circle radius.

[0135] Finally, Pair Loss is composed of the above two constraint terms (4) and (5):

[0136] Pair Loss = λ p L radius + λ q L cons (6)

[0137] In the experiment, the two constant terms λq and λ q are both set to 1.

[0138] The Pair Loss proposed by the present invention has two advantages: (1) It is differentiable and supports backpropagation. (2) It makes an overall prediction for the regression targets (the inner circle and the outer circle).

[0139] 1.4 Total loss function

[0140] The total loss function of the model is:

[0141] L = w loc *L loc + w reg *L reg + w mask *L mask (7)

[0142] where L reg is the Pair Loss, L loc is the Modified Focal Loss (Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar, “Focal loss for dense object detection,” ′in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2980–2988.), and L mask is the cross-entropy loss. Although using the Modified Focal Loss for the center point loss can better train the model, this will lead to an unbalanced loss contribution of the segmentation part. To make up for the segmentation loss, in the experiment, the three hyperparameters w loc , w reg , and w mask are respectively set to 1, 1, and 10.

[0143] 2. Experiments

[0144] 2.1 Experimental details

[0145] The model is implemented based on the pytorch1.4 framework and accelerated using CUDA11.4 and cuDNN7.6.0. All experiments use the same optimizer and learning strategy: the optimizer uses Adam, the initial learning rate is 0.001, the weight decay is set to 5e-4, and an adaptive learning rate adjustment strategy is adopted, with a maximum of 150 epochs. All models are trained using 1 Nvidia GPU.

[0146] 2.2 Dataset

[0147] Four publicly available iris image datasets were used in the experiment:

[0148] 1) CASIA-Distance.v4 (B.I.Test.Casia.v4 Database. Accessed: Feb. 2020. [Online]. Available: http: / / www.idealtest.org / dbDetailForUser.do?id=4) dataset (hereinafter referred to as CASIA-D4). This dataset was captured in an infrared environment. The present invention uses the same 400 non-repeating iris images. The first 300 images are for the training set, and the last 100 images are for the test set. The resolution of each image is 640×480. See the example in Figure 5 (a).

[0149] 2) CASIA-M (Q. Zhang, H. Li, M. Zhang, Z. He, Z. Sun, and T. Tan, “Fusion of face and iris biometrics on mobile devices using near-infrared images,” in Proc. Chin. Conf. Biometric Recognit. Springer, 2015, pp. 569–578) dataset. This dataset was captured in an infrared environment. The resolution of each image is 400×400, and there are a total of 3000 non-repeating images, including 1500 images for the training set and 1500 images for the test set. See the example in Figure 5 (b).

[0150] 3) NICEII (H. Proenca, S. Filipe, R. Santos, J. Oliveira, and L. A. Alexandre, “The UBIRIS.v2: A database of visible wavelength iris images captured On-the-Move and At-a-Distance,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 32, no. 8, pp. 1529–1535, Aug. 2010) dataset. This dataset was captured in a visible light environment and has many noises such as illumination, reflection, and eyelids. See the example in Figure 5(c). There are a total of 2,000 non-repeating images, among which 1,000 are in the training set and 1,000 are in the test set. The image resolution is 300×400. See the example in Figure 5 (c).

[0151] 4) I-Social-DB dataset. This dataset was released by the literature (Labati, R. Donida, Angelo Genovese, Vincenzo Piuri, Fabio Scotti, and Sarvesh Vishwakarma. "I-SOCIAL-DB: A labeled database of images collected from websites and social media for iris recognition." Image and Vision Computing 105: 104058 (2021).) and is collected from social media (Twitter and Facebook). There are a total of 3,000 non-repeating images, among which 1,500 are in the training set and 1,500 are in the test set. The image resolution is 300×350.

[0152] The above four datasets are collected from different environments (infrared environment, visible light environment), different shooting conditions (cooperative and non-cooperative), and different shooting devices (infrared cameras and smartphones). The images in these databases contain various noise factors such as defocus, motion blur, gaze deviation, occlusion, specular reflection, illuminance change, iris deformation, and rotation. Therefore, it is convincing and reasonable to use these datasets to evaluate the performance of the proposed method.

[0153] 2.3 Evaluation Metrics

[0154] 1) Iris Localization Evaluation Metrics

[0155] For the iris localization task, we use the intersection over union (IoU) of the bounding box to measure the accuracy of localization. Figure 6 Shows an example diagram of the IoU of the bounding box represented by a circle.

[0156] In the following text, we use and to represent the inner circle IoU, the outer circle IoU, and the average IoU of the inner and outer circles respectively. The IoU ranges from [0, 1], and the closer it is to 1, the better the localization result.

[0157] 2) Iris Segmentation Evaluation Metrics

[0158] The present invention uses the mask IoU, F1, and E1 metrics to measure the segmentation effect of the model. The calculation methods of each metric are as follows.

[0159] The calculation formula for the E1 metric is as follows:

[0160]

[0161] where r and c represent rows and columns, G and M represent the ground truth mask and the predicted mask, respectively, represents the exclusive OR operation.

[0162] The calculation formula for the mask IoU metric is as follows:

[0163]

[0164] where P ij represents the number of pixels with the iris category predicted as the background, and P ji represents the number of pixels with the background category predicted as the iris, and P ii represents the number of correctly predicted pixels.

[0165] The calculation formula for the F1 metric is as follows:

[0166]

[0167] where TP represents the number of pixels predicted as the iris category and labeled as the iris category, FP represents the number of pixels predicted as the iris category but labeled as the background, and FN represents the number of pixels predicted as the background category but labeled as the iris.

[0168] 3) Model resource occupancy and inference time evaluation metrics

[0169] The present invention uses two metrics, the number of floating-point operations per second (FLOPs) and the number of frames processed per second (FPS), to measure the processing speed of the model. In addition, the average processing time of the model is also recorded.

[0170] 2.4 Ablation experiments

[0171] 1) Center point analysis

[0172] Table 1 shows the sample ratio and size statistics of the overlapping sample centers in the dataset. The multi-level feature pyramid (Multi-level FPN) can detect targets with the same center point but different sizes. However, this method requires clear size division. It can be found that it is difficult to find appropriate scale values to divide the inner circle size and the outer circle size on all datasets. Therefore, the multi-scale feature target prediction method cannot be used for iris localization tasks. In addition, the proportion of samples with the same center point is not very high. Using the method of dense detection (one center point corresponding to multiple target sizes) for iris localization will introduce a large number of redundant detections, and processing these redundant detections will increase the model processing time and seriously affect the processing speed. Therefore, the dual center point localization method proposed by the present invention is most suitable for iris localization tasks.

[0173] Table 1 Dataset Sample Analysis

[0174]

[0175] Note: Radius range refers to the size range of the radius of the inner circle or the outer circle of the iris in the dataset; Samecenter refers to the samples where the center points of the inner circle and the outer circle fall at the same position; the ratio refers to the proportion of the number of samples with the same center point of the inner / outer circle in the dataset.

[0176] 2) Comparison between Pair loss and smooth-l1

[0177] To test the effect of Pair loss, the loss changes and localization results during the training processes of smooth-l1-loss and pair loss were tested in the proposed model. In the experiment, except for the loss function of the regression part, the other settings were the same.

[0178] From Figure 7 it can be found that under the epochs before the model converges, the loss value after using smooth-l1 Loss is basically greater than the loss value after using Pair loss. Moreover, from the loss change curve, it can be found that the loss curve using Pair Loss drops faster than the loss curve using smooth-l1, indicating that the proposed loss function can make the model converge faster. At the same time, it can be found from Table 2 that except for the NICEII dataset, the model localization accuracy has been improved on other datasets. The inner and outer circle localization accuracy has been improved by a total of 1.4% on the CAISA-M dataset, and the inner and outer circle localization accuracy has been improved by a total of 2.16% on the I-Social-DB dataset, indicating the effectiveness of Pair loss in radius regression.

[0179] Table 2 Comparison of Localization Results between Pair loss and smooth-l1 loss

[0180]

[0181]

[0182] 2.5 Method Comparison

[0183] In multiple aspects, the proposed method is compared with FCDNN, FCEDN, IrisDenseNet, and IrisParseNet. These methods are all iris segmentation networks for infrared iris images and visible light iris images, and have excellent performance in segmentation.

[0184] FCDNN is an end-to-end fully convolutional iris segmentation method, and its network structure adopts a semi-parallel design; FCEDN uses an encoder-decoder network design; IrisDenseNet constructs an iris segmentation network through dense connection blocks to obtain deep features. It should be noted that these segmentation methods only output the iris mask. For the iris localization task, parametric processing methods such as circular Hough transform are used on the output iris mask to fit the parameters of the inner and outer circles of the iris.

[0185] IrisParseNet designs an intermediate network module with an attention mechanism to improve iris segmentation performance, and it is a multi-task network. In addition to predicting the iris mask, it also predicts the pupil contour mask and the iris outer contour mask at the same time. Iris localization is achieved by performing parametric post-processing on the predicted contour mask. It has the current optimal localization accuracy on infrared images.

[0186] For better fair comparison, we record the localization results, segmentation results, inference time, post-processing time, and FPS of each model on each dataset on two infrared datasets and two visible light datasets.

[0187] 1) Comparison of Iris Segmentation Effects

[0188] Table 3 shows the segmentation results of different models on the infrared datasets CASIA-D4 and CASIA-M. It can be found that FCEDN has the highest mIoU and F1 on the CASIA-D4

[26] dataset, reaching 86.32% and 92.49% respectively, and the E1 index of IrisParseNet is the best, which is 0.005312, indicating that the encoder-decoder type network structure is more suitable for the iris segmentation task than the fully convolutional network structure. On the CASIA-M dataset, the method proposed in the present invention achieves the best performance in terms of mIoU and F1 metrics, which are 88.37% and 93.76% respectively.

[0189] Table 3 Segmentation Results on the Infrared Dataset

[0190]

[0191]

[0192] Table 4 shows the segmentation results of the model on visible light iris images. It can be found that on the NICEII dataset, the method proposed in the present invention achieved an mIoU of 88.82% and an F1 of 93.85%, and on the I-Social-DB dataset, it achieved an mIoU of 86.96% and an F1 of 92.85%, both higher than other methods. In conclusion, although there are other methods that perform optimally in certain metrics, the method proposed in the present invention is higher than other methods in most metrics, indicating that the proposed method has high effectiveness in the iris segmentation task.

[0193] Table 4 Segmentation Results on the Visible Light Dataset

[0194]

[0195] 2) Comparison of Iris Localization Effects

[0196] Table 5 shows the localization results of each method on the infrared iris dataset. It can be found that among the three iris localization methods that fit the inner and outer circle parameters of the iris mask using methods such as Hough circle transformation, on the CASIA-D4 dataset and the CASIA-M dataset, the FCEDN method and the FCDNN method have the highest average localization accuracy, which are 89.87% and 85.04% respectively. The IrisParseNet method that uses post-processing on the contour mask predicted by the model to achieve iris localization achieved 89.96% and 89.13% on the two infrared datasets respectively, both higher than the former, indicating that the method of using the idea of segmentation for iris localization is better than the parameter fitting method. However, the method proposed in the present invention achieved 93.13% and 93.27% on the two infrared datasets respectively, both higher than the above methods, indicating that the method of regarding the localization problem as a regression problem and making the localization dependent on the feature map is better than other methods.

[0197] Table 5 Localization Results on the Infrared Dataset

[0198]

[0199] Table 6 shows the localization results of each method on the iris visible light dataset. It can be found that the localization results of each model are lower than those on the infrared images, indicating that the iris localization task is more challenging on visible light images with high noise. Moreover, from the localization results, it can be seen that for the three methods that perform inner and outer circle fitting on the segmentation mask to achieve localization, compared with the results on infrared images, their localization performance has severely degraded (the average box mIoU is lower than 80%), mainly because visible light images are accompanied by a large amount of noise, which causes the segmented iris mask to lose its edge regularity, and circular fitting is easily affected by irregular edges. IrisParseNet, due to its localization method of performing refined post-processing on the contour mask predicted by the model, has much higher results than the former. Although IrisParseNet achieved average However, it is still lower than the method proposed in the present invention (82.85% and 86.09%), indicating that the proposed method has better localization robustness.

[0200] Table 6 Localization Results on the Visible Light Dataset

[0201]

[0202] 3) Comparison of Model Inference Time and Post-Processing Time

[0203] As mentioned above, except for IrisParseNet, other comparison methods only output the iris mask, and for iris localization, it is obtained by performing parametric fitting on the iris mask. The IrisParseNet model performs parametric fitting based on the output contour mask. In the following text, the inference time refers to the calculation time for the model to obtain the output from the input, and the post-processing time refers to the time for performing preprocessing and parametric fitting on the iris mask or contour mask to obtain the inner and outer circle localization of the iris (i.e., the center point coordinates and radii of the inner and outer circles).

[0204] Note: Since the method RISNet proposed in the present invention is a multi-task network that integrates the iris localization task and the iris segmentation task, the model directly outputs the iris mask and the inner and outer circle coordinates and radii, so there is no post-processing step. Therefore, the total processing time of the model is only the inference time.

[0205] The processing times of each method on different datasets are shown in Tables 7 and 8. It can be found that due to the differences in image sizes among the four datasets, the inference times of the same model on different datasets are also different. In terms of the model inference time, the FCDNN method takes the least time on the four datasets, and then comes the method RISNet proposed by the present invention. In terms of the post-processing time, since factors such as the regularity of the interior or edge of the iris mask will affect parameter fitting, although the FCDNN, FCEDN, and IrisDenseNet methods all use the same parametric fitting method based on circular Hough transform, there are some slight deviations in the post-processing time. The post-processing of IrisParseNet is performed on the contour mask, and the post-processing time is affected by the shape of the contour edge, and the post-processing times on CASIA-M and the two visible light datasets are higher than those of other methods. In terms of the total time, due to the advantage of no post-processing in the method proposed by the present invention, the total time occupied is much lower than that of other methods. Correspondingly, the FPS is also higher than that of other methods.

[0206] Table 7 Model Processing Times on Infrared Datasets

[0207]

[0208] Table 8 Model Processing Times on Visible Light Datasets

[0209]

[0210] 4) Comparison of Model Parameters and FLOPs

[0211] Table 9 lists the network structure parameter sizes and FLOPs of different models. It is not difficult to find that for the method proposed by the present invention, the model file size is 12.4M and the number of network structure parameters is 3,142,346. At an input resolution of 320×320, the FLOPs of the proposed method are only 22.31G, which is lower than that of other methods, indicating that the method proposed by the present invention requires less computing resources compared to other methods.

[0212] Table 9 Comparison of Model GFLOPs and Parameter Quantities

[0213]

[0214]

[0215] 5) Visualization Comparison

[0216] Figure 8The segmentation and localization results of the methods FCEDN, IrisDenseNet, IrisParseNet and the method RISNet proposed by the present invention on each group of test images are shown. Among them, the first two groups are infrared images, and the last two groups are visible light images. The first column on the left is the ground truth, and the last column is the detection result of the method RISNet proposed by the present invention. It can be seen that the iris localization of the FCEDN and IrisDenseNet methods is easily affected by the irregular edges of the iris mask, and there is also a deviation in the localization of the inner circle (see the second column of the third row). The IrisParseNet method fails to segment the pupil area for the visible light iris image example with a gaze deviation (see the fourth column of the third row), resulting in the failure of post-processing and the inability to obtain the localization result. Compared with the ground truth, the proposed method RISNet has better localization effect and better segmentation effect on the edges.

[0217] 3. Summary

[0218] Aiming at the problems of the existing localization methods based on iris masks or contour masks, such as weak localization effect, low robustness, and over-reliance on refined post-processing on visible light images, a center point-based iris localization method is proposed according to the characteristics of the iris area, making the localization completely dependent on the feature map; aiming at the problem of a large number of model parameters in the existing models, the present invention constructs a lightweight backbone extraction network; in addition, a multi-task processing network is designed to perform iris localization and iris segmentation tasks simultaneously. Compared with other methods, it has improved in terms of localization accuracy and processing speed, and the model occupies less computing resources.

[0219] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.

Claims

1. A real-time iris localization and segmentation system based on a deep neural network, characterized in that, The system includes: A backbone network for extracting image features from the original image to obtain an image feature map; A semantic segmentation subnet for performing iris segmentation on the image feature map; A center point localization subnet for locating the center points of the inner and outer circles of the iris; A size regression subnet for predicting the radii and offsets of the inner and outer circles of the iris; The loss calculation of the center point localization subnet is as follows: Let the loss of the center point localization subnet be Pair Loss; Given a circular bounding box, generate the ground truth heatmap G using a Gaussian kernel H*W*1 , and the non-zero regions in the heatmap are called Gaussian regions; Since the iris pupil region must be located within the iris region, the generated inner circle non-zero Gaussian region is always located within the outer circle non-zero Gaussian region. The overlapping region of the two is defined as A m , and the points within this region are defined as p; Since the inner and outer circles of the iris always appear in pairs and there is an obvious constraint relationship between the radii, there are the following optimization problems and constraint conditions: Q: min f(x) p belongs to A m , where Q represents the optimization objective, represents the model prediction value, and ∈ is the true difference between the inner circle radius and the outer circle radius; Convert the original optimization problem into a Lagrangian function: L(x, λ p , λ q ) = f(x) + λ p L radius + λ q L constraint Among them, L radius The formula is expressed as follows: Among them, λ k is a hyperparameter, which is set to 1 in the experiment. (i, j, k) represents the position and the category of this point. || is the L1 norm, and G(i, j) is the Gaussian value corresponding to this point. Only the samples falling within the Gaussian region are regarded as positive samples, and they are all responsible for predicting the target radius. The samples outside the Gaussian region are regarded as negative samples, and they do not make radius predictions; L constraint The formula is expressed as follows: Among them, p i,j,inner is the coordinate with the predicted category of the inner circle, is the predicted radius of the inner circle corresponding to the coordinate; p i,j,outer is the center point coordinate with the predicted category of the outer circle, is the predicted radius of the corresponding outer circle; ∈ is the distance constraint term, which is the true distance between the inner circle radius and the outer circle radius; Pair Loss=λ p L radius +λ q L cons λ p and λ q are two constant terms.

2. The real-time iris localization and segmentation system based on a deep neural network according to claim 1, characterized in that, The center point localization subnet includes: 3 convolutional modules, a 1×1 convolutional layer, and a sigmoid function; The convolutional module includes: a 3×3 convolutional layer, a BatchNorm normalization layer, and a ReLU activation function.

3. The real-time iris localization and segmentation system based on a deep neural network according to claim 1, characterized in that, The size regression subnet includes: 3 convolutional modules, a convolutional layer, and a sigmoid function; The convolutional module consists of: a 3×3 convolutional layer, a BatchNorm normalization layer, and a ReLU activation function; The input and output of the convolutional module have the same size; The sigmoid function normalizes the predicted value to [0, 1].

4. The real-time iris localization and segmentation system based on a deep neural network according to claim 1, characterized in that, The size regression subnet includes: An outer circle size regression branch and an inner circle size regression branch; Both the outer circle size regression branch and the inner circle size regression branch include: A radius prediction subnet and an offset prediction subnet; The radius prediction subnet includes: a 3×3 convolutional layer, a BatchNorm normalization layer, a ReLU activation function, and a 1×1 convolutional layer; the input and output of the convolutional layer have the same size, and the channel dimension of the input of the radius prediction subnet is 32; The offset prediction subnet includes: a 3×3 convolutional layer, a BatchNorm normalization layer, a ReLU activation function, a 1×1 convolutional layer, and a sigmoid function.

5. A real-time iris localization and segmentation system based on a deep neural network according to claim 1, characterized in that, The backbone network includes: An Encoder module and a Decoder module; The Encoder module includes 2 ConvX layers and 2 Conv2X layers; The Decoder module includes 4 ConvX layers; There is also a Conv4X layer arranged between the Encoder module and the Decoder module; Each layer structure of the Encoder module reduces the width and height of the image by a factor of two and doubles the depth; Each layer structure of the Decoder module doubles the width and height of the image and halves the depth.

6. A real-time iris localization and segmentation system based on a deep neural network according to claim 1, characterized in that, The total loss function of the system is: L = w loc *L loc +w reg *L reg +w mask *L mask Among them, L reg is the Pair Loss, L loc is the Modified Focal Loss, L mask is the cross-entropy loss; w loc , w reg , w mask These are three hyperparameters, which are set to 1, 1, and 10 respectively.

7. A real-time iris localization and segmentation method for the system according to any one of claims 1-6, characterized in that, It includes the following steps: Use the backbone network to extract image features from the original image to obtain an image feature map; Input the image feature map into the semantic segmentation subnet to perform iris segmentation and predict the iris mask; Input the image feature map into the center point localization subnet and the size regression subnet, and perform inner and outer circle localization in a dual center point confidence prediction manner and a dual-branch regression manner to achieve iris localization.

8. A real-time iris localization and segmentation method according to claim 7, characterized in that, The center point localization subnet realizes iris localization in a dual center point confidence prediction manner, including: For a heatmap, the value distribution satisfies a Gaussian distribution, and the coordinates where the maximum Gaussian value is located correspond to the center point of the circle; just sort the predicted values of the heatmap by size and only take the coordinates where the maximum value is located to directly obtain the unique center point of the circle. The predicted value of the heatmap is equivalent to the confidence level of the coordinate point where the predicted value is located becoming the center point of the circle. The higher the confidence level, the greater the possibility that the coordinate point becomes the center point of the circle; therefore, the problem of circle center point localization is transformed into the problem of point confidence prediction. Since the inner circle and the outer circle appear in pairs, the network always predicts two heatmaps to correspond to the inner circle and the outer circle. The problem of localizing the center points of the iris inner circle and outer circle is transformed into a problem of confidence prediction for a pair of points, that is, sort the two output heatmaps by value respectively, and take the coordinate points where the maximum values are located as the center points of the inner circle and the outer circle. Suppose the predicted result is 2 heatmaps Then the value of the heatmap is the confidence level, and the coordinates corresponding to the maximum value of the heatmap are the final positioning results of the inner circle center point and the outer circle center point; The heatmap is output by the center point localization subnet and is in the form of two tensors. The width and height of the tensors are H and W, and each position has a predicted value. The range of the predicted value is 0-1, which satisfies a Gaussian distribution. The position where the maximum value is located corresponds to the center point of the circle.

9. A real-time iris localization and segmentation method according to claim 7, characterized in that, The size regression subnet realizes iris localization in a dual-branch regression manner, including: Let p k =(x k , y k ) denote the center point of the circular bounding box, where k is the inner circle class or the outer circle class; Heat-map obtained by using the central point to locate the subnet to determine the central point position pk; Based on the center point position, predict the target radius through the radius prediction subnet. The steps include: The defined radius prediction subnet outputs a tensor whose width, height, and center point are the same as those of the heatmap output by the location subnet; let the predicted value in the tensor be r x,y , where (x, y) are coordinates, then the predicted value r x,y means: when (x, y) is the center point of the circle, the radius value corresponding to the center point; The center point coordinates p of the circle output based on the center point positioning subnet k =(x k , y k ). According to this coordinate, in the tensor predicted by the radius prediction subnet, find the value r x,y corresponding to this coordinate. This value is the radius value of the circle corresponding to the center point output by the center point positioning subnet; Perform offset prediction through the offset prediction subnet. The steps include: Since the coordinate values of the real circle may be floating-point numbers, while the center point coordinates output by the radius prediction subnet are integers, to compensate for the coordinate accuracy loss caused by rounding of floating-point numbers, the prediction tensor of the offset prediction subnet The tensor The predicted value in is the offset value of the predicted coordinate, ranging from 0 to 1. Based on the center point coordinates output by the center point localization subnet, in the tensor find the offset value corresponding to the center point.