Elliptic Iris Location Method, System and Device Based on Deep Convolutional Neural Network

Through the elliptical iris positioning method based on deep convolutional neural network, the problems of low anti-interference capability and insufficient accuracy of the iris positioning algorithm in the prior art are solved, and high-precision iris positioning is achieved, which improves the accuracy and robustness of iris recognition technology.

CN120126205BActive Publication Date: 2025-08-01BEIJING JILIAN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510622369.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-01
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The existing iris positioning algorithm has low anti-interference capability and insufficient accuracy, making it difficult to achieve high-precision iris positioning in complex environments.

Method used

The elliptical iris positioning method based on deep convolutional neural network is adopted. By acquiring the iris image training data set, the deep neural network model is designed, training and inference is performed, and high-precision iris positioning is achieved by combining data preprocessing, candidate box IOU calculation, location information coding and loss design.

Benefits of technology

It improves the accuracy and anti-interference ability of iris positioning, and can accurately identify the boundaries between iris and pupils in complex environments, improving the accuracy and robustness of iris recognition technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126205B_ABST
    Figure CN120126205B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an elliptical iris localization method, system, and device based on a deep convolutional neural network. Applied to the field of iris recognition technology, the method includes marking a training data set to obtain elliptical parameter information of elliptical frames in the training data set; designing elements of a deep neural network model and constructing the deep neural network model; training the deep neural network model using the training data set; using the trained deep neural network model to perform forward inference on an iris image to be processed to obtain an inference result, where the inference result includes elliptical parameter information of a prediction box in the iris image; and performing post-processing on the inference result, and parsing elliptical parameters corresponding to the boundaries between the iris and the pupil in the iris image according to the elliptical parameter information of the prediction box to achieve high-precision iris localization. In this way, the technical problems of low anti-interference ability and the need for further improvement in iris localization accuracy in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of iris recognition technology, and in particular, to an elliptical iris localization method, system and device based on a deep convolutional neural network. Background Art

[0002] Iris recognition technology is a technology that uses the iris in the human eye for identity authentication and belongs to a type of human biometric recognition technology. Iris recognition has characteristics such as uniqueness, stability, non-contact, and high security, and is recognized as the most accurate and convenient biometric recognition technology. It has been widely used in various scenarios that require precise identity authentication, such as finance, security, checkpoints, access control, and insurance. Iris localization refers to accurately determining the outer boundary of the iris and the outer boundary of the pupil in the iris image of the human eye, and is one of the key technologies in iris recognition technology. Generally, it is considered that the boundary shapes of the pupil and the iris are close to circular, and many iris localization algorithms are also implemented based on the circular hypothesis. In fact, the more realistic boundary shapes of the iris and the pupil are elliptical, especially when the human eye is in a non-frontal gaze state. Therefore, an iris localization method based on the assumption that the boundary shapes of the iris and the pupil are elliptical can more realistically reflect the boundary shapes of the iris and the pupil, thereby improving the iris localization accuracy.

[0003] However, the existing technology has obvious deficiencies. In the early stage, iris localization algorithms based on the elliptical hypothesis all used image processing means to first detect the outer boundaries of the iris and the pupil, and then fit the elliptical parameters by means of approximation similar to the least squares method; this method has very low anti-interference ability, and the actual effect is worse than that of the iris localization algorithm based on the circular hypothesis. Therefore, it has not been actually applied and promoted. In addition, with the development of deep learning technology, object detection models have become mature, making it possible to implement a detection model with an elliptical shape as the target. Summary of the Invention

[0004] The present disclosure provides an elliptical iris localization method, system and device based on a deep convolutional neural network, which can achieve iris localization based on the elliptical hypothesis in a high-precision and high-efficiency manner, greatly promoting the development of iris localization technology. It solves the technical problems such as low anti-interference ability and insufficient iris localization accuracy in the existing technology.

[0005] According to a first aspect of the present disclosure, there is provided an elliptical iris localization method based on a deep convolutional neural network, including:

[0006] Obtain a training data set containing iris images, mark the training data set, and obtain the elliptical parameter information of the elliptical frames in the training data set;

[0007] Design the structure of the deep neural network model, data preprocessing, candidate boxes, calculation of the IOU of elliptical boxes, encoding of position information, selection of positive and negative samples, and loss design elements, and construct a deep neural network model. Use the training dataset to train the deep neural network model;

[0008] Use the trained deep neural network model to perform forward inference on the iris image to be processed, and obtain the inference result. The inference result contains the elliptical parameter information of the prediction box in the iris image;

[0009] Perform post-processing on the inference result. According to the elliptical parameter information of the prediction box, parse out the elliptical parameter corresponding to the boundary between the iris and the pupil in the iris image, and achieve high-precision iris localization.

[0010] According to the second aspect of the present disclosure, there is provided an elliptical iris localization system based on a deep convolutional neural network, including:

[0011] A marking module for marking the training dataset to obtain the ground truth. The training dataset contains the iris images to be trained, and the ground truth contains the elliptical parameter information of the elliptical boxes in the training dataset;

[0012] A construction module for constructing a deep neural network model;

[0013] A training module for training the deep neural network model using the training dataset;

[0014] An inference module for performing forward inference on the iris image to be processed using the trained deep neural network model to obtain the inference result; the inference result contains the elliptical parameter information of the prediction box in the iris image;

[0015] A post-processing module for performing post-processing on the inference result and parsing the elliptical parameter information corresponding to the boundary between the iris and the pupil in the iris image according to the position information of the prediction box.

[0016] According to the third aspect of the present disclosure, there is provided an electronic device. The electronic device includes: a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the method according to the first aspect of the present disclosure is implemented.

[0017] Compared with the prior art, the advantages and positive effects obtained by the present disclosure are:

[0018] The present disclosure proposes an effective method for detecting the outer boundaries of the iris and the pupil based on an elliptical hypothesis. The elliptical boxes are marked on the iris images in the training dataset to complete the model training, and the elliptical parameter information of the boundary between the iris and the pupil is parsed according to the elliptical parameter information of the prediction box given by the model, which greatly improves the accuracy and anti-interference ability of iris localization.

[0019] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation of the present disclosure. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0021] Figure 1 A flowchart of an elliptical iris localization method based on a deep convolutional neural network according to an embodiment of the present disclosure is shown;

[0022] Figure 2 A schematic diagram of information marking according to an embodiment of the present disclosure is shown;

[0023] Figure 3 A schematic structural diagram of a deep neural network model according to the present disclosure is shown;

[0024] Figure 4 A schematic structural diagram of a convolutional block ConvBlock according to the present disclosure is shown;

[0025] Figure 5 A schematic structural diagram of a basic block BaseBlock according to the present disclosure is shown;

[0026] Figure 6 A schematic structural diagram of a downsampling block DownBlock according to the present disclosure is shown;

[0027] Figure 7 A schematic diagram of a CTEBlock module according to the present disclosure is shown;

[0028] Figure 8 A schematic diagram of an InceptionBlock module according to the present disclosure is shown;

[0029] Figure 9 A block diagram of an elliptical iris localization system based on a deep convolutional neural network according to an embodiment of the present disclosure is shown;

[0030] Figure 10 A block diagram of an exemplary electronic device capable of implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0032] In addition, the term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after.

[0033] Figure 1 The flowchart of a method 100 for elliptical iris localization based on a deep convolutional neural network in the embodiments of the present disclosure is shown, as Figure 1 shown, the method 100 includes:

[0034] S110: Obtain a training data set containing iris images, mark the training data set, and obtain the elliptical parameter information of the elliptical frames in the training data set.

[0035] Optionally, in some embodiments, the elliptical parameter information of the elliptical frames in the training data set is the ground truth (GT); the information for marking the training data set includes the elliptical frame parameter information (loc) corresponding to the boundary shapes of the iris and the pupil in an iris image and the classification information (class) of the iris and the pupil.

[0036] It should be noted that in the embodiments, the elliptical frame parameter information corresponds to the elliptical parameters corresponding to the boundary shapes of the iris and the pupil, and the elliptical parameters include , representing the elliptical center coordinates corresponding to the iris or the pupil, representing the lengths of the two axes of the ellipse, being the axis of the ellipse and the axis in the clockwise direction angle.

[0037] The classification information of the iris and the pupil, 1 represents the iris, and 0 represents the pupil.

[0038] According to the above description, the ground truth expression of a complete eye is , where represents the elliptical parameters of the iris, and '1' represents that the category is the iris; The elliptical parameters representing the pupil, where '0' indicates that this category is the pupil; a complete eye should contain and only contain one valid iris information and one valid pupil information. The information marking method of the truth refers to the appendix Figure 2 , appendix Figure 2 In which, marking 1 represents the iris boundary, and marking 2 represents the pupil boundary.

[0039] In the embodiments of the present application, the accuracy of data annotation and model training, by accurately annotating the iris and pupil boundaries in the iris image and recording the elliptical parameter information (such as center coordinates, axis lengths, rotation angles, etc.), can provide high-quality annotation data for machine learning or deep learning models. These annotation data are the basis of model training and directly affect the performance of the model. Accurate annotation data can help the model better learn the features of the iris and pupil, thereby improving the accuracy and robustness of the model in tasks such as iris recognition and pupil detection. This is particularly important for biometric technologies (such as iris recognition) because iris recognition usually requires extremely high precision. Standardization of elliptical parameter information, the standardized annotation of elliptical parameter information (such as center coordinates, axis lengths, rotation angles, etc.) enables the iris and pupil boundaries in different images to be represented in a unified manner. Standardization helps the model maintain consistency when processing different images and avoid errors caused by different annotation formats. The standardized elliptical parameter information not only simplifies the model training process but also improves the interpretability of the model; researchers and developers can more intuitively understand the output of the model, facilitating optimization and debugging. Clarity of classification information, by marking the iris and pupil as 1 and 0 respectively, clearly distinguishes the category information of the two. The classification information helps the model distinguish the features of the iris and pupil during the learning process and avoid confusion. Clear classification information can improve the classification accuracy of the model in iris recognition tasks, especially in complex image environments (such as lighting changes, occlusions, etc.), and the model can more accurately identify the boundaries of the iris and pupil. Expression of complete eye information, by combining the elliptical parameter information of the iris and pupil into a complete GT (Ground Truth) expression method, ensures that the annotation information of each eye contains and only contains one iris and one pupil; this expression method makes the annotation information of each eye more complete and standardized. The complete GT expression method helps the model consider the features of both the iris and pupil simultaneously when processing a single eye, thereby improving the comprehensive performance of the model in tasks such as iris recognition and pupil detection; in addition, this expression method also provides convenience for subsequent multi-task learning (such as performing iris recognition and pupil detection simultaneously).

[0040] In summary, the iris images in the training dataset of this embodiment are accurately labeled, and the standardized expressions of the elliptical parameter information and classification information provide a high-quality data basis for model training. The labeled data can help the model better learn the features of the iris and pupil, thereby improving the performance of the model in tasks such as iris recognition and pupil detection. It not only improves the accuracy and robustness of iris recognition technology, but also provides strong support for the development of biometric technology. Accurate iris recognition technology has broad application prospects in fields such as security authentication and identity recognition, and can bring higher security and convenience to society. Generally speaking, the above steps lay a solid foundation for the further development of iris recognition technology through accurate data annotation and standardized information expression, and have important technical significance and application value.

[0041] S120: Design elements such as the structure of the deep neural network model, data preprocessing, candidate boxes, calculation of the ellipse box IOU, encoding of position information, selection of positive and negative samples, and loss design, and construct the deep neural network model.

[0042] Optionally, in some embodiments, the construction of the deep neural network model structure refers to Appendix Figure 3 ; For specific content, refer to Table 1, Figures 4 - 6 ; where Figure 4 Convolution block ConvBlock, Figure 5 Basic block BaseBlock, Figure 6 Downsampling block DownBlock; Figure 7 is the (multi-scale) fusion output block CTEBlock; Figure 8 is the feature fusion block InceptionBlock;

[0043] Table 1 Deep neural network model structure

[0044]

[0045] It should be noted that in the embodiment, the structure of the deep neural network model specifically includes:

[0046] The first stage: includes multiple consecutive downsampling processes. Its function is to extract features during the rapid image degradation process. In this embodiment, 4 downsampling processes are used, the DownBlock structure is used, and the downsampling rate is 2.

[0047] The second stage: includes multiple consecutive feature fusion processes. Its function is to fuse the initially extracted features. In this embodiment, 3 fusion processes are used, the InceptionBlock structure is used, and no downsampling is performed during the fusion process.

[0048] The third stage: a process that includes multi-scale output prediction box information, whose function is to output prediction results, including the position information of the prediction box, the classification information of the prediction box, etc. In this embodiment, 3 different scale channels are used.

[0049] The first stage (Stage 1): downsampling and feature extraction. The purpose of this stage is to accurately capture the iris pupil boundary features in an efficient manner. Therefore, in the design of the model structure, it is necessary to balance the relationship between accuracy and efficiency at the same time.

[0050] This stage uses the downsampling module DownBlock (such as Figure 6 ) which includes 2 consecutive BaseBlocks (such as Figure 5 ), and each BaseBlock includes 2 consecutive basic ConvBlocks (such as Figure 4 ). Its expression is:

[0051] ConvBlock = Conv + BN + RELU

[0052] BaseBlock = ConvBlock(3 3) + ConvBlock(1 )

[0053] DownBlock = BaseBlock(s = 1) + BaseBlock(s = 2)

[0054] It constitutes a Quad-Conv convolutional structure containing 4 convolutions. Among them, BaseBlock is composed of 3 3 + 1 1 convolutions, which can ensure the accuracy of feature extraction, increase the depth of the model, and at the same time, the 1 1 convolution is also beneficial to improving efficiency. DownBlock first uses a convolution with step = 1, and then uses a convolution with setp = 2 to achieve downsampling, which can further deepen the fusion degree between neighborhood features and improve the accuracy of feature extraction.

[0055] Considering that the Quad-Conv convolutional structure has a strong fine-grained feature extraction ability, in order to improve the execution efficiency of the model, the number of channels of the convolution is appropriately restricted. The initial channels are quickly expanded to 64 channels to ensure that enough original information enters, but the subsequent channels do not adopt a 2-fold expansion strategy, but are restricted within 128 channels to improve the execution efficiency of the model.

[0056] The second stage (Stage 2): feature fusion and receptive field adjustment. In this example, 3 consecutive InceptionBlocks (such as Figure 8 ) are used to achieve the required functions.

[0057] The InceptionBlock used includes 4 different sub-channels, and each sub-channel can perform feature extraction. After 3 consecutive combinations of InceptionBlocks, 64 different feature fusion methods can be formed.

[0058] At the same time, each InceptionBlock can form 3 different receptive fields. After 3 consecutive combinations of InceptionBlocks, 7 different receptive fields can be formed, significantly improving the model's object detection ability for different granularity scales.

[0059] In addition, there is a 1 ×1 convolutional channel in the InceptionBlock, which plays a role similar to a residual jump connection and can more effectively transmit gradient information to ensure the trainability of the deep model.

[0060] The third stage (Stage 3): multi-scale output and ellipse parameter prediction, including the ellipse prediction result and classification prediction result of the iris image. The prediction box of the iris image needs to be represented in the form of an ellipse, and the output layer is specially designed; it also includes:

[0061] Ellipse parameter regression head: Design an ellipse parameter regression head to directly predict the five parameters of the ellipse:

[0062]

[0063] In the formula, is the ellipse regularization module, is the regularization parameter.

[0064] Multi-scale prediction module: In this example, 3 different scales of feature fusion outputs are used to support target predictions of multiple different scales. The output points of the 3 output scales are respectively after the InceptionBlock, after one downsampling, and after another downsampling. Then, the different scale outputs of the 3 output points are fused through the CTEBlock to obtain the final multi-scale output result.

[0065]

[0066] Among them, are the feature maps of 3 scales respectively.

[0067] It should be noted that in the embodiments of this application, data preprocessing includes preprocessing of iris image data and preprocessing of labeled data.

[0068] The preprocessing of iris image data includes:

[0069] The iris image data undergoes various data augmentation processes. Operations such as deformation, cropping, horizontal flipping, and rotation, which may affect the position information of the labeled data, also adjust the corresponding position information of the labels.

[0070] The image after data augmentation contains at least one complete labeled eye image.

[0071] The final input image is scaled to a single-channel grayscale image of a specified size (width 640 × height 480), and after normalization processing, it is input into the deep neural network model to be trained.

[0072] Normalization processing formula:

[0073]

[0074] In the formula, represents the pixel value of the iris image after normalization processing; represents the j-th pixel value in the iris image before processing; represents the statistical average value of the j-th pixel value in the training dataset; represents the standard deviation of the j-th pixel value in the training dataset; N represents the total number of pixels in the iris image; represents the natural exponential function, which is used to enhance the smoothness of the pixel value distribution; represents the normalization constant, which ensures the mathematical integrity of the formula. First, each pixel value is normalized by subtracting its average value and dividing by the standard deviation. This step ensures that all pixel values are centered around zero; a Gaussian weighting function is introduced. By calculating the difference between each pixel value and the average value and using it as the input to the Gaussian function, a smooth weight value can be obtained; all weighted pixel values are divided by the normalization constant to ensure the mathematical integrity of the entire formula and make the normalized pixel values conform to the standard normal distribution.

[0075] The preprocessing of the labeled data includes the following steps:

[0076] In coordination with the augmentation processing operations that may affect the position information of the labeled data, the corresponding position information is adjusted.

[0077] After the position information is finally normalized, it is used as the new Ground Truth information and input into the network model to be trained.

[0078] The normalization processing method is ; where is the new coordinate point after normalization processing, is the coordinate point before processing; is the new ellipse axis length after normalization processing, The length of the elliptical axis for processing money; width is the pixel width of the corresponding image, and height is the pixel height of the corresponding image.

[0079] It should be noted that in the embodiments of this application, for the candidate elliptical bounding box design: there are 3 width sizes, 1 aspect ratio, and 3 densities for the candidate elliptical bounding box design, which are represented by different step sizes.

[0080] The marking method of the candidate elliptical bounding box is , where is the center coordinate of the candidate elliptical bounding box after normalization, is the axis length of the elliptical bounding box after normalization, is the elliptical bounding box axis and axis clockwise angle.

[0081] It should be noted that in the embodiments of this application, the position information encoding includes:

[0082] The input is the normalized elliptical parameter information Loc of an elliptical bounding box , where represents the center point coordinate of the elliptical bounding box, represents the axis length of the ellipse, is the elliptical bounding box axis and axis clockwise angle, expressed in radians.

[0083] The output is , where represents the center encoding result of the elliptical bounding box, represents the axis length encoding result of the elliptical bounding box, is the rotation angle encoding result of the elliptical bounding box.

[0084] Among them, , , is the normalized center coordinate of the candidate elliptical bounding box, is the normalized elliptical axis length of the candidate elliptical bounding box, is the encoding coefficient.

[0085] ,

[0086] is the encoding coefficient; , is the normalized elliptical rotation angle of the candidate bounding box.

[0087] The decoding process is the inverse process of the encoding process.

[0088] It should be noted that in the embodiments of this application, the calculation of the IOU of the elliptical box includes the following steps:

[0089] Use a polygon (such as a 16-sided polygon) to represent an ellipse, greatly reducing redundant information and improving accuracy.

[0090] Calculate the intersection polygon of the two polygons.

[0091] Calculate the area of the intersection polygon.

[0092] Calculate the intersection over union of the two polygons to obtain the IOU value.

[0093] In iris image processing, the calculation of the IOU (intersection over union) of the elliptical box is a key step for evaluating the similarity of two elliptical regions; a 16-sided polygon will be used to approximate the ellipse.

[0094] Approximate the ellipse with a 16-sided polygon (regular hexadecagon), and let the center of the ellipse be , the major axis radius be , the minor axis radius be , and the rotation angle be , then the parametric equation of the ellipse is:

[0095] ;

[0096]

[0097] In the formula, is the parameter angle, and the value range is ; Divide into 16 equal parts to obtain 16 vertex coordinates:

[0098]

[0099] The coordinates of each vertex are:

[0100]

[0101] For the calculation of the intersection polygon, assume there are two ellipses, represented by 16-sided polygons and respectively; it is necessary to calculate the intersection polygon of these two polygons;

[0102] Find all the intersection points of and ;

[0103] According to the intersection points, divide the polygon into multiple sub-polygons;

[0104] Select the sub-polygons located inside both polygons to form .

[0105] Calculation of the area of the intersection polygon :

[0106]

[0107] Wherein are the vertex coordinates of the polygon, is the number of vertices, and , .

[0108] The calculation of the area of the union polygon is the sum of the areas of the two polygons minus the intersection area:

[0109]

[0110] Wherein and are the areas of the two polygons respectively.

[0111] The final IOU value is the ratio of the intersection area to the union area:

[0112]

[0113] Through the above steps, the IOU value of the ellipse box can be accurately calculated in iris image processing, providing a reliable basis for matching and recognition.

[0114] It should be noted that in the embodiments of the present application, the selection of positive and negative samples includes the following steps:

[0115] The samples with IOU>0 between the prediction box and the candidate box and ranked among the top N are positive samples; the others are candidate negative samples.

[0116] Calculate the classification loss of all candidate negative samples and sort them.

[0117] Select the most difficult negative samples, that is, the M candidate negative samples with the largest loss, as the finally selected negative samples.

[0118] Wherein , α is the negative sample magnification factor, α = 10.

[0119] It should be noted that in the embodiments of the present application, in the design of the loss function, the training loss of the deep neural network model is designed as a combined loss, using a combination of 2 loss functions with different weights, including the loss function and the classification function.

[0120] The loss function is used to regress the position information, using smoothL1, and the calculation formula is:

[0121] ;

[0122] Among them, is the normalized encoded numerical difference value between the predicted box and the ground truth box.

[0123] The classification function uses the cross-entropy loss function to confirm whether the predicted box is an iris or a pupil, and the calculation method is as follows:

[0124]

[0125] Among them, is the total number of samples, is the total number of categories; represents the ground truth sample the th label of the category, using one-shot encoding; represents the predicted probability of the sample the th category. In this embodiment, K = 3, representing the iris, pupil, and background categories respectively.

[0126] In summary, the above two losses are superimposed according to different weights, and the calculation method is:

[0127]

[0128] Among them, is the calculation weight of the th loss; is the calculation value of the th loss; in this embodiment , .

[0129] In the embodiment of the present application, in the model structure design, the design of each module is to better extract image features while reducing the computational complexity. The convolutional block usually consists of a convolutional layer, an activation function (such as ReLU), and a normalization layer (such as BatchNorm), and its main role is to extract local features. Through the sliding window operation of the convolutional kernel, it captures details such as textures and edges in the image; the design of the convolutional block enables the model to efficiently process high-dimensional data, and at the same time, through multi-layer stacking, gradually extracts higher-level semantic features; extracts multi-level features, gradually abstracting from low-level to high-level; reduces the number of parameters and improves the computational efficiency; introduces non-linearity through the activation function to enhance the expression ability of the model. The convolutional block is the basis of the deep neural network and determines the feature extraction ability of the model; by reasonably designing the convolutional kernel size, stride, and number of channels, the performance of the model can be significantly improved. The basic block (BaseBlock, Figure 5 ) is an extension of the convolutional block. The downsampling block (DownBlock, Figure 6(1) It consists of convolutional layers with a stride, which is used to reduce the resolution of the feature map, reduce the computational amount, and at the same time expand the receptive field to capture more global features; reduce the size of the feature map and the computational complexity; enhance the robustness of the model to input size changes. The downsampling block is the key to the efficient operation of the model. By reasonably designing the downsampling strategy, the computational cost can be reduced while ensuring performance. The multi-branch structure (Inception module) uses a variety of different branch structures, which can adaptively adjust the weights of different branches and enhance the adaptability of feature extraction; combining multiple multi-branch structures can not only expand the receptive field, but also form more combined branches. Each combined branch channel has a different receptive field, which is more suitable for multi-scale object detection.

[0130] Data preprocessing is an important part of model training, which directly affects the convergence speed and final performance of the model; it is the basis of model training, and the preprocessing strategy can significantly improve the performance and stability of the model.

[0131] Calculate the IOU between the candidate box and the ellipse box, which is the ratio of the intersection and union of the rectangular boxes, and is used to evaluate the accuracy of the predicted box; the IOU calculation of the ellipse box is more complex and usually requires numerical integration or approximation algorithms, which is applicable to special scenarios. IOU is the core index of the object detection task and directly affects the localization accuracy of the model. By optimizing the IOU calculation, the detection performance of the model can be improved.

[0132] Position information encoding is to convert the position of the target box (such as the center coordinates, width and height) into a format that can be learned by the model. It is the key to the object detection task and can improve the localization accuracy and generalization ability of the model.

[0133] Positive and negative sample selection In the object detection task, the selection of positive and negative samples directly affects the training effect of the model; the selection of positive and negative samples is the core of model training, which can accelerate the convergence of the model and improve the detection performance.

[0134] The loss function is the objective of model optimization, which usually includes two parts: classification loss and regression loss. The classification loss measures the difference between the predicted category and the true category (such as cross-entropy loss); the regression loss measures the position difference between the predicted box and the true box (such as Smooth L1 loss). The design of the loss function directly affects the optimization direction of the model and can improve the detection accuracy and robustness of the model.

[0135] In summary, in this embodiment, by reasonably designing the structure of the deep neural network model, data preprocessing, candidate boxes, IOU calculation of ellipse boxes, position information encoding, positive and negative sample selection, and loss function, an efficient object detection model can be constructed. Attached Figure 3 、Table 1、 Figures 4 - 7 Each component in realizes different functions, collaborates together, and finally realizes the optimization and performance improvement of the model.

[0136] S130: Train a deep neural network model using a training data set.

[0137] Optionally, in some embodiments, split the labeled iris images and data into two sets: a training data set and a validation data set, with a ratio of 9:1.

[0138] Use the training data set for training and the validation data set for validating the deep neural network model and screening parameters.

[0139] Input the preprocessed iris images and labeled data of the training data set into the deep neural network. The deep neural network performs forward inference on the input iris images to obtain the current inference result.

[0140] Use the current inference result and the corresponding labeled data to calculate the classification loss and regression loss, and obtain the current combined loss.

[0141] The current loss is propagated back to the deep neural network through reverse gradient propagation, and together with other parameters such as the learning rate, adjust the weights of the current deep neural network.

[0142] Repeat the above process until the training ends.

[0143] The condition for the end of training is: complete the specified number of rounds (e.g., epoch = 120).

[0144] The conditions for screening the parameters of the deep neural network model are: the metrics of the validation data set are the best results of the current training and meet the expectations. The metrics in this embodiment include the IOU metric [IOU > 0.95] and the F1Score metric [F1Score > 0.99].

[0145] It should be noted that in the embodiments of this application, during the process of the deep neural network performing forward inference on the input iris images, the core lies in gradually extracting discriminative features from the iris images through multi-level feature extraction and information conversion, and finally generating the inference result, which specifically includes the following steps:

[0146] After the input iris images are preprocessed, they enter the primary feature extraction stage, and basic structural information such as texture, edges, and local details is extracted from the iris images; the specific process is as follows:

[0147] Local detail capture, through pixel-by-pixel local calculations, extract the minute changes and detail features in the iris images, manifested as the bright or color mutation regions in the iris texture;

[0148] Spatial relationship modeling, based on local details, further captures the spatial relationships between adjacent pixels to form preliminary structured features; for example, the spatial distribution between stripes and spots in iris texture.

[0149] Intermediate feature extraction, based on the primary features, further models the complex patterns in the iris image; extracts more semantically meaningful features, such as texture combinations, shapes and spatial distributions of local regions, etc.; the specific process is as follows:

[0150] Feature combination and abstraction: Combines and abstracts the primary features to form a higher-level feature representation; describes more complex visual patterns in the iris image, such as periodic changes in texture or shape features of local regions.

[0151] Context information fusion: Enhances the representational ability of features by introducing context information of local regions; for example, correlates the texture features of a certain region with the features of its surrounding regions to form a richer feature description.

[0152] Advanced feature extraction, based on the intermediate features, further models the global information in the iris image; extracts features that can describe the overall semantics of the iris, such as the overall texture distribution, structural relationships and category information of the iris, etc. The specific process is as follows:

[0153] Global information aggregation: Globally aggregates the intermediate features to form a feature representation that can describe the overall semantics of the iris, with high abstraction and discriminability.

[0154] Semantic relationship modeling: Based on the global information, captures the semantic relationships between different regions in the iris image to form a complete semantic description; for example, the relationship between the global distribution and local details of iris texture.

[0155] Inference result generation, after the advanced feature extraction is completed, the inference result generation stage converts the advanced features into the final inference results; generates task-related outputs (such as classification labels or bounding boxes) according to the extracted features. The specific process is as follows:

[0156] Feature mapping and transformation: Maps the advanced features to the task-related output space to generate preliminary inference results.

[0157] Result optimization and screening: Optimizes and screens the preliminary inference results to ensure they meet the task requirements; for example, removes redundant or low-confidence results through methods such as threshold screening or non-maximum suppression.

[0158] Hierarchical information transfer and feedback, the specific process is as follows:

[0159] Information transmission: The feature extraction results at each level will be used as the input for the next level to ensure the gradual transmission and accumulation of information.

[0160] Feedback mechanism: After the inference result is generated, the result is compared with the labeled data through the feedback mechanism to generate a loss signal, and the feature extraction process of each level is adjusted through backpropagation to further improve the accuracy of the inference result.

[0161] In the embodiment of the present application, the forward inference process gradually extracts discriminative features from the iris image through hierarchical information extraction and transformation, and finally generates an inference result. It has high creativity and hierarchy, can effectively improve the performance of the deep neural network in iris image tasks, and at the same time meets the definition of creativity in the patent law.

[0162] S140: Use the trained deep neural network model to perform forward inference on the iris image to be processed to obtain an inference result, and the inference result includes the elliptical parameter information of the prediction box in the iris image.

[0163] S150: Post-process the inference result, and parse out the elliptical parameters corresponding to the boundaries of the iris and the pupil in the iris image according to the elliptical parameter information of the prediction box to achieve high-precision iris localization.

[0164] Optionally, in some embodiments, first perform decoding processing on the inference result; after the decoding processing, the data is processed by non-maximum suppression to screen out the best-matched prediction box. The IOU threshold used for non-maximum suppression is 0.7, and the classification threshold is 0.8; the selected prediction box is de-normalized to obtain the elliptical parameter information of the prediction box.

[0165] In the embodiments of the present application, post-processing is performed on the inference results, specifically through steps such as decoding, non-maximum suppression (NMS), and inverse normalization, to parse out the elliptical parameters corresponding to the boundaries between the iris and the pupil in the iris image. This process has important technical effects and practical significance in the iris localization task. Through decoding processing and inverse normalization operations, the predicted bounding box information output by the deep neural network can be converted into actual elliptical parameters (such as center coordinates, major axis, minor axis, and rotation angle, etc.), so as to accurately describe the boundary between the iris and the pupil; non-maximum suppression (NMS) can effectively eliminate redundant or low-confidence prediction results by screening out the best-matching predicted bounding boxes, further improving the localization accuracy. In complex scenarios (such as light changes, blurred iris textures, or occlusions, etc.), NMS and threshold screening can filter out inaccurate predicted bounding boxes to ensure that the finally output elliptical parameters have high reliability; through the dual screening of the IOU threshold (0.7) and the classification threshold (0.8), false detection and missed detection problems can be effectively avoided, improving the robustness of the model. NMS processing can significantly reduce the number of redundant predicted bounding boxes, reduce the subsequent calculation overhead, and thus improve the overall processing efficiency; the inverse normalization operation converts the predicted bounding box parameters from the normalized space to the actual image space, avoiding unnecessary calculations and further optimizing the calculation efficiency; by parsing the elliptical parameter information of the predicted bounding box, the boundary between the iris and the pupil can be accurately fitted, providing high-quality input data for tasks such as iris recognition and identity verification. Achieved significance: High-precision iris localization is a key step in the iris recognition system. Through post-processing operations, the accuracy of iris boundary extraction can be significantly improved, thereby improving the overall performance of the iris recognition system; accurate iris localization can reduce errors in subsequent feature extraction and matching processes, improving the recognition rate and reliability; in the fields of security, finance, and healthcare, the iris recognition technology has attracted much attention due to its high security and uniqueness. High-precision iris localization technology can promote the wide application of iris recognition technology in these fields; for example, in scenarios such as mobile device unlocking, access control systems, and identity verification, high-precision iris localization can significantly improve the user experience and system security. In practical applications, iris images may be affected by factors such as light, occlusion, and blur. Through post-processing operations such as NMS and threshold screening, stable iris localization can be achieved in complex scenarios, providing reliable technical support for practical applications; the above post-processing method combines the advantages of deep learning and traditional image processing technologies, providing an innovative solution for the iris localization task. It can provide reference for other similar tasks (such as face key point detection, target tracking, etc.), promoting the further development of deep learning technology in the field of computer vision.

[0166] In summary, through post - processing operations such as decoding, non - maximum suppression, and inverse normalization on the inference results, this embodiment can achieve high - precision iris localization, improving the accuracy, robustness, and computational efficiency of the model. It can not only promote the development of iris recognition technology but also provide a reliable solution for iris localization in complex scenarios, facilitating the innovation and application of deep learning technology in the field of computer vision.

[0167] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present disclosure.

[0168] The above is the introduction of method embodiments. The following further illustrates the solution of the present disclosure through system embodiments.

[0169] Figure 9 FIG. shows a block diagram of an elliptical iris localization system 200 based on a deep convolutional neural network according to an embodiment of the present disclosure. As Figure 9 shown, the system 200 includes:

[0170] A marking module 210 for marking the training data set to obtain the ground truth. The training data set contains iris images to be trained, and the ground truth contains the elliptical parameter information of the elliptical frames in the training data set.

[0171] A construction module 220 for constructing a deep neural network model;

[0172] A training module 230 for training the deep neural network model using the training data set;

[0173] An inference module 240 for performing forward inference on the iris image to be processed using the trained deep neural network model to obtain an inference result; the inference result contains the elliptical parameter information of the predicted box in the iris image;

[0174] A post - processing module 250 for post - processing the inference result and parsing the elliptical parameter information corresponding to the iris and pupil boundaries in the iris image according to the position information of the predicted box.

[0175] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0176] Optionally, in some embodiments, the marking information of the marking module 210 for the training data set includes: the elliptical parameter information corresponding to the boundary between the iris and the pupil in the iris image, and the classification information indicating whether it is the iris or the pupil.

[0177] The construction module 220 includes: a preprocessing unit, a design unit, an encoding unit, a selection unit, and a calculation unit.

[0178] The preprocessing unit is configured to preprocess the image and the marking data.

[0179] The design unit is configured to design an elliptical box for the eye images in the training data set.

[0180] The encoding unit is configured to encode the position information of the elliptical box.

[0181] The selection unit is configured to select positive and negative samples from the training data set.

[0182] The calculation unit is configured to calculate the combined loss.

[0183] It should be noted that, in the embodiments of the present application, the calculation unit is specifically configured to: calculate the classification loss; calculate the regression loss; perform weighted averaging on the obtained classification loss and regression loss to obtain the combined loss.

[0184] In some embodiments, the construction module 220 further includes: a construction unit.

[0185] The construction unit is configured to construct a network model structure including three stages, and the three stages include: a downsampling process, a feature fusion process, and a multi-scale prediction box information output process.

[0186] In some embodiments, the training module 230 includes: an inference unit, a loss calculation unit, an adjustment unit, and a repeated execution unit.

[0187] The inference unit is configured to input the preprocessed image and marking data of the training data set into the network, and the network performs forward inference on the input image to obtain the current inference result.

[0188] The loss calculation unit is configured to calculate the classification loss and the regression loss using the current inference result and the corresponding marking data to obtain the current combined loss.

[0189] The adjustment unit is configured to backpropagate the loss into the network through reverse gradients and adjust the current network weights together with other parameters.

[0190] The repeated execution unit is configured to repeat the above process until the training ends.

[0191] It should be noted that, in the embodiments of the present application, the current inference result includes multiple adjacent result values.

[0192] In some embodiments, the post - processing module 250 includes: a decoding unit, a suppression unit, an inverse - normalization unit, and a judgment unit.

[0193] The decoding unit is used to first perform decoding processing on the inference result.

[0194] The suppression unit is used to perform non - maximum suppression processing on the data after decoding processing to screen out the best - matching prediction boxes.

[0195] The inverse - normalization unit is used to perform inverse - normalization on the screened - out prediction boxes to obtain the elliptical parameter information of the prediction boxes.

[0196] The judgment unit is used to perform iris and pupil judgment according to the set classification threshold.

[0197] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device.

[0198] Figure 10 The schematic block diagram of an electronic device 300 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0199] The electronic device 300 includes a computing unit 301, which can execute various appropriate actions and processes according to the computer program stored in the ROM 302 or the computer program loaded from the storage unit 308 into the RAM 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. The I / O interface 305 is also connected to the bus 304.

[0200] A plurality of components in the electronic device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a disk, an optical disc, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0201] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 executes the various methods and processes described above, such as the elliptical iris localization method based on a deep convolutional neural network. For example, in some embodiments, the elliptical iris localization method based on a deep convolutional neural network can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the elliptical iris localization method based on a deep convolutional neural network described above can be executed. Alternatively, in other embodiments, the computing unit 301 can be configured to execute the elliptical iris localization method based on a deep convolutional neural network by any other suitable means (e.g., by means of firmware).

[0202] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0203] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0204] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0205] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0206] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0207] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0208] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0209] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An elliptical iris localization method based on a deep convolutional neural network, characterized in that, Including: Obtain a training data set containing iris images, and label the training data set to obtain the elliptical parameter information of the elliptical frames in the training data set; Among them, the elliptical parameter information of the elliptical frames in the training data set is the ground truth; the information for labeling the training data set includes the elliptical frame parameter information corresponding to the boundary shapes of the iris and the pupil in an iris image, and the classification information of the iris and the pupil; The elliptical frame parameter information corresponds to the elliptical parameters corresponding to the shape of the iris and pupil boundaries. The elliptical parameters include , representing the corresponding elliptical center coordinates of the iris or pupil, representing the lengths of the two axes of the ellipse, where t is the clockwise angle between the a-axis of the ellipse and the x-axis; For the classification information of the iris and the pupil, 1 represents the iris and 0 represents the pupil; The truth expression of a complete eye is , where represents the elliptical parameters of the iris, and 1 represents the category of the iris; represents the elliptical parameters of the pupil, and 0 represents the category of the pupil; a complete eye should contain and only contain one valid iris information and one valid pupil information; Design the structure of the deep neural network model, data preprocessing, candidate frames, elliptical frame IOU calculation, position information encoding, positive and negative sample selection, and loss design elements, and construct a deep neural network model to train the deep neural network model using the training data set; Use the trained deep neural network model to perform forward inference on the iris image to be processed, and obtain an inference result, where the inference result includes the elliptical parameter information of the predicted frames in the iris image; Perform post-processing on the inference result, and based on the elliptical parameter information of the predicted frames, parse out the elliptical parameters corresponding to the boundaries of the iris and the pupil in the iris image to achieve high-precision iris localization.

2. The elliptical iris localization method based on a deep convolutional neural network according to claim 1, wherein The structure of the deep neural network model includes: The first stage: includes multiple consecutive downsampling processes, and its function is to extract features during the fast image degradation process; The second stage: includes multiple consecutive feature fusion processes, and its function is to fuse the initially extracted features; The third stage: includes a process of outputting prediction frame information at multiple scales, and its function is to output prediction results, including the position information of the prediction frames and the classification information of the prediction frames.

3. The elliptical iris localization method based on a deep convolutional neural network according to claim 2, wherein The first stage contains: Use a specially designed Quad-Conv structure.

4. The elliptical iris localization method based on a deep convolutional neural network according to claim 2, characterized in that The second stage contains: feature fusion. The feature fusion of the iris image needs to balance local details and global structure, specifically including: Use multiple InceptionBlocks to achieve feature fusion and receptive field expansion.

5. The elliptical iris localization method based on a deep convolutional neural network according to claim 2, wherein The third stage multi-scale output and elliptical parameter prediction. The prediction frames of the iris image need to be represented in the form of an ellipse, and the output layer is specially designed; including: An elliptical parameter regression head, used to design an elliptical parameter regression head to directly predict the five parameters of the ellipse; A multi-scale prediction module, used to combine feature maps of different scales.

6. The elliptical iris localization method based on a deep convolutional neural network according to claim 1, wherein, The preprocessing of the iris image data includes: The iris image data undergoes various data augmentation processes, among which operations such as deformation, cropping, horizontal flipping, and rotation affect the position information of the labeled data, and at the same time adjust the corresponding labeled position information; The image after data augmentation contains at least one complete labeled eye image; Finally, the input image is scaled to a single-channel grayscale image of a specified size, and after being normalized, it is input into the deep neural network model to be trained; The preprocessing of the labeled data includes: adjusting the corresponding position information; After the position information is finally normalized, it is used as the new Ground Truth information and input into the network model to be trained; The marking method of the candidate ellipse box is , where is the center coordinate of the candidate ellipse box after normalization, a and b are the axis lengths of the ellipse box after normalization, and t is the clockwise angle between the a-axis of the ellipse box and the x-axis.

7. The elliptical iris localization method based on a deep convolutional neural network according to claim 1, characterized in that The training of the deep neural network model using the training data set includes: Split the labeled iris images and data into two sets: a training data set and a validation data set; Train using the training dataset and validate and screen parameters of the deep neural network model using the validation dataset; Input the preprocessed iris images and labeled data of the training dataset into the deep neural network. The deep neural network performs forward inference on the input iris images to obtain the current inference result; Use the current inference result and the corresponding labeled data to calculate the classification loss and the regression loss, and obtain the current combined loss; The current loss is propagated back to the deep neural network through reverse gradient and, together with the learning rate parameter, adjusts the weights of the current deep neural network; Repeat until the training ends; The condition for the end of training is: finishing the specified number of rounds; During the process of the deep neural network performing forward inference on the input iris images, the core lies in gradually extracting discriminative features from the iris images through multi-level feature extraction and information transformation, and finally generating the inference result; First, perform decoding processing on the inference result; For the data after decoding processing, perform non-maximum suppression processing to screen out the best-matching prediction boxes; Perform inverse normalization on the screened prediction boxes to obtain the elliptical parameter information of the prediction boxes.

8. An elliptical iris localization system based on a deep convolutional neural network, characterized in that Including: A labeling module for labeling the training dataset to obtain the ground truth. The training dataset contains the iris images to be trained, and the ground truth contains the elliptical parameter information of the elliptical boxes in the training dataset; Among them, the elliptical parameter information of the elliptical boxes in the training dataset is the ground truth; the information for labeling the training dataset includes the elliptical box parameter information corresponding to the boundary shapes of the iris and the pupil in an iris image, and the classification information of the iris and the pupil; The elliptical frame parameter information corresponds to the elliptical parameters corresponding to the boundary shape of the iris and the pupil. The elliptical parameters include , representing the corresponding elliptical center coordinates of the iris or the pupil, representing the lengths of the two axes of the ellipse, and t is the clockwise angle between the a-axis of the ellipse and the x-axis; For the classification information of the iris and the pupil, 1 represents the iris and 0 represents the pupil; The truth expression of a complete eye is , where represents the elliptical parameters of the iris, and 1 represents the category of the iris; represents the elliptical parameters of the pupil, and 0 represents the category of the pupil; a complete eye should contain and only contain one valid iris information and one valid pupil information; A construction module for constructing a deep neural network model; A training module for training the deep neural network model using the training dataset; An inference module for using the trained deep neural network model to perform forward inference on the iris images to be processed and obtain the inference result; the inference result contains the elliptical parameter information of the prediction boxes in the iris images; A post-processing module for post-processing the inference result and parsing the elliptical parameter information corresponding to the boundary between the iris and the pupil in the iris image according to the position information of the prediction boxes.

9. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Iris image segmentation and positioning method, system and device based on deep learning

    CN109815850A

  • Segmentation method of iris region in iris image based on Mask R-CNN neural network

    CN110059589A