Iris real-time tracking positioning method, device and equipment and storage medium

By combining multi-level convolution and mask prediction models with iris concentricity constraint relationships, the problem of insufficient iris boundary positioning accuracy in complex environments is solved, and high-precision and robust iris positioning is achieved.

CN121259900BActive Publication Date: 2026-07-03SHENZHEN INTERFACE COGNITIVE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INTERFACE COGNITIVE TECH CO LTD
Filing Date
2025-10-28
Publication Date
2026-07-03

Smart Images

  • Figure CN121259900B_ABST
    Figure CN121259900B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, device, and storage medium for real-time iris tracking and localization. The method includes: performing multi-level convolution processing on a real-time input iris image to extract iris features at different scales, obtaining a feature map set; inputting the feature map set into a preset mask prediction model for mask prediction, obtaining a pupil boundary mask and an outer iris boundary mask; extracting the boundary contours of the pupil boundary mask and the outer iris boundary mask, obtaining a pupil boundary contour point set and an outer iris boundary contour point set; performing geometric fitting based on the pupil boundary contour point set and the outer iris boundary contour point set, and optimizing and correcting the fitting parameters according to the concentric constraint relationship between the iris and the pupil to obtain the iris localization result. This invention, by combining boundary enhancement feature fusion with geometric constraint optimization, can accurately locate the inner and outer boundaries of the iris in non-cooperative environments, improving the accuracy and robustness of iris localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, and storage medium for real-time iris tracking and positioning. Background Technology

[0002] Iris recognition, as a high-precision biometric identification technology, has wide applications in fields such as identity authentication and security monitoring. Iris localization, as a crucial preprocessing step in an iris recognition system, directly impacts the performance of subsequent feature extraction and recognition. Traditional iris localization methods primarily rely on geometric fitting algorithms such as the Daugman integral-differential operator or the Hough transform. These methods achieve good localization results under ideal acquisition conditions, but require manual feature design and setting of numerous parameters, and also demand high image quality.

[0003] However, iris images acquired in non-cooperative environments often face complex interference factors such as eyelid and eyelash occlusion, eyeglass reflections, uneven lighting, defocus blur, and angular deviation, leading to insufficient robustness and a significant decrease in localization accuracy of traditional methods. In recent years, deep learning-based iris localization methods have demonstrated better generalization ability in complex scenes by automatically learning features through neural networks. However, existing deep learning methods often directly output localization results using end-to-end black-box models, lacking targeted enhancement of boundary regions and explicit modeling of geometric constraints, making them prone to localization errors in cases of blurred boundaries or severe occlusion. Summary of the Invention

[0004] The main objective of this invention is to solve the technical problems of insufficient boundary positioning accuracy and lack of geometric constraint optimization mechanism in existing iris localization methods in complex non-cooperative environments.

[0005] This invention provides a real-time iris tracking and positioning method, the real-time iris tracking and positioning method comprising:

[0006] Multi-level convolution processing is performed on the real-time input iris image to extract iris features at different scales and obtain a set of feature maps;

[0007] The feature maps in the feature map set are input into a preset mask prediction model to perform mask prediction, and the pupil boundary mask and the outer iris boundary mask are obtained.

[0008] The boundary contours of the pupil boundary mask and the outer iris boundary mask are extracted to obtain the pupil boundary contour point set and the outer iris boundary contour point set;

[0009] Geometric fitting is performed based on the pupil boundary contour point set and the iris outer boundary contour point set, and the fitting parameters are optimized and corrected according to the concentric constraint relationship between the iris and the pupil to obtain the iris localization result.

[0010] The present invention also provides an iris real-time tracking and positioning device, the iris real-time tracking and positioning device comprising:

[0011] The feature extraction module is used to perform multi-level convolution processing on the real-time input iris image to extract iris features at different scales and obtain a set of feature maps.

[0012] The mask prediction module is used to input the feature maps in the feature map set into a preset mask prediction model to perform mask prediction and obtain the pupil boundary mask and the outer boundary mask of the iris.

[0013] The contour extraction module is used to extract the boundary contours of the pupil boundary mask and the outer iris boundary mask to obtain the pupil boundary contour point set and the outer iris boundary contour point set.

[0014] The parameter optimization module is used to perform geometric fitting based on the pupil boundary contour point set and the iris outer boundary contour point set, and to optimize and correct the fitting parameters based on the concentric constraint relationship between the iris and the pupil to obtain the iris localization result.

[0015] The present invention also provides an iris real-time tracking and positioning device, comprising: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a line; the at least one processor invokes the instructions in the memory to cause the iris real-time tracking and positioning device to perform the steps of the above-described iris real-time tracking and positioning method.

[0016] The present invention also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the steps of the above-described real-time iris tracking and positioning method.

[0017] The aforementioned real-time iris tracking and localization method, device, equipment, and storage medium extract iris features at different scales by performing multi-level convolution processing on real-time input iris images, obtaining a feature map set. The feature map set is then input into a preset mask prediction model for mask prediction, yielding a pupil boundary mask and an outer iris boundary mask. Boundary contours are extracted from the pupil boundary mask and the outer iris boundary mask, resulting in a pupil boundary contour point set and an outer iris boundary contour point set. Geometric fitting is performed based on the pupil boundary contour point set and the outer iris boundary contour point set, and the fitting parameters are optimized and corrected according to the concentric constraint relationship between the iris and the pupil to obtain the iris localization result. This invention, through a combination of boundary enhancement feature fusion and geometric constraint optimization, can accurately locate the inner and outer boundaries of the iris in non-cooperative environments, improving the accuracy and robustness of iris localization.

[0018] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the first embodiment of the iris real-time tracking and positioning method in this invention;

[0021] Figure 2 This is a schematic diagram of a second embodiment of the iris real-time tracking and positioning method in this invention;

[0022] Figure 3 This is a schematic diagram of one embodiment of the iris real-time tracking and positioning device in this invention;

[0023] Figure 4 This is a schematic diagram of one embodiment of the iris real-time tracking and positioning device in this invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0026] To facilitate understanding of this embodiment, a real-time iris tracking and positioning method disclosed in this invention will first be described in detail. For example... Figure 1 As shown, this method includes the following steps:

[0027] 101. Perform multi-level convolution processing on the real-time input iris image to extract iris features at different scales and obtain a set of feature maps;

[0028] In this embodiment, the step of performing multi-level convolution processing on the real-time input iris image to extract iris features at different scales and obtain a feature map set includes: dividing the iris image into blocks and performing initial feature mapping through convolutional layers to obtain an initial feature map; using the initial feature map as the first layer input feature map, performing N layers of feature extraction sequentially, wherein the i-th layer feature extraction involves spatial downsampling and channel expansion of the i-th layer input feature map to obtain an i-th layer downsampled feature map; performing depthwise separable convolution processing on the i-th layer downsampled feature map to obtain an i-th layer intermediate feature map; adding the residuals of the i-th layer intermediate feature map and the i-th layer downsampled feature map to obtain an i-th layer output feature map, and using the i-th layer output feature map as the input feature map of the (i+1)-th layer; and collecting the initial feature map and the feature maps of each layer to obtain the feature map set.

[0029] Specifically, after receiving the iris image transmitted by the acquisition module, the iris recognition device can perform multi-level convolution processing on the iris image to extract iris features at different scales. It should be noted that the iris image can be a grayscale image acquired by a near-infrared camera, typically with a resolution of 1280×720 or 1920×1080 pixels. In non-cooperative environments, the iris image may contain complex interference factors such as eyelid occlusion, eyelash interference, eyeglass reflections, and uneven lighting; therefore, multi-level feature extraction is necessary to capture multi-scale information from local texture to global structure.

[0030] In one specific embodiment, the iris recognition device first performs block processing on the input iris image. Specifically, the iris image can be divided into multiple image blocks, each of which can be set to a size of 3×3 or 4×4 pixels. Then, the iris recognition device performs feature mapping on these image blocks through an initial convolutional layer. This initial convolutional layer uses a 3×3 convolutional kernel to expand the number of channels in the input image from a single channel to a preset initial number of channels, such as 64 or 128 channels, resulting in an initial feature map. This initial feature map retains a spatial resolution similar to the input image and contains shallow feature information of the iris image, such as edges and textures.

[0031] Based on this, the iris recognition device uses the initial feature map as the input feature map for the first layer and begins a multi-level feature extraction process. In this embodiment, the number of feature extraction layers N can be set to 4 or 5 layers. The feature extraction process for each layer includes operations such as spatial downsampling, channel expansion, depthwise separable convolution processing, and residual connection.

[0032] For feature extraction at layer i, the iris recognition device first performs spatial downsampling and channel expansion operations on the input feature map of layer i. Specifically, spatial downsampling is achieved through a convolution operation with a stride of 2, reducing the height and width of the feature map to half of their original values ​​while doubling the number of channels. For example, if the size of the input feature map of layer i is H×W×C, then after downsampling and channel expansion, the size of the downsampled feature map of layer i is H / 2×W / 2×2C. Through this downsampling operation, the receptive field of the features can be expanded layer by layer, enabling deeper features to capture information from a wider range of iris regions.

[0033] Furthermore, the iris recognition device performs depthwise separable convolution processing on the downsampled feature map of the i-th layer. It should be noted that depthwise separable convolution consists of two parts: depthwise convolution and pointwise convolution. Depthwise convolution performs convolution operations independently on each input channel, using a 3×3 kernel to extract the spatial features of each channel; pointwise convolution uses a 1×1 kernel to linearly combine the output of the depthwise convolution between channels. Compared to standard convolution, depthwise separable convolution can significantly reduce the computational load and parameter count while maintaining similar feature extraction capabilities, which is of great significance for achieving real-time inference on embedded SoC platforms. Through depthwise separable convolution processing, the iris recognition device obtains the intermediate feature map of the i-th layer.

[0034] Understandably, in deep neural networks, the vanishing gradient problem tends to occur as the number of network layers increases. To address this issue, this embodiment employs a residual connection mechanism. Specifically, the iris recognition device adds the intermediate feature map of layer i to the downsampled feature map of layer i element-wise to obtain the output feature map of layer i. This residual connection method establishes a direct path from input to output, allowing the gradient to propagate directly forward, effectively mitigating the vanishing gradient problem, while preserving shallow feature information, enabling deep features to simultaneously contain high-level semantic information and low-level detailed information.

[0035] After completing feature extraction at layer i, the iris recognition device uses the output feature map of layer i as the input feature map for layer (i+1) to continue feature extraction at the next layer. Through this progressive layer-by-layer approach, the iris recognition device sequentially completes feature extraction at N layers, constructing a feature pyramid structure from shallow to deep, from local to global. In this feature pyramid, shallow feature maps have high spatial resolution and fewer channels, containing rich local texture details; deep feature maps have lower spatial resolution and more channels, containing more abstract semantic information and a wider range of contextual information. Finally, the iris recognition device collects this initial feature map and the output feature maps from each layer to form a multi-scale feature map set.

[0036] 102. Input the feature maps in the feature map set into a preset mask prediction model to perform mask prediction, and obtain the pupil boundary mask and the outer boundary mask of the iris;

[0037] In this embodiment, after acquiring a multi-scale feature map set, the iris recognition device can input this feature map set into a preset mask prediction model for mask prediction to obtain the pupil boundary mask and the outer iris boundary mask. It should be noted that the mask prediction model can adopt an encoder-decoder structure. The encoder part corresponds to the aforementioned multi-level feature extraction process, while the decoder part is responsible for progressively restoring the extracted feature maps to the original image resolution and generating the mask prediction result.

[0038] Specifically, in a common implementation, iris recognition devices can employ a mask prediction model based on the U-Net architecture. This model upsamples and fuses the multi-scale feature maps extracted by the encoder through a decoder, ultimately outputting a pixel-level segmentation mask. Furthermore, the iris recognition device first upsamples the deepest feature map in the feature map set, which can double the spatial resolution of the feature map through bilinear interpolation or transposed convolution. For example, if the size of the deepest feature map is H / 16×W / 16×C, after upsampling, a feature map with a size of H / 8×W / 8×C / 2 can be obtained.

[0039] Building upon this, the iris recognition device concatenates and fuses the upsampled feature map with shallow feature maps of the corresponding scale from the feature map set. Specifically, the feature map of the current layer of the decoder can be concatenated with feature maps of the same spatial size in the encoder along the channel dimension to form a fused feature map. This cross-layer connection method combines deep semantic information with shallow detail information, enabling the decoder to preserve boundary details while restoring spatial resolution. Furthermore, the iris recognition device processes the fused feature map using convolutional layers. A 3×3 convolutional kernel can be used to perform a nonlinear transformation on the fused features to obtain the output feature map of the current decoding layer.

[0040] The iris recognition device repeats the upsampling, feature fusion, and convolution processes described above, restoring the spatial resolution of the feature map to the same size as the input image layer by layer. In the final layer of the decoder, the iris recognition device generates a pupil boundary mask and an outer iris boundary mask through two independent 1×1 convolutional layers. Specifically, each 1×1 convolutional layer compresses the number of channels in the feature map to 2, representing the probabilities of the background and target categories, respectively. These are then normalized using the Softmax activation function to obtain a pixel-level classification probability map. The iris recognition device extracts the target category channel from the probability map to obtain the pupil boundary mask and the outer iris boundary mask. The value of each pixel in the mask represents the probability that the pixel belongs to either the pupil boundary or the outer iris boundary.

[0041] 103. Extract the boundary contours of the pupil boundary mask and the outer iris boundary mask to obtain the pupil boundary contour point set and the outer iris boundary contour point set;

[0042] In this embodiment, the step of extracting the boundary contours of the pupil boundary mask and the outer iris boundary mask to obtain the pupil boundary contour point set and the outer iris boundary contour point set includes: performing binarization processing on the pupil boundary mask and the outer iris boundary mask respectively to obtain a pupil binary mask and an outer iris boundary binary mask; performing morphological processing and connected component filtering on the pupil binary mask and the outer iris boundary binary mask respectively to obtain an effective pupil region and an effective outer iris boundary region; and performing boundary extraction on the effective pupil region and the effective outer iris boundary region respectively to obtain the pupil boundary contour point set and the outer iris boundary contour point set.

[0043] Specifically, after acquiring the pupil boundary mask and the outer iris boundary mask, the iris recognition device needs to extract a precise set of boundary contour points from these probabilistic masks for subsequent geometric parameter fitting. It should be noted that the mask output by the mask prediction model is typically a probability map, where the value of each pixel represents the probability that the pixel belongs to the target region, ranging from 0 to 1. To determine a clear boundary region from the probability map, the iris recognition device first needs to binarize the mask.

[0044] In one specific embodiment, the iris recognition device performs threshold segmentation processing on the pupil boundary mask and the outer iris boundary mask respectively. Specifically, the iris recognition device can set a preset threshold, such as 0.5, marking pixels with pixel values ​​greater than the threshold as foreground pixels and assigning them a value of 1, and marking pixels with pixel values ​​less than or equal to the threshold as background pixels and assigning them a value of 0. Through this binarization processing, the iris recognition device can obtain a binary mask of the pupil and a binary mask of the outer iris boundary, which clearly distinguish the target area and the background area. It is understood that the choice of threshold will affect the binarization result; an excessively high threshold may lead to a shrinkage of the target area, while an excessively low threshold may introduce more noise. Therefore, an appropriate threshold can be selected according to the actual application scenario and data characteristics.

[0045] Furthermore, the iris recognition device performs morphological processing on the binary mask of the pupil and the binary mask of the outer boundary of the iris to remove noise and smooth the boundaries. It should be noted that morphological processing is an image processing technique based on mathematical morphology, mainly including basic operations such as erosion, dilation, opening, and closing. In this embodiment, the iris recognition device can first perform a closing operation on the binary mask, followed by an opening operation. Specifically, the closing operation consists of a dilation operation followed by an erosion operation. The dilation operation uses structuring elements to expand the binary image, filling small holes and broken boundaries within the target region; the erosion operation uses structuring elements to shrink the dilated image, restoring the approximate size of the target region. Through the closing operation, the iris recognition device can effectively fill the holes inside the mask, making the target region more continuous and complete.

[0046] Based on this, the iris recognition device performs an opening operation on the result of the closing operation. The opening operation consists of an erosion operation followed by a dilation operation, which can remove small noise points and edge burrs in the image. Specifically, the erosion operation first removes small noise areas and protrusions, and then the dilation operation restores the size of the main area. Through the opening operation, the iris recognition device can make the mask boundaries smoother and remove small interference areas introduced by inaccurate prediction or image noise. It should be noted that the structuring element used in morphological operations can be circular, rectangular, or cross-shaped, etc. The size of the structuring element affects the processing effect; it is usually set to a 3×3 or 5×5 kernel.

[0047] After morphological processing, the iris recognition device further filters the processed binary masks for the pupil and iris outer boundary. Specifically, the device labels the binary masks with connected components, grouping interconnected foreground pixels in the image into the same connected component, each representing an independent region. In non-cooperative environments, due to interference factors such as occlusion and reflection, multiple connected components may exist in the mask. Only the largest connected component corresponds to the true pupil or iris outer boundary region; other smaller connected components are usually noise or false detections. The device then calculates the area of ​​each connected component, determining its size by counting the number of pixels it contains. The device selects the largest connected component as the effective region, obtaining the effective pupil region and the effective iris outer boundary region. This connected component filtering method effectively eliminates scattered noise regions and retains the main iris boundary regions.

[0048] Finally, the iris recognition device extracts the boundaries of the effective pupil region and the effective outer boundary region of the iris. It should be noted that boundary extraction refers to extracting the coordinate information of boundary pixels from a binary region. In this embodiment, the iris recognition device can employ a boundary tracing algorithm to extract the contour. Specifically, a chain code-based contour tracing algorithm can be used. This algorithm starts from a boundary point of the region, checks neighboring pixels in a predetermined direction, finds the next boundary point, and repeats this process until it returns to the starting point, thus forming a closed boundary contour. Through boundary tracing, the iris recognition device can sequentially extract the coordinates of all pixels on the outer contour of the effective region, obtaining the pupil boundary contour point set and the iris outer boundary contour point set. This contour point set is an ordered set of two-dimensional coordinate points, such as {(x1,y1),(x2,y2),...,(xM,yM)}, where M is the number of pixels on the boundary. These points are arranged according to the continuity of the boundary, completely describing the geometry of the boundary.

[0049] 104. Perform geometric fitting based on the pupil boundary contour point set and the iris outer boundary contour point set, and optimize and correct the fitting parameters based on the concentric constraint relationship between the iris and the pupil to obtain the iris localization result.

[0050] In this embodiment, the step of performing geometric fitting based on the pupil boundary contour point set and the iris outer boundary contour point set, and optimizing and correcting the fitting parameters according to the concentric constraint relationship between the iris and the pupil to obtain the iris localization result includes: performing geometric fitting on the pupil boundary contour point set and the iris outer boundary contour point set respectively to obtain the initial geometric parameters of the pupil and the initial geometric parameters of the iris outer boundary; performing a weighted average of the center coordinates of the initial geometric parameters of the pupil and the center coordinates of the initial geometric parameters of the iris outer boundary to obtain shared center coordinates; calculating the radius of the pupil boundary contour point set and the iris outer boundary contour point set respectively based on the shared center coordinates to obtain the pupil radius and the iris outer boundary radius; and combining the shared center coordinates, the pupil radius, and the iris outer boundary radius to form the iris localization result.

[0051] Specifically, after acquiring the pupil boundary contour point set and the iris outer boundary contour point set, the iris recognition device can perform geometric fitting on these discrete contour points to obtain the geometric parameters of the inner and outer boundaries of the iris. It should be noted that the iris and pupil typically appear as approximately circular or elliptical structures in images. Geometric fitting can transform discrete boundary points into a continuous geometric model, thereby obtaining parameters such as the center coordinates and radius of the boundary. In this embodiment, the iris recognition device uses a circular fitting method to fit the boundary points into the form of a circle.

[0052] In one specific embodiment, the iris recognition device performs independent geometric fitting on the pupil boundary contour point set and the iris outer boundary contour point set. Specifically, the iris recognition device can use a least-squares circle fitting algorithm to solve for the optimal circle parameters. The least-squares method is a classic parameter estimation method, the basic idea of ​​which is to minimize the sum of squared distances from the fitted circle to all contour points. For the pupil boundary contour point set, the iris recognition device needs to solve for the center coordinates and radius, minimizing the sum of squared deviations between the distances from all points to the center and the radius. Specifically, this can be achieved by constructing an objective function and taking the partial derivatives with respect to the center coordinates and radius, setting the partial derivatives to zero, obtaining a system of linear equations, and solving this system of equations to obtain the initial geometric parameters of the pupil. Similarly, the iris recognition device performs least-squares circle fitting on the iris outer boundary contour point set to obtain the initial geometric parameters of the iris outer boundary, including the center coordinates and radius. Through this independent fitting method, the preliminary geometric parameters of the pupil and the outer boundary of the iris can be obtained separately.

[0053] Understandably, ideally, the pupil and iris are concentric, meaning the center of the pupil and the center of the iris should coincide or be very close. However, in practical applications, due to factors such as occlusion, noise, and shooting angle, the two centers obtained by independent fitting often have a certain offset, which affects the consistency and accuracy of positioning. To solve this problem, this embodiment utilizes the concentric constraint relationship between the iris and pupil to optimize and correct the initial fitting parameters.

[0054] Next, the iris recognition device calculates the offset distance between the center coordinates of the initial geometric parameters of the pupil and the center coordinates of the initial geometric parameters of the outer iris boundary. This offset distance can be obtained by calculating the Euclidean distance between the two center coordinates. The iris recognition device evaluates the concentricity of the two centers based on this offset distance; the smaller the offset distance, the better the concentricity. Based on this, the iris recognition device performs a weighted average of the center coordinates of the initial geometric parameters of the pupil and the center coordinates of the initial geometric parameters of the outer iris boundary to obtain shared center coordinates. Specifically, the weights can be set according to the number of contour point sets or the reliability of the fit. Generally, the outer iris boundary has a larger number of contour points and a clearer boundary, so a larger weight can be assigned to the center of the outer iris boundary, for example, the weight ratio can be set to 30% to 70% or 40% to 60%. Through weighted averaging, the iris recognition device obtains the shared center coordinates, which integrates the positional information of the pupil and the outer iris boundary, satisfying the concentricity constraint.

[0055] It should be noted that after determining the shared center, the radius parameters obtained from the previously independently fitted points are no longer applicable, as these radii were calculated based on their respective center points. Therefore, the iris recognition device needs to recalculate the radii of the pupil and the outer boundary of the iris based on the coordinates of the shared center. Specifically, for the pupil boundary contour point set, the iris recognition device calculates the distance from each contour point to the shared center, obtaining multiple distance values. Then, the iris recognition device performs statistical analysis on these distance values, determining the pupil radius using methods such as the median, average, or weighted average. In this embodiment, to reduce the influence of outliers, the median can be used as the pupil radius; that is, the median value is selected from all distance values ​​as the optimized pupil radius. Similarly, the iris recognition device calculates the distance from each point in the iris outer boundary contour point set to the shared center, obtaining the optimized outer boundary radius of the iris through statistical analysis. Through this radius recalculation method based on the shared center, the iris recognition device obtains optimized radius parameters that satisfy the concentricity constraint.

[0056] As can be seen, compared with traditional joint optimization methods, this embodiment avoids complex nonlinear optimization processes by using center-weighted averaging and radius recalculation based on shared centers. This results in simple and efficient computation, making it particularly suitable for real-time applications. Furthermore, this method fully utilizes prior anatomical knowledge of the iris and pupil, i.e., concentricity constraints, significantly improving the geometric consistency of the localization results.

[0057] Ultimately, the iris recognition device generates an iris localization result based on a combination of shared center coordinates, pupil radius, and iris outer boundary radius. Specifically, this iris localization result includes the geometric parameters of the inner pupil boundary and the outer iris boundary. The inner pupil boundary's geometric parameters are the shared center coordinates and the optimized pupil radius, while the outer iris boundary's geometric parameters are the shared center coordinates and the optimized outer iris boundary radius. The iris recognition device can output these geometric parameters in a structured format, such as JSON, containing the center coordinates and radius information of both the pupil and the outer iris boundary. This iris localization result can be directly used in subsequent steps such as iris normalization, feature extraction, and recognition matching.

[0058] In this embodiment, multi-level convolution processing is performed on the real-time input iris image to extract iris features at different scales, resulting in a feature map set. This feature map set is then input into a preset mask prediction model for mask prediction, yielding a pupil boundary mask and an outer iris boundary mask. Boundary contours are extracted from the pupil boundary mask and the outer iris boundary mask, resulting in a pupil boundary contour point set and an outer iris boundary contour point set, respectively. Geometric fitting is performed based on these two sets, and the fitting parameters are optimized and corrected according to the concentric constraint relationship between the iris and the pupil, resulting in the iris localization result. This invention, through a combination of boundary enhancement feature fusion and geometric constraint optimization, can accurately locate the inner and outer boundaries of the iris in non-cooperative environments, improving the accuracy and robustness of iris localization.

[0059] Please see Figure 2 The second embodiment of the iris real-time tracking and positioning method in this application includes:

[0060] 201. Perform multi-level convolution processing on the real-time input iris image to extract iris features at different scales and obtain a set of feature maps;

[0061] In this embodiment, step 201 is similar to step 101 in the first embodiment, and will not be described again here.

[0062] 202. Input the feature maps from the feature map set into the preset mask prediction model;

[0063] 203. The feature maps in the feature map set are fused across scales using the mask prediction model, and the spatial difference information of the feature maps is calculated. The fused feature maps are then weighted according to the spatial difference information to obtain the boundary enhancement feature map.

[0064] In this embodiment, the step of performing cross-scale fusion of feature maps in the feature map set using the mask prediction model, calculating the spatial difference information of the feature maps, and weighting the fused feature maps according to the spatial difference information to obtain a boundary enhancement feature map includes: upsampling high-level feature maps in the multi-scale feature map set using the mask prediction model, connecting the upsampled high-level feature maps with low-level feature maps of the corresponding scale across scales to obtain a connected feature map; performing dual-path parallel processing on the connected feature map to obtain a local enhancement feature map and a global perception feature map, and concatenating them in the channel dimension to obtain a fused feature map; calculating the average pooling result in the spatial dimension of the fused feature map, performing a difference operation between the fused feature map and the average pooling result to obtain spatial difference information; generating a boundary weight map by convolution and normalization processing on the spatial difference information, and weighting the fused feature map pixel by pixel according to the boundary weight map to obtain the boundary enhancement feature map.

[0065] Specifically, after acquiring a multi-scale feature map set, the iris recognition device can process the feature map set using a preset mask prediction model to generate boundary enhancement feature maps. It should be noted that iris images acquired in non-cooperative environments often face problems such as blurred iris boundaries and severe occlusion, making accurate boundary location difficult to pinpoint solely through feature extraction. Therefore, this embodiment utilizes feature fusion and boundary enhancement mechanisms to specifically enhance features in the inner and outer boundary regions of the iris, thereby improving the accuracy of boundary localization.

[0066] In one specific embodiment, the iris recognition device first performs cross-scale fusion of feature maps from a multi-scale feature map set using a mask prediction model. Specifically, this feature map set contains feature maps at different levels; deep feature maps contain rich semantic information but have lower spatial resolution, while shallow feature maps contain rich detail information but have lower semantic abstraction. To utilize both deep semantics and shallow details simultaneously, the iris recognition device performs upsampling on the high-level feature maps. Upsampling is an operation that increases the spatial resolution of a feature map, which can be achieved through bilinear interpolation or transposed convolution. For example, for the deepest feature map with size H divided by 16 multiplied by W divided by 16, upsampling can yield a feature map with size H divided by 8 multiplied by W divided by 8, effectively doubling its spatial resolution.

[0067] Furthermore, the iris recognition device performs a cross-scale connection between the upsampled high-level feature map and the corresponding low-level feature map. Specifically, this connection operation can be a channel-level concatenation, superimposing the two feature maps in the channel direction to form a thicker feature map. For example, if the upsampled high-level feature map has C1 channels and the corresponding low-level feature map has C2 channels, then the concatenated connected feature map will have C1 plus C2 channels. Through this cross-scale connection method, the iris recognition device can fuse high-level semantic information with low-level detail information to obtain a connected feature map. This connected feature map simultaneously contains global semantic understanding and local boundary details, providing a rich information foundation for subsequent boundary enhancement.

[0068] Building upon this, the iris recognition device performs dual-path parallel processing on the connected feature map. It's important to note that dual-path parallel processing means simultaneously enhancing the feature map from both local and global perspectives. The first path focuses on enhancing local texture, while the second path focuses on modeling global spatial relationships. Specifically, in the first path, the iris recognition device enhances the local texture of the connected feature map using gated convolution. Gated convolution is a selective feature enhancement mechanism that controls the importance of different feature channels by generating gate weights. The iris recognition device uses one set of convolutional layers to generate the gate signal, and another set of convolutional layers to generate the input signal. The gate signal is then normalized to weights between zero and one using a sigmoid activation function. Finally, these weights are multiplied element-wise with the input signal to obtain the locally enhanced feature map. Through this gating mechanism, important local texture features are preserved and enhanced, while less important features are suppressed.

[0069] Simultaneously, in the second path, the iris recognition device calculates the spatial position weights of the connectivity feature map using a spatial attention mechanism. Spatial attention is a mechanism for modeling the importance of spatial positions; its core idea is to evaluate the contribution of each spatial position in the feature map to the task. Specifically, the iris recognition device first performs max pooling and average pooling operations on the connectivity feature map. Max pooling extracts the most salient feature response at each position, while average pooling provides the average contextual information for that position. Then, the iris recognition device concatenates the max pooling and average pooling results along the channel dimension, compresses the concatenated features into a single-channel feature map through a convolutional layer, and normalizes it to a spatial position weight map between zero and one using a sigmoid activation function. The iris recognition device then multiplies this spatial position weight map element-wise with the connectivity feature map to obtain a globally perceptual feature map. Through this spatial attention mechanism, important positions such as the iris boundary receive higher weights, while irrelevant regions such as the background receive lower weights.

[0070] Understandably, the local enhanced feature map and the global perceived feature map enhance the connected feature map from different perspectives, and they contain complementary information. Furthermore, the iris recognition device concatenates the local enhanced feature map and the global perceived feature map along the channel dimension to obtain a fused feature map. This fused feature map simultaneously contains local texture details and global spatial relationships, providing rich feature representations for accurate boundary localization.

[0071] After obtaining the fused feature map, the iris recognition device further highlights boundary information through a spatial difference mechanism. It's important to note that the iris boundary is the transition region between the iris and non-iris regions, which manifests as abrupt changes in feature values ​​on the feature map. To extract this abrupt change information, the iris recognition device calculates the average pooling result across the spatial dimension of the fused feature map. Average pooling averages the value of each channel across the entire spatial dimension, obtaining the global average response for each channel. Then, the iris recognition device performs a difference operation between the fused feature map and its average pooling result, subtracting the average value of the corresponding channel from the pixel value of the fused feature map to obtain spatial difference information. This spatial difference information reflects the degree of deviation of each location from the global average; boundary locations, due to drastic feature changes, have larger difference values, while flat areas have smaller difference values. Through this spatial difference mechanism, the iris recognition device effectively highlights the boundary transition region.

[0072] Finally, the iris recognition device further processes the spatial difference information to generate a boundary weight map. Specifically, the device performs feature transformation on the spatial difference information through convolutional layers, then normalizes it using batch normalization layers to stabilize the feature distribution, and finally maps the feature values ​​to between zero and one using a sigmoid activation function, generating the boundary weight map. The value at each position in this boundary weight map represents the probability that the position belongs to the boundary; the weight for boundary regions is close to one, and the weight for non-boundary regions is close to zero. Furthermore, the iris recognition device performs pixel-by-pixel weighting on the fused feature map based on this boundary weight map, that is, multiplying the boundary weight map and the fused feature map element-by-element to obtain a boundary enhancement feature map. In the boundary enhancement feature map, the features of boundary regions are significantly enhanced, while the features of non-boundary regions are relatively suppressed, thereby improving the sensitivity and accuracy of subsequent mask predictions to boundaries.

[0073] Furthermore, the dual-path parallel processing of the connection feature map to obtain the local enhancement feature map and the global awareness feature map includes: in the first path, generating a gated signal feature map from the connection feature map through a first set of convolutional layers, generating an input signal feature map from the connection feature map through a second set of convolutional layers, normalizing the gated signal feature map using a Sigmoid activation function to obtain a gated weight map, and multiplying the gated weight map element-wise with the input signal feature map to obtain the local enhancement feature map; in the second path, performing max pooling and average pooling on the connection feature map in the spatial dimension to obtain max pooling feature maps and average pooling feature maps, concatenating them in the channel dimension to obtain a pooled concatenated feature map, compressing the pooled concatenated feature map through convolutional layers to obtain a single-channel feature map, normalizing the single-channel feature map using a Sigmoid activation function to obtain a spatial position weight map, and multiplying the spatial position weight map element-wise with the connection feature map to obtain the global awareness feature map.

[0074] Specifically, further, the dual-path parallel processing of the connection feature map to obtain the local enhancement feature map and the global awareness feature map includes: in the first path, generating a gated signal feature map from the connection feature map through a first set of convolutional layers, generating an input signal feature map from the connection feature map through a second set of convolutional layers, normalizing the gated signal feature map using a Sigmoid activation function to obtain a gated weight map, and multiplying the gated weight map element-wise with the input signal feature map to obtain the local enhancement feature map; in the second path, performing max pooling and average pooling on the connection feature map in the spatial dimension to obtain max pooling feature maps and average pooling feature maps, concatenating them in the channel dimension to obtain a pooled concatenated feature map, compressing the pooled concatenated feature map through convolutional layers to obtain a single-channel feature map, normalizing the single-channel feature map using a Sigmoid activation function to obtain a spatial position weight map, and multiplying the spatial position weight map element-wise with the connection feature map to obtain the global awareness feature map.

[0075] Specifically, after obtaining the connectivity feature map, the iris recognition device can perform dual-path parallel processing on the map to enhance the features from both local and global dimensions. It should be noted that accurate localization of the iris boundary requires both local texture details and constraints from global spatial relationships; therefore, this embodiment employs a dual-path parallel processing mechanism to allow both feature enhancement methods to function simultaneously.

[0076] In the first path processing, the iris recognition device employs a gated convolution mechanism to enhance the local texture of the connected feature map. Specifically, the iris recognition device simultaneously inputs the connected feature map into two sets of parallel convolutional layers. The first set of convolutional layers uses a 3x3 convolutional kernel to generate a gated signal feature map, which represents the importance of features in each channel. The second set of convolutional layers also uses a 3x3 convolutional kernel to generate an input signal feature map, used to extract the feature representation to be enhanced. It should be noted that gated convolution is a feature selection technique based on a gating mechanism, which dynamically controls the activation level of different feature channels through learned gating weights.

[0077] Furthermore, the iris recognition device normalizes the gated signal feature map using a sigmoid activation function, mapping the feature values ​​to weights between zero and one, thus obtaining a gated weight map. Each element in this gated weight map reflects the importance of the feature at its corresponding location; a weight close to one indicates that the feature at that location is very important and should be retained, while a weight close to zero indicates that the feature at that location is unimportant and should be suppressed. The iris recognition device then performs element-wise multiplication between the gated weight map and the input signal feature map. This weighting operation preserves and enhances important local texture features while weakening unimportant features, ultimately resulting in a locally enhanced feature map.

[0078] In the processing of the second path, the iris recognition device employs a spatial attention mechanism to model global spatial relationships. Specifically, the iris recognition device first performs max pooling and average pooling operations on the connection feature map in the spatial dimension. Max pooling takes the maximum value for each channel of the feature map across the entire space, obtaining the most significant response for that channel; average pooling takes the average value for each channel across the entire space, obtaining the global average response for that channel. The iris recognition device obtains a max-pooled feature map through max pooling and an average-pooled feature map through average pooling.

[0079] Furthermore, the iris recognition device concatenates the max-pooling feature map and the average-pooling feature map along the channel dimension to obtain a pooled concatenated feature map. This pooled concatenated feature map contains complementary information extracted by the two pooling methods. The iris recognition device then performs channel compression on this pooled concatenated feature map using a 7x7 convolutional layer, compressing the dual-channel features into a single-channel feature map. This convolutional operation integrates information from both max-pooling and average-pooling, learning a comprehensive spatial importance representation.

[0080] Based on this, the iris recognition device normalizes the single-channel feature map using the Sigmoid activation function to obtain a spatial location weight map. This spatial location weight map has the same spatial size as the connection feature map, where the value of each spatial location represents the importance weight of that location. Iris boundary locations often have higher weights due to significant feature changes, while flat areas such as the background have lower weights. Finally, the iris recognition device performs element-wise multiplication of the spatial location weight map and the connection feature map to obtain a globally perceptual feature map. During the multiplication process, the spatial location weight map is broadcast to the same number of channels as the connection feature map, ensuring that each location in each channel is weighted according to the same spatial weight.

[0081] Understandably, the gated convolution in the first path focuses on channel-level feature selection, excelling at capturing local texture patterns; while the spatial attention in the second path focuses on spatial-level position selection, excelling at modeling global spatial relationships. The two paths process the same connected feature map in parallel, enhancing features from different dimensions, resulting in locally enhanced feature maps and globally perceived feature maps that contain complementary information.

[0082] 204. Predict the pupil boundary mask and the outer iris boundary mask based on the boundary enhancement feature map;

[0083] In this embodiment, the iris recognition device first performs upsampling and convolutional decoding on the boundary enhancement feature map to restore it to the same spatial resolution as the input iris image. The upsampling operation can be achieved through bilinear interpolation or transposed convolution, gradually increasing the spatial size of the feature map. For example, if the size of the boundary enhancement feature map is H divided by 8 multiplied by W divided by 8, then after multiple upsampling operations, it can be restored to the original image size of H multiplied by W. During the upsampling process, the iris recognition device uses convolutional layers to decode the features, refining the upsampled features through convolutional operations to reduce the coarse edges caused by upsampling, thus obtaining the decoded feature map.

[0084] Furthermore, the iris recognition device adjusts the channel dimensions of the decoded feature maps using convolutional layers to generate pupil classification feature maps and iris outer boundary classification feature maps. Specifically, the iris recognition device can use two independent convolutional layers to process the decoded feature maps separately. Each convolutional layer adjusts the number of channels in the feature map to two, representing the background category channel and the target category channel, respectively. For the pupil classification feature map, the target category channel represents the probability that each pixel belongs to the pupil boundary; for the iris outer boundary classification feature map, the target category channel represents the probability that each pixel belongs to the iris outer boundary.

[0085] Based on this, the iris recognition device normalizes the pupil classification feature map and the iris outer boundary classification feature map using the Softmax activation function. The Softmax function is a multi-class normalization function that converts the score of each pixel in different categories into a probability distribution, ensuring that the sum of the probabilities of the same pixel in all categories is one. The iris recognition device obtains the pupil probability map after applying Softmax processing to the pupil classification feature map, and the iris outer boundary probability map after applying Softmax processing to the iris outer boundary classification feature map.

[0086] Finally, the iris recognition device extracts the pupil boundary category channel from the pupil probability map as the pupil boundary mask, and extracts the iris outer boundary category channel from the iris outer boundary probability map as the iris outer boundary mask. These two masks are pixel-level probability maps, where the value of each pixel represents the probability that the pixel belongs to the corresponding boundary region, with a value ranging from zero to one.

[0087] 205. Extract the boundary contours of the pupil boundary mask and the outer iris boundary mask to obtain the pupil boundary contour point set and the outer iris boundary contour point set;

[0088] 206. Perform geometric fitting based on the pupil boundary contour point set and the iris outer boundary contour point set, and optimize and correct the fitting parameters based on the concentric constraint relationship between the iris and the pupil to obtain the iris localization result.

[0089] In this embodiment, steps 205-206 are similar to steps 103-104 in the first embodiment, and will not be described again here.

[0090] In this embodiment, multi-level convolution processing is performed on the real-time input iris image to extract iris features at different scales, resulting in a feature map set. This feature map set is then input into a preset mask prediction model for mask prediction, yielding a pupil boundary mask and an outer iris boundary mask. Boundary contours are extracted from the pupil boundary mask and the outer iris boundary mask, resulting in a pupil boundary contour point set and an outer iris boundary contour point set, respectively. Geometric fitting is performed based on these two sets, and the fitting parameters are optimized and corrected according to the concentric constraint relationship between the iris and the pupil, resulting in the iris localization result. This invention, through a combination of boundary enhancement feature fusion and geometric constraint optimization, can accurately locate the inner and outer boundaries of the iris in non-cooperative environments, improving the accuracy and robustness of iris localization.

[0091] The above describes the real-time iris tracking and positioning method in the embodiments of the present invention. The following describes the real-time iris tracking and positioning device in the embodiments of the present invention. Please refer to [link to relevant documentation] for details on this real-time iris tracking and positioning device. Figure 3 One embodiment of the iris real-time tracking and positioning device in this invention includes:

[0092] The feature extraction module 301 is used to perform multi-level convolution processing on the real-time input iris image to extract iris features at different scales and obtain a set of feature maps.

[0093] The mask prediction module 302 is used to input the feature maps in the feature map set into a preset mask prediction model to perform mask prediction, and obtain the pupil boundary mask and the iris outer boundary mask.

[0094] The contour extraction module 303 is used to extract the boundary contours of the pupil boundary mask and the outer iris boundary mask to obtain the pupil boundary contour point set and the outer iris boundary contour point set.

[0095] The parameter optimization module 304 is used to perform geometric fitting based on the pupil boundary contour point set and the iris outer boundary contour point set, and to optimize and correct the fitting parameters based on the concentric constraint relationship between the iris and the pupil to obtain the iris positioning result.

[0096] In this embodiment of the invention, the real-time iris tracking and positioning device operates the aforementioned real-time iris tracking and positioning method. The device performs multi-level convolution processing on the real-time input iris image to extract iris features at different scales, obtaining a feature map set. The feature map set is then input into a preset mask prediction model for mask prediction, resulting in a pupil boundary mask and an outer iris boundary mask. Boundary contours are extracted from the pupil boundary mask and the outer iris boundary mask, yielding a pupil boundary contour point set and an outer iris boundary contour point set. Geometric fitting is performed based on the pupil boundary contour point set and the outer iris boundary contour point set, and the fitting parameters are optimized and corrected according to the concentric constraint relationship between the iris and the pupil to obtain the iris positioning result. This invention, through a combination of boundary enhancement feature fusion and geometric constraint optimization, can accurately locate the inner and outer boundaries of the iris in non-cooperative environments, improving the accuracy and robustness of iris positioning.

[0097] above Figure 3 The iris real-time tracking and positioning device in this embodiment of the invention will be described in detail from the perspective of unitized functional entities. The iris real-time tracking and positioning device in this embodiment of the invention will be described in detail from the perspective of hardware processing.

[0098] Figure 4This is a schematic diagram of the structure of a real-time iris tracking and positioning device 400 provided in an embodiment of the present invention. The real-time iris tracking and positioning device 400 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 410 (e.g., one or more processors) and a memory 420, and one or more storage media 430 (e.g., one or more mass storage devices) for storing application programs 333 or data 432. The memory 420 and storage media 430 can be temporary or persistent storage. The program stored in the storage media 430 may include one or more units (not shown in the diagram), each unit may include a series of instruction operations on the real-time iris tracking and positioning device 400. Furthermore, the processor 410 may be configured to communicate with the storage media 430 and execute the series of instruction operations in the storage media 430 on the real-time iris tracking and positioning device 400 to implement the steps of the aforementioned real-time iris tracking and positioning method.

[0099] The iris real-time tracking and positioning device 400 may also include one or more power supplies 440, one or more wired or wireless network interfaces 450, one or more input / output interfaces 460, and / or one or more operating systems 431, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 4 The illustrated iris real-time tracking and positioning device structure does not constitute a limitation on the iris real-time tracking and positioning device provided by the present invention. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0100] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the real-time iris tracking and positioning method.

[0101] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0102] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0103] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A real-time iris tracking and positioning method, characterized in that, The real-time iris tracking and positioning method includes: Multi-level convolution processing is performed on the real-time input iris image to extract iris features at different scales and obtain a set of feature maps; The feature maps in the feature map set are input into a preset mask prediction model. The high-level feature maps in the feature map set are upsampled using the mask prediction model. The upsampled high-level feature maps are then connected across scales with the corresponding low-level feature maps to obtain connection feature maps. In the first path, the connection feature maps are processed through a first set of convolutional layers to generate gated signal feature maps. The connection feature maps are then processed through a second set of convolutional layers to generate input signal feature maps. The gated signal feature maps are normalized using a Sigmoid activation function to obtain gated weight maps. The gated weight maps are then multiplied element-wise with the input signal feature maps to obtain local enhancement feature maps. In the second path, the connection feature maps are subjected to max pooling and average pooling in the spatial dimension to obtain max pooling feature maps and average pooling feature maps, which are then concatenated in the channel dimension to obtain... The pooled feature map is concatenated, and then channel-compressed through a convolutional layer to obtain a single-channel feature map. This single-channel feature map is then normalized using a sigmoid activation function to obtain a spatial position weight map. The spatial position weight map is then element-wise multiplied with the connection feature map to obtain a global perception feature map. The local enhancement feature map and the global perception feature map are concatenated along the channel dimension to obtain a fused feature map. The average pooling result in the spatial dimension is calculated on the fused feature map, and the difference between the fused feature map and the average pooling result is calculated to obtain spatial difference information. This spatial difference information is then processed through convolution and normalization to generate a boundary weight map. The fused feature map is then weighted pixel-wise based on the boundary weight map to obtain a boundary enhancement feature map. Finally, the pupil boundary mask and the outer iris boundary mask are predicted based on the boundary enhancement feature map. The boundary contours of the pupil boundary mask and the outer iris boundary mask are extracted to obtain the pupil boundary contour point set and the outer iris boundary contour point set; Geometric fitting is performed based on the pupil boundary contour point set and the iris outer boundary contour point set, and the fitting parameters are optimized and corrected according to the concentric constraint relationship between the iris and the pupil to obtain the iris localization result.

2. The iris real-time tracking and positioning method according to claim 1, characterized in that, The process involves performing multi-level convolutional processing on the real-time input iris image to extract iris features at different scales, resulting in a feature map set including: The iris image is segmented and initial feature mapping is performed through a convolutional layer to obtain an initial feature map; The initial feature map is used as the first layer input feature map, and N layers of feature extraction are performed sequentially. The i-th layer feature extraction is to perform spatial downsampling and channel expansion on the i-th layer input feature map to obtain the i-th layer downsampled feature map. The i-th layer downsampled feature map is then subjected to depthwise separable convolution to obtain the i-th layer intermediate feature map. The residual of the i-th layer intermediate feature map and the i-th layer downsampled feature map is added to obtain the i-th layer output feature map, and the i-th layer output feature map is used as the input feature map of the (i+1)-th layer. The initial feature map and the feature maps of each layer are collected to obtain the feature map set.

3. The iris real-time tracking and positioning method according to claim 1, characterized in that, The step of extracting the boundary contours of the pupil boundary mask and the outer iris boundary mask to obtain the pupil boundary contour point set and the outer iris boundary contour point set includes: The pupil boundary mask and the iris outer boundary mask are binarized to obtain a pupil binary mask and an iris outer boundary binary mask. Morphological processing and connected component filtering are performed on the binary mask of the pupil and the binary mask of the outer boundary of the iris to obtain the effective region of the pupil and the effective region of the outer boundary of the iris. The effective region of the pupil and the effective region of the outer boundary of the iris are extracted to obtain the pupil boundary contour point set and the outer boundary contour point set.

4. The iris real-time tracking and positioning method according to claim 1, characterized in that, The step of performing geometric fitting based on the pupil boundary contour point set and the iris outer boundary contour point set, and optimizing and correcting the fitting parameters based on the concentric constraint relationship between the iris and the pupil to obtain the iris localization result includes: Geometric fitting is performed on the pupil boundary contour point set and the iris outer boundary contour point set to obtain the initial geometric parameters of the pupil and the initial geometric parameters of the iris outer boundary. The shared center coordinates are obtained by taking a weighted average of the center coordinates of the initial geometric parameters of the pupil and the initial geometric parameters of the outer boundary of the iris. Based on the shared center coordinates, the radii of the pupil boundary contour point set and the outer boundary contour point set of the iris are calculated respectively to obtain the pupil radius and the outer boundary radius of the iris. The iris localization result is formed by combining the shared center coordinates, pupil radius, and outer boundary radius of the iris.

5. A real-time iris tracking and positioning device, characterized in that, For performing the real-time iris tracking and positioning method according to claim 1, the real-time iris tracking and positioning device includes: The feature extraction module is used to perform multi-level convolution processing on the real-time input iris image to extract iris features at different scales and obtain a set of feature maps. The mask prediction module is used to input the feature maps in the feature map set into a preset mask prediction model to perform mask prediction and obtain the pupil boundary mask and the outer boundary mask of the iris. The contour extraction module is used to extract the boundary contours of the pupil boundary mask and the outer iris boundary mask to obtain the pupil boundary contour point set and the outer iris boundary contour point set. The parameter optimization module is used to perform geometric fitting based on the pupil boundary contour point set and the iris outer boundary contour point set, and to optimize and correct the fitting parameters based on the concentric constraint relationship between the iris and the pupil to obtain the iris localization result.

6. A real-time iris tracking and positioning device, characterized in that, The real-time iris tracking and positioning device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the iris real-time tracking and positioning device to perform the steps of the iris real-time tracking and positioning method as described in any one of claims 1-4.

7. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is executed by the processor, it implements the steps of the real-time iris tracking and positioning method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Iris segmentation method and device based on full convolutional neural network, medium and equipment

    CN114445904A