Iris segmentation method, device, medium and equipment based on fully convolutional neural network

Through the method based on a fully convolutional neural network, the iris image is preprocessed and feature extraction, combined with attention module and expansion model, the problems of large time-consuming and poor robustness of traditional iris segmentation algorithms are solved, and high-precision and high-rootability iris segmentation are achieved, which improves the accuracy of iris recognition.

CN114445904BActive Publication Date: 2025-05-23BEIJING INST OF RADIO METROLOGY & MEASUREMENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111561511.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2025-05-23
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

The traditional iris segmentation algorithm takes a lot of time to calculate, has a huge parameter space, and can only process high-quality iris pictures. It is easily affected by occlusion, reflection and shadow, and it is difficult to meet high requirements in terms of real-time and robustness.

Method used

Using an iris segmentation method based on a fully convolutional neural network, the acquired iris images are preprocessed, and the trained compression model is input for feature extraction. The attention module is used to assign weights to different pixels, and the dilation model is combined for upsampling and feature fusion. Finally, the iris mask, pupil mask and outer iris boundary are obtained through channel separation, and noise is removed through post-processing optimization.

Benefits of technology

It realizes high-precision and robust iris segmentation, which is suitable for all kinds of iris images, improves the accuracy of overall iris recognition and reduces calculation time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114445904B_ABST
    Figure CN114445904B_ABST
Patent Text Reader

Abstract

The present invention discloses an iris segmentation method and apparatus, medium and equipment based on a fully convolutional neural network. In a specific embodiment, the method includes inputting a preprocessed iris image into a trained compression model, extracting features from the preprocessed iris image to obtain a compressed iris feature P; using the compressed iris feature P as the input of an attention module, assigning different weights W to different pixels of the iris image, and finally outputting an iris feature P'; using the iris feature P' as the input of a trained expansion model, upsampling and fusing it with the corresponding features in the compression model in the channel dimension to expand the iris image to the original input size; performing channel separation on the output result of the expansion model to obtain an iris mask, a pupil mask and an iris outer boundary; optimizing and removing noise from the iris mask, the pupil mask and the iris outer boundary through post-processing to obtain the final iris mask and iris inner and outer boundary information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet application technology, and more specifically, to an iris segmentation method and apparatus, medium and device based on a fully convolutional neural network. Background Art

[0002] In recent years, identity authentication through biometrics has become increasingly common in security fields such as banks, borders, and access control. Biometric technologies such as fingerprint recognition, face recognition, and iris recognition have been widely used in our lives. Among various types of biometric information, iris has attracted the attention of many researchers due to its high reliability, high stability, and non-contact nature. Iris recognition mainly includes five steps: iris acquisition, iris segmentation, normalization, iris encoding, and iris matching. Iris segmentation refers to segmenting the iris area from the acquired iris image. As the basis for the subsequent recognition process, iris segmentation has an important impact on the speed and accuracy of the entire iris recognition process.

[0003] Traditional iris segmentation algorithms mainly rely on digital image processing technology to perform segmentation based on the grayscale and gradient information of iris images, with differential-difference algorithm and Hough transform algorithm as the main representatives. The main problem with traditional methods is that the parameter space is very large, which leads to a lot of computational time. At the same time, traditional methods can only process high-quality iris images, and require users to cooperate highly in the iris acquisition process. Otherwise, occlusions such as eyelashes and eyelids, as well as reflections and shadows will affect the calculation results.

[0004] With the development of artificial intelligence, deep learning technology has achieved performance that surpasses traditional methods in various computer vision tasks. As one of the tasks of computer vision, iris segmentation is also being actively explored for the possibility of using deep learning technology. However, unlike general image segmentation tasks, iris segmentation has extremely high requirements for the accuracy, robustness, and real-time performance of the algorithm. In addition, most iris segmentation methods based on deep learning can only obtain iris masks. In order to perform subsequent normalization, other methods are needed to obtain the information of the inner and outer boundaries of the iris based on the iris mask. This method is different from the previous algorithm that first obtains the inner and outer boundaries and then obtains the iris mask. It is easy to cause positioning errors in the case of severe occlusion. Summary of the invention

[0005] In order to solve the above problems, the object of the present invention is to provide an iris segmentation method and apparatus, medium and equipment based on a fully convolutional neural network.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A first aspect of the present invention provides an iris segmentation method based on a fully convolutional neural network, the method comprising:

[0008] Preprocessing the collected iris image;

[0009] Inputting the preprocessed iris image into the trained compression model, extracting features from the preprocessed iris image to obtain compressed iris features P;

[0010] Using the compressed iris feature P as the input of the attention module, assigning different weights W to different pixels of the iris image, and finally outputting the iris feature P';

[0011] The iris feature P' is used as the input of the trained expansion model, up-sampled and fused with the corresponding feature in the compression model in the channel dimension to expand the iris image to the original input size;

[0012] Performing channel separation on the output result of the dilation model to obtain an iris mask, a pupil mask and an iris outer boundary;

[0013] The iris mask, pupil mask and iris outer boundary are optimized and noise is removed by post-processing to obtain the final iris mask and iris inner and outer boundary information.

[0014] Furthermore, the preprocessing of the collected iris image includes:

[0015] The collected iris image is preprocessed by using a mean subtraction method to highlight iris features, and the size of the iris image is adjusted by center cropping and / or scaling to obtain an iris image of a preset size.

[0016] Furthermore, the structure of the trained compression model is as follows:

[0017] The first layer is the convolution layer, which includes 64 convolution kernels of size 3*3*3 and a step size of 1. This layer uses the SAME mode filling to ensure that the feature size remains unchanged before and after convolution, and only the number of channels changes;

[0018] The second layer is the convolution layer, which includes 64 convolution kernels of size 3*3*64 and a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output0;

[0019] The third layer is the maximum pooling layer, with a step size of 2 and a size of 2*2;

[0020] The fourth layer is the convolution layer, which includes 128 convolution kernels of size 3*3*64 and a step size of 1. This layer is filled with the SAME mode.

[0021] The fifth layer is the convolution layer, which includes 128 3*3*128 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output1.

[0022] The sixth layer is the maximum pooling layer, with a step size of 2 and a size of 2*2;

[0023] The seventh layer is the convolution layer, which includes 256 3*3*128 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0024] The eighth layer is the convolution layer, which includes 256 3*3*256 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0025] The ninth layer is the convolution layer, which includes 256 3*3*256 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output2;

[0026] The tenth layer is the maximum pooling layer, with a step size of 2 and a size of 2*2;

[0027] The eleventh layer is the convolution layer, which includes 512 3*3*256 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0028] The twelfth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode;

[0029] The thirteenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output3;

[0030] The fourteenth layer is the maximum pooling layer, with a step size of 2 and a size of 2*2;

[0031] The fifteenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0032] The sixteenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0033] The seventeenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​the compressed iris feature P;

[0034] Each convolutional layer is followed by a ReLU activation function.

[0035] Furthermore, the step of using the compressed iris feature P as an input of an attention module and assigning different weights to different pixels of the iris image includes:

[0036] Use global average pooling to obtain the global average value AdaptiveAvgPool(P) of the compressed iris feature P, and then use a convolution layer composed of 256 convolution kernels of size 1*1*512 for convolution, and then use linear interpolation for upsampling to restore the global average value AdaptiveAvgPool(P) to its original size. The calculation formula is:

[0037] G(P)=Up(Conv 1×1 (AdaptiveAvgPool(P)))……(1)

[0038] Among them, Conv 1×1 (AdaptiveAvgPool(P)) represents the output result of the convolutional layer;

[0039] Up(Conv 1×1 (AdaptiveAvgPool(P))) represents the upsampling output result;

[0040] The compressed iris feature P is convolved with 256 convolution kernels of size 1*1*512, and the output of the last convolution layer is D 1 (P):

[0041] D 1 (P) = Conv 1×1 (P)……(2)

[0042] The compressed iris feature P is convolved using dilated convolutions with dilation rates of 6, 12, and 18, respectively. The three dilated convolutions each include 256 convolution kernels with a size of 3*3*512:

[0043]

[0044]

[0045]

[0046] The output results G(P), D of each convolutional layer 1 (P), D 2 (P), D 3 (P) and D 4 (P) Concatenate in the channel dimension, then perform two convolutions and activate with the sigmoid() activation function to obtain the output weight W of the attention mechanism:

[0047]

[0048] W = Sigmoid(Conv 3×3 (Conv 3×3 (H)))……(7)

[0049] in The symbol represents the concatenation in the channel dimension, and H represents the concatenation result;

[0050] The compressed iris feature P is element-wise multiplied with the output weight W and concatenated with itself to obtain the final output iris feature P' of the attention mechanism:

[0051]

[0052] in The symbol represents concatenation in the channel dimension, Represents element-wise dot product.

[0053] The structure of the trained expansion model is as follows:

[0054] The first layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size. The number of output channels is 1024, and its output matrix is ​​recorded as up0;

[0055] The second layer is the concatenation layer, which concatenates up0 and output3 according to the number of channels;

[0056] The third layer is the convolution layer, which includes 768 convolution kernels of size 3*3*1536;

[0057] The fourth layer is the convolution layer, which includes 256 3*3*786 convolution kernels;

[0058] The fifth layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size, and its output matrix is ​​recorded as up1;

[0059] The sixth layer is the concatenation layer, which concatenates up1 and output2 according to the number of channels;

[0060] The seventh layer is the convolution layer, which includes 256 convolution kernels of size 3*3*512;

[0061] The eighth layer is the convolution layer, which includes 128 convolution kernels of size 3*3*256;

[0062] The ninth layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size, and its output matrix is ​​recorded as up2;

[0063] The tenth layer is the concatenation layer, which concatenates up2 and output1 according to the number of channels;

[0064] The eleventh layer is the convolution layer, which includes 128 convolution kernels of size 3*3*256;

[0065] The twelfth layer is the convolution layer, which includes 64 convolution kernels of size 3*3*128;

[0066] The thirteenth layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size, and its output matrix is ​​recorded as up3;

[0067] The fourteenth layer is a splicing layer, which splices up3 and output0 according to the number of channels;

[0068] The fifteenth layer is the convolution layer, which includes 64 convolution kernels of size 3*3*128;

[0069] The sixteenth layer is a convolution layer, which includes 32 convolution kernels with a size of 3*3*64, and its output result is recorded as the output result of the expansion model;

[0070] Each convolutional layer is followed by a batch normalization and ReLU activation function.

[0071] Furthermore, the step of performing channel separation on the output result of the dilation model to obtain an iris mask, a pupil mask and an outer boundary of the iris includes:

[0072] The output result of the dilation model is input into a convolution layer composed of three convolution kernels of size 1*1*32, the number of channels of the output result of the convolution layer is set to 3, the output result is channel segmented, the sigmoid function is used to set the value of the output result to between [0,1], and then multiplied by 255 to be an integer, and the final output result of the network, namely, the iris mask, pupil mask and iris outer boundary, is obtained.

[0073] Furthermore, the post-processing is performed to optimize and remove noise from the iris mask, pupil mask and iris outer boundary to obtain the final iris mask and iris inner and outer boundary information, including:

[0074] Using fixed thresholds to binarize the iris mask, pupil mask and iris outer boundary respectively to obtain valid pixels;

[0075] Use morphological closing operation to connect the tiny breakpoints in the iris boundary to make the outer boundary of the iris more complete;

[0076] Extracting 8-connected domains in the iris mask, pupil mask and iris outer boundary respectively;

[0077] The Chebyshev distances between the iris mask connected domain and the pupil mask connected domain, as well as the iris mask connected domain and the iris outer boundary mask connected domain are calculated pairwise to obtain a candidate triplet set;

[0078] Select the largest triplet from the candidate triplet set, whose members are the candidate iris mask, pupil mask and iris outer boundary;

[0079] Extracting a contour from the candidate pupil mask and the outer boundary of the iris, and performing least squares circle fitting with the points on the contour to obtain the center coordinates and radius of the optimized inner and outer boundaries of the iris;

[0080] The candidate iris mask is optimized according to the optimized inner and outer boundaries of the iris, and the portion outside the inner and outer boundaries is removed to obtain the optimized iris mask.

[0081] A second aspect of the present invention provides an iris segmentation device based on a fully convolutional neural network, the device comprising:

[0082] An iris image acquisition module, used for capturing iris images;

[0083] An iris image preprocessing module is used to perform a preprocessing operation on the iris image obtained by the iris image acquisition module to obtain a preprocessed iris image;

[0084] An iris image compression module is used to input the preprocessed iris image obtained by the iris image preprocessing module into a trained compression model to obtain corresponding features and compressed iris features P;

[0085] An iris image processing module, used for assigning different weights to different pixels of the compressed iris feature P output by the iris image compression module through an attention module, and obtaining an iris feature P';

[0086] An iris image expansion module, used to upsample the compressed iris feature P' to expand it to its original size, and to merge it with the corresponding feature output by the iris image compression module in the channel dimension;

[0087] An iris image segmentation module, used for performing channel separation on the output result of the iris image expansion module to obtain an iris mask, a pupil mask and an iris outer boundary;

[0088] The iris image post-processing module is used to optimize and remove noise from the iris mask, pupil mask and iris outer boundary through post-processing to obtain the final iris mask and iris inner and outer boundary information.

[0089] The third aspect of the present invention provides a computer-readable storage medium, which stores an application program, wherein when the application program is executed by a processor, an iris image segmentation method based on a fully convolutional neural network provided in the first aspect of the present invention is implemented.

[0090] The fourth aspect of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the iris image segmentation method based on a fully convolutional neural network provided in the first aspect of the present invention can be implemented.

[0091] Beneficial effects of the present invention:

[0092] The technical solution provided by the present invention performs iris segmentation based on a fully convolutional neural network, has high precision and high robustness, is suitable for the segmentation of various types of iris images, and helps to improve the accuracy of overall iris recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0094] Figure 1 A step diagram of an iris segmentation method based on a fully convolutional neural network provided by an embodiment of the present invention is shown.

[0095] Figure 2 A schematic diagram of a fully convolutional neural network provided by an embodiment of the present invention is shown.

[0096] Figure 3 A step diagram showing an image post-processing method provided by the present embodiment

[0097] Figure 4 A schematic diagram of an iris segmentation device based on a fully convolutional neural network provided by an embodiment of the present invention is shown.

[0098] Figure 5 A schematic diagram showing the structure of a computer system for implementing the apparatus provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0099] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only an embodiment of a part of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in the field without making creative work should fall within the scope of protection of the present invention.

[0100] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish between objects of a category, and are not necessarily used to describe a specific order or sequence. It should be understood that the objects used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include those steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0101] First embodiment - an iris segmentation method based on a fully convolutional neural network, such as Figure 1 As shown, the iris segmentation method provided in this embodiment includes the following steps:

[0102] S1: preprocessing the collected iris image;

[0103] S2: Inputting the preprocessed iris image into the trained compression model, extracting features from the preprocessed iris image to obtain compressed iris features P;

[0104] S3: using the compressed iris feature P as the input of the attention module, assigning different weights W to different pixels of the iris image, and finally outputting the iris feature P';

[0105] S4: taking the iris feature P' as the input of the trained expansion model, upsampling it and fusing it with the corresponding feature in the compression model in the channel dimension to expand the iris image to the original input size;

[0106] S5: performing channel separation on the output result of the dilation model to obtain an iris mask, a pupil mask and an iris outer boundary;

[0107] S6: Optimize and remove noise from the iris mask, pupil mask and iris outer boundary through post-processing to obtain the final iris mask and iris inner and outer boundary information.

[0108] In a possible implementation manner, the preprocessing of the collected iris image includes:

[0109] The collected iris image is preprocessed by using a mean subtraction method to highlight iris features, and the size of the iris image is adjusted by center cropping and / or scaling to obtain an iris image of a preset size.

[0110] Those skilled in the art will understand that the iris mask is a black and white image, where white pixels represent iris pixels and black pixels represent non-iris pixels. The inner and outer boundaries of the iris are represented by a circle on a black background image. Here, since the width of the boundary output by the iris segmentation network is greater than 1, morphological dilation is used to widen the width of the marked circle.

[0111] When training a fully convolutional neural network, the collected iris images need to be manually annotated and preprocessed. The iris mask and the inner and outer boundaries of the iris need to be manually annotated on the iris image. In the preprocessing part, data enhancement is used to solve the problem of insufficient iris image data. In deep learning of images, in order to enrich the image training set, better extract image features, and generalize the model, data enhancement is generally performed on the data image. The data enhancement methods used in the examples of the present invention include: random image resizing, random image blurring, random translation, random rotation, random flipping, and random cropping.

[0112] The fully convolutional neural network is trained using the manually annotated and preprocessed iris images as a training set. The fully convolutional neural network structure includes: a compression model, an attention module and an expansion model.

[0113] In a specific embodiment, the compression model is divided into five stages, which include multiple convolutional layers and pooling layers. The convolution kernel size used in the convolutional layer is 3×3, and the number of channels of the convolution output is 64, 128, 256, 512, and 512 respectively. Each convolutional layer is followed by a ReLU activation function. At the end of the first four stages, a 2×2 maximum pooling layer is used for downsampling, while expanding the receptive field of the convolutional layer. It can be seen that as the network deepens, the number of channels of the iris feature map continues to increase, while the size gradually decreases. That is to say, in the first few stages, the convolution extracts the lower-level spatial information in the image, while the higher-level semantic information is extracted in the later stage.

[0114] There is a transition stage between the contraction module and the expansion model. The contraction module gradually extracts high-dimensional features from the iris image, and then gradually expands the segmentation results in the expansion model based on the compressed features. Therefore, the processing in the transition stage is particularly important for the segmentation results of the network. The attention mechanism is a mechanism in deep learning. It is inspired by the fact that when humans observe a picture, they tend to focus on the important parts of the picture and ignore other unimportant parts. In the deep learning network, a weight value can also be trained for each pixel in the image, so that pixels with larger weights and more critical pixels play a more important role in subsequent processing.

[0115] In one possible implementation, the attention module has a total of 5 parallel processes:

[0116] Use global average pooling and 1×1 convolution to obtain global information, and then upsample to the original compressed feature size; perform 1×1 convolution on the compressed features; use 3×3 convolution with dilation rates of 6, 12, and 18 to convolve the compressed features;

[0117] Then the output results of the five parallel modules are concatenated in the channel dimension, convolved twice, and adjusted with the sigmoid function to obtain the weight of the attention mechanism output. The compressed feature is dot-multiplied with the weight and then concatenated with itself to obtain the output of the attention module, which is also the input of the expansion model.

[0118] Corresponding to the contraction module, the expansion model consists of four stages. In each stage, the features are first upsampled using bilinear interpolation, then concatenated with the corresponding features in the contraction module, and then convolved by two 3×3 convolutional layers. Each convolutional layer is followed by batch normalization (BN) and ReLU activation function. The number of output channels in each stage is 256, 128, 64, and 32 respectively.

[0119] The manually annotated and preprocessed iris image is input as a training set into the fully convolutional neural network to obtain a trained fully convolutional neural network, including: a trained compression model, an expansion model and an attention module.

[0120] In one possible implementation, Figure 2 As shown, the structure of the trained compression model is as follows:

[0121] The first layer is the convolution layer, which includes 64 convolution kernels of size 3*3*3 and a step size of 1. This layer uses the SAME mode filling to ensure that the feature size remains unchanged before and after convolution, and only the number of channels changes;

[0122] The second layer is the convolution layer, which includes 64 convolution kernels of size 3*3*64 and a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output0;

[0123] The third layer is the maximum pooling layer, with a step size of 2 and a size of 2*2;

[0124] The fourth layer is the convolution layer, which includes 128 convolution kernels of size 3*3*64 and a step size of 1. This layer is filled with the SAME mode.

[0125] The fifth layer is the convolution layer, which includes 128 3*3*128 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output1.

[0126] The sixth layer is the maximum pooling layer, with a step size of 2 and a size of 2*2;

[0127] The seventh layer is the convolution layer, which includes 256 3*3*128 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0128] The eighth layer is the convolution layer, which includes 256 3*3*256 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0129] The ninth layer is the convolution layer, which includes 256 3*3*256 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output2;

[0130] The tenth layer is the maximum pooling layer, with a step size of 2 and a size of 2*2;

[0131] The eleventh layer is the convolution layer, which includes 512 3*3*256 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0132] The twelfth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode;

[0133] The thirteenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output3;

[0134] The fourteenth layer is the maximum pooling layer, with a step size of 2 and a size of 2*2;

[0135] The fifteenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0136] The sixteenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode.

[0137] The seventeenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​the compressed iris feature P;

[0138] Each convolutional layer is followed by a ReLU activation function.

[0139] In a possible implementation, taking the compressed iris feature P as the input of the attention module and assigning different weights to different pixels of the iris image includes:

[0140] Use global average pooling to obtain the global average value AdaptiveAvgPool(P) of the compressed iris feature P, and then use a convolution layer composed of 256 convolution kernels of size 1*1*512 for convolution, and then use linear interpolation for upsampling to restore the global average value AdaptiveAvgPool(P) to its original size. The calculation formula is:

[0141] G(P)=Up(Conv 1×1 (AdaptiveAvgPool(P)))......(1)

[0142] Among them, Conv 1×1 (AdaptiveAvgPool(P)) represents the output result of the convolutional layer;

[0143] Up(Conv 1×1 (AdaptiveAvgPool(P))) represents the upsampling output result;

[0144] The compressed iris feature P is convolved with 256 convolution kernels of size 1*1*512, and the output of the last convolution layer is D 1 (P):

[0145] D 1 (P) = Conv 1×1 (P)……(2)

[0146] The compressed iris feature P is convolved using dilated convolutions with dilation rates of 6, 12, and 18, respectively. The three dilated convolutions each include 256 convolution kernels with a size of 3*3*512:

[0147]

[0148]

[0149]

[0150] The output results G(P), D of each convolutional layer 1 (P), D 2 (P), D 3 (P) and D 4 (P) Concatenate in the channel dimension, then perform two convolutions and activate with the sigmoid() activation function to obtain the output weight W of the attention mechanism:

[0151]

[0152] W = Sigmoid(Conv 3×3 (Conv 3×3 (H)))……(7)

[0153] in The symbol represents the concatenation in the channel dimension, and H represents the concatenation result;

[0154] The compressed iris feature P is element-wise multiplied with the output weight W and concatenated with itself to obtain the final output iris feature P' of the attention mechanism:

[0155]

[0156] in The symbol represents concatenation in the channel dimension, Represents element-wise dot product.

[0157] In a specific embodiment, Figure 2 As shown, the structure of the trained expansion model is as follows:

[0158] The first layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size. The number of output channels is 1024, and its output matrix is ​​recorded as up0;

[0159] The second layer is the concatenation layer, which concatenates up0 and output3 according to the number of channels;

[0160] The third layer is the convolution layer, which includes 768 convolution kernels of size 3*3*1536;

[0161] The fourth layer is the convolution layer, which includes 256 3*3*786 convolution kernels;

[0162] The fifth layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size, and its output matrix is ​​recorded as up1;

[0163] The sixth layer is the concatenation layer, which concatenates up1 and output2 according to the number of channels;

[0164] The seventh layer is the convolution layer, which includes 256 convolution kernels of size 3*3*512;

[0165] The eighth layer is the convolution layer, which includes 128 convolution kernels of size 3*3*256;

[0166] The ninth layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size, and its output matrix is ​​recorded as up2;

[0167] The tenth layer is the concatenation layer, which concatenates up2 and output1 according to the number of channels;

[0168] The eleventh layer is the convolution layer, which includes 128 convolution kernels of size 3*3*256;

[0169] The twelfth layer is the convolution layer, which includes 64 convolution kernels of size 3*3*128;

[0170] The thirteenth layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size, and its output matrix is ​​recorded as up3;

[0171] The fourteenth layer is a splicing layer, which splices up3 and output0 according to the number of channels;

[0172] The fifteenth layer is the convolution layer, which includes 64 convolution kernels of size 3*3*128;

[0173] The sixteenth layer is a convolution layer, which includes 32 convolution kernels with a size of 3*3*64, and its output result is recorded as the output result of the expansion model;

[0174] Each convolutional layer is followed by a batch normalization and ReLU activation function.

[0175] In a specific embodiment, the step of performing channel separation on the output result of the dilation model to obtain the iris mask, the pupil mask and the outer boundary of the iris includes:

[0176] The output result of the dilation model is input into a convolution layer composed of three convolution kernels of size 1*1*32, the number of channels of the output result of the convolution layer is set to 3, the output result is channel segmented, the sigmoid function is used to set the value of the output result to between [0,1], and then multiplied by 255 to be an integer, and the final output result of the network, namely, the iris mask, pupil mask and iris outer boundary, is obtained.

[0177] In one possible implementation, Figure 3As shown, the iris mask, pupil mask and iris outer boundary are optimized and noise is removed by post-processing to obtain the final iris mask and iris inner and outer boundary information including:

[0178] Binarization: using a fixed threshold to binarize the iris mask, pupil mask and iris outer boundary respectively to obtain valid pixels;

[0179] Perform morphological closing operation on the outer boundary of the iris: Use morphological closing operation to connect the tiny breakpoints in the iris boundary to make the outer boundary of the iris more complete;

[0180] Extracting 8-connected domains: extracting 8-connected domains in the iris mask, pupil mask and iris outer boundary respectively;

[0181] Calculate the Chebyshev distance to form a candidate triple set: Calculate the Chebyshev distance between the iris mask connected domain and the pupil mask connected domain, and the Chebyshev distance between the iris mask connected domain and the iris outer boundary mask connected domain, and set the connected domains with both distances less than the threshold as a candidate triple. When calculating the connected domain distance, first extract the outer contours of the two connected domains to be calculated, and then calculate the Chebyshev distance between each pixel in the contour, and use the closest distance between all pixels as the distance between the two connected domains. In this way, a candidate triple set can be obtained.

[0182] Select the largest triplet: Select the largest triplet in the candidate triplet set, whose members are the candidate iris mask, pupil mask and iris outer boundary;

[0183] Calculate the inner boundary of the iris based on the candidate pupil mask: Extract the inner boundary information of the iris from the candidate pupil mask. Extract the outer contour of the pupil mask, which is the contour of the inner boundary. Fit a circle based on all the pixels in the contour using the least squares circle fitting method to obtain the center coordinates and radius of the inner boundary of the iris.

[0184] Calculate the iris outer boundary based on the candidate iris outer boundary: Extract the iris outer boundary information from the candidate iris outer boundary. Since the iris outer boundary obtained in the above process is in the form of a contour, all the pixels therein are directly used to perform least squares circle fitting to obtain the center coordinates and radius of the iris outer boundary.

[0185] Optimizing the iris mask: optimizing the candidate iris mask according to the optimized inner and outer boundaries of the iris, and removing the parts outside the inner and outer boundaries to obtain the optimized iris mask.

[0186] Second embodiment - an iris segmentation device based on a fully convolutional neural network, such as Figure 4 As shown, the device comprises:

[0187] An iris image acquisition module, used for capturing iris images;

[0188] An iris image preprocessing module is used to perform a preprocessing operation on the iris image obtained by the iris image acquisition module to obtain a preprocessed iris image;

[0189] An iris image compression module is used to input the preprocessed iris image obtained by the iris image preprocessing module into a trained compression model to obtain corresponding features and compressed iris features P;

[0190] An iris image processing module, used for assigning different weights to different pixels of the compressed iris feature P output by the iris image compression module through an attention module, and obtaining an iris feature P';

[0191] An iris image expansion module, used to upsample the compressed iris feature P' to expand it to its original size, and to merge it with the corresponding feature output by the iris image compression module in the channel dimension;

[0192] An iris image segmentation module is used to perform channel separation on the output result of the iris image expansion module to obtain an iris mask, a pupil mask and an iris outer boundary;

[0193] The iris image post-processing module is used to optimize and remove noise from the iris mask, pupil mask and iris outer boundary through post-processing to obtain the final iris mask and iris inner and outer boundary information.

[0194] It should be noted that the principle and workflow of the iris segmentation device based on a fully convolutional neural network provided in this embodiment are similar to the above-mentioned iris segmentation method based on a fully convolutional neural network. The relevant parts can be referred to the above description and will not be repeated here.

[0195] Third embodiment - a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program:

[0196] Preprocessing the collected iris image;

[0197] Inputting the preprocessed iris image into the trained compression model, extracting features from the preprocessed iris image to obtain compressed iris features P;

[0198] Using the compressed iris feature P as the input of the attention module, assigning different weights W to different pixels of the iris image, and finally outputting the iris feature P';

[0199] The iris feature P' is used as the input of the trained expansion model, up-sampled and fused with the corresponding feature in the compression model in the channel dimension to expand the iris image to the original input size;

[0200] Performing channel separation on the output result of the dilation model to obtain an iris mask, a pupil mask and an iris outer boundary;

[0201] The iris mask, pupil mask and iris outer boundary are optimized and noise is removed by post-processing to obtain the final iris mask and iris inner and outer boundary information.

[0202] Fourth embodiment - A computer readable storage medium having a computer program stored thereon, which when executed by a processor implements:

[0203] Preprocessing the collected iris image;

[0204] Inputting the preprocessed iris image into the trained compression model, extracting features from the preprocessed iris image to obtain compressed iris features P;

[0205] The compressed iris feature P is used as the input of the attention module, different weights W are assigned to different pixels of the iris image, and finally an iris feature P' is output;

[0206] The iris feature P' is used as the input of the trained expansion model, up-sampled and fused with the corresponding feature in the compression model in the channel dimension to expand the iris image to the original input size;

[0207] Performing channel separation on the output result of the dilation model to obtain an iris mask, a pupil mask and an iris outer boundary;

[0208] The iris mask, pupil mask and iris outer boundary are optimized and noise is removed by post-processing to obtain the final iris mask and iris inner and outer boundary information.

[0209] In practical applications, the computer-readable storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present embodiment, the computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.

[0210] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0211] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0212] Computer program code for performing the operations of the present invention may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0213] Fifth embodiment - a computer device. Figure 5 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0214] like Figure 5 As shown, the computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).

[0215] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0216] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0217] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 5 not shown, usually called a "hard drive"). Although Figure 5 Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.

[0218] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described herein.

[0219] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed through an input / output (I / O) interface 22. Furthermore, the computer device 12 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through a network adapter 20. Figure 5 As shown, the network adapter 20 communicates with other modules of the computer device 12 via the bus 18. It should be understood that although Figure 5 Not shown, other hardware and / or software modules may be used in conjunction with computer device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0220] The processor unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing an iris segmentation method based on a fully convolutional neural network provided in an embodiment of the present invention.

[0221] It should be noted that, in the description of the present invention, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusions.

[0222] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not limitations on the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is impossible to list all the embodiments here. All obvious changes or modifications derived from the technical solution of the present invention are still within the protection scope of the present invention.

Claims

1. An iris segmentation method based on fully convolutional neural network, It is characterized in that The method comprises: Preprocessing the collected iris image; Inputting the preprocessed iris image into the trained compression model, extracting features from the preprocessed iris image to obtain compressed iris features P; Using the compressed iris feature P as the input of the attention module, assigning different weights W to different pixels of the iris image, and finally outputting the iris feature P'; The iris feature P' is used as the input of the trained expansion model, up-sampled and fused with the corresponding feature in the compression model in the channel dimension to expand the iris image to the original input size; Performing channel separation on the output result of the dilation model to obtain an iris mask, a pupil mask and an iris outer boundary; The iris mask, pupil mask and iris outer boundary are optimized and noise is removed by post-processing to obtain the final iris mask and iris inner and outer boundary information.

2. The method according to claim 1, It is characterized in that The preprocessing of the collected iris image comprises: The collected iris image is preprocessed by using a mean subtraction method to highlight iris features, and the size of the iris image is adjusted by center cropping and / or scaling to obtain an iris image of a preset size.

3. The method according to claim 1, It is characterized in that The structure of the trained compression model is as follows: The first layer is the convolution layer, which includes 64 convolution kernels of size 3*3*3 and a step size of 1. This layer uses the SAME mode filling to ensure that the feature size remains unchanged before and after convolution, and only the number of channels changes; The second layer is the convolution layer, which includes 64 convolution kernels of size 3*3*64 and a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output0; The third layer is the maximum pooling layer, with a step size of 2 and a size of 2*2; The fourth layer is the convolution layer, which includes 128 convolution kernels of size 3*3*64 and a step size of 1. This layer is filled with the SAME mode. The fifth layer is the convolution layer, which includes 128 3*3*128 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output1. The sixth layer is the maximum pooling layer, with a step size of 2 and a size of 2*2; The seventh layer is the convolution layer, which includes 256 3*3*128 convolution kernels with a step size of 1. This layer is filled with the SAME mode. The eighth layer is the convolution layer, which includes 256 3*3*256 convolution kernels with a step size of 1. This layer is filled with the SAME mode. The ninth layer is the convolution layer, which includes 256 3*3*256 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output2; The tenth layer is the maximum pooling layer, with a step size of 2 and a size of 2*2; The eleventh layer is the convolution layer, which includes 512 3*3*256 convolution kernels with a step size of 1. This layer is filled with the SAME mode. The twelfth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode; The thirteenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​recorded as output3; The fourteenth layer is the maximum pooling layer, with a step size of 2 and a size of 2*2; The fifteenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode. The sixteenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode. The seventeenth layer is the convolution layer, which includes 512 3*3*512 convolution kernels with a step size of 1. This layer is filled with the SAME mode, and its output matrix is ​​the compressed iris feature P; Each convolutional layer is followed by a ReLU activation function.

4. The method according to claim 1, It is characterized in that The step of using the compressed iris feature P as the input of the attention module and assigning different weights to different pixels of the iris image comprises: Use global average pooling to obtain the global average value AdaptiveAvgPool(P) of the compressed iris feature P, and then use a convolution layer composed of 256 convolution kernels of size 1*1*512 for convolution, and then use linear interpolation for upsampling to restore the global average value AdaptiveAvgPool(P) to its original size. The calculation formula is: G(P)=Up(Conv 1×1 (AdaptiveAvgPool(P)))……(1) Among them, Conv 1×1 (AdaptiveAvgPool(P)) represents the output result of the convolutional layer; Up(Conv 1×1 (AdaptiveAvgPool(P))) represents the upsampling output result; The compressed iris feature P is convolved with 256 convolution kernels of size 1*1*512, and the output of the last convolution layer is D 1 (P): D 1 (P)=Conv 1×1 (P)……(2) The compressed iris feature P is convolved using dilated convolutions with dilation rates of 6, 12, and 18, respectively. The three dilated convolutions each include 256 convolution kernels with a size of 3*3*512: The output results G(P), D of each convolutional layer 1 (P), D 2 (P), D 3 (P) and D 4 (P) Concatenate in the channel dimension, then perform two convolutions and activate with the sigmoid() activation function to obtain the output weight W of the attention mechanism: W=Sigmoid(Conv 3×3 (Conv 3×3 (H)))……(7) in The symbol represents the concatenation in the channel dimension, and H represents the concatenation result; The compressed iris feature P is element-wise multiplied with the output weight W and concatenated with itself to obtain the final output iris feature P' of the attention mechanism: in The symbol represents concatenation in the channel dimension, Represents element-wise dot product.

5. The method according to claim 1, It is characterized in that The structure of the trained expansion model is as follows: The first layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size. The number of output channels is 1024, and its output matrix is ​​recorded as up0; The second layer is the concatenation layer, which concatenates up0 and output3 according to the number of channels; The third layer is the convolution layer, which includes 768 convolution kernels of size 3*3*1536; The fourth layer is the convolution layer, which includes 256 3*3*786 convolution kernels; The fifth layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size, and its output matrix is ​​recorded as up1; The sixth layer is the concatenation layer, which concatenates up1 and output2 according to the number of channels; The seventh layer is the convolution layer, which includes 256 convolution kernels of size 3*3*512; The eighth layer is the convolution layer, which includes 128 convolution kernels of size 3*3*256; The ninth layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size, and its output matrix is ​​recorded as up2; The tenth layer is the concatenation layer, which concatenates up2 and output1 according to the number of channels; The eleventh layer is the convolution layer, which includes 128 convolution kernels of size 3*3*256; The twelfth layer is the convolution layer, which includes 64 convolution kernels of size 3*3*128; The thirteenth layer is the upsampling layer, which uses bilinear interpolation to expand the upsampled input matrix to twice its size, and its output matrix is ​​recorded as up3; The fourteenth layer is a splicing layer, which splices up3 and output0 according to the number of channels; The fifteenth layer is the convolution layer, which includes 64 convolution kernels of size 3*3*128; The sixteenth layer is a convolution layer, which includes 32 convolution kernels with a size of 3*3*64, and its output result is recorded as the output result of the expansion model; Each convolutional layer is followed by a batch normalization and ReLU activation function.

6. The method according to claim 1, It is characterized in that The step of performing channel separation on the output result of the dilation model to obtain an iris mask, a pupil mask and an outer boundary of the iris includes: The output result of the dilation model is input into a convolution layer composed of three convolution kernels of size 1*1*32, the number of channels of the output result of the convolution layer is set to 3, the output result is channel segmented, the sigmoid function is used to set the value of the output result to between [0,1], and then multiplied by 255 to be an integer, and the final output result of the network, namely, the iris mask, pupil mask and iris outer boundary, is obtained.

7. The method according to claim 1, It is characterized in that The post-processing is performed to optimize and remove noise from the iris mask, pupil mask and iris outer boundary to obtain the final iris mask and iris inner and outer boundary information, including: Using fixed thresholds to binarize the iris mask, pupil mask and iris outer boundary respectively to obtain valid pixels; Use morphological closing operation to connect the tiny breakpoints in the iris boundary to make the outer boundary of the iris more complete; Extracting 8-connected domains in the iris mask, pupil mask and iris outer boundary respectively; The Chebyshev distances between the iris mask connected domain and the pupil mask connected domain, as well as the iris mask connected domain and the iris outer boundary mask connected domain are calculated pairwise to obtain a candidate triplet set; Select the largest triplet from the candidate triplet set, whose members are the candidate iris mask, pupil mask and iris outer boundary; Extracting a contour from the candidate pupil mask and the outer boundary of the iris, and performing least squares circle fitting with the points on the contour to obtain the center coordinates and radius of the optimized inner and outer boundaries of the iris; The candidate iris mask is optimized according to the optimized inner and outer boundaries of the iris, and the portion outside the inner and outer boundaries is removed to obtain the optimized iris mask.

8. An iris segmentation device based on a fully convolutional neural network, It is characterized in that include: An iris image acquisition module, used for capturing iris images; An iris image preprocessing module is used to perform a preprocessing operation on the iris image obtained by the iris image acquisition module to obtain a preprocessed iris image; An iris image compression module is used to input the preprocessed iris image obtained by the iris image preprocessing module into a trained compression model to obtain corresponding features and compressed iris features P; An iris image processing module, used for assigning different weights to different pixels of the compressed iris feature P output by the iris image compression module through an attention module, and obtaining an iris feature P'; An iris image expansion module, used to upsample the compressed iris feature P' to expand it to its original size, and to merge it with the corresponding feature output by the iris image compression module in the channel dimension; An iris image segmentation module, used for performing channel separation on the output result of the iris image expansion module to obtain an iris mask, a pupil mask and an iris outer boundary; The iris image post-processing module is used to optimize and remove noise from the iris mask, pupil mask and iris outer boundary through post-processing to obtain the final iris mask and iris inner and outer boundary information.

9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores an application program, wherein when the application program is executed by a processor, the iris image segmentation method based on a fully convolutional neural network as described in any one of claims 1 to 7 is implemented.

10. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the iris image segmentation method based on a fully convolutional neural network as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Face recognition method based on convolutional neural network and attention model

    CN111582044A

  • Iris automatic segmentation method and system based on multi-model voting mechanism

    CN113706469A