A Finger Vein Recognition Method and System Based on Convolutional Neural Network and SIFT Algorithm

Through the combination of convolutional neural network and SIFT algorithm, the problems of difficulty in obtaining images and noise interference in finger vein recognition are solved, and efficient and accurate finger vein recognition is achieved, improving the system's recognition performance and speed.

CN112597812BActive Publication Date: 2025-07-11西安国创防务科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011407380.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-03
Publication Date
2025-07-11
Estimated Expiration
2040-12-03

AI Technical Summary

Technical Problem

In the prior art, the finger vein recognition method is affected by a variety of factors, which leads to difficulty in obtaining images, easy to be disturbed by noise, lack of specificity of finger vein characteristics, difficult to effectively extract trace information, and limited recognition performance.

Method used

The method based on convolutional neural network and SIFT algorithm is adopted to train the convolutional neural network through image preprocessing, deep feature extraction and SIFT feature extraction, and combine ternary loss function to improve feature extraction capabilities, simplify the recognition process, and speed up the recognition speed.

Benefits of technology

It greatly improves the recognition capability and accuracy of the finger vein recognition system, simplifies the recognition process, and improves the recognition speed and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112597812B_ABST
    Figure CN112597812B_ABST
Patent Text Reader

Abstract

The present invention proposes a finger vein recognition method and system based on a convolutional neural network and the SIFT (Scale-invariant feature transform) algorithm. The implementation process is as follows: First, obtain the finger vein image of the user to be recognized through data acquisition; Second, preprocess the received original vein image information and extract the region of interest of the image; Finally, extract the vein image feature information of the finger of the user to be recognized through the neural network in combination with the SIFI algorithm, and further perform recognition processing on the obtained finger vein image feature information of the user to be recognized, so as to achieve the purpose of recognizing the identity information corresponding to the finger vein information. The present invention can effectively extract finger vein features, improve the redundancy of noise, and significantly improve the recognition accuracy of the finger vein recognition system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a finger vein recognition method and system based on a convolutional neural network and the SIFT algorithm, and particularly relates to the technical field of biometric recognition. Background Art

[0002] Finger vein recognition is a type of biometric recognition technology that uses the superficial veins on the palm side of the finger for identity recognition. As an emerging biometric recognition method, finger vein recognition has attracted wide attention due to its high security, advantages in live body recognition, and high specificity.

[0003] The principle of finger vein biometric recognition is to obtain finger vein images formed by the absorption of near-infrared light by deoxyhemoglobin in the blood through a special camera. Then, the images are converted and stored as human biometric authentication data templates. During authentication, the finger vein images of the user to be recognized are obtained, biometric data is generated and compared with the user's stored template to identify the user's identity, which has the characteristics of security, live body recognition, stability, anti-counterfeiting, and non-contact.

[0004] In the prior art, during the finger vein recognition process, traditional finger vein image recognition methods are affected by various factors, and there are problems such as difficulty in obtaining finger vein images, easy interference of vein images by noise, and lack of specificity of finger vein features. Therefore, it is difficult to effectively extract finger vein texture information based on traditional manual features at present, resulting in limited recognition performance of the authentication system. Summary of the Invention

[0005] Object of the Invention: One object is to propose a finger vein recognition method based on a convolutional neural network and the SIFT algorithm to solve the above problems existing in the prior art. A further object is to propose an identity recognition system for implementing the above method.

[0006] Technical Solution: A finger vein recognition method based on a convolutional neural network and the SIFT algorithm includes the following steps:

[0007] Step 1: Obtain the original finger vein image of the user to be recognized through a finger vein information collection device;

[0008] Step 2: Receive the finger vein image of the user to be recognized and perform image data preprocessing;

[0009] Step 3: Extract the feature information of the image data;

[0010] Step 4: Identify the user identity information.

[0011] In a further embodiment, step two is further to receive the finger vein image collected in step one, and preprocess the received finger vein image data through grayscale conversion, edge extraction, image enhancement, curve extraction, image fusion, and normalization operations.

[0012] Among them, the grayscale conversion is further to perform grayscale conversion on the original finger vein image of the user to be recognized.

[0013] The edge extraction is further to obtain the finger edge area in the original finger vein image and extract the finger vein region of interest image.

[0014] The image enhancement is further to enhance the finger vein information image data in the original image by performing exposure, filtering, and contrast enhancement on the original finger vein image, thereby improving the visual effect of the image, converting it into a form more suitable for human or machine analysis and processing, and highlighting the meaningful information in the analysis process. Among them, the filtering further includes Gaussian filtering, median filtering, and Gabor filtering, which are used to highlight the spatial information of the image, suppress other irrelevant information, or remove certain information in the image and restore other information.

[0015] The curve extraction is further to perform a binary operation on the finger vein information image data after image enhancement to extract the finger vein curve according to the valley shape analysis method.

[0016] The image fusion is further to fuse the finger vein image after image enhancement and the finger vein image after curve extraction.

[0017] The normalization further includes pixel normalization and size normalization; the pixel normalization is to perform pixel normalization processing on the extracted finger vein region of interest image using the maximum-minimum method, and make the value range of the processed pixel points within [0,1]; the size normalization is to perform normalization processing on the image size of the finger vein image information after pixel normalization using the bilinear interpolation method, and make the processed image sizes the same, improving the processing efficiency of subsequent steps.

[0018] In a further embodiment, step three is further to first receive the finger vein image after image preprocessing in step two; secondly, input the region of interest in the received finger vein image into a convolutional neural network to extract the deep feature information of the finger vein image; thirdly, combine the SIFT algorithm to perform SIFT feature extraction on the region of interest in the received finger vein image.

[0019] In a further embodiment, the convolutional neural network is further trained for the accuracy of the network's feature extraction ability through a triplet loss function. Meanwhile, for the trained convolutional neural network, its encoder is used to process the received finger vein image, and the extracted features are stored in the database as deep features; the database is used to store the feature information extracted from the finger veins of existing users, including the deep feature information extracted by the convolutional neural network and the SIFT feature information obtained after processing the region of interest of the finger vein by the SIFT algorithm.

[0020] In a further embodiment, step four is further to read the corresponding existing identity information stored in the finger vein feature database, and compare the finger vein image feature information extracted in step three with the read information, so as to further obtain the identity information of the user to be identified. The comparison of the feature information uses the cosine distance. Specifically, the distance between the deep features of the finger vein image of the user to be identified and the deep features of the finger vein images of the users in the database is calculated, and the users in the database corresponding to the top predetermined number of features with smaller cosine distances are selected. The SIFT features of the finger vein image of the user to be identified are used to perform feature matching with the SIFT features of the finger vein images of the selected predetermined number of users in the database, so as to obtain the identity information of the user to be identified.

[0021] A finger vein recognition system based on a convolutional neural network and the SIFT algorithm can be used to implement the above method, and specifically includes:

[0022] An image acquisition module for acquiring finger vein images;

[0023] An image preprocessing module for preprocessing the original finger vein image data of the acquired user;

[0024] An image feature extraction module for extracting finger vein image features;

[0025] An image training module for training the image feature extraction ability;

[0026] An image recognition module for confirming the identity of the user to be identified.

[0027] In a further embodiment, the image preprocessing module further includes a grayscale module, an image region of interest extraction module, an image enhancement module, a vein curve extraction module, an image fusion module, and a normalization module; wherein the grayscale module performs grayscale processing on the image by setting the numerical values of the channel colors; the image region of interest extraction module performs edge detection on the image data grayed by the grayscale module using the image gradient difference to obtain the finger edge region in the image; the image enhancement module includes an exposure module, a filtering module, and a contrast enhancement module, which are used to enhance the finger vein image; the filtering module includes a Gaussian filtering module, a median filtering module, and a Gabor filtering module; the vein curve extraction module extracts the finger vein curve from the enhanced finger vein image by means of valley shape analysis and using binary operation; the image fusion module is used to fuse the finger vein images processed by the image enhancement module and the vein curve extraction module together; the normalization module includes a pixel normalization module and a size normalization module; the pixel normalization module is used to perform pixel normalization processing on the image of the finger vein region of interest extracted; the size normalization module is used to perform size normalization processing on the image size of the image information after pixel normalization.

[0028] In a further embodiment, the image feature extraction module further includes a deep feature extraction module and a SIFT feature extraction module; the deep feature extraction module further performs deep feature extraction on the finger vein image preprocessed by the image preprocessing module through a convolutional neural network; the SIFT feature extraction module further uses the scale-invariant feature transform algorithm to extract SIFT features from the finger vein image data after passing through the image preprocessing module by means of scale space generation, spatial extreme point detection, extreme point localization, feature pointing direction determination, key point descriptor generation, etc.

[0029] In a further embodiment, the image recognition module is used to identify the user's identity information based on the user's finger vein information, and further obtains, processes, and extracts the finger vein image information to be identified for the user's identity information by the image acquisition module, the image preprocessing module, and the image feature extraction module, and then performs feature comparison with the user's finger vein image feature information stored in the database by the image training module, and uses the feature data with a high degree of similarity as the user's feature information, and reads the identity information corresponding to the feature information, so as to realize the comparison and recognition of the identity information.

[0030] Advantages: The present invention provides a finger vein recognition method and system based on a convolutional neural network and the SIFT algorithm. By implementing the finger vein recognition method through the proposed system, the recognition ability of the finger vein recognition system is significantly improved. For the collected finger vein images, when extracting image features, the convolutional neural network involved is trained under the guidance of a triplet loss function, which improves the feature extraction ability for finger veins and enhances the security of the finger vein recognition system. At the same time, after the proposed method based on the convolutional neural network and the SIFT algorithm extracts the depth features and SIFT features of the finger veins, it first performs a preliminary screening through the depth features and then determines the user information to be recognized through the matching of the SIFT features, thus effectively simplifying the finger vein recognition process, accelerating the finger vein recognition speed, and improving the accuracy of vein recognition. Brief Description of the Drawings

[0031] Figure 1 is a schematic structural diagram of the finger vein recognition system based on the convolutional neural network and the SIFT algorithm of the present invention.

[0032] Figure 2 is a schematic flowchart of the finger vein recognition method based on the convolutional neural network and the SIFT algorithm of the present invention.

[0033] Figure 3 is a structural block diagram of the image preprocessing module in the method of the present invention. Detailed Embodiments

[0034] The present invention realizes the purpose of recognizing the user identity by a finger vein recognition method and system based on a convolutional neural network and the SIFT algorithm.

[0035] The applicant believes that using the superficial veins on the palmar side of the finger for identity recognition has the following defects while having the advantages of security, liveness recognition, stability, anti-counterfeiting, and non-contact:

[0036] It is difficult to obtain finger vein images, and it is difficult to establish an effective mathematical model and extract appropriate finger vein features in the current database.

[0037] Vein images are easily affected by noise, which in turn affects the distribution related to vein features in the image.

[0038] The scarcity of finger vein samples results in the lack of specificity of currently manually designed finger vein features.

[0039] Therefore, it is difficult to effectively extract finger vein pattern information by the current method based on manual features, resulting in limited recognition performance of the authentication system.

[0040] Meanwhile, the current application of deep learning in finger vein recognition also has great limitations and is often only applicable to the recognition of large-sample finger vein data and used as an image processing tool. Traditional deep learning methods based on the softmax function are often limited to increasing the inter-class distance of samples for classification. However, this approach ignores the intra-class distance of samples, making it often require large-sample finger vein data for training and not fully leveraging its powerful feature learning ability.

[0041] The following will further specifically describe this solution through embodiments in conjunction with the accompanying drawings. The described embodiments are partial embodiments of the present invention, not all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0042] As Figure 1 shown, a finger vein recognition system based on a convolutional neural network and the SIFT algorithm provided by the present invention includes: an image acquisition module, an image preprocessing module, an image feature extraction module, an image training module, and an image recognition module; wherein, the image acquisition module is used to acquire the original finger vein image of a user; the image preprocessing module is used to preprocess the original finger vein image of the user; the image feature extraction module is used to extract the feature information of the preprocessed original finger vein image; the image training module is used to train based on the original finger vein image of the user to be trained to obtain training parameters; the image recognition module is used to recognize the identity information of the user to be recognized according to the extracted finger vein features of the user to be recognized. Among them, the finger vein features specifically include depth features and SIFT features.

[0043] As Figure 3As shown in the figure, the image preprocessing module includes a grayscale unit, an ROI unit, an image enhancement unit, a binarization unit, a normalization unit, and an image fusion unit connected in sequence. Among them, the grayscale unit is used to grayscale the original user finger vein image; the ROI unit uses image gradient difference for edge detection to obtain the finger edge area in the image, so as to determine the finger image position, that is, to extract the finger vein region of interest image; the image enhancement unit uses methods such as image equalization exposure, Gaussian filtering, median filtering, contrast limited enhancement, and Gabor filtering to enhance the finger vein information in the original image; the binarization unit uses valley shape analysis to binarize the enhanced finger vein image and extract the vein curve; the image fusion unit is used to perform image fusion on the enhanced finger vein image and the binarized finger vein image; the normalization unit uses the MAX-MIN method to perform pixel normalization processing on the extracted finger vein region of interest image, and after processing, the pixel of each pixel point is in the range of [0,1]; and uses bilinear interpolation method to normalize the image size of the image information after pixel normalization, and the processed image sizes are the same, so as to facilitate the next step of processing. Among them, the processed image size is preferably 280 120.

[0044] The image feature extraction module includes a depth feature extraction unit and a SIFT feature extraction unit, which includes: a depth feature extraction unit for extracting depth features, which is extracted by using a convolutional neural network trained with a triplet loss function; a SIFT feature extraction unit for extracting SIFT features, which extracts feature values by using the SIFT algorithm.

[0045] The image training module is a convolutional neural network model, which includes: an image acquisition unit for acquiring the finger vein image information of all users; an image preprocessing unit for performing image preprocessing on the finger vein image information of all users to obtain the preprocessed ROI image; a training unit for inputting the preprocessed image information to be recognized into the convolutional neural network for training, and the training function is a triplet loss function to obtain the network parameters of the trained convolutional neural network.

[0046] Among them, the convolutional neural network includes an input layer, a convolutional layer, a max pooling layer, a batch normalization layer, a fully connected layer, and an output layer, and the loss function is a triplet loss function.

[0047] An image recognition module, which includes: a recognition unit based on depth features, calculating the distance between the depth features of the user's finger vein image and the depth features of the user's finger vein image in the database using cosine distance, and selecting the top 10 users in the database corresponding to the features with smaller cosine distance; a recognition unit based on SIFT features, performing feature matching on the SIFT features of the user's finger vein image and the SIFT features of the finger vein images of the 10 selected users in the database to obtain the identity information of the user to be recognized.

[0048] In another embodiment, as Figure 2 shown, an embodiment of the present invention provides a finger vein recognition method based on a convolutional neural network and the SIFT algorithm, which specifically includes the following steps:

[0049] Step 1: Obtain the original finger vein image of the user to be recognized through a finger vein information acquisition device;

[0050] Step 2: Receive the finger vein image of the user to be recognized and perform image data preprocessing; the specific processing method of this step is: first, grayscale the image; use the image gradient difference for edge detection to obtain the finger edge area in the image to determine the finger image position, that is, extract the region of interest (ROI) of the finger vein; adopt methods such as image equalization exposure, Gaussian filtering, median filtering, contrast-limited enhancement, and Gabor filtering to enhance the finger vein information in the original image; perform binarization on the enhanced finger vein image using valley shape analysis to extract the vein curve; perform image fusion on the enhanced finger vein image and the binarized finger vein image; perform pixel normalization on the extracted finger vein region of interest image using the MAX-MIN method, and after processing, the pixel of each pixel point is in the range of [0,1]; and perform normalization on the image size of the image information after pixel normalization using the bilinear interpolation method, and the processed image sizes are the same to obtain the final image.

[0051] Among them, in the process of grayscaling the image, a grayscaling method that conforms to the characteristics of the human eye is adopted, and let:

[0052]

[0053] In the formula, R, G, B are the three-channel values of the original color input image, and represents the grayscale image.

[0054] Among them, edge detection is performed to obtain the finger edge region in the image to determine the finger image position, that is, to extract the region of interest of finger veins, and the edge gradient difference is adopted. In this embodiment, since the number of pixel points in the horizontal axis and vertical axis of the image are 120 and 280 respectively, and the outside of the finger is a weak background, that is, the pixel value is close to 0. Therefore, the point with the fastest average increasing gradient in the horizontal direction from pixel point 0 to 60 is selected as the starting point of the ROI in the horizontal direction, and the point with the fastest average decreasing gradient is searched for from pixel point 60 to 120 as the ending point of the ROI in the horizontal direction. In the vertical direction, the point with the fastest average increasing gradient in the range from pixel point 0 to 140 is selected as the starting point of the ROI in the vertical direction, and the point with the fastest average decreasing gradient is searched for from pixel point 140 to 280 as the ending point of the ROI in the horizontal direction.

[0055] Among them, in the process of performing equalized exposure processing on the image information, contrast-limited adaptive histogram equalization is adopted. Specifically: First, considering an image of the region of interest, a threshold is set. If a certain gray level in the histogram exceeds the threshold, it is clipped, and the part exceeding the threshold is evenly distributed to the remaining gray levels, making the image of the region of interest smoother. Second, the image is divided into blocks and the histogram of each block is calculated. For each pixel point, its four adjacent windows are found, and the mapping values of the four window histograms to this pixel point are calculated respectively. Finally, bilinear interpolation is used to obtain the final mapping value of this pixel point.

[0056] Among them, the filtering and noise reduction is further realized by using Gaussian low-pass filtering and median filtering on the grayscale image information. The Gaussian low-pass filtering is defined as:

[0057]

[0058] In the formula, G represents the grayscale image, represents the image after Gaussian low-pass noise reduction, I represents the two-dimensional Gaussian kernel used, and at the same time satisfies , where represents the standard deviation, x and y represent the horizontal and vertical coordinates of the image respectively.

[0059] The median filtering is defined as:

[0060]

[0061] In the formula, represents the grayscale image, represents the image after median filtering noise reduction, Med represents the median of the selection matrix, x and y represent the horizontal and vertical coordinates of the image respectively, k and l represent the selected coordinate ranges respectively.

[0062] The Gabor filtering method is defined as follows:

[0063]

[0064] In the formula, \(G\) represents the grayscale image, represents the image after Gabor noise reduction, represents the two-dimensional Gaussian kernel used, and it satisfies , where , , , in the formula The value of represents the direction of the Gabor kernel function, The value of determines the wavelength of the Gabor filter, k represents the total number of directions, determines the size of the Gaussian window, and is preferably , x and y respectively represent the horizontal and vertical coordinates of the image.

[0065] Among them, valley shape analysis is to analyze the image by calculating the direction field of the image. In this embodiment, a 9*9 eight-direction mask template is defined, as shown in Table 1. Among them, 1 to 8 respectively represent 8 discrete directions. To accurately determine the valley shape direction, the local finger vein image is compared with the eight-direction mask template. If the pixel value of the local finger vein image is greater than the pixel value of the mask template, the direction of the central pixel point is taken as , otherwise take .

[0066] Table 1: 9*9 eight-direction mask template

[0067] 2 0 3 0 4 0 5 0 6 0 0 0 0 0 0 0 0 0 1 0 2 3 4 5 6 0 7 0 0 1 0 0 0 7 0 0 0 0 0 0 0 0 0 0 0 0 0 7 0 0 0 1 0 0 7 0 6 5 4 3 2 0 1 0 0 0 0 0 0 0 0 0 6 0 5 0 4 0 3 0 2

[0068] Among them, for pixel grayscale value normalization, the grayscale value of each pixel point is normalized to between 0 and 1. The normalization method is the MAX-MIN method, and the normalization formula is as follows:

[0069]

[0070] In the formula, is the normalized grayscale value of each pixel point, represents the grayscale value of the original image. Among them, bilinear interpolation is used for scale normalization processing. Each pixel point on the image information is traversed, and interpolation and adjustment are performed in the x , y direction for scale normalization processing. In this embodiment, the image size is normalized to 280*120.

[0071] Step 3: Extract the feature information of the image data. In this step, first, receive the finger vein image after image preprocessing in Step 2. Secondly, input the region of interest in the received finger vein image into a convolutional neural network to extract the deep feature information of the finger vein image. Thirdly, combine the SIFT algorithm to extract the SIFT features of the finger vein image for the region of interest in the received finger vein image.

[0072] Specifically, select the finger vein region of interest image as the training image, input the training image set into the convolutional neural network based on triplet loss for training. The input of the convolutional neural network is the finger vein image, and the output is the image deep feature value and label value of the image. Among them, the convolutional neural network consists of an input layer, a convolutional layer, a max-pooling layer, a fully connected layer, a batch normalization layer, and an output layer. Use the combination of triplet loss as the training index to train the convolutional neural network to obtain the network weights of the convolutional neural network. After completion of training, select the convolutional neural network model to process the input image and obtain the deep features. Use the triplet loss function to reduce the intra-class distance of the finger vein image and increase the inter-class distance of the finger vein image, so that the finger vein features of the same class are closer in the feature space; at the same time, make the finger vein features of different classes farther away in the feature space, so as to improve the recognition effect of the finger vein deep learning recognition model based on small-sample finger vein data.

[0073] In a specific embodiment of the present invention, select a sample triplet, which is composed as follows: randomly select a sample from the training data set, this sample is called the "anchor sample", and then randomly select a sample that belongs to the same class as the "anchor sample", called the positive sample, and a sample of a different class is called the negative sample, thus forming a triplet. For each element sample in the triplet, train a neural network to obtain the feature expressions of the three samples, which are respectively denoted as: , , . The purpose of the triplet loss is to make the distance between and as small as possible through learning, while the distance between and is as large as possible, and also make the distance between and have a minimum interval from the distance between and . The formula for the triplet loss is:

[0074]

[0075] In the formula, To represent the interval between two distances. In fact, to achieve better convergence and reduce parameter settings, a soft margin based on the softplus function is introduced into the triplet loss function. The improved triplet loss function is as follows:

[0076]

[0077] The formula of the softplus function is:

[0078]

[0079] The main role of the triplet loss is to guide the feature extraction ability of the convolutional neural network, making the intra-class distance between samples greater than the inter-class distance, so that the features of the same class of samples are more compactly distributed in the feature space. At the same time, to achieve better convergence, a soft margin based on the softplus function is also introduced into the loss function to achieve better training results.

[0080] In this embodiment, the convolutional neural network used consists of an input layer, a convolutional layer, a depthwise separable convolutional layer, an average pooling layer, a fully connected layer, and an output layer. Among them, a convolutional layer consists of a 3×3 convolution, a batch normalization layer, and a Relu activation function connected in sequence. Similarly, a depthwise separable convolutional layer consists of a 3×3 depthwise separable convolution, a batch normalization layer, a Relu activation function, a 1×1 convolution, a batch normalization layer, and a Relu activation function connected in sequence.

[0081] The specific model construction is as follows:

[0082] The first layer is the input layer, and the size of the input layer is the size of the finger vein image, 280 120 1;

[0083] The second layer is the convolutional layer, with a kernel size of 3×3, a stride of 2, and 32 feature maps are output;

[0084] The third layer is the depthwise separable convolutional layer, with a kernel size of 3×3, a stride of 1, and 32 feature maps are output;

[0085] The fourth layer is the convolutional layer, with a kernel size of 1×1, a stride of 1, and 64 feature maps are output;

[0086] The fifth layer is the depthwise separable convolutional layer, with a kernel size of 3×3, a stride of 2, and 64 feature maps are output;

[0087] The sixth layer is the convolutional layer, with a kernel size of 1×1, a stride of 1, and 128 feature maps are output;

[0088] The seventh layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 1, and outputs 128 feature maps;

[0089] The eighth layer is a convolutional layer with a kernel size of 1×1, a stride of 1, and outputs 128 feature maps;

[0090] The ninth layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 2, and outputs 128 feature maps;

[0091] The tenth layer is a convolutional layer with a kernel size of 1×1, a stride of 1, and outputs 256 feature maps;

[0092] The eleventh layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 1, and outputs 256 feature maps;

[0093] The twelfth layer is a convolutional layer with a kernel size of 1×1, a stride of 1, and outputs 256 feature maps;

[0094] The thirteenth layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 2, and outputs 256 feature maps;

[0095] The fourteenth layer is a convolutional layer with a kernel size of 1×1, a stride of 1, and outputs 512 feature maps;

[0096] The fifteenth layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 1, and outputs 512 feature maps;

[0097] The sixteenth layer is a convolutional layer with a kernel size of 1×1, a stride of 1, and outputs 512 feature maps;

[0098] The seventeenth layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 1, and outputs 512 feature maps;

[0099] The eighteenth layer is a convolutional layer with a kernel size of 1×1, a stride of 1, and outputs 512 feature maps;

[0100] The nineteenth layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 1, and outputs 512 feature maps;

[0101] The twentieth layer is a convolutional layer with a kernel size of 1×1, a stride of 1, and outputs 512 feature maps;

[0102] The twenty-first layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 1, and outputs 512 feature maps;

[0103] The twenty-second layer is a convolutional layer with a kernel size of 1×1, a stride of 1, and outputs 512 feature maps;

[0104] The twenty-third layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 1, and outputs 512 feature maps;

[0105] The twenty-fourth layer is a convolutional layer with a convolution kernel size of 1×1, a stride of 1, and outputs 512 feature maps;

[0106] The twenty-fifth layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 2, and outputs 512 feature maps;

[0107] The twenty-sixth layer is a convolutional layer with a convolution kernel size of 1×1, a stride of 1, and outputs 1024 feature maps;

[0108] The twenty-seventh layer is a depthwise separable convolutional layer with a kernel size of 3×3, a stride of 2, and outputs 1024 feature maps;

[0109] The twenty-eighth layer is a convolutional layer with a convolution kernel size of 1×1, a stride of 1, and outputs 1024 feature maps;

[0110] The twenty-ninth layer is an average pooling layer with a kernel size of 7×7, a stride of 1, and outputs 1024 feature maps;

[0111] The thirtieth layer is a fully connected layer with 1024 nodes.

[0112] Table 2: The convolutional neural network model adopted

[0113] Neural network layer Stride Convolution kernel size Input Convolution layer 2 3#timg#3 280#timg#120#timg#1 Depthwise separable convolution layer 1 3#timg#3 140#timg#60#timg#32 Convolution layer 1 1#timg#1 140#timg#60#timg#32 Depthwise separable convolution layer 2 3#timg#3 140#timg#60#timg#64 Convolution layer 1 1#timg#1 70#timg#30#timg#64 Depthwise separable convolution layer 1 3#timg#3 70#timg#30#timg#128 Convolution layer 1 1#timg#1 70#timg#30#timg#128 Depthwise separable convolution layer 2 3#timg#3 70#timg#30#timg#128 Convolution layer 1 1#timg#1 35#timg#15#timg#128 Depthwise separable convolution layer 1 3#timg#3 35#timg#15#timg#256 Convolution layer 1 1#timg#1 35#timg#15#timg#256 Depthwise separable convolution layer 2 3#timg#3 35#timg#15#timg#256 Convolution layer 1 1#timg#1 18#timg#8#timg#512 Depthwise separable convolution layer 1 3#timg#3 18#timg#8#timg#512 Convolution layer 1 1#timg#1 18#timg#8#timg#512 Depthwise separable convolution layer 1 3#timg#3 18#timg#8#timg#512 Convolution layer 1 1#timg#1 18#timg#8#timg#512 Depthwise separable convolution layer 1 3#timg#3 18#timg#8#timg#512 Convolution layer 1 1#timg#1 18#timg#8#timg#512 Depthwise separable convolution layer 1 3#timg#3 18#timg#8#timg#512 Convolution layer 1 1#timg#1 18#timg#8#timg#512 Depthwise separable convolution layer 1 3#timg#3 18#timg#8#timg#512 Convolution layer 1 1#timg#1 18#timg#8#timg#512 Depthwise separable convolution layer 2 3#timg#3 18#timg#8#timg#512 Convolution layer 1 1#timg#1 9#timg#4#timg#512 Depthwise separable convolution layer 2 3#timg#3 9#timg#4#timg#1024 Convolution layer 1 1#timg#1 4#timg#2#timg#1024 Average pooling 1 7#timg#7 4#timg#2#timg#1024 Fully connected layer / / 1024

[0114] During the feature extraction process, for SIFT features, key point descriptors of the finger vein image need to be selected first. The coordinate axes are determined as the directions of the key points to ensure rotational invariance. A 5×5 window is taken with the key point as the center, and the gradient of each pixel in the 10×10 window around the key point is calculated. Moreover, a Gaussian decay function is used to reduce the weights far from the center. When calculating the features, a 128-dimensional descriptor is formed for each feature, and each dimension can represent the scale in 4×4 grids. The number of SIFT key point descriptors is the first 100 high-quality descriptors, and the two-dimensional horizontal and vertical coordinates of the key points are recorded. Therefore, the number of SIFT features for each finger vein image is 100×(128 + 2) = 13000.

[0115] Step 4: Identify the user's identity information. In this step, the existing identity information corresponding to the stored finger vein feature database is read, and the finger vein image feature information extracted in Step 3 is received. The two are compared in terms of feature information to further obtain the identity information of the user to be identified. The comparison of feature information uses the cosine distance. Specifically, the distance between the depth features of the finger vein image of the user to be identified and the depth features of the finger vein images of the users in the database is calculated, and the top ten features corresponding to the smaller cosine distance in the database are selected. The SIFT features of the finger vein image of the user to be identified are matched with the SIFT features of the finger vein images of the 10 selected users in the database to obtain the identity information of the user to be identified.

[0116] In a specific embodiment, that is, if it is necessary to match the key points of finger vein image 1 and image 2, each key point descriptor of the former is taken one by one, and the two key points with the closest Euclidean distance in image 2 are found. Among these two key points, if the closest distance divided by the second-closest distance is less than 0.85, and the Euclidean distance of their two-dimensional horizontal and vertical coordinates is less than 25 pixels, then this pair of matching points is accepted. The proportion of the accepted key points in image 2 among all key points is counted. If this proportion value is greater than 0.16, it is considered that finger vein image 1 and finger vein image 2 match, otherwise they are excluded.

[0117] To prove the superiority of this method, this embodiment compares the performance of traditional texture feature extraction methods: the method based on Gabor filter and 2DPCA, the method based on sliding window filtering, the method based on double sliding window filtering, and the results of the convolutional neural network method using the softmax loss function. Experiments are carried out on the self-owned dataset, and the results are shown in Table 3. Among them, the neural network structure is shown in Table 1.

[0118] Table 3. Comparison of equal error rates

[0119] Method Equal error rate (%) Based on Gabor filter and 2DPCA 8.22 Based on sliding window filtering 4.67 Based on double sliding window filtering 3.35 Convolutional neural network based on softmax loss function 1.32 Method of this example 0.75

[0120] On the self-owned finger vein database, based on a total of 2036 types of finger vein images, 6 samples are collected for each type. The equal error rate of the recognition result of the fusion of Gabor filter and 2DPCA is 8.22%, the equal error rate of the recognition result based on sliding window filtering is 4.67%, the equal error rate of the recognition result based on double sliding window filtering is 3.35%. Under the same convolutional neural network structure, the equal error rate of the recognition result using the softmax loss function is 1.32%, and the equal error rate of the recognition using the features extracted in this embodiment is 0.75%. The features extracted in this example method can better express the fundamental information of finger veins. Therefore, it is more effective to use the method in this example to extract finger vein image features.

[0121] As described above, although the present invention has been shown and described with reference to particular preferred embodiments, it should not be construed as a limitation on the invention itself. Various changes in form and detail may be made therein without departing from the spirit and scope of the invention as defined by the appended claims.

Claims

1. A finger vein recognition method based on a convolutional neural network and the SIFT algorithm, characterized in that It includes the following steps: Step 1: Obtain the original finger vein image of the user to be recognized through a finger vein information acquisition device; Step 2: Receive the finger vein image of the user to be recognized and perform preprocessing on the image data; Step 3: Extract the region of interest from the preprocessed image data and input it into a convolutional neural network for extracting the depth feature information of the finger vein image. At the same time, use the SIFT algorithm to extract the SIFT features of the region of interest in the received finger vein image to obtain the total image data feature information; The convolutional neural network uses a triplet loss function to train the accuracy of the network's feature extraction ability; Step 4: Calculate the distance between the depth features of the finger vein image of the user to be recognized and the depth features of the finger vein images of the users in the database, select the users in the database corresponding to the top predetermined number of features with smaller cosine distances, and then use the SIFT features of the finger vein image of the user to be recognized and the SIFT features of the finger vein images of the predetermined number of users selected from the database for feature matching to determine the user identity information.

2. The finger vein recognition method based on convolutional neural network and SIFT algorithm according to claim 1, characterized in that, The further details of Step 2 are to receive the finger vein image collected in Step 1 and perform preprocessing on the received finger vein image data through operations such as grayscale conversion, edge extraction, image enhancement, curve extraction, image fusion, and normalization; Among them, the grayscale conversion is further to perform grayscale conversion on the original finger vein image of the user to be recognized; The edge extraction is further to obtain the finger edge region in the original finger vein image and extract the finger vein region of interest image; The image enhancement is further to enhance the finger vein information image data in the original image by performing exposure, filtering, and contrast enhancement on the original finger vein image; the filtering further includes Gaussian filtering, median filtering, and Gabor filtering; the Gaussian filtering is further defined as: G σ = I * G where G represents the grayscale image, and G σ represents the image after Gaussian low-pass noise reduction, I represents the two-dimensional Gaussian kernel used, and simultaneously satisfies where σ represents the standard deviation, and x and y respectively represent the horizontal and vertical coordinates of the image; The median filtering is defined as: G m (x,y) = Med[G(x - k,y - l)], (k,l ∈ W) where G(x, y) represents the grayscale image, G m (x, y) represents the image after median filtering for noise reduction, Med represents the median of the selected matrix, x and y respectively represent the horizontal and vertical coordinates of the image, and k and l respectively represent the selected coordinate ranges; The Gabor filtering method is defined as: G g = I g * G In the formula, G represents the grayscale image, G g represents the image after Gabor noise reduction, I g represents the two-dimensional Gaussian kernel used, and satisfies where In the formula, the value of μ represents the direction of the Gabor kernel function, the value of determines the wavelength of the Gabor filter, k represents the total number of directions, determines the size of the Gaussian window, and x and y respectively represent the horizontal and vertical coordinates of the image; The curve extraction is further to perform binary operation to extract the finger vein curve on the finger vein information image data after image enhancement according to the valley shape analysis method; The image fusion is further to fuse the finger vein image after image enhancement and the finger vein image after curve extraction; The normalization further includes pixel normalization and size normalization; the pixel normalization is to perform pixel normalization processing on the extracted finger vein region of interest image, and it is further defined as: In the formula, Z(x,y) is the grayscale value of each pixel point after normalization, and g(x,y) represents the grayscale value of the original image; the size normalization is to perform normalization processing on the image size of the finger vein image information after pixel normalization.

3. A finger vein recognition method based on a convolutional neural network and the SIFT algorithm according to claim 1, characterized in that, The convolutional neural network further uses a triplet loss function to train the accuracy of the network's feature extraction ability. At the same time, for the trained convolutional neural network, use it to process the received finger vein image and store the extracted features as depth features in the database; The database is used to store the feature information extracted from the finger veins of existing users, including the deep feature information extracted by the convolutional neural network and the SIFT feature information obtained after processing the region of interest of the finger veins by the SIFT algorithm.

4. A finger vein recognition system based on a convolutional neural network and the SIFT algorithm, for implementing the method according to any one of claims 1 to 3, characterized in that, It includes: An image acquisition module for acquiring finger vein images; An image preprocessing module for preprocessing the original finger vein image data of the acquired user; An image feature extraction module for extracting finger vein image features; An image training module for training the image feature extraction ability; An image recognition module for confirming the identity of the user to be recognized; The image feature extraction module further includes a deep feature extraction module and a SIFT feature extraction module; the deep feature extraction module further extracts deep features from the finger vein image preprocessed by the image preprocessing module through a convolutional neural network; the SIFT feature extraction module further uses the SIFT algorithm to extract SIFT features from the finger vein image data after passing through the image preprocessing module through scale space generation, spatial extreme point detection, extreme point localization, feature pointing direction determination, and key point descriptor generation; The image recognition module is used to identify the identity information of the user according to the finger vein information of the user, further obtain, process, and extract the finger vein image information to be recognized for the user identity information by the image acquisition module, the image preprocessing module, and the image feature extraction module, and then compare the features with the finger vein image feature information of the user stored in the database by the image training module. The feature data with a high degree of similarity is used as the feature information of the user, and the identity information corresponding to the feature information is read, so as to realize the comparison and recognition of the identity information.

5. The finger vein recognition system based on a convolutional neural network and the SIFT algorithm according to claim 4, wherein The image preprocessing module further includes a grayscale module, an image region of interest extraction module, an image enhancement module, a vein curve extraction module, an image fusion module, and a normalization module; wherein the grayscale module performs grayscale processing on the image; the image region of interest extraction module uses the image gradient difference to perform edge detection on the grayscale image data to obtain the finger edge region in the image; the image enhancement module includes an exposure module, a filtering module, and a contrast enhancement module, which are used to enhance the finger vein image; the vein curve extraction module extracts the finger vein curve from the enhanced finger vein image by using the valley shape analysis method and performing binary operation; the image fusion module is used to fuse the finger vein images processed by the image enhancement module and the vein curve extraction module together; the normalization module includes a pixel normalization module and a size normalization module; The pixel normalization module is used to perform pixel normalization processing on the extracted finger vein region of interest image; The size normalization module is used to perform size normalization processing on the image size of the image information after pixel normalization.

6. The finger vein recognition system based on a convolutional neural network and the SIFT algorithm according to claim 4, wherein The image training module is further configured to train the feature extraction ability of the convolutional neural network. In combination with the image acquisition module and the image preprocessing module, it first reads the finger vein image data collected by the image acquisition module and inputs it into the image preprocessing module for image preprocessing, and obtains the region of interest image in the preprocessed finger vein image. Subsequently, it is input into the convolutional neural network, and the triplet loss function is used to train the deep feature extraction ability.

Citation Information

Patent Citations

  • Convolutional variational auto-encoder neural network-based finger vein identification method and system

    CN108009520A

  • Finger vein recognition method and system based on cosine center loss

    CN110390282A