A finger vein feature extraction method based on deep metric learning

Through the finger vein feature extraction method based on deep metric learning, the problem of poor robustness of the algorithm and inability to handle new categories in the prior art is solved, and the robustness and performance improvement of data set expansion and feature extraction are achieved.

CN114973336BActive Publication Date: 2025-05-06UNIV OF ELECTRONICS SCI & TECH OF CHINA ZHONGSHAN INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210543640.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-19
Publication Date
2025-05-06
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

The existing venous feature extraction methods have poor algorithm robustness and are unable to effectively deal with the problems of new categories and backbone networks being unable to represent rich and deep features.

Method used

A finger vein feature extraction method based on deep metric learning is adopted. By pre-processing the original image of the finger vein, the region of interest is extracted, and the triplets composed of positive and negative sample pairs are transferred to the feature extraction network to obtain the feature triplet. Loss is calculated in the Euro-style measurement space and cosine measurement space, and backpropagation optimization ensures the intra-class compactness and inter-class dispersion of features in the measurement space.

Benefits of technology

The data set is expanded, significant features are extracted, with good robustness and performance, and has excellent performance on large-scale data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973336B_ABST
    Figure CN114973336B_ABST
Patent Text Reader

Abstract

The present invention provides a finger vein feature extraction method based on deep metric learning, comprising the following steps: step 1, preprocessing the original finger vein image to extract the region of interest of the image; step 2, passing the triplet (a, p, n) consisting of positive and negative sample pairs into the feature extraction network to obtain the feature triplet; step 3, calculating the loss in the Euclidean metric space, then passing the feature triplet into the classifier to obtain the category probability, and calculating the loss in the cosine metric space, and optimizing them together through back propagation. Compared with the prior art, the present invention has the following beneficial effects: (1) Based on the training method of sample pairing, the data set is expanded in another form. (2) It can extract significant features and achieve the goal of small distance between the same type and large distance between different types in the metric space. (3) It has good robustness and still has good performance on large-scale data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of finger vein recognition, and in particular relates to a finger vein feature extraction method based on deep metric learning. Background Art

[0002] Feature extraction is a key stage in the finger vein recognition process. In this stage, the finger vein image collected by the infrared sensor is mapped to a certain feature measurement space, and the quantifiable properties of the finger vein are represented by feature vectors. These feature vectors will be used as templates for finger vein recognition and for matching or recognition in the next stage.

[0003] There are roughly three types of traditional finger vein feature extraction methods: vein pattern-based methods, local binary code-based methods, and dimensionality reduction-based methods. These methods extract features by manually designing feature parameters (such as the geometry and topology of vein patterns, local binary codes, etc.) by human experts. They require images to have clear vein patterns and require careful manual adjustment of parameters, which is complex to implement and has poor algorithm robustness. With the development of deep learning, it has become a feasible approach to allow computational autonomous learning of feature parameters. Feature extraction methods based on deep learning, such as FV-Net and FV-GAN, have achieved good results on some public finger vein datasets, but these methods have two problems: First, the algorithm test is based on closed set recognition, which may not meet the needs of continuously adding new categories in finger vein scenarios. Second, the backbone networks of these methods usually use small CNNs, which cannot represent rich deep features, and the optimization objectives only focus on feature separability without considering the intra-class compactness of features. Summary of the invention

[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a finger vein feature extraction method based on deep metric learning to solve the problems raised in the above background technology.

[0005] The present invention is implemented by the following technical solution: A finger vein feature extraction method based on deep metric learning, comprising the following steps:

[0006] Step 1, preprocessing the original finger vein image to extract the region of interest of the image;

[0007] Step 2: Pass the triplet (a, p, n) consisting of positive and negative sample pairs into the feature extraction network to obtain a feature triplet;

[0008] Step 3: Calculate the loss in the Euclidean metric space, then pass the feature triples into the classifier to obtain the category probability, calculate the loss in the cosine metric space, and optimize them together through back propagation.

[0009] As a preferred implementation, in step 1, first obtain the upper and lower boundaries of the finger to determine the position of the finger, then calculate the central axis of the finger based on the boundaries of the finger, and inversely calculate the angle of the finger plane offset to perform rotational correction on the finger, and finally find the position of the second joint of the finger to capture the area of ​​interest.

[0010] After adopting the above technical solution, the beneficial effects of the present invention are as follows: (1) Based on the training method of sample pairing, the data set is expanded in another form. (2) It can extract significant features and achieve the goal of small distance between similar types and large distance between different types in the metric space. (3) It has good robustness and still has good performance on large-scale data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0012] Figure 1 Schematic diagram of the overall framework for finger vein feature extraction.

[0013] Figure 2 Schematic diagram of finger boundary detection.

[0014] Figure 3 Schematic diagram of finger rotation correction.

[0015] Figure 4 Schematic diagram of ROI extraction for finger images.

[0016] Figure 5 Schematic diagram of residual learning.

[0017] Figure 6 Schematic diagram of the residual module.

[0018] Figure 7 Schematic representation of the FV-FE-Net network structure.

[0019] Figure 8 Schematic diagram of the problems in metric learning.

[0020] Fig. 9 Schematic diagram of sample update direction.

[0021] Fig.10 A schematic diagram for comparison with other methods. DETAILED DESCRIPTION

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] The present invention provides a technical solution: a finger vein feature extraction method based on deep metric learning, comprising the following steps:

[0024] Step 1, preprocess the original finger vein image and extract the region of interest of the image;

[0025] Step 2, input the triplet (a, p, n) composed of positive and negative sample pairs into the feature extraction network to obtain a feature triplet;

[0026] Step 3, calculate the loss in the Euclidean metric space. Then, input the feature triplet into the classifier to obtain the class probability, and calculate the loss in the cosine metric space, and optimize them together through backpropagation.

[0027] Please refer to Figure 1 , as an embodiment of the present invention:

[0028] Regarding Step 1.

[0029] Preprocess the original finger vein image and extract the region of interest (Region Of Interest) of the image. Secondly, input the triplet (a, p, n) composed of positive and negative sample pairs into the feature extraction network to obtain a feature triplet, and calculate the loss in the Euclidean metric space. In the figure, ||a - p|| < ||a - n|| represents the target that the distance between the positive sample pair is less than the distance between the negative sample pair. Then, input the feature triplet into the classifier to obtain the class probability, and calculate the loss in the cosine metric space. In the figure, cos(a, p) < cos(a, n) represents the target that the angle between the positive sample pair is less than the angle between the negative sample pair, and optimize them together through backpropagation. Doing so can make the network be doubly constrained by distance and angle, ensuring the effective learning of the network.

[0030] Such as Figure 2 , Figure 2In the figure, (a) original image; (b) gradient in X direction; (c) gradient in Y direction; (d) finger boundary. In order to ensure the consistency of data distribution and remove the interference of external factors, the region of interest of the finger needs to be extracted before finger vein recognition. The interference comes from two aspects: 1) The finger vein imaging device collects images under near-infrared light. The collected images include the finger area and the background area. There are many noises in these background areas and the edge areas of the fingers, which will affect the recognition of finger veins; 2) Although the finger vein collection device has a groove to limit the range of movement of the finger, due to the differences in thickness, length, etc. of the user's fingers, it is inevitable that there will be problems such as front-to-back offset, left-right offset and axial posture change of the finger. Therefore, before recognition, it is usually necessary to preprocess the collected image, the most important of which is the extraction of the region of interest.

[0031] G=G y -G x (3-1)

[0032]

[0033] Extracting the region of interest of the finger is actually to obtain a relative part of the finger of different users. It is similar to face alignment in face recognition, which is to obtain the head area in the image and then compare the similarity. The region of interest of the finger should contain as many vein patterns as possible, exclude the interference of irrelevant information, and ensure that the image has rich information, clear patterns and uniform size. The idea of ​​extracting the region of interest of the finger is: first obtain the upper and lower boundaries of the finger to determine the position of the finger, then calculate the central axis of the finger according to the boundary of the finger, and inversely calculate the angle of the plane offset of the finger to correct the rotation of the finger, and finally find the position of the second joint of the finger and intercept the region of interest. In order to obtain the upper and lower boundaries of the finger, the Sobel operator is used for edge detection. The edge of the object will show a large gradient value change in the image. The Sobel operator obtains the gradient value of the image through the convolution kernel operation (3-2), and the image gradient value in each direction is weighted to obtain the edge of the object. In order to avoid the interference of noise in the image, a Gaussian low-pass filter is required before edge detection to obtain the image A after filtering the noise. The definition of the Sobel operator in the x-direction and y-direction is as follows:

[0034] In finger vein images, because of the interference of natural light when collecting images, the natural aperture also has obvious edges in the image, such as Figure 2 (a). We hope to obtain only the edge information of the finger and avoid some irrelevant environmental interference. Considering that the upper and lower boundaries of the finger are parallel to the x-axis, the image gradient in the y-axis direction minus the gradient in the x-axis direction (3-1) is used to obtain the image without aperture interference. Since other environmental noise points are relatively scattered, a small number of connected areas are filtered through the union-find algorithm to obtain the upper and lower boundaries of the finger, such as Figure 2 (d).

[0035] The rotation of a finger at any angle in the plane may cause a finger to have different postures when collected multiple times, so it is necessary to perform rotation correction on the finger. After obtaining the upper and lower boundaries of the finger, the position of the midpoint of the finger can be calculated based on the boundary coordinates. However, the finger boundary is discontinuous after filtering the noise, so the midpoint of the finger is fitted by the least squares method to obtain a straight line, which is the central axis of the finger. Finally, based on the slope k of the finger midline, the angle of finger rotation θ = arctank is inversely calculated, and the original image is rotated -θ to complete the correction, such as Figure 3 As shown, (a) fitting centerline; (b) rotating original image.

[0036] See also Figure 4 , Figure 4 In the figure, (a) the inscribed rectangle is intercepted; (b) the grayscale difference of the central axis; (c) the ROI is extracted. In order to remove the interference of the background such as the device, it is necessary to intercept the inscribed rectangle of the upper and lower boundaries of the finger. Figure 4 (a). However, radial translation of the finger will still lead to large differences in the extracted image area, so it is necessary to further locate the stable area of ​​the finger image. Finger vein image acquisition usually includes the joints between two phalanges, and there is synovial fluid in the finger joint cavity, which will appear as a brighter area in the finger vein image. Therefore, it is hoped to locate the lateral coordinates of the finger joint cavity by changing the grayscale difference of the central axis of the finger image. Figure 4 (b), we can find that the finger joint cavity presents two peaks on the corresponding horizontal axis. Therefore, by determining the x-axis coordinates of the peaks, we can locate the area of ​​interest of the finger. Figure 4 (c).

[0037] For step 2.

[0038] We want to build a deep neural network, which can perform more complex feature extraction and achieve better results in finger vein recognition tasks. However, as the number of network layers increases, network performance will degrade (when the deep network begins to converge, the accuracy reaches saturation and then degrades rapidly, and the training error increases). ResNet solves the network degradation problem through residual learning.

[0039] like Figure 5As shown in the figure, x represents the network input and F(x) represents the network output. Assuming H(x) is an ideal mapping function, the traditional network is to let F(x) approximate H(x), while the goal of residual learning is to let F(x) approximate H(x)-x. The idea of ​​residual learning is to establish an identity mapping between network input and output, and let F(x)+x=x, that is, F(x)=0, so that the deep parameters of the network can be easily satisfied after being initialized to zero. In addition, the gradient of the deep layer of the network can be directly transmitted back to the shallow layer through shortcut connections, which can avoid the problem of gradient disappearance.

[0040] Learning the salient features of finger vein images is the key to improving the network's discriminative ability. Convolutional networks construct image features by fusing the spatial and channel information of each layer. These image features are not relatively isolated. This paper hopes to explicitly model the interdependencies between features.

[0041] like Figure 6 Based on ResNet34, the network is redesigned according to the finger vein scenario to obtain the finger vein feature extraction network FV-FE-Net (Finger Vein Feature Extraction Network), as shown in Table 3-1. ResNet34 contains a 7×7 convolution layer and four network layers with different numbers of residual modules stacked, and the output is average pooled and fully connected with 1000 dimensions. Since the finger vein recognition task focuses more on the details of the vein, the 7×7 convolution layer is replaced with a small 3×3 convolution kernel, so that the network will retain more detailed information during the training process and reduce the number of model parameters. Batch normalization is performed before each convolution to ensure that the distribution of each layer of data does not change too much and affect the feature extraction ability of the network. At the same time, the ReLU activation function is replaced with the Leaky ReLU activation function to avoid mean shift and make the network more stable. After the output of the basic network, an extended network needs to be constructed to further extract discriminative vein features. After the extended network passes the output of the basic network through the batch normalization layer, DropOut is used to avoid network overfitting, and then a 256-dimensional feature vector is obtained through the fully connected layer. In order to ensure that the feature distribution of the network output does not change too much, the output feature vector is L2 regularized, which can also ensure the accuracy of the similarity measurement.

[0042] For step 3.

[0043] The loss function plays an important role in the performance of neural networks. A reasonable loss function can constrain the network's search space and make the network's output closer to the optimization goal.

[0044] like Figure 7 , in order to make the features extracted by FV-FE-Net satisfy the characteristics of compactness within the class and dispersion between classes.

[0045] By making the feature vectors satisfy the constraint of ||a - p|| < ||a - n|| in the Euclidean metric space and cos(a, p) < cos(a, n) in the cosine metric space. This is because, in the next stage of finger vein recognition: feature matching and recognition, the Euclidean distance or cosine similarity is usually used as the discrimination criterion. Moreover, the network is trained in both metric spaces, which also helps it learn the representation ability of high-level features.

[0046] For the above goal, the objective form of the loss function is defined as Lfv = α·L1 + β·L2, where L1 is the distance loss in the Euclidean metric space, and the triplet loss is adopted. L2 is the angular loss in the cosine metric space, and arcface is adopted. Arcface is an improved version of the softmax loss, which adds a margin in radians to expand the decision boundary. In the formula, α and β are adjustment factors to prevent the loss gap between the two parts from being too large. The specific form of the loss function is as follows:

[0047]

[0048] The first term of the loss function is the triplet loss, which represents the sample distance metric in the Euclidean space. Among them, γ is the distance margin, f(·) represents the feature extraction function, xai represents the anchor sample, xpi represents the sample of the same class as the anchor, and xni represents the sample of a different class from the anchor. The difference between the second term and softmax is that the bias is set to zero, and the feature x and the weight w are normalized respectively. The dot product result of the two vectors is cosθ, where θ is the angle between the feature x and the weight w, s is the scaling factor, and m is the class margin. The triplet loss constrains the position of the samples in the Euclidean metric space, while arcface adjusts the samples in the cosine metric space. This method provides a richer metric standard, so that the network has better generalization ability.

[0049] An equation like Equation (3-3) seems to meet the goal, but there will be difficulties in convergence during actual training. Because in metric learning, if the distance between selected samples of the same class and the distance between samples of different classes are too small, it will cause a large oscillation amplitude in the network training process and the phenomenon that the network is difficult to converge. Taking the triplet loss as an example, Figure 8 is the process of a sample in one training iteration. This method only considers pulling the positive sample p closer to the anchor a and pushing the negative sample n away from the anchor a. This will cause the negative sample n to be pushed into the category where the anchor sample a is located during the training process. After this iteration is completed, the distance between the negative sample n and other samples in the category where the anchor sample a is located becomes smaller. This will not only cause difficulties in network convergence but also lead to a deterioration of the network's robustness. Therefore, it is necessary to constrain the update direction of the samples.

[0050] In order to explore the way to update samples, the triple loss is derived to obtain the following formula. a represents the embedding of the anchor point in the metric space, f p represents the embedding of positive samples of the same type as the anchor point in the metric space, f n Represents the embedding of negative samples of different classes from the anchor in the metric space.

[0051]

[0052] By taking the derivative, we can find that during the gradient descent process, the positive sample will be subjected to a force in the fa-fp direction, which pulls the positive sample toward the anchor point, and the negative sample will be subjected to a force in the fn-fa direction, which pushes the negative sample away from the anchor point, while the anchor point sample will be subjected to a force in the fp-fn direction, as shown in the following figure. Fig. 9 shown.

[0053] To avoid Figure 3-8 In the problem shown in the figure, we hope that the negative samples are as far away from the center of the category where the anchor point is located as possible. The midpoint between the anchor point a and the positive sample point p is called point c. In each batch, point c can be approximately considered as the center of the category where the anchor point is located, and f can be obtained. c =f a+ f n / 2. Therefore, the optimization goal becomes to make the negative sample points radially away from the approximate class center point c, that is, to make NC orthogonal to AP.

[0054]

[0055] It can be known that NC = f c -f n , and PA = f a -f p , so the above formula becomes

[0056]

[0057] Because ||f a ||=||f n ||=||f p ||=1, it is easy to know f c ⊥P A , that is, f c ·(f a -f p )=0, expand the denominator of the above formula, and we can get

[0058] f n ·f p -f n ·f a =0(3-7)

[0059] The above formula is the direction in which we want the sample to be updated, so by adding a direction constraint to the loss function, we can get

[0060]

[0061] Formula (3-8) adds a directional constraint regularization term to the joint loss function, where λ is a hyperparameter. The joint loss with directional constraints converges faster during training, and the resulting network accuracy is also higher.

[0062] As an embodiment of the present invention: In order to make a fair comparison with other finger vein recognition methods, a trial method proposed by other algorithms is adopted here. Half of the samples of all categories in the public data set are selected for training, and the remaining samples are used as test sets. Fig.10 The comparison results are shown in Figure 2. The proposed method achieves the best performance on the public dataset. This is because: the metric learning using sample pairing, to some extent, expands the data scale and provides data driving force for the deep network. In addition, this method can fully explore the relative relationship between data, which is conducive to the network learning richer feature expression.

[0063] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A finger vein feature extraction method based on deep metric learning, characterized in that: The following steps are involved: Step 1, preprocessing the original finger vein image to extract the region of interest of the image; Step 2: Pass the triplet (a, p, n) consisting of positive and negative sample pairs into the feature extraction network to obtain a feature triplet; Step 3, calculate the loss in the Euclidean metric space, then pass the feature triplet into the classifier to obtain the category probability, and calculate the loss in the cosine metric space, and optimize it together through back propagation. In step 3, in order to make the feature extraction network training subject to the constraints of the Euclidean metric space and the cosine metric space, a joint loss function is designed, and a regularization term with directional constraints is proposed. The specific formula is as follows: In the formula, α and β are adjustment factors used to adjust the size of the loss in the Euclidean metric space and the cosine metric space; the first term of the loss function is the triple loss, which represents the sample distance metric in the Euclidean space, where γ is the distance interval, f(·) represents the feature extraction function, represents the anchor point sample, represents samples of the same type as the anchor point, represents samples of different classes from the anchor point; It is a regular term of directional constraint, which is used to constrain the direction of sample update. In the second term of the loss function, the sample feature x and weight w are normalized respectively, and the vector dot product of the two is cosθ, which represents the cosine distance between the two, where θ is the angle between the feature x and the weight w, s is the scaling factor, and m is the category interval.

2. The finger vein feature extraction method based on deep metric learning according to claim 1, characterized in that: In step 1, first obtain the upper and lower boundaries of the finger to determine the position of the finger. Secondly, calculate the central axis of the finger based on the boundaries of the finger, and inversely calculate the angle of the finger plane offset to correct the finger rotation. Finally, find the position of the second joint of the finger and intercept the area of ​​interest.

3. The finger vein feature extraction method based on deep metric learning according to claim 1, characterized in that: In step 2, the feature extraction network is based on ResNet34 and is improved according to the characteristics of finger vein images. Specifically, the 7×7 convolution kernel is replaced by a 3×3 small convolution kernel, so that the network can retain more local vein detail information during the feature extraction process, while reducing the number of model parameters. Batch normalization is performed before each convolution layer to ensure that the data distribution of each layer of the network does not change too much. The ReLU activation function is replaced by the Leaky ReLU activation function, and a channel attention mechanism is introduced after each residual structure block.