Hand multi-feature recognition system based on graph convolutional network

CN120032398APending Publication Date: 2025-05-23新疆维吾尔自治区科技项目服务中心(新疆维吾尔自治区对外科技交流服务中心) +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510127019.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-28
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional identity authentication technology has security risks such as physical loss, forgery and password forgetting, and cannot meet the needs of modern society for personal privacy protection.

Method used

A hand multi-feature recognition system based on graph convolution network is adopted, and multi-modal feature acquisition and pre-processing of fingerprint and finger vein is achieved by combining convolutional neural networks and traditional image enhancement algorithms to recognize and verify multi-features of the hand.

Benefits of technology

Provides a secure, accurate and convenient authentication solution, improving the security and reliability of identity authentication and reducing the risk of password forgetting and forging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032398A_ABST
    Figure CN120032398A_ABST
Patent Text Reader

Abstract

The invention discloses a hand multi-feature recognition system based on a graph convolutional network. The hand multi-feature recognition system comprises a new user registration module, a user identity verification module, an enhanced recognition module and a logout function module. According to the hand multi-feature recognition system based on the graph convolutional network, an enhanced recognition module is used for collection and preprocessing of finger multi-modal features, and the collection and preprocessing of the finger multi-modal features utilize collection equipment to collect fingerprints and finger veins respectively; the method comprises the steps of collecting a fingerprint image, performing denoising, binaryzation, normalization and region-of-interest enhancement on the collected image, combining a traditional feature extraction mode with an emerging convolutional neural network, and selecting the convolutional neural network to replace a manual design feature extraction algorithm in consideration of the complexity of fingerprint image preprocessing. And a traditional image enhancement algorithm is selected for the vein image with strong directivity. Therefore, the model identification speed is improved, that is, when the mode is converted, a new mode can be quickly adapted only by retraining network parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biometric technology, and in particular to a hand multi-feature recognition system based on a graph convolutional network. Background Art

[0002] At present, there are two main types of traditional identity authentication technologies. One is to rely on physical objects for personal identity authentication, such as keys, certificates, etc., but the physical objects for authentication may be lost due to personal negligence, or may be forged by criminals through technical means; the second is chip-based identity authentication technology, such as student cards and bank cards embedded with chips, which may cause the password to be forgotten or leaked due to personal negligence. These drawbacks of traditional identity recognition technology have brought many security risks to people's lives and cannot meet the current society's demand for personal privacy protection; at the same time, the emergence of biometric technology has improved the convenience and security of people's lives, and has therefore been widely used. With the continuous investment of countries around the world in the research of biometric technology, related research on biometric identity authentication technology has become more and more extensive. Biometric-based authentication technology can well solve the defects of traditional identification technology; at the same time, due to the uniqueness of biometrics, not easy to lose or forget, this technology has also been recognized by the society and has been widely used in our daily lives. For example: Alipay fingerprint recognition authentication, Alipay face recognition authentication, attendance system based on palm features, and mobile phone fingerprint unlocking, etc. Biometric recognition technology can be divided into single-modal recognition technology and multi-modal recognition technology according to the number of biometric features.

[0003] In the daily life of modern society, personal information is very important to everyone. The leakage of identity information often brings inestimable losses. Personal information generally includes basic identity information, account information, privacy information, network behavior information, etc. In real life, when criminals obtain someone's personal information, they can not only steal the individual's property through the Internet, causing property and reputation losses to the victim; but also spread the victim's private information through the Internet, causing irreparable harm to the victim's family, life and other aspects. Therefore, identity authentication technology is becoming more and more important in today's society. How to provide safe, reliable and efficient identity authentication protection for personal life has also become a very important social issue.

[0004] Although biometric recognition technology has been widely used in various fields such as financial payment, personal identity authentication and Internet of Things security, it cannot meet the requirements of social security due to the limitations of the characteristics of a single organism. Summary of the invention

[0005] The purpose of the present invention is to provide a hand multi-feature recognition system based on a graph convolutional network to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A hand multi-feature recognition system based on a graph convolutional network includes a new user registration module, a user identity authentication module, an enhanced recognition module and a logout function module, wherein the enhanced recognition module is used for the collection and preprocessing of finger multi-modal features, and the collection and preprocessing of finger multi-modal features utilizes a collection device to collect fingerprints and finger veins respectively, and performs denoising, binarization, normalization, and enhancement of the region of interest on the collected images, and the collection and preprocessing of finger multi-modal features include fingerprint region of interest interception and image enhancement, finger vein image enhancement, feature extraction, filtering, principal component analysis, histogram series fusion calculation of LGBP features, feature matching calculation, feature-level fusion calculation of multi-set discriminant correlation analysis, convolutional neural network multi-modal recognition, convolutional neural network multi-modal feature recognition, multi-modal feature recognition network and deep network and score fusion multi-modal biometric recognition;

[0008] Fingerprint region of interest interception and image enhancement, first of all, the region of interest ROI is intercepted from the collected image, and the fingerprint ROI is enhanced by the adaptive local contrast enhancement algorithm ACE. The essence of the ACE method is to use the unsharp mask and contrast gain function to achieve local contrast enhancement of the image. In the ACE algorithm, the gray value x(i, j) of a point is taken, and the local mean is calculated in a rectangular window with a side length of 2n+1 with the pixel point (i, j) as the center:

[0009]

[0010] Compute the local variance:

[0011]

[0012] Subtract the original image from the blurred image, retain the high-frequency part and multiply it by the amplification gain coefficient to get the enhanced pixel value:

[0013] f(i,j)=m x (i, j) + G (i, j) [x (k, l) - m x (i, j)] (3)

[0014] Where G(i,j) is the contrast gain function, the contrast gain function is CG, CG>1, the spatially adaptive CG function is proportional to the global mean and inversely proportional to the local mean square error:

[0015]

[0016] Finger vein image enhancement, using a multi-directional filtering finger vein image enhancement method. In the finger vein recognition process, due to the non-permeability of skin tissue, interference from fingerprints and other reasons, the quality of the image will directly affect the recognition accuracy. Traditional filtering methods do not take into account the obvious directionality of finger veins, so the enhancement effect is not good. Use a finger vein image enhancement method based on multi-directional filtering. Fully consider the direction of the finger veins, based on the direction of the finger veins, use 8 different directional filters to construct a fingerprint extraction filter. The size of the 8 filters is 5×5 to avoid missing fingerprints in any direction;

[0017] Taking pixel (i, j) as the center, the gray value of the pixel is I(i, j), the pixel value after filtering is Ie(i, j), and the pixel value after enhancement is:

[0018] I' t0 (i, j) = I e (i, j-2)+I e (i,j+2)-I e (i,j) (5)

[0019] Select the minimum value of the 8 directional filters and adjust the image grayscale range to [0,255]:

[0020]

[0021] Feature extraction refers to the feature extraction method based on statistical learning. In the feature extraction stage, the focus is on the feature extraction method based on statistical learning, using linear and nonlinear features to extract image information. Four methods, Gabor filtering, local binary pattern LBP, linear discriminant analysis LDA and principal component analysis PCA, are used to select texture and scale features. After discussing the experimental results, an algorithm with better comprehensive performance can be selected for application.

[0022] Since veins have unique ridge texture characteristics in images, the selection of parameters for Gabor filters should conform to the special texture characteristics of these two veins. Gabor filtering uses the Gabor filtering algorithm. Aiming at the unique ridge texture characteristics of vein images, the Gabor filter uses parameters that conform to the special texture of veins to construct an even-symmetric Gabor filter EGF in the spatial domain. EGF is defined as:

[0023]

[0024] Where k = γ / (2πσ 2 ), x θ =xcosθ+ysinθ,yθ =-xsinθ+ycosθ, θ represents the filter direction, f c represents the center frequency of the sine wave, σ and γ are the standard deviation and aspect ratio of the Gaussian curve respectively. Selecting appropriate Gabor filter parameters can well extract the texture features in the image;

[0025] As a multimodal biometric recognition system, compared with a single-modal recognition system, its input data volume will certainly increase exponentially. If different biometric features are simply and directly superimposed, it is easy to cause problems such as insufficient computer memory and dimensionality disaster. Therefore, dimensionality reduction is the key to feature extraction in the pattern recognition process. Selecting a suitable dimensionality reduction algorithm can enable the system to extract more significant features from the original input image, thereby reducing the amount of data during classification and enhancing the recognition system's ability to distinguish between modalities. During principal component analysis, a dimensionality reduction algorithm is used to extract significant features from the original input image, reduce the amount of data during classification, and enhance the recognition system's ability to distinguish between modalities. The dimensionality reduction algorithm uses the principal component analysis PCA algorithm, which uses a linear projection matrix to project high-dimensional data into a low-dimensional space. The projection principle is to ensure that the dimension of the low-dimensional space is small enough while maximizing the variance of the reduced-dimensional data in the projected direction, so as to maximize the retention of the characteristics of the original data. The PCA algorithm flow is as follows:

[0026] First, center all samples:

[0027] Compute the covariance matrix:

[0028] Perform eigenvalue decomposition on the covariance matrix: S·p i =λ i ·p i ;

[0029] According to the energy percentage of the principal component Select the eigenvectors corresponding to the first i eigenvalues ​​to form a matrix W, where:

[0030] According to the above process, the matrix W is calculated for feature extraction: Z = W T x;

[0031] The histogram serial fusion calculation of LGBP features adopts the histogram serial fusion algorithm based on LGBP features to perform feature encoding of LGBP, combining multi-scale and multi-directional Gabor filtering with local binary pattern LBP to effectively describe the global direction information and local features of finger veins, thereby reducing the influence of changes in external illumination and finger posture on finger feature fusion. LBP adopts a rotationally invariant circular operator. The LBP operator replaces the square neighborhood with a circular neighborhood, allowing P pixels in a circular neighborhood with a radius of R, R = 1.0, P = 8, and taking the central pixel as the center, the LBP features in the circular neighborhood are rotated at 0°, 45°, 90°, 135°, 180°, 225°, 270°, and 315°, respectively, to calculate 8 LBP feature encoding values, and take the smallest value of these LBP feature encoding values ​​as the final LBP feature encoding value of the central pixel. The rotationally invariant LBP formula is as follows:

[0032]

[0033] Among them, x c ,y c is the center pixel, i c is the gray value of the center pixel, i n is the gray value of adjacent pixels, and s is a sign function:

[0034] Histogram concatenation fusion based on LGBP features: According to the existing method ideas, LBP encoding is performed on the enhanced ROI image, and then the image is divided into N*N blocks, and the grayscale histogram of each block is calculated. Finally, the grayscale histograms of all blocks are connected end to end to form the feature histograms of fingerprints, finger veins and knuckle prints, which are then connected in series to form the final feature histogram;

[0035] Different fusion algorithms require appropriate matching algorithms to correspond to them. The feature matching calculation uses the feature matching algorithm. The feature histogram after the series fusion of the LGBP feature histogram uses the matching algorithm based on the histogram intersection method to match. The matching algorithm uses the histogram intersection method as the similarity measurement criterion. The formula is as follows:

[0036]

[0037] Among them, H 1 and H 2 are the two LGBP feature histograms to be matched, L is the feature histogram H 1 , H 2 The dimension of Sim(H 1 ,H 2 ) is the similarity of the two histograms, Sim(H 1 ,H2 ) The larger, H 1 and H 2 The higher the similarity, the more likely they are to match successfully. Set a similarity decision threshold th, Sim(H 1 , H 2 ) > th, then the test sample and the template sample belong to the same category; Sim(H 1 , H 2 ) < th, then they belong to different categories;

[0038] Multi-modal recognition of finger veins is performed using ImageNet and PalmNet. The encoding fusion method of local Gabor binary pattern (LGBP) features is used to achieve the fusion of fingerprint and finger vein multi-modal features. The encoding of local Gabor binary pattern (LGBP) features and the feature extraction method of convolutional neural network are adopted, and a spatial pyramid pooling layer is added to the network to solve the problem that the input image size in the convolutional neural network must be fixed due to the different sizes of images in different fingerprint databases;

[0039] The feature-level fusion calculation of multi-set discriminant correlation analysis adopts the feature-level fusion algorithm based on multi-set discriminant correlation analysis. The two simplest and typical feature-level fusion (feature-level fusion) methods are serial feature fusion and parallel feature fusion: serial feature fusion is equivalent to the superposition of multiple feature vectors along the vertical axis, and parallel feature fusion refers to the horizontal superposition of multiple groups of feature vectors. However, simply superposing in any direction will cause redundancy of feature information, resulting in an increase in computational complexity and a decrease in the specificity of feature vectors at the same time. Since different feature vectors come from different biometric modalities, in order to solve the problem of incompatibility between features, some other feature-level fusion strategies are needed to help achieve it. Researchers have conducted a large number of studies on feature fusion using canonical correlation analysis to explore the internal relationship between two sets of variables. Canonical correlation analysis (CCA) is used to help with two types of feature-level fusion. CCA is an analysis algorithm that uses the correlation relationship between comprehensive variable pairs to reflect the overall correlation between two sets of indicators. In the problem of feature fusion, when faced with two sets of feature vector data and hoping to study the relationship between these two sets of feature vectors, canonical correlation analysis will be used. The algorithm is as follows:

[0040] The input matrices are X = [x 1 , …, x n , Y = [y 1 , …, y n , where the within-class and between-class variance matrices of X and Y are denoted as S xx ∈R p×q and S yy ∈R q×q The difference matrix is denoted as Sxy ∈R p×q , and there exists For the overall variance matrix S, which contains information within and between classes, the formula is:

[0041]

[0042] According to the formula of correlation coefficient:

[0043]

[0044] in,

[0045] And there is

[0046] The purpose of canonical correlation analysis is to solve the following formula:

[0047]

[0048] Calculate a pair of projection matrices W X and W y , and then use the projection matrix to get the new eigenvector and

[0049] As a common linear subspace method, CCA has the following disadvantages: when the relationship between two sets of features is a complex nonlinear correlation, it is difficult for CCA to accurately reflect the relationship between the two sets of feature vectors. In addition, as an unsupervised feature extraction method, CCA does not take into account the category information of the sample.

[0050] Discriminant correlation analysis DCA is set on the basis of CCA to increase the category information of training samples by improving the objective function, so as to solve the problem when the relationship between two sets of features is a complex nonlinear correlation relationship;

[0051] The DCA algorithm is as follows: Input vector, X = [x 1 ,…,x n ],Y=[y 1 ,…,y n ], the intra-class variance matrices of X and Y are still recorded as and. Different from the objective function of CCA, the objective function of DCA adds a diagonal matrix A, which contains the label information of the classification category:

[0052]

[0053] After obtaining the projection matrix, the two sets of vectors are fused in serial and parallel modes to form a new feature vector, as shown below:

[0054]

[0055] Discriminant correlation analysis increases the category information of training samples by improving the objective function in canonical correlation analysis.

[0056] Convolutional neural network multimodal recognition, the classic convolutional neural network is mainly composed of basic units such as convolution layer, pooling layer, linear rectification layer (Relu layer). The original image is input into the convolutional neural network composed of convolution layer, pooling layer, and linear rectification layer. Through the arrangement and combination of convolution layer, pooling layer, and linear rectification layer, the feature map containing the semantic features of the image is gradually extracted. After the feature map is input into the classifier, the predicted probability value of the classification is obtained;

[0057] The pooling layer is also an important component of the convolutional neural network. The advantage of the pooling layer is that it does not need to update the weight parameters, and at the same time it can flexibly reduce the size of the previous level feature map, thereby reducing the multiplication operation of the next level of convolution;

[0058] The linear rectifier layer uses an activation function. The activation function is another very important part of the neural network. It is used to perform nonlinear mapping on the output features of the convolutional layer after the convolutional layer, thereby enhancing the network's ability to express features. The activation function selects the linear rectifier function ReLU, and the formula is: ReLU(x) = max(0,x), indicating that everything with an x ​​value less than zero is mapped to a y value of 0, but everything with an x ​​value greater than zero is mapped to itself;

[0059] Its differential equation is:

[0060]

[0061] The advantage of the ReLU function is that it can avoid the problem of gradient disappearance, because in the process of updating weights, its update rule is:

[0062]

[0063] Among them, w (L) For all weights in the last layer L, when partial derivatives appear When the number of weights and biases is very small, the update of many weights and biases is very small, that is, the gradient vanishing problem occurs. When the ReLU activation function is used, its differential equation shows that its result is either 0 or 1, also known as unilateral inhibition, which introduces good sparsity into the neural network and alleviates the problem of gradient vanishing during training. While solving the problem of gradient vanishing, it also improves the efficiency in terms of time and space complexity.

[0064] Softmax is a linear multi-classification model whose function is to convert the scores of each category into reasonable probability values;

[0065]

[0066] Where W, b are model parameters, K is the total number of categories, y is the probability of predicting that the sample belongs to category i. The higher the probability, the greater the possibility of belonging to the category. The sum of all probabilities is normalized to 1.

[0067] The multimodal feature recognition of convolutional neural networks uses the Inception-ResNet module. Convolutional neural networks containing residual connection ResNet and Inception modules are selected for fingerprint feature recognition. Since neural networks are not sensitive to noise, the parameters of many early classic models are actually very redundant. Therefore, early deep learning models have high requirements for computer computing performance and memory, and general equipment is difficult to meet the above hardware requirements. In order to solve this problem, researchers began to study the use of efficient network structure design to reduce the amount of calculation. For example, GoogleNet uses the Inception module instead of simply stacking network layers to reduce the amount of calculation. ResNet has also achieved excellent image recognition results by introducing residual connections. On the basis of the above, a convolutional neural network including residual connection and Inception module is selected for fingerprint feature recognition. The characteristic of residual connection is that a copy of the data of the previous layer of network is added to the output result and then input into the next layer of network; the feature map of the original image is superimposed with the output, so as to help the training error gradually decrease as the depth of the network increases. Through residual connection, the neural network changes from fitting the output F(x) to fitting the residual F(x)-x. The residual structure is easier to learn than the original function and is more suitable for deep model iteration. At present, most of the deep convolutional neural networks have a very obvious feature of deepening the network by stacking multiple residual blocks, so there will be no serious overfitting when training very deep networks. When designing the network structure, learn the advantages of the deep residual network ResNet and use the residual structure for jump transmission in the middle layer of the network. On the one hand, it is conducive to the downward propagation of gradients, so that the network can still directly use more original information after multiple convolution operations, and at the same time, it is conducive to better upward transmission of gradients.

[0068] The characteristic of the Inception-ResNet module lies first in the Inception structure. The ordinary single node between two activation functions is expanded into a new neural network, also known as a "network in network". It uses multiple convolution kernels of different sizes to obtain receptive fields of different scales, and then fuses features of different scales. As the number of network layers increases, the extracted features will become more abstract. Therefore, the Inception structure uses this dense network structure to obtain a convolutional visual network that is closer to reality.

[0069] The Inception-ResNet module adds the classic residual connection structure in ResNet based on the original Inception module, which can retain the original information of the previous item and is more suitable for deep model iteration. The Relu activation function is connected after the 1x1 convolution. By introducing nonlinear excitation, the network's ability to express nonlinear functions is improved;

[0070] The improved Inception module reduces the number of channels through 1×1 convolution, gathers information and then calculates it. While effectively utilizing computing power, it makes the model nonlinear and improves cross-channel information integration and information interaction. At the end of each module, several branches are finally merged through aggregation operations and aggregated on the last dimension of the output channel. The Inception module contains three convolutions of different sizes and one maximum pooling, which increases the adaptability of the network to different scales and increases the width of the network, avoiding the problem of training gradient dispersion due to the network being too deep.

[0071] Multimodal feature recognition network, the Inception series of networks originated from the classic image classification network GoogLeNet, proposed by Google, and is a convolutional neural network with excellent network depth and width. Its characteristic is that it designs a network with excellent topological structure, so it is often used as a feature processor in the field of computer vision. The Inception-ResNet-V2 network structure is selected for fingerprint feature recognition. Its network structure uses different Inception-ResNet modules one by one from shallow to deep on the main path. The first module on the main path uses the Stem module. The Inception-ResNet modules are named Inception-ResNet-A / B / C respectively. Each Inception-ResNet module performs multiple convolution operations or pooling operations on the input image in parallel, and splices all output results into a very deep feature map. In the process of extracting image features, convolution kernels of 1*1, 3*3 and 7*7 scales are used to extract image features at the same time, and then the outputs of these convolutions are stacked and passed to the next layer of the network. The nonlinear activation function ReLu is used, and then the fully connected layer in the traditional convolutional neural network is replaced by the global average pooling layer GAP to remove the spatial relationship in the image features and obtain features with high-level semantic information.

[0072] Adding residual connection structure makes the network structure more context-aware, which is beneficial for the network to express the deep features of fingerprints;

[0073] Use smaller 3x3 convolutions to make up for the smaller receptive field by deepening the network; decompose the large convolution kernel into a combination of a series of small convolution kernels. The combination of small convolution kernels improves the adaptability of the network to different scales and improves the generalization ability of the network.

[0074] Applying 1x1 convolution before the large convolution kernel can flexibly reduce the number of featuremap channels, reduce the amount of calculation by reducing the number of parameters and the number of multiplication operations, and ultimately achieve the purpose of improving fingerprint recognition efficiency;

[0075] The role of the loss function is to achieve network convergence by calculating the difference between the predicted value and the true value in the training set, with the goal of minimizing the loss function. The smaller the value of the loss function, the smaller the difference between the predicted value and the true value, and the more accurate the model's prediction. Under the supervision of the loss function, the convolutional neural network can learn better image features and speech information. Depending on the type of task, choosing an appropriate loss function can shorten the network training time, accelerate convergence, and improve recognition accuracy.

[0076] Combine center loss and softmax loss as a joint loss function to train the convolutional neural network:

[0077] Center loss function constrains the distance between the category centers of sample features:

[0078]

[0079] where x i is the eigenvalue, c yi is the center point of all sample features of the category corresponding to sample i, and N is the total number of samples;

[0080] Softmax loss cross entropy loss function constrains the distance between the actual output and the expected output:

[0081]

[0082] Among them, yi represents the prediction result, N and K are the total number of samples and the total number of categories respectively. The smaller the value of Softmax loss is, the closer the predicted value is to the true value, and the more accurate the prediction result is.

[0083] Joint loss function to train convolutional neural network:

[0084] L=L s +λL c (twenty three)

[0085] Among them, λ is the balance parameter of center loss and softmax loss. The larger λ is, the higher the weight of center loss is. In order to avoid the problem of reducing the intra-class distance and increasing the difficulty of model optimization during the calculation of the joint loss function, λ=0.01 is selected;

[0086] Multimodal biometric recognition with deep network and score fusion, based on the finger multimodal feature recognition network, uses convolutional neural network to automatically classify fingerprint images, adopts the fingerprint and finger vein multimodal recognition score fusion method and fingerprint and finger vein multimodal recognition feature fusion method, uses the deep network Inception-ResNet-V2, after the softmax classifier, obtains the comprehensive matching score through score-level fusion to obtain the final recognition result;

[0087] The characteristic of the model of the present invention is that it combines the traditional feature extraction method with the emerging convolutional neural network. Considering the complexity of fingerprint image preprocessing, the convolutional neural network is selected to replace the manually designed feature extraction algorithm, and the traditional image enhancement algorithm is selected for the vein image with strong directionality. The appropriate algorithm is selected according to the characteristics of the modality, which can combine the advantages of the two, saving the time of manually extracting fingerprint features, and does not require complex preprocessing of fingerprint images, thereby improving the model recognition speed, and enhancing the generalization ability of the model, that is, when the modality is converted, only the network parameters need to be retrained to quickly adapt to the new modality.

[0088] There are several commonly used score-level fusion methods: concatenation, weighted sum, mean fusion, and maximum fusion. For the simple concatenation fusion method, if the modality increases or the input category increases, it will cause dimensionality explosion, insufficient computing memory, and other problems. For the weighted sum fusion method, the choice of weights will directly affect the recognition accuracy of the system. The fingerprint feature recognition network uses a softmax classifier. In order to ensure the superposition of prediction scores of different modalities, the traditional method of vein feature recognition is also converted into a reasonable probability value (the value is between [0,1] and the sum of the probabilities is 1) through a softmax regression model after obtaining the matching score. For example, a sample may belong to N categories. Assume that the matching score score = {d1, d2,..., dN}. After the softmax function, the corresponding category probability value is:

[0089]

[0090] Then, different score fusion models are selected to fuse the category probabilities to obtain the final prediction result. Next, these three different score-level fusion models will be introduced respectively.

[0091] Mean Fusion

[0092] When the classification capabilities of different modalities are similar, the mean model can be used for fusion. Assuming that the results of M softmax classifiers are {h1,h2,...,hM}, the prediction result after fusion is:

[0093]

[0094] Weighted sum fusion

[0095] When different classifiers have different effects on the final prediction results, a weighted sum model is selected to assign different weights to different modality classifiers according to their prediction accuracy, where the results of the M softmax classifiers are {h1,h2,...,hM}:

[0096]

[0097] Maximum Fusion

[0098] When the prediction results of different modality classifiers are different, the maximum model can be used for fusion. Assume that the results of M softmax classifiers are {h1,h2,...,hM}, and select the category with the largest prediction probability as the final prediction result:

[0099] H(x)=max(h i )(29)

[0100] The deep network is combined with the multimodal recognition system of feature fusion and fractional fusion to effectively identify the multimodal features of fingers.

[0101] As a further solution of the present invention: the convolutional neural network is composed of multiple convolutional layers stacked together. This is because the features learned by a single-layer convolutional layer are local and limited. By continuously superimposing multiple layers of convolutional layers, more comprehensive and deeper image features can be learned. The convolution operation is to extract matrix blocks from the input feature matrix according to certain rules, and perform the same transformation operation on these matrix blocks to generate a new output feature matrix. The convolution kernel is equivalent to a sliding window function acting on the matrix, and the convolution kernel is multiplied one by one with the corresponding block matrix elements and then summed.

[0102] As a further solution of the present invention: the pooling operation of the pooling layer compresses the input feature matrix, reduces the size of the feature matrix by downsampling, and thereby simplifies the calculation complexity. The pooling operation is divided into average pooling and maximum pooling. Average pooling is to take the regional mean, thereby retaining the overall data characteristics, and maximum pooling is to take the regional maximum value, thereby retaining the texture characteristics of the matrix.

[0103] As a further solution of the present invention: the characteristic of the Inception-ResNet module lies first in the Inception structure, which expands the ordinary single node between the two activation functions into a new neural network, adopts multiple convolution kernels of different sizes, obtains receptive fields of different scales, and then performs feature fusion of different scales. As the number of network layers increases, the extracted features will become more abstract. Therefore, the Inception structure utilizes this dense network structure to obtain a convolutional visual network that is closer to reality.

[0104] As a further solution of the present invention: the simultaneous 1x1 convolution is also a way to fuse channel information.

[0105] As a further solution of the present invention: the basic network selects Inception-ResNet-V2, and the Inception-ResNet-V2 network on the ImageNet dataset is used for pre-training. Since the ImageNet dataset and the fingerprint dataset are not similar, the low-level features learned by the previous convolutional layer are retained, and the terminal layer is replaced with a customized output layer, and the terminal layer is retrained. In terms of initial parameter selection, the input batch size of the convolutional neural network is 28, and the learning rate is 0.01 at the beginning. After training 80 epochs, it is reduced to 0.0001 to fine-tune the network. When fine-tuning the pre-trained network, the learning rate is as small as possible to avoid too much impact on the original training weights. The probability of random inactivation of neurons is 0.8, and the optimizer selects adam. By adjusting the network optimizer and parameters, it is observed whether it is overfitting, and the optimal training model is selected.

[0106] As a further solution of the present invention: the Stem module stacks 1x1, 3x3, 1x7, 7x1 conv and 3x3 pooling together, adds convolution parameters to the network according to the width of the network and the adaptability of the network to the scale. After the network adds the convolution parameters, it independently selects the required convolution filter by changing the weights, introduces a special reduction block to help reduce the size of the feature map from 35x35 to 17x17, and from 17x17 to 8x8, solves the problem of too high dimension of feature data, selects two 3*3 convolutions instead of large convolution kernel convolution, reduces the number of parameters while accelerating the calculation, and uses 1*3 and 3*1 convolution kernels instead of 3*3 convolution kernels in the middle layer of the network structure.

[0107] As a further solution of the present invention: the new user registration module is used to provide a new user registration interface, requiring the new user to input basic information and biometric data (such as palm prints, palm veins and finger veins) to ensure the accuracy and reliability of the registration process, and store the information and biometric data of the newly registered user, and update the user database; the core function of the user identity authentication module is to identify the user, by loading the user's biometric data information, and calling the trained model to compare and identify the input information, and give a verification result; the enhanced recognition module is used to provide a data enhancement algorithm to increase the diversity and reliability of biometric data. It supports image preprocessing and enhancement to improve the accuracy of biometric recognition. The logout function module allows the user to actively request to log out of his account and related information, and verifies and confirms the user's logout request. After confirming the user's logout, the user's personal information and biometric data in the system are deleted, including the corresponding records in the database.

[0108] Compared with the prior art, the present invention has the following beneficial effects:

[0109] 1. The present invention provides a safe, accurate and convenient identity authentication solution. The system is implemented based on a variety of technologies, such as image processing, pattern recognition and machine learning, etc. The maturity, accuracy and availability of related technologies are investigated to ensure that the implementation of the system is feasible. Considering the privacy issues of identity identification, appropriate security measures are taken to protect the security and privacy of users' personal information and biometric data.

[0110] 2. The characteristic of the model proposed in the present invention is that it combines the traditional feature extraction method with the emerging convolutional neural network. Considering the complexity of fingerprint image preprocessing, the convolutional neural network is selected instead of the manually designed feature extraction algorithm, and the traditional image enhancement algorithm is selected for the vein image with strong directionality. Selecting a suitable algorithm according to the characteristics of the modality can combine the advantages of both, which not only saves the time of manually extracting fingerprint features, but also does not require complex preprocessing of fingerprint images, thereby improving the model recognition speed, and enhancing the generalization ability of the model, that is, when the modality is converted, only the network parameters need to be retrained to quickly adapt to the new modality. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] Figure 1 A diagram showing the direction filter a of the hand multi-feature recognition system based on graph convolutional network.

[0112] Figure 2 Figure 2 shows the directional filter b of the hand multi-feature recognition system based on graph convolutional network.

[0113] Figure 3 Figure 2 shows the directional filter c of the hand multi-feature recognition system based on graph convolutional network.

[0114] Figure 4 Figure 2 shows the directional filter d of the hand multi-feature recognition system based on graph convolutional network.

[0115] Figure 5 Figure 2 shows the directional filter e of the hand multi-feature recognition system based on graph convolutional network.

[0116] Figure 6 Figure 2 shows the directional filter f for the multi-feature hand recognition system based on graph convolutional network.

[0117] Figure 7 Display diagram of the direction filter g for the hand multi-feature recognition system based on graph convolutional network

[0118] Figure 8 A diagram showing the direction filter h of the hand multi-feature recognition system based on graph convolutional network.

[0119] Fig. 9 This is the first demonstration of the ridge texture finger vein in the hand multi-feature recognition system based on graph convolutional network.

[0120] Fig.10 This is the second demonstration image of the ridge texture finger vein in the hand multi-feature recognition system based on graph convolutional network.

[0121] Fig.11 This is the third demonstration diagram of the ridge texture finger vein in the hand multi-feature recognition system based on graph convolutional network.

[0122] Fig.12 A diagram showing the circular LBP operator in the hand multi-feature recognition system based on graph convolutional networks.

[0123] Fig.13 Schematic diagram of rotation-invariant LBP in the multi-feature hand recognition system based on graph convolutional network.

[0124] Fig.14 This is the network structure diagram of pyramid pooling in the hand multi-feature recognition system based on graph convolutional network.

[0125] Fig.15 Schematic diagram of the convolution calculation process in the hand multi-feature recognition system based on graph convolutional network.

[0126] Fig.16 Schematic diagram of the pooling layer input in the hand multi-feature recognition system based on graph convolutional network

[0127] Fig.17 Figure 2 is a diagram of the residual connection structure in the hand multi-feature recognition system based on graph convolutional network.

[0128] Fig.18 This diagram shows the Inception-ResNet-A module in the multi-feature hand recognition system based on graph convolutional networks.

[0129] Fig.19 A diagram showing the Inception-ResNet-B module in the multi-feature hand recognition system based on graph convolutional networks.

[0130] Fig. 20 This diagram shows the Inception-ResNet-C module in the hand multi-feature recognition system based on graph convolutional network.

[0131] Fig.21 This is the network structure diagram of fingerprint recognition in the hand multi-feature recognition system based on graph convolutional network.

[0132] Fig. 22 A diagram showing the stem module in the hand multi-feature recognition system based on graph convolutional networks.

[0133] Fig.23This diagram shows the Reduction module in the multi-feature hand recognition system based on graph convolutional networks.

[0134] Fig.24 This diagram shows the structure of fingerprint and vein multimodal recognition score fusion in the hand multi-feature recognition system based on graph convolutional network.

[0135] Fig.25 This diagram shows the fingerprint and vein multimodal recognition feature fusion structure in the hand multi-feature recognition system based on graph convolutional network. DETAILED DESCRIPTION

[0136] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0137] See also Figures 1 to 25 In an embodiment of the present invention, a hand multi-feature recognition system based on a graph convolutional network includes a new user registration module, a user identity authentication module, an enhanced recognition module and a logout function module. The enhanced recognition module is used for collecting and preprocessing finger multimodal features. The collection and preprocessing of finger multimodal features utilizes a collection device to collect fingerprints and finger veins respectively, and performs denoising, binarization, normalization, and enhancement of the region of interest on the collected images. The collection and preprocessing of finger multimodal features includes fingerprint region of interest interception and image enhancement, finger vein image enhancement, feature extraction, filtering, principal component analysis, histogram series fusion calculation of LGBP features, feature matching calculation, feature-level fusion calculation of multi-set discriminant correlation analysis, convolutional neural network multimodal recognition, convolutional neural network multimodal feature recognition, multimodal feature recognition of multimodal feature recognition network and deep network and score fusion of multimodal biometric feature recognition;

[0138] Fingerprint region of interest interception and image enhancement, first of all, the region of interest ROI is intercepted from the collected image, and the fingerprint ROI is enhanced by the adaptive local contrast enhancement algorithm ACE. The essence of the ACE method is to use the unsharp mask and contrast gain function to achieve local contrast enhancement of the image. In the ACE algorithm, the gray value x(i, j) of a point is taken, and the local mean is calculated in a rectangular window with a side length of 2n+1 with the pixel point (i, j) as the center:

[0139]

[0140] Compute the local variance:

[0141]

[0142] Subtract the original image from the blurred image, retain the high-frequency part and multiply it by the amplification gain coefficient to get the enhanced pixel value:

[0143] f(i,j)=m x (i,j)+G(i,j)[x(k,l)-m x (i,j)] (3)

[0144] Where G(i,j) is the contrast gain function, the contrast gain function is CG, CG>1, the spatially adaptive CG function is proportional to the global mean and inversely proportional to the local mean square error:

[0145]

[0146] Finger vein image enhancement, using a multi-directional filtering finger vein image enhancement method. In the finger vein recognition process, due to the non-permeability of skin tissue, interference from fingerprints and other reasons, the quality of the image will directly affect the recognition accuracy. Traditional filtering methods do not take into account the obvious directionality of finger veins, so the enhancement effect is not good. Use a finger vein image enhancement method based on multi-directional filtering. Fully consider the direction of the finger veins, based on the direction of the finger veins, use 8 different directional filters to construct a fingerprint extraction filter. The size of the 8 filters is 5×5 to avoid missing fingerprints in any direction;

[0147] Taking pixel (i, j) as the center, the gray value of the pixel is I(i, j), the pixel value after filtering is Ie(i, j), and the pixel value after enhancement is:

[0148] I' t0 (i, j) = I e (i, j-2)+I e (i, j+2)-I e (i,j) (5)

[0149] Select the minimum value of the 8 directional filters and adjust the image grayscale range to [0,255]:

[0150]

[0151] Feature extraction refers to the feature extraction method based on statistical learning. In the feature extraction stage, the focus is on the feature extraction method based on statistical learning, using linear and nonlinear features to extract image information. Four methods, Gabor filtering, local binary pattern LBP, linear discriminant analysis LDA and principal component analysis PCA, are used to select texture and scale features. After discussing the experimental results, an algorithm with better comprehensive performance can be selected for application.

[0152] Since veins have unique ridge texture characteristics in images, the selection of parameters for Gabor filters should conform to the special texture characteristics of these two veins. Gabor filtering uses the Gabor filtering algorithm. Aiming at the unique ridge texture characteristics of vein images, the Gabor filter uses parameters that conform to the special texture of veins to construct an even-symmetric Gabor filter EGF in the spatial domain. EGF is defined as:

[0153]

[0154] Where k = γ / (2πσ 2 ), x θ =xcosθ+ysinθ,y θ =-xsinθ+ycosθ, θ represents the filter direction, f c represents the center frequency of the sine wave, σ and γ are the standard deviation and aspect ratio of the Gaussian curve respectively. Selecting appropriate Gabor filter parameters can well extract the texture features in the image;

[0155] As a multimodal biometric recognition system, compared with a single-modal recognition system, its input data volume will certainly increase exponentially. If different biometric features are simply and directly superimposed, it is easy to cause problems such as insufficient computer memory and dimensionality disaster. Therefore, dimensionality reduction is the key to feature extraction in the pattern recognition process. Selecting a suitable dimensionality reduction algorithm can enable the system to extract more significant features from the original input image, thereby reducing the amount of data during classification and enhancing the recognition system's ability to distinguish between modalities. During principal component analysis, a dimensionality reduction algorithm is used to extract significant features from the original input image, reduce the amount of data during classification, and enhance the recognition system's ability to distinguish between modalities. The dimensionality reduction algorithm uses the principal component analysis PCA algorithm, which uses a linear projection matrix to project high-dimensional data into a low-dimensional space. The projection principle is to ensure that the dimension of the low-dimensional space is small enough while maximizing the variance of the reduced-dimensional data in the projected direction, so as to maximize the retention of the characteristics of the original data. The PCA algorithm flow is as follows:

[0156] First, center all samples:

[0157] Compute the covariance matrix:

[0158] Perform eigenvalue decomposition on the covariance matrix: S·p i =λ i ·p i ;

[0159] According to the energy percentage of the principal component Select the eigenvectors corresponding to the first i eigenvalues ​​to form a matrix W, where:

[0160] According to the above process, the matrix W is calculated for feature extraction: Z = W T x;

[0161] The histogram serial fusion calculation of LGBP features adopts the histogram serial fusion algorithm based on LGBP features to perform feature encoding of LGBP, combining multi-scale and multi-directional Gabor filtering with local binary pattern LBP to effectively describe the global direction information and local features of finger veins, thereby reducing the influence of changes in external illumination and finger posture on finger feature fusion. LBP adopts a rotationally invariant circular operator. The LBP operator replaces the square neighborhood with a circular neighborhood, allowing P pixels in a circular neighborhood with a radius of R, R = 1.0, P = 8, and taking the central pixel as the center, the LBP features in the circular neighborhood are rotated at 0°, 45°, 90°, 135°, 180°, 225°, 270°, and 315°, respectively, to calculate 8 LBP feature encoding values, and take the smallest value of these LBP feature encoding values ​​as the final LBP feature encoding value of the central pixel. The rotationally invariant LBP formula is as follows:

[0162]

[0163] Among them, x c ,y c is the center pixel, i c is the gray value of the center pixel, i n is the gray value of adjacent pixels, and s is a sign function:

[0164] Histogram concatenation fusion based on LGBP features: According to the existing method ideas, LBP encoding is performed on the enhanced ROI image, and then the image is divided into N*N blocks, and the grayscale histogram of each block is calculated. Finally, the grayscale histograms of all blocks are connected end to end to form the feature histograms of fingerprints, finger veins and knuckle prints, which are then connected in series to form the final feature histogram;

[0165] Different fusion algorithms require appropriate matching algorithms to correspond to them. The feature matching calculation uses the feature matching algorithm. The feature histogram after the series fusion of the LGBP feature histogram uses the matching algorithm based on the histogram intersection method to match. The matching algorithm uses the histogram intersection method as the similarity measurement criterion. The formula is as follows:

[0166]

[0167] Among them, H 1and H 2 are two LGBP feature histograms to be matched. L is the dimension of the feature histogram H 1 and H 2 . Sim(H 1 , H 2 ) is the similarity between the two histograms. The larger Sim(H 1 , H 2 ) is, the higher the similarity between H 1 and H 2 . Set a similarity decision threshold th. If Sim(H 1 , H 2 ) > th, then the test sample and the template sample belong to the same category; if Sim(H 1 , H 2 ) < th, then they belong to different categories;

[0168] Use ImageNet and PalmNet for multi-modal recognition of finger veins. Use the encoding fusion method of local Gabor binary pattern (LGBP) features to achieve the fusion of fingerprint and finger vein multi-modal features. Adopt the encoding of local Gabor binary pattern (LGBP) features and the feature extraction method of convolutional neural network, and then add a spatial pyramid pooling layer to the network to solve the problem that the input image size in the convolutional neural network must be fixed due to the different sizes of images in different fingerprint databases;

[0169] The feature-level fusion calculation of multi-set discriminant correlation analysis adopts the feature-level fusion algorithm based on multi-set discriminant correlation analysis. The two simplest and typical feature-level fusion (feature-level fusion) methods are serial feature fusion and parallel feature fusion respectively: serial feature fusion is equivalent to the superposition of multiple feature vectors along the vertical axis, and parallel feature fusion refers to the horizontal superposition of multiple groups of feature vectors. However, simply superposing in any direction will cause redundancy of feature information, and the result is that while increasing the computational complexity, the specificity of the feature vector is also reduced. Since different feature vectors come from different biometric modalities, in order to solve the problem of incompatibility between features, some other feature-level fusion strategies are needed to help achieve this. Researchers have conducted a large number of studies on feature fusion using canonical correlation analysis to explore the internal relationship between two sets of variables. Use canonical correlation analysis (CCA) to help with two types of feature-level fusion. CCA is an analysis algorithm that uses the correlation relationship between comprehensive variable pairs to reflect the overall correlation between two sets of indicators. In the problem of feature fusion, when faced with two sets of feature vector data and hoping to study the relationship between these two sets of feature vectors, canonical correlation analysis will be used. The algorithm is as follows:

[0170] The input matrix is X = [x 1,…,x n ],Y=[y 1 ,…,y n ], where the intra-class and inter-class variance matrices of X and Y are denoted by S xx ∈R p×q and S yy ∈R q×q The difference matrix is ​​denoted as S xy ∈R p×q , and there exists For the overall variance matrix S, which contains information within and between classes, the formula is:

[0171]

[0172] According to the formula of correlation coefficient:

[0173]

[0174] in,

[0175] And there is

[0176] The purpose of canonical correlation analysis is to solve the following formula:

[0177]

[0178] Calculate a pair of projection matrices W X and W y , and then use the projection matrix to get the new eigenvector and

[0179] As a common linear subspace method, CCA has the following disadvantages: when the relationship between two sets of features is a complex nonlinear correlation, it is difficult for CCA to accurately reflect the relationship between the two sets of feature vectors. In addition, as an unsupervised feature extraction method, CCA does not take into account the category information of the sample.

[0180] Discriminant correlation analysis DCA is set on the basis of CCA to increase the category information of training samples by improving the objective function, so as to solve the problem when the relationship between two sets of features is a complex nonlinear correlation relationship;

[0181] The DCA algorithm is as follows: Input vector, X = [x 1 ,…,x n ],Y=[y 1 ,…,y n], the intra-class variance matrices of X and Y are still recorded as and. Different from the objective function of CCA, the objective function of DCA adds a diagonal matrix A, which contains the label information of the classification category:

[0182]

[0183] After obtaining the projection matrix, the two sets of vectors are fused in serial and parallel modes to form a new feature vector, as shown below:

[0184]

[0185] Discriminant correlation analysis increases the category information of training samples by improving the objective function in canonical correlation analysis.

[0186] Convolutional neural network multimodal recognition, the classic convolutional neural network is mainly composed of basic units such as convolution layer, pooling layer, linear rectification layer (Relu layer). The original image is input into the convolutional neural network composed of convolution layer, pooling layer, and linear rectification layer. Through the arrangement and combination of convolution layer, pooling layer, and linear rectification layer, the feature map containing the semantic features of the image is gradually extracted. After the feature map is input into the classifier, the predicted probability value of the classification is obtained;

[0187] The pooling layer is also an important component of the convolutional neural network. The advantage of the pooling layer is that it does not need to update the weight parameters, and at the same time it can flexibly reduce the size of the previous level feature map, thereby reducing the multiplication operation of the next level of convolution;

[0188] The linear rectifier layer uses an activation function. The activation function is another very important part of the neural network. It is used to perform nonlinear mapping on the output features of the convolutional layer after the convolutional layer, thereby enhancing the network's ability to express features. The activation function selects the linear rectifier function ReLU, and the formula is: ReLU(x) = max(0,x), indicating that everything with an x ​​value less than zero is mapped to a y value of 0, but everything with an x ​​value greater than zero is mapped to itself;

[0189] Its differential equation is:

[0190]

[0191] The advantage of the ReLU function is that it can avoid the problem of gradient disappearance, because in the process of updating weights, its update rule is:

[0192]

[0193] Among them, w (L)For all weights in the last layer L, when partial derivatives appear When the number of weights and biases is very small, the update of many weights and biases is very small, that is, the gradient vanishing problem occurs. When the ReLU activation function is used, its differential equation shows that its result is either 0 or 1, also known as unilateral inhibition, which introduces good sparsity into the neural network and alleviates the problem of gradient vanishing during training. While solving the problem of gradient vanishing, it also improves the efficiency in terms of time and space complexity.

[0194] Softmax is a linear multi-classification model whose function is to convert the scores of each category into reasonable probability values;

[0195]

[0196] Where W, b are model parameters, K is the total number of categories, y is the probability of predicting that the sample belongs to category i. The higher the probability, the greater the possibility of belonging to the category. The sum of all probabilities is normalized to 1.

[0197] The multimodal feature recognition of convolutional neural networks uses the Inception-ResNet module. Convolutional neural networks containing residual connection ResNet and Inception modules are selected for fingerprint feature recognition. Since neural networks are not sensitive to noise, the parameters of many early classic models are actually very redundant. Therefore, early deep learning models have high requirements for computer computing performance and memory, and general equipment is difficult to meet the above hardware requirements. In order to solve this problem, researchers began to study the use of efficient network structure design to reduce the amount of calculation. For example, GoogleNet uses the Inception module instead of simply stacking network layers to reduce the amount of calculation. ResNet has also achieved excellent image recognition results by introducing residual connections. On the basis of the above, a convolutional neural network including residual connection and Inception module is selected for fingerprint feature recognition. The characteristic of residual connection is that a copy of the data of the previous layer of network is added to the output result and then input into the next layer of network; the feature map of the original image is superimposed with the output, so as to help the training error gradually decrease as the depth of the network increases. Through residual connection, the neural network changes from fitting the output F(x) to fitting the residual F(x)-x. The residual structure is easier to learn than the original function and is more suitable for deep model iteration. At present, most of the deep convolutional neural networks have a very obvious feature of deepening the network by stacking multiple residual blocks, so there will be no serious overfitting when training very deep networks. When designing the network structure, learn the advantages of the deep residual network ResNet and use the residual structure for jump transmission in the middle layer of the network. On the one hand, it is conducive to the downward propagation of gradients, so that the network can still directly use more original information after multiple convolution operations, and at the same time, it is conducive to better upward transmission of gradients.

[0198] The characteristic of the Inception-ResNet module lies first in the Inception structure. The ordinary single node between two activation functions is expanded into a new neural network, also known as a "network in network". It uses multiple convolution kernels of different sizes to obtain receptive fields of different scales, and then fuses features of different scales. As the number of network layers increases, the extracted features will become more abstract. Therefore, the Inception structure uses this dense network structure to obtain a convolutional visual network that is closer to reality.

[0199] The Inception-ResNet module adds the classic residual connection structure in ResNet based on the original Inception module, which can retain the original information of the previous item and is more suitable for deep model iteration. The Relu activation function is connected after the 1x1 convolution. By introducing nonlinear excitation, the network's ability to express nonlinear functions is improved;

[0200] The improved Inception module reduces the number of channels through 1×1 convolution, gathers information and then calculates it. While effectively utilizing computing power, it makes the model nonlinear and improves cross-channel information integration and information interaction. At the end of each module, several branches are finally merged through aggregation operations and aggregated on the last dimension of the output channel. The Inception module contains three convolutions of different sizes and one maximum pooling, which increases the adaptability of the network to different scales and increases the width of the network, avoiding the problem of training gradient dispersion due to the network being too deep.

[0201] Multimodal feature recognition network, the Inception series of networks originated from the classic image classification network GoogLeNet, proposed by Google, and is a convolutional neural network with excellent network depth and width. Its characteristic is that it designs a network with excellent topological structure, so it is often used as a feature processor in the field of computer vision. The Inception-ResNet-V2 network structure is selected for fingerprint feature recognition. Its network structure uses different Inception-ResNet modules one by one from shallow to deep on the main path. The first module on the main path uses the Stem module. The Inception-ResNet modules are named Inception-ResNet-A / B / C respectively. Each Inception-ResNet module performs multiple convolution operations or pooling operations on the input image in parallel, and splices all output results into a very deep feature map. In the process of extracting image features, convolution kernels of 1*1, 3*3 and 7*7 scales are used to extract image features at the same time, and then the outputs of these convolutions are stacked and passed to the next layer of the network. The nonlinear activation function ReLu is used, and then the fully connected layer in the traditional convolutional neural network is replaced by the global average pooling layer GAP to remove the spatial relationship in the image features and obtain features with high-level semantic information.

[0202] Adding residual connection structure makes the network structure more context-aware, which is beneficial for the network to express the deep features of fingerprints;

[0203] Use smaller 3x3 convolutions to make up for the smaller receptive field by deepening the network; decompose the large convolution kernel into a combination of a series of small convolution kernels. The combination of small convolution kernels improves the adaptability of the network to different scales and improves the generalization ability of the network.

[0204] Applying 1x1 convolution before the large convolution kernel can flexibly reduce the number of featuremap channels, reduce the amount of calculation by reducing the number of parameters and the number of multiplication operations, and ultimately achieve the purpose of improving fingerprint recognition efficiency;

[0205] The role of the loss function is to achieve network convergence by calculating the difference between the predicted value and the true value in the training set, with the goal of minimizing the loss function. The smaller the value of the loss function, the smaller the difference between the predicted value and the true value, and the more accurate the model's prediction. Under the supervision of the loss function, the convolutional neural network can learn better image features and speech information. Depending on the type of task, choosing an appropriate loss function can shorten the network training time, accelerate convergence, and improve recognition accuracy.

[0206] Combine center loss and softmax loss as a joint loss function to train the convolutional neural network:

[0207] Center loss function constrains the distance between the category centers of sample features:

[0208]

[0209] where x i is the eigenvalue, c yi is the center point of all sample features of the category corresponding to sample i, and N is the total number of samples;

[0210] Softmax loss cross entropy loss function constrains the distance between the actual output and the expected output:

[0211]

[0212] Among them, yi represents the prediction result, N and K are the total number of samples and the total number of categories respectively. The smaller the value of Softmax loss is, the closer the predicted value is to the true value, and the more accurate the prediction result is.

[0213] Joint loss function to train convolutional neural network:

[0214] L=L s +λL c (twenty three)

[0215] Among them, λ is the balance parameter of center loss and softmax loss. The larger λ is, the higher the weight of center loss is. In order to avoid the problem of reducing the intra-class distance and increasing the difficulty of model optimization during the calculation of the joint loss function, λ=0.01 is selected;

[0216] Multimodal biometric recognition with deep network and score fusion, based on the finger multimodal feature recognition network, uses convolutional neural network to automatically classify fingerprint images, adopts the fingerprint and finger vein multimodal recognition score fusion method and fingerprint and finger vein multimodal recognition feature fusion method, uses the deep network Inception-ResNet-V2, after the softmax classifier, obtains the comprehensive matching score through score-level fusion to obtain the final recognition result;

[0217] The characteristic of the model of the present invention is that it combines the traditional feature extraction method with the emerging convolutional neural network. Considering the complexity of fingerprint image preprocessing, the convolutional neural network is selected to replace the manually designed feature extraction algorithm, and the traditional image enhancement algorithm is selected for the vein image with strong directionality. The appropriate algorithm is selected according to the characteristics of the modality, which can combine the advantages of the two, saving the time of manually extracting fingerprint features, and does not require complex preprocessing of fingerprint images, thereby improving the model recognition speed, and enhancing the generalization ability of the model, that is, when the modality is converted, only the network parameters need to be retrained to quickly adapt to the new modality.

[0218] There are several commonly used score-level fusion methods: concatenation, weighted sum, mean fusion, and maximum fusion. For the simple concatenation fusion method, if the modality increases or the input category increases, it will cause dimensionality explosion, insufficient computing memory, and other problems. For the weighted sum fusion method, the choice of weights will directly affect the recognition accuracy of the system. The fingerprint feature recognition network uses a softmax classifier. In order to ensure the superposition of prediction scores of different modalities, the traditional method of vein feature recognition is also converted into a reasonable probability value (the value is between [0,1] and the sum of the probabilities is 1) through a softmax regression model after obtaining the matching score. For example, a sample may belong to N categories. Assume that the matching score score = {d1, d2,..., dN}. After the softmax function, the corresponding category probability value is:

[0219]

[0220] Then, different score fusion models are selected to fuse the category probabilities to obtain the final prediction result. Next, these three different score-level fusion models will be introduced respectively.

[0221] Mean Fusion

[0222] When the classification capabilities of different modalities are similar, the mean model can be used for fusion. Assuming that the results of M softmax classifiers are {h1,h2,...,hM}, the prediction result after fusion is:

[0223]

[0224] Weighted sum fusion

[0225] When different classifiers have different effects on the final prediction results, a weighted sum model is selected to assign different weights to different modality classifiers according to their prediction accuracy, where the results of the M softmax classifiers are {h1,h2,...,hM}:

[0226]

[0227] Maximum Fusion

[0228] When the prediction results of different modality classifiers are different, the maximum model can be used for fusion. Assume that the results of M softmax classifiers are {h1,h2,...,hM}, and select the category with the largest prediction probability as the final prediction result:

[0229] H(x)=max(h i )(29)

[0230] The deep network is combined with the multimodal recognition system of feature fusion and fractional fusion to effectively identify the multimodal features of fingers.

[0231] A convolutional neural network is composed of multiple stacked convolutional layers. This is because the features learned by a single convolutional layer are local and limited. By continuously stacking multiple convolutional layers, more comprehensive and in-depth image features can be learned. The convolution operation extracts matrix blocks from the input feature matrix according to certain rules, performs the same transformation operation on these matrix blocks, and generates a new output feature matrix. The convolution kernel is equivalent to a sliding window function acting on the matrix. The convolution kernel is multiplied by the corresponding block matrix elements one by one and then summed.

[0232] The pooling operation of the pooling layer compresses the input feature matrix, reduces the size of the feature matrix by downsampling, and simplifies the computational complexity. The pooling operation is divided into average pooling and maximum pooling. Average pooling takes the regional mean to retain the overall data characteristics, and maximum pooling takes the regional maximum value to retain the texture characteristics of the matrix.

[0233] The characteristic of the Inception-ResNet module lies first in the Inception structure, which expands the ordinary single node between the two activation functions into a new neural network, uses multiple convolution kernels of different sizes to obtain receptive fields of different scales, and then fuses features of different scales. As the number of network layers increases, the extracted features will become more abstract. Therefore, the Inception structure uses this dense network structure to obtain a convolutional visual network that is closer to reality.

[0234] At the same time, 1x1 convolution is also a way to fuse channel information.

[0235] The basic network selects Inception-ResNet-V2, and the Inception-ResNet-V2 network on the ImageNet dataset is used for pre-training. Since the ImageNet dataset and the fingerprint dataset are not similar, the low-level features learned by the previous convolutional layer are retained, and the terminal layer is replaced with a custom output layer, and the terminal layer is retrained. In terms of initial parameter selection, the input batch size of the convolutional neural network is 28, and the learning rate is 0.01 at the beginning. After training 80 epochs, it is reduced to 0.0001. Fine-tune the network. When fine-tuning the pre-trained network, the learning rate is as small as possible to avoid too much impact on the original training weights. The probability of random inactivation of neurons is 0.8, and the optimizer selects adam. By adjusting the network optimizer and parameters, observe whether it is overfitting, and select the optimal training model.

[0236] The Stem module stacks 1x1, 3x3, 1x7, 7x1 conv and 3x3 pooling together, and adds convolution parameters to the network based on the width of the network and its adaptability to scale. After the network adds convolution parameters, it independently selects the required convolution filter by changing the weights. A special reduction block is introduced to help reduce the size of feature maps from 35x35 to 17x17, and from 17x17 to 8x8, to solve the problem of too high dimension of feature data. Two 3*3 convolutions are selected instead of large convolution kernel convolutions to speed up calculations while reducing the number of parameters. In the middle layer of the network structure, 1*3 and 3*1 convolution kernels are used instead of 3*3 convolution kernels.

[0237] The new user registration module is used to provide a new user registration interface, requiring new users to enter basic information and biometric data (such as palm prints, palm veins and finger veins) to ensure the accuracy and reliability of the registration process, and store the information and biometric data of newly registered users, and update the user database; the core function of the user identity authentication module is to identify the user, by loading the user's biometric data information, and calling the trained model to compare and identify the input information, and give the verification result; the enhanced recognition module is used to provide a data enhancement algorithm to increase the diversity and reliability of biometric data. It supports image preprocessing and enhancement to improve the accuracy of biometric recognition. The logout function module allows users to actively request to log out their accounts and related information, and verifies and confirms the user's logout request. After confirming the user's logout, delete the user's personal information and biometric data in the system, including the corresponding records in the database.

[0238] The present invention collects finger multimodal biometric identification data and establishes a fingerprint and finger vein feature sample database; establishes a sample collection volunteer information database; completes the finger multimodal biometric image sample preprocessing task; completes the fingerprint and finger vein feature image sample partial global feature extraction task. Research deep neural network theory and improve CNN.

[0239] The present invention can complete the task of extracting global features of fingerprints and finger veins; complete the task of extracting local features of fingerprints and finger veins; establish a fingerprint and finger vein feature database; select some commonly used classifiers and decision-level fusion methods to conduct experiments and obtain preliminary experimental results; develop a fingerprint and finger vein recognition experimental platform; and conduct fingerprint and finger vein recognition experiments on networks such as improved CNN.

[0240] The present invention can extract special structural point features of fingerprints and finger veins; study and implement a neural network classifier based on a selection mechanism; implement a recognition system based on the fusion of fractional fingerprints and finger veins, and try multiple fusion methods; divide all the extracted features into 2-3 groups, and use artificial neural network feature training and selection; combine the neural network classifier based on the selection mechanism to improve the fingerprint and finger vein recognition experimental platform; and implement an improved CNN combined fingerprint and finger vein fusion recognition system.

[0241] Compared with the prior art, the present invention has higher recognition accuracy. Multimodal biometric recognition technology is not a simple superposition of single biometric features, but a combination of multiple biometrics such as face, iris, fingerprint, palm vein, voiceprint, etc. through a carefully designed fusion algorithm, so that the characteristics and advantages of different biometric features are fully utilized, thereby making the identity authentication and recognition process more accurate. Taking face recognition and iris recognition as examples, the structure and features of the human face contain unique identity information, and the texture details of the iris are also rich and diverse. The fusion of face recognition and iris recognition can achieve information complementarity and further improve the accuracy of recognition.

[0242] Compared with the prior art, the present invention is more secure. With the continuous development of technology, the theft, theft and duplication of information are becoming increasingly easy. Multimodal biometric technology uses multiple biometric features in combination to make up for the vulnerability of certain biometric features to forgery. For attackers, it is more difficult to forge multiple biometric features than a single biometric feature, which greatly increases the cost of forgery.

[0243] Compared with the prior art, the present invention has a wider scope of application. Due to interference from the external environment and changes in individual conditions, some users may suffer from biometric damage, which may affect the integrity of the biometrics and further affect the accuracy and reliability of the biometric system. For example, the user's hands are incomplete or the innate fingerprint patterns are not obvious. Therefore, a single biometric technology often has certain application limitations, or is only applicable to some application scenarios, while multimodal biometrics can largely make up for the application defects of a single biometric and broaden the scope of application scenarios.

[0244] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A hand multi-feature recognition system based on graph convolutional network, including a new user registration module, a user identity authentication module, an enhanced recognition module and a logout function module, characterized in that: The enhanced recognition module is used for the collection and preprocessing of finger multimodal features. The collection and preprocessing of finger multimodal features utilizes collection equipment to collect fingerprints and finger veins respectively, and performs denoising, binarization, normalization, and enhancement of the region of interest on the collected images. The collection and preprocessing of finger multimodal features include fingerprint region of interest interception and image enhancement, finger vein image enhancement, feature extraction, filtering, principal component analysis, histogram series fusion calculation of LGBP features, feature matching calculation, feature-level fusion calculation of multi-set discriminant correlation analysis, convolutional neural network multimodal recognition, convolutional neural network multimodal feature recognition, multimodal biometric recognition of multimodal feature recognition network and deep network and score fusion; Fingerprint region of interest interception and image enhancement, first intercept the region of interest ROI of the collected image, and enhance the fingerprint ROI image through the adaptive local contrast enhancement algorithm ACE. In the ACE algorithm, the gray value x(i, j) of a point is taken, and the local mean is calculated in a window with a side length of 2n+1 with the pixel point (i, j) as the center: Compute the local variance: Subtract the original image from the blurred image, retain the high-frequency part and multiply it by the amplification gain coefficient to get the enhanced pixel value: f(i,j)=m x (i,j)+G(i,j)[x(k,l)-m x (i,j)] (3) Where G(i,j) is the contrast gain function, the contrast gain function is CG, CG>1, the spatially adaptive CG function is proportional to the global mean and inversely proportional to the local mean square error: Finger vein image enhancement, using a multi-directional filtering method for finger vein image enhancement. Based on the direction of the finger veins, 8 different directional filters are used to construct fingerprint extraction filters. The size of the 8 filters is 5×5. Taking pixel (i, j) as the center, the gray value of the pixel is I(i, j), the pixel value after filtering is Ie(i, j), and the pixel value after enhancement is: I’ t0 (i,j)=I e (i,j-2)+I e (i,j+2)-I e (i,j) (5) Select the minimum value of the 8 directional filters and adjust the image grayscale range to [0,255]: Feature extraction refers to a feature extraction method based on statistical learning, which uses linear and nonlinear features to extract image information, and uses four methods: Gabor filtering, local binary pattern LBP, linear discriminant analysis LDA, and principal component analysis PCA to select texture and scale features; Gabor filtering uses the Gabor filtering algorithm. Aiming at the unique ridge texture characteristics of the vein image, the Gabor filter uses parameters that conform to the special vein texture and constructs an even-symmetric Gabor filter EGF in the spatial domain. EGF is defined as: Where k = γ / (2πσ 2 ), x θ =xcosθ+ysinθ,y θ =-xsinθ+ycosθ, θ represents the filter direction, f c represents the center frequency of the sine wave, σ and γ are the standard deviation and aspect ratio of the Gaussian curve respectively; During principal component analysis, a dimensionality reduction algorithm is used to extract significant features from the original input image, reduce the amount of data during classification, and enhance the recognition system's ability to distinguish between modalities. The dimensionality reduction algorithm uses the principal component analysis PCA algorithm. The PCA algorithm flow is as follows: First, center all samples: Compute the covariance matrix: Perform eigenvalue decomposition on the covariance matrix: S·p i =λ i ·p i ; According to the energy percentage of the principal component Select the eigenvectors corresponding to the first i eigenvalues ​​to form a matrix W, where: According to the above process, the matrix W is calculated for feature extraction: Z = W T x; The concatenated fusion calculation of LGBP features adopts the concatenated fusion algorithm of LGBP feature histograms to perform feature encoding of LGBP, combining multi-scale and multi-directional Gabor filtering with the local binary pattern LBP, effectively describing the global direction information and local features of finger veins, thereby weakening the influence of external illumination and finger posture changes on finger feature fusion. LBP uses a rotation-invariant circular operator. The LBP operator replaces the square neighborhood with a circular neighborhood, allowing P pixel points in a circular neighborhood with a radius of R. R = 1.0 and P = 8. Centered on the central pixel point, the LBP features in the circular neighborhood are rotated at 0°, 45°, 90°, 135°, 180°, 225°, 270°, and 315° respectively, and 8 LBP feature encoding values are calculated. The smallest value among these LBP feature encoding values is taken as the final LBP feature encoding value of the central pixel point. The rotation-invariant LBP formula is as follows: Among them, x c ,y c is the center pixel, i c is the gray value of the center pixel, i n is the gray value of adjacent pixels, and s is a sign function: LBP encoding is performed on the enhanced ROI image, and then the image is evenly divided into N*N blocks respectively. The grayscale histogram of each block is calculated, and finally the grayscale histograms of all blocks are concatenated end to end to form the feature histograms of fingerprints, finger veins, and knuckle prints, and then they are concatenated to form the final feature histogram; The feature matching calculation adopts the feature matching algorithm. The feature histogram after the concatenated fusion of LGBP features is selected to use the matching algorithm based on the histogram intersection method for matching. The matching algorithm uses the histogram intersection method as the similarity measurement criterion, and the formula is as follows: Among them, H1 and H2 are two LGBP feature histograms to be matched, L is the dimension of the feature histograms H1 and H2, Sim(H1, H2) is the similarity of the two histograms. The larger Sim(H1, H2) is, the higher the similarity between H1 and H2, and the more likely they are to match successfully. A similarity decision threshold th is set. If Sim(H1, H2) > th, the test sample and the template sample belong to the same category; if Sim(H1, H2) < th, they belong to different categories; ImageNet and PalmNet are used for multi-modal recognition of finger veins. The encoding fusion method of local Gabor binary pattern LGBP features is used to realize the fusion of fingerprint and finger vein multi-modal features. The encoding of local Gabor binary pattern (LGBP) features and the feature extraction method of convolutional neural network are adopted, so as to add a spatial pyramid pooling layer to the network to solve the problem that the input image size in the convolutional neural network must be fixed due to the different sizes of different fingerprint database images; The feature-level fusion calculation of multi-set discriminant correlation analysis adopts the feature-level fusion algorithm based on multi-set discriminant correlation analysis. Since different feature vectors come from different biometric modalities, in order to solve the problem of incompatibility between features, canonical correlation analysis CCA is used to help perform two kinds of feature-level fusion. CCA is an analysis algorithm that uses the correlation relationship between comprehensive variable pairs to reflect the overall correlation between two groups of indicators. The algorithm is as follows: The input matrix is ​​X = [x1,…,x n ],Y=[y1,…,y n ], where the intra-class and inter-class variance matrices of X and Y are denoted by S xx ∈R p×q and S yy ∈R q×q The difference matrix is ​​denoted as S xy ∈R p×q , and there exists For the overall variance matrix S, which contains information within and between classes, the formula is: According to the formula of the correlation coefficient: in, And there is The purpose of canonical correlation analysis is to solve the following formula: Calculate a pair of projection matrices W X and W y , and then use the projection matrix to get the new eigenvector and Discriminant correlation analysis DCA is set on the basis of CCA to increase the category information of training samples by improving the objective function, so as to solve the problem when the relationship between two sets of features is a complex nonlinear correlation relationship; The DCA algorithm is as follows: Input vector, X = [x1,…,x n ],Y=[y1,…,y n ], the intra-class variance matrices of X and Y are still recorded as and. Different from the objective function of CCA, the objective function of DCA adds a diagonal matrix A, which contains the label information of the classification category: After obtaining the projection matrix, a new feature vector is formed by fusing these two sets of vectors in serial and parallel modes, as shown below: Discriminant correlation analysis increases the category information of training samples by improving the objective function in canonical correlation analysis. Convolutional neural network multimodal recognition: the original image is input into a convolutional neural network composed of a convolutional layer, a pooling layer, and a linear rectification layer. Through the arrangement and combination of the convolutional layer, the pooling layer, and the linear rectification layer, a feature map containing the semantic features of the image is gradually extracted. After the feature map is input into the classifier, the predicted probability value of the classification is obtained; The advantage of the pooling layer is that it does not need to update the weight parameters, and at the same time it can flexibly reduce the size of the feature map of the previous level, thereby reducing the multiplication operation of the next level of convolution; The linear rectification layer uses an activation function to perform nonlinear mapping on the output features of the convolutional layer after the convolutional layer, thereby enhancing the network's ability to express features. The activation function selects the linear rectification function ReLU, and the formula is: ReLU(x) = max(0,x), indicating that everything with an x ​​value less than zero is mapped to a y value of 0, but everything with an x ​​value greater than zero is mapped to itself; Its differential equation is: The advantage of the ReLU function is that it can avoid the problem of gradient disappearance, because in the process of updating weights, its update rule is: Among them, w (L) For all weights in the last layer L, when partial derivatives appear When the number of weights and biases is very small, the update of many weights and biases is very small, that is, the gradient vanishing problem occurs. When the ReLU activation function is used, its differential equation shows that its result is either 0 or 1, also known as unilateral inhibition, which introduces good sparsity into the neural network and alleviates the problem of gradient vanishing during training. While solving the problem of gradient vanishing, it also improves the efficiency in terms of time and space complexity. Softmax is a linear multi-classification model whose function is to convert the scores of each category into reasonable probability values; Where W, b are model parameters, K is the total number of categories, y is the probability of predicting that the sample belongs to category i. The higher the probability, the greater the possibility of belonging to the category. The sum of all probabilities is normalized to 1. The multimodal feature recognition of convolutional neural network uses Inception-ResNet module. The convolutional neural network including residual connection ResNet and Inception module is selected for fingerprint feature recognition. The characteristic of residual connection is that a copy of the data of the previous layer of network is added to the output result and then input into the next layer of network; the feature map of the original image is superimposed with the output, so as to help the training error gradually reduce as the depth of the network increases. Through residual connection, the neural network changes from fitting the output F(x) to fitting the residual F(x)-x. The residual structure is used for jump transmission in the middle layer of the network. On the one hand, it is conducive to the downward propagation of gradients, so that the network can still directly use more original information after multiple convolution operations, and at the same time, it is conducive to better upward transmission of gradients. The Inception-ResNet module adds the classic residual connection structure in ResNet based on the original Inception module, which can retain the original information of the previous item and is more suitable for deep model iteration. The Relu activation function is connected after the 1x1 convolution. By introducing nonlinear excitation, the network's ability to express nonlinear functions is improved; The number of channels is reduced through 1×1 convolution, and information is gathered for calculation. While effectively utilizing computing power, the model is nonlinearized, which improves cross-channel information integration and information interaction. At the end of each module, several branches are finally merged through aggregation operations and aggregated on the last dimension of the output channel. The Inception module contains three convolutions of different sizes and one maximum pooling, which increases the adaptability of the network to different scales and increases the width of the network, avoiding the problem of training gradient dispersion due to the network being too deep. Multimodal feature recognition network, select Inception-ResNet-V2 network structure for fingerprint feature recognition. Its network architecture uses different Inception-ResNet modules one by one from shallow to deep on the main path. The first module on the main path uses the Stem module. The Inception-ResNet modules are named Inception-ResNet-A / B / C respectively. Each Inception-ResNet module performs multiple convolution operations or pooling operations on the input image in parallel, and splices all output results into a very deep feature map. In the process of extracting image features, convolution kernels of 1*1, 3*3 and 7*7 scales are used to extract image features at the same time, and then the outputs of these convolutions are stacked and passed to the next layer of the network. The nonlinear activation function ReLu is used, and then the fully connected layer in the traditional convolutional neural network is replaced by the global average pooling layer GAP to remove the spatial relationship in the image features and obtain features with high-level semantic information. Combine center loss and softmax loss as a joint loss function to train the convolutional neural network: Center loss function constrains the distance between the category centers of sample features: where x i is the eigenvalue, c yi is the center point of all sample features of the category corresponding to sample i, and N is the total number of samples; Softmax loss cross entropy loss function constrains the distance between the actual output and the expected output: Among them, yi represents the prediction result, N and K are the total number of samples and the total number of categories respectively. The smaller the value of Softmaxloss is, the closer the predicted value is to the true value, and the more accurate the prediction result is. Joint loss function to train convolutional neural network: L=L s +λL c (23) Among them, λ is the balance parameter of center loss and softmax loss. The larger λ is, the higher the weight of center loss is. In order to avoid the problem of reducing the intra-class distance and increasing the difficulty of model optimization during the calculation of the joint loss function, λ=0.01 is selected; Multimodal biometric recognition based on deep network and score fusion, based on the finger multimodal feature recognition network, uses convolutional neural network to automatically classify fingerprint images, adopts the fingerprint and finger vein multimodal recognition score fusion method and the fingerprint and finger vein multimodal recognition feature fusion method, uses the deep network Inception-ResNet-V2, after the softmax classifier, obtains the comprehensive matching score through score-level fusion to obtain the final recognition result.

2. The hand multi-feature recognition system based on graph convolutional network according to claim 1, characterized in that: The convolutional neural network is composed of multiple stacked convolutional layers. The convolution operation extracts matrix blocks from the input feature matrix according to certain rules, performs the same transformation operation on these matrix blocks, and generates a new output feature matrix. The convolution kernel is equivalent to a sliding window function acting on the matrix. The convolution kernel is multiplied by the corresponding block matrix elements one by one and then summed.

3. The hand multi-feature recognition system based on graph convolutional network according to claim 1, characterized in that: The pooling operation of the pooling layer compresses the input feature matrix, reduces the size of the feature matrix by downsampling, and thus simplifies the computational complexity. The pooling operation is divided into average pooling and maximum pooling. Average pooling takes the regional mean to retain the overall data features, and maximum pooling takes the regional maximum to retain the texture features of the matrix.

4. The hand multi-feature recognition system based on graph convolutional network according to claim 1, characterized in that: The characteristic of the Inception-ResNet module lies first in the Inception structure, which expands the ordinary single node between the two activation functions into a new neural network, adopts multiple convolution kernels of different sizes to obtain receptive fields of different scales, and then performs feature fusion of different scales. As the number of network layers increases, the extracted features will become more abstract. Therefore, the Inception structure uses this dense network structure to obtain a convolutional visual network that is closer to reality.

5. The hand multi-feature recognition system based on graph convolutional network according to claim 1, characterized in that: The simultaneous 1x1 convolution is also a way to fuse channel information.

6. The hand multi-feature recognition system based on graph convolutional network according to claim 1, characterized in that: The basic network selects Inception-ResNet-V2, and the Inception-ResNet-V2 network on the ImageNet dataset is used for pre-training. Since the ImageNet dataset and the fingerprint dataset are not similar, the low-level features learned by the previous convolutional layer are retained, and the terminal layer is replaced with a custom output layer, and the terminal layer is retrained. In terms of initial parameter selection, the input batch size of the convolutional neural network is 28, and the learning rate is 0.01 at the beginning. After training 80 epochs, it is reduced to 0.0001. The fine-tuning network is used. When fine-tuning the pre-trained network, the learning rate is as small as possible to avoid too much impact on the original training weights. The probability of random inactivation of neurons is 0.8, and the optimizer selects adam. By adjusting the network optimizer and parameters, whether it is overfitting is observed, and the optimal training model is selected.

7. The hand multi-feature recognition system based on graph convolutional network according to claim 1, characterized in that: The Stem module stacks 1x1, 3x3, 1x7, 7x1 conv and 3x3 pooling together, and adds convolution parameters to the network according to the width of the network and the adaptability of the network to the scale. After the network adds the convolution parameters, it independently selects the required convolution filter by changing the weights. A special reduction block is introduced to help reduce the size of the feature map from 35x35 to 17x17, and from 17x17 to 8x8, to solve the problem of too high dimension of feature data, and selects two 3*3 convolutions instead of large convolution kernel convolutions to speed up calculations while reducing the number of parameters. In the middle layer of the network structure, two convolution kernels of 1*3 and 3*1 are used instead of the 3*3 convolution kernel.

Citation Information

Cited By

  • Finger vein recognition method and device based on multi-domain feature enhancement and fusion

    CN121545193A