Palm print identification method based on multi-scale parallel hybrid network

By adopting a parallel hybrid network under multi-scale vision in palm print recognition technology, integrating local and global features, the problem of poor recognition quality in the existing technology under noise and different lighting environments is solved, and the palm print recognition effect with high accuracy and robustness is achieved.

CN119964206APending Publication Date: 2025-05-09BEIJING UNIV OF TECH

Patent Information

Application Number
CN202510064993.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing palm print recognition technology has poor recognition quality when facing environmental noise and different lighting environments, making it difficult to achieve high accuracy and robustness.

Method used

A parallel hybrid network under multi-scale vision is adopted, and a parallel hybrid feature extraction module combining convolutional neural network and self-attention network is fused to integrate local and global features, and the network adaptively learns the distribution imbalance of feature distribution through multi-scale features.

Benefits of technology

It realizes high accuracy and high robust palm print recognition under different data acquisition methods, and is suitable for application scenarios in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964206A_ABST
    Figure CN119964206A_ABST
Patent Text Reader

Abstract

The invention discloses a palmprint recognition method based on a multi-scale parallel hybrid network, and belongs to the field of computer vision. Firstly, a parallel mixed feature extraction module is designed, and the parallel mixed feature extraction module comprises a branch based on a convolutional neural network and a branch based on a self-attention network to extract local and global features in parallel; for a branch based on a convolutional neural network, the invention provides a comprehensive attention module fusing a space attention mechanism and a pixel attention mechanism to further strengthen the important feature discrimination capability. Secondly, a multi-scale feature fusion network is designed, and local and global features are fused under different scales and different depths through a sampling mechanism and jump connection in combination with a parallel mixed feature extraction module. And finally, mapping the image into an information-intensive feature vector for classification and matching. Experiments prove that the method has high precision and high robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Based on deep learning technology, this paper studies a high-precision palmprint recognition method that mixes local and global features under multi-scale vision. Specifically, it involves feature extraction and feature fusion technology for palmprint images, so as to obtain the most representative and discriminative texture features in the palmprint image, and further realize high-precision and high-efficiency palmprint recognition. Background Art

[0002] Palmprint recognition, as a biometric technology, uses the unique pattern of an individual's palm to provide an important solution for personal authentication and security protection. It is valued for its high individual uniqueness and contactless operation, eliminating the need for any physical contact or invasive methods. As an important branch of the field of biometrics, palmprint recognition technology plays an increasingly critical role in today's society. With its unique advantages, it plays an irreplaceable role in security verification and identity recognition. It has the following advantages: (1) Uniqueness. Every individual's palmprint is unique. Even the palmprints of identical twins are different, which makes palmprint recognition have a high degree of individual recognition; (2) Stability. Palmprints are relatively stable since birth and change very little with age, which provides a reliable basis for long-term identity verification; (3) Non-invasiveness. Palmprint recognition does not require the collection of blood or other biological samples. Users only need to place their palms on the recognition device. This non-invasive method is more easily accepted by users. (4) Palmprint recognition technology can be easily integrated into existing security systems, such as access control systems and attendance systems, enhancing the security of these systems. In the past decade, palmprint recognition technology has played an important role in many practical application scenarios, including but not limited to electronic payment, smart office and other financial and transportation intelligent services. Despite this, the current palmprint recognition technology still needs to overcome some difficulties in actual deployment, such as adaptability to environmental noise and robustness of palmprint recognition quality under different lighting environments. Therefore, developing a technology that can resist noise interference and can extract palmprint features with high accuracy is crucial to promote the development of palmprint recognition technology.

[0003] Palmprint recognition technology can be mainly divided into the following categories: 1) Methods based on manual feature extraction. These methods rely on prior knowledge and achieve recognition by extracting features such as texture, direction, and frequency in palmprint images. For example, methods based on histograms and dimensional mapping can describe the texture distribution of palmprints. 2) Methods based on deep learning. With the development of deep learning technology, palmprint recognition methods based on deep learning can automatically extract features without manually designing feature extractors. These methods learn complex feature representations of palmprint images through structures such as convolutional neural networks (CNN). The advantage of deep learning methods is that they can process large-scale data sets and can learn more abstract and discriminative feature representations.

[0004] The advantage of traditional methods based on manual feature extraction is that they do not require a large amount of training data, but the disadvantage is that they are susceptible to noise and outliers, and the feature expression of palmprint images of different qualities is not always effective. Although deep learning-based technologies can automatically extract features, these methods often ignore the interaction between local and global features, limiting the further improvement of recognition performance. The present invention aims to propose a palmprint recognition method based on a multi-scale parallel hybrid network, effectively fuse local and global features from a multi-scale vision, adaptively learn the imbalanced distribution of important features, and effectively improve the accuracy and robustness of palmprint recognition under different data collection methods to address the limitations of the existing technology. Summary of the invention

[0005] The present invention first designs a parallel hybrid feature extraction module, which first preliminarily extracts texture and direction features through a convolutional layer and a learnable Gabor filter, and on this basis uses a dual-branch structure of a convolutional neural network and a self-attention network to extract local features and global features of the image respectively. For the branch based on the convolutional neural network, the present invention designs a comprehensive attention mechanism module, which uses a spatial attention mechanism and a pixel attention mechanism to enhance the ability to distinguish the importance of features; for the branch based on the self-attention network, a simplified Transformer Encoder block is used to extract global dependencies. Secondly, a multi-scale feature fusion network is designed, which is divided into an encoder part and a decoder part. The encoder part uses a downsampling mechanism to obtain shallow contour features at different scales; the decoder part uses an upsampling mechanism to obtain deep abstract detail features. Each layer of the encoder and decoder uses a parallel hybrid feature extraction module and a jump connection to fuse local and global features at different scales and depths. Finally, a fully connected layer is used to map the image into an information-intensive feature vector, and the feature vector is matched and classified using cosine similarity. The present invention belongs to the field of computer vision, and specifically relates to technologies such as biometric recognition and image processing.

[0006] The present invention provides a palmprint recognition method based on a multi-scale parallel hybrid network. Different from the existing palmprint recognition methods, the present invention designs a parallel hybrid attention module to fuse local and global features, and designs a comprehensive attention module to adaptively learn the distribution imbalance of important features. On this basis, a multi-scale feature extraction network architecture is designed, which combines the parallel hybrid attention module to fuse local and global features from different scales and depths. The high accuracy and high robustness of palmprint recognition under different data collection methods are achieved, which has certain significance and value for palmprint recognition application scenarios in complex environments.

[0007] To implement the above method, the specific steps are as follows:

[0008] Step 100: normalize all palmprint image ROIs so that the pixel values ​​of the palmprint images are normalized to the range of [0, 1] and perform data preprocessing.

[0009] Step 200: construct a parallel hybrid feature extraction block (PHEB), which includes two parallel branches based on a convolutional neural network (CNN) and a self-attention network (Transformer), which are used to extract local and global features respectively.

[0010] Step 210: For the convolutional neural network branch, a comprehensive attention block (CAB) is constructed to integrate the advantages of spatial attention (SA) and pixel attention (PA) to adaptively learn the imbalance of feature importance distribution.

[0011] Step 220: For the self-attention network branch, use the global dependency extraction capability of the simplified Transformer to achieve global feature extraction without increasing the computational complexity.

[0012] Step 300, construct a multi-scale feature fusion network, and realize local and global feature extraction and fusion under different scale vision through sampling mechanism, jump connection and parallel hybrid feature extraction module in step 200. This multi-scale feature extraction method can capture multi-level details in palmprint images, thereby improving the accuracy and robustness of recognition.

[0013] Step 400: Select five data sets collected in different ways, and randomly divide the input data sets into training sets and test sets. In the training phase, the model is trained using the training data and the sample labels, and the cross-entropy loss function is used to guide the model parameter update.

[0014] Step 500: In the testing phase, the trained model is used to generate information-intensive feature vectors for the input palmprint image, and the cosine similarity of the feature vectors is calculated for feature matching and classification, and the feature vectors with the highest cosine similarity are assigned to the same class.

[0015] Step 600, using the equal error rate (EER) and the receiver operating characteristic curve (ROC) as the criteria for evaluating the model results, a variety of classic palmprint recognition methods are compared under data sets of different acquisition methods to prove the high accuracy and high robustness of the method.

[0016] The step 100 of normalizing the palm print image to achieve image preprocessing includes:

[0017] The palmprint ROI image is normalized so that its pixel value is normalized to the range of [0,1], which is convenient for subsequent feature extraction and model training. And efficient data preprocessing methods such as random selection, color jitter, and random rotation are used for data enhancement to enhance the robustness of the model under different environmental interferences.

[0018] The step 200 of extracting features from the normalized palmprint image using a parallel hybrid feature extraction module includes:

[0019] A parallel hybrid feature extraction module is constructed. This module first uses convolutional layers and a learnable Gabor filter to perform preliminary texture feature extraction, and then combines two different types of deep learning architectures - convolutional neural networks and self-attention networks to extract local and global features of palmprint images in parallel. This design fully utilizes the advantages of convolutional neural networks in capturing local features (such as textures and edges) and the ability of Transformer in capturing global dependencies and long-distance relationships.

[0020] Sub-step 210, the branch based on the convolutional neural network is composed of a comprehensive attention extraction module. The comprehensive attention extraction module comprehensively utilizes the spatial attention mechanism and the pixel attention mechanism to adaptively learn the imbalance of the distribution of important features, and further improves the local feature extraction capability of the branch.

[0021] Sub-step 220, the self-attention network-based branch captures the global features in the palmprint image through a multi-head self-attention mechanism. This mechanism allows the model to consider the information of the entire image when processing each pixel, thereby capturing long-range dependencies. In order to provide the model with spatial location information, each input feature map is encoded with a position, which helps the branch understand the relative position of pixels in the image. In this process, we use a standard TransformerEncoder block, but reduce the number of heads of its multi-head attention mechanism and the hidden layer dimension.

[0022] The step 300 of constructing a multi-scale feature fusion network using a sampling mechanism includes:

[0023] A network architecture is designed that uses a sampling mechanism to segment feature maps of different scales, and down-samples and up-samples feature maps at different depth levels to capture information at different scales. The down-sampling process is in the encoder part of the network, and the spatial dimension of the feature map is gradually reduced through continuous convolution layers and pooling layers, while increasing the number of channels. This process not only extracts local features, but also gradually abstracts more advanced representations, so that the network can capture the details and texture information in the palm print image. The up-sampling process is in the decoder part of the network, and the spatial dimension of the feature map is gradually restored through transpose convolution, while reducing the number of channels. This process helps to restore the spatial information of the image and ensure that important positioning information is not lost during the feature extraction process. Each layer uses the parallel hybrid feature extraction block of step 200 to achieve local and global feature extraction at that scale. Finally, the shallow features in the encoder are fused with the deep features in the decoder through jump connections. This fusion strategy enables the network to simultaneously utilize shallow detail information and deep semantic information, enhancing the expressiveness of the features.

[0024] The step 400 of randomly dividing the input data set into a training set and a test set, and training the model using the training data and the sample labels includes:

[0025] Given the training data and the sample labels corresponding to the training data, the goal of the model is to map the image into an information-dense feature vector with high discrimination. First, the input data set is randomly divided into a training set and a test set. The specific operation is to randomly select three palmprint images from each category as the training set, and the remaining palmprint images in the category as the test set. The cross entropy loss is used as the loss function to guide the learning process of the model parameters.

[0026] The step 500 of using cosine similarity to perform feature matching and classification includes:

[0027] For each palmprint image in the test set, the feature vector is calculated, and the cosine similarity between it and the feature vectors of all palmprint images in the training set is calculated, and the samples to which the feature vector with the highest cosine similarity belongs are matched to the same category. Cosine similarity is a measure of the angular difference between two vectors, and its value ranges from -1 (completely dissimilar) to 1 (completely similar).

[0028] The step 600 of testing and evaluating the model results includes:

[0029] The EER of the feature vector is calculated and the ROC curve is drawn. These two key indicators are used to test and judge the model effect. The present invention is compared with multiple classic palmprint recognition methods under multiple different acquisition method data sets to verify the extremely high accuracy of the present invention and its robustness under multi-scenario data sets.

[0030] The present invention has the following advantages:

[0031] (1) The method proposed in the present invention is an adaptive feature learning method based on deep learning. It does not need to rely on any prior knowledge, is completely data-driven, and has strong adaptability.

[0032] (2) Compared with other palmprint recognition methods, the method proposed in the present invention has extremely high accuracy and is particularly suitable for scenarios requiring strict recognition accuracy;

[0033] (3) The method of the present invention has extremely high accuracy under a variety of different data collection methods, proving its strong robustness and suitability for complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a model architecture diagram of a multi-scale parallel hybrid network.

[0035] Figure 2 Examples of palmprint images from five different databases.

[0036] Figure 3 The ROC curves for the comparison of various palmprint recognition methods under five different databases. DETAILED DESCRIPTION

[0037] The method will be described in detail below with reference to the accompanying drawings. It should be noted that, in general, the method can be applied to palmprint image recognition with different image qualities and different application scenarios.

[0038] Step 100: Palmprint image normalization and preprocessing

[0039] Sub-step 110: crop all palmprint image ROIs to a standard size of 128×128, and normalize the pixel values ​​of the palmprint images to a range of [0, 1].

[0040] Step 120: The normalized image is enhanced using data preprocessing methods such as random selection, random cropping, random rotation, and Gaussian blur to improve the versatility and anti-interference ability of the model.

[0041] Step 200: Construct a parallel hybrid feature extraction module

[0042] The existing palmprint recognition methods focus on the mining of detailed features, ignoring the influence of the correlation between local features and global features on the palmprint recognition effect, which limits the recognition accuracy to a certain extent. In addition, the current research lacks sufficient research on the imbalanced distribution of important features, which limits the robustness of diverse data under different acquisition environments. In order to solve the above problems, the present invention proposes a parallel hybrid feature extraction module (such as Figure 1 The local and global features are extracted in parallel through convolutional neural networks and self-attention networks. It can be divided into the following sub-steps:

[0043] Sub-step 210, first use a convolution layer with a convolution kernel size of 3×3, a step size of 2, and a padding of 1, and a learnable Gabor filter with a size of 7×7, a step size of 1, and a padding of 3 to perform preliminary texture and direction feature extraction. In this step, the input channel and output channel of the convolution layer and the learnable Gabor filter are equal and remain unchanged. The learnable Gabor filter is a commonly used method in the field of palmprint recognition. The influencing factors such as direction, phase, and wavelength are set as learnable parameters, which are automatically adjusted during the training process, as shown in the following formula:

[0044]

[0045] x′=x Cos(θ)+y sin(θ)

[0046] y′=-x sin(θ)+y cos(θ)

[0047] Where ω, θ, ψ and σ represent the wavelength, direction, phase shift and standard deviation of the Gabor function respectively. Coordinates x and y refer to the spatial coordinates in the input image, and x′ and y′ are the transformed coordinates. These factors are automatically adjusted by model training. The features after preliminary extraction are extracted by the branch based on convolutional neural network and the branch based on self-attention network to extract local and global features respectively.

[0048] Sub-step 220, for the branch based on convolutional neural network. The present invention proposes a comprehensive attention module (such as Figure 1As shown in CAB on the right, this module integrates spatial attention (SA) and pixel attention (PA) to adaptively learn the imbalanced distribution of features of different importance and further enhance the ability to extract local features. Spatial attention SA includes X-direction attention (XA), Y-direction attention (YA) and channel attention (CA), which focus on the correlation between rows, columns and channels of feature maps respectively to enhance the model's ability to recognize horizontal and vertical features and important channel information in the image. Among them, the XA module uses global average pooling to compress the width dimension of the feature map to 1, and successively passes through a convolution layer with a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a ReLU activation function, a second convolution kernel size of 1×1, a stride of 1, and a padding of 0, and a Sigmoid activation function to obtain the weight proportion of each row and multiply it with the original input. The calculation ideas of the YA and CA modules are consistent with those of the XA module. They compress the height dimension and channel dimension into 1 through global average pooling, and then pass through a 1×1, step size 1, fill 0 convolution layer, ReLU activation function, the second 1×1, step size 1, fill 0 convolution layer, and Sigmoid activation function to obtain the weight proportion of the corresponding column and channel and perform product operations on the original input. The number of input channels and output channels of the XA, YA, CA, and PA modules in this step are consistent, and the value is related to the number of channels required for the location where the module is used. The location where the module is used in the subsequent step 300 will be further explained. The calculation steps of XA, YA, CA, and PA are shown in the following formula:

[0049]

[0050] PA=Sig(Conv(Max(Conv(CA out ),0)))·CA out

[0051] Among them, XA, YA, CA, and PA represent the outputs of the X-direction attention mechanism, the Y-direction attention mechanism, the channel attention mechanism, and the pixel attention mechanism, respectively. in , Y in , X c and CA out Respectively represent the X-direction attention mechanism, the Y-direction attention mechanism, the input of the channel attention mechanism, and the output of the channel attention mechanism. Conv represents the convolution operation, sig represents the sigmoid function, and Max represents the maximum value. H and W represent the height and width of the feature map, respectively. in (h), Y in(w) represents the hth row of information input by the X-direction attention mechanism and the wth column of information input by the Y-direction attention mechanism, and h and w are traversed from 1 to H and 1 to W respectively. i and j represent the coordinates of the pixel points, traversing every pixel point on the feature map.

[0052] Sub-step 230, for the self-attention network branch. We use a simplified TransformerEncoder module for global feature extraction. This branch first cuts the input image into 16 small pixel aggregation area blocks in sequence, then uses a four-head multi-head attention mechanism to capture global dependencies and performs the first step of feature processing through element addition and layer normalization. Then a multilayer perceptron (MLP) with a hidden layer size of 256 is used for feature mapping, and then a second feature processing is performed through element addition and layer normalization. This step is a standard Transformer Encoder block step. We only reduce the number of attention heads in the multi-head attention mechanism to 4 and the hidden layer size in the MLP to 256. The number of input and output channels of this module is consistent, and is consistent with the number of output channels in step 210.

[0053] Sub-step 240: We concatenate the local and global features obtained in sub-steps 220 and 230 at the channel level. This step will double the number of feature channels. Then, we use a shared weight convolution kernel with a size of 1×1, a step size of 1, and a padding of 0 to perform feature fusion. The number of output channels is halved, which is consistent with the number of output channels in steps 220 and 230. Finally, we perform a batch normalization operation.

[0054] Step 300: Design a multi-scale parallel hybrid network architecture.

[0055] The existing palmprint recognition methods have not made sufficient research on the correlation between global and local features at different scales, which limits the multi-scale feature extraction capability of the method and is one of the bottlenecks for improving recognition accuracy. The network architecture of the present invention is divided into two parts: encoder and decoder (such as Figure 1 The PMSHNet part in the upper left corner is shown), which includes two layers of up and down sampling mechanisms. Each layer uses the parallel hybrid feature extraction module proposed in step 200 to extract local and global features at the current scale, which can be divided into the following sub-steps:

[0056] In sub-step 310, the encoder part first uses a 3×3 convolutional layer with a stride of 1 and padding of 1 to change the number of channels of the input image from 1 to 32. Then, feature extraction is performed through a parallel hybrid feature extraction module in step 200. This step does not change the size and number of channels of the input and output feature maps. The input and output channels of the parallel hybrid feature module used are both 64. Then, two downsamplings are performed. Both downsamplings use a 3×3 convolutional layer with a stride of 2 and padding of 1 to halve the feature map size and double the number of channels. The shallow features obtained after each sampling operation will use the parallel hybrid feature extraction module in step 200 to extract local and global features at the current scale and depth. The first downsampling changes the feature map size from 128×128×32 to 64×64×64, so the input and output channels of the parallel hybrid feature extraction block used after the first downsampling are 64. The second downsampling changes the feature map size from 64×64×64 to 32×32×128, so the input and output channels of the parallel hybrid feature extraction module used after the second downsampling are 128.

[0057] In sub-step 320, the decoder part uses an upsampling mechanism to restore the feature map size. The upsampling uses transposed convolution to double the feature map size and halve the number of channels. The decoder part uses two upsampling operations. Both upsampling operations use transposed convolution with a convolution kernel size of 3×3, a step size of 1, and a padding of 1 to obtain the deep features of the network as the abstract detail features of the image. The input feature map size obtained by the first upsampling operation is 32×32×128, and the output feature map size after upsampling is 64×64×64. It is spliced ​​at the channel level with the feature map output by the decoder of the same layer, and the resulting feature map size is 64×64×128. Then, a parallel hybrid feature extraction module of step 200 is used, and the input feature and output feature channels of this module are both 128. Next, the second upsampling operation is performed, the input feature map size is 64×64×128, the output feature map size is 128×128×64, and the channel-level feature concatenation is performed with the output of the encoder at the same layer, and the result is 128×128×128. Finally, a convolution layer with a kernel size of 3×3, a stride of 2, and a padding of 1 and a maximum pooling layer with a size of 2×2 and a stride of 2 is sampled. The result is flattened and passed through two fully connected layers (the hidden layer size is 4096) to obtain a final feature vector of size 1500 bits.

[0058] Step 400: On five datasets collected in different ways, the input dataset is randomly divided into a training set and a test set. Specifically, three images are used as the training set and the remaining images are used as the test set. The sample data of the five datasets are as follows: Figure 2As shown. In the training phase, the sample labels of the training data are known, and the model is trained using the training data and the sample labels. The cross-entropy loss function is used to guide the model parameter update. The cross-entropy loss function is defined as follows:

[0059]

[0060] where y c Represents the one-hot encoding of the true label, p c is the probability that the model predicts that it belongs to class c.

[0061] Step 500: In the test phase, each palm print image in the test set is mapped into an information-intensive feature vector (1500 bits in size) through the network in step 300, and the cosine similarity between the feature vectors of all palm print images in the training set is calculated, and the samples to which the feature vector with the highest cosine similarity belongs are matched to the same category. Cosine similarity is a measure of the angular difference between two vectors, and its value ranges from -1 (completely dissimilar) to 1 (completely similar). The formula is as follows:

[0062]

[0063] Where A·B represents the inner product of vector A and vector B, and ||A|| and ||B|| represent the modulus of vector A and vector B.

[0064] Step 600: Compare the data sets with five different data collection methods and multiple classic palmprint recognition technologies, and use EER and ROC to measure the recognition effect.

[0065] Sub-step 610, calculate EER. EER calculation is an important indicator in biometric recognition system, which represents the threshold when the false acceptance rate (FAR) and the false rejection rate (FRR) are equal. In the test set, we adjust the classification threshold, record the FAR and FRR under different thresholds, and find the point where FAR equals FRR, that is, EER. The lower the EER, the better the recognition performance of the model, because it means that while maintaining a low false rejection rate, the false acceptance rate can also be effectively reduced.

[0066] Sub-step 620, draw the ROC curve as a tool to evaluate the performance of the classification model. It is represented by drawing a curve graph of the true positive rate (True Positive Rate, TPR) against the false positive rate (False Positive Rate, FPR). The true positive rate is the proportion of successfully identified positive samples, while the false positive rate is the proportion of negative samples mistakenly identified as positive. In the test set, the true positive rate and FPR are calculated for each possible threshold, and then the ROC curve is drawn. The closer the curve is to the upper left, the stronger the recognition ability of the model is.

[0067] Through these two indicators, we can comprehensively evaluate the performance of the model and perform subsequent optimization accordingly.

[0068] The palmprint images in the present invention come from five public palmprint databases with different collection methods, including: Tongji, IITD, MS Green, MS Blue, MS Red and MS NIR. Among them, the Tongji dataset contains 12,000 images from 600 palms, which are divided into two sessions, with 10 images for each palm, resulting in 60,000 real matches and 35,940,000 fake matches. The IITD dataset contains 2,601 images, with 5 to 7 images for each palm, with a total of 1,266,840 real matches and 2,760 fake matches. Each multispectral (MS) dataset contains 12 images for each palm, with a total of 18,000 real matches and 8,982,000 fake matches.

[0069] The experimental environment is a PC, NVIDIA GeForce RTX 4090 GPU and PyTorch 2.3.1 environment. The learning rate is set to 0.0001 and the batch size is set to 32.

[0070] First, the EER comparison experiment was conducted between the method of the present invention and various classic palmprint recognition methods. Table 1 shows the EER of each method on five public palmprint databases with different collection methods: Tongji, IITD, MS Red, MS Green, MS Blue, MS Red and MS NIR. The experimental results show that the method of the present invention achieves the best results on all five databases, proving the high accuracy of the present invention and its high robustness in dealing with different collection methods.

[0071] Table 1 Equal error rates of different recognition methods on five databases (%)

[0072]

[0073]

[0074] Second, the ROC curves of the method of the present invention and the above classic palmprint recognition methods are compared. The ROC curves of different recognition methods are as follows: Figure 3 As shown in the figure, the closer the ROC curve is to the upper left, the better the recognition performance. Figure 3 We can see that the method of the present invention has the best ROC curve and the best recognition performance.

[0075] The palmprint recognition method based on a multi-scale parallel hybrid network in the present invention has the following advantages:

[0076] (1) This technical solution proposes an efficient palmprint recognition model based on deep learning, which can flexibly adapt to different data environments and does not require any pre-set knowledge as support during the execution process;

[0077] (2) Improved recognition accuracy, suitable for scenarios requiring high recognition accuracy;

[0078] (3) It is highly robust to data collected in different ways and is suitable for complex industrial production environments.

Claims

1. A palmprint recognition method based on a multi-scale parallel hybrid network, characterized in that The following steps are involved: Firstly, a parallel hybrid feature extraction module is designed. The module firstly extracts texture and direction features through convolutional layers and learnable Gabor filters. On this basis, a dual-branch structure of convolutional neural network and self-attention network is used to extract local features and global features of the image respectively. For the branch based on convolutional neural network, the present invention designs a comprehensive attention mechanism module, which uses spatial attention mechanism and pixel attention mechanism to enhance the ability to distinguish the importance of features. For the branch based on self-attention network, a multi-head self-attention encoder is used. Secondly, a multi-scale feature fusion network is designed, which is divided into an encoder part and a decoder part. The encoder part uses a downsampling mechanism to obtain shallow contour features at different scales. The decoder part uses an upsampling mechanism to obtain deep abstract detail features; each layer of the encoder and decoder uses a parallel hybrid feature extraction module and jump connections to fuse local and global features at different scales and depths; finally, the fully connected layer is used to map the image into an information-dense feature vector, and the feature vectors are matched and classified using cosine similarity.

2. The method according to claim 1, characterized in that: Step 100: Palmprint image normalization and preprocessing Sub-step 110, cropping all palmprint image ROIs to a standard size of 128×128, and normalizing the pixel values ​​of the palmprint images to a range of [0,1]; Step 120: The normalized image is enhanced by using data preprocessing methods such as random selection, random cropping, random rotation and Gaussian blur to improve the versatility and anti-interference ability of the model; Step 200: Construct a parallel hybrid feature extraction module It is divided into the following sub-steps: Sub-step 210, first use a convolution layer with a convolution kernel size of 3×3, a step size of 2, and a padding of 1 and a learnable Gabor filter with a size of 7×7, a step size of 1, and a padding of 3 to perform preliminary texture and direction feature extraction; set the influencing factors such as direction, phase, wavelength, etc. as learnable parameters, which are automatically adjusted during the training process, as shown in the following formula: x′=x cos(θ)+y sin(θ) y′=-x sin(θ)+y cos(θ) Where ω, θ, ψ and σ represent the wavelength, direction, phase shift and standard deviation of the Gabor function respectively; the coordinates x and y refer to the spatial coordinates in the input image, and x′ and y′ are the transformed coordinates; these factors are automatically adjusted by model training; the features after preliminary extraction are extracted by the branch based on convolutional neural network and the branch based on self-attention network to extract local and global features respectively; Sub-step 220, for the branch based on convolutional neural network; a comprehensive attention module is proposed, the spatial attention SA includes X-direction attention XA, Y-direction attention YA and channel attention CA, which respectively focus on the correlation between rows, columns and channels of the feature map to enhance the model's ability to recognize horizontal and vertical features and important channel information in the image; The XA module uses global average pooling to compress the width dimension of the feature map into 1, and passes through a convolution layer with a convolution kernel size of 1×1, a step size of 1, and a padding of 0, a ReLU activation function, a second convolution kernel size of 1×1, a step size of 1, and a padding of 0, and a Sigmoid activation function to obtain the weight ratio of each row and multiply it with the original input; the YA and CA modules have the same calculation ideas as the XA module, and compress the height dimension and channel dimension into 1 through global average pooling respectively, and pass through a 1×1, step size of 1, and a padding of 0 convolution layer, a ReLU activation function, a second 1×1, step size of 1, and a padding of 0 convolution layer, and a Sigmoid activation function to obtain the weight ratio of the corresponding column and channel and multiply it with the original input; the number of input channels and output channels of the XA, YA, CA, and PA modules in this step are consistent, and the calculation steps of XA, YA, CA, and PA are shown in the following formulas: PA=Sig(Conv(Max(Conv(CA) out ),0)))·CA out Among them, XA, YA, CA, and PA represent the outputs of the X-direction attention mechanism, the Y-direction attention mechanism, the channel attention mechanism, and the pixel attention mechanism, respectively; in , Y in , X c and CA out Respectively represent the input of X-direction attention mechanism, Y-direction attention mechanism, channel attention mechanism and the output of channel attention mechanism; Conv represents convolution operation, sig represents sigmoid function, and Max represents the maximum value; H and W represent the height and width of the feature map respectively. in (h), Y in (w) represents the h-th row of the X-direction attention mechanism input and the w-th column of the Y-direction attention mechanism input, h and w are traversed from 1 to H and 1 to W respectively; i and j represent the coordinates of the pixel points, traversing every pixel point on the feature map; Sub-step 230, for the self-attention network branch; We use a simplified Transformer Encoder module for global feature extraction; this branch first cuts the input image into 16 small pixel aggregation area blocks in sequence, then uses a four-head multi-head attention mechanism to capture global dependencies and performs the first feature processing through element addition and layer normalization; then uses a multi-layer perceptron MLP with a hidden layer size of 256 for feature mapping, and then performs the second feature processing through element addition and layer normalization; this step is a standard Transformer Encoder block step, we only reduce the number of attention heads in the multi-head attention mechanism to 4, and the hidden layer size in the MLP to 256; the number of input and output channels of this module is consistent, and is consistent with the number of output channels in step 210; Sub-step 240: For the local and global features obtained in sub-steps 220 and 230, they are concatenated at the channel level. This step will double the number of feature channels. Then, the features are fused through a shared weight convolution kernel with a size of 1×1, a step size of 1, and a padding of 0. The number of output channels is halved, which is consistent with the number of output channels in steps 220 and 230. Finally, a batch normalization operation is performed. Step 300: design a multi-scale parallel hybrid network architecture; The network architecture is divided into two parts: encoder and decoder, including two layers of up and down sampling mechanisms. Each layer uses the parallel hybrid feature extraction module proposed in step 200 to extract local and global features at the current scale, which is divided into the following sub-steps: In sub-step 310, the encoder part first uses a 3×3, stride 1, padding 1 convolution layer to change the number of channels of the input image from 1 to 32; then, feature extraction is performed through a parallel hybrid feature extraction module in step 200. This step does not change the size and number of channels of the input and output feature maps. The input and output channels of the parallel hybrid feature module used are both 64; then, two downsamplings are performed. Both downsamplings use a 3×3, stride 2, padding 1 convolution layer to halve the feature map size and double the number of channels; each sampling operation The shallow features obtained after the first downsampling will use the parallel hybrid feature extraction module of step 200 to extract local and global features at the current scale and depth; the first downsampling changes the feature map size from 128×128×32 to 64×64×64, so the input channel and output channel of the parallel hybrid feature extraction block used after the first downsampling are 64; the second downsampling changes the feature map size from 64×64×64 to 32×32×128, so the input and output channels of the parallel hybrid feature extraction module used after the second downsampling are 128; Sub-step 320, the decoder part uses an upsampling mechanism to restore the feature map size. The upsampling uses a transposed convolution to double the feature map size and halve the number of channels. The decoder part uses two upsampling operations. Both upsampling operations use a transposed convolution with a convolution kernel size of 3×3, a step size of 1, and a padding of 1 to obtain the deep features of the network as the abstract detail features of the image. The input feature map size obtained by the first upsampling operation is 32×32×128, and the output feature map size after upsampling is 64×64×64, which is spliced ​​at the channel level with the feature map output by the decoder of the same layer, and the resulting feature map size is 64×64×128; then a parallel hybrid feature extraction module of step 200 is passed, and the input feature and output feature channels of this module are both 128; then a second upsampling operation is performed, the input feature map size is 64×64×128, the output feature map size is 128×128×64, and the channel level feature splicing is performed with the output result of the encoder of the same layer, and the result is 128×128×128; finally, a convolution layer with a convolution kernel size of 3×3, a stride of 2, and a padding of 1 and a maximum pooling layer with a size of 2×2 and a stride of 2 are sampled, and the result is flattened and passed through two fully connected layers, with a hidden layer size of 4096, to obtain a feature vector with a final size of 1500 bits; Step 400: on five datasets acquired in different ways, the input dataset is randomly divided into a training set and a test set. Specifically, three images are used as the training set and the remaining images are used as the test set. In the training phase, the sample labels of the training data are known, and the model is trained using the training data and the sample labels. The cross entropy loss function is used to guide the update of the model parameters. The cross entropy loss function is defined as follows: where y c Represents the one-hot encoding of the true label, p c is the probability that the model predicts that it belongs to category c; Step 500, in the test phase, each palm print image in the test set is mapped into an information-intensive feature vector through the network calculation in step 300, and the cosine similarity between the feature vector and the feature vectors of all palm print images in the training set is calculated, and the samples to which the feature vector with the highest cosine similarity belongs are matched to the same category; cosine similarity is a measure of the angular difference between two vectors, and its value ranges from -1, i.e., completely dissimilar, to 1, i.e., completely similar; the formula is as follows: Where A·B represents the inner product of vector A and vector B, and ||A|| and ||B|| represent the modulus of vector A and vector B.

Citation Information

Patent Citations

  • Injection pump for dual fuel engine

    CA490050A

Cited By

  • Data registration method based on parallel variable window convolutional neural network

    CN120411179A

  • A data registration method based on parallel variable window convolutional neural network

    CN120411179B

  • Visual inspection method and system for intelligent blank carrying line

    CN121392214A

  • Visual inspection method and system for a blank intelligent handling line

    CN121392214B

  • Direct positioning method based on multi-scale feature fusion neural network

    CN122085209A