Non-contact voiceprint and palmprint palm vein multi-modal identity recognition system and method
By collecting voiceprint, palmprint, and palm vein features non-contactly, and combining deep learning and feature fusion algorithms, the problems of environmental sensitivity and low recognition accuracy in existing technologies are solved, achieving highly secure and robust identity authentication.
Patent Information
- Application Number
- CN202210927661.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-03
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2042-08-03
AI Technical Summary
In existing technologies, fingerprint recognition has high environmental requirements, is sensitive to humidity and cleanliness, has a low recognition rate, is difficult to recognize fingerprints with scratches, has high requirements for operating procedures, and fingerprint imprints may remain on the device; palm print and palm vein recognition are affected by factors such as light and temperature, have difficulty in ROI positioning and segmentation, have low accuracy in non-contact methods, and are easily interfered with by counterfeiting methods.
Non-contact acquisition of voiceprint, palmprint, and palm vein features is adopted. Combined with the ResNet network and SE module of deep learning, the system performs image preprocessing, feature extraction, feature fusion, and comparison. Corner points are detected using improved FAST and Shi-Tomasi algorithms, and identity authentication is performed using a joint discriminative sparse coding algorithm.
It improves the security and accuracy of identity authentication, is suitable for scenarios with high hygiene requirements, enhances the robustness and portability of the system, reduces noise interference, and improves the security and robustness of identification.
Smart Images

Figure CN115188084B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of biometric recognition, and in particular to a non-contact multi-modal identity recognition system and method for voiceprint and palmprint palm vein. BACKGROUND
[0002] With the rapid development of global information industry, how to carry out fast, accurate and secure identity recognition and verification in a digital environment is a hot topic of concern in recent years. The traditional identity authentication is prone to loss, forgetfulness and forgery, so that biometric recognition technology is attracting more and more attention. Biometric recognition is a process of identifying the authenticity of identity information by collecting physiological and behavioral characteristics of human body after systematic processing. At present, the more mature or widely used biometric recognition technologies include face, voice, fingerprint, iris, finger vein, DNA, signature and gait, etc. However, single-modal biometric recognition may have a decreased accuracy due to sensor noise, inappropriate feature extraction or matching method, and may have security problems due to the forgery of features, such as fake fingerprints. Further, multi-modal biometric recognition has entered the sight of people. Different description methods or perspectives of the same object are called modalities, and multi-modal representation is to use information from multiple entities to jointly realize the representation of a specific task. Generally, a multi-modal biometric recognition system fuses two or more biological characteristics at different levels, which can be divided into sensor layer, feature layer, score layer and decision layer. The difficulty of multi-modal fusion authentication research is how to effectively collect, extract and compare the features of multi-source heterogeneous data.
[0003] Representation learning technology refers to a set of technologies that convert original complex data distribution into a form that can be effectively recognized and applied by a machine according to a task, i.e. to extract useful information from data to learn data representation, thereby greatly improving the effectiveness of algorithm models and the accuracy of predictors. The research on representation learning technology based on multi-modal data environment enables representation learning to establish a model for processing and associating multiple modal information, and to perform multi-modal information fusion, thereby improving the accuracy and security of identity authentication. The goal of multi-modal representation learning is to extract the representation of data objects (users) from multiple heterogeneous modal data. A typical method is to concatenate independent representations of each modality to form a joint representation, and then learn subsequent tasks on this joint representation. The fusion of data representation and the unification of data from multiple data sources not only overcomes the heterogeneity between data, but also extracts complementary information from multiple data sources, so that the fused representation has more abundant and effective information than single-modal representation.
[0004] Prior art one
[0005] Fingerprint identification is to use the uneven lines on the skin of the end of the finger to identify the identity. Fingerprint has uniqueness and stability. By comparing the fingerprint with the fingerprint saved in the database, the real identity can be verified. In various biometric identification technologies, fingerprint identification is still the most mature identity identification technology. Fingerprint identification has been accepted by the official in many countries and has become an effective means for the judicial department to identify identity. It has been widely used in many other industries and has become a synonym for biological identification and a de facto standard. Fingerprint identification technology mainly involves fingerprint image acquisition, fingerprint image preprocessing, fingerprint feature extraction, establishment of fingerprint image database, comparison and matching of fingerprint feature values and other processes. After years of research by relevant personnel, various fingerprint identification methods have been developed, among which the most mature and widely used is the fingerprint identification method based on minutia points. The image used in the laboratory is collected by using the existing devices in the laboratory. The specific content includes: Obtaining a fingerprint image direction map; Segmentation of the fingerprint image; Fingerprint image enhancement; Binarization and post-processing of the fingerprint image; Fingerprint image thinning; Feature extraction of the fingerprint image; Matching of the fingerprint image.
[0006] Disadvantages of prior art one
[0007] High requirement for environment, sensitive to humidity and cleanliness of fingers, dirty, oil and water can cause identification failure or affect the identification result;
[0008] There are difficulties in identifying low-quality fingerprints with scars, exuviae, etc., and the identification rate is low;
[0009] High operation specification requirement for fingerprint identification;
[0010] Fingerprint marks may be left on the device, and these marks may be used to copy fingerprints.
[0011] Prior art two
[0012] Palmprint and palm vein fusion identification technology belongs to a kind of multi-modal identification technology. Compared with other biometric identification technologies, palmprint and palm vein fusion identification has higher identification accuracy, convenience and stability, which helps to improve the convenience of people's life and improve the security of personal information to a certain extent.
[0013] The palm print and palm vein pattern do not change with age, and palm print feature recognition has the advantages of rich texture features, easy acceptance by users, high security and stability, etc.
[0014] Disadvantages of Prior Art Two
[0015] (1) Influence of palm vein and palm print image acquisition environment. Palm vein acquisition mainly includes contact acquisition and non-contact acquisition. Regardless of which acquisition method is used, the acquisition process will be affected by factors such as light, acquisition background, and temperature.
[0016] (2) Influence of palm vein key region positioning and segmentation. In order to obtain a region with rich vein features, the hand palm region of interest (ROI) image needs to be positioned and segmented. Researchers generally use palm vein images from the Hong Kong Polytechnic University database for vein recognition research. The palm vein images in the database have a hardware device installed between the middle finger and the ring finger to fix the palm during acquisition, making it difficult to position and segment the palm vein ROI image. Due to the lack of appropriate ROI positioning and segmentation methods, the accuracy of feature extraction is low, and the recognition rate is not high.
[0017] (3) Interference between palm vein and palm print. Palm vein images have palm prints, and existing algorithms still cannot completely remove the interference of palm prints. For example, using fuzzy threshold judgment and global gray value matching to improve the robustness of the algorithm, but without better removing the interference of palm prints, the recognition effect of palm veins is not good.
[0018] (4) Non-contact acquisition methods mainly have problems such as position offset, distance drift, image defocus, and brightness fluctuation of palm print sample images. For palm print recognition anti-counterfeiting, there are mainly silicone prostheses, palm print films, and other counterfeit means. These factors are the main reasons why non-contact palm print recognition systems have lower accuracy than contact palm print recognition systems, and are the main reasons limiting the practical application of non-contact palm print recognition systems.
[0019] References
[0020] [1] Liu Qianying, Liu Ji. Development Trend of Biometric Recognition Technology in Identity Verification Field [J]. Electronic World, 2020 (05): 23-24;
[0021] [2] Xie Lu, Yu Fei. Security Identity Authentication Technology Based on Multi-modal Biometric Recognition [J]. Security Science and Technology, 2016 (01): 36-40;
[0022] [3] Zhou Chenyi. Multi-modal Biometric Recognition Based on Fusion Algorithm and Deep Learning [D]. Southern Medical University, 2020;
[0023] [4] Zhang L, Wang HB, Tao L, Zhou J. Adaptive multi-modal biometric fusion based on classification distance score[J]. Journal of Computer Research and Development, 2018, 55(1): 151-162;
[0024] [5] Ma Z. Research on dual-mode identity authentication based on fingerprint and electrocardio signal[D]. Tianjin University of Technology, 2021;
[0025] [6] Zhang Y. Algorithm research on multi-modal biometric recognition technology[D]. Changchun University of Technology, 2017;
[0026] [7] Ding X. Multi-modal biometric recognition technology and its standardization dynamic[J]. Computer Knowledge and Technology, 2017, 13(36): 153-154;
[0027] [8] Jiansong, Lu Kai. Review of representation learning for complex heterogeneous data[J]. Computer Science, 2020, 47(02): 1-9. SUMMARY
[0028] The present application provides a non-contact voiceprint and palmprint palm vein multi-modal identity recognition system and method to overcome the defects of the prior art. The multi-modal identity authentication based on intelligent data representation theory is the core content, and the related technology is integrated in the network security scene. It is an identity authentication multi-modal biometric recognition method with high security, convenience and reliability.
[0029] In order to achieve the above invention purpose, the technical scheme adopted by the present application is as follows:
[0030] A non-contact voiceprint and palmprint palm vein multi-modal identity recognition method, comprising:
[0031] Step 1, image preprocessing; preprocessing mainly includes three steps, first, low-pass filtering is used to denoise the infrared collected palm image, second, Sauvola algorithm is used to extract the binary image of the palm area in the image enhancement part, and finally, ROI positioning part is used to perform gray scale transformation on the palmprint and palm vein, so that the palm edge is prominent, then Canny operator is used to detect the palm edge, and finally the image is cropped to obtain the palm area image of interest;
[0032] Step 2, feature extraction; feature extraction is divided into two parts, the first part is to extract speech features, and the second part is to extract palm print and palm vein two hand features; ResNet is used as the main structure, SE module is introduced, and SE-ResNet network structure is constructed, the preprocessed picture is input into the SE-ResNet network structure, a global pooling layer is added to generate feature distribution, and information coding extraction is completed; in order to obtain the correlation between channels, ReLU activation function and sigmoid gate control mechanism are combined to complete feature rescaling;
[0033] Step 3, feature fusion; a multi-layer feature fusion mechanism is adopted, the interaction between hand and audio different modalities is obtained by decomposing the bilinear model, and the paired audio and hand features are input into the fusion model, and the final result is output through softmax on the full connection layer;
[0034] Step 4, feature comparison; the feature points preliminarily extracted by the improved FAST corner detection algorithm are calculated by Shi-Tomasi algorithm, and the corner response function of each point is calculated, and the points with the maximum response value are determined as feature points according to the corner response function; for the matching of binary feature description vector, hamming distance is used as the similarity measure between descriptors;
[0035] Step 5, output interaction; joint discriminative sparse coding algorithm is used to judge the sample feature points in three modalities, so that the distance in the class is minimized, and the distance between classes is maximized; according to the actual scene requirement, a suitable threshold is set, if two matched samples belong to the same class and the voiceprint, palm print and palm vein are matched successfully, the interface displays authentication success, otherwise, it prompts authentication failure.
[0036] Further, step 2 is specifically: for any given information entering the network module, conversion is performed as shown in formula (1):
[0037] (1)
[0038] is the input picture, is the extracted feature;
[0039] SE compresses global spatial information into a channel descriptor, which contains the global distribution of feature response in channel dimension, and obtains statistical data in channel dimension by using global average pooling layer; the statistical value is obtained by formula (2) compression with spatial dimension :
[0040] (2)
[0041] transformed output The statistics of the channel descriptors, interpreted as a collection of local descriptors, are able to express the whole image;
[0042] The aggregated information obtained by the compression operation fully captures the dependencies in the channel dimension; a simple gating mechanism (3) with sigmoid activation function is chosen:
[0043] (3)
[0044] where, denotes the ReLU activation function, and ;
[0045] To limit the complexity of the model and help the model generalize, two fully connected layers (FC) around the nonlinearity are chosen to parameterize the gating mechanism in a bottleneck structure; the final output of the block is rescaled using the activation function (4) on the transformed output resulting in:
[0046] (4)
[0047] where, , denotes the product of the corresponding channel of the feature map and the scalar ; the role of this activation function s is to assign a weight to each channel depending on the descriptor of the input feature.
[0048] Further, step 3 is specified as follows:
[0049] Decomposing the bilinear model considers each feature pair by a linear transformation:
[0050] (5)
[0051] where x e R n and y e R m are the input feature vectors from different modalities of hand and audio, W i is the weight matrix, and b i is the bias;
[0052] The weight matrix W i is decomposed into two low-rank matrices, i.e., W = UV , where U e R and V e R impose a constraint d < min(n,m) on the dimension d; the equation (5) is further rewritten as:
[0053] (6)
[0054] Capture the inherent correlation between the two isomeric modes, formula (7):
[0055] (7)
[0056] Where 1∈R d Represents the column vector of 1, and represents the Hadamard or element-wise product;
[0057] In order to obtain the output feature vector z, two three-order tensors are needed: U=[U1,…,U O ]∈R n×d×o And V= [V1,…,V o ]∈R m×d×o ; Replace the column vector with linear projection P∈R d×o , the vector z is represented as:
[0058] (8)
[0059] Where b∈R o Is the bias vector;
[0060] Add a nonlinear activation function after each linear mapping, and the vector z is further represented as:
[0061] (9)
[0062] Where σ represents any nonlinear activation function, and x and y represent hand attention vectors and audio feature vectors respectively, then the value of x is greater than 0, and y is in the range of [-1, 1];
[0063] Further add Relu function to normalize the output of the network, and the final vector z is represented as:
[0064] (10)
[0065] Input the paired audio and hand features into the fusion model, and output the final result through softmax on the fully connected layer.
[0066] Further, step 4 is specifically as follows:
[0067] The improved FAST algorithm is adopted, and the specific improvement is: taking 24 pixel points around a pixel point P as the detection template, the gray value of P point is IP, and setting threshold T, if the gray value of 14 continuous pixel points in 24 pixel points is greater than IP+T or less than IP-T, then P is the corner point;
[0068] The feature points are optimized using Shi-Tomasi algorithm, the Shi-Tomasi algorithm compares the smaller one of two feature values with a given minimum threshold, if greater than the minimum threshold, a strong corner point will be obtained;
[0069] The Shi-Tomasi algorithm calculates the local small window The corner points are detected after moving in each direction and detecting the gray scale condition; the window is translated The gray scale changes For
[0070] (11)
[0071] In the formula, M is 2 The autocorrelation matrix of the derivative of the image is calculated
[0072] (12)
[0073] For the matching of the binary feature description vector, the Hamming distance is used as the similarity measure between the descriptors; assuming that the two feature vectors of the descriptor are , The Hamming distance of , is:
[0074] (13)
[0075] By determining the threshold of the Hamming distance, it is judged whether the feature vectors match.
[0076] Further, step 5 is specifically as follows:
[0077] The joint discriminant sparse coding algorithm is specifically as follows: given the feature matrices X, Y and Z of three modalities, three projection matrices P x , P y and P z are learned jointly, three peak features are mapped to sparse matrices V x ∈R d ×N, V y ∈R d ×N and V z ∈R d ×N, which can accurately approximate the original matrices X, Y and Z;
[0078] After obtaining the feature representations V x , V y and V z from the three modalities, they are quantized as
[0079] ; (14)
[0080] where sgn() is the sign function, and C x = [c 1 , c 2 , …, c N ]∈R l×N , c i denotes the learned sparse binary code of the i-th class, and l = (1, 2, …, 12) is the length of the sparse binary code.
[0081] The sparse constraint is imposed on the projected feature representation V x and V y to reduce the projection error, and the Frobenius norm is used as the cost function, and the engineering error is represented as
[0082] (15)
[0083] where a, b > 0, a + b ∈ (0, 1) is the balance parameter of the three different modalities.
[0084] Two constraints are imposed on the projected sparse feature: 1) for each modal intra-class sample, the distance within the class is minimized, and the distance between classes is maximized; 2) for intra-class samples, the information correlation between feature points is maximized, so the distance is minimized; By constraint, the projected sparse feature has stronger resolution and compactness.
[0085] A non-contact voiceprint and palmprint palm vein multi-modal identity recognition system for the multi-modal identity recognition method, characterized by comprising: a power supply module, a fixed wavelength infrared LED light source module, an image acquisition CCD module, a voice acquisition module and a storage module.
[0086] The power supply module is used to supply power to the entire multi-modal identity recognition system.
[0087] The fixed wavelength infrared LED light source module irradiates the human hand through the infrared LED light source, and assists the image acquisition CCD module to collect the human palmprint and palm vein information features.
[0088] The image acquisition CCD module collects the human palmprint and palm vein information features.
[0089] The voice acquisition module extracts voice information using MFCC features.
[0090] The storage module is used to store the data collected by the voice acquisition module and the image acquisition CCD module.
[0091] The multi-modal identity recognition module outputs the results through picture preprocessing, picture feature extraction, feature fusion comparison.
[0092] Compared with the prior art, the present application has the advantages that:
[0093] (1) The voiceprint, palmprint and palm vein features are collected in a non-contact manner, improving the security of authentication, and being suitable for scenes with high requirements for health environment in epidemic situations.
[0094] (2) The feature extraction adopts a deep learning method, reducing the tediousness of manual feature extraction, enhancing the anti-noise interference ability, and improving the robustness and portability of the system.
[0095] (3) The voiceprint recognition is integrated into the palmprint and palm vein recognition, and three modalities of feature fusion are used for identity authentication, improving the security, accuracy and robustness of authentication. BRIEF DESCRIPTION OF DRAWINGS
[0096] Figure 1 is the multi-modal identity recognition system architecture of the embodiment of the present application;
[0097] Figure 2 is the multi-modal identity recognition system workflow diagram of the embodiment of the present application;
[0098] Figure 3 is the SE-ResNet network structure diagram of the embodiment of the present application;
[0099] Figure 4 is the multi-layer feature fusion model structure diagram of the embodiment of the present application;
[0100] Figure 5 is the feature matching flowchart of the embodiment of the present application. DETAILED DESCRIPTION
[0101] In order to make the purpose, technical scheme and advantages of the present application clearer, the following will further describe the present application according to the drawings and examples.
[0102] The voice, hand multi-modal information collection device is an important device for human identity recognition, and its collection principle is as shown in Figure 1 The voice, hand multi-modal identity recognition designed by the present system places the human hand in an infrared LED light source environment, uses a CCD device to collect human palmprint and palm vein information features, extracts voice information using MFCC features, compares the extracted voice, hand multi-modal features with verification images, and achieves the purpose of human identity recognition.
[0103] The overall architecture design of the present system is as shown in Figure 1 , mainly including a hardware part and a software part. The hardware part is mainly used for collecting multi-modal voice, hand feature information, and the software part is mainly used for multi-modal information processing and recognition. The system flowchart is as shown in Figure 2The hardware part specifically comprises a power supply module, a fixed-wavelength infrared LED light source module, an image acquisition CCD module, a voice acquisition module and a storage module; and the software part comprises image preprocessing, feature extraction algorithm, feature fusion comparison and user interaction interface.
[0104] The feature extraction is divided into two parts, the first part is to extract voice features, and the second part is to extract palm print and palm vein two hand features. Figure 3 As shown in the figure. A global pooling layer is added to generate a feature distribution, and information coding extraction is completed. In order to obtain the correlation between channels, the ReLU activation function and the sigmoid gate control mechanism are combined to complete the feature re-labeling. In addition, in order to simplify the complexity of model parameters, full connection layers are used at both ends of the ReLU function.
[0105] The SE (Squeeze-and-Excitation Networks) module is a calculation unit composed of any given transformation, which transforms the information entering the network module as shown in (1):
[0106] (1)
[0107] is the input picture, is the extracted feature. In order to make the lower level layer also obtain the information of the global receptive field from the network, the SE compresses the global spatial information into a channel descriptor, which contains the global distribution of feature response in the channel dimension. A statistical data in the channel dimension is obtained by using the global average pooling layer. The statistical value is obtained by compressing the with spatial dimension by (2):
[0108] (2)
[0109] The transformation output can be interpreted as a set of local descriptors, and the statistical information of these descriptors can express the whole image. In order to be able to utilize the aggregated information obtained by the compression operation, the next goal is to completely capture the dependence in the channel dimension. A simple gating mechanism (3) with sigmoid activation function is selected:
[0110] (3)
[0111] In the formula, denotes a ReLU activation function, and To limit the complexity of the model and help the model generalize, the two fully connected layers (FC) around the nonlinearity are composed into a bottleneck structure to parameterize the gating mechanism, e.g., one dimension reduction layer with a parameter , one ReLU activation function and one dimension increase layer with a parameter The final output of the block is obtained by rescaling the transformed output using the activation function (4):
[0112] (4)
[0113] where , denotes the product of the corresponding channels of the feature map and the scalar s. The role of this activation function is to assign a weight to each channel according to the descriptor of the input feature.
[0114] Technical route and implementation of multi-layer feature fusion mechanism
[0115] Generally, concatenation or element-wise summation is the most common scheme for heterogeneous feature fusion. Since the distribution of audio and hand features usually varies greatly, and their feature size usually differs, the representation ability of these simple fusion schemes can not be sufficient to achieve reliable speaker naming performance. By decomposing the bilinear model (FBM) for fusion, the interaction between the two different modalities can be better captured, and it is usually superior to simple fusion methods (e.g., concatenation), as shown in Figure 4
[0116] The decomposed bilinear model considers each feature pair through a linear transformation:
[0117] (5)
[0118] where x e R n and y e R m are the input feature vectors from two different modalities (e.g., high-level features of hand and audio), W i is the weight matrix, and b i is the bias. Although the bilinear model can capture the pairwise relationship between the two modalities, it usually introduces a large number of parameters, which can lead to high computational cost. To solve this problem, an effective method is to factorize the weight matrix W i decomposed into two low-rank matrices, i.e., X = UVT and Y = U'VT where and A constraint d ≤ min(n, m) is imposed on the dimension d. Therefore, formula (5) can be further rewritten as:
[0119] (6)
[0120] Generally, the first term on the right side of the equation can be further converted with Hadamard product or element-wise multiplication to capture the inherent correlation between the two heterogeneous modalities:
[0121] (7)
[0122] where 1 ∈ R d denotes a column vector of 1, and denotes Hadamard or element-wise multiplication. To obtain the output feature vector z, two third-order tensors are required: U = [U1, …, UO] ∈ R n×d×o and V = [V1, …, V o ] ∈ R m×d×o . Replace the column vector with a linear projection P ∈ R d ×o , so the vector z can be represented as:
[0123] (8)
[0124] where b ∈ R o is a bias vector. The application of a nonlinear activation function generally helps to increase the representation capacity of the bilinear model. Therefore, a nonlinear activation function is added after each linear mapping, so the vector z can be further represented as:
[0125] (9)
[0126] where σ denotes an arbitrary nonlinear activation function, such as ReLU, Sigmoid, or tanh. Assuming x and y represent the hand attention vector and the audio feature vector, respectively, the values of x are all greater than 0, while y is in the range of [-1, 1]. In order to avoid information loss, different nonlinear activation functions can be used to map the values to a finite interval. Since element-wise multiplication is introduced to obtain the correlation between the two modalities, the size of the output neuron can change greatly. In order to reduce the impact of this change, a Relu function is further added to normalize the output of the network, and the final vector z can be represented as:
[0127] (10)
[0128] During the training process, the fusion parameters of the FBM can be updated and optimized through back propagation. The paired audio and hand feature inputs are input into the fusion model, and the final result is output through softmax on the fully connected layer.
[0129] Technical route and implementation of feature matching
[0130] The FAST algorithm is a relatively fast corner detection algorithm at present, but it will cause false detection of some edge points, resulting in the existence of some pseudo corner points. In order to exclude the interference of the edge points on the detection result, the improved FAST algorithm is adopted in the application, and the specific improvement is that 24 pixel points around a pixel point P are taken as a detection template, the gray value of P is IP, a threshold T is set, if the gray values of 14 continuous pixel points in the 24 pixel points are greater than IP+T or less than IP-T, then P is a corner point. The Shi-Tomasi algorithm is used for optimization of the feature points in the application, and the Shi-Tomasi algorithm compares the smaller one of two eigenvalues with a given minimum threshold value, if the smaller one is greater than the minimum threshold value, then a strong corner point is obtained.
[0131] The Shi-Tomasi algorithm detects a corner point by calculating the local small window The gray scale of the pixel points after moving in each direction is detected. The gray scale changes are generated
[0132] (10)
[0133] In the formula, M is a 2 correlation matrix, which can be calculated from the derivative of the image
[0134] (11)
[0135] Two characteristics of the matrix M are analyzed and , because the uncertainty of the curvature is dependent on the small corner point, the corner response function is defined as . The Shi-Tomasi algorithm is used to calculate the corner response function of each point of the feature points preliminarily extracted by the improved FAST corner detection algorithm, and the first N points with the maximum response value are determined as the feature points according to . The surrounding of the screened feature points at least exists 2 strong boundaries in different directions, and such feature points are easy to identify and stable.
[0136] For the matching of the binary feature description vector (such as Figure 5As shown in FIG. 2, the Hamming distance is generally used as a similarity measure between descriptors. The Hamming distance is the minimum number of substitutions needed to change one binary string to another. Given two feature vectors of a descriptor , , , the Hamming distance between them is:
[0137] (12)
[0138] The threshold of Hamming distance is determined to judge whether the feature vectors match.
[0139] The method according to the present application described above can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium such as a CD ROM, a RAM, a floppy disk, a hard disk or a magneto-optical disk, or be downloaded through a network originally stored in a remote recording medium or a non-transitory machine readable medium and stored in a local recording medium, so that the method described herein can be processed by such software using a general purpose computer, a special purpose processor, or programmable or special purpose hardware such as an ASIC or an FPGA. It is understood that the computer, the processor, the microprocessor controller, or the programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code that, when accessed and executed by the computer, the processor, or the hardware, implements the processing method described herein. In addition, when a general purpose computer accesses the code for implementing the processing shown herein, the execution of the code will convert the general purpose computer into a special purpose computer for executing the processing shown herein.
[0140] Those skilled in the art will understand that the embodiments described herein are for the purpose of helping the reader to understand the method of implementing the present application, and should be understood as the scope of protection of the present application is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific modifications and combinations according to the technical inspiration disclosed in the present application without departing from the essence of the present application, and these modifications and combinations are still within the scope of protection of the present application.
Claims
1. A non-contact multi-modal identity recognition method of voiceprint and palmprint palm vein, characterized in that, Comprise: Step 1, image preprocessing; Preprocessing mainly includes three steps, first, the palm image collected by infrared is denoised by low-pass filter, second, the binary image of palm region is extracted by Sauvola algorithm in image enhancement part, and finally, ROI positioning part is used to make the palm edge prominent by gray scale transformation of palmprint and palm vein, then Canny operator is used to detect the palm edge, and finally the image is cropped to obtain the image of the interested palm region; Step 2, feature extraction; Feature Extraction is divided into two parts, the first part is to extract speech features, and the second part is to extract palmprint and palm vein two hand features; ResNet is used as the main structure, SE module is introduced, and SE-ResNet network structure is constructed, the preprocessed picture is input into SE-ResNet network structure, a global pooling layer is added to generate feature distribution, and information coding extraction is completed; In order to obtain the correlation between channels, ReLU activation function and sigmoid gate control mechanism are combined to complete feature rescaling; Step 3, feature fusion; Multi-layer feature fusion mechanism is adopted, the interaction between hand and audio of different modalities is obtained by decomposing bilinear model, and the paired audio and hand features are input into fusion model, and the final result is output through softmax on full connection layer; Step 4, feature comparison; The feature points preliminarily extracted by improved FAST corner detection algorithm are calculated by Shi-Tomasi algorithm, and the corner response function of each point is calculated, and the feature points are determined according to the top N response values of the corner response function; For binary feature description vector matching, hamming distance is used as the similarity measure between descriptors; Step 5, output interaction; Joint discriminative sparse coding algorithm is used to judge the sample feature points in the class, so that the distance in the class is minimized, and the distance between classes is maximized; According to the actual scene requirements, set a suitable threshold, if two matched samples belong to the same class and are matched successfully in voiceprint, palmprint and palm vein, the interface displays authentication success, otherwise, it prompts authentication failure.
2. The multi-modal identity recognition method of claim 1, wherein: Step 2 is as follows: for any given information entering the network module, the conversion shown in formula (1) is carried out: (1) is an input picture, is an extracted feature; The SE compresses the global spatial information into a channel descriptor, which contains the global distribution of feature responses in the channel dimension, and obtains a statistical data in the channel dimension by using a global average pooling layer; the statistical value is obtained by compressing the feature responses with the spatial dimension of the formula (2) : (2) Transform output The statistics of the channel descriptors, interpreted as a collection of local descriptors, are able to express the whole image. The aggregated information obtained by compression operation completely captures the dependence on channel dimension; A simple threshold mechanism (3) with sigmoid activation function is selected: (3) wherein denotes a ReLU activation function, and ; To limit the complexity of the model and help the model generalize, two fully connected layers (FC) in the nonlinear part are combined into a bottleneck structure to parameterize the gating mechanism, and the final output of the block is rescaled by the activation function (4) The result is: (4) wherein , denotes the product of the corresponding channels of the feature map and the scalar ; s the activation function has the effect of assigning a weight to each channel depending on the descriptor of the input feature.
3. The multi-modal identity recognition method of claim 1, wherein: Step 3 is as follows: The decomposition bilinear model considers each feature pair through linear transformation: (5) where x e R n and y e R m are input feature vectors from different modalities of hand and audio, W i is a weight matrix, b i is a bias term; The weight matrix W i is decomposed into two low-rank matrices, i.e., W = UVT where and imposing a constraint d ≤ min(n, m) on the dimension d; equation (5) is further rewritten as: (6) Capture the inherent correlation between the two heterogeneous modes, formula (7): (7) where 1 ∈ R d denotes a column vector of ones, and denotes the Hadamard or element-wise product; To obtain the output feature vector z, two third-order tensors are needed: U = [U1,..., U O ] ∈ R n×d×o and V = [V1,..., V o ] ∈ R m×d×o ; with the linear projection P e R d×o The column vector is replaced by the vector z is represented as: (8) where b e R o is the bias vector; Add a nonlinear activation function after each linear mapping, and the vector z is further represented as: (9) Where sigma represents any nonlinear activation function, and x and y represent hand attention vector and audio feature vector respectively, then the value of x is greater than 0, and y is in the range of [-1, 1]; Further add Relu function to regulate the output of network, and the final vector z is represented as: (10) The paired audio and hand features are input into the fusion model, and the final result is output through softmax on full connection layer.
4. The multi-modal identity recognition method of claim 1, wherein: Step 4 is as follows: The improved FAST algorithm is adopted, and the specific improvement is that 24 pixel points around a pixel point P are taken as a detection template, the gray value of P is IP, a threshold T is set, if the gray values of 14 continuous pixel points in the 24 pixel points are greater than IP+T or less than IP-T, then P is a corner point; The Shi-Tomasi algorithm is used for optimization of the feature points, the Shi-Tomasi algorithm compares the smaller one of two eigenvalues with a given minimum threshold, if the smaller one is greater than the minimum threshold, then a strong corner point is obtained; The Shi-Tomasi algorithm detects corners by computing local window grayscale changes after moving in various directions; the window is translated to produce grayscale changes to (11) where M is 2 autocorrelation matrix calculated from the derivative of the image (12) For matching of binary feature descriptor vectors, Hamming distance is used as the similarity measure between descriptors; let two feature vectors of a descriptor be , then the Hamming distance between them is , (13) The threshold of Hamming distance is determined to judge whether the feature vectors are matched.
5. The multi-modal identity recognition method of claim 1, wherein: Step 5 is specifically as follows: The joint discriminative sparse coding algorithm is as follows: Given feature matrices X, Y, and Z of three modalities, three projection matrices P are jointly learned. x P y and P z Mapping the three peak features to a sparse matrix V x ∈R d ×N、V y ∈R d ×N and V z ∈R d ×N can accurately approximate the original matrices X, Y, Z; The feature representation V is obtained from the three modalities x , y , z After that, it is quantized into ; (14) where sgn() is the element-wise sign function, resulting in sparse binary codes, C x = [c 1 , c 2 ,..., c N ] e R l×N , c i denotes the learned sparse binary code for the i-th class, and l = (1, 2,..., 12) is the length of the sparse binary code. The projection feature representation V is projected using two projection matrices x and V y The sparsity constraint is applied to reduce the projection error, using the Frobenius norm as the cost function, and the engineering error is represented as (15) Wherein a, b > 0, a+b∈(0,1) is a balance parameter for balancing three different modes; Two constraints are performed on the projected sparse feature: 1) for the intra-class samples of each mode, the distance in the class is minimized, and the distance between classes is maximized; 2) for the intra-class samples, the information correlation between feature points is maximized, so the distance is minimized; through the constraints, the projected sparse feature has stronger resolution and compactness.
6. A non-contact multi-modal identity recognition system of voiceprint and palmprint palm vein, for implementing the multi-modal identity recognition method of any one of claims 1 to 5, characterized in that, It comprises: A power supply module, a fixed wavelength infrared LED light source module, an image acquisition CCD module, a voice acquisition module and a storage module; The power supply module is used for power supply of the whole multi-modal identity recognition system The fixed wavelength infrared LED light source module irradiates the human hand through the infrared LED light source, and assists the image acquisition CCD module to collect the palm print and palm vein information features of the human body; The image acquisition CCD module collects the palm print and palm vein information features of the human body; The voice acquisition module extracts the voice information by using the MFCC feature; The storage module is used for storing the data collected by the voice acquisition module and the image acquisition CCD module; The multi-modal identity recognition module outputs the result through picture preprocessing, picture feature extraction, feature fusion comparison.
Citation Information
Patent Citations
Non-contact gesture, palm print and palm vein fused identity recognition system and method
CN114220130A