Palm vein recognition method and system based on multi-spectral imaging and deep learning, and storage medium
By integrating near-infrared, short-wave infrared, and visible light band images through multispectral imaging and deep learning, high-quality vein feature maps are generated. Combined with a dual-branch deep learning network, the problems of image susceptibility to interference, unstable recognition, and weak anti-forgery ability in existing technologies are solved, achieving high-precision and fast identity recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 浙江微特电子信息有限公司
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-19
AI Technical Summary
Existing palm vein recognition technology is susceptible to ambient light interference, resulting in poor image quality, unstable recognition, weak anti-spoofing capabilities, and difficulty in balancing recognition speed and accuracy, especially in large-scale databases.
Multispectral imaging technology is used to simultaneously acquire images in the near-infrared, short-wave infrared, and visible light bands. High-quality vein feature maps are generated through image fusion and enhancement processing. A deep learning dual-branch network is combined to extract vein structure and dynamic features, and a multi-branch deep learning recognition model is constructed for liveness detection and identity recognition.
It significantly improves the quality and robustness of vein images, effectively defends against forgery attacks, balances recognition speed and accuracy, and is suitable for authentication in complex environments and large-scale databases.
Smart Images

Figure CN121600560B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biometric recognition, and is applicable to high-security identity authentication scenarios. Specifically, it relates to a palm vein recognition method, system, and storage medium based on multispectral imaging and deep learning. Background Technology
[0002] Currently, palm vein recognition technology has been widely used in finance, security and other fields due to its liveness detection characteristics, non-contact data acquisition and high security. However, existing technologies still have the following problems:
[0003] Single-wavelength imaging is susceptible to interference from ambient light, leading to a decrease in image quality;
[0004] Vein images are easily affected by temperature and blood circulation status, resulting in unstable recognition.
[0005] Traditional feature extraction methods are insufficient in defending against complex forgery attacks (such as silicone models and 3D printed hand molds);
[0006] It is difficult to balance recognition speed and accuracy, especially when dealing with large-scale databases.
[0007] Existing patents mostly focus on single imaging methods or simple image processing, lacking a systematic design for multimodal information fusion and anti-attack capabilities. For example, the published document with application number CN202410628111.0, entitled "A method and system for dynamic identification of living palm veins," collects a sequence of palm vein images of the subject at various time frames and introduces image processing and analysis algorithms in the backend to perform temporal feature analysis of the sequence of palm vein images of the subject. It uses the dynamic information in the palm vein image sequence to determine whether the subject is an authorized user. This one-way information judgment is poor in terms of recognition speed and corresponding recognition security, and has insufficient defense capabilities. Summary of the Invention
[0008] The purpose of this invention is to provide a palm vein recognition method, system, and storage medium based on multispectral imaging and deep learning, so as to solve the problems of poor image quality, weak anti-spoofing ability, and low recognition accuracy in the prior art.
[0009] To achieve the above objectives, the present invention provides the following technical solution: a palm vein recognition method based on multispectral imaging and deep learning, comprising the following steps:
[0010] S1: Acquire a multispectral image sequence of the palm, the multispectral image sequence including near-infrared band images, short-wave infrared band images and visible light band images acquired synchronously in time;
[0011] S2: Image fusion, performing image fusion and enhancement processing on the multispectral image sequence to obtain a high-quality vein feature map;
[0012] S3: Deep learning recognition, constructing a multi-branch deep learning recognition model, including a vein structure feature extraction branch and a dynamic feature extraction branch;
[0013] The vein structure feature extraction branch receives the input of the vein feature map and obtains a structural feature vector with vein topology information attached.
[0014] The dynamic feature extraction branch includes a blood flow dynamic feature extraction branch and a muscle micro-motion feature extraction branch. It receives time-series image data input from the multispectral image sequence. The blood flow dynamic feature extraction branch obtains a blood flow feature vector with a blood flow fraction distribution map with spatial alignment of the venous structure. The muscle micro-motion feature extraction branch obtains a muscle micro-motion feature vector with spatial distribution data of palm micro-motion amplitude.
[0015] The feature vectors output from different branches of the multi-branch deep learning recognition model are fused to obtain a fused feature vector, and liveness detection and identity recognition are performed based on the fused feature vector.
[0016] Preferably, the specific method for image fusion and enhancement processing of the multispectral image sequence includes the following steps:
[0017] The near-infrared band image, short-wave infrared band image and visible light band image are preprocessed, and the preprocessing includes noise suppression, non-uniform illumination correction and high-precision geometric registration.
[0018] Based on the preprocessed image, an adaptive weighted fusion algorithm is used to perform pixel-level fusion to generate a multispectral fused palm vein image;
[0019] The multispectral fused palm vein image is subjected to vein network enhancement processing to obtain the high-quality vein feature map.
[0020] As a preferred method, the pixel-level fusion method using an adaptive weighted fusion algorithm includes the following steps:
[0021] For each pixel in the registered image, its local characteristic index is calculated on the near-infrared band image, short-wave infrared band image and visible light band image respectively. The local characteristic index includes local contrast, gradient magnitude and local information entropy.
[0022] Based on the local characteristic index, a fusion weight is dynamically assigned to each pixel in each band image using a preset weight function.
[0023] The pixel value of each pixel in the fused image is obtained by weighted summation of the pixel values of each pixel in each band image and their corresponding fusion weights.
[0024] Preferably, the specific method for extracting blood flow feature vectors by the dynamic feature extraction branch includes the following steps:
[0025] Receive vein spatial guidance information provided by the vein structure feature extraction branch, and divide the palm image into multiple regions based on the guidance information;
[0026] For each region, the original photoplethysmography (PPG) signal for that region is extracted from the time-series image data;
[0027] The original photoplethysmography (PPG) signal was enhanced and denoised to obtain a pure physiological signal.
[0028] The physiological signal is input into a lightweight neural network, which analyzes the morphology, amplitude and dynamic characteristics of the physiological signal and outputs a blood flow fraction that characterizes the blood flow vitality and microcirculation perfusion efficiency in the region.
[0029] Based on the blood flow fraction in all regions, a blood flow fraction distribution map spatially aligned with the venous structure is generated, and the blood flow fraction distribution map is encoded into the blood flow feature vector.
[0030] Preferably, the specific method for extracting muscle micro-motion feature vectors via the muscle micro-motion feature extraction branch includes the following steps:
[0031] Motion amplification and signal separation are performed on the time-series image data to extract physiological tremor micromotion signals regulated by the central nervous system;
[0032] Multidimensional features are extracted from the micro-motion signals, including: frequency domain signature features of the dominant tremor features, spatial coherence topological features characterizing the correlation between micro-motion signals in different areas of the palm, and spatial distribution pattern features of micro-motion amplitude.
[0033] The multidimensional features are encoded into muscle micro-movement feature vectors with spatial distribution data of palm micro-movement amplitude.
[0034] As a preferred embodiment, the specific method for performing liveness detection and identity recognition based on the fused feature vector includes inputting the fused feature vector into a parallel liveness detection classification head and an identity recognition classification head, respectively.
[0035] The liveness detection classification head determines whether a person is a real live body based on the temporal coupling relationship between the blood flow feature vector and the muscle micro-movement feature vector extracted by the dynamic feature extraction branch, and outputs a binary classification result.
[0036] The identity recognition classification head integrates structural features and dynamic physiological features to achieve identity matching and output identity recognition results.
[0037] Preferably, the lightweight neural network includes a hybrid structure of a one-dimensional convolutional neural network and a long short-term memory network. The one-dimensional convolutional neural network is used to extract local morphological features of the physiological signal, and the long short-term memory network is used to model the long-range temporal dependencies of the signal.
[0038] Preferably, the identification results are also verified based on a security verification mechanism, which includes time consistency verification and micro-motion detection.
[0039] The temporal consistency check is used to verify whether the extracted physiological signals conform to biological rhythms; the micro-motion detection is used to analyze whether there are biological micro-motions during the imaging process;
[0040] The final result of successful identity recognition is only output when the liveness detection is true and the timing consistency check and micro-motion detection pass.
[0041] To address the aforementioned technical problems, this invention also provides a palm vein recognition system based on multispectral imaging and deep learning, comprising:
[0042] Memory, used to store computer programs;
[0043] A processor for executing the computer program, which, when executed by the processor, implements the steps of the palm vein recognition method based on multispectral imaging and deep learning as described in any of the preceding claims.
[0044] To address the aforementioned technical problems, the present invention also provides a readable storage medium having a computer program stored thereon.
[0045] When the computer program is executed by the processor, it implements the steps of the palm vein recognition method based on multispectral imaging and deep learning as described in any of the above.
[0046] In summary, the beneficial effects of this invention are:
[0047] This invention integrates multispectral simultaneous imaging technologies of near-infrared, short-wave infrared, and visible light, combined with a dual-branch deep learning network specifically designed for palm veins. This systematically addresses the shortcomings of existing technologies, such as image susceptibility to interference, unstable recognition, weak anti-counterfeiting capabilities, and the difficulty in balancing efficiency and accuracy. Its beneficial effects include: significantly improving the quality and robustness of vein images in complex environments through complementary and adaptive fusion of multi-band information; extracting and fusing multi-dimensional biometric features from static vein structure, dynamic blood flow characteristics, and muscle micro-movement features to achieve high-precision identity recognition while completely eliminating forgery attacks using photos, silicone molds, etc., through three-dimensional vein texture modeling and dynamic liveness detection anti-counterfeiting mechanisms; furthermore, the method, through efficient network design and feature fusion mechanisms, balances recognition speed and accuracy, making it particularly suitable for fast and secure identity verification scenarios with large-scale databases.
[0048] The present invention also provides a palm vein recognition system and storage medium based on multispectral imaging and deep learning, which has the above-mentioned beneficial effects, and will not be elaborated here. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of the invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is a schematic diagram of the workflow framework of the palm vein recognition method based on multispectral imaging and deep learning of the present invention.
[0051] Figure 2 This is a schematic diagram of a superficial vein image in the palm vein recognition method based on multispectral imaging and deep learning of the present invention.
[0052] Figure 3 This is a schematic diagram of a deep vein image in the palm vein recognition method based on multispectral imaging and deep learning of the present invention.
[0053] Figure 4 This is a comparative schematic diagram of the vein structure in the palm vein recognition method based on multispectral imaging and deep learning of this invention. Detailed Implementation
[0054] The present invention will now be described in further detail with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. These drawings are simplified schematic diagrams, which are only used to illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0055] To facilitate understanding of the present invention, a more complete description of the invention will be given below with reference to the accompanying drawings, which illustrate several embodiments of the invention. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of the invention will be more thorough and complete.
[0056] All features disclosed in this specification, or steps in all methods or processes disclosed herein, may be combined in any way, except for mutually exclusive features and / or steps.
[0057] Any feature disclosed in this specification (including any appended claims, abstract, and drawings) may be replaced by other equivalent or similar features for a similar purpose, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.
[0058] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection, a direct connection, or an indirect connection through an intermediate medium; they can refer to the internal communication of at least two elements or the interaction relationship of at least two elements, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0059] The following is combined Figures 1-4 The present invention will be described in detail below;
[0060] An overall process reference for an embodiment of the present invention is provided. Figure 1 A palm vein recognition method based on multispectral imaging and deep learning includes the following steps:
[0061] Step 1: Acquisition of multi-dimensional image information data of the palm;
[0062] refer to Figure 2 and Figure 3The high-performance data acquisition front end is responsible for acquiring raw palm images containing rich physiological and structural information. Its core objective is to simultaneously acquire high-resolution, high signal-to-noise ratio multi-dimensional optical information of the palm in complex environments, providing the best data source for subsequent vein feature extraction and liveness detection.
[0063] Specifically, simultaneous imaging in three bands—near-infrared (NIR), short-wave infrared (SWIR), and visible light (VIS)—is employed, wherein:
[0064] NIR bands: Primarily capture the structural shadows of the superficial subcutaneous venous network and are a major source of venous mainline features. (Reference) Figure 2 ;
[0065] Hemoglobin, especially deoxyhemoglobin, has a significant absorption peak in this band, while the absorption of surrounding tissues is relatively low. Therefore, blood-filled veins absorb NIR light strongly and reflect / transmit weakly, appearing as dark continuous lines in the image.
[0066] SWIR band: Sensitive to moisture, it can penetrate deeper tissues to obtain information on deep veins and microvessels, and complements the NIR band, reducing the impact of differences in skin condition. (Reference) Figure 3 ;
[0067] This wavelength band is the absorption band of water, and water in biological tissues strongly absorbs SWIR light. There is a difference in water content between venous blood and surrounding tissues. Simultaneously, SWIR light has stronger tissue penetration. Therefore, SWIR images can reveal deeper and finer veins, are less affected by skin surface conditions, provide complementary structural information to NIR, and enhance the richness and stability of venous features.
[0068] VIS band: Acquires information on the texture, shape, and color of the palm surface, providing key surface biological features for liveness detection.
[0069] The sensors corresponding to the three wavebands mentioned above receive the light signals at the same time and are triggered by the same clock to ensure that the three images of NIR, SWIR and VIS are completely consistent in time, perfectly capturing the physiological state of the palm at the same instant, and providing a basis for subsequent analysis of instantaneous blood flow distribution and accurate image fusion.
[0070] Simultaneously, a built-in feedback algorithm automatically and independently adjusts the light source intensity and sensor gain of the three bands based on the brightness of the initially acquired image. This ensures that properly exposed and detailed original images are obtained under various ambient lighting conditions and on users' palms with different skin tones. Ultimately, a set of spatially precisely registered and temporally synchronized "three-band palm image sequences" is output. This set of data not only includes static vein structure information (NIR+SWIR) for identification, but also dynamic temporal information (multi-frame NIR / SWIR sequences) for liveness detection and blood flow analysis, as well as surface anti-counterfeiting information (VIS).
[0071] The three-band synchronous / quasi-synchronous imaging mechanism ensures the consistency of different spectral information in time and space, laying a solid foundation for subsequent multi-dimensional information fusion and dynamic analysis, and effectively overcoming the defect that a single wavelength is susceptible to interference from ambient light.
[0072] Step 2: Image fusion and enhancement;
[0073] The core task is to transform the raw, heterogeneous, and multi-dimensional image data obtained in the first step into a high-quality feature map with ultra-high vein contrast, rich details, and spatial alignment, as well as a set of temporally aligned multispectral image sequences that can be used for dynamic analysis.
[0074] Specifically, it includes the following steps:
[0075] 1. Preprocessing: "Purify" and "align" the raw data to create optimal conditions for fusion.
[0076] Noise suppression: Adaptive filtering algorithms are used to address the noise characteristics of sensors in different bands, such as salt-and-pepper noise in CMOS, thermal noise in InGaAs, and imaging depth.
[0077] For NIR and VIS images (with high signal-to-noise ratio), bilateral filtering or nonlocal mean filtering can be used to smooth noise while preserving the edge vein contours relatively well. For SWIR images with low signal-to-noise ratio, more powerful denoising methods such as wavelet transform or block matching 3D filtering may be used to extract weak deep vascular signals.
[0078] Non-uniform illumination correction:
[0079] Due to the influence of the ring light source layout and the curved surface of the palm, the image often presents an uneven illumination field with a bright center and dark edges, which will seriously interfere with the extraction of vein features.
[0080] By using large-scale Gaussian filtering or morphological top-hat transformation, the slowly varying illumination background components are estimated from the original image.
[0081] Divide the original image by the estimated background image, or perform homomorphic filtering to transform the multiplicative / additive uneven illumination model into an additive model and filter it out, thereby obtaining a palm image with uniform illumination.
[0082] High-precision geometric registration: NIR, SWIR, and VIS images acquired from different sensors or at different time points have slight differences in translation, rotation, and scaling. Subpixel-level alignment is necessary to achieve effective pixel-level fusion. The specific process is as follows:
[0083] Feature point detection and matching: Detect stable feature points, such as SIFT, ORB, or deep learning keypoints, on VIS and NIR / SWIR images, typically using high-resolution VIS or NIR images as a benchmark.
[0084] Transformation model estimation: Using matching point pairs, calculate an affine or projective transformation matrix to describe the spatial relationships between images.
[0085] Image resampling and alignment: Based on the transformation matrix, SWIR and another image are resampled to ensure complete pixel-level alignment with the reference image. This step ensures that features such as vein structures and palm print contours correspond strictly in space.
[0086] 2. Adaptive weighted fusion
[0087] Regional characteristic analysis and weight map generation:
[0088] Measurement index calculation: For each registered pixel (i,j), calculate a set of local characteristic indices that reflect its "information value" on the three band images respectively:
[0089] Local contrast: measures the difference in brightness between the area surrounding the point. At the edges of veins, the local contrast in NIR / SWIR images is high.
[0090] Gradient magnitude: Reflects the edge intensity of an image at that point. Vein boundaries have high gradient values.
[0091] Local information entropy: measures the complexity and information content of the texture in a region. VIS image regions with rich palm prints have higher entropy values.
[0092] Dynamic weight allocation: Based on these real-time calculated metrics, a fusion weight W_band(i,j) is assigned to each pixel in each band through a predefined weight function, such as rule-based or trained small neural networks.
[0093] The core logic is: in areas where veins are clearly visible (high NIR / SWIR contrast), NIR and SWIR images are given higher weights; in areas with palm surface texture, contours, or reflective areas of suspected counterfeit materials, the weight of VIS images is increased; and in areas with blurred images or high noise, the weight of that band is reduced.
[0094] Weight map optimization: The generated initial weight map may not be smooth, so guided filtering or small-scale Gaussian smoothing is required to avoid blocky artifacts in the fusion result and ensure a natural transition.
[0095] Multispectral pixel-level fusion:
[0096] For each output pixel I_fused(i,j), its value is obtained by weighted summation of the input pixel values of the three bands:
[0097] I_fused(i,j)=W_NIR(i,j)*I_NIR(i,j)+W_SWIR(i,j)*I_SWIR(i,j)+W_VIS(i,j)*I_VIS(i,j)
[0098] And W_NIR+W_SWIR+W_VIS=1.
[0099] The final generated "multispectral fusion palm vein image" integrates the clear outline of superficial veins, supplementary information of deep vessels, and surface context for liveness detection, resulting in a global and adaptive improvement in the contrast of the vein network relative to the background.
[0100] 3. Enhanced venous network: to maximize the prominence of the target feature—veins.
[0101] Multi-scale Gabor filtering is employed.
[0102] The Gabor filter is a direction- and scale-adjustable bandpass filter whose kernel function is highly matched to the strip-like, tubular structure and local directionality of blood vessels.
[0103] The fused image is convolved using a set of Gabor filters at multiple scales, corresponding to blood vessels of different thicknesses, and at multiple directions, such as 0°, 45°, 90°, and 135°.
[0104] For each pixel, the maximum value among the responses of all orientation and scale filters is taken as the enhanced output value for that pixel. This ensures that blood vessels of different orientations and thicknesses can be effectively enhanced.
[0105] The output is a vein enhancement response map, in which blood vessels are presented as bright lines and background tissue is significantly suppressed.
[0106] Finally, post-processing and optimization are performed, mainly by applying morphological operations such as closing operations, connecting blood vessel breakpoints, and refining lines.
[0107] Finally, output the following two key data points:
[0108] High-quality vein feature map: This refers to the enhanced vein image after the complete processing described above. It will serve as the direct input to the vein structure feature extraction branch in the deep learning recognition model. (See reference...) Figure 4 The figure shows the results of vein image extraction for different age groups. In each sub-figure, the age gradually increases from (a) to (d), and the circles mark the vein feature structures of specific areas.
[0109] Preprocessed and aligned multispectral time-series image sequences: The preprocessed NIR, SWIR, and VIS image sequences, such as those that have undergone denoising, correction, and registration, are packaged and output in time frames.
[0110] Step 3: Deep learning for recognition;
[0111] 1. Construct a dual-branch convolutional neural network (CNN), with one branch processing vein structure features and the other extracting dynamic features of the palm, including blood flow dynamic features and muscle micro-movement features, to achieve integrated liveness detection and identity recognition.
[0112] Specifically, for branch one: processing venous structural features;
[0113] The input mainly comes from the high-quality vein feature map generated in the second step, and uses a mature, pruned and optimized deep convolutional neural network (CNN) as the backbone, such as a variant of ResNet34 / 50 or EfficientNet-B3;
[0114] Customization was made for the characteristics of vein images: large-scale downsampling layers designed for ImageNet classification in the original network were removed to prevent the loss of fine vein texture; convolutional kernels with smaller receptive fields were used in the shallow network to capture the fine direction and bifurcation points of blood vessels;
[0115] Introducing attention mechanisms, such as the Squeeze-and-Excitation module or CBAM, allows the network to automatically focus on vascularized ROI regions and suppress irrelevant background.
[0116] Finally, the network output features: at the end of the network, a global average pooling layer is used to compress the extracted two-dimensional spatial feature map into a high-dimensional, such as 512 or 1024-dimensional, "structural feature vector" (Fs), which encodes the unique venous topology information of an individual.
[0117] Specifically, the blood flow dynamic feature extraction branch of branch two can be divided into three sub-stages:
[0118] Phase A: ROI localization and spatiotemporal signal preprocessing;
[0119] Input: A preprocessed and aligned multispectral temporal image sequence from step 2, for example, 30 frames / second, lasting 2 seconds, for a total of 60 frames.
[0120] ROI delineation guided by vascular region:
[0121] Structural Prior Injection: To perform physiologically meaningful region analysis, this branch receives a venous spatial attention heatmap or coarse segmentation map from an intermediate layer of branch one, rather than the final output, as guidance. This ensures that blood flow analysis focuses on real vascular regions, rather than random image patches.
[0122] Meshing and Adaptive Partitioning: Guided by the spatial structure of the main veins, the palm image is divided into a series of overlapping or non-overlapping grid units. The grid density can be adaptively adjusted according to the vein density, using a finer grid in areas with dense blood vessels.
[0123] Spatiotemporal cube construction: For each divided spatial unit, extract the pixel values of all frames in that region from the entire temporal image sequence to form a spatiotemporal data cube of [T,H,W], where T is the number of time frames, and H and W are the height and width of the unit.
[0124] Phase B: Physiological signal generation and purification;
[0125] Original signal extraction reflecting blood volume pulse wave (PPG): For each spatiotemporal cube, calculate the average gray value of all pixels inside it at each time point, thereby obtaining an original time series signal of length T - that is, the original photoplethysmography (PPG) signal of that region.
[0126] Signal enhancement and noise reduction: Use bandpass filtering, such as 0.5Hz~5Hz, corresponding to a heart rate of 30-300BPM, to filter out high-frequency sensor noise and extremely low-frequency baseline drift, such as that caused by breathing or slow movement.
[0127] By employing blind source separation technology or multi-channel adaptive filtering, interference signals related to global light source fluctuations and rigid hand movements are separated, while pure physiological PPG components caused by local blood volume changes are preserved.
[0128] Phase C: Blood flow fraction map generation, assigning each region a quantified "blood flow vitality" score;
[0129] The overall architecture is as follows:
[0130] Input: The purified PPG signal fragments, which may be padded with zeros or interpolated to a fixed length, are usually spliced together with the static texture features of the region, which may be shallow CNN features extracted from a single frame image of the corresponding spatial unit, to form a fused input.
[0131] Network Design: This sub-network is a lightweight one-dimensional convolutional neural network (1D-CNN) combined with a recurrent neural network (RNN / LSTM) or a Transformer encoder.
[0132] 1D-CNN: Quickly extracts local morphological features of PPG signals, such as peaks, troughs, and rising slope.
[0133] LSTM / Transformer: Models the long-range temporal dependencies of signals, capturing the dynamic characteristics and periodic rhythm within the heartbeat cycle.
[0134] Output and supervision: The network finally outputs a scalar value, Score_flow, which represents the microcirculation perfusion efficiency and blood flow vitality of the region. For example, normalized to the [0,1] interval, it represents the "blood flow fraction" of the region. This score is a composite index that is theoretically positively correlated with local blood flow perfusion rate, vasodilation degree, and pulse wave intensity.
[0135] The monitoring signal is acquired through one of the following methods:
[0136] Joint calibration: During the data collection phase, medical equipment such as laser Doppler flowmeters were used to measure the corresponding areas of the volunteers' palms to obtain approximate values as supervisory labels.
[0137] Self-supervised / weakly supervised: Using a large amount of unlabeled time-series data, design self-supervised tasks, such as next frame prediction and heart rate estimation, to pre-train the network, and then fine-tune the regression head using a small amount of labeled data.
[0138] Physiological model derivation: Based on the PPG signal waveform, some alternative indicators, such as perfusion index and vascular elasticity index, were calculated using classical physiological models, such as second-order differential wave analysis, as pseudo-labels.
[0139] Atlas formation: The above process is performed on all grid cells of the entire palm to obtain a blood flow fraction distribution atlas that corresponds one-to-one with the spatial position of the palm. Then, a lightweight CNN encoder encodes the atlas into a compact blood flow feature vector (Fd).
[0140] For the muscle micromovement feature extraction branch in branch two, the aim is to capture and quantify a frequently overlooked but highly valuable living biomarker: involuntary physiological tremors of the hand regulated by the central nervous system. This feature forms an orthogonal and complementary validation dimension with blood flow dynamics regulated by the cardiovascular system. Specifically, it includes the following steps:
[0141] Input: Shares the same high-frame-rate multispectral time-series image sequence as the blood flow dynamics feature extraction branch, but focuses on analyzing different bands:
[0142] VIS band sequence: It is most sensitive to minute displacements of surface texture and is an ideal channel for observing skin surface tremors.
[0143] NIR band sequences can penetrate the epidermis, capture minute movements of subcutaneous tissue and veins, and provide motion information related to internal physiological activities.
[0144] Step 1: Subpixel-level micro-motion signal separation and extraction;
[0145] Motion magnification: First, Euler video magnification or a deep learning-based motion magnification algorithm is applied to the input VIS / NIR time-series image sequence. This technique can magnify micro-motions invisible to the naked eye to an analyzable amplitude, explicitly revealing tremor patterns.
[0146] Signal decoupling: The amplified video sequence contains mixed signals.
[0147] Frequency domain filtering: Design a digital bandpass filter to directly retain the target frequency band, the main distribution area of physiological tremor, and filter out low-frequency components such as heartbeat and breathing, as well as high-frequency noise.
[0148] Blind Source Separation (BSS): A better approach is to use blind source separation techniques such as independent component analysis (ICA) to treat the temporal signal of each pixel as a linear mixture of multiple source signals and separate them using the ICA algorithm. Then, the independent component that best represents physiological tremor is automatically identified and selected.
[0149] Step 2: Multi-dimensional micro-motion feature engineering;
[0150] Three core features were extracted from the isolated pure "micro-motion signal field":
[0151] Frequency domain signature features:
[0152] Perform short-time Fourier transform or continuous wavelet transform on the micro-motion signal of each pixel region or analysis unit.
[0153] The dominant tremor frequency, frequency band energy ratio, and frequency stability are extracted. These characteristics are related to the individual's neurophysiological basis, have long-term stability, and can undergo characteristic drift due to states such as anxiety and fatigue, and can themselves serve as a basis for assessing the in vivo state.
[0154] Spatial coherence topological features:
[0155] Divide the palm into M regions of interest.
[0156] Calculate the coherence coefficient or phase lock value between the micro-motion signals of any two regions ROI_i and ROI_j to form an M×M coherence matrix.
[0157] The tremors of a real, living hand originate from a common neural drive source, so the micro-motion signals in different muscle groups will exhibit high spatial coherence and specific phase delay patterns. In contrast, the "tremors" of different parts of a prosthetic hand, simulated by a mechanical vibrator or driven by an independent motor, are independent or uncorrelated, and their coherence matrices will exhibit drastically different patterns.
[0158] Amplitude distribution pattern characteristics:
[0159] Calculate the root mean square amplitude of the entire micro-motion signal field in the time dimension to generate a heat map of the spatial distribution of the micro-motion amplitude.
[0160] This heat map characterizes the intensity distribution of tremors in different parts of the palm, which is related to an individual's muscle tension distribution, hand grip habits, and even neural control characteristics, forming another unique spatial pattern feature.
[0161] Step 3: Encoding of micro-motion feature vectors;
[0162] Feature aggregation: Aggregate the frequency domain feature vectors, spatial coherence matrix and amplitude distribution heatmap extracted above.
[0163] Encoding Network: A lightweight convolutional-fully-encoding network or graph neural network is used to process these aggregate features, and GNNs are used to process the inter-region relationship graph represented by the coherence matrix.
[0164] Output: The encoding network outputs a fixed-length "micro-motion feature vector FM", which reflects the user's neuromuscular control.
[0165] 2. Feature fusion;
[0166] The structural feature vector FS obtained from branch one is fused with the blood flow feature vector Fd and the micromotion feature vector FM obtained from branch two. The fusion method is to use cross-attention fusion or gating fusion mechanism.
[0167] For example, one learning logic is: "When the uncertainty of blood flow features is high, refer more to micromotion features," and "locally align and fuse the venous structure region with the corresponding blood flow and micromotion features" to generate a fused feature F_fused. This fusion can uncover deep correlations between the three modalities.
[0168] The shared fusion feature F_fused is fed into two parallel fully connected layer classification heads:
[0169] Liveness detection head: a binary classification output layer;
[0170] Rule 1 (Cardiovascular Validation): The blood flow feature vector Fd must conform to the physiological laws of living cardiovascular disease (such as the existence of a reasonable periodic PPG waveform).
[0171] Rule 2 (Neuromuscular Validation): The micromotion feature vector FM must exhibit characteristics of a living neuromuscular system, show physiological tremors in a specific frequency band, and have biologically reasonable spatial coherence.
[0172] Rule 3 (Cross-system Coupling Verification - Ultimate Defense): This is the highest level of verification, analyzing whether there is a weak cross-frequency coupling or correlation between blood flow signals (originating from heartbeats) and micro-motion signals (originating from nerve oscillations).
[0173] For example, is there a phenomenon where "the amplitude of micro-movements is slightly modulated with a specific phase of the heartbeat cycle or blood pressure wave"? This spontaneous, non-artificial coupling relationship between different physiological systems is a manifestation of the high complexity of living organisms and is almost impossible to replicate or simulate by any external mechanical system.
[0174] Identity Recognition Head: Integrating static structure and dynamic physiological features to achieve high-precision identity matching. The features used to distinguish individuals include anatomical structure, cardiovascular function atlas, and neuromuscular control signature. This three-dimensional integrated biometric template has an exponentially increasing uniqueness and complexity, enabling it to distinguish users with extreme accuracy in large-scale databases and resist the most extreme "cooperative forgery attacks," which simultaneously attempt to mimic the target's veins, blood flow, and micro-motion patterns.
[0175] Step 4: Security verification mechanism, introducing time-series consistency checks and micro-motion detection to prevent static image or model forgery attacks. Specifically, one or more of the following mechanisms can be used:
[0176] Timing consistency verification: Verify whether the collected time-series physiological signals, such as PPG waveforms, exhibit natural and coherent rhythmic characteristics that conform to the laws of life, in order to defend against video playback attacks.
[0177] Micro-motion detection: Analyze whether there are unconscious, sub-pixel-level biological micro-movements in the palm during the imaging process to defend against attacks from high-precision static models.
[0178] Synergy with deep learning models: This mechanism forms cross-validation with the liveness detection branch in the recognition model. The system is only successfully authenticated when the deep learning model determines that the subject is alive and this mechanism passes the verification. This builds a two-layer active defense system.
[0179] The foregoing detailed the embodiments of the palm vein recognition method based on multispectral imaging and deep learning. Based on this, the present invention also discloses a palm vein recognition system and storage medium based on multispectral imaging and deep learning corresponding to the above method.
[0180] A palm vein recognition system based on multispectral imaging and deep learning includes:
[0181] Memory, used to store computer programs;
[0182] A processor is configured to execute the computer program, which, when executed by the processor, is capable of implementing the relevant steps in the palm vein recognition method based on multispectral imaging and deep learning disclosed in any of the foregoing embodiments.
[0183] The processor may include one or more processing cores, such as a core processor. The processor can be implemented using at least one of the following hardware forms: Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor may also include a main processor and coprocessors. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessors are low-power processors used to process data in the standby state.
[0184] In some embodiments, the processor may integrate a Graphics Processing Unit (GPU) responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor may also include an Artificial Intelligence (AI) processor for handling computational operations related to machine learning.
[0185] The memory may include one or more readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory is used to store at least the following computer program, which, after being loaded and executed by the processor, is capable of implementing the relevant steps in the palm vein recognition method based on multispectral imaging and deep learning disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory may also include an operating system and data, and the storage method may be temporary or permanent storage. The operating system may be Windows. The data may include, but is not limited to, the data involved in the above methods.
[0186] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules can be implemented in hardware or as software functional modules. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in the various embodiments of the present invention.
[0187] To this end, embodiments of the present invention also provide a readable storage medium storing a computer program, which, when executed by a processor, implements steps such as the palm vein recognition method based on multispectral imaging and deep learning.
[0188] The readable storage medium may include: USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program code.
[0189] The computer program contained in the readable storage medium provided in this embodiment can implement the steps of the palm vein recognition method based on multispectral imaging and deep learning as described above when executed by a processor, with the same effect.
[0190] The foregoing has provided a detailed description of the palm vein recognition method, system, and storage medium based on multispectral imaging and deep learning provided by this invention. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus, devices, and readable storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from the principles of the invention, and these improvements and modifications also fall within the protection scope of the claims of this invention.
[0191] The above description is merely a specific embodiment of the invention, but the scope of protection of the invention is not limited thereto. Any variations or substitutions conceived without inventive effort should be included within the scope of protection of the invention. Therefore, the scope of protection of the invention should be determined by the scope defined in the claims.
Claims
1. A palm vein recognition method based on multispectral imaging and deep learning, characterized in that: Includes the following steps: S1: Acquire a multispectral image sequence of the palm; S2: Image fusion, performing image fusion and enhancement processing on the multispectral image sequence to obtain a vein feature map; S3: Deep learning recognition, constructing a multi-branch deep learning recognition model, including a vein structure feature extraction branch and a dynamic feature extraction branch; The vein structure feature extraction branch receives the input of the vein feature map and obtains a structural feature vector with vein topology information attached. The dynamic feature extraction branch includes a blood flow dynamic feature extraction branch and a muscle micro-motion feature extraction branch. It receives time-series image data input from the multispectral image sequence. The blood flow dynamic feature extraction branch obtains a blood flow feature vector with a blood flow fraction distribution map with spatial alignment of the venous structure. The muscle micro-motion feature extraction branch obtains a muscle micro-motion feature vector with spatial distribution data of palm micro-motion amplitude. The muscle micro-motion feature extraction branch performs motion amplification and signal separation on the time-series image data to extract physiological tremor micro-motion signals regulated by the central nervous system. It also extracts frequency domain signature features, spatial coherence topological features characterizing the correlation between micro-motion signals in different areas of the palm, and spatial distribution pattern features of micro-motion amplitude from the micro-motion signals, and encodes them to obtain a muscle micro-motion feature vector with spatial distribution data of palm micro-motion amplitude. The feature vectors output from different branches of the multi-branch deep learning recognition model are fused to obtain a fused feature vector. The fused feature vector is then input into a parallel liveness detection classification head and an identity recognition classification head. The liveness detection classification head determines whether the subject is a real live subject based on the temporal cross-physiological system coupling relationship between the blood flow feature vector and the muscle micro-movement feature vector, and outputs a binary classification result. The identity recognition classification head integrates structural features and dynamic physiological features to achieve identity matching and output an identity recognition result.
2. The palm vein recognition method based on multispectral imaging and deep learning according to claim 1, characterized in that: The multispectral image sequence includes near-infrared images, short-wave infrared images, and visible light images acquired synchronously in time. The specific method for image fusion and enhancement processing of the multispectral image sequence includes the following steps: The near-infrared band image, short-wave infrared band image and visible light band image are preprocessed, and the preprocessing includes noise suppression, non-uniform illumination correction and high-precision geometric registration. Based on the preprocessed image, an adaptive weighted fusion algorithm is used to perform pixel-level fusion to generate a multispectral fused palm vein image; The multispectral fused palm vein image is subjected to vein network enhancement processing to obtain a high-quality vein feature map.
3. The palm vein recognition method based on multispectral imaging and deep learning according to claim 2, characterized in that: The method for pixel-level fusion using an adaptive weighted fusion algorithm includes the following steps: For each pixel in the registered image, its local characteristic index is calculated on the near-infrared band image, short-wave infrared band image and visible light band image respectively. The local characteristic index includes local contrast, gradient magnitude and local information entropy. Based on the local characteristic index, a fusion weight is dynamically assigned to each pixel in each band image using a preset weight function. The pixel value of each pixel in the fused image is obtained by weighted summation of the pixel values of each pixel in each band image and their corresponding fusion weights.
4. The palm vein recognition method based on multispectral imaging and deep learning according to claim 1, characterized in that: The specific method for extracting blood flow feature vectors via the blood flow dynamic feature extraction branch includes the following steps: Receive vein spatial guidance information provided by the vein structure feature extraction branch, and divide the palm image into multiple regions based on the guidance information; For each region, the original photoplethysmography (PPG) signal for that region is extracted from the time-series image data; The original photoplethysmography (PPG) signal was enhanced and denoised to obtain a pure physiological signal. The physiological signal is input into a lightweight neural network, which analyzes the morphology, amplitude and dynamic characteristics of the physiological signal and outputs a blood flow fraction that characterizes the blood flow vitality and microcirculation perfusion efficiency in the region. Based on the blood flow fraction in all regions, a blood flow fraction distribution map spatially aligned with the venous structure is generated, and the blood flow fraction distribution map is encoded into the blood flow feature vector.
5. The palm vein recognition method based on multispectral imaging and deep learning according to claim 4, characterized in that: The lightweight neural network includes a hybrid structure of a one-dimensional convolutional neural network and a long short-term memory network. The one-dimensional convolutional neural network is used to extract local morphological features of the physiological signal, and the long short-term memory network is used to model the long-range temporal dependencies of the signal.
6. The palm vein recognition method based on multispectral imaging and deep learning according to claim 1, characterized in that: It also includes verifying the identification results based on a security verification mechanism, which includes timing consistency verification and micro-motion detection; The temporal consistency check is used to verify whether the extracted physiological signals conform to biological rhythms; the micro-motion detection is used to analyze whether there are biological micro-motions during the imaging process. The final result of successful identity recognition will only be output when the liveness detection is true and the timing consistency check and micro-motion detection pass.
7. A palm vein recognition system based on multispectral imaging and deep learning, characterized in that: include Memory, used to store computer programs; A processor for executing the computer program, wherein the computer program, when executed by the processor, implements the steps of the palm vein recognition method based on multispectral imaging and deep learning as described in any one of claims 1-6.
8. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the palm vein recognition method based on multispectral imaging and deep learning as described in any one of claims 1-6.