An ultrasonic, infrared combined palm print and palm vein recognition method, device and apparatus

By combining ultrasound and infrared dual-sensor fusion technology with deep learning and traditional methods, palm print and palm vein features are extracted, solving the problem of existing systems being vulnerable to fake hand attacks and achieving higher recognition accuracy and security.

CN120220196BActive Publication Date: 2025-12-30NINGBO XINRAN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510273336.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-12-30
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Existing palm print and palm vein recognition systems are vulnerable to attacks from high-precision 3D prosthetic hands, resulting in insufficient recognition accuracy.

Method used

A method combining ultrasound and infrared was used to acquire palm print and palm vein images through dual-sensor fusion technology. Regions of interest (ROIs) were extracted using deep learning and traditional methods. Multimodal feature extraction and classification were then performed by combining a feature encoder and a feature fusion network.

Benefits of technology

It improves the security and accuracy of the identification system, reduces the risk of fake hand attacks, enhances the ability to distinguish real biometric features, and adapts to different environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220196B_ABST
    Figure CN120220196B_ABST
Patent Text Reader

Abstract

The application discloses an ultrasonic and infrared combined palm print and palm vein identification method, equipment and device, relates to the field of biological information identification and identity authentication, and comprises the following steps: step 1, collecting original images, including a palm vein image and an ultrasonic palm print image; step 2, performing ROI extraction, identifying the region of palm biological feature information contained in the original image, and obtaining two ROIs, namely the ROI of the palm vein image and the ROI of the ultrasonic palm print image; step 3, inputting the ROI of the palm vein image into a feature encoder 1 and the ROI of the ultrasonic palm print image into a feature encoder 2, and respectively extracting palm vein features and palm print features; step 4, in the process of feature extraction, synchronously performing bidirectional feature fusion, and finally obtaining multi-scale features; and step 5, inputting the multi-scale features into a classifier for classification identification, and outputting a classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biometric identification and identity authentication, and in particular to a method, device, and apparatus for palm print and palm vein recognition using a combination of ultrasound and infrared. Background Technology

[0002] With the rapid development of information technology, people have a more urgent need for security and convenience. Biometric technology is gradually entering people's lives and work, and is being applied in important areas such as financial payments, smart security, and information security. Palmprint recognition and palm vein recognition are both biometric technologies.

[0003] Palmprint recognition utilizes the unique texture features of the human palm, such as wrinkles, lines, and protrusions. A camera captures images of the palm, and specialized algorithms extract and analyze these features. The captured palmprint features are then compared to existing palmprint templates to identify an individual. However, palmprints are epidermal features of the palm, making them easily obtained and forged by criminals.

[0004] Palm vein recognition is a biometric technology that uses the uniqueness of palm veins for identification. Its main principle is to take advantage of the difference between the absorption characteristics of deoxygenated hemoglobin in veins and the absorption of near-infrared light by other physiological tissues (skin, fat). This makes the palm appear darker in the subcutaneous veins when illuminated by near-infrared light, thereby obtaining information about the veins.

[0005] Combining palm prints and palm veins as a biometric technology offers advantages such as high security, high accuracy, immunity to environmental influences, high-speed recognition, and biometric diversity, making it suitable for various identity verification fields that require high security and high accuracy.

[0006] 1. High Security: Combining palm print and palm vein biometrics can improve the security of the identification system. Palm print and palm vein features are not easily imitated or forged, and are difficult to steal, increasing the identification system's resistance to attacks and fraud.

[0007] 2. High accuracy: Combining palmprint and palm vein recognition technologies improves recognition accuracy. The combination of these two biometric features provides a more comprehensive description of an individual's body structure, reducing false recognition rates and enhancing the precision and reliability of the identification process.

[0008] 3. Unaffected by environmental factors: Palm print and palm vein biometrics are unaffected by environmental factors during the recognition process. Whether in humid, dry, or low-light conditions, palm print and palm vein recognition technology can reliably perform identification.

[0009] 4. Biometric Diversity: Palm prints and palm veins are different biometric features. Combining them increases the feature diversity of the biometric system. This improves the system's robustness, allowing identification to be completed even if some features are damaged or incomplete, using other biometric features.

[0010] While existing palmprint and palm vein recognition systems have achieved high accuracy in acquiring palmprint and palm vein images using a single sensor with binocular cameras, they still remain vulnerable to attacks by highly sophisticated 3D prosthetic hands. Such attacks could exploit this vulnerability by generating prosthetic hands with palmprints and palm veins that closely resemble those of a real hand using high-fidelity 3D printing or simulation techniques, thus deceiving the recognition system.

[0011] Therefore, those skilled in the art are dedicated to developing a new palm print and palm vein recognition method, device, and apparatus to solve the aforementioned problems in the prior art. Summary of the Invention

[0012] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is how to improve the accuracy of palm print and palm vein recognition methods and reduce the risk of being attacked by high-precision 3D fake hands.

[0013] To achieve the above objectives, the present invention provides a palm vein recognition method combining ultrasound and infrared technologies, comprising the following steps:

[0014] Step 1: Acquire raw images, including palm vein images and ultrasound palm print images;

[0015] Step 2: Extract the Region of Interest (ROI) and identify the areas containing palm biometric information in the original image to obtain two ROIs: the ROI of the palm vein image and the ROI of the ultrasonic palmprint image.

[0016] Step 3: Input the ROI of the palm vein image into feature encoder 1 and the ROI of the ultrasonic palm print image into feature encoder 2 to extract palm vein features and palm print features respectively.

[0017] Step 4: During the feature extraction process, bidirectional feature fusion is performed simultaneously. On one hand, the palm vein features and palm print features extracted from different layers are input into the feature fusion network for fusion to obtain fused features. On the other hand, the obtained fused features are fed back into the feature extraction process to optimize the extraction of palm vein features and palm print features, and finally obtain multi-scale features.

[0018] Step 5: Input the multi-scale features into the classifier for classification and recognition, and output the classification result;

[0019] In step 2, the ROI extraction is performed using a deep learning method.

[0020] The deep learning method includes: 1) performing convolution operations on the original image using the Canny edge detection operator to find edge pixels in the original image, and then extracting the contour information of the palm according to a threshold segmentation algorithm; 2) using a localization network to obtain the normalized coordinates of the palm and detect key landmarks of the palm, wherein the normalized coordinates are used to define the position and shape of the ROI; 3) transforming the original image using a spatial transformation network based on the key landmarks and extracting the ROI.

[0021] Furthermore, the ROI extraction in step 2 can also be performed using traditional ROI extraction methods;

[0022] The traditional ROI extraction method includes: 1) performing convolution operation on the original image using the Canny edge detection operator to find edge pixels in the original image, and then extracting the contour information of the palm according to the threshold segmentation algorithm; 2) calculating the geometric shape parameters of the palm and locating the region of the palm biometric information; 3) using the seed filling method to collect all points in the region of the palm biometric information and extracting the ROI.

[0023] Furthermore, the feature extraction strategies adopted by the feature encoder 1 and the feature encoder 2 in step 3 are dilated convolution, depthwise separable convolution, residual network + feature pyramid network, or U-Net.

[0024] Furthermore, the feature fusion network in step 4 adopts fusion methods including: element-wise addition or concatenation, adaptive pooling, or attention mechanism.

[0025] Furthermore, the classifier in step 5 includes a convolutional layer, a normalization layer, a pooling layer, and a fully connected layer. The fully connected layer is the last layer. Through multiple fully connected operations, the multi-scale features are unfolded into a vector, and then a non-linear transformation is performed through an activation function to map the high-dimensional features to the dimension of the classification label, and finally the classification result is output.

[0026] The present invention also provides a palm print and palm vein recognition device combining ultrasound and infrared, including the palm print and palm vein recognition method combining ultrasound and infrared as described above, and also including hardware and software algorithm components.

[0027] The hardware component includes an optical module and an ultrasonic module; the software algorithm component includes a palm detection and localization module, a feature extraction module, a bidirectional feature fusion module, and a backend classifier module.

[0028] The optical module is responsible for obtaining palm vein images, and the ultrasound module is responsible for obtaining ultrasound palm print images. The obtained palm vein images and ultrasound palm print images are sent to the software algorithm part for processing. After the palm detection and localization module detects the presence of a palm, the feature extraction module and the bidirectional feature fusion module perform feature recognition. Finally, the final classification result is obtained through the back-end classifier module.

[0029] Furthermore, the optical module includes several optical modules, including a camera, an infrared lamp, an indicator light, and a CMOS; the ultrasonic module is composed of an ultrasonic array, which is an array structure composed of several ultrasonic transducers.

[0030] Furthermore, the optical module and the ultrasonic module are arranged such that the optical module is located on one side of the device and the ultrasonic module is located on the other side of the device.

[0031] Furthermore, the optical module and the ultrasonic module are arranged such that the optical module is located at the bottom of the device and the ultrasonic module is located at the top of the device; or the ultrasonic array in the ultrasonic module surrounds the optical module.

[0032] The present invention also provides a palm print and palm vein recognition device combining ultrasound and infrared, including any of the aforementioned palm print and palm vein recognition devices combining ultrasound and infrared.

[0033] The present invention provides a method, device, and apparatus for palm vein and fingerprint recognition combining ultrasound and infrared technologies, which has at least the following technical advantages:

[0034] 1. The technical solution provided by this invention uses ultrasonic technology, which has the advantages of being unaffected by light interference, unaffected by surface contamination, and non-contact, thereby avoiding the influence of lighting conditions on the clarity and contrast of the image. Excessive or insufficient light may lead to recognition failure. At the same time, ultrasonic technology can detect blood flow in the palm, and use this information to determine whether the palm information is real, thus improving the security of liveness detection.

[0035] 2. The ultrasonic palmprint recognition technology used in the technical solution provided by this invention can acquire high-resolution palmprint images. Compared with traditional optical palmprint recognition technology, ultrasonic palmprint recognition has better anti-counterfeiting and adaptability, and can obtain accurate palmprint images in humid and dry environments. The palmprint image information acquired by the ultrasonic module can be combined with the surface palmprint image information acquired by the optical module to achieve multimodal palmprint recognition and improve recognition accuracy.

[0036] 3. The technical solution provided by this invention is aimed at the multimodal recognition task of palm print and palm vein. It proposes a multimodal recognition network, including a feature encoder, bidirectional feature fusion and a back-end classifier, to achieve more accurate multimodal recognition, improve the uniqueness and forgery resistance of biometric features. The back-end classifier can achieve efficient feature extraction and classification, reduce the number of model parameters and computational complexity, and improve the system's operating efficiency and response speed.

[0037] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description

[0038] Figure 1 This is a flowchart of a preferred embodiment of the present invention;

[0039] Figure 2 This is a schematic diagram of the device composition according to a preferred embodiment of the present invention;

[0040] Figure 3 This is a schematic diagram of the arrangement of the optical module and the ultrasonic module in a preferred embodiment of the present invention. Detailed Implementation

[0041] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0042] This invention utilizes an ultrasonic module and an optical module to achieve multimodal palmprint and palm vein recognition technology through multi-sensor fusion. Compared to existing technologies and equipment, this multi-sensor fusion technology, combining ultrasonic and optical modules, can more comprehensively acquire and identify palmprint and palm vein features, improving the system's ability to distinguish between real biometric features and spoofs, and effectively preventing spoofing attacks. By employing multimodal fusion, ultrasonic and optical feature information can be comprehensively utilized, improving the accuracy and precision of the palmprint and palm vein recognition system, reducing the false recognition rate, and enhancing recognition performance. The multi-sensor fusion technology increases the complexity of recognition, further improving system security and making the recognition results more reliable and secure. This device can be applied in fields such as smart locks, financial payments, and access control and attendance systems.

[0043] Example 1

[0044] like Figure 1 The diagram shown is a flowchart of a palm vein recognition method combining ultrasound and infrared technology. The method includes the following steps:

[0045] Step 1: Acquire raw images, including palm vein images and ultrasound palm print images;

[0046] Step 2: Extract the Region of Interest (ROI) and identify the areas containing palm biometric information in the original image to obtain two ROIs: the ROI of the palm vein image and the ROI of the ultrasound palmprint image. The ROI is the region of interest image.

[0047] Step 3: Input the ROI of the palm vein image into feature encoder 1 and the ROI of the ultrasonic palmprint image into feature encoder 2 to extract palm vein features and palmprint features respectively.

[0048] Step 4: During feature extraction, bidirectional feature fusion is performed simultaneously. On one hand, palm vein and palmprint features extracted from different layers are input into the feature fusion network for fusion to obtain fused features. On the other hand, the obtained fused features are fed back into the feature extraction process to optimize the extraction of palm vein and palmprint features. This feedback mechanism helps optimize the performance of feature extraction, allowing the feature extraction process to be adjusted and optimized based on the feedback information from the fused features, thereby better achieving palmprint and palm vein feature extraction. The extracted palmprint, palm vein, and fused features are then fused to obtain multi-scale features, further fully integrating palmprint and palm vein features, improving feature representation ability and model performance.

[0049] Step 5: Input the multi-scale features into the classifier for classification and recognition, and output the classification results.

[0050] Specifically, the palm vein image in step 1 is acquired using an infrared lamp and a camera. An ultrasonic palmprint image is formed by ultrasonic signals propagating along the surface of the palm and within the skin in a certain way. Due to the different structures of the skin and the underlying tissues, ultrasonic waves travel at different speeds in different tissues, causing reflection and refraction as they pass through the palm print and skin. The reflected and refracted ultrasonic signals contain information about different tissue structures, such as palm creases (ridges and valleys) and blood vessels under the skin. The received ultrasonic signals are transmitted to a signal processing unit for signal processing and analysis. By processing the amplitude, frequency, and other characteristics of the ultrasonic signals, detailed information such as the texture and even the distribution of blood vessels under the skin can be extracted from the palmprint image, forming an ultrasonic palmprint image similar to an optical palmprint image.

[0051] Specifically, before ROI extraction in step 2, the acquired raw images (ultrasound palmprint images and palm vein images) are preprocessed, mainly including image denoising, grayscale processing, and edge detection, to improve image clarity and accuracy. ROI extraction refers to identifying and extracting regions containing palm biometric information from the obtained raw images (ultrasound palmprint images and palm vein images). This information is crucial for the recognition module. The ROI extraction process helps improve the accuracy and efficiency of the recognition system while reducing resource consumption. There are generally two approaches to ROI extraction: traditional methods and deep learning-based methods, which can be flexibly chosen according to actual needs.

[0052] Specifically, feature encoder 1 and feature encoder 2 in step 3 can use different feature extraction strategies for feature extraction, such as DCNN (dilated convolution), MobileNets (depthseparable convolution), ResNet (residual network) + FPN (feature pyramid network), or U-Net. Additional skip connections can also be added to feature extraction networks of different depths to enhance the network's ability to capture image details, such as residual connections and dense connections. Depending on the actual application requirements, the feature extraction network can flexibly adjust the pooling layers, batch normalization layers, and convolutional layers to achieve better results. For example, changing the pooling or normalization strategy, using the ReLU activation function and its variants (LeakyReLU, PReLU), adjusting the convolutional kernel size, and changing the convolutional kernel shape (dilated convolution, deformable convolution).

[0053] Specifically, in step 4, the simultaneous design of bidirectional feature fusion during feature extraction effectively reduces computational and memory overhead, while also helping to better utilize the correlation between different features. The feature fusion network can employ a deep neural network structure, including convolutional layers, pooling layers, and fully connected layers, to effectively integrate and fuse these features. The feature fusion network can perform fusion in several ways: element-wise addition or concatenation, adaptive pooling, and attention mechanisms. Element-wise addition or concatenation: High-level feature maps are upsampled or convolved to the same size as low-level feature maps, and then they are added or concatenated element-wise to achieve feature fusion at different scales. Adaptive pooling: Adaptive pooling is applied to feature maps at different scales, compressing them to the same size, and then the compressed feature maps are fused. Adaptive pooling can retain important information in the feature maps, helping to improve the representational power of the fused features. Attention mechanisms: Introducing attention mechanisms helps the network dynamically learn the importance between features at different scales and fuse features based on importance, such as introducing attention modules like SENet, STN, or CBAM. By introducing an attention mechanism, features from different modalities, channels, and spatial locations of the input are adaptively fused. This allows the network to focus more on important feature regions, thereby better utilizing the correlation between palmprint and palm vein information and improving the ability to perceive targets.

[0054] Specifically, in step 5, the fused features are input into the classifier for the final recognition and classification process. This classifier contains convolutional layers, normalization layers, pooling layers, fully connected layers, etc., and can use structures such as residual connections or dense connections internally. Appropriate structures and layer types can be selected based on the specific task and resource constraints to build an efficient and accurate model. The last layer of the network is usually a fully connected layer. These layers are used to integrate and summarize the multi-scale features extracted by the network, ultimately outputting the classification result. In the fully connected layer, the network unfolds the multi-scale feature maps into a vector through multiple fully connected operations, and then performs a non-linear transformation through an activation function to map the high-dimensional features to the dimension of the classification label, ultimately outputting the classification result.

[0055] Step 2, ROI extraction, employs a deep learning method. This method includes the following steps: 1) Convolutional operation is performed on the original image using the Canny edge detection operator to find edge pixels in the original image, and then the contour information of the palm is extracted using a threshold segmentation algorithm; 2) Normalized coordinates of the palm are obtained using a localization network, and key landmarks of the palm are detected, wherein the normalized coordinates are used to define the position and shape of the ROI; 3) Based on the key landmarks, a spatial transformation network is used to transform the original image to extract the ROI.

[0056] The deep learning method uses a localization network and a spatial transformation network to output the ROI of the image. The localization network is responsible for detecting key landmarks in the hand image and transforming the image based on these landmarks to correct the elastic deformation and non-affine transformation of the hand image. The transformed ROI image is then fed into subsequent modules for feature extraction and recognition.

[0057] Example 2

[0058] Based on Example 1, the deep learning method used for ROI extraction in step 2 specifically includes:

[0059] 1) Contour Extraction: The Canny operator, an edge detection operator, is used to perform convolution operations on the image to find edge pixels. The Canny operator is a multi-stage algorithm that includes Gaussian filtering, gradient calculation, non-maximum suppression, and double thresholding to obtain the final edge image. The edge image obtained by the edge detection operator is then used to extract the palm contour using a thresholding algorithm. Additional steps are then used to remove small linear objects or blemishes caused by thresholding errors. Commonly used contour extraction algorithms include the findContours function and contour approximation methods.

[0060] 2) Localization Network: This includes a feature extraction network and a fully connected regression network, used to output normalized coordinates of the palm print. These coordinates define the location and shape of the palm print's Region of Interest (ROI). The feature extraction network's hyperparameters and backbone network can be adjusted according to specific needs, such as using ResNet50 or VGG-16. The extracted features are connected to a fully connected regression network, whose hyperparameters can also be flexibly adjusted. Typically, hidden layers are followed by a ReLU activation function and a Dropout layer to prevent overfitting.

[0061] 3) Spatial Transformation Network: This includes a mesh generator and a bilinear sampler, capable of performing effective spatial image transformations such as affine transformations, projection transformations, or thin-plate spline transformations. The mesh generator creates a deformed sampling network based on the normalized landmark coordinates output from the localization network. The bilinear sampler receives the deformed sampling network and the original image, transforming the original image into a hand ROI image with a regular mesh through sampling.

[0062] Example 3

[0063] Based on Example 1, the ROI extraction in step 2 adopts the traditional ROI extraction method.

[0064] Traditional ROI extraction methods include:

[0065] 1) Contour Extraction: The Canny operator, an edge detection operator, is used to perform convolution operations on the image to find edge pixels. The Canny operator is a multi-stage algorithm that includes Gaussian filtering, gradient calculation, non-maximum suppression, and double thresholding to obtain the final edge image. The edge image obtained by the edge detection operator is then used to extract the palm contour using a thresholding algorithm. Additional steps are then used to remove small linear objects or blemishes caused by thresholding errors. Commonly used contour extraction algorithms include the findContours function and contour approximation methods.

[0066] 2) Locating the Region of Interest (ROI): Based on the extracted contour information, the geometric parameters of the palm can be calculated, such as area, perimeter, and centroid coordinates. Using these parameters, a virtual circular region is located, which covers most of the discriminative area of ​​the palm information. Through multiple scans, the size and shape of this region are adjusted, and the ROI is refined to better fit the actual shape, thus locating the area of ​​palm biometric information.

[0067] 3) Extract the final ROI: Use the seed filling method to collect all points in the area of ​​palm biometric information and extract the ROI.

[0068] Example 4

[0069] like Figure 2 As shown, this embodiment of the invention provides a palmprint and palm vein recognition device combining ultrasound and infrared, including the palmprint and palm vein recognition method combining ultrasound and infrared as described in embodiments 1, 2, or 3 above, and further including a hardware part and a software algorithm part; wherein, the hardware part includes an optical module and an ultrasound module; the software algorithm part includes a palm detection and localization module, a feature extraction module, a bidirectional feature fusion module, and a back-end classifier module.

[0070] The optical module is responsible for obtaining palm vein images, and the ultrasound module is responsible for obtaining ultrasound palmprint images. The obtained palm vein images and ultrasound palmprint images are sent to the software algorithm part for processing. After the palm detection and localization module detects the presence of a palm, it enters the feature extraction module and the bidirectional feature fusion module for feature recognition. Finally, the final classification result is obtained through the back-end classifier module.

[0071] Specifically, the core circuit board also includes a computing and communication module, comprising computing chips, communication chips, power interfaces, transistors, resistors, capacitors, etc., responsible for performing functions such as signal processing, control circuits, data storage, and calculation and recognition. The communication between the hardware components and the calculation of the software algorithm together enable palmprint and palm vein recognition.

[0072] Example 5

[0073] Based on Example 4, the optical module refers to the module that performs imaging, acquisition or sensing through optical technology, including several optical modules, mainly composed of cameras, infrared lights, indicator lights, CMOS, etc., which work simultaneously with the ultrasound module to acquire palm vein images through infrared lights and cameras; the ultrasound module is composed of an ultrasound array, which wakes up the entire module when an object is detected to approach, emits and receives ultrasound waves, and obtains an ultrasound palmprint image through signal processing algorithms.

[0074] Specifically, the camera is used to capture image or video data; it can be a regular color camera or an infrared camera, the appropriate type chosen based on specific application requirements. The infrared light provides an infrared light source, illuminating the palm to reveal veins. Indicator lights indicate the device's operating status and camera position, such as whether it is capturing, connected, or in an abnormal state, improving user experience and facilitating operation.

[0075] Specifically, CMOS (Complementary Metal-Oxide-Semiconductor) is an integrated circuit chip technology used in image sensors for optical modules. CMOS image sensors can convert captured light signals into electrical signals, achieving the sensing and acquisition of light signals through photoelectric conversion. They integrate a series of image processing functions, such as white balance, automatic exposure, and noise reduction, and can perform real-time processing while acquiring images to improve image quality and accuracy.

[0076] Through the collaborative work of these components, the optical module enables the imaging, acquisition, and analysis of target objects or scenes, providing support for applications such as recognition, monitoring, and measurement. In palmprint and palm vein recognition technology, the optical module is mainly used to capture image data of the palm and extract palm vein information, providing data support for subsequent feature extraction and recognition.

[0077] Specifically, an ultrasonic array, or ultrasonic transducer array, is an array structure composed of several ultrasonic transducers arranged in a specific pattern. An ultrasonic transducer is a device that converts electrical energy into acoustic energy or vice versa, used to transmit and receive ultrasonic signals. Each ultrasonic transducer generates or receives ultrasonic pulses. By controlling the operating time, intensity, and phase of each transducer, specific beam shapes and directions can be formed, enabling different scanning methods and imaging modes, thereby achieving the localization and imaging of the target area. Because multiple transducers operate simultaneously, rapid imaging and real-time monitoring can be achieved. Furthermore, by controlling the density and layout of the transducers in the array, the spatial resolution of the imaging can be improved.

[0078] Specifically, the main functions and principles of the ultrasonic module are as follows: 1) Detecting an approaching object and waking up the entire module; One transducer in the ultrasonic transducer array remains operational while the rest of the module components remain dormant until an approaching object is detected, waking up the entire module, including the entire ultrasonic transducer array and optical modules. The principle is that when an object moves within a certain area, it causes changes in air density and reflection, which in turn cause corresponding changes in the ultrasonic echo. Therefore, by analyzing the ultrasonic echo, the presence and movement of an object can be determined. 2) Detecting hand-width distance: The ultrasonic ranging principle utilizes the fixed speed of ultrasonic waves in air. The distance between an object and the sensor is measured by sending and receiving ultrasonic signals. The emitted ultrasonic signal propagates in the air, and when it encounters an object's surface, a portion is reflected back. The sensor calculates the time it takes for the ultrasonic signal to travel from the sensor to the object's surface and back by measuring the time interval between the transmission and reception of the ultrasonic waves. Based on the speed of sound in air (approximately 343 meters per second), the distance between the object and the sensor can be calculated using the simple formula: distance = speed * time. 3) Liveness Detection: Ultrasonic technology can detect blood flow within the palm, using this information to determine the authenticity of the palm data, thus improving the security of liveness detection. Since blood movement causes minute changes in tissue structure, these changes are reflected in the returned ultrasound signal. Processing and analyzing the received ultrasound signal can capture these minute fluctuations caused by blood flow. By detecting blood flow information within the palm, the system can verify whether the subject is a real living being. Fake palm models typically do not have normal blood flow patterns, and ultrasonic technology can detect these differences. 4) Acquiring Palm Images: After being activated, the ultrasound module emits ultrasound signals, which propagate along the palm surface and within the skin in a certain way. Due to the different structures of the skin and the tissues beneath it, ultrasound waves travel at different speeds in different tissues, causing reflection and refraction as they pass through the palm prints and skin. Simultaneously, the ultrasound module receives the reflected and refracted ultrasound signals. The received ultrasound signals contain information about different tissue structures, such as palm lines (ridges, valleys) and blood vessels under the skin. The received ultrasound signals are transmitted to the signal processing unit for signal processing and analysis. By processing the amplitude, frequency and other characteristics of ultrasonic signals, detailed information such as texture and even blood vessel distribution under the skin can be extracted from palm print images to form ultrasonic palm print images that are similar to optical palm print images.

[0079] Example 6

[0080] Based on embodiment 4 or 5, the ultrasound module and optical module need to be installed in a location where data can be directly acquired, and must be unobstructed (the camera or ultrasound array must not be blocked).

[0081] like Figure 3 As shown, here are some possible arrangements:

[0082] The optical module is located on one side of the device, and the ultrasonic module is located on the other side.

[0083] The optical module is located at the bottom of the device, and the ultrasonic module is located at the top of the device;

[0084] The ultrasonic array in the ultrasonic module surrounds the optical module.

[0085] Example 7

[0086] Building upon embodiments 4, 5, or 6, the palm detection and localization module performs Region of Interest (ROI) extraction, identifying regions containing palm biometric information within the original image to obtain two ROIs: the ROI of the palm vein image and the ROI of the ultrasonic palmprint image. ROI extraction refers to identifying and extracting regions containing palm biometric information from the obtained original images (ultrasonic palmprint image, palm vein image). This information is crucial for the recognition module. The ROI extraction process helps improve the accuracy and efficiency of the recognition system while reducing resource consumption.

[0087] Specifically, the palm detection and localization module includes a palm state detection module and a ROI localization module. When the user places their palm, palm state detection is performed. Once the state is verified to be normal, the ROI localization module extracts the ROI images of the palm print and palm veins.

[0088] Hand position detection uses images captured by multiple sensors (infrared, RGB camera, or ultrasound) within the system to detect the user's hand position, ensuring the hand is correctly positioned on the recognition device. If the hand is not detected, is incomplete, or is misaligned during the hand detection phase, the system can provide corresponding prompts (text, voice, or image) to guide the user to re-collect the data. Prompts may include instructions to straighten the hand, adjust the hand to a suitable distance, or ensure the hand is appropriately open, thus improving the reliability and accuracy of the collected data. Once the status is normal, the high-quality biometric information collected is entered into the ROI location process.

[0089] There are generally two approaches to ROI extraction: traditional methods and deep learning-based methods. The appropriate approach can be chosen based on specific needs. Traditional ROI extraction methods include:

[0090] 1) Contour Extraction: The Canny operator, an edge detection operator, is used to perform convolution operations on the image to find edge pixels. The Canny operator is a multi-stage algorithm that includes Gaussian filtering, gradient calculation, non-maximum suppression, and double thresholding to obtain the final edge image. The edge image obtained by the edge detection operator is then used to extract the palm contour using a thresholding algorithm. Additional steps are then used to remove small linear objects or blemishes caused by thresholding errors. Commonly used contour extraction algorithms include the findContours function and contour approximation methods.

[0091] 2) Locating the Region of Interest (ROI): Based on the extracted contour information, the geometric parameters of the palm can be calculated, such as area, perimeter, and centroid coordinates. Using these parameters, a virtual circular region is located, which covers most of the discriminative area of ​​the palm information. Through multiple scans, the size and shape of this region are adjusted, and the ROI is refined to better fit the actual shape, thus locating the area of ​​palm biometric information.

[0092] 3) Extract the final ROI: Use the seed filling method to collect all points in the area of ​​palm biometric information and extract the ROI.

[0093] Deep learning methods for ROI extraction utilize localization networks and spatial transformation networks to output the image ROI. The localization network detects key landmarks in the hand image and transforms the image based on these landmarks to correct elastic deformation and non-affine transformations of the hand. The transformed ROI image is then fed into subsequent modules for feature extraction and recognition. Specifically, deep learning methods include:

[0094] 1) Contour Extraction: The Canny operator, an edge detection operator, is used to perform convolution operations on the image to find edge pixels. The Canny operator is a multi-stage algorithm that includes Gaussian filtering, gradient calculation, non-maximum suppression, and double thresholding to obtain the final edge image. The edge image obtained by the edge detection operator is then used to extract the palm contour using a thresholding algorithm. Additional steps are then used to remove small linear objects or blemishes caused by thresholding errors. Commonly used contour extraction algorithms include the findContours function and contour approximation methods.

[0095] 2) Localization Network: This includes a feature extraction network and a fully connected regression network, used to output normalized coordinates of the palm print. These coordinates define the location and shape of the palm print's Region of Interest (ROI). The feature extraction network's hyperparameters and backbone network can be adjusted according to specific needs, such as using ResNet50 or VGG-16. The extracted features are connected to a fully connected regression network, whose hyperparameters can also be flexibly adjusted. Typically, hidden layers are followed by a ReLU activation function and a Dropout layer to prevent overfitting.

[0096] 3) Spatial Transformation Network: This includes a mesh generator and a bilinear sampler, capable of performing effective spatial image transformations such as affine transformations, projection transformations, or thin-plate spline transformations. The mesh generator creates a deformed sampling network based on the normalized landmark coordinates output from the localization network. The bilinear sampler receives the deformed sampling network and the original image, transforming the original image into a hand ROI image with a regular mesh through sampling.

[0097] The feature extraction module inputs the ROI of the palm vein image into feature encoder 1 and the ROI of the ultrasonic palm print image into feature encoder 2 to extract palm vein features and palm print features respectively.

[0098] Specifically, feature encoder 1 and feature encoder 2 can use different feature extraction strategies, such as DCNN (dilated convolution), MobileNets (depthseparable convolution), ResNet (residual network) + FPN (feature pyramid network), or U-Net. Additional skip connections can also be added to feature extraction networks of different depths to enhance the network's ability to capture image details, such as residual connections and dense connections. Depending on the specific application requirements, the feature extraction network can flexibly adjust pooling layers, batch normalization layers, and convolutional layers to achieve better results. For example, changing the pooling or normalization strategy, using the ReLU activation function and its variants (LeakyReLU, PReLU), adjusting the convolutional kernel size, and changing the convolutional kernel shape (dilated convolution, deformable convolution).

[0099] The bidirectional feature fusion module performs feature fusion simultaneously during feature extraction. On one hand, it inputs palm vein and palmprint features extracted from different layers into the feature fusion network for fusion to obtain fused features. On the other hand, it feeds back the obtained fused features to the feature extraction process, optimizing the extraction of palm vein and palmprint features. This feedback mechanism helps optimize feature extraction performance, allowing the feature extraction process to be adjusted and optimized based on the feedback information from the fused features, thereby better achieving palmprint and palm vein feature extraction. By fusing the extracted palmprint, palm vein, and fused features, multi-scale features are obtained, further fully integrating palmprint and palm vein features, improving feature representation ability and model performance. This design effectively reduces computational and memory overhead, while also helping to better utilize the correlations between different features.

[0100] Specifically, feature fusion networks can employ deep neural network structures, including convolutional layers, pooling layers, and fully connected layers, to effectively integrate and fuse these features. Feature fusion networks can perform fusion in several ways: element-wise addition or concatenation, adaptive pooling, and attention mechanisms. Element-wise addition or concatenation: High-level feature maps are upsampled or convolved to the same size as low-level feature maps, and then element-wise added or concatenated to achieve feature fusion at different scales. Adaptive pooling: Adaptive pooling is applied to feature maps at different scales, compressing them to the same size, and then the compressed feature maps are fused. Adaptive pooling can retain important information in the feature maps, helping to improve the representational power of the fused features. Attention mechanisms: Introducing attention mechanisms helps the network dynamically learn the importance of features at different scales and fuse features based on importance, such as introducing attention modules like SENet, STN, or CBAM. By introducing attention mechanisms, features from different modalities, channels, and spatial locations of the input can be adaptively fused. This allows the network to focus more on important feature regions, thereby better utilizing the correlation between palm print information and palm vein information and improving the ability to perceive targets.

[0101] The fused features are then input into the classifier in the backend classifier module for the final recognition and classification process.

[0102] Specifically, this classifier includes convolutional layers, normalization layers, pooling layers, and fully connected layers, and can use structures such as residual connections or dense connections internally. Appropriate structures and layer types can be selected based on the specific task and resource constraints to build an efficient and accurate model. The last layer of the network is usually a fully connected layer. These layers are used to integrate and summarize the multi-scale features extracted by the network, ultimately outputting the classification result. In the fully connected layer, the network unfolds the multi-scale feature maps into a vector through multiple fully connected operations, and then performs a non-linear transformation through an activation function to map the high-dimensional features to the dimension of the classification label, ultimately outputting the classification result.

[0103] Example 8

[0104] This invention also provides a palm print and palm vein recognition device combining ultrasound and infrared, including the palm print and palm vein recognition equipment combining ultrasound and infrared as described above.

[0105] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. An ultrasonic, infrared combined palmprint and palm vein recognition method, characterized in that, The method comprises the following steps: Step 1, collecting original images, including palm vein images and ultrasonic palmprint images; Step 2, performing ROI extraction to identify the region of palm biometric information contained in the original image, obtaining two ROIs, which are the ROI of the palm vein image and the ROI of the ultrasonic palmprint image respectively; Step 3, inputting the ROI of the palm vein image into feature encoder 1 and the ROI of the ultrasonic palmprint image into feature encoder 2 to extract palm vein features and palmprint features respectively; Step 4, during feature extraction, simultaneously performing bidirectional feature fusion, on one hand, inputting the palm vein features and the palmprint features extracted from different layers into a feature fusion network for fusion to obtain fusion features, on the other hand, feeding back the obtained fusion features to the feature extraction process to optimize the extraction of the palm vein features and the palmprint features, and finally obtaining multi-scale features; Step 5, inputting the multi-scale features into a classifier for classification and recognition to output a classification result; In the step 2, the ROI extraction adopts a deep learning method; The deep learning method comprises: 1) performing convolution operation on the original image using an edge detection Canny operator to find edge pixels in the original image, and then extracting contour information of the palm according to a threshold segmentation algorithm; 2) using a positioning network to obtain normalized coordinates of the palm to detect key landmarks of the palm, wherein the normalized coordinates are used to define the position and shape of the ROI; 3) using a spatial transformation network to transform the original image according to the key landmarks to extract the ROI; The feature extraction strategies adopted by the feature encoder 1 and the feature encoder 2 in the step 3 are dilated convolution, depth separable convolution, residual network + feature pyramid network, or U-Net; The fusion mode adopted by the feature fusion network in the step 4 includes element-wise addition or splicing, adaptive pooling, or attention mechanism.

2. The ultrasonic, infrared combined palmprint and palm vein recognition method according to claim 1, characterized in that, In the step 2, the ROI extraction can also adopt a traditional ROI extraction method; The traditional ROI extraction method comprises: 1) performing convolution operation on the original image using an edge detection Canny operator to find edge pixels in the original image, and then extracting contour information of the palm according to a threshold segmentation algorithm; 2) calculating geometric shape parameters of the palm and positioning the region of palm biometric information; 3) using a seed filling method to collect all points in the region of palm biometric information to extract the ROI.

3. The ultrasonic, infrared combined palmprint and palm vein recognition method according to claim 1, characterized in that, The classifier in the step 5 comprises a convolution layer, a normalization layer, a pooling layer and a fully connected layer, wherein the fully connected layer is the last layer, the multi-scale features are unfolded into a vector through multi-layer fully connected operation, then a non-linear transformation is performed through an activation function to map high-dimensional features to the dimension of classification labels, and finally the classification result is output.

4. An ultrasonic, infrared combined palmprint and palmvein recognition device, characterized by, The hardware part and the software algorithm part are further included for implementing the ultrasonic and infrared combined palmprint and palm vein recognition method in any one of claims 1-3. The hardware part comprises an optical module and an ultrasonic module; the software algorithm part comprises a palm detection and positioning module, a feature extraction module, a bidirectional feature fusion module and a backend classifier module; The optical module is responsible for obtaining a palm vein image, and the ultrasonic module is responsible for obtaining an ultrasonic palmprint image; the obtained palm vein image and ultrasonic palmprint image are sent to the software algorithm part for processing; after the palm detection and positioning module detects a palm, the feature extraction module and the bidirectional feature fusion module are entered for feature recognition, and finally the backend classifier module is used to obtain a final classification result; The optical module comprises a plurality of optical modules, including a camera, an infrared lamp, an indicator lamp and a CMOS; the ultrasonic module is composed of an ultrasonic array, and the ultrasonic array is an array structure composed of a plurality of ultrasonic transducers arranged in an array; The arrangement mode of the optical module and the ultrasonic module is that the optical module is located on one side of the device, and the ultrasonic module is located on the other side of the device; The arrangement mode of the optical module and the ultrasonic module is that the optical module is located at the bottom of the device, and the ultrasonic module is located at the top of the device; or the ultrasonic array in the ultrasonic module surrounds the optical module.

5. An ultrasonic, infrared combined palmprint and palmvein recognition device, characterized by, The ultrasonic and infrared combined palmprint and palm vein recognition device of claim 4.

Citation Information

Patent Citations

  • Identity recognition verification method and module

    CN108846273A

  • Palm print image recognition method and device, equipment and storage medium

    CN115984906A

  • Non-contact multi-modal fusion biological recognition system, method and device

    CN119007252A