Hand image verification method with mixed deep learning palm fusion model
By using a hybrid deep learning palm fusion model, which combines visible light imaging equipment with multi-channel feature fusion technology, the problem of low accuracy and efficiency in palm verification in existing technologies has been solved, achieving low-cost, safe and efficient biometric identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ASIA UNIVERSITY
- Filing Date
- 2025-05-13
- Publication Date
- 2026-06-05
AI Technical Summary
Existing biometric identification technologies for palm verification methods struggle to effectively utilize visible light imaging devices in conjunction with palm prints and vein features. Furthermore, infrared sensors are expensive, raise health concerns, and have limited computing resources, resulting in low recognition efficiency.
A hybrid deep learning palm fusion model is adopted to acquire palm images through a visible light image capture device, detect the region of interest of the palm, extract multi-channel image features, and use the hybrid deep learning model to perform feature fusion and comparison to achieve the recognition of palm print and vein features.
It reduces device costs, avoids the health hazards of infrared radiation, improves recognition efficiency and accuracy, is suitable for mobile devices such as smartphones, and enhances the popularity and security of biometric identification.
Smart Images

Figure CN122157316A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a hand image verification method with a hybrid deep learning hand fusion model, and more particularly to a hand image verification method that uses visible light images and multi-channel image fusion technology for hand image verification. Background Technology
[0002] Biometric technology is currently the preferred technology for identity recognition and authentication. Among various biometric technologies, multimodal biometric systems offer superior performance compared to unimodal systems. In multimodal biometric systems, palm prints and palm veins each possess unique characteristics; the combination of these two features provides unique recognizability, making them commonly used identification targets in biometric technology. For access control operations, such as electronic device login, smartphone login, transportation system access, gatehouse access, or customs clearance, palm-based biometric identification technology is an ideal choice for these devices or systems.
[0003] While current palm recognition technology can capture palm prints using visible light imaging devices, identifying palm veins requires infrared sensors to detect their images and extract their features. Although infrared sensors are effective at capturing palm vein images, their high cost hinders widespread adoption, and prolonged exposure to infrared radiation raises health concerns. Therefore, replacing infrared sensors with visible light imaging devices has become a key research focus. Given the widespread availability of visible light imaging devices in smartphones and similar handheld devices, capturing palm images and effectively combining palm print and vein features would significantly increase the adoption rate of such devices. Furthermore, the limited computing resources of handheld devices pose a significant challenge for biometric identification. Requiring complex computational models for recognition would likely hinder efficiency and accuracy, reducing user willingness to use such devices.
[0004] In view of this, existing biometric technologies still have considerable shortcomings in the verification devices and methods for palm prints and palm veins, making it difficult to complete the identification process accurately and quickly. To address this, the inventors of this invention have conceived and designed a palm image verification method with a hybrid deep learning palm fusion model, improving upon the deficiencies of existing technologies and thereby enhancing its industrial application. Summary of the Invention
[0005] In view of the problems of the prior art, the purpose of this invention is to provide a hand image verification method with a hybrid deep learning hand fusion model, so as to solve the problem that existing hand image verification methods are difficult to improve in terms of accuracy and verification efficiency.
[0006] According to an objective of the present invention, a method for verifying hand images with a hybrid deep learning hand fusion model is proposed, comprising the following steps: acquiring a visible light image of the subject's hand using an image capture device, storing it in an image processing device, and performing Region of Interest (ROI) detection processing on the visible light image of the hand using the image processing device to extract the ROI image; performing channel extraction processing on the ROI image of the hand using the image processing device to obtain a saturation image, a red image, a green image, and a blue image; performing image enhancement processing on the saturation image, red image, green image, and blue image of the hand using the image processing device to obtain an enhanced saturation image, an enhanced red image, an enhanced green image, and an enhanced blue image; and executing a hybrid deep learning hand fusion model using the image processing device to transfer the enhanced saturation image to the blue image. Inputting saturation, red-enhanced, green-enhanced, and blue-enhanced images, saturation feature vectors, red feature vectors, green feature vectors, and blue feature vectors are extracted. The image processing device accesses and stores the hand feature vector, and combines it with the saturation, red, green, and blue feature vectors to calculate saturation feature scores, red feature scores, green feature scores, and blue feature scores. These are then fused according to channel weights to obtain a channel fusion score. The image processing device compares the channel fusion score with a default threshold. If the comparison result matches, the verification is considered successful; otherwise, the verification is considered unsuccessful.
[0007] Optionally, the hybrid deep learning palm fusion model may include a pre-trained palm feature extraction module and a hybrid feature fusion module. The pre-trained palm feature extraction module extracts features from the enhanced saturation image, enhanced red image, enhanced green image, and enhanced blue image to obtain saturation feature map, red feature map, green feature map, and blue feature map, respectively. The hybrid feature fusion module combines the saturation feature map, red feature map, green feature map, and blue feature map to extract saturation feature vector, red feature vector, green feature vector, and blue feature vector.
[0008] Optionally, the pre-trained palm feature extraction module may include a standard convolutional module, four inverse residual modules, and four separable convolutional modules.
[0009] Optionally, the hybrid feature fusion module may include a channel connection stage, a bottleneck control stage, a spatial pyramid pooling stage, and a feature vector calculation stage.
[0010] Optionally, the channel connection stage may include pairing an enhanced saturation image with an enhanced red image and connecting them by weight to generate a first paired image, pairing an enhanced green image with an enhanced blue image and connecting them by weight to generate a second paired image, and connecting the first paired image with the second paired image by weight to generate a connected image.
[0011] Optionally, the bottleneck control stage may include performing a first pointwise convolution operation on the connected images to reduce the number of channels, performing a second pointwise convolution operation after separation and expansion, and then establishing channel shortcut connections with the saturation feature map, red feature map, green feature map and blue feature map respectively to create four branch channel feature maps.
[0012] Optionally, the spatial pyramid pooling stage may include performing spatial pyramid pooling (SPPB) operations and linear pointwise convolution operations on the feature maps of the four branch channels, and then weighting and fusing the feature maps of the results to generate a fused feature map.
[0013] Optionally, the feature vector calculation stage may include a training stage and a testing stage. In the training stage, a loss function is performed on the fused feature map to calculate the saturation feature vector, red feature vector, green feature vector, and blue feature vector respectively.
[0014] Optionally, during the testing phase, the cosine distances between the saturation feature vector, red feature vector, green feature vector, and blue feature vector and the stored palm feature vector are calculated respectively to generate saturation feature scores, red feature scores, green feature scores, and blue feature scores.
[0015] Optionally, if the channel fusion score is not greater than the default threshold, the comparison result is considered to be consistent and the verification is considered successful; if the channel fusion score is greater than the default threshold, the comparison result is considered to be inconsistent and the verification is considered to be failed.
[0016] As described above, the hand image verification method with a hybrid deep learning hand fusion model according to the present invention may have one or more of the following advantages:
[0017] (1) This hand image verification method with a hybrid deep learning hand fusion model can identify the visible light image of the hand captured by a smartphone. It does not require a near-infrared sensing device, which reduces the cost of the device and increases the popularity of the verification device. It also avoids the health hazards to users caused by irradiating near-infrared light.
[0018] (2) This hand image verification method with a hybrid deep learning hand fusion model can reduce the amount of system computation and reduce hardware costs by using a lightweight hybrid deep learning hand fusion model, and is applicable to existing handheld or mobile devices, thus improving ease of use.
[0019] (3) This palm image verification method with a hybrid deep learning palm fusion model can detect a variety of business features such as palm prints and palm veins. It extracts image features through multi-channel feature extraction and obtains the final evaluation score by combining weighted connection and fusion, thereby improving the system recognition rate and enhancing the security and stability of identity verification. Attached Figure Description
[0020] To make the technical features, content, advantages, and effects of the present invention more apparent, the present invention will be described in detail below with reference to the accompanying drawings and embodiments:
[0021] Figure 1 This is a flowchart of a hand image verification method with a hybrid deep learning hand fusion model, according to an embodiment of the present invention.
[0022] Figure 2 This is a schematic diagram illustrating the acquisition of a visible light image of a palm according to an embodiment of the present invention.
[0023] Figure 3 This is a schematic diagram of the model input image in an embodiment of the present invention.
[0024] Figure 4 This is an architecture diagram of the hybrid deep learning palm fusion model according to an embodiment of the present invention.
[0025] Explanation of reference numerals in the attached figures:
[0026] 11: Smartphones
[0027] 12: Palm area
[0028] 13: Visible light image of the palm
[0029] 14: Region of Interest
[0030] 15: Image of the region of interest on the palm
[0031] 15B: Enhanced Blue Imagery
[0032] 15G: Enhancing Green Imagery
[0033] 15R: Enhanced Red Imagery
[0034] 15S: Enhanced Saturation Image
[0035] 20: Hybrid Deep Learning Hand Fusion Model
[0036] 21: Pre-trained palm feature extraction module
[0037] 22: Hybrid Feature Fusion Module
[0038] 22A: Channel Connection Phase
[0039] 22B: Bottleneck Control Phase
[0040] 22C: Spatial Pyramid Pooling Stage
[0041] 22D: Eigenvector Calculation Stage
[0042] CBR: Standard Convolutional Module
[0043] Conv: Linear pointwise convolutional layer
[0044] IR1~IR4: Inverse Residual Modules
[0045] S1~S8: Steps
[0046] SB1~SB4: Separable Convolutional Modules
[0047] Score1: Saturation Feature Score
[0048] Score2: Red Feature Score
[0049] Score3: Green Feature Score
[0050] Score4: Blue Feature Score
[0051] SPPB: Spatial Pyramid Pooling
[0052] V1: Saturation feature vector
[0053] V2: Red Feature Vector
[0054] V3: Green Feature Vector
[0055] V4: Blue Feature Vector Detailed Implementation
[0056] To facilitate understanding of the technical features, content, advantages, and effects of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and embodiments. The drawings used are for illustrative purposes only and to assist in the description. They may not represent the actual proportions and precise configurations of the present invention after implementation. Therefore, the proportions and configurations of the accompanying drawings should not be used to interpret or limit the scope of the present invention in actual implementation.
[0057] Please see Figure 1 This is a flowchart of a hand image verification method with a hybrid deep learning hand fusion model according to an embodiment of the present invention. As shown in the figure, the wrist vein verification method includes the following steps (S1~S8):
[0058] Step S1: Acquire a visible light image of the subject's palm using an image capture device, store it in an image processing device, and then perform region of interest (ROI) detection processing on the visible light image of the palm using the image processing device to extract the ROI image. When performing biometric identification, there are various options for which features to use as the identification target. Among various organ sites, the palm, which simultaneously possesses palm prints and palm veins, is an ideal option for biometric identification if its features can be effectively identified by capturing images of the palm. As previously described, acquiring images of palm veins using an infrared camera can easily cause harm to human health and consumes a large amount of equipment costs. In this embodiment, a visible light image capture device, such as a camera or camcorder on a smartphone, is used to capture a visible light image of the subject's palm, which is then used as the analysis image for palm prints and palm veins.
[0059] Please also refer to Figure 2 This is a schematic diagram illustrating the acquisition of a visible light image of the palm according to an embodiment of the present invention. As shown in the figure, the subject takes a picture of the palm portion 12 of the subject using the camera or video camera of the smartphone 11, and obtains a visible light image 13 of the palm. This visible light image 13 of the palm is a visible light image taken using a general camera or video camera. The captured image can be stored in an image processing device. In this embodiment, the visible light image 13 of the palm can be stored in the memory of the smartphone 11, and the processor of the smartphone 11 accesses the image in the memory for analysis and verification.
[0060] Since the visible light image 13 of the palm is captured by the subject or the examiner using an image capture device, such as a smartphone 11, in a non-contact environment, the distance, angle, and direction of the palm are not fixed. Therefore, the original image is processed by an image processing device to detect the region of interest 14 of the visible light image 13 of the palm, and the region of interest image 15 of the palm is extracted. The detection and processing of the region of interest 14 may, for example, crop the original image to a predetermined size, detect the outline of the palm, use the attributes between points on the outline to determine the palm location, and extract a square region of interest 14 of the palm according to the settings, outputting the palm region of interest image 15, which is also stored in the memory of the image processing device.
[0061] Step S2: The image processing device performs channel extraction processing on the palm region of interest image to obtain palm saturation image, palm red image, palm green image, and palm blue image. After obtaining the palm region of interest image 15, the image processing device can perform image conversion processing to extract different image channels. In this embodiment, a channel refers to the red, green, and blue values (R, G, B values) of the hue attribute and the saturation value of the saturation attribute in the HSL color space. Since palm prints and palm veins can be captured simultaneously under visible light, image analysis of different channels can enhance the extraction of complementary features, thereby improving the final recognition rate. The palm saturation image, palm red image, palm green image, and palm blue image obtained by the image processing device are also stored in the memory of the image processing device.
[0062] Step S3: The image processing device performs image enhancement processing on the saturation image, red image, green image, and blue image of the palm to obtain enhanced saturation image, enhanced red image, enhanced green image, and enhanced blue image. After obtaining the saturation image, red image, green image, and blue image of the palm, the image processing device can further perform image enhancement processing to enhance the images of different channels as input images for subsequent analysis of the hybrid deep learning palm fusion model.
[0063] Please see Figure 3This is a schematic diagram of the model input image in an embodiment of the present invention. As shown in the figure, the saturation image, red image, green image, and blue image of the palm can be enhanced by local ridge enhancement (LRE), which includes automatic gamma correction (AGC) on the region of interest image of the wrist, limiting contrast-limited adaptive histogram equalization (CLAHE) to increase image contrast, eliminating noise with a low-pass Gaussian filter, and highlighting edge details with a high-pass Laplacian operator. Finally, the enhanced saturation image 15S, enhanced red image 15R, enhanced green image 15G, and enhanced blue image 15B are obtained and stored in the memory of the image processing device.
[0064] Step S4: Execute the hybrid deep learning palm fusion model through the image processing device, inputting the enhanced saturation image, enhanced red image, enhanced green image, and enhanced blue image, and extracting saturation feature vectors, red feature vectors, green feature vectors, and blue feature vectors. As described in the previous step, the enhanced saturation image 15S, enhanced red image 15R, enhanced green image 15G, and enhanced blue image 15B, after image enhancement, can be input into the hybrid deep learning palm fusion model to extract image features. In order to perform calculations on mobile devices such as smartphones or tablets, this disclosure discloses a lightweight model that allows lightweight computing devices with limited computing power to successfully perform palm image analysis and recognition offline, without needing to upload data online to execute the palm image analysis and recognition model, thus improving ease of use.
[0065] Please see Figure 4 This is an architectural diagram of the hybrid deep learning palm fusion model according to an embodiment of the present invention. As shown in the figure, the hybrid deep learning palm fusion model 20 may include a pre-trained palm feature extraction module 21 and a hybrid feature fusion module 22. The enhanced saturation image 15S, enhanced red image 15R, enhanced green image 15G, and enhanced blue image 15B are respectively input as four channel images into the pre-trained palm feature extraction module 21 for feature extraction to obtain saturation feature map, red feature map, green feature map, and blue feature map. These feature maps are then further input into the hybrid feature fusion module 22 to extract saturation feature vector V1, red feature vector V2, green feature vector V3, and blue feature vector V4.
[0066] In the pre-trained palm feature extraction module 21, the image of each channel is processed through an architecture consisting of a standard convolutional block (CBR), four inverted residual blocks (IR1~IR4), and four separable convolutional blocks (SB1~SB4) interspersed between the standard convolutional block (CBR) and the inverted residual blocks (IR1~IR4). This results in the generation of saturation feature maps, red feature maps, green feature maps, and blue feature maps for the four channels. For example, the saturation-enhanced image 15S, the red-enhanced image 15R, the green-enhanced image 15G, and the blue-enhanced image 15B are 128×128×1 images. The pre-trained palm feature extraction module 21 generates 8×8×128 saturation feature maps, red feature maps, green feature maps, and blue feature maps, respectively.
[0067] In the hybrid feature fusion module 22, the saturation feature map, red feature map, green feature map and blue feature map go through the channel connection stage 22A, the bottleneck control stage 22B, the spatial pyramid pooling stage 22C and the feature vector calculation stage 22D to obtain the saturation feature vector V1, the red feature vector V2, the green feature vector V3 and the blue feature vector V4.
[0068] Channel connection stage 22A includes pairing the enhanced saturation image with the enhanced red image and connecting them by weight to generate a first paired image of 16×8×128, pairing the enhanced green image with the enhanced blue image and connecting them by weight to generate a second paired image of 16×8×128, and connecting the first paired image with the second paired image by weight to generate a connected image of 16×16×128. The calculation of feature map fusion can be expressed by formula (1):
[0069] (1)
[0070] in, F It is a fused channel feature map. n It is the number of feature maps. w i It is a feature map f i Weights, feature maps f i weight w i It is learned during the model's training process.
[0071] To adjust model complexity and enhance its generalization, a gating mechanism was employed. In the bottleneck control stage, two 1×1 pointwise convolutional layers were used to establish the bottleneck. The first pointwise convolutional layer (Conv) was applied to the connected images to reduce the number of channels, resulting in a 16×16×64 feature map. This was then separated, and a second pointwise convolutional layer (Conv) was used to expand the number of channels, resulting in an 8×8×128 feature map. This feature map was then connected via channel shortcuts with the saturation, red, green, and blue feature maps to mitigate degradation, creating four branch channel feature maps.
[0072] In the spatial pyramid pooling stage 22C, spatial pyramid pooling (SPPB) and linear pointwise convolutional layer (Conv) operations are performed on the feature maps of the four branch channels. The feature maps of the operation results are then weighted and fused to produce a fused feature map with 512 values.
[0073] Finally, the feature vector calculation stage 22D may include a training stage and a testing stage. During the training stage, a loss function is applied to the fused feature map to calculate the saturation feature vector V1, the red feature vector V2, the green feature vector V3, and the blue feature vector V4, respectively. In this embodiment, the AdaCos loss function can be used to maximize the angular margins between different identities, thus optimizing its generalization. The saturation feature vectors V1, V2, V3, and V4 of different users can be stored in the memory of the image processing device through the registration program to form a stored palm feature vector.
[0074] Step S5: Access the stored palm feature vector through the image processing device, combine the saturation feature vector, red feature vector, green feature vector, and blue feature vector to calculate the saturation feature score, red feature score, green feature score, and blue feature score, and fuse them according to the channel weights to obtain the channel fusion score. During the testing phase, after calculating the saturation feature vector V1, red feature vector V2, green feature vector V3, and blue feature vector V4, calculate the cosine distance with the stored palm feature vector to generate the saturation feature score Score1, red feature score Score2, green feature score Score3, and blue feature score Score4. The calculation of the cosine distance is shown in formula (2):
[0075] (2)
[0076] in, si,j There are two eigenvectors v i and v j The cosine distance between them m It is the length of the feature vector. Then, weighted score level fusion is used to combine the scores from all channels, as shown in formula (3):
[0077] (3)
[0078] in, S It is a fused score. n It is the number of channels. w i It is the first i The weight of the channel score, s i It is the first i Channel score.
[0079] Step S6: Compare the channel fusion score with the default threshold using the image processing device. If the comparison result matches, proceed to Step S7: Verification is successful. If the comparison result does not match, proceed to Step S8: Verification fails. The channel fusion score calculated above can be compared with the system's default threshold value. If the comparison result matches, for example, if the channel fusion score is greater than or less than the default threshold, or within a preset range of the preset threshold, the comparison result is considered to match, and verification is successful. Taking a real smartphone as an example, when verification is successful, it can serve as proof of user identity recognition or verification, thereby enabling the execution of applications requiring identity verification, such as unlocking and processing operations. Conversely, if the comparison result does not match, verification fails, indicating that the visible light image of the subject's palm captured is not the user's palm registered in the device's memory, and identity verification cannot be passed.
[0080] The hand image verification method using the hybrid deep learning hand fusion model described above allows users to capture visible light images of their hands with their smartphones, and then perform verification processing directly on the phone's processing device and memory. The lightweight model reduces the amount of hardware computation, lowers hardware requirements, and improves verification processing efficiency, making it more convenient for users to verify their hands and thus achieve the effect of identity verification.
[0081] The above description is illustrative only and not restrictive. Any equivalent modifications or alterations made without departing from the spirit and scope of this invention should be included in the appended claims.
Claims
1. A method for verifying hand images using a hybrid deep learning hand fusion model, comprising the following steps: The visible light image of the subject's palm is acquired by an image capture device, stored in an image processing device, and the visible light image of the palm is processed by the image processing device to detect the region of interest of the palm and extract the image of the region of interest of the palm. The image processing device performs channel extraction processing on the region of interest image of the palm to obtain a saturation image of the palm, a red image of the palm, a green image of the palm, and a blue image of the palm. The image processing device performs image enhancement processing on the saturation image of the palm, the red image of the palm, the green image of the palm, and the blue image of the palm to obtain enhanced saturation image, enhanced red image, enhanced green image, and enhanced blue image. The image processing device executes a hybrid deep learning palm fusion model, inputting the enhanced saturation image, the enhanced red image, the enhanced green image, and the enhanced blue image, and extracting saturation feature vectors, red feature vectors, green feature vectors, and blue feature vectors. The image processing device accesses the stored palm feature vector, combines the saturation feature vector, the red feature vector, the green feature vector, and the blue feature vector to calculate the saturation feature score, the red feature score, the green feature score, and the blue feature score, and then fuses them according to the channel weights to obtain the channel fusion score. The image processing device compares the channel fusion score with the default threshold. If the comparison result matches, the verification is considered successful; otherwise, the verification is considered unsuccessful.
2. The hand image verification method with a hybrid deep learning hand fusion model as described in claim 1, wherein the hybrid deep learning hand fusion model includes a pre-trained hand feature extraction module and a hybrid feature fusion module, wherein the pre-trained hand feature extraction module extracts features from the enhanced saturation image, the enhanced red image, the enhanced green image, and the enhanced blue image to obtain saturation feature maps, red feature maps, green feature maps, and blue feature maps, respectively, and the hybrid feature fusion module combines the saturation feature maps, the red feature maps, the green feature maps, and the blue feature maps to extract the saturation feature vector, the red feature vector, the green feature vector, and the blue feature vector.
3. The hand image verification method with a hybrid deep learning hand fusion model as described in claim 2, wherein the pre-trained hand feature extraction module includes a standard convolutional module, four inverse residual modules, and four separable convolutional modules.
4. The hand image verification method with a hybrid deep learning hand fusion model as described in claim 2, wherein the hybrid feature fusion module includes a channel connection stage, a bottleneck control stage, a spatial pyramid pooling stage, and a feature vector calculation stage.
5. The hand image verification method with a hybrid deep learning hand fusion model as described in claim 4, wherein the channel connection stage includes pairing the enhanced saturation image with the enhanced red image and connecting them by weights to generate a first paired image, pairing the enhanced green image with the enhanced blue image and connecting them by weights to generate a second paired image, and connecting the first paired image with the second paired image by weights to generate a connected image.
6. The hand image verification method with a hybrid deep learning hand fusion model as described in claim 5, wherein the bottleneck control stage includes performing a first pointwise convolution operation on the connected image to reduce the number of channels, performing a second pointwise convolution operation after separation expansion, and then establishing channel shortcut connections with the saturation feature map, the red feature map, the green feature map and the blue feature map respectively to establish four branch channel feature maps.
7. The hand image verification method with a hybrid deep learning hand fusion model as described in claim 6, wherein the spatial pyramid pooling stage includes performing spatial pyramid pooling operation and linear pointwise convolution operation on the feature maps of the four branch channels, and weighting and fusing the feature maps of the operation results to generate a fused feature map.
8. The hand image verification method with a hybrid deep learning hand fusion model as described in claim 7, wherein the feature vector calculation stage includes a training stage and a testing stage, wherein in the training stage, a loss function is performed on the fused feature map to calculate the saturation feature vector, the red feature vector, the green feature vector, and the blue feature vector, respectively.
9. The hand image verification method with a hybrid deep learning hand fusion model as described in claim 8, wherein in the testing phase, the cosine distances between the saturation feature vector, the red feature vector, the green feature vector, and the blue feature vector and the stored hand feature vector are calculated respectively to generate the saturation feature score, the red feature score, the green feature score, and the blue feature score.
10. The hand image verification method with a hybrid deep learning hand fusion model as described in claim 1, wherein if the channel fusion score is not greater than the default threshold, the comparison result is considered to be consistent and the verification is deemed successful; if the channel fusion score is greater than the default threshold, the comparison result is considered to be inconsistent and the verification is deemed to be unsuccessful.