Palm vein recognition method based on enhanced irednet model
By using the enhanced Iresnet model, line feature enhancement module and distillation training method in palm vein recognition, the problems of different picture quality, palm posture deformation and background interference in palm vein recognition are solved, and the recognition accuracy and anti-counterfeiting ability are improved.
Patent Information
- Application Number
- CN202510164992.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-13
AI Technical Summary
The existing palm vein recognition algorithm model faces problems such as different picture quality, palm posture deformation and background interference during training, resulting in insufficient recognition accuracy and anti-counterfeiting capabilities.
The palm vein recognition method based on the enhanced iresnet model was adopted to collect palm vein pictures through a near-infrared camera, and the palm area was determined using the SCRFD detection algorithm, and the ROI area rotation and threshold screening were performed. The line feature enhancement module and distillation training method were combined to extract and identify palm vein features.
It effectively eliminates the interference factors of palm vein pictures, improves recognition accuracy and anti-counterfeiting capabilities, and solves the problem of insufficient generalization performance of small-scale models.
Smart Images

Figure CN120148079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of palm vein recognition, and particularly relates to a palm vein recognition method based on an enhanced iresnet model. Background Art
[0002] The methods adopted in traditional identity authentication mainly include the following two types: authentication methods based on identity markers and authentication methods based on identity marker knowledge. The former mostly uses physical objects such as certificates, keys, magnetic cards, etc., and the latter mostly requires the authenticator to remember, such as passwords, passwords, etc. However, both recognition methods have insurmountable disadvantages, such as the forgotten marker knowledge, the loss of markers, etc.
[0003] Compared with traditional authentication methods, biometric authentication has the advantages of being not easily forgotten, good anti-counterfeiting performance, not being easily forged or stolen, and being "carried" with the body. Among biometric authentication, palm vein recognition is a biometric recognition technology that uses the distribution information of the venous blood vessels in the human palm for personal identity authentication. The palm vein is located under the skin epidermis, has in-vivo effectiveness, and the human hand is usually in a semi-fist state, so the palm vein information is not easily stolen, having high security; at the same time, the palm vein contains rich personal information and has high identity discrimination ability, and is suitable for occasions with high security requirements such as public security, commercial finance, etc. When using the palm vein for identity authentication, the image features of the palm vein are obtained, and a live detection will be performed during the recognition process, increasing the difficulty of forgery; furthermore, because the obtained are the venous image features inside the palm, rather than the image features on the palm surface, it will not cause recognition obstacles due to damage, wear, dryness or excessive wetness on the palm surface.
[0004] However, the quality of the palm vein pictures used for training by the existing palm vein recognition algorithm models is uneven, it is difficult to judge the deformation caused by different postures of the palm and eliminate the problem of background interference in live detection, the classification scale of the palm vein area is single, and at the same time, the palm vein features extracted during model training are not obvious, resulting in insufficient generalization performance of the application of small-scale models and poor anti-counterfeiting accuracy of model recognition.
[0005] Therefore, designing a palm vein recognition method based on an enhanced iresnet model that can eliminate the interference factors of palm vein pictures and improve the recognition accuracy has become an urgent technical problem to be solved. Summary of the Invention
[0006] The present invention provides a palm vein recognition method based on an enhanced iresnet model to solve the above problems.
[0007] The technical solution of the present invention, a palm vein recognition method based on an enhanced iresnet model, includes the following steps, s1, collecting venous pictures of the human palm through a near-infrared camera; s2. Use the detection algorithm SCRFD to process the images obtained in S1, determine the approximate area of the palm, and predict the four key points between the five fingers; s3. Normalize and crop the image to obtain the palm vein ROI area. Use the key points predicted in S2 to rotate the palm with an inclined angle to the vertical direction; s4. Set the thresholds for palm inclination, clarity, brightness, inclination angle, and the proportion of the palm in the image in the image, and eliminate the palm vein images that cannot meet the thresholds simultaneously; s5. Perform palm vein detection on the image, perform palm vein detection and dilation processing, and use the multi-scale determination results of a single image for liveness judgment; s6. Establish a palm vein recognition model, train the iresnet series model by distillation, input the palm vein images determined to be live by liveness judgment, combine the line feature enhancement module in the convolutional layer of the model, extract the palm vein features in the image, and use the trained palm vein recognition model for palm vein recognition and authentication.
[0008] As a further improvement of the present invention, in step s4, by using the mobilenetV3-Large model, calculate the multi-branch output prediction values of the image as the determination values of palm inclination and clarity, and compare them with the thresholds of palm inclination and clarity; calculate the pixel values of the image through OpenCV as the determination value of brightness, and compare it with the threshold of brightness; calculate through the key point computer detection rectangle to obtain the determination values of the inclination angle and the proportion of the palm in the image, and compare them with the thresholds of the inclination angle and the proportion of the palm in the image; regard the palm vein images with the determination values of palm inclination, clarity, brightness, inclination angle, and the proportion of the palm in the image all greater than the thresholds as qualified quality images, and input them into step s5 for liveness judgment.
[0009] As a further improvement of the present invention, the liveness judgment includes the following sub-steps s5.1. Set a palm vein rectangle x with width and height of w and h respectively on the palm vein image, crop and add black borders according to the rectangle x to obtain image I1; s5.2. According to the set dilation ratio r, perform dilation processing on the palm vein rectangle x to obtain a dilated image, and then add black borders to the dilated image to obtain the corresponding image I2; s5.3. Use the liveness classification model to classify image I1 and the corresponding image I2, perform weighted average calculation on the classification results of the two images, and then compare with the preset liveness threshold Th to obtain the classification result. Output the palm vein images that reach the liveness threshold as 1, indicating liveness, and output other results as 0, indicating a prosthesis.
[0010] As a further improvement of the present invention, the step s6 includes a training stage and an inference and recognition stage. The training stage includes the following sub-steps: s6.1, using palm vein pictures for training, inputting into the iresnet50 model and the iresnet18 model. The original training loss of the models adopts Arcface Loss, and both models output 512-dimensional feature vectors. s6.2, taking the 512-dimensional feature vector f output by the iresnet50 model t as the label for training the feature vector f output by the iresnet18 model s The distillation loss adopts the mean square error loss, and the expression is where n represents the number of picture samples in a training batch. Combining with the original training loss of the model, the total training loss L is obtained, and the expression is L = L distill + L arcface , where L_distill represents the distillation loss and L_arcface represents the original training loss of the model. The inference and recognition stage includes the following sub-steps: s6.3, using the iresnet18 model trained in step s6.2 to extract the 512-dimensional features of the palm vein picture. s6.4, extracting the 512-dimensional features of the existing palm vein pictures through step s6.3, establishing a filing database and storing the extracted features; s6.5, extracting the 512-dimensional features of the palm vein picture to be recognized and authenticated through step 6.3, comparing them one by one with the features stored in the filing database to obtain a similarity score table, and using the highest similarity score in the similarity score table as S max , comparing with the set recognition success threshold S th , when S max ≥ S th , it represents that the recognition and authentication are successful. On the contrary, when S max < S th , it represents that the recognition and authentication fail.
[0011] As a further improvement of the present invention, the structure of the IResNet series model includes a 3x3 convolution and multiple IBasicBlock modules. The line feature enhancement module passes through the feature map HxWxC output by the input 3x3 convolution and the IBasicBlock module, where H and W represent the height and width of the feature map, and C represents the number of channels of the feature map; the head and tail of this line feature enhancement module include two 1x1 convolutions to reduce and increase the number of channels of the feature map. First, a 5x5 convolution is used in the middle to smooth the image, then 6 directional Gaussian filters with a size of 11x11 are used to obtain the line feature response results in 6 directions, and finally the maximum value is taken by channel to synthesize a channel result.
[0012] After adopting the above method, through steps S1 - 4, palm vein images with poor quality are screened out, improving the effect of subsequent model training; by expanding the palm vein image and adding black borders on the basis of palm vein detection, more palm vein features such as palm vein deformation can be included. Expanding by a certain proportion can also include some background information other than palm veins. In an indoor and live body situation, the background is usually all black, but a non - live body will contain information about the attack means, such as the edge of a piece of paper, etc.; in an outdoor situation, both live and non - live bodies will contain a certain amount of background information, which helps to more accurately determine whether the palm vein in the image belongs to a live body, realizing a multi - scale model for comprehensively judging a live body and coping with palm vein background interference; by enhancing the IResNet model through a line feature enhancement module, after the feature map HxWxC passes through the line feature enhancement module, the output size remains HxWxC. Using the edge - finding effect of the line feature enhancement module, the edges in the feature map can be found, highlighting the lines in the palm vein image, making the enhanced IResNet model more delicate in feature ability for palm vein recognition and more able to pay attention to the differences in line styles between palm veins. By using Arcface Loss for the original training loss of the model and adopting the distillation method to train the model, and using the distillation learning method from a large - scale model to train, the problem of insufficient generalization performance of applying a small - scale model is solved, and the anti - counterfeiting accuracy of the model recognition is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 The following shows a schematic diagram of the overall framework of palm vein recognition according to the present invention.
[0014] Figure 2 The following shows a schematic diagram of the palm vein image alignment process according to the present invention.
[0015] Figure 3 The following shows a schematic diagram of the training method of the recognition model according to the present invention.
[0016] Figure 4 The following shows a schematic diagram of the line feature enhancement module and six - direction Gaussian filters according to the present invention.
[0017] Figure 5 The following shows a schematic diagram of the structural comparison before and after adding the line feature enhancement module to the convolutional layer of the iresnet model according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] As Figures 1-5 shown, a palm vein recognition method based on an enhanced iresnet model includes the following steps. S1, collecting palm vein images of a human palm through a near - infrared camera; s2. Use the detection algorithm SCRFD to process the images obtained in S1, determine the approximate area of the palm, and predict the four key points between the five fingers; s3. Normalize and crop the image to obtain the ROI area of the palm vein. Use the key points predicted in S2 to rotate the palm with an inclined angle to the vertical direction; s4. Set the thresholds for palm inclination, clarity, brightness, inclination angle, and the proportion of the palm in the image in the image. Eliminate the palm vein images that cannot meet the thresholds simultaneously; s5. Perform palm vein detection on the image, perform palm vein detection and dilation processing, and use the multi-scale determination results of a single image to perform liveness judgment; s6. Establish a palm vein recognition model, train the iresnet series models by distillation. Input the palm vein images determined to be live by liveness judgment, combine the line feature enhancement module in the convolutional layer of the model, extract the palm vein features in the images, and use the trained palm vein recognition model to perform palm vein recognition and authentication.
[0019] In step s4, by using the mobilenetV3-Large model, calculate the multi-branch output prediction values of the image as the determination values of palm inclination and clarity, and compare them with the thresholds of palm inclination and clarity; calculate the pixel values of the image through OpenCV as the determination value of brightness, and compare it with the threshold of brightness; calculate through the key point computer detection rectangle to obtain the determination values of the inclination angle and the proportion of the palm in the image, and compare them with the thresholds of the inclination angle and the proportion of the palm in the image; regard the palm vein images with the determination values of palm inclination, clarity, brightness, inclination angle, and the proportion of the palm in the image all greater than the thresholds as qualified images in terms of quality, and input them into step s5 for liveness judgment.
[0020] The liveness judgment includes the following sub-steps, s5.1. Set a palm vein rectangle x with width and height w and h on the palm vein image, crop it according to the rectangle x and add black borders to obtain the image I1; s5.2. According to the set dilation ratio r, perform dilation processing on the palm vein rectangle x to obtain a dilated image, and then add black borders to the dilated image to obtain the corresponding image I2; s5.3. Use the liveness classification model to classify the image I1 and the corresponding image I2, perform weighted average calculation on the classification results of the two images, and then compare with the preset liveness threshold Th to obtain the classification result. Output the palm vein images that reach the liveness threshold as 1, indicating live, and output other results as 0, indicating prosthesis.
[0021] Step s6 includes a training stage and an inference and recognition stage. The training stage includes the following sub-steps, S6.1. Use the palm vein image as training, input it into the iresnet50 model and the iresnet18 model. The original training loss of the model uses Arcface Loss, and both models output 512-dimensional feature vectors. S6.2. The 512-dimensional feature vector f output by the iresnet50 model t is used as the label for training the feature vector f output by the iresnet18 model. The distillation loss uses the mean squared error loss, and the expression is s where n represents the number of image samples in a training batch. Combining with the original training loss of the model, the total training loss L is obtained, and the expression is L = L + L distill where L_distill represents the distillation loss, and L_arcface represents the original training loss of the model. arcface The inference and recognition stage includes the following sub-steps. S6.3. Use the iresnet18 model trained in step S6.2 to extract the 512-dimensional features of the palm vein image. S6.4. Extract the 512-dimensional features of the existing palm vein images through step S6.3, establish a filing database and store the extracted features. S6.5. Extract the 512-dimensional features of the palm vein image to be recognized and authenticated through step 6.3, compare them one by one with the features stored in the filing database to obtain a similarity score table, and use S max to represent the highest similarity score in the similarity score table, and compare it with the set recognition success threshold St h . When S max ≥ S th , it represents successful recognition and authentication. On the contrary, when S max < S th , it represents failed recognition and authentication.
[0022] The structure of the IResNet series model includes a 3x3 convolution and multiple IBasicBlock modules. The line feature enhancement module passes through the feature map HxWxC output by the input 3x3 convolution and the IBasicBlock module, where H and W represent the height and width of the feature map, and C represents the number of channels of the feature map. The head and tail of this line feature enhancement module contain two 1x1 convolutions to reduce and increase the number of channels of the feature map. First, use a 5x5 convolution in the middle to smooth the image, then use 6 directional Gaussian filters with a size of 11x11 to obtain the line feature response results in 6 directions, and finally take the maximum value by channel to synthesize a channel result.
[0023] Through steps S1-4, poor-quality palm vein images are screened out, improving the effect of subsequent model training. By expanding and adding black borders to palm vein images based on palm vein detection, more palm vein features such as palm vein deformation can be included. Expanding by a certain ratio can also include some background information other than palm veins. In an indoor living body situation, the background is usually all black, but a non-living body will contain information about attack means, such as the edge of a piece of paper, etc. In an outdoor situation, both living and non-living bodies will contain a certain amount of background information, which helps to more accurately determine whether the palm vein in the image belongs to a living body, realizing a multi-scale model to comprehensively judge a living body and coping with palm vein background interference. By enhancing the IResNet model through the line feature enhancement module, after the feature map HxWxC passes through the line feature enhancement module, the output size remains HxWxC. Using the edge-finding effect of the line feature enhancement module, the edges in the feature map can be found, highlighting the lines in the palm vein image, making the enhanced IResNet model more delicate in feature ability for palm vein recognition and more able to pay attention to the differences in line styles between palm veins. By using Arcface Loss for the original training loss of the model and training the model in a distillation manner, using the distillation learning method from a large-scale model, the problem of insufficient generalization performance of applying a small-scale model is solved, and the anti-counterfeiting accuracy of the model recognition is high.
Claims
1. A palm vein recognition method based on an enhanced IRESNET model, characterized in that: The following steps are included: s1, collects the vein images of human palms through near-infrared cameras; s2, use the detection algorithm SCRFD to process the image obtained in S1, determine the approximate area of the palm, and predict the four key points between the five fingers; S3, normalize and crop the image to obtain the palm vein ROI area, and use the key points predicted in S2 to rotate the tilted palm to the vertical direction; s4, set the thresholds of palm tilt, clarity, brightness, tilt angle and palm-to-image ratio in the image, and remove palm vein images that cannot meet the thresholds at the same time; s5, perform palm vein detection on the image, perform palm vein detection and expansion processing, and use the multi-scale judgment results of a single image to perform liveness judgment; S6, build a palm vein recognition model, use the distillation method to train the IRESNet series model, input the palm vein picture of the living body judged as living, combine the line feature enhancement module in the convolution layer of the model to extract the palm vein features in the picture, and use the trained palm vein recognition model to perform palm vein recognition and authentication.
2. According to claim 1, a palm vein recognition method based on an enhanced IRESNET model is characterized in that: In step s4, the multi-branch output prediction value of the image is calculated by using the mobilenetV3-Large model as the judgment value of palm tilt and clarity, and compared with the thresholds of palm tilt and clarity; the pixel value of the image is calculated by OpenCV as the judgment value of brightness, and compared with the threshold of brightness; The judgment values of the tilt angle and the proportion of the palm to the image are obtained by computer detection of key points and rectangular frame calculation, and compared with the thresholds of the tilt angle and the proportion of the palm to the image; the palm vein image whose judgment values of palm tilt, clarity, brightness, tilt angle and palm proportion to the image are all greater than the threshold is regarded as a qualified quality image and input into step s5 for liveness judgment.
3. According to claim 1, a palm vein recognition method based on an enhanced IRESNET model is characterized in that: The living body determination comprises the following sub-steps: s5.1, set a palm vein rectangular frame x with width w and height h on the palm vein image, crop according to the rectangular frame x and add a black border to obtain image I1; s5.2, according to the set expansion ratio r, the palm vein rectangular frame x is expanded to obtain an expanded image, and then a black border is added to the expanded image to obtain the corresponding image I2; s5.3, use the liveness classification model to classify image I1 and the corresponding image I2, perform weighted average calculation on the classification results of the two images, and then compare them with the preset liveness threshold Th to obtain the classification result. The palm vein image that reaches the liveness threshold is output as 1, indicating liveness, and the other results are output as 0, indicating prosthesis.
4. The palm vein recognition method based on the enhanced IRESNET model according to claim 1, characterized in that: The step s6 includes a training phase and an inference recognition phase, and the training phase includes the following sub-steps: s6.1, using palm vein images as training, input iresnet50 model and iresnet18 model, the original training loss of the model uses Arcface Loss, and both models output 512-dimensional feature vectors; s6.2, the 512-dimensional feature vector f output by the iresnet50 model t , as the output feature vector f of the training iresnet18 model s The distillation loss uses the mean square error loss, which is expressed as Where n represents the number of image samples in a training batch. Combined with the original training loss of the model, the total training loss L is obtained, which is expressed as L = L distill +L arcface, Among them, L_distill represents the distillation loss, and L_arcface represents the original training loss of the model; The reasoning and identification phase includes the following sub-steps: s6.3, use the iresnet18 model trained in step s6.2 to extract the 512-dimensional features of the palm vein image; s6.4, extracting 512-dimensional features of the existing palm vein image through step s6.3, and establishing a filing database to store the extracted features; s6.5, extract the 512-dimensional features of the palm vein image to be identified and authenticated through step 6.3, compare the cosine similarity with the features stored in the record database one by one, and obtain a similarity score table. The highest similarity score in the similarity score table is expressed as S max , and the recognition success threshold S th Compare, when S max ≥S th When S max <S th , it means that the identification authentication failed.
5. The palm vein recognition method based on the enhanced IRESNET model according to claim 1, characterized in that: The structure of the IResNet series model includes a 3x3 convolution and multiple IBasicBlock modules. The line feature enhancement module inputs the 3x3 convolution and the feature map HxWxC output by the IBasicBlock module, where H and W represent the height and width of the feature map, and C represents the number of channels of the feature map. The head and tail of the line feature enhancement module contain two 1x1 convolutions to reduce and increase the number of channels of the feature map. In the middle, a 5x5 convolution is used to smooth the image, and then 6 11x11 directional Gaussian filters are used to obtain the line feature response results in 6 directions. Finally, the maximum value of the channel is taken to synthesize a channel result.