Palm living body recognition method and device and electronic equipment
By performing key point detection and local image cropping on palm images, and analyzing local features using a target palm liveness detection model, the problem of low palm recognition accuracy in existing technologies has been solved, thereby improving the accuracy of palm recognition. In particular, by applying local features, the problem of low accuracy in palm liveness detection has been solved, achieving higher recognition accuracy.
Patent Information
- Application Number
- CN202511070853.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
AI Technical Summary
The accuracy of hand liveness detection in existing technologies is relatively low, mainly because the global frequency domain features are relatively simple and cannot make full use of the characteristic differences of different areas of the hand.
By detecting key points of the palm in the input image, key points of the palm are identified, and local image cropping is performed based on these key points. The target palm liveness detection model is then used to perform feature analysis on the local image to obtain the target palm liveness detection result.
It improves the accuracy of hand liveness detection, makes full use of the feature information of different areas of the hand, and avoids the problem of ignoring feature differences when processing global features.
Smart Images

Figure CN120977020A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of biometric identification, and particularly relates to a palm living body recognition method and device and electronic equipment. BACKGROUND
[0002] In palm living body recognition, it is usually necessary to analyze spatial texture, geometric structure or frequency domain features in a single frame image to determine the authenticity of a palm.
[0003] At present, in the related art, global frequency spectrum processing is generally performed on an input image to obtain global frequency domain features, and palm living body recognition is performed by using the difference between the global frequency domain features of a living palm and the global frequency domain features of a non-living palm.
[0004] However, palm living body recognition by using global frequency domain features has a single feature basis, and has the problem of low accuracy of palm living body recognition. SUMMARY
[0005] Embodiments of the present application provide a palm living body recognition method, device and electronic equipment to solve the problem of low accuracy of palm living body recognition in the related art.
[0006] In a first aspect, embodiments of the present application provide a palm living body recognition method, which comprises:
[0007] Performing palm key point detection on an input image to identify palm key points in the input image;
[0008] Performing cropping on the input image based on the palm key points to obtain at least one local image corresponding to the input image; each local image contains at least one palm key point;
[0009] Using a target palm living body recognition model to obtain a target palm living body recognition result corresponding to the input image based on at least one local image.
[0010] Optionally, before the step of using a target palm living body recognition model to obtain a target palm living body recognition result corresponding to the input image based on at least one local image, the method further comprises:
[0011] Obtaining at least one local image sample of an image sample and local actual information corresponding to each local image sample; the local actual information is consistent with global actual information of the image sample;
[0012] Using a palm living body recognition model to be trained to obtain predicted information corresponding to each local image sample based on at least one local image sample;
[0013] obtaining target loss values corresponding to the local image samples based on the prediction information and the actual information;
[0014] training the palm liveness recognition model to be trained based on the target loss values to obtain the target palm liveness recognition model.
[0015] Optionally, the prediction information includes prediction probabilities corresponding to the local image samples and prediction spectrums corresponding to the local image samples.
[0016] The obtaining, by the palm liveness recognition model to be trained, prediction information corresponding to at least one local image sample includes:
[0017] The obtaining, by the palm liveness recognition model to be trained, prediction information corresponding to at least one local image sample includes:
[0018] The prediction probabilities corresponding to the local image samples are obtained according to the sample time domain features, and the prediction spectrums corresponding to the local image samples are obtained according to the sample frequency domain features.
[0019] Optionally, the prediction information includes prediction probabilities corresponding to the local image samples and prediction spectrums corresponding to the local image samples; and the local actual information includes local actual labels corresponding to the local image samples and local actual spectrums corresponding to the local image samples.
[0020] The obtaining, by the palm liveness recognition model to be trained, prediction information corresponding to at least one local image sample includes:
[0021] The first loss values corresponding to the local image samples are obtained based on differences between the prediction probabilities and the local actual labels.
[0022] The second loss values corresponding to the local image samples are obtained based on the sample time domain features.
[0023] The third loss values corresponding to the local image samples are obtained based on differences between the prediction spectrums and the local actual spectrums.
[0024] The target loss values corresponding to the local image samples are obtained according to the first loss values, the second loss values and the third loss values.
[0025] Optionally, the obtaining, by the palm liveness recognition model to be trained, second loss values corresponding to the local image samples based on the sample time domain features includes:
[0026] A current local image sample in the at least one local image sample is obtained.
[0027] Based on the similarity of the time domain features of any one of the current local image samples and the time domain features of other local image samples, a second loss value corresponding to each of the local image samples is obtained.
[0028] Optionally, in the target palm living body recognition model, the target palm living body recognition result corresponding to the input image is obtained based on at least one of the local images, and the method comprises the following steps:
[0029] The palm living body recognition model is used to obtain an initial palm living body recognition result corresponding to each of the at least one local image based on the features of each of the at least one local image.
[0030] The target palm living body recognition result corresponding to the input image is obtained based on the initial palm living body recognition results corresponding to the at least one local image.
[0031] Optionally, the target palm living body recognition result corresponding to the input image is obtained based on the initial palm living body recognition results corresponding to the at least one local image, and the method comprises the following steps:
[0032] The target palm living body recognition result corresponding to the input image is obtained based on the average value of the initial palm living body recognition results corresponding to the at least one local image.
[0033] Alternatively,
[0034] The target palm living body recognition result corresponding to the input image is obtained based on the weighted average value of the initial palm living body recognition results corresponding to the at least one local image.
[0035] Alternatively,
[0036] The target palm living body recognition result corresponding to the input image is obtained based on the maximum value of the initial palm living body recognition results corresponding to the at least one local image.
[0037] Optionally, the palm key point detection on the input image comprises the following steps:
[0038] The palm key point detection model is used to obtain the palm key points in the input image based on the input image, and the palm key point detection model is obtained based on palm image samples.
[0039] In a second aspect, the embodiments of the present application provide a palm living body recognition device, and the device comprises:
[0040] The recognition module is configured to perform palm key point detection on the input image to recognize palm key points in the input image.
[0041] The image processing module is configured to crop the input image based on the palm key points to obtain at least one local image corresponding to the input image, wherein each of the local images contains at least one palm key point.
[0042] The calculation module is configured to obtain a target palm liveness recognition result corresponding to the input image based on at least one of the local images by using a target palm liveness recognition model.
[0043] Optionally, the calculation module is specifically configured to obtain at least one local image sample of an image sample and local actual information corresponding to each of the local image samples, wherein the local actual information is consistent with global actual information of the image sample; obtain predicted information corresponding to each of the local image samples based on at least one of the local image samples by using a palm liveness recognition model to be trained; obtain a target loss value corresponding to each of the local image samples based on the predicted information and the actual information; and train the palm liveness recognition model to be trained based on the target loss value to obtain the target palm liveness recognition model.
[0044] Optionally, the predicted information includes a predicted probability corresponding to each of the local image samples and a predicted frequency spectrum corresponding to each of the local image samples; and the calculation module is specifically configured to obtain a sample time domain feature and a sample frequency domain feature corresponding to each of the local image samples based on at least one of the local image samples by using the palm liveness recognition model to be trained; obtain the predicted probability corresponding to each of the local image samples according to the sample time domain feature; and obtain the predicted frequency spectrum corresponding to each of the local image samples according to the sample frequency domain feature.
[0045] Optionally, the predicted information includes a predicted probability corresponding to each of the local image samples and a predicted frequency spectrum corresponding to each of the local image samples; and the local actual information includes a local actual label corresponding to each of the local image samples and a local actual frequency spectrum corresponding to each of the local image samples; and the calculation module is specifically configured to obtain a first loss value corresponding to each of the local image samples based on a difference between the predicted probability and the local actual label; obtain a second loss value corresponding to each of the local image samples based on the sample time domain feature; obtain a third loss value corresponding to each of the local image samples based on a difference between the predicted frequency spectrum and the local actual frequency spectrum; and obtain the target loss value corresponding to each of the local image samples according to the first loss value, the second loss value, and the third loss value.
[0046] Optionally, the calculation module is specifically configured to obtain any one of the current local image samples in the at least one local image sample; and obtain a second loss value corresponding to each of the local image samples based on the similarity between the time domain feature of any one of the current local image samples and the time domain features of the other local image samples.
[0047] Optionally, the calculation module is specifically configured to obtain an initial palm liveness recognition result corresponding to each of the at least one local image based on the feature of each of the at least one local image by using a palm liveness recognition model; and obtain the target palm liveness recognition result corresponding to the input image based on the initial palm liveness recognition result corresponding to each of the at least one local image.
[0048] Optionally, the calculation module is specifically configured to obtain the target palm liveness recognition result corresponding to the input image based on the average value of the initial palm liveness recognition result corresponding to each of the at least one local image; or obtain the target palm liveness recognition result corresponding to the input image based on the weighted average value of the initial palm liveness recognition result corresponding to each of the at least one local image; or obtain the target palm liveness recognition result corresponding to the input image based on the maximum value of the initial palm liveness recognition result corresponding to each of the at least one local image.
[0049] Optionally, the recognition module is further configured to obtain the palm key point in the input image based on the input image by using a palm key point detection model; and the palm key point detection model is obtained based on palm image samples.
[0050] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory, the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the palm liveness recognition method according to any one of the first aspect.
[0051] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, the readable storage medium stores programs or instructions, and the programs or instructions are executed by a processor to implement the steps of the palm liveness recognition method according to any one of the first aspect.
[0052] In a fifth aspect, an embodiment of the present application provides a computer program product, the program product is executed by a processor of a vehicle or a cloud server to implement the steps of the palm liveness recognition method according to any one of the first aspect.
[0053] The palm living body recognition method, device and electronic equipment provided by the embodiments of the present application can recognize palm key points in an input image by detecting the palm key points in the input image, crop the input image based on the palm key points, obtain at least one local image corresponding to the input image, and each local image contains at least one palm key point, and obtain a target palm living body recognition result corresponding to the input image based on at least one local image by using a target palm living body recognition model. The palm living body recognition is realized based on the features of the local region of the palm, so that the target palm living body recognition result corresponding to the input image is obtained, and the feature difference of different regions of the palm cannot be distinguished when the living body recognition is performed based on the features of the global image of the palm image, the feature information of different regions of the palm can be fully utilized, and the accuracy of the palm living body recognition is improved. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 A flowchart of a palm living body recognition method provided by the embodiments of the present application is shown in the figure;
[0055] Figure 2 A flowchart of another palm living body recognition method provided by the embodiments of the present application is shown in the figure;
[0056] Figure 3 A structure diagram of a palm living body recognition model network to be trained provided by the embodiments of the present application is shown in the figure;
[0057] Figure 4 A flowchart of another palm living body recognition method provided by the embodiments of the present application is shown in the figure;
[0058] Figure 5 A flowchart of another palm living body recognition method provided by the embodiments of the present application is shown in the figure;
[0059] Figure 6 A flowchart of another palm living body recognition method provided by the embodiments of the present application is shown in the figure;
[0060] Figure 7 A flowchart of another palm living body recognition method provided by the embodiments of the present application is shown in the figure;
[0061] Figure 8 A flowchart of another palm living body recognition method provided by the embodiments of the present application is shown in the figure;
[0062] Figure 9 An example diagram of palm key point detection provided by the embodiments of the present application is shown in the figure;
[0063] Figure 10 A flowchart of another palm living body recognition method provided by the embodiments of the present application is shown in the figure;
[0064] Figure 11 FIG. 1 is a structural schematic diagram of a palm living body recognition device according to an embodiment of the present application. DETAILED DESCRIPTION
[0065] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.
[0066] The terms "first", "second", and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second" are generally of a kind and do not limit the number of objects.
[0067] In the related art, the input image is globally processed, the global frequency domain features of the extracted palm image are compared with the global frequency domain features of the living body sample palm image and the global frequency domain features of the non-living body sample palm image, and the palm living body recognition result is obtained by globally processing the input palm image. However, different regions of the palm (such as fingertips, palm centers, finger joints, etc.) have different structural and texture characteristics, and global processing cannot distinguish the characteristic differences of different regions of the palm, making it difficult to fully utilize the information of different regions of the palm. For example, the high-frequency texture features of the fingertips may be diluted or covered in the global frequency spectrum, resulting in the key information of the fingertips not being fully utilized, thereby resulting in low accuracy of palm living body recognition.
[0068] The palm living body recognition method provided by the present application can be applied to palm swiping products. When swiping the palm, it is necessary to first identify whether the collected input image is a living body or a non-living body. After determining that it is a living body palm, subsequent recognition operations are performed. The input image can be collected by a camera, for example, using a monocular RGB camera to collect the input image.
[0069] To improve the accuracy of palm living body recognition, the present application provides a palm living body recognition method, which can utilize the features of local images of the palm in the input image for living body recognition to obtain the target palm living body recognition result of the input image. The features of each local image are processed separately, which is beneficial to extracting and utilizing the local features of different regions of the palm, avoiding the problem that the features of the palm cannot be distinguished when the features of the palm are processed, and thus the key information of the local features is ignored and cannot be fully utilized, thereby improving the accuracy of palm living body recognition.
[0070] The embodiments of the palm living body recognition method will be described below with reference to several specific embodiments.Figure 1 A flowchart of a palm living body recognition method provided for an embodiment of the present application is shown in FIG. 1. Figure 1
[0071] S11: Perform palm key point detection on the input image to identify palm key points in the input image.
[0072] The input image collected is subjected to palm key point detection using a palm key point detection model to detect a preset number of key points of the palm, for example, 21 key points of the palm are detected using a target detection model YOLOv11, which includes finger tip key points, knuckle key points and wrist key points.
[0073] S12: Crop the input image based on the palm key points to obtain at least one local image corresponding to the input image.
[0074] Each local image contains at least one palm key point.
[0075] The input image with labeled palm key points is cropped, and the coordinates of the palm key points and the coordinates and size of the cropped region are used to determine that the cropped image contains palm key points, and the cropped image is a local image.
[0076] If the candidate region contains at least one key point, the candidate region is determined to be the local region. For example, assuming that the size of the input image is WxH, the size of the cropped region is 112x112, and the top-left corner coordinates (x, y) of the cropped region can be randomly generated, where x = random(0, W-112) and y = random(0, H-112), which means that x is a random number between 0 and W-112, and y is a random number between 0 and H-112. Assuming that the coordinates of the key points are (k x , k y ), whether the cropped region contains palm key points can be determined by the following formula:
[0077] ((k x ≥x)∧(k x <x+112))∧((k y ≥y)∧(k y <y+112))
[0078] If the coordinates of the cropped region and the coordinates of the palm key point satisfy the above formula, the cropped region contains the palm key point, which means that the cropped region can be used as a local image. For example, assuming that the size of the input image is 512x512, the top-left corner coordinates of the cropped region are (5, 16), the size of the cropped region is 112x112, and the coordinates of a palm key point are (55, 55), it is determined that the coordinates of the cropped region and the coordinates of the palm key point satisfy the above formula, and the cropped region can be used as a local image.
[0079] S13: Using the target palm liveness recognition model, a target palm liveness recognition result corresponding to the input image is obtained based on at least one local image.
[0080] The features of each local image are obtained, and an initial palm liveness recognition result corresponding to the local image is obtained according to the features of the local image. If the input image has one local image, the initial palm recognition result corresponding to the local image is used as the target palm recognition result of the input image. If the input image has two or more local images, the average of the initial palm liveness recognition results corresponding to the multiple local images is obtained as the target palm liveness recognition result of the input image, or the weighted average of the initial palm liveness recognition results corresponding to the multiple local images is obtained as the target palm liveness recognition result of the input image, or the maximum value of the initial palm liveness recognition results corresponding to the multiple local images is obtained as the target palm liveness recognition result of the input image.
[0081] In this embodiment, palm key points in the input image are identified by detecting the palm key points in the input image. At least one local image corresponding to the input image is obtained by cropping the input image based on the palm key points. Each local image contains at least one palm key point. A target palm liveness recognition result corresponding to the input image is obtained by using a target palm liveness recognition model based on at least one local image. The palm liveness recognition is realized based on the features of the local region of the palm, so that the target palm liveness recognition result corresponding to the input image is obtained. The characteristics of different regions of the palm cannot be distinguished when the global image features of the palm image are used for liveness recognition, so the features of different regions of the palm can be fully utilized, thereby improving the accuracy of palm liveness recognition.
[0082] Figure 2 Another flowchart of a palm liveness recognition method provided by the embodiment of the present application is shown in FIG. 6. Figure 2 As shown in FIG. 6, Figure 2 In the embodiment shown in FIG. 6, Figure 1 Before S13 is performed, the embodiment shown in FIG. 6 can further include:
[0083] S21: Obtain at least one local image sample of the image sample, and local actual information corresponding to each local image sample.
[0084] The local actual information is consistent with global actual information of the image sample. The local actual information includes local actual labels corresponding to each local image sample and local actual spectrums corresponding to each local image sample.
[0085] The image sample includes an image sample with an actual label of "live body" and an image sample with an actual label of "non-live body". For the image sample, a large amount of data set containing different condition hand images needs to be collected, the hand images cover various shooting scenes, light conditions and hand postures, for each collected hand image, a preset number of hand key points are marked using a marking tool (assuming 21 hand key points are marked), and a local image sample is obtained for the hand image with the marked key points. Optionally, in order to improve the consistency and stability of the local image sample, a preprocessing operation can be performed on the collected image, for example, size normalization.
[0086] For each image sample, at least one local image sample corresponding to the image sample is obtained according to the method of S12, the local actual information of the obtained local image sample is consistent with the global actual information of the image sample, which means that if the actual label of the image sample is "live body", the local actual label of the local image sample corresponding to the image sample is also "live body"; similarly, if the actual label of the image sample is "non-live body", the local actual label of the local image sample corresponding to the image sample is also "non-live body".
[0087] Each local image sample obtains a local actual spectrum corresponding to each local image sample through Fourier transform. Specifically, for a local image sample, it is first represented as a two-dimensional matrix f(x, y), where each element f(x, y) represents the pixel value of the image at coordinates (x, y). The formula for calculating the frequency domain information F(u, v) is as follows:
[0088]
[0089] Where F(u, v) represents the frequency domain information of the input image, u and v represent the frequency domain coordinates, x and y represent the pixel coordinates, W represents the width of the input image, H represents the height of the input image, f(x, y) represents the pixel value of the local image at coordinates (x, y), and j is the imaginary unit.
[0090] By visualizing F(u,v) as a two-dimensional image, a local actual spectrum F is obtained. Optionally, an amplitude spectrum can be calculated according to the obtained F(u,v). The amplitude spectrum is the amplitude part of the frequency domain information, which can be obtained by taking the modulus of a complex number. Then, a logarithmic amplitude spectrum is calculated according to the amplitude spectrum. The logarithmic transformation can compress the dynamic range in the amplitude spectrum, so that the difference between low-frequency and high-frequency components is more obvious. Secondly, the spectrum is centralized and normalized to obtain a normalized spectrum graph. Finally, the normalized spectrum graph is converted into a two-dimensional image. Optionally, a grayscale image or a pseudo-color image can be used for display. The grayscale image can directly map the normalized value to the grayscale value. The pseudo-color image can enhance the visual effect through color mapping.
[0091] S22: Using the palm living body recognition model to be trained, obtaining prediction information corresponding to each local image sample based on the at least one local image sample.
[0092] The prediction information includes a prediction probability corresponding to each local image sample and a prediction spectrum corresponding to each local image sample.
[0093] Figure 3 A structure diagram of a palm living body recognition model network to be trained is provided for the embodiments of the present application, as shown in Figure 3 The palm living body recognition model includes a main branch, a self-supervised similarity branch, and a Fourier spectrum graph supervision branch. The main branch is a binary classification branch. In the main branch, the feature map passes through a first fully connected layer, and then the feature further passes through a second fully connected layer. The second fully connected layer is used to generate a prediction probability y'. The self-supervised similarity branch is used to measure the similarity between two different local image samples. The Fourier spectrum graph supervision branch uses frequency domain information to enhance the living body recognition capability. Because the prosthesis (for example, a printed photo or a screen flip) and the real palm show different characteristics in the frequency domain, the frequency domain supervision can learn to capture these differences in the frequency domain. Through the convolution block of the Fourier spectrum graph supervision branch, the prediction spectrum corresponding to the local image sample can be obtained.
[0094] In the figure, S represents the predicted local image feature corresponding to the local image sample, S a represents the predicted local image feature corresponding to the other local image sample with the same actual label as S, FC represents a fully connected layer, y' represents the prediction probability corresponding to the local image sample, F' represents the prediction spectrum, F represents the local actual spectrum, L c represents the loss function of the main branch, L s represents the loss function of the self-supervised similarity branch, and L freq represents the loss function of the Fourier spectrum graph supervision branch. The convolution block in the figure is used to obtain the prediction spectrum corresponding to the local image sample.
[0095] First, the local image passes through an encoder, which can be a pre-trained convolutional neural network (CNN) such as ResNet, VGG or MobileNet, or a custom feature extraction network. The encoder converts the original image into a high-dimensional feature representation rich in semantic information. The encoder output is then processed by a first FC, which performs feature dimension reduction and nonlinear transformation to better suit the classification task.
[0096] S23: Based on the prediction information and the actual information, obtain a target loss value corresponding to each local image sample.
[0097] Based on the prediction probability corresponding to the local image sample, a first loss value is calculated by the loss function of the main branch; based on the local actual label corresponding to any one of the local image samples in the image sample and the local actual labels corresponding to the other local image samples in the image sample, a second loss value is calculated by the loss function of the self-supervised similarity branch; based on the prediction spectrum corresponding to each local image sample and the local actual spectrum corresponding to each local image sample, a third loss value is calculated by the loss function of the Fourier spectrum supervision branch; the weighted sum of the first loss value, the second loss value and the third loss value is determined as the target loss value.
[0098] S24: Based on the target loss value, the palm living body recognition model to be trained is trained to obtain a target palm living body recognition model.
[0099] The parameters of the palm living body recognition model to be trained are updated based on the target loss value. Optionally, the parameters of the palm living body recognition model to be trained include weights and biases.
[0100] A preset condition is set. If the target loss value meets the preset condition, it is determined that the palm living body recognition model to be trained converges, and the converged living body recognition model is determined as the target living body recognition model. For example, the target loss value less than 0.3 is set as the preset condition, and if the obtained target loss value is 0.27, which is less than 0.3, the preset condition is met, which means that the living body recognition model converges.
[0101] If it is determined that the palm living body recognition model to be trained has not converged, return to execute S21, i.e., return to execute the step of obtaining at least one local image sample of the image sample and the local actual information corresponding to each local image sample.
[0102] In this embodiment, at least one local image sample and corresponding local information of each local image sample are acquired. The local information is consistent with the global information of the image sample. Using the hand liveness detection model to be trained, prediction information corresponding to each local image sample is obtained based on at least one local image sample. Based on the prediction and actual information, a target loss value is obtained for each local image sample. Based on the target loss value, the hand liveness detection model to be trained is trained to obtain the target hand liveness detection model. This achieves training the hand liveness detection model based on local image samples and corresponding actual information, allowing the obtained target hand liveness detection model to focus on learning local features, thereby improving the accuracy of hand liveness detection.
[0103] Figure 4 A flowchart illustrating another palm liveness detection method provided in this application embodiment is shown below. Figure 4 As shown, Figure 4 Is Figure 2 Based on the illustrated embodiment, one possible implementation of S22 is as follows:
[0104] S221: Using the hand liveness recognition model to be trained, based on at least one local image sample, obtain the sample time-domain features and sample frequency-domain features corresponding to each local image sample.
[0105] The cropped local image samples are input into the hand liveness detection model to be trained. The encoder of the hand liveness detection model to be trained is used to extract features to obtain feature maps, which are the temporal features of the samples corresponding to the local image samples.
[0106] Using the above Figure 3 The convolutional block shown performs a Fourier transform on the feature map to obtain the frequency domain features of each local image sample.
[0107] S222: Based on the temporal characteristics of the samples, obtain the predicted probability corresponding to each local image sample, and based on the frequency characteristics of the samples, obtain the predicted spectrum corresponding to each local image sample.
[0108] The feature map extracted from the input local image sample is fed into one or more fully connected layers (FC). The FC calculates a score (x) based on the weights learned during training. The score x is then fed into the sigmoid function to calculate the prediction probability y′.
[0109] The formula for calculating the Sigmoid function is shown below:
[0110]
[0111] wherein x is an input real number, e is the base of the natural logarithm, and e is approximately equal to 2.71828.
[0112] The sigmoid function maps the input real number to the interval (0, 1), that is, the prediction probability y i of the i-th local image. i In a binary classification problem, if the output y i is greater than 0.5, it is classified as a positive class; if the output y i is less than 0.5, it is classified as a negative class. For example, if the output of the sigmoid function is 0.8, and the positive class is a living body and the negative class is a non-living body, it can be understood that the palm living body recognition model considers that the probability of the local image belonging to a living body is 80%, that is, the initial palm living body recognition result of the i-th local image is a living body.
[0113] The convolution block can convert the feature map of the local image sample into a sample frequency domain feature through Fourier transform, learn the frequency domain feature, and output the predicted frequency spectrum corresponding to the local image sample.
[0114] In this embodiment, based on at least one local image sample, the palm living body recognition model to be trained is used to obtain the sample time domain feature and the sample frequency domain feature corresponding to each local image sample, obtain the prediction probability corresponding to each local image sample according to the sample time domain feature, and obtain the predicted frequency spectrum corresponding to each local image sample according to the sample frequency domain feature. The characteristics of the palm image are described by using the time domain and frequency domain features, and the palm living body recognition is performed by using the time domain and frequency domain features. The palm living body recognition based on the local area features and the frequency domain features of the palm provides a basis for obtaining the target palm living body recognition result of the input image, avoids the difference in characteristics of different regions of the palm when using the global image features and the global frequency domain features of the palm image, can fully utilize the time domain and frequency domain information of different regions of the palm, and thus improves the accuracy of the palm living body recognition.
[0115] Figure 5 Another flowchart of a palm living body recognition method provided by the embodiment of the present application is shown in FIG. 8. Figure 5 Based on the embodiment shown in FIG. 7, a possible implementation of S23 is shown as follows: Figure 5 Figure 2
[0116] S231: Based on the difference between the prediction probability and the local actual label, obtain the first loss value corresponding to each local image sample.
[0117] The loss function of the main branch is a binary cross-entropy loss function, which measures the difference between the predicted probability and the local actual label. By minimizing the loss function of the main branch, the palm living body recognition model can learn more accurate feature representation and classification boundary. The output result obtained according to the loss function of the main branch is a first loss value.
[0118] The loss function of the main branch includes:
[0119] L c =-[ylog(y′)+(1-y)log(1-y′)]
[0120] Wherein, L c represents the loss function of the main branch, y = 1 represents that the local actual label is "living body", y = 0 represents that the local actual label is "non-living body", and y' represents the predicted probability of the local image sample.
[0121] S232: Based on the sample time domain feature, the second loss value corresponding to each local image sample is obtained.
[0122] Specifically, any one of the at least one local image sample is obtained. Based on the similarity of the time domain features of any one of the current local image sample and the time domain features of other local image samples, the second loss value corresponding to each local image sample is obtained.
[0123] S233: Based on the difference between the predicted spectrum and the local actual spectrum, the third loss value corresponding to each local image sample is obtained.
[0124] The Fourier spectrum graph supervision branch calculates the third loss value according to the difference between the predicted spectrum and the local actual spectrum. The loss function of the Fourier spectrum graph supervision branch supervises the output of the palm living body recognition model in the frequency domain, so that the palm living body recognition model can learn the frequency domain features of the image. The living body and non-living body images are distinguished by the spectral energy distribution, spectral center and spectral texture features. The high-frequency components (such as pores and fine lines) of the living palm are relatively rich, while the high-frequency information of the non-living palm is often lost or periodic noise. In the low-frequency part, the illumination distribution of the living palm is more natural, while the non-living palm may present uniform or conspicuous reflection characteristics. For example, common counterfeit materials such as silica gel and plastic often show high energy concentration in the high-frequency band, while the skin of the real palm shows higher low-frequency continuity and specific frequency domain patterns.
[0125] The loss function of the Fourier spectrum graph supervision branch includes:
[0126] L freq =||F′-F||
[0127] Wherein, L freqLet F' represent the loss function of the supervised branch of the Fourier spectrogram, F' represent the predicted spectrum, F represents the local actual spectrum, which is the visualized two-dimensional image corresponding to the frequency domain information of the local image, and the symbol "||||" represents the norm.
[0128] S234: Based on the first loss value, the second loss value, and the third loss value, obtain the target loss value corresponding to each local image sample.
[0129] The target loss function can be obtained from the loss functions of the three branches in the palm liveness detection model. Based on the target loss function, the target loss value is obtained. Through the target loss function, the palm liveness detection model learns not only features in the time domain but also features in the frequency domain. The palm liveness detection model can analyze local images from multiple perspectives, thereby improving the accuracy and robustness of palm liveness detection. The target loss function includes:
[0130] L=λL c +αL s +βL freq
[0131] Where L represents the target loss function, λ represents the weight factor corresponding to the loss function of the main branch, α represents the weight factor corresponding to the loss function of the self-supervised similarity branch, and β represents the weight factor corresponding to the loss function of the Fourier spectrogram supervised branch.
[0132] In this embodiment, a first loss value is obtained for each local image sample based on the difference between the predicted probability and the actual local label; a second loss value is obtained for each local image sample based on the temporal features of the sample; a third loss value is obtained for each local image sample based on the difference between the predicted spectrum and the actual local spectrum; and a target loss value is obtained for each local image sample based on the first, second, and third loss values. By combining multiple loss functions, the model can simultaneously optimize temporal features, frequency domain features, and self-supervised similarity, thereby improving the performance of the palm liveness recognition model and thus increasing the accuracy of palm liveness recognition.
[0133] Optional, in Figure 5 In the illustrated embodiment, the target loss value can be obtained based on the loss function of any one of the main branches, the self-supervised similarity branches, or the Fourier spectrum supervision branches. That is, the first loss value is determined as the target loss value, or the second loss value is determined as the target loss value, or the third loss value is determined as the target loss value.
[0134] Optionally, the target loss value can also be obtained based on the loss function of any two of the main branch, the self-supervised similarity branch or the Fourier spectrum supervision branch, that is, the weighted sum of the first loss value and the second loss value is determined as the target loss value, or the weighted sum of the first loss value and the third loss value is determined as the target loss value, or the weighted sum of the second loss value and the third loss value is determined as the target loss value.
[0135] Figure 6 A flowchart of another palm living body recognition method provided by an embodiment of the present application is shown in FIG. 6. Figure 6 Figure 6 Figure 5 Based on the embodiment shown in FIG. 6, a possible implementation of S232 is shown as follows.
[0136] S2321: Obtain any one current local image sample in the at least one local image sample.
[0137] From the at least one local image sample obtained from the image sample, any one current local image sample is determined as the input of the model, and the time domain feature is extracted; then the time domain features of other local image samples corresponding to the image sample are input, and the time domain features of the two samples are kept similar.
[0138] S2322: Obtain the second loss value corresponding to each local image sample based on the similarity between the time domain feature of any one current local image sample and the time domain feature of other local image samples.
[0139] The input of the self-supervised similarity branch is the time domain features corresponding to two different local images corresponding to the same image sample. The features between different local images of a living body palm have consistency, because the texture, light scattering characteristics and the like of real skin remain relatively stable in different local areas. Similarly, different local images of a fake palm may contain similar fake patterns (such as printed texture, pixel structure, reflection pattern, etc.), which means that different local areas of a fake palm should also exhibit consistency in the feature space.
[0140] The loss function of the self-supervised similarity branch includes:
[0141]
[0142] wherein L s represents the loss function of the self-supervised similarity branch, S1 and S2 represent the time domain features of different local image samples in a living body sample or the time domain features of two different local image samples in a non-living body sample, the symbol “.” represents the dot product operation, and the symbol “|||” represents the norm.
[0143] The output of the loss function of the self-supervised similarity branch is the second loss value.
[0144] In the embodiment, by obtaining any one current local image sample in the at least one local image sample, and obtaining the second loss value corresponding to each local image sample based on the similarity of the time domain feature of any one current local image sample and the time domain feature of other local image samples, the loss function of the self-supervised similarity branch is used to keep the similarity of different local images from the living body in the feature space, or to keep the similarity of different local images from the non-living body in the feature space, so that the model learns the similarity of the local image samples in the feature space, thereby paying more attention to the consistency of the living body features in the feature extraction process, or paying more attention to the consistency of the living body features, avoiding misjudgment, and thereby improving the accuracy of the palm living body recognition model.
[0145] Figure 7 Another flowchart of a palm living body recognition method provided by the embodiment of the application is shown in Figure 7 , as shown in Figure 7 , which is based on the embodiment shown in Figure 1 , a possible implementation of S13 is as follows:
[0146] S131: using the palm living body recognition model, obtaining the initial palm living body recognition result corresponding to each local image based on the feature of each local image in the at least one local image.
[0147] By using the palm living body recognition model, the initial palm living body recognition result corresponding to the input local image is obtained. For the input local image, classification can be performed according to the extracted features. If the similarity of the features of the local region and the features of the local region of the living body sample is higher than the preset threshold, the initial palm living body recognition result corresponding to the local region is living body; if the similarity of the features of the local region and the features of the local region of the non-living body sample is higher than the preset threshold, the initial palm living body recognition result corresponding to the local region is non-living body. For example, assuming that the preset threshold is 0.8, if the similarity of the frequency domain features of the local region and the frequency domain features of the local region of the living body sample is higher than 0.8, the initial palm living body recognition result corresponding to the local region is living body.
[0148] S132: obtaining the target palm living body recognition result corresponding to the input image based on the initial palm living body recognition result corresponding to each local image of the at least one local image.
[0149] A possible implementation is as follows:
[0150] Based on the average value of the initial palm living body recognition result corresponding to each local image of the at least one local image, the target palm living body recognition result corresponding to the input image is obtained.
[0151] Assuming that there are P local images from an input image, the target palm liveness recognition result of the input image is the average of the initial palm liveness recognition results corresponding to the P local images, and the calculation formula of the target palm liveness recognition result of the input image is as follows:
[0152]
[0153] Wherein, LiveProb is the target palm liveness recognition result of the input image, i represents the i-th local image, y i ′ represents the predicted probability corresponding to the i-th local image output in the main branch of the palm liveness recognition model.
[0154] Another possible implementation is:
[0155] Based on the weighted average of the initial palm liveness recognition results corresponding to at least one local image, the target palm liveness recognition result of the input image is obtained.
[0156] Assuming that there are P local images from an input image, the target palm liveness recognition result of the input image is the weighted average of the initial palm liveness recognition results corresponding to the P local images, and the calculation formula of the target palm liveness recognition result of the input image is as follows
[0157]
[0158] Wherein, LiveProb is the target palm liveness recognition result of the input image, a represents the weight of the initial palm liveness recognition result corresponding to the first local image, y1′ represents the initial palm liveness recognition result corresponding to the first local image, b represents the weight of the initial palm liveness recognition result corresponding to the second local image, y′2 represents the initial palm liveness recognition result corresponding to the second local image, z represents the weight of the initial palm liveness recognition result corresponding to the p-th local image, and y′ p represents the initial palm liveness recognition result corresponding to the p-th local image.
[0159] Another possible implementation is:
[0160] Based on the maximum value of the initial palm liveness recognition results corresponding to at least one local image, the target palm liveness recognition result of the input image is obtained.
[0161] Assuming that there are P local images from an input image, the target palm liveness recognition result of the input image is the maximum value of the probabilities in the initial palm liveness recognition results corresponding to the P local images.
[0162] For example, assuming that there are 9 local images, and the probability of the initial palm liveness recognition result corresponding to the 5th local image is the maximum value 0.7, the target palm liveness recognition result of the input image is the initial palm liveness recognition result corresponding to the 5th local image, that is, the target palm liveness recognition result of the input image is 0.7, which means that the target palm liveness recognition result of the input image is a living body.
[0163] In the embodiment, by using the palm liveness recognition model, the initial palm liveness recognition result corresponding to each of the at least one local image is obtained based on the feature of each of the at least one local image, and the target palm liveness recognition result corresponding to the input image is obtained based on the initial palm liveness recognition result corresponding to each of the at least one local image, so that the initial palm liveness recognition result corresponding to the local image is analyzed to obtain the target palm liveness recognition result of the input image, and even if some local images are affected by noise interference or occlusion, other local images can still provide reliable recognition information, thereby improving the accuracy and reliability of the palm liveness recognition method.
[0164] Figure 8 Another flowchart of a palm liveness recognition method provided by the embodiment of the present application is shown in Figure 8 , as shown in Figure 8 , as shown in Figure 1 , on the basis of the embodiment shown in
[0165] S111: using a palm key point detection model, obtaining palm key points in the input image based on the input image.
[0166] The palm key point detection model is trained based on palm image samples.
[0167] The target detection model can be selected as the palm key point detection model, for example, YOLOv11 is used as the palm key point detection model, the YOLOv11 model is based on a deep convolutional neural network, and the feature representation of the palm image is automatically extracted through multiple convolutional layers, pooling layers and activation functions. The palm key point detection model is trained by palm image samples with labeled palm key points, and in the training process, the model continuously adjusts the internal parameters to minimize the error between the predicted key point position and the real key point position. As shown in Figure 9 , as shown in Figure 9This is an example schematic diagram of palm key point detection provided in an embodiment of this application. The palm key points cover the key structural locations of the palm, including: wrist point (point 0): the center of the wrist; four key points of the thumb (points 1-4): the base of the thumb (point 1), the third joint of the thumb (point 2), the second joint of the thumb (point 3), and the tip of the thumb (point 4); four key points of the index finger (points 5-8): the base of the index finger (point 5), the third joint of the index finger (point 6), the second joint of the index finger (point 7), and the tip of the index finger (point 8); and four key points of the middle finger. Points 9-12: Base of the middle finger (point 9), third joint of the middle finger (point 10), second joint of the middle finger (point 11), fingertip of the middle finger (point 12); Four key points of the ring finger (points 13-16): Base of the ring finger (point 13), third joint of the ring finger (point 14), second joint of the ring finger (point 15), fingertip of the ring finger (point 16); Four key points of the little finger (points 17-20): Base of the little finger (point 17), third joint of the little finger (point 18), second joint of the little finger (point 19), fingertip of the little finger (point 20). A dataset of key points is constructed by assigning specific coordinates to each key point.
[0168] Optionally, a multi-scale training strategy can be used to train the palm keypoint detection model, which means that the palm keypoint detection model is trained at different image scales, so that the palm keypoint detection model can adapt to palm images of different sizes. Data augmentation techniques (such as random rotation, translation, flipping, etc.) can also be introduced to expand the diversity of the dataset and further improve the model's generalization ability, thereby improving the detection accuracy and robustness of the palm keypoint detection model.
[0169] In this embodiment, by utilizing a palm keypoint detection model, the keypoints of the palm in the input image are obtained, providing a foundation for obtaining local images using palm keypoints. The keypoint regions of the palm typically contain rich biometric information, such as fingertips and knuckles. By detecting and locating these regions through palm keypoint detection and extracting their local images, features for liveness detection can be obtained more effectively, improving the accuracy of the initial palm liveness detection results corresponding to the obtained local images, thereby enhancing the accuracy of palm liveness detection.
[0170] Figure 10 A flowchart illustrating another palm liveness detection method provided in this application embodiment is shown below. Figure 10 As shown, the input image is first processed by detecting key points on the palm, identifying 21 key points. Then, a local image is cropped based on these 21 key points and input into the palm liveness detection model for recognition, outputting y. i ′, according to y iOutput the target palm living body recognition result, that is, output "living body" or "non-living body".
[0171] Figure 11 A structure diagram of a palm living body recognition device provided by an embodiment of the present application includes an identification module 1101, an image processing module 1102, and a calculation module 1103. The identification module 1101 is configured to detect palm key points in an input image and identify the palm key points in the input image. The image processing module 1102 is configured to crop the input image based on the palm key points and obtain at least one local image corresponding to the input image. Each of the local images contains at least one palm key point. The calculation module 1103 is configured to obtain a target palm living body recognition result corresponding to the input image based on at least one local image by using a target palm living body recognition model.
[0172] Optionally, the calculation module 1103 is specifically configured to obtain at least one local image sample of an image sample and local actual information corresponding to each local image sample. The local actual information is consistent with global actual information of the image sample. The calculation module 1103 is specifically configured to obtain predicted information corresponding to each local image sample based on at least one local image sample by using a palm living body recognition model to be trained. The calculation module 1103 is specifically configured to obtain a target loss value corresponding to each local image sample based on the predicted information and the actual information. The calculation module 1103 is specifically configured to obtain the target palm living body recognition model by training the palm living body recognition model to be trained based on the target loss value.
[0173] Optionally, the predicted information includes predicted probability corresponding to each local image sample and predicted frequency spectrum corresponding to each local image sample. The calculation module 1103 is specifically configured to obtain sample time domain features and sample frequency domain features corresponding to each local image sample based on at least one local image sample by using a palm living body recognition model to be trained. The calculation module 1103 is specifically configured to obtain the predicted probability corresponding to each local image sample according to the sample time domain features, and obtain the predicted frequency spectrum corresponding to each local image sample according to the sample frequency domain features.
[0174] Optionally, the prediction information comprises prediction probabilities corresponding to the local image samples and prediction spectrums corresponding to the local image samples; the local actual information comprises local actual labels corresponding to the local image samples and local actual spectrums corresponding to the local image samples; the calculation module 1103 is specifically configured to obtain first loss values corresponding to the local image samples based on differences between the prediction probabilities and the local actual labels; obtain second loss values corresponding to the local image samples based on the sample time domain features; obtain third loss values corresponding to the local image samples based on differences between the prediction spectrums and the local actual spectrums; and obtain the target loss values corresponding to the local image samples according to the first loss values, the second loss values and the third loss values.
[0175] Optionally, the calculation module 1103 is specifically configured to obtain any one of the current local image samples in the at least one local image sample; and obtain second loss values corresponding to the local image samples based on similarities between a time domain feature of any one of the current local image samples and time domain features of the other local image samples.
[0176] Optionally, the calculation module 1103 is specifically configured to obtain initial palm liveness recognition results corresponding to the at least one local image respectively based on features of each of the at least one local image by using a palm liveness recognition model; and obtain the target palm liveness recognition result corresponding to the input image based on the initial palm liveness recognition results corresponding to the at least one local image respectively.
[0177] Optionally, the calculation module 1103 is specifically configured to obtain the target palm liveness recognition result corresponding to the input image based on an average value of the initial palm liveness recognition results corresponding to the at least one local image respectively; or obtain the target palm liveness recognition result corresponding to the input image based on a weighted average value of the initial palm liveness recognition results corresponding to the at least one local image respectively; or obtain the target palm liveness recognition result corresponding to the input image based on a maximum value of the initial palm liveness recognition results corresponding to the at least one local image respectively.
[0178] Optionally, the recognition module 1101 is further configured to obtain the palm key points in the input image based on the input image by using a palm key point detection model; and the palm key point detection model is obtained based on palm image samples.
[0179] The device of the embodiment can be used to execute the technical solutions of the above method embodiments, and has similar implementation principles and technical effects, which will not be described here.
[0180] The embodiment of the present application further provides an electronic device, which comprises a processor and a memory. The memory stores programs or instructions which can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the palm living body identification method shown in the foregoing embodiment are implemented. Figures 1 to 10 The embodiment of the present application further provides an electronic device, which comprises a processor and a memory. The memory stores programs or instructions which can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the palm living body identification method shown in the foregoing embodiment are implemented.
[0181] The embodiment of the present application further provides a computer readable storage medium, which stores programs or instructions, and when the programs or instructions are executed by a processor, the steps of the palm living body identification method shown in the foregoing embodiment are implemented. Figures 1 to 10 The embodiment of the present application further provides a computer readable storage medium, which stores programs or instructions, and when the programs or instructions are executed by a processor, the steps of the palm living body identification method shown in the foregoing embodiment are implemented.
[0182] The embodiment of the present application further provides a computer program product, and when the program product is executed by a processor of a vehicle or a cloud server, the steps of the palm living body identification method shown in the foregoing embodiment are implemented. Figures 1 to 10 The embodiment of the present application further provides a computer program product, and when the program product is executed by a processor of a vehicle or a cloud server, the steps of the palm living body identification method shown in the foregoing embodiment are implemented.
[0183] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of computer software products and necessary general hardware platforms, and of course, can also be realized by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disc, optical disc, etc.), and comprises a plurality of instructions for making a terminal or network side device execute the method described in each embodiment of the present application.
[0184] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative, but not restrictive. Those skilled in the art can make many forms of embodiments under the inspiration of the present application without departing from the scope of the present application and the protection scope of the claims.
Claims
1. A method for palm liveness detection, characterized in that, include: Hand key point detection is performed on the input image to identify the hand key points in the input image; The input image is cropped based on the palm key points to obtain at least one local image corresponding to the input image; each local image contains at least one of the palm key points. Using a target hand liveness detection model, based on at least one of the local images, the target hand liveness detection result corresponding to the input image is obtained.
2. The method according to claim 1, characterized in that, Before obtaining the target hand liveness recognition result corresponding to the input image based on at least one of the local images using the target hand liveness recognition model, the method further includes: At least one local image sample of an image sample is obtained, along with local actual information corresponding to each local image sample; the local actual information is consistent with the global actual information of the image sample. Using the hand liveness recognition model to be trained, prediction information corresponding to each of the local image samples is obtained based on at least one of the local image samples. Based on the predicted information and the actual information, the target loss value corresponding to each of the local image samples is obtained; Based on the target loss value, the hand liveness recognition model to be trained is trained to obtain the target hand liveness recognition model.
3. The method according to claim 2, characterized in that, The prediction information includes the prediction probability corresponding to each of the local image samples and the prediction spectrum corresponding to each of the local image samples; The step of using the hand liveness detection model to be trained, based on at least one of the local image samples, to obtain prediction information corresponding to each of the local image samples includes: Using the hand liveness recognition model to be trained, based on at least one of the local image samples, the sample time-domain features and sample frequency-domain features corresponding to each of the local image samples are obtained; Based on the temporal features of the samples, the predicted probability corresponding to each of the local image samples is obtained, and based on the frequency domain features of the samples, the predicted spectrum corresponding to each of the local image samples is obtained.
4. The method according to claim 3, characterized in that, The prediction information includes the prediction probability and the prediction spectrum corresponding to each local image sample; the local actual information includes the local actual label and the local actual spectrum corresponding to each local image sample. The step of obtaining the target loss value corresponding to each of the local image samples based on the predicted information and the actual information includes: Based on the difference between the predicted probability and the actual local label, a first loss value is obtained for each local image sample. Based on the temporal features of the samples, a second loss value corresponding to each of the local image samples is obtained; Based on the difference between the predicted spectrum and the local actual spectrum, a third loss value is obtained for each local image sample. Based on the first loss value, the second loss value, and the third loss value, the target loss value corresponding to each of the local image samples is obtained.
5. The method according to claim 4, characterized in that, The step of obtaining the second loss value corresponding to each local image sample based on the temporal features of the sample includes: Obtain any one of the current local image samples from at least one local image sample; Based on the similarity between the temporal features of any current local image sample and the temporal features of other local image samples, a second loss value is obtained for each local image sample.
6. The method according to claim 1, characterized in that, In the target hand liveness detection model, based on at least one of the local images, the target hand liveness detection result corresponding to the input image is obtained, including: Using a palm liveness detection model, based on the features of each of the at least one of the local images, initial palm liveness detection results are obtained for each of the at least one of the local images. Based on the initial hand liveness detection results corresponding to at least one of the local images, the target hand liveness detection result corresponding to the input image is obtained.
7. The method according to claim 6, characterized in that, The step of obtaining the target hand liveness detection result corresponding to the input image based on the initial hand liveness detection result corresponding to at least one of the local images includes: Based on the average value of the initial hand liveness detection results corresponding to at least one of the local images, the target hand liveness detection result corresponding to the input image is obtained; or, Based on the weighted average of the initial hand liveness recognition results corresponding to at least one of the local images, the target hand liveness recognition result corresponding to the input image is obtained; or, Based on the maximum value of the initial hand liveness detection result corresponding to at least one of the local images, the target hand liveness detection result corresponding to the input image is obtained.
8. The method according to claim 1, characterized in that, The step of detecting key hand points in the input image and identifying key hand points in the input image includes: Using a palm keypoint detection model, the palm keypoints in the input image are obtained based on the input image; the palm keypoint detection model is trained based on palm image samples.
9. A palm liveness detection device, characterized in that, The device includes: The recognition module is used to detect key points of the palm in the input image and identify key points of the palm in the input image. The image processing module is used to crop the input image based on the key points of the palm to obtain at least one local image corresponding to the input image; each local image contains at least one key point of the palm. The calculation module is used to obtain the target hand liveness recognition result corresponding to the input image based on at least one of the local images using the target hand liveness recognition model.
10. An electronic device, characterized in that, include: A processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the palm liveness detection method as described in any one of claims 1 to 8.