End-side palm vein representation attack detection method and device based on key point guidance
By using a key-point-guided end-side palm vein representation attack detection method, the problems of unstable ROI localization and insufficient generalization ability of multiple attack types in the existing technology are solved. Stable ROI localization and efficient attack detection are achieved under complex acquisition conditions, which can meet the real-time inference requirements of embedded devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTH CHINA AGRICULTURAL UNIVERSITY
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-01
AI Technical Summary
Existing palm vein recognition systems are highly dependent on ROI localization and standardization, are easily affected by posture and acquisition conditions, have difficulty adapting to different hand shapes, acquisition distances and placement postures, and lack generalization ability against multiple types of attacks, making it difficult to achieve real-time inference on embedded devices.
A keypoint-guided end-palm vein representation attack detection method is adopted. The palm target bounding box and keypoints are obtained through a palm detection and keypoint localization network model. The ROI is normalized by using the geometric relationship of keypoints. Combined with a lightweight network structure and texture-sensitive enhancement mechanism, pose correction and ROI clipping are realized to adapt to different acquisition poses and attack types.
It improves the stability and anti-attack capability of the palm vein system under complex acquisition conditions, reduces feature drift caused by posture changes, enhances robustness against multiple types of attacks and real-time performance of edge deployment, and reduces false alarm rate and computational load.
Smart Images

Figure CN121963265A_ABST
Abstract
Description
Keypoint-guided method and device for detecting attacks on end-palmar vein representation Technical Field
[0001] This invention belongs to the field of biometric recognition and information security technology, specifically relating to a method and device for detecting end-palm vein representation attacks based on key point guidance. Background Technology
[0002] Palm veins, as a type of internal vascular texture feature, possess a certain degree of concealment and stability, attracting widespread attention in scenarios such as access control and attendance, financial payments, and identity verification. Palm veins are typically identified using near-infrared (NIR) imaging. This involves using the differences in penetration and absorption of subcutaneous vascular structures by NIR light to create texture contrast, thereby obtaining palm vein images that can be used for comparison and identification. Compared to externally visible features such as fingerprints and faces, palm veins have the advantage of being difficult to directly observe and replicate.
[0003] However, with the development of technologies such as printing and copying, 3D molding, material simulation, and generative models, palm vein recognition systems may still face various representation attack risks. For example, attackers can use physical carriers such as paper-printed palm vein images, bionic molds, and gloves to create forgeries, or use generative methods to synthesize forged images with realistic textures, in order to bypass the authenticity verification of the recognition system. To address this, the industry typically introduces representation attack detection (PAD) as a pre-processing step in the palm vein recognition workflow to determine whether the input sample comes from a real, live hand or from a forged representation, thereby improving system security and resistance to attacks.
[0004] Existing palm vein PAD solutions generally suffer from the following problems:
[0005] (1) ROI localization and normalization are highly dependent and easily affected by posture and acquisition conditions. Due to factors such as acquisition distance, hand rotation, placement angle, occlusion and background interference, the position, scale and orientation of the effective texture area in palm vein images vary greatly; if ROI extraction is unstable, it is easy for subsequent feature extraction to focus on non-critical areas, resulting in missed detection or false detection. Some methods rely on fixed template cropping or manually set geometric rules to achieve ROI extraction, which is difficult to adapt to different hand shapes, acquisition distances and placement postures, and has significantly insufficient robustness;
[0006] (2) Schemes that rely solely on bounding boxes or pixel-level segmentation are difficult to balance accuracy and efficiency in edge deployment. Using only bounding boxes for cropping often fails to ensure consistent ROI orientation and can easily introduce irrelevant areas such as wrists and backgrounds; while pixel-level segmentation / keypoint models, if designed in a complex manner, may result in a large number of parameters and computational load, which is not conducive to real-time inference on embedded devices or mobile devices.
[0007] (3) Insufficient generalization ability for multiple types of attacks. Different attack vectors (such as printing, molds, gloves, generated images, etc.) differ significantly in terms of texture details, edge artifacts, and reflection characteristics; if the PAD model only learns one type of feature of the overall appearance or local texture, it may misjudge when faced with "unseen attack types";
[0008] (4) Edge deployment requires a unified input specification and a stable engineering link. In practical applications, PAD modules are usually used as online detection modules on terminal devices. They need to complete high frame rate inference under the conditions of limited computing power and limited storage space. If there is a lack of adaptation design and model compression strategy with edge inference frameworks (such as ONNX / NCNN), it is difficult to meet the requirements of real-time performance and deployment stability.
[0009] Therefore, there is an urgent need for a palm vein representation attack detection method and device that can achieve stable ROI localization and standardization in near-infrared palm vein acquisition scenarios, while taking into account both end-side deployment efficiency and robustness against multiple types of representation attacks, in order to overcome the shortcomings of existing technologies in terms of pose changes, scale changes, complex attacks, and terminal deployment. Summary of the Invention
[0010] In order to overcome the shortcomings of the existing technology, the present invention provides a method for detecting end-side palm vein representation attacks based on key point guidance. The method for detecting end-side palm vein representation attacks can improve the stability and anti-attack capability of the palm vein system under complex acquisition conditions.
[0011] The second objective of this invention is to provide a device for detecting attacks on the palmar vein representation based on key point guidance.
[0012] The technical solution of the present invention to solve the above-mentioned technical problems is:
[0013] A key-point-guided method for detecting end-palmar vein representation attacks includes the following steps:
[0014] S1. Acquire single-frame images or continuous video frames of the palm using a near-infrared imaging device, and preprocess the acquired palm images.
[0015] S2. Input the preprocessed palm image into the palm detection and key point localization network model, output the palm target box and the coordinates of four key points K1 to K4 located at the finger gap valley of the palm, and obtain the confidence score corresponding to the palm target box and the confidence score corresponding to each key point.
[0016] S3. Determine the validity of the input sample based on the confidence scores of the target bounding box and the key points. If the validity determination result is an invalid sample, output an invalid flag and end the process. If the validity determination result is a valid sample, proceed to step S4.
[0017] S4. Calculate the palm posture correction parameters based on the coordinates of the key points, and perform rotation correction on the palm image based on the palm posture correction parameters;
[0018] S5. On the rotated and corrected palm image, first determine whether the key points used to construct the ROI region meet the preset legality constraints: if the key points meet the preset legality constraints, then determine the ROI region based on the geometric relationship of the key points and crop it to obtain the ROI image. Then scale the ROI image to the preset input specifications to obtain the normalized ROI image; if the key points do not meet the preset legality constraints, then use the ROI backtracking strategy to generate a normalized ROI image that meets the preset input specifications; subsequently, input the generated normalized ROI image into the palm vein representation attack detection network model to obtain the attack probability of the corresponding input sample.
[0019] S6. Compare the attack probability with a preset threshold, output the judgment result based on the comparison result, and output the attack probability and the preset threshold at the same time.
[0020] Preferably, in step S1, the near-infrared imaging device includes a near-infrared camera module and a near-infrared supplementary light source; the preprocessing includes one or more combinations of grayscale conversion, normalization, and scaling.
[0021] Preferably, in step S2, the four key points K1 to K4 are arranged in the following order from thumb to little finger: thumb-index finger valley point K1, index finger-middle finger valley point K2, middle finger-ring finger valley point K3, and ring finger-little finger valley point K4.
[0022] Preferably, in step S3, when the confidence level of the hand target box is less than the first preset threshold, it is determined to be an invalid sample, a non-hand / invalid sample flag is output and the detection is terminated; when the confidence level of the key point used for ROI construction is less than the second preset threshold or the geometric relationship of the key point does not meet the legality constraint, the ROI rollback strategy in step S5 is triggered; wherein, the legality constraint includes one or more combinations of key point order abnormality, key point spacing abnormality, and ROI out-of-bounds risk.
[0023] Preferably, in step S4, the attitude correction parameters are determined by key point A and key point B, wherein key point A is the index-middle finger valley point K2, and key point B is the ring-little finger valley point K4; based on the obtained vector Angle with the horizontal direction Rotate the palm image to make the line connecting key point A and key point B parallel to the horizontal direction.
[0024] Preferably, in step S5, the step of determining the ROI region is as follows:
[0025] A local coordinate system is established with the line connecting key point A and key point B as the X-axis and the perpendicular bisector of that line as the Y-axis. The origin O of the local coordinate system is the midpoint of the line connecting key point A and key point B. A distance of [missing information] from the origin O is selected on the Y-axis. Point C; with point C as the center and side length as... The square region is used as the ROI region and then cropped; where, and This is a preset proportional coefficient. Take a value of 0.5 to 1.2; Take a value of 0.8 to 1.8.
[0026] Preferably, in step S5, the ROI rollback strategy includes any one or more combinations of the following strategies 1 to 4:
[0027] Strategy 1: Using the center of the palm target bounding box as the center of the ROI region, generate the ROI region according to the width and height of the palm target bounding box and a preset scaling factor;
[0028] Strategy 2: Replace the line connecting the index finger-middle finger valley point K2 and the middle finger-ring finger valley point K3, or the line connecting the thumb-index finger valley point K1 and the ring finger-little finger valley point K4, with the line connecting the index finger-middle finger valley point K2 and the ring finger-little finger valley point K3, and calculate the palm posture correction parameters and scale parameters accordingly.
[0029] Strategy 3: When the input is a series of video frames, reuse the affine transformation parameters corresponding to the effective ROI region of the previous frame, and perform smooth update processing on the affine transformation parameters.
[0030] Strategy 4: Skip the rotation correction step of the palm image and directly perform ROI region cropping operation based on the palm target bounding box in the original image coordinate system.
[0031] Preferably, in step S5, the end-palm vein representation attack detection network model includes a feature extraction backbone network and a classification head, wherein the feature extraction backbone network includes depthwise separable convolutional units, Ghost feature generation units, and channel attention units.
[0032] Preferably, in step S6, if the attack probability is greater than a preset threshold, it is determined to be an attack sample; if the attack probability is less than or equal to the preset threshold, it is determined to be a real sample.
[0033] A key-point guided end-palm vein representation attack detection device includes:
[0034] The image acquisition and preprocessing module is used to acquire single-frame images or continuous video frames of the palm captured by the near-infrared imaging device, and to preprocess the acquired palm images.
[0035] The palm detection and key point localization module is used to input the preprocessed palm image into the palm detection and key point localization network model, and output the palm target box, the coordinates of four key points located at the palm finger gap valley, the confidence score of the palm target box, and the confidence score of each key point.
[0036] The sample validity determination module is used to determine the validity of the input sample based on the confidence level of the palm target box and the confidence level of each key point; if the sample is determined to be invalid, an invalid flag is output and the detection is terminated; if the sample is determined to be valid, the palm posture correction module is triggered to work.
[0037] The hand posture correction module is used to calculate hand posture correction parameters based on the coordinates of the key points, and to perform rotation correction on the hand image based on the hand posture correction parameters.
[0038] The ROI processing and attack probability acquisition module is used to first determine whether the key points used to construct the ROI region on the rotated and corrected palm image meet the preset legality constraints. If they do, the ROI region is determined and cropped based on the geometric relationship of the key points to obtain the ROI image. Then, the ROI image is scaled to the preset input specifications to obtain the ROI normalized image. If it does not meet the requirements, an ROI backtracking strategy is used to generate an ROI normalized image that meets the preset input specifications. Subsequently, the ROI normalized image is input to the palm vein representation attack detection network model to obtain the attack probability of the corresponding input sample.
[0039] The attack detection and judgment module is used to compare the attack probability with a preset threshold, output the final attack detection and judgment result based on the comparison result, and simultaneously output the attack probability and the preset threshold.
[0040] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0041] 1. This invention obtains the target bounding box and key points K1 to K4 of the palm through a palm detection and key point localization network model, and normalizes the ROI by using the geometric relationship of the key points as a constraint, so that the direction and scale of the ROI remain consistent or approximately consistent under different acquisition postures, different distances and different placement angles, reducing feature drift caused by posture changes and improving the stability of subsequent palm vein representation attack detection.
[0042] 2. This invention employs a pose correction and ROI construction method based on key point connection AB, utilizes affine transformation to achieve rotation correction, and introduces a configurable scaling factor using |AB| as the scale reference. and By determining the center and side length of the ROI, multi-scale ROI region extraction can be achieved based on different distance rules after determining the location of key points, adapting to different hand shapes and different acquisition distances, thereby improving the versatility and transferability of the end-palm vein representation attack detection method of this invention.
[0043] 3. Before edge-side inference, the present invention uniformly scales the input to a preset size (e.g., 128×128) and prioritizes the completion of key point-guided ROI clipping and normalization before scaling. This can effectively reduce the interference of background, wrist and non-target areas on edge-side palm vein representation attack determination. Under the constraint of fixed input size, it can still maximize the preservation of effective palm vein texture information, thereby avoiding texture dilution and information loss caused by direct full-image scaling.
[0044] 4. This invention sets up a sample validity determination and ROI rollback strategy to improve fault tolerance in complex scenarios and reduce false alarms caused by false triggers and invalid samples.
[0045] 5. The end-side palm vein representation attack detection network model in this invention adopts a lightweight network structure design, integrating modules such as depthwise separable convolution, Ghost feature generation, and channel attention. While ensuring detection performance, it significantly reduces the number of parameters and computational load, adapts to the resource-constrained characteristics of end-side devices, meets the requirements of real-time inference, and improves deployment feasibility and operational stability.
[0046] 6. This invention introduces a texture sensitivity enhancement mechanism in the feature extraction stage of the end-side palm vein representation attack detection network model. For example, it uses central difference convolution (CDC) to enhance the response to edge artifacts and fine-grained texture differences. In the training stage, it introduces texture energy proxy supervision based on Sobel and Laplacian to improve the model's sensitivity to texture differences in various types of attacks such as printing, molds, gloves, and generated classes, and enhance the robustness of distinguishing complex forgeries and unseen attack types.
[0047] 7. This invention uses a probability score and threshold-based decision method to output the attack probability, compares it with a preset threshold, and outputs the comparison result. This method facilitates threshold configuration and dynamic adjustment of participation strategies in different business scenarios, improving the configurability and maintainability of engineering implementation.
[0048] 8. This invention targets video / continuous frame detection scenarios, and performs temporal smoothing and fusion processing on ROI affine parameters and attack probabilities to reduce the impact of key point jitter and single-frame noise on detection results, improve the output stability of continuous frames, and optimize user experience.
[0049] 9. The network model in this invention supports exporting to ONNX format and further converting it to the format for edge inference. With the FP16 and INT8 quantization strategies, lightweight edge deployment and efficient inference can be achieved. Attached Figure Description
[0050] Figure 1 is an overall flowchart of the end-palm vein representation attack detection method based on key point guidance of the present invention.
[0051] Figure 2 is a schematic diagram of the structure of the terminal palmar vein representation attack detection network model.
[0052] Figure 3 is a schematic diagram of ROI extraction and normalization based on keypoint guidance, which shows the selection of keypoint A and keypoint B, and the rotation correction angle. The calculation, establishment of the coordinate system, determination of point C, and ROI interception and scaling centered at point C.
[0053] Figure 4 is a schematic diagram of the near-infrared imaging module.
[0054] Figure 5 is a schematic diagram of the data acquisition and system operation process.
[0055] Figure 6 shows examples of palm vein images in different scenarios and under different representation attack conditions in the self-built dataset, where (a) to (h) are examples of near-infrared palm vein images under different sample conditions.
[0056] Figure 7 is a schematic diagram of the real-time inference output from the edge, showing the target bounding box of the palm, the localization results of key points, and the judgment results and probability information. Detailed Implementation
[0057] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0058] Example 1
[0059] Referring to Figures 1-7, the key-point-guided end-side palm vein representation attack detection method of the present invention is applicable to real-time anti-counterfeiting scenarios on embedded end-side platforms (e.g., RV1106), and specifically includes the following steps:
[0060] S1. Acquire palm images using a near-infrared imaging device and preprocess the acquired palm images. The acquisition method can be single-frame acquisition or continuous video stream acquisition. When it is a video stream, it can be input into the subsequent process frame by frame according to the frame rate, and a continuous frame smoothing strategy can be enabled to improve stability. The near-infrared imaging device includes a near-infrared imaging module (as shown in Figures 4 and 5), a near-infrared supplementary light source (which can be integrated into the near-infrared imaging module or external), an embedded processing platform (e.g., RV1106), and a display / communication interface, etc. The center wavelength of the near-infrared supplementary light source can be selected as 850nm or 940nm, and the imaging method can be transmission or reflection imaging. This embodiment uses reflection near-infrared palm imaging.
[0061] The preprocessing includes one or more combinations of grayscale conversion, normalization, and scaling. Before inputting the palm image into the model for inference, it is necessary to uniformly scale it to a preset size, which is 128×128 pixels.
[0062] S2. Input the preprocessed palm image into the palm detection and key point localization network model, output the palm target box and the coordinates of four key points K1 to K4 located at the finger gap valley of the palm, and obtain the confidence score corresponding to the palm target box and the confidence score corresponding to each key point.
[0063] In this embodiment, the hand detection and keypoint localization network model can adopt a lightweight structure of the YOLO pose estimation network. This hand detection and keypoint localization network model is a detection-pose joint network that outputs the hand target box and keypoints K1 to K4 simultaneously in one forward inference, or it is a cascaded network composed of a hand target detection network and a keypoint localization network. The four keypoints K1 to K4 are as follows from thumb to little finger: thumb-index finger valley point K1, index finger-middle finger valley point K2, middle finger-ring finger valley point K3, and ring finger-little finger valley point K4. The keypoint output shape is kpt_shape=[4,3], that is, each keypoint contains a (x,y,conf) triple. Among them, the index finger-middle finger valley point K2 and the ring finger-little finger valley point K4 are used for pose and scale calculation, while the thumb-index finger valley point K1 and the middle finger-ring finger valley point K3 are used for auxiliary localization and validity verification.
[0064] S3. Determine the validity of the input sample based on the confidence level of the target bounding box and the confidence level of the key points. If the validity determination result is an invalid sample, output an invalid flag and end the process. If the validity determination result is a valid sample, proceed to step S4.
[0065] In this embodiment, when the confidence level of the hand target bounding box is less than the first preset threshold τ_det, it is determined to be an invalid sample, a non-hand / invalid sample flag is output, and the detection is terminated; when the confidence level of the key point used for ROI construction is less than the second preset threshold τ_kpt, or the geometric relationship of the key point (e.g., the distance between key points, the relative position relationship) does not meet the legality constraints, the ROI backtracking strategy in step S5 is triggered, so as to ensure that a valid and usable ROI region can still be output even when the key point positioning is unstable or occluded; wherein, the legality constraints include one or more combinations of key point order abnormality, key point spacing abnormality, and ROI out-of-bounds risk.
[0066] S4. Calculate the palm posture correction parameters based on the coordinates of the key points, and perform rotation correction on the palm image based on the palm posture correction parameters;
[0067] In this embodiment, the attitude correction parameters are determined by key point A and key point B, wherein key point A is the index-middle finger valley point K2, and key point B is the ring-little finger valley point K4; based on the obtained vector... Angle with the horizontal direction Rotation correction is applied to the palm image; the palm image is rotated. This ensures that the line connecting keypoints A and B is parallel or approximately parallel to the horizontal direction. The rotation correction is achieved using an affine transformation, meaning the coordinates of the keypoints are simultaneously subjected to the same rotation transformation to obtain the rotated keypoint coordinates A and B. The rotated image boundaries are then filled to avoid missing effective areas.
[0068] Horizontal angle for:
[0069] ;
[0070] In the formula: These are the pixel coordinates of two target key points in a Cartesian coordinate system.
[0071] S5. On the rotated and corrected palm image, first determine whether the key points used to construct the ROI region meet the preset legality constraints: if the key points meet the preset legality constraints, then determine the ROI region based on the geometric relationship of the key points and crop it to obtain the ROI image. Then, scale the ROI image to the preset input specifications to obtain the ROI normalized image; if the key points do not meet the preset legality constraints, then use the ROI backtracking strategy to generate an ROI normalized image that conforms to the preset input specifications; subsequently, input the generated ROI normalized image into the palm vein representation attack detection network model to obtain the attack probability of the corresponding input sample; in this embodiment, the steps for determining the ROI region are as follows:
[0072] A local coordinate system is established with the line connecting key point A and key point B as the X-axis and the perpendicular bisector of that line as the Y-axis. The origin O of this local coordinate system is the midpoint of the line connecting key point A and key point B. A distance of [missing information] from the origin O is selected on the Y-axis. Point C; with point C as the center and side length as... The square region is used as the ROI and then cropped. and This is a preset proportional coefficient. Take a value of 0.5 to 1.2; The value is taken as 0.8 to 1.8. In this embodiment, Take 0.8; The preset scaling factor is 1.2; it can also be adjusted according to factors such as acquisition distance, lens field of view, and palm image size to achieve multi-scale ROI region extraction.
[0073] Furthermore, the cropped palm vein ROI image is first scaled to a preset size of 128×128 or 224×224 pixels and converted into a single-channel grayscale image. During edge implementation, the single-channel grayscale image is normalized to map pixel values to the [-1,1] range, forming a palm vein representation attack detection network model input tensor in a 1×128×128 format. If the vein representation attack detection network model requires RGB image input, the grayscale ROI can be copied into a three-channel format, or a single-channel / pseudo-color three-channel near-infrared image can be directly used as input to match the input constraints of the vein representation attack detection network model.
[0074] When the confidence level of key points is insufficient or the geometric relationship is abnormal, an ROI rollback strategy is adopted to ensure the robustness of the process. The ROI rollback strategy includes any one or more combinations of the following strategies 1 to 4:
[0075] Strategy 1: Target bounding box backtracking: Using the center of the hand target bounding box as the center of the ROI region, generate and crop the ROI region according to the width and height of the hand target bounding box and a preset scaling factor.
[0076] Strategy 2, Key Point Substitution and Regression: Replace the original key point A and the original key point B with the line connecting the index finger-middle finger valley point K2 and the middle finger-ring finger valley point K3, or the line connecting the thumb-index finger valley point K1 and the ring finger-little finger valley point K4, and calculate the palm posture correction parameters and scale parameters.
[0077] Strategy 3, Temporal Multiplexing Backoff: When the input is a series of video frames, reuse the affine transformation parameters corresponding to the effective ROI region of the previous frame, and perform smooth update processing on the affine transformation parameters.
[0078] Strategy 4, Skip Rotation Backtracking: Skip the rotation correction step of the palm image and directly perform ROI region cropping operation based on the palm target bounding box in the original image coordinate system.
[0079] In this embodiment, the end-side palm vein representation attack detection network model can adopt a lightweight convolutional neural network structure (as shown in Figure 2), including a feature extraction backbone network and a classification head. The feature extraction backbone network integrates lightweight convolution operators (e.g., depthwise separable convolution or BSConvU), efficient feature generation structures (e.g., Ghost modules), and channel attention structures (e.g., ECA). This lightweight design can significantly reduce the number of model parameters, adapting to the real-time inference requirements of the RV1106 end-side device. During the training phase, surrogate supervision / texture enhancement branches can be selected (e.g., texture-sensitive feature construction based on central difference convolution CDC, and auxiliary supervision based on gradient or Laplacian response) to improve the model's feature learning ability. During the deployment and inference phase, all auxiliary branches are turned off, and only the feature extraction backbone network and the classification head are retained, thereby further improving the end-side inference speed and operational stability.
[0080] The ROI-normalized image is input into the end palmar vein representation attack detection network model, which outputs classification logits or category probabilities. The attack probability p is obtained by taking the softmax output of the attack category, i.e., p = softmax(attack).
[0081] S6. The attack probability is compared with a preset threshold. If the attack probability is greater than the preset threshold, it is determined to be an attack sample; if the attack probability is less than or equal to the preset threshold, it is determined to be a real sample. The system outputs the attack probability and the preset threshold simultaneously for log recording, interface display, or secondary decision-making by upper-layer business. When the detection input is a continuous video frame, one or both of the ROI affine parameters and attack probabilities corresponding to the continuous frame are smoothly updated. Through dynamic iterative optimization of inter-frame parameters and probabilities, the random fluctuations of single-frame detection are suppressed, and the stable output of palm vein attack detection results under continuous video frames is achieved.
[0082] In this embodiment, the preset threshold can be set as a fixed threshold, or the threshold size can be adaptively adjusted according to the image quality evaluation results. The quality evaluation results include indicators such as brightness statistics, sharpness statistics, and the proportion of reflection / overexposure areas. Alternatively, the quality score calculated using the texture energy index corresponding to the surrogate supervision branch can be used to maintain a balance and stability between the false rejection rate and the false acceptance rate under different lighting conditions and imaging quality.
[0083] In this embodiment, the palm detection and keypoint localization network model and / or the end-palm vein representation attack detection network model can both be trained using a self-built dataset. The self-built dataset can include real near-infrared palm samples and various representation attack samples, such as printing attacks, screen replays, silicone / resin molds, glove occlusion, etc., and cover variations in pose, distance, lighting, skin color, background, etc., to improve generalization ability (examples of attack samples can be found in Figure 6). Furthermore, the palm detection and keypoint localization network model and the end-palm vein representation attack detection network model can be individually or simultaneously derived as ONN models from PyTorch models. The X format is then converted to NCNN format and deployed to the RV1106 edge platform. During the debugging phase, image acquisition, image preprocessing, model inference, and display logic verification are first completed on the PC (e.g., VS2022 environment) (as shown in Figures 6 and 7), and then migrated to the RV1106 edge platform for real-time operation. On the RV1106 edge platform, FP16 and / or INT8 quantization methods are used to accelerate inference and compress the model. Finally, the target bounding box of the hand, key points, and judgment results are superimposed on the screen (as shown in Figure 7), and the processing time of each stage such as detection and attack detection can be statistically analyzed to verify the real-time inference performance of the system.
[0084] In addition, the ROI affine parameters of consecutive frames can be smoothed by using moving average or exponential moving average, and the attack probability of consecutive frames can be fused in a temporal manner to improve the decision stability of consecutive frames.
[0085] Example 2
[0086] The present invention provides a key-point-guided end-palm vein representation attack detection device, comprising:
[0087] The image acquisition and preprocessing module is used to acquire single-frame images or continuous video frames of the palm captured by the near-infrared imaging device, and to preprocess the acquired palm images.
[0088] The palm detection and key point localization module is used to input the preprocessed palm image into the palm detection and key point localization network model, and output the palm target box, the coordinates of four key points located at the palm finger gap valley, the confidence score of the palm target box, and the confidence score of each key point.
[0089] The sample validity determination module is used to determine the validity of the input sample based on the confidence level of the palm target box and the confidence level of each key point; if the sample is determined to be invalid, an invalid flag is output and the detection is terminated; if the sample is determined to be valid, the palm posture correction module is triggered to work.
[0090] The hand posture correction module is used to calculate hand posture correction parameters based on the coordinates of the key points, and to perform rotation correction on the hand image based on the hand posture correction parameters.
[0091] The ROI processing and attack probability acquisition module is used to first determine whether the key points used to construct the ROI region on the rotated and corrected palm image meet the preset legality constraints. If they do, the ROI region is determined and cropped based on the geometric relationship of the key points to obtain the ROI image. Then, the ROI image is scaled to the preset input specifications to obtain the ROI normalized image. If it does not meet the requirements, an ROI backtracking strategy is used to generate an ROI normalized image that meets the preset input specifications. Subsequently, the ROI normalized image is input to the palm vein representation attack detection network model to obtain the attack probability of the corresponding input sample.
[0092] The attack detection and judgment module is used to compare the attack probability with a preset threshold, output the final attack detection and judgment result based on the comparison result, and simultaneously output the attack probability and the preset threshold.
[0093] The above are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above content. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A method for detecting end-palmar vein representation attacks based on keypoint guidance, characterized in that, Includes the following steps: S1. Acquire single-frame images or continuous video frames of the palm using a near-infrared imaging device, and preprocess the acquired palm images. S2. Input the preprocessed palm image into the palm detection and key point localization network model, output the palm target box and the coordinates of four key points K1 to K4 located at the finger gap valley of the palm, and obtain the confidence score corresponding to the palm target box and the confidence score corresponding to each key point. S3. Determine the validity of the input samples based on the confidence scores of the target bounding box and key points. If the validity determination result is an invalid sample, then output an invalid flag and end; If the validity determination result is a valid sample, proceed to step S4; S4: Calculate the palm posture correction parameters based on the coordinates of the key points, and perform rotation correction on the palm image based on the palm posture correction parameters; S5: On the rotated and corrected palm image, first determine whether the key points used to construct the ROI region meet the preset legality constraints: if the key points meet the preset legality constraints, determine the ROI region based on the geometric relationship of the key points and perform cropping to obtain the ROI image, and then scale the ROI image to the preset input specifications to obtain the normalized ROI image; If the key point does not meet the preset legality constraints, the ROI backtracking strategy is used to generate a normalized ROI image that conforms to the preset input specifications; then, the generated normalized ROI image is input into the palm vein representation attack detection network model to obtain the attack probability of the corresponding input sample; S6, the attack probability is compared with the preset threshold, and the judgment result is output according to the comparison result, while the attack probability and the preset threshold are output.
2. The method for detecting end-palmar vein representation attacks based on key point guidance according to claim 1, characterized in that, In step S1, the near-infrared imaging device includes a near-infrared camera module and a near-infrared supplementary light source; the preprocessing includes one or more combinations of grayscale conversion, normalization, and scaling.
3. The method for detecting end-palmar vein representation attacks based on key point guidance according to claim 2, characterized in that, In step S2, the four key points K1 to K4 are arranged in the following order from thumb to little finger: thumb-index finger valley point K1, index finger-middle finger valley point K2, middle finger-ring finger valley point K3, and ring finger-little finger valley point K4.
4. The method for detecting attacks on end-palmar vein representation based on key point guidance according to claim 3, characterized in that, In step S3, when the confidence level of the hand target box is less than the first preset threshold, it is determined to be an invalid sample, a non-hand / invalid sample flag is output, and the detection is terminated; when the confidence level of the key point used for ROI construction is less than the second preset threshold or the geometric relationship of the key point does not meet the legality constraint, the ROI rollback strategy in step S5 is triggered; wherein, the legality constraint includes one or more combinations of key point order abnormality, key point spacing abnormality, and ROI out-of-bounds risk.
5. The method for detecting end-palmar vein representation attacks based on key point guidance according to claim 4, characterized in that, In step S4, the attitude correction parameters are determined by key point A and key point B, wherein key point A is the index-middle finger valley point K2, and key point B is the ring-little finger valley point K4; based on the obtained vector... Angle with the horizontal direction Rotate the palm image to make the line connecting key point A and key point B parallel to the horizontal direction.
6. The method for detecting end-palmar vein representation attacks based on key point guidance according to claim 5, characterized in that, In step S5, the steps for determining the ROI region are as follows: A local coordinate system is established with the line connecting key point A and key point B as the X-axis and the perpendicular bisector of that line as the Y-axis; the origin O of the local coordinate system is the midpoint of the line connecting key point A and key point B; a distance of [missing information] from the origin O is selected on the Y-axis. Point C; with point C as the center and side length as... The square region is used as the ROI region and then cropped; where, and This is a preset proportional coefficient. Take a value of 0.5 to 1.2; Take a value of 0.8 to 1.
8.
7. The method for detecting end-palmar vein representation attacks based on key point guidance according to claim 6, characterized in that, In step S5, the ROI rollback strategy adopts any one or more combinations of the following strategies 1 to 4: Strategy 1: Using the center of the palm target box as the center of the ROI region, the ROI region is generated according to the width and height of the palm target box and a preset scaling factor; Strategy 2: The line connecting the index finger-middle finger valley point K2 and the middle finger-ring finger valley point K3, or the line connecting the thumb-index finger valley point K1 and the ring finger-little finger valley point K4, is used to replace the line connecting the original key point A and the original key point B, and the palm pose correction parameters and scale parameters are calculated accordingly; Strategy 3: When the input is a continuous video frame, the affine transformation parameters corresponding to the effective ROI region of the previous frame are reused, and the affine transformation parameters are smoothly updated; Strategy 4: Skip the rotation correction step of the palm image, and directly perform the ROI region cropping operation based on the palm target box in the original image coordinate system.
8. The method for detecting end-palmar vein representation attacks based on key point guidance according to claim 7, characterized in that, In step S5, the end-palm vein representation attack detection network model includes a feature extraction backbone network and a classification head, wherein the feature extraction backbone network includes depthwise separable convolutional units, Ghost feature generation units, and channel attention units.
9. The method for detecting end-palmar vein representation attacks based on key point guidance according to claim 7, characterized in that, In step S6, if the attack probability is greater than a preset threshold, it is determined to be an attack sample; if the attack probability is less than or equal to the preset threshold, it is determined to be a real sample.
10. A device for detecting attacks on end-palmar vein representation based on key point guidance, characterized in that, include: The image acquisition and preprocessing module is used to acquire single-frame images or continuous video frames of the palm captured by the near-infrared imaging device, and to preprocess the acquired palm images. The palm detection and key point localization module is used to input the preprocessed palm image into the palm detection and key point localization network model, and output the palm target box, the coordinates of four key points located at the palm finger gap valley, the confidence score of the palm target box, and the confidence score of each key point. The sample validity determination module is used to determine the validity of the input sample based on the confidence level of the target bounding box of the palm and the confidence level of each key point; If a sample is determined to be invalid, an invalid flag is output and the detection is terminated; if a sample is determined to be valid, the hand pose correction module is triggered. The hand pose correction module is used to calculate hand pose correction parameters based on the coordinates of the key points, and to perform rotation correction on the hand image based on the hand pose correction parameters. The ROI processing and attack probability acquisition module is used to first determine whether the key points used to construct the ROI region meet the preset legality constraints on the rotated and corrected hand image; if they meet the constraints, the ROI region is determined and cropped based on the geometric relationship of the key points to obtain the ROI image, and then the ROI image is scaled to the preset input specifications to obtain the ROI normalized image; if the constraints do not meet the constraints, an ROI backtracking strategy is used to generate an ROI normalized image that meets the preset input specifications. Subsequently, the palmar vein representation of the input end of the ROI normalized image is used to represent the attack detection network model and obtain the attack probability of the corresponding input sample. The attack detection and judgment module is used to compare the attack probability with a preset threshold, output the final attack detection and judgment result according to the comparison result, and simultaneously output the attack probability and the preset threshold.