Embedded palm vein image recognition method and system based on YOLO lightweight model
Through the simplification of the YOLO model, pruning, knowledge distillation, mixing and quantization, multi-scale feature extraction and anatomical attention optimization, the contradiction between complexity and accuracy in palm vein recognition of the YOLO model is solved, and efficient and accurate palm vein recognition is achieved, which is suitable for embedded devices.
Patent Information
- Application Number
- CN202510733603.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
There is a contradiction between model complexity and recognition accuracy in palm vein recognition in existing YOLO models, which makes it difficult to achieve efficient and low-latency recognition on resource-constrained embedded devices. In addition, the palm vein image contrast is low, and the venous depth changes affect the recognition effect. How to effectively extract multi-level vein features and maintain recognition accuracy is a challenge.
By analyzing the contributions of each layer of the YOLO model, the network is streamlined by deep separable convolution and pruning technology, combining knowledge distillation and mixing precision quantization, multi-scale feature extraction and anatomical attention mechanism are introduced, specific loss functions are designed, and network structure is optimized using adaptive feature enhancement modules.
It significantly improves the accuracy and efficiency of palm vein recognition, reduces the computing resource requirements, is suitable for resource-constrained embedded devices, improves the accuracy of low-contrast image recognition, and enhances the performance of the system in complex environments.
Smart Images

Figure CN120260084A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biometric image recognition, and particularly relates to an embedded palm vein image recognition method and system based on a YOLO lightweight model. Background Art
[0002] Palm vein recognition technology has important applications in the field of biometrics. However, it faces the contradiction between model complexity and recognition accuracy. How to achieve high-precision and low-latency recognition on resource-constrained embedded devices remains a major challenge. Although the traditional YOLO model performs well in object detection, directly applying it to palm vein recognition will result in an overly large model, which cannot meet the computing power and storage limitations of embedded devices. At the same time, palm vein images usually have low contrast and contain multi-level structures such as main blood vessels, secondary branches, and capillaries. How to effectively extract these features and maintain recognition accuracy has become a thorny problem. In addition, the change in vein depth will also affect the recognition effect. How to introduce anatomical knowledge into the model to enhance the attention to veins of different depths, balance the contradiction between model size, computational complexity, and recognition accuracy during the model optimization process, and design a suitable loss function to guide the model to learn hierarchical vein feature representations are all key issues that need to be deeply studied. Finally, how to deploy the optimized model to an embedded device and achieve real-time recognition while ensuring recognition accuracy is the core challenge faced by existing technical solutions. Summary of the Invention
[0003] To solve the above technical problems, the present invention provides an embedded palm vein image recognition method and system based on a YOLO lightweight model. Among them, an embedded palm vein image recognition method based on a YOLO lightweight model includes:
[0004] Obtain the network structure of the original YOLO model, determine the redundant layers and the number of filters by analyzing the contribution of each convolutional layer to palm vein recognition, and replace the standard convolution with depthwise separable convolution to obtain a refined first network structure;
[0005] Based on the first network structure, use pruning technology to calculate the importance of each connection for extracting vein texture features. If the importance is lower than a preset threshold, remove the corresponding connection to obtain a pruned second network structure;
[0006] Extract the feature representation of the palm vein image from a pre-established teacher network, transfer the features to the second network structure through knowledge distillation, and adjust the parameters of the student network to obtain a third network structure with preliminary recognition ability;
[0007] Analyze the sensitivity of each layer in the third network structure to low-contrast palm vein images, and adopt a mixed-precision quantization strategy to allocate 16-bit width representation to sensitive layers and 8-bit width representation to non-sensitive layers, obtaining a quantized fourth network structure;
[0008] By designing a multi-level feature extraction module, add a feature pyramid network to the fourth network structure to capture the features of main blood vessels, secondary branches, and capillary networks at different resolutions, obtaining a fifth network structure containing multi-scale information;
[0009] According to the vein hierarchical features in the fifth network structure, introduce an anatomy-guided attention mechanism, and obtain the attention of the model to veins at different depths by calculating the correlation between each layer of feature maps and vein depth, obtaining an optimized sixth network structure;
[0010] Design a loss function according to the output features of the sixth network structure, and adjust the network weights by measuring the hierarchical representation error of main blood vessels, secondary branches, and capillary features, obtaining a seventh network structure that accurately extracts hierarchical features;
[0011] Obtain the running data of the seventh network structure on the embedded device, magnify the texture details of the low-contrast palm vein image through an adaptive feature enhancement module, and judge whether the enhanced features meet the recognition accuracy requirements, obtaining a final eighth network structure; perform palm vein image recognition based on the eighth network structure to obtain an image recognition result.
[0012] Preferably, the process of obtaining the simplified first network structure includes:
[0013] Obtain the network structure of the original YOLO model, decompose each layer's parameters and output features through convolutional layer decomposition, and determine the preliminary analysis result;
[0014] By calculating the contribution of the output features of each convolutional layer to palm vein recognition, judge redundant layers and unnecessary filters, obtaining a redundancy determination set;
[0015] Use statistical methods to analyze the filter usage rate in the redundancy determination set, determine the redundant layers and the number of filters to be removed, obtaining a simplified layer configuration;
[0016] Extract the remaining convolutional layers from the simplified layer configuration, and replace the standard convolution with depthwise separable convolution, obtaining a converted first network structure;
[0017] Run the palm vein dataset through the first network structure, obtain the recognition rate and computational complexity, and judge whether they meet the preset threshold, obtaining a performance evaluation result;
[0018] If the performance evaluation result is lower than the preset threshold, adjust the parameters of the depthwise separable convolution to obtain an optimized network structure; otherwise, directly use it as the first network structure.
[0019] Preferably, the process of obtaining the second pruned network structure includes:
[0020] Obtain the weight data of each connection in the first network structure through pruning technology, calculate the contribution degree of each connection to the vein texture using a convolutional neural network, and obtain the importance value;
[0021] Compare the importance value with the preset threshold. If the importance value is lower than the preset threshold, mark the corresponding connection as a to-be-removed state to obtain a preliminary list of removed connections;
[0022] Based on the preliminary list of removed connections, perform a connection deletion operation on the first network structure using a matrix operation tool to obtain a temporarily pruned network structure;
[0023] For the temporarily pruned network structure, extract the feature data of the vein texture through texture analysis technology, judge the integrity of feature extraction. If the integrity meets the preset conditions, extract connection evaluation data from the temporarily pruned network structure;
[0024] According to the connection evaluation data, classify the effectiveness of the connections using a support vector machine to obtain an optimized connection distribution;
[0025] Based on the optimized connection distribution, perform a final adjustment on the temporarily pruned network structure to obtain the second pruned network structure.
[0026] Preferably, the process of obtaining the third network structure with preliminary recognition ability includes:
[0027] Extract the feature representation of the palm vein image through a pre-established teacher network to obtain an initial feature set;
[0028] Adopt the knowledge distillation method to transfer features from the initial feature set to the second network structure to obtain a transferred feature set;
[0029] Adjust the parameters of the second network structure according to the transferred feature set to obtain optimized network parameters;
[0030] Train the second network structure with the optimized network parameters to obtain the second network structure with preliminary recognition ability;
[0031] Obtain the output result of the second network structure. If the recognition ability is lower than the preset threshold, update and adjust the parameters through a feedback mechanism to obtain the updated third network structure with preliminary recognition ability.
[0032] Preferably, the process of obtaining the quantized fourth network structure includes:
[0033] Analyze the responses of each layer in the third network to the low-contrast palm vein image through a convolutional neural network to obtain the sensitivity distribution;
[0034] According to the sensitivity distribution, judge the sensitive layer and the non-sensitive layer, and determine the layer assignment result;
[0035] According to the layer assignment result, adopt a mixed-precision quantization strategy, assign a 16-bit width representation to the sensitive layer to obtain the first quantization parameter; assign an 8-bit width representation to the non-sensitive layer to obtain the second quantization parameter;
[0036] According to the first quantization parameter and the second quantization parameter, fuse and generate the fourth network structure;
[0037] If there is a precision loss in the fourth network structure, adjust the quantization parameters through the gradient descent algorithm to obtain the quantized fourth network structure.
[0038] Preferably, the process of obtaining the fifth network structure containing multi-scale information includes:
[0039] Extract feature sets from the palm vein image data through a multi-level module to obtain a preliminary feature set;
[0040] Use a feature pyramid network to process the preliminary feature set, capture features of different resolutions, and generate multi-scale feature maps;
[0041] Separate the features of the main blood vessels from the multi-scale feature maps to obtain the first blood vessel feature subset;
[0042] Analyze the secondary branches according to the first blood vessel feature subset to obtain the second blood vessel feature subset;
[0043] Detect the capillary network through the second blood vessel feature subset to determine the third blood vessel feature subset;
[0044] Obtain the fusion information of the third blood vessel feature subset and the multi-scale feature maps, and judge the output features of the fifth network structure;
[0045] Optimize the network structure according to the output features of the fifth network structure to obtain the final multi-scale blood vessel network representation.
[0046] Preferably, the process of obtaining the optimized sixth network structure includes:
[0047] Obtain hierarchical feature data according to the venous network structure image, and extract the venous depth information;
[0048] Use a convolutional neural network to extract features from the venous network structure image to obtain multi-layer feature maps;
[0049] Construct a depth correlation matrix by calculating the correlation coefficient between each layer of feature maps and the vein depth information;
[0050] If the correlation coefficient is greater than a preset threshold, the weight of the corresponding layer of feature maps is increased;
[0051] If the correlation coefficient is less than a preset threshold, the weight of the corresponding layer of feature maps is decreased;
[0052] Design an attention mechanism module according to the adjusted feature map weights to generate an attention map;
[0053] Perform weighted fusion of the attention map and the original feature map to obtain an enhanced feature representation;
[0054] Based on the enhanced feature representation, use a fully connected layer for vein depth classification and output an optimized sixth network structure.
[0055] Preferably, the process of obtaining a seventh network structure for accurately extracting hierarchical features includes:
[0056] Obtain output features through the sixth network, and use convolution operations to extract the spatial distributions of major blood vessels, secondary branches, and capillaries to obtain an initial hierarchical representation;
[0057] For the initial hierarchical representation, construct a loss function, calculate the representation errors of major blood vessels, secondary branches, and capillaries, and determine the error distribution;
[0058] If the error distribution exceeds a preset threshold, adjust the network weights through backpropagation to obtain updated weight parameters;
[0059] Optimize the sixth network structure using the updated weight parameters to generate a seventh network and extract accurate hierarchical representations;
[0060] According to the hierarchical representation of the seventh network, judge the feature consistency of major blood vessels, secondary branches, and capillaries to obtain a feature extraction result;
[0061] According to the feature extraction result, obtain a hierarchical description of the blood vessel structure and determine the final output of the seventh network;
[0062] Extract the optimized feature distribution from the final output of the seventh network to obtain an accurate hierarchical representation.
[0063] Preferably, the process of obtaining a final eighth network structure includes:
[0064] Obtain the operation data of the seventh network structure on an embedded device, including computing resource occupancy, processing speed, and recognition accuracy;
[0065] Preprocess the low-contrast palm vein image through an adaptive feature enhancement module to extract the texture detail information in the image, and use the histogram equalization algorithm to enhance the contrast of the texture detail information to generate an enhanced palm vein feature image;
[0066] Judge whether the image features of the enhanced palm vein feature image meet the recognition accuracy requirements according to the preset threshold. If not, return to the adaptive feature enhancement module for parameter adjustment;
[0067] For the features that meet the accuracy requirements, use a convolutional neural network to extract deep features, construct a palm vein feature vector, and use a support vector machine classifier to classify and identify the feature vector to obtain the palm vein identity recognition result;
[0068] Optimize and adjust the network structure according to the recognition result and operation data to obtain the final eighth network structure.
[0069] The present invention also provides an embedded palm vein image recognition system based on the YOLO lightweight model, including:
[0070] A network simplification module, which is used to obtain the network structure of the original YOLO model, determine the redundant layers and the number of filters by analyzing the contribution of each convolutional layer to palm vein recognition, and use depthwise separable convolution to replace the standard convolution to obtain the simplified first network structure;
[0071] A pruning and optimization module, which is used to calculate the importance of each connection for vein texture feature extraction based on the first network structure by using pruning technology. If the importance is lower than the preset threshold, remove the corresponding connection to obtain the pruned second network structure;
[0072] A knowledge distillation module, which is used to extract the feature representation of palm vein images from a pre-established large teacher network, transfer the features to the second network structure through the knowledge distillation method, and adjust the parameters of the student network to obtain the third network structure with preliminary recognition ability;
[0073] A hybrid quantization module, which is used to analyze the sensitivity of each layer in the third network structure to low-contrast palm vein images, adopt a mixed-precision quantization strategy, allocate 16-bit width representation to sensitive layers and 8-bit width representation to non-sensitive layers to obtain the quantized fourth network structure;
[0074] A multi-scale feature module, which is used to capture the features of the main blood vessels, secondary branches and capillary networks at different resolutions by designing a multi-level feature extraction module and adding a feature pyramid network to the fourth network structure to obtain the fifth network structure containing multi-scale information;
[0075] An attention enhancement module, which is used to introduce an anatomy-guided attention mechanism according to the vein hierarchical features in the fifth network structure, obtain the attention of the model to veins at different depths by calculating the correlation between each layer of feature maps and the vein depth, and obtain an optimized sixth network structure;
[0076] A loss optimization module, which is used to design a loss function according to the output features of the sixth network structure, adjust the network weights by measuring the hierarchical representation errors of the main blood vessels, secondary branches and capillary features, and obtain a seventh network structure that accurately extracts hierarchical features;
[0077] An adaptive enhancement module, which is used to obtain the operation data of the seventh network structure on the embedded device, magnify the texture details of the low-contrast palm vein image through the adaptive feature enhancement module, judge whether the enhanced features meet the recognition accuracy requirements, and obtain a final eighth network structure; perform palm vein image recognition based on the eighth network structure to obtain an image recognition result.
[0078] Compared with the prior art, the present invention has the following advantages and technical effects:
[0079] The present invention first analyzes the contribution of each layer of the YOLO model to palm vein recognition, and uses depthwise separable convolution and pruning techniques to streamline the network structure. Subsequently, knowledge distillation is used to transfer features from a large teacher network, and a mixed-precision quantization strategy is introduced to adapt to low-contrast images. Through the multi-level feature extraction module and the anatomy-guided attention mechanism, the present invention enhances the recognition ability of veins at different depths. Finally, the hierarchical feature representation is optimized through a specific loss function, and the performance is verified on an embedded device. This method significantly improves the accuracy and efficiency of palm vein recognition, while reducing the computational resource requirements, providing an effective solution for biometric recognition on embedded devices.
[0080] The present invention realizes the lightweight, high-efficiency and precision of the palm vein recognition model, significantly improves the recognition accuracy and operation efficiency, and is particularly suitable for resource-constrained embedded devices.
[0081] The present invention realizes adaptive feature enhancement on an embedded device, effectively improving the recognition accuracy of low-contrast palm vein images. This method significantly improves the performance and efficiency of the palm vein recognition system in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0083] Figure 1 is a schematic flowchart of the method according to the embodiment of the present invention;
[0084] Figure 2 This is a schematic diagram of the system structure according to an embodiment of the present invention. Detailed implementation manners
[0085] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0086] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0087] Embodiment 1
[0088] As Figure 1 shown, in this embodiment, an embedded palm vein image recognition method based on the YOLO lightweight model is provided, including:
[0089] Obtain the network structure of the original YOLO model, determine the redundant layers and the number of filters by analyzing the contribution of each convolutional layer to palm vein recognition, and replace the standard convolution with depthwise separable convolution to obtain a refined first network structure;
[0090] Based on the first network structure, use pruning technology to calculate the importance of each connection for extracting vein texture features. If the importance is lower than a preset threshold, remove the corresponding connection to obtain a pruned second network structure;
[0091] Extract the feature representation of the palm vein image from a pre-established teacher network, transfer the features to the second network structure through knowledge distillation, and adjust the parameters of the student network to obtain a third network structure with preliminary recognition ability;
[0092] Analyze the sensitivity of each layer in the third network structure to low-contrast palm vein images, adopt a mixed-precision quantization strategy, allocate 16-bit width representation to sensitive layers and 8-bit width representation to non-sensitive layers to obtain a quantized fourth network structure;
[0093] By designing a multi-level feature extraction module, add a feature pyramid network to the fourth network structure to capture the features of the main blood vessels, secondary branches, and capillary networks at different resolutions, and obtain a fifth network structure containing multi-scale information;
[0094] According to the vein hierarchical features in the fifth network structure, introduce an anatomy-guided attention mechanism, calculate the correlation between each layer of feature maps and the vein depth, and obtain the attention of the model to veins at different depths to obtain an optimized sixth network structure;
[0095] Design a loss function based on the output features of the sixth network structure. By measuring the hierarchical representation error of the features of major blood vessels, secondary branches, and capillaries, adjust the network weights to obtain the seventh network structure that can accurately extract hierarchical features.
[0096] Obtain the operation data of the seventh network structure on the embedded device. Use the adaptive feature enhancement module to magnify the texture details of the low-contrast palm vein image, and determine whether the enhanced features meet the recognition accuracy requirements to obtain the final eighth network structure. Perform palm vein image recognition based on the eighth network structure to obtain the image recognition result.
[0097] Furthermore, the process of obtaining the streamlined first network structure includes:
[0098] Obtain the network structure of the original YOLO model. Through convolutional layer decomposition, obtain the parameters and output features of each layer to determine the preliminary analysis result.
[0099] By calculating the contribution degree of the output features of each convolutional layer to palm vein recognition, judge the redundant layers and unnecessary filters to obtain the redundancy determination set.
[0100] Use statistical methods to analyze the filter utilization rate of the redundancy determination set, determine the redundant layers and the number of filters to be removed to obtain the streamlined layer configuration.
[0101] Extract the remaining convolutional layers from the streamlined layer configuration and use depthwise separable convolutions to replace the standard convolutions to obtain the converted first network structure.
[0102] Run the palm vein dataset through the first network structure, obtain the recognition rate and computational complexity, and judge whether they meet the preset threshold to obtain the performance evaluation result.
[0103] If the performance evaluation result is lower than the preset threshold, adjust the parameters of the depthwise separable convolution to obtain the optimized network structure; otherwise, directly use it as the first network structure.
[0104] Specifically, the network structure of the original YOLO model usually consists of multiple convolutional layers for feature extraction and object detection. In the palm vein recognition scenario, its network structure can be analyzed to decompose the parameters and output features of each convolutional layer.
[0105] For example, assume that the original YOLO model contains 20 convolutional layers, and each layer has a different number of filters. For example, the first layer has 32 filters, and the output feature map is 416×416×32. Through layer-by-layer analysis, it can be found that the output features of some layers contribute less to the discrimination of palm vein textures. Therefore, the model structure can be streamlined without affecting the recognition result.
[0106] Specifically, when calculating the contribution of the output features of each convolutional layer to palm vein recognition, this embodiment adopts a feature visualization method to observe whether the feature map of each layer effectively captures the texture details of the palm vein.
[0107] For example, if the feature map of layer 5 shows a large amount of low-frequency redundant information, while layer 10 can highlight the edges of palm veins, then layer 5 may be determined as a redundant layer.
[0108] Preferably, by counting the activation frequencies of the filters, assuming that only 10 of the 32 filters in the fifth layer are frequently activated, the remaining 22 filters can be regarded as non-essential filters to form a redundant decision set.
[0109] In one possible implementation, the filter usage rate is analyzed for the redundant decision set, and the average activation value of each layer of filters on the training set can be counted. For example, among the 64 filters in the 8th layer, 40 activation values are lower than 0.1, then these 40 are removed and 24 are retained to obtain a streamlined layer configuration. This method can effectively reduce the amount of calculation while retaining the recognition ability as much as possible.
[0110] It should be noted that the number of layers after pruning may be reduced from 20 to 15, and the total number of filters is reduced from 512 to 300. After extracting the retained convolutional layers from the streamlined layer configuration, the standard convolution is replaced by the depthwise separable convolution to form the first network structure. For example, the original 3×3 standard convolution requires 9 multiplication and addition operations, while the depthwise separable convolution is split into 3×3 depthwise convolution and 1×1 point convolution, and the total amount of operations is reduced to about 3 multiplication and addition. This conversion significantly reduces the computational complexity while retaining the feature extraction capability.
[0111] In one embodiment, the palm vein data set is processed by the first network structure. Assuming that the recognition rate reaches 95%, the computational complexity is reduced from 100 GFLOPs to 30 GFLOPs. If the preset thresholds are 90% recognition rate and 50 GFLOPs, the performance meets the requirements.
[0112] It can be understood that if the recognition rate is only 85%, which is lower than the threshold, the depthwise separable convolution parameters are adjusted.
[0113] For example, by increasing the number of channels of the deep convolution from 16 to 24, the recognition rate increased to 92% after optimization, forming an optimized network structure. This adjustment can improve the model's ability to perceive palm vein details. Finally, by processing the input data through the optimized network structure, a streamlined palm vein recognition network is obtained. Its advantage is that the computing efficiency is improved by about 70%, and the recognition rate remains at a high level, which is suitable for deployment in embedded devices.
[0114] For example, in practical applications, the lightweight network can achieve palm vein recognition at 10 frames per second on low-power devices, which is a significant improvement in efficiency compared to 2 frames per second of the original YOLO model. This method not only reduces hardware requirements but also extends the device's battery life, demonstrating the practical value of technology optimization.
[0115] Furthermore, the process of obtaining the second pruned network structure includes:
[0116] Obtain the weight data of each connection in the first network structure through pruning technology, and use a convolutional neural network to calculate the contribution of each connection to the vein texture to obtain an importance value;
[0117] Compare the importance value with a preset threshold. If the importance value is lower than the preset threshold, mark the corresponding connection as to-be-removed status to obtain a preliminary list of removed connections;
[0118] Based on the preliminary list of removed connections, use a matrix operation tool to perform connection deletion operations on the first network structure to obtain a temporarily pruned network structure;
[0119] For the temporarily pruned network structure, extract the feature data of the vein texture through texture analysis technology, and judge the integrity of feature extraction. If the integrity meets the preset conditions, extract connection evaluation data from the temporarily pruned network structure;
[0120] According to the connection evaluation data, use a support vector machine to classify the effectiveness of the connections to obtain an optimized connection distribution;
[0121] Based on the optimized connection distribution, perform final adjustments on the temporarily pruned network structure to obtain the second pruned network structure.
[0122] Specifically, obtaining the weight data of each connection through pruning technology based on the first network structure, it can be understood that the weight data reflects the influence degree of each connection in the network on the result.
[0123] Exemplarily, in the palm vein recognition scenario, assume that the first network contains 1000 connections. After analysis by a pruning tool, each connection is assigned a weight value ranging from 0 to 1. Connections with high weights may correspond to key features of the texture edge, while connections with low weights may only process noise or redundant information.
[0124] In a possible implementation, when using a convolutional neural network to calculate the contribution of each connection to the vein texture, the sensitivity of each connection to the output features can be extracted through a combination of forward propagation and backward propagation. For example, a connection with a weight of 0.8 has a significant impact on the boundary detection of the vein texture; while another connection with a weight of 0.2 may only have a weak effect on the smooth area. The importance value is thus generated for subsequent screening.
[0125] Specifically, when comparing the importance value with a preset threshold, in this embodiment, the threshold is set to 0.3. If the importance value of a certain connection is 0.25, it is marked as the status to be removed. This way ensures that the connections that substantially contribute to texture recognition are retained. For example, the preliminary list may contain 300 low-importance connections, which are mostly distributed in the shallow convolutional layers and may only capture low-level features.
[0126] It should be noted that when performing connection deletion through a matrix operation tool, the network can be regarded as a sparse matrix. For example, the original network has 5 million parameters. After removing 300 connections, the number of parameters drops to 4.8 million. The temporary network structure is thus generated, and its computational burden is significantly reduced.
[0127] In one embodiment, when judging the integrity of feature extraction through texture analysis technology, the trunk and branch features of vein texture can be extracted. Assuming that the integrity requirement is to retain more than 95% of the trunk features and more than 80% of the branch features, if the temporary network meets this condition, it is confirmed as the second network. This method ensures the retention of core information.
[0128] For example, when extracting connection evaluation data from the second network, the distribution density of connections in each layer can be counted. Suppose a certain layer originally had 200 connections, and after pruning, there are 150 left. The support vector machine can classify them into two categories: "efficient" and "inefficient" to optimize the connection distribution. The efficient connections may be concentrated in the texture detail areas, and the inefficient connections are further removed.
[0129] In one possible implementation, when finally adjusting the second network, the connection weights can be fine-tuned.
[0130] For example, the weight of a key connection in a certain layer is increased from 0.7 to 0.85 to enhance the response to texture features. A stable network structure is thus formed, and its recognition efficiency and feature expression ability are both improved, which helps the stable operation of subsequent palm vein recognition tasks.
[0131] Furthermore, the process of obtaining the third network structure with preliminary recognition ability includes:
[0132] Extracting the feature representation of the palm vein image through a pre-established teacher network to obtain an initial feature set;
[0133] Adopting the knowledge distillation method to transfer features from the initial feature set to the second network structure to obtain a transferred feature set;
[0134] Adjusting the parameters of the second network structure according to the transferred feature set to obtain optimized network parameters;
[0135] Training the second network structure with the optimized network parameters to obtain the second network structure with preliminary recognition ability;
[0136] Obtain the output result of the second network structure. If the recognition ability is lower than the preset threshold, update and adjust the parameters through a feedback mechanism to obtain an updated third network structure with preliminary recognition ability.
[0137] Specifically, for the palm vein recognition task, the teacher network in this embodiment is a pre-trained large-scale deep learning model. For example, this embodiment can adopt ResNet-152 as the teacher network, which is pre-trained on the ImageNet dataset and has strong feature extraction capabilities. By processing palm vein images through this teacher network, rich feature representations can be obtained, including low-level texture features and high-level semantic features.
[0138] Knowledge distillation, as a model compression technique, can transfer the knowledge of the teacher network to a smaller student network. Specifically, the KL divergence can be used as the distillation loss function to guide the student network to learn the output distribution of the teacher network. In a possible implementation, the temperature parameter is set to 3 and the soft label weight is 0.7, which helps to balance the contributions of the distilled knowledge and the true labels.
[0139] Network parameter optimization is a key link to improve the model performance. Exemplarily, this embodiment can adopt the Adam optimizer, with the initial learning rate set to 0.001 and dynamically adjusted using the cosine annealing strategy. At the same time, L2 regularization is introduced, and the regularization coefficient is set to 0.0001, which helps to alleviate the overfitting problem. During the training process, cross-validation can be used to evaluate the recognition ability of the model.
[0140] It should be noted that if the recognition accuracy is lower than the preset threshold of 95%, the feedback mechanism is triggered.
[0141] In one embodiment, the gradient clipping technique is adopted to limit the gradient value within the range of [-1, 1], which helps to stabilize the training process.
[0142] Preferably, after the model training is completed, the test set can be used to evaluate the final recognition performance. For example, the ROC curve and AUC metrics are used to comprehensively evaluate the recognition ability of the model.
[0143] It can be understood that a high-quality recognition model can not only accurately recognize known palm vein images but also have good generalization ability for newly acquired image data, providing reliable technical support for practical applications.
[0144] Furthermore, the process of obtaining the quantized fourth network structure includes:
[0145] Analyze the responses of each layer in the third network to low-contrast palm vein images through a convolutional neural network to obtain the sensitivity distribution;
[0146] According to the sensitivity distribution, determine the sensitive layer and the non-sensitive layer, and determine the layer allocation result;
[0147] According to the layer allocation result, adopt a mixed-precision quantization strategy, allocate 16-bit width representation to the sensitive layer to obtain the first quantization parameter; allocate 8-bit width representation to the non-sensitive layer to obtain the second quantization parameter;
[0148] According to the first quantization parameter and the second quantization parameter, fuse and generate the fourth network structure;
[0149] If there is a precision loss in the fourth network structure, adjust the quantization parameter through the gradient descent algorithm to obtain the quantized fourth network structure.
[0150] Specifically, by analyzing the response of each layer in the third network to the low-contrast palm vein image through a convolutional neural network, obtaining the sensitivity distribution is a key link.
[0151] For example, by inputting a group of low-contrast palm vein images, observe the activation of each convolutional kernel layer, record which layers respond more strongly to detailed textures, and which layers output more smoothly. Suppose the third network has 10 convolutional layers. After analysis, the first 3 layers may be more sensitive to image edges and fine textures, while the latter layers tend to extract global features. This distribution reflects the network's processing characteristics for low-contrast data and helps with subsequent optimization.
[0152] In a possible implementation, when determining the sensitive layer and the non-sensitive layer, a threshold can be set. For example, layers with an average activation value exceeding 0.7 are classified as sensitive layers, and the rest are non-sensitive layers.
[0153] Exemplarily, assume that the average activation values of the first layer and the second layer are 0.8 and 0.75, and the fifth layer is only 0.3. Then the first two are classified as sensitive layers. This classification logic is clear and can provide a basis for subsequent quantization. Adopting a mixed-precision quantization strategy for the layer allocation result is an efficient resource allocation method.
[0154] Specifically, allocating 16-bit width representation to the sensitive layer can retain more detailed features.
[0155] For example, when the first layer processes low-contrast palm vein images, it may need to capture weak texture changes, and 16-bit precision can ensure that this information is not lost. Using 8-bit width representation for the non-sensitive layer can reduce the computational amount. Suppose the fifth layer mainly extracts large-scale contours, and 8 bits are sufficient to meet the requirements. This differential allocation ensures both precision and efficiency. After obtaining the first quantization parameter and the second quantization parameter, fusing and generating the fourth network structure is an important step.
[0156] In one embodiment, the feature maps represented by 16 bits and 8 bits can be integrated through parameter splicing.
[0157] For example, after adjusting the output of the sensitive layer to a unified scale, it is spliced with the output of the non-sensitive layer to generate a new feature set. This method can maintain information integrity.
[0158] It should be noted that if the accuracy drops after fusion, for example, the recognition rate drops from 90% to 85%, further adjustment is required. If there is an accuracy loss in the fourth network structure, adjusting the quantization parameters through the gradient descent algorithm is a feasible solution. For example, for the quantization parameters of the sensitive layer, the quantization step size can be gradually reduced from 0.01 to 0.005 to observe whether the accuracy recovers.
[0159] In one embodiment, after adjustment, it is found that the details of the sensitive layer are more completely retained, and the recognition rate is increased to 88%. This shows that fine-tuning the quantization parameters can effectively make up for the loss. Verifying the feature extraction ability of low-contrast palm vein images through the optimized network structure is the ultimate goal.
[0160] It can be understood that by inputting a set of test images, such as 50 low-contrast samples, it can be observed whether the network can correctly extract key texture features.
[0161] For example, in a certain image, the palm vein pattern is blurred, and the optimized network can still identify the main branches, while without optimization, only noise may be output. This improvement in ability is of great significance for practical applications.
[0162] Specifically, the implementation of the finally quantized fourth network structure can be regarded as a balance between resources and accuracy. Exemplarily, when deployed on an embedded device, the 8-bit non-sensitive layer reduces the memory occupancy by about 30%, while the 16-bit sensitive layer ensures that key features are not lost. This design not only reduces power consumption but also maintains the recognition effect, and is particularly suitable for scenarios such as palm vein recognition that are sensitive to details.
[0163] Furthermore, the process of obtaining the fifth network structure containing multi-scale information includes:
[0164] Performing feature extraction on the palm vein image data through a multi-level module to obtain a preliminary feature set;
[0165] Using a feature pyramid network to process the preliminary feature set, capturing features of different resolutions, and generating multi-scale feature maps;
[0166] Separating the features of the main blood vessels from the multi-scale feature maps to obtain a first blood vessel feature subset;
[0167] Analyzing the secondary branches according to the first blood vessel feature subset to obtain a second blood vessel feature subset;
[0168] Detect the capillary network through the second blood vessel feature subset to determine the third blood vessel feature subset;
[0169] Obtain the fusion information of the third blood vessel feature subset and the multi-scale feature map, and judge the output features of the fifth network structure;
[0170] Optimize the network structure according to the output features of the fifth network structure to obtain the final multi-scale blood vessel network representation.
[0171] Specifically, feature extraction is performed on the input data through a multi-level module to obtain a preliminary feature set. This process can be understood as decomposing the original low-contrast palm vein image into multiple basic features.
[0172] For example, in a possible implementation, edge information and texture details can be extracted through convolution operations. Assuming the input image resolution is 512×512 and the pixel value range is between 0-255, the multi-level module may contain 3 layers of convolution, each using different-sized convolution kernels, such as 3×3 and 5×5, to capture local features in different ranges. This hierarchical extraction helps to retain the key structural information in the image.
[0173] Use a feature pyramid network to process the preliminary feature set, capture features at different resolutions, and generate a multi-scale feature map. The core of this design is to solve the problem of insufficient feature expression at a single resolution. Specifically, the feature pyramid network can combine high-level semantic information with low-level detail information.
[0174] For example, the bottom-layer feature map may maintain a resolution of 256×256 to capture fine textures, while the top-layer feature map is reduced to 64×64 to extract more abstract structural features. Through the fusion operation of upsampling and downsampling, the multi-scale feature map can simultaneously reflect the main trunk and branch details of the palm vein. Separate the features of the main blood vessels from the multi-scale feature map to obtain the first blood vessel feature subset. This process requires targeted analysis of the feature map.
[0175] In one embodiment, the main blood vessel region can be distinguished through threshold segmentation technology. Assuming that the pixel intensity of the main blood vessel is 1.5 times higher than the average value, morphological operations are combined to remove noise points, thereby extracting a clear main blood vessel contour. The advantage of this method is that it can quickly locate the core structure in the image. Analyze the secondary branches for the first blood vessel feature subset to obtain the second blood vessel feature subset. This link is a further refinement of the main blood vessels.
[0176] Preferably, a direction filter can be introduced to analyze the extension direction of blood vessels. For example, a set of predefined angle templates can be used to match the orientation of secondary branches. By comparing the response intensities in different directions, the secondary branch features can be screened out. This refinement helps to improve the description ability of complex blood vessel networks. The capillary network is detected through the second blood vessel feature subset to determine the third blood vessel feature subset. This stage focuses on the recognition of tiny structures.
[0177] It can be understood that capillaries are difficult to extract directly due to low contrast. Therefore, high-resolution feature maps can be combined, and local enhancement techniques can be used to amplify weak signals.
[0178] For example, the regions in the feature map below a certain threshold are amplified, and then the capillary distribution is determined through clustering analysis. This method can effectively supplement the network's perception of fine features.
[0179] The fusion information of the third blood vessel feature subset and the multi-scale feature maps is obtained to judge the output features of the fifth network structure. This process aims to integrate multi-level information. In a possible implementation, the third blood vessel feature subset and the multi-scale feature maps can be added element by element through weighted fusion, and the weights are dynamically adjusted according to the feature importance. For example, the weight of the main blood vessel is set to 0.6, and the weight of the capillary is set to 0.3. This fusion can generate a more comprehensive feature representation. The network structure is optimized according to the output features of the fifth network structure to obtain the final multi-scale blood vessel network representation. This optimization process reflects the self-adaptability of the network. For example, the fusion weights can be adjusted through backpropagation to make the output feature map closer to the real blood vessel distribution.
[0180] Exemplarily, if there are breaks at the branch connections in the initial feature map, the network's learning of continuity can be strengthened by increasing the number of training iterations or adjusting the loss function. The advantage of this optimization is to improve the robustness and integrity of feature extraction, providing reliable support for subsequent applications.
[0181] Furthermore, the process of obtaining the optimized sixth network structure includes:
[0182] Hierarchical feature data is obtained from the venous network structure image, and the venous depth information is extracted;
[0183] A convolutional neural network is used to extract features from the venous network structure image to obtain multi-layer feature maps;
[0184] By calculating the correlation coefficients between each layer of feature maps and the venous depth information, a depth correlation matrix is constructed;
[0185] If the correlation coefficient is greater than the preset threshold, the weight of the corresponding layer of feature maps is increased;
[0186] If the correlation coefficient is less than the preset threshold, the weight of the corresponding layer feature map is reduced;
[0187] According to the adjusted feature map weights, design an attention mechanism module to generate an attention map;
[0188] Perform weighted fusion of the attention map and the original feature map to obtain an enhanced feature representation;
[0189] Based on the enhanced feature representation, use a fully connected layer for vein depth classification to output an optimized sixth network structure.
[0190] Specifically, for the vein network structure image, hierarchical feature data is obtained, and vein depth information is extracted. The convolutional neural network is a core tool. In a possible implementation, through multi-layer convolutional operations, basic features such as edges and textures can be extracted from the input vein image, gradually transitioning to more abstract deep features. Assuming the input image is a grayscale image with pixel values ranging from 0 to 255, shallow convolutions may focus on the gray-scale changes of vein edges, while deep convolutions capture the overall pattern of vein distribution. This hierarchical extraction method can gradually construct a feature representation from local to global, which helps the subsequent analysis of depth information.
[0191] Specifically, the correlation coefficient reflects the matching degree between the feature map and the depth information. Exemplarily, assume there are three layers of feature maps. The first layer captures vein edges with a correlation coefficient of 0.8; the second layer captures vein branches with a coefficient of 0.6; the third layer captures the overall structure with a coefficient of only 0.3. If the preset threshold is 0.5, the weights of the first and second layers will increase, while the weight of the third layer will decrease. This method adjusts the feature importance through data-driven means to ensure that the depth information depends more on strongly correlated features.
[0192] In one embodiment, when designing the attention mechanism module according to the adjusted feature map weights, spatial attention can be introduced, and the attention map can highlight the key regions of vein depth. For example, for a deeper vein, the attention map may generate a higher weight value, such as 0.9, in the corresponding region, while the weight in the shallow layer region is lower, such as 0.2. The attention map generated in this way can effectively guide the network to focus on the depth-related regions, thereby improving the quality of the features.
[0193] It should be noted that the weighted fusion of the attention map and the original feature map is a process of enhancing the feature representation. For example, during fusion, the weights of the attention map can be multiplied element-wise with the feature map, and the feature values in the deep vein regions are amplified while the irrelevant regions are suppressed. This fusion method makes the enhanced feature representation more focused on the depth information, providing a more reliable basis for subsequent classification. Based on the enhanced feature representation, using a fully connected layer for vein depth classification is the final output optimization step.
[0194] In a possible implementation, the fully connected layer can be designed to have three types of outputs: superficial, middle, and deep veins. Exemplarily, assuming an input image containing multiple veins, the network may output that the proportion of superficial veins is 40%, middle veins is 30%, and deep veins is 30%. Such classification results not only clearly reflect the distribution of vein depths but also provide intuitive data support for subsequent applications.
[0195] Specifically, by examining the implementation of the attention mechanism from multiple aspects, the solution can be further extended.
[0196] For example, channel attention can be combined to adjust the weights of different feature channels respectively, highlighting the feature dimensions most relevant to depth. Preferably, if a certain channel is strongly correlated with deep veins, its weight may be increased to 0.85, while the weight of a weakly correlated channel is decreased to 0.1. This multi-dimensional attention design can jointly optimize feature representation from both spatial and channel aspects.
[0197] It can be understood that each step of the above method closely focuses on the extraction of vein depth information.
[0198] For example, the calculation of the correlation coefficient ensures the scientific nature of feature selection, the introduction of the attention mechanism improves the pertinence of features, and weighted fusion and fully connected classification guarantee the accuracy of the output. These links, through a progressive logical relationship, jointly support the complete process from image to depth classification.
[0199] In one embodiment, if the resolution of the input image is 512x512, after processing, the network can effectively distinguish vein regions of different depths, providing high-quality basic data for subsequent analysis.
[0200] Furthermore, the process of obtaining the seventh network structure for accurately extracting hierarchical features includes:
[0201] Obtaining output features through the sixth network, using convolutional operations to extract the spatial distribution of main blood vessels, secondary branches, and capillaries, and obtaining an initial hierarchical representation;
[0202] For the initial hierarchical representation, constructing a loss function, calculating the representation errors of the main blood vessels, secondary branches, and capillaries, and determining the error distribution;
[0203] If the error distribution exceeds a preset threshold, then adjust the network weights through backpropagation to obtain updated weight parameters;
[0204] Optimize the sixth network structure using the updated weight parameters to generate the seventh network and extract accurate hierarchical representations;
[0205] According to the hierarchical representation of the seventh network, judge the feature consistency of the main blood vessels, secondary branches and capillaries, and obtain the feature extraction result;
[0206] According to the feature extraction result, obtain the hierarchical description of the blood vessel structure and determine the final output of the seventh network;
[0207] Extract the optimized feature distribution from the final output of the seventh network to obtain an accurate hierarchical representation.
[0208] Specifically, for the sixth network to obtain output features, the process of using convolution operations to extract the spatial distribution of the main blood vessels, secondary branches and capillaries can be understood as performing a sliding window process on the input image through a convolution kernel to capture blood vessel features at different scales.
[0209] For example, when processing venous network images, the convolution kernel size can be set to 3x3 or 5x5 to extract the outlines of thick blood vessels and the details of thin branches respectively.
[0210] Exemplarily, if the input image resolution is 256x256 and the pixel value range is between 0-255, a multi-channel feature map can be generated after convolution operations. The number of channels is set to 64 for example, and each channel highlights the spatial characteristics of a certain type of blood vessel. This method can effectively retain the spatial position information of the blood vessels and provide a basis for subsequent hierarchical representation. After constructing the initial hierarchical representation, a loss function is designed for the main blood vessels, secondary branches and capillaries, and the representation error is calculated. The core of this link lies in quantifying the accuracy of feature extraction.
[0211] Specifically, the difference between the predicted distribution and the true distribution can be measured by the mean squared error or cross-entropy loss.
[0212] For example, if the true distribution ratio of the main blood vessels is 40%, the secondary branches are 35%, and the capillaries are 25%, while the model prediction values are 38%, 36%, and 26% respectively, the error distribution can be obtained by comparing item by item. This error calculation helps to identify the bias of the model towards different blood vessel types and provides a direction for subsequent optimization.
[0213] It should be noted that if the error distribution exceeds a preset threshold, such as set to 0.05, the network weights are adjusted through backpropagation. In one possible implementation, backpropagation can be based on the gradient descent method with a learning rate set to 0.001 to gradually update the weights of the convolutional layer.
[0214] For example, when the error of the main blood vessel feature is large, the weights of the relevant convolution kernels will be adjusted more significantly to enhance the attention to this type of feature. This dynamic adjustment enables the model to better adapt to the distribution characteristics of different blood vessel types.
[0215] Preferably, the process of optimizing the sixth network structure with the updated weight parameters to generate the seventh network is similar to a structural iteration. For example, the original sixth network may have 10 convolutional layers. After optimization, the seventh network can add a feature fusion layer, and the number of channels can be increased from 64 to 128 to further integrate multi-scale features. This structural upgrade can enhance the model's ability to represent complex vascular networks.
[0216] In one embodiment, when judging the feature consistency of main blood vessels, secondary branches, and capillaries, cosine similarity can be introduced as a measurement criterion. For example, after extracting the feature vector of the main blood vessel and comparing it with the standard template vector, if the similarity is higher than 0.9, it is considered that the consistency is good. This method ensures the reliability of the hierarchical representation by quantifying the relationship between features.
[0217] Specifically, when obtaining a hierarchical description of the vascular structure from the feature extraction results and determining the final output, the feature map can be divided into three types of regions.
[0218] For example, in a 512x512 feature map, the pixel values of the main blood vessel region are concentrated in 200 - 255, those of the secondary branch are 100 - 199, and those of the capillary are 0 - 99. This hierarchical description intuitively reflects the spatial hierarchical distribution of blood vessels.
[0219] For example, when extracting the optimized feature distribution from the final output, the feature map can be observed through a visualization tool.
[0220] For example, in a certain experiment, the optimized feature map shows that the edges of the main blood vessels are sharper and the details of the capillaries are richer. This precise hierarchical representation can provide a higher-quality data basis for subsequent analysis, helping to improve the depth and breadth of vascular network research.
[0221] It can be understood that the entire process from feature extraction to optimized output forms a closed-loop iterative mechanism. The adjustment of each link focuses on improving the accuracy of the vascular hierarchical representation. The finally generated feature distribution has significant advantages in both spatial resolution and class discrimination, laying a solid foundation for further research on the venous network.
[0222] Furthermore, the process of obtaining the final eighth network structure includes:
[0223] Obtaining the operation data of the seventh network structure on the embedded device, including computing resource occupancy, processing speed, and recognition accuracy;
[0224] Preprocessing the low-contrast palm vein image through an adaptive feature enhancement module, extracting the texture detail information in the image, and enhancing the contrast of the texture detail information using the histogram equalization algorithm to generate an enhanced palm vein feature image;
[0225] Judge whether the image features of the enhanced palm vein feature image meet the recognition accuracy requirements according to the preset threshold. If not, return to the adaptive feature enhancement module for parameter adjustment;
[0226] For the features that meet the accuracy requirements, use a convolutional neural network to extract deep features, construct a palm vein feature vector, and use a support vector machine classifier to classify and recognize the feature vector to obtain the palm vein identity recognition result;
[0227] Optimize and adjust the network structure according to the recognition result and running data to obtain the final eighth network structure.
[0228] Specifically, the adaptive feature enhancement module in this implementation dynamically adjusts the enhancement parameters by analyzing the gray distribution of the input image, effectively improving the details in the low-contrast area.
[0229] Exemplarily, for a palm vein image with low contrast, the module may first apply the adaptive histogram equalization algorithm to redistribute the image gray values to the full range of 0-255, highlighting the vein texture.
[0230] It should be noted that histogram equalization may introduce noise. Therefore, in a possible implementation, the system will combine noise reduction methods such as Gaussian filtering to keep the image smooth while enhancing the contrast.
[0231] Specifically, convolution can be performed using a 3x3 or 5x5 Gaussian kernel, and the σ value can be dynamically adjusted according to the image noise level, usually between 0.5-2. In the feature extraction stage, the design of the convolutional neural network is crucial.
[0232] In one embodiment, the network adopts a structure similar to VGG, including multiple convolutional layers, pooling layers, and fully connected layers.
[0233] Preferably, the first convolutional layer uses a larger convolutional kernel (such as 7x7) to capture large-scale features, and the convolutional kernel size gradually decreases to 3x3 in subsequent layers to extract finer texture information.
[0234] It can be understood that this progressive feature extraction helps to construct a multi-scale palm vein representation.
[0235] For example, for an input image of 300x300 pixels, the network may include 5 convolutional blocks, each block containing 2-3 convolutional layers and a max pooling layer. Finally, a fixed-length (such as 512-dimensional) feature vector is obtained through global average pooling. This design can not only ensure the discriminability of the features but also control the computational complexity, making it suitable for deployment on embedded devices. In the classification stage, the choice of the kernel function of the support vector machine (SVM) directly affects the recognition performance.
[0236] In one embodiment, a Radial Basis Function (RBF) kernel can be adopted, and its parameter γ is determined by cross-validation, usually in the range of 0.001 - 0.1. This non-linear kernel function can effectively handle complex decision boundaries in high-dimensional feature spaces and improve the recognition accuracy.
[0237] It should be further noted that the construction of the hardware platform in this embodiment includes:
[0238] The control platform selects STM32F407VET6 as the control platform, which is responsible for the control logic and data transmission of the system. This platform has the characteristics of high performance and low power consumption, and is suitable for the control of embedded devices.
[0239] Main control chip: Select Jetson Xavier NX as the main control chip to run the palm vein image recognition algorithm. Jetson Xavier NX has powerful computing capabilities and can efficiently run the lightweight YOLO model.
[0240] Camera module: Connect the camera module to collect palm vein images.
[0241] Specifically, the optimized lightweight YOLO model is deployed to the Jetson Xavier NX main control chip.
[0242] The model is run on the embedded device to perform real-time recognition of the collected palm vein images.
[0243] The recognition result is output to external devices, such as a display screen or a communication module, through the STM32F407VET6 control platform.
[0244] In this embodiment, through the lightweight design of the model, the computational complexity and storage requirements of the model are significantly reduced, enabling it to run efficiently on resource-constrained embedded devices. By using technical means such as multi-level feature extraction, attention mechanism, and specific loss functions, the recognition accuracy of palm vein images is effectively improved, and different-depth vein features can be accurately distinguished. The optimized model can quickly complete the recognition of palm vein images on the embedded device, meeting the real-time requirements and being applicable to actual application scenarios. Through techniques such as data augmentation and adaptive feature enhancement, the adaptability of the model to different lighting conditions, palm postures, and low-contrast images is improved, enhancing the robustness of the system.
[0245] Embodiment Two
[0246] As Figure 2 shown, based on the same inventive concept, this embodiment also provides an embedded palm vein image recognition system based on the lightweight YOLO model, including:
[0247] The network simplification module is used to obtain the network structure of the original YOLO model. By analyzing the contribution of each convolutional layer to palm vein recognition, redundant layers and the number of filters are determined. Standard convolutions are replaced with depthwise separable convolutions to obtain the first simplified network structure;
[0248] The pruning optimization module is used to calculate the importance of each connection for vein texture feature extraction based on the first network structure using pruning techniques. If the importance is lower than a preset threshold, the corresponding connection is removed to obtain the second pruned network structure;
[0249] The knowledge distillation module is used to extract the feature representation of palm vein images from a pre-established large teacher network. The features are migrated to the second network structure through the knowledge distillation method to adjust the parameters of the student network and obtain the third network structure with preliminary recognition ability;
[0250] The hybrid quantization module is used to analyze the sensitivity of each layer in the third network structure to low-contrast palm vein images. A mixed-precision quantization strategy is adopted to allocate 16-bit width representation to sensitive layers and 8-bit width representation to non-sensitive layers to obtain the fourth quantized network structure;
[0251] The multi-scale feature module is used to capture the features of major blood vessels, secondary branches, and capillary networks at different resolutions by designing a multi-level feature extraction module and adding a feature pyramid network to the fourth network structure to obtain the fifth network structure containing multi-scale information;
[0252] The attention enhancement module is used to introduce an anatomy-guided attention mechanism according to the vein hierarchical features in the fifth network structure. By calculating the correlation between each layer of feature maps and the vein depth, the attention of the model to veins at different depths is obtained to obtain the optimized sixth network structure;
[0253] The loss optimization module is used to design a loss function based on the output features of the sixth network structure. By measuring the hierarchical representation error of major blood vessels, secondary branches, and capillary features, the network weights are adjusted to obtain the seventh network structure that can accurately extract hierarchical features;
[0254] The adaptive enhancement module is used to obtain the operation data of the seventh network structure on the embedded device. The texture details of low-contrast palm vein images are amplified through the adaptive feature enhancement module, and it is judged whether the enhanced features meet the recognition accuracy requirements to obtain the final eighth network structure. Palm vein image recognition is performed based on the eighth network structure to obtain the image recognition result.
[0255] An embedded palm vein image recognition system based on the YOLO lightweight model provided in this embodiment has all the advantages of the embedded palm vein image recognition method based on the YOLO lightweight model provided in Embodiment 1.
[0256] Example 3
[0257] This embodiment also discloses a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method described in Embodiment 1.
[0258] Example 4
[0259] This embodiment also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method described in Embodiment 1.
[0260] Example 5
[0261] This embodiment also discloses a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the method described in Embodiment 1.
[0262] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An embedded palm vein image recognition method based on the YOLO lightweight model, characterized in that, Including: Obtain the network structure of the original YOLO model. By analyzing the contribution degree of each convolutional layer to palm vein recognition, determine the redundant layers and the number of filters. Replace the standard convolution with depthwise separable convolution to obtain the first streamlined network structure. Based on the first network structure, use pruning technology to calculate the importance of each connection for palm vein texture feature extraction. If the importance is lower than the preset threshold, remove the corresponding connection to obtain the second pruned network structure. Extract the feature representation of palm vein images from a pre-established teacher network. Transfer the features to the second network structure through knowledge distillation, and adjust the parameters of the student network to obtain the third network structure with preliminary recognition ability. Analyze the sensitivity of each layer in the third network structure to low-contrast palm vein images. Adopt a mixed-precision quantization strategy, allocate 16-bit width representation to sensitive layers and 8-bit width representation to non-sensitive layers to obtain the fourth quantized network structure. By designing a multi-level feature extraction module, add a feature pyramid network to the fourth network structure to capture the features of major blood vessels, secondary branches, and capillary networks at different resolutions, and obtain the fifth network structure containing multi-scale information. According to the vein hierarchical features in the fifth network structure, introduce an anatomy-guided attention mechanism. By calculating the correlation between each layer of feature maps and vein depth, obtain the attention of the model to veins at different depths, and obtain the optimized sixth network structure. Design a loss function according to the output features of the sixth network structure. By measuring the hierarchical representation error of major blood vessels, secondary branches, and capillary features, adjust the network weights to obtain the seventh network structure that can accurately extract hierarchical features. Obtain the operation data of the seventh network structure on the embedded device. Amplify the texture details of low-contrast palm vein images through an adaptive feature enhancement module, and judge whether the enhanced features meet the recognition accuracy requirements to obtain the final eighth network structure. Perform palm vein image recognition based on the eighth network structure to obtain the image recognition result.
2. The method according to claim 1, wherein The process of obtaining the first streamlined network structure includes: Obtain the network structure of the original YOLO model. Decompose each layer's parameters and output features through convolutional layer decomposition to determine the preliminary analysis result. By calculating the contribution degree of the output features of each convolutional layer to palm vein recognition, judge the redundant layers and unnecessary filters to obtain a redundant determination set. Use statistical methods to analyze the filter usage rate of the redundant determination set, determine the redundant layers and the number of filters to be removed, and obtain the streamlined layer configuration. Extract the remaining convolutional layers from the streamlined layer configuration, and replace the standard convolution with depthwise separable convolution to obtain the converted first network structure. Run the palm vein dataset through the first network structure, obtain the recognition rate and computational complexity, and judge whether they meet the preset threshold to obtain the performance evaluation result. If the performance evaluation result is lower than the preset threshold, adjust the parameters of the depthwise separable convolution to obtain an optimized network structure; otherwise, directly use it as the first network structure.
3. The method according to claim 1, wherein The process of obtaining the second network structure after pruning includes: Obtain the weight data of each connection in the first network structure through pruning technology, calculate the contribution degree of each connection to the vein texture using a convolutional neural network, and obtain the importance value; Compare the importance value with a preset threshold. If the importance value is lower than the preset threshold, mark the corresponding connection as a to-be-removed state to obtain a preliminary list of removed connections; Execute a connection deletion operation on the first network structure using a matrix operation tool through the preliminary list of removed connections to obtain a temporarily pruned network structure; For the temporarily pruned network structure, extract the feature data of the vein texture through texture analysis technology, and judge the integrity of feature extraction. If the integrity meets the preset conditions, extract connection evaluation data from the temporarily pruned network structure; Classify the effectiveness of the connections using a support vector machine according to the connection evaluation data to obtain an optimized connection distribution; Execute a final adjustment on the temporarily pruned network structure through the optimized connection distribution to obtain the second network structure after pruning.
4. The method according to claim 1, wherein The process of obtaining the third network structure with preliminary recognition ability includes: Extract the feature representation of the palm vein image through a pre-established teacher network to obtain an initial feature set; Adopt a knowledge distillation method to transfer features from the initial feature set to the second network structure to obtain a transferred feature set; Adjust the parameters of the second network structure according to the transferred feature set to obtain optimized network parameters; Train the second network structure through the optimized network parameters to obtain the second network structure with preliminary recognition ability; Obtain the output result of the second network structure. If the recognition ability is lower than the preset threshold, update and adjust the parameters through a feedback mechanism to obtain the updated third network structure with preliminary recognition ability.
5. The method according to claim 1, wherein The process of obtaining the fourth network structure after quantization includes: Analyze the response of each layer in the third network to the low-contrast palm vein image through a convolutional neural network to obtain a sensitivity distribution; Judge the sensitive layer and the non-sensitive layer according to the sensitivity distribution to determine the layer allocation result; According to the layer allocation result, adopt a mixed-precision quantization strategy to allocate a 16-bit width representation to the sensitive layer to obtain the first quantization parameter; allocate an 8-bit width representation to the non-sensitive layer to obtain the second quantization parameter; Fuse and generate the fourth network structure according to the first quantization parameter and the second quantization parameter; If there is a precision loss in the fourth network structure, adjust the quantization parameters through a gradient descent algorithm to obtain the fourth network structure after quantization.
6. The method according to claim 1, wherein The process of obtaining the fifth network structure containing multi-scale information includes: Extract features from the palm vein image data through a multi-level module to obtain a preliminary feature set; Process the preliminary feature set using a feature pyramid network to capture features of different resolutions and generate multi-scale feature maps; Separate the features of the main blood vessels from the multi-scale feature maps to obtain a first blood vessel feature subset; Analyze secondary branches based on the first subset of vascular features to obtain a second subset of vascular features; Detect a capillary network through the second subset of vascular features to determine a third subset of vascular features; Obtain the fusion information of the third subset of vascular features and the multi-scale feature map, and judge the output features of the fifth network structure; Optimize the network structure according to the output features of the fifth network structure to obtain the final multi-scale vascular network representation.
7. The method according to claim 1, wherein The process of obtaining the optimized sixth network structure includes: Obtain hierarchical feature data from the venous network structure image and extract venous depth information; Use a convolutional neural network to extract features from the venous network structure image to obtain a multi-layer feature map; Construct a depth correlation matrix by calculating the correlation coefficient between each layer of the feature map and the venous depth information; If the correlation coefficient is greater than a preset threshold, the weight of the corresponding layer of the feature map is increased; If the correlation coefficient is less than a preset threshold, the weight of the corresponding layer of the feature map is decreased; Design an attention mechanism module according to the adjusted feature map weights to generate an attention map; Perform weighted fusion of the attention map and the original feature map to obtain an enhanced feature representation; Based on the enhanced feature representation, use a fully connected layer for venous depth classification and output the optimized sixth network structure.
8. The method according to claim 1, wherein The process of obtaining the seventh network structure for accurately extracting hierarchical features includes: Obtain output features through the sixth network, and use convolution operations to extract the spatial distributions of major blood vessels, secondary branches, and capillaries to obtain an initial hierarchical representation; For the initial hierarchical representation, construct a loss function, calculate the representation errors of major blood vessels, secondary branches, and capillaries, and determine the error distribution; If the error distribution exceeds a preset threshold, adjust the network weights through backpropagation to obtain updated weight parameters; Optimize the sixth network structure using the updated weight parameters to generate a seventh network and extract accurate hierarchical representations; Based on the hierarchical representation of the seventh network, judge the feature consistency of major blood vessels, secondary branches, and capillaries to obtain a feature extraction result; According to the feature extraction result, obtain a hierarchical description of the vascular structure and determine the final output of the seventh network; Extract the optimized feature distribution from the final output of the seventh network to obtain an accurate hierarchical representation.
9. The method according to claim 1, wherein The process of obtaining the final eighth network structure includes: Obtain the operation data of the seventh network structure on the embedded device, including computing resource occupancy, processing speed, and recognition accuracy; Preprocess the low-contrast palm vein image through an adaptive feature enhancement module, extract the texture detail information in the image, and use the histogram equalization algorithm to enhance the contrast of the texture detail information to generate an enhanced palm vein feature image; Judge whether the image features of the enhanced palm vein feature image meet the recognition accuracy requirements according to a preset threshold. If not, return to the adaptive feature enhancement module for parameter adjustment; For features that meet the accuracy requirements, deep features are extracted using a convolutional neural network, a palm vein feature vector is constructed, and a support vector machine classifier is used to classify and identify the feature vector to obtain a palm vein identity recognition result; According to the recognition result and operation data, the network structure is optimized and adjusted to obtain the final eighth network structure.
10. An embedded palm vein image recognition system based on the YOLO lightweight model, characterized in that, It includes: A network simplification module, which is used to obtain the network structure of the original YOLO model, determine the redundant layers and the number of filters by analyzing the contribution of each convolutional layer to palm vein recognition, and replace the standard convolution with depthwise separable convolution to obtain the first simplified network structure; A pruning optimization module, which is used to calculate the importance of each connection for vein texture feature extraction based on the first network structure using pruning technology. If the importance is lower than a preset threshold, the corresponding connection is removed to obtain the second pruned network structure; A knowledge distillation module, which is used to extract the feature representation of palm vein images from a pre-established large teacher network, transfer the features to the second network structure through the knowledge distillation method, and adjust the parameters of the student network to obtain the third network structure with preliminary recognition ability; A mixed quantization module, which is used to analyze the sensitivity of each layer in the third network structure to low-contrast palm vein images, adopt a mixed-precision quantization strategy, allocate 16-bit width representation to sensitive layers and 8-bit width representation to non-sensitive layers to obtain the fourth quantized network structure; A multi-scale feature module, which is used to capture the features of main blood vessels, secondary branches and capillary networks at different resolutions by designing a multi-level feature extraction module and adding a feature pyramid network to the fourth network structure to obtain the fifth network structure containing multi-scale information; An attention enhancement module, which is used to introduce an anatomy-guided attention mechanism according to the vein hierarchical features in the fifth network structure, and obtain the attention of the model to veins at different depths by calculating the correlation between each layer of feature maps and vein depth to obtain the optimized sixth network structure; A loss optimization module, which is used to design a loss function according to the output features of the sixth network structure, adjust the network weights by measuring the hierarchical representation error of main blood vessels, secondary branches and capillary features to obtain the seventh network structure that accurately extracts hierarchical features; An adaptive enhancement module, which is used to obtain the operation data of the seventh network structure on an embedded device, magnify the texture details of low-contrast palm vein images through an adaptive feature enhancement module, and judge whether the enhanced features meet the recognition accuracy requirements to obtain the final eighth network structure; Based on the eighth network structure, palm vein image recognition is performed to obtain an image recognition result.
Citation Information
Patent Citations
Vein authentication method based on network pruning, medium and equipment
CN118230369A
Palm vein authentication method combining CNN (Convolutional Neural Network) and Transformer
CN118658185A
Method and system for evaluating consistency of multiple / time blood vessel segmentation results
CN119107327A