Palm vein recognition method based on domain adaptation strategy

By using a palm vein recognition method based on domain adaptation strategy, an adversarial domain adaptation network is constructed, which solves the problem of performance degradation of palm recognition technology under data set mismatch, realizes stable recognition and low-cost biometric recognition, adapts to different acquisition devices and environments, and reduces computational complexity and privacy risks.

CN120656210AActive Publication Date: 2025-09-16NINGBO JINGXINHAO NETWORK TECH CO LTD

Patent Information

Application Number
CN202510798800.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-16
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Existing deep learning palm recognition technology suffers from degraded recognition performance when the source and target datasets do not match. The dataset differences lead to model overfitting and poor generalization, high data collection and labeling costs, and difficulty in sharing private data.

Method used

A palm vein recognition method based on domain adaptation strategy is adopted to construct a dataset and design an adversarial domain adaptation network, including a palm feature extraction module, a classification module and a domain discrimination module. Through adversarial training and iterative learning, domain-invariant features are extracted, reducing data dependence and computational complexity.

Benefits of technology

Maintain stable recognition performance in different application environments, reduce manual labeling and computing costs, enhance model generalization capabilities, reduce privacy leakage risks, adapt to small amounts of labeled data, and meet real-time application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656210A_ABST
    Figure CN120656210A_ABST
Patent Text Reader

Abstract

The invention discloses a palm vein recognition method based on a domain adaptation strategy, and relates to the technical field of biological information recognition, identity authentication and palm recognition, and the method comprises the steps: 1, constructing a data set which comprises a source data set and a target data set, and employing all data in the source data set and part of data in the target data set as an initial training set, using other data in the target data set as an initial test set, then performing ROI extraction, segmenting the initial training set into a formal training set, and segmenting the initial test set into a formal test set; 2, designing an adversarial domain adaptation network; 3, pre-training is carried out on the source data set, so that the classification module can reach preset classification performance on the source data set; 4, adversarial training is carried out on the formal training set, so that the palm feature extraction module learns domain-invariant features, and when the training reaches a preset iteration frequency or a model threshold value, the training is ended; and 5, testing by using the formal test set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of biometric information recognition, identity authentication and palm recognition, and in particular to a palm vein recognition method based on a domain adaptation strategy. Background Art

[0002] Biometric recognition is a technology that uses the body's inherent physiological characteristics to identify individuals. Currently, in-depth research is underway both domestically and internationally on biometric recognition technologies such as fingerprints, faces, palm prints, irises, and signatures. Due to its high security, anti-counterfeiting, and convenience, palm recognition technology is widely used in various fields, such as finance, security, healthcare, and access control systems, bringing greater safety and convenience to people's lives and work.

[0003] Existing deep learning palm recognition technologies have achieved some progress in their respective applications. However, these methods primarily focus on single datasets, requiring that the training and test sets be collected under identical conditions. However, in real-world applications, datasets are often collected using different devices. Due to dataset disparity, a model trained on one dataset will experience significant performance degradation when tested on a different dataset.

[0004] When faced with a data mismatch between the source and target datasets, retraining the target dataset can adapt the model to the target dataset's data distribution, but it faces numerous drawbacks. Firstly, collecting and annotating the target dataset requires significant human, material, and time resources. This is especially true for biometric identification, such as palm recognition, where acquiring high-quality palm images and then accurately annotating them is a tedious and time-consuming process. Secondly, if the target dataset is limited, retraining can lead to model overfitting and poor generalization.

[0005] Therefore, those skilled in the art are committed to developing a new palm vein recognition method to solve the above-mentioned defects in the prior art. Summary of the Invention

[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to avoid degradation of recognition performance when the source dataset data and the target dataset data of the palm print / palm vein data do not match.

[0007] To achieve the above object, the present invention provides a palm vein recognition method based on a domain adaptation strategy, characterized in that the method comprises the following steps: Step 1: Construct a dataset, including a source dataset and a target dataset. Use all the data in the source dataset and part of the data in the target dataset as the initial training set, and use the remaining data in the target dataset as the initial test set. Then perform ROI extraction to obtain image data of the palm ROI area. Split the initial training set into a formal training set, and split the initial test set into a formal test set. Step 2: Design an adversarial domain adaptation network, which includes a palm feature extraction module, a classification module, and a domain discrimination module; Step 3: Pre-training is performed on the source dataset. The palm feature extraction module extracts palm features from the source dataset and inputs the palm features into the classification module for classification prediction. Based on the classification loss, the back propagation algorithm is used to update the parameters of the palm feature extraction module and the classification module. Through continuous iterative training, the classification module is able to achieve a preset classification performance on the source dataset. Step 4: Conduct adversarial training on the formal training set, input the formal training set into the pre-trained palm feature extraction module to obtain the corresponding feature representation, then input the obtained feature representation into the domain discrimination module to determine whether the feature representation comes from the source dataset or the target dataset, and calculate the adversarial loss; at the same time, input the obtained feature representation into the classification module, calculate the classification loss based on the classification prediction result in the classification module; calculate the joint loss based on the adversarial loss and the classification loss, and then use the back propagation algorithm to update the parameters of the palm feature extraction module; through continuous iterative training, the palm feature extraction module gradually learns the domain-invariant features, and the training ends when the preset number of iterations or model threshold is reached; Step 5: During the testing phase, the trained palm feature extraction module and the classification module are tested using the formal test set to obtain the palm ID and complete the evaluation of the model performance.

[0008] Furthermore, the step 1 includes the following sub-steps: Step 1.1: Use two or more different palm datasets as the source dataset and the target dataset, respectively. The source dataset is historical data collected by an existing device, and the target dataset is data collected by a new palm collection device. Use all the data in the source dataset and part of the data in the target dataset as the initial training set, and use the remaining data in the target dataset as the initial test set. Step 1.2: performing data preprocessing on the source dataset and the target dataset to unify the data formats. The data preprocessing methods include but are not limited to normalization, enhancement, or resizing. Step 1.3: Add palm segmentation and key point detection tasks to the target detection model, perform palm location and segmentation through the ROI extraction, determine the positions of key feature points in the palm, split the initial training set into the formal training set, and split the initial test set into the formal test set.

[0009] Furthermore, the target detection model in step 1.3 is YOLO series, Faster R-CNN or SSD.

[0010] Furthermore, the key feature points in the palm in step 1.3 include 7, namely the base point of the index finger, the base point of the middle finger, the base point of the thumb, the base point of the little finger, the center point of the palm, and the two key points on the left and right where the wrist and palm are connected.

[0011] Furthermore, the palm feature extraction module in step 2 includes a depthwise separable convolution module, a multi-level feature extraction module and a sub-graph Transformer module, wherein the depthwise separable convolution module performs preliminary feature extraction, the multi-level feature extraction module stacks multiple improved inverse residual structures and linear Transformer encoders, and the sub-graph Transformer module divides the image of the palm ROI area into multiple sub-graphs of equal size, and uses the Transformer model to mine and output a global feature vector of fixed dimension.

[0012] Furthermore, the improved inverse residual structure in the multi-level feature extraction module first expands and groups the number of channels through grouped point-by-point convolution, performs convolution operations within each group, and then performs 3×3 depth convolution on the expanded channel dimension to extract the local spatial features of the channel. At the same time, the channel attention module is combined to perform global average pooling on the features of each channel to obtain the global information of the channel, and then generate a weight coefficient for each channel. Finally, the number of channels is compressed back to the dimension before expansion and grouping through the grouped point-by-point convolution to achieve feature fusion and linear transformation.

[0013] Furthermore, the multi-level feature extraction module replaces the ReLU activation function with a Mish function or a Swish function, including but not limited to a Mish function or a Swish function, wherein the Mish function has a self-gating characteristic, and its expression is Mish(x)=x*tanh(softplus(x)); the expression of the Swish function is Swish(x)=x*sigmoid(x).

[0014] Furthermore, the improved inverse residual structure of the multi-level feature extraction module extracts feature X, and divides the feature map of feature X into multiple overlapping small patches of equal size. Each patch is regarded as an element in a sequence. The linear Transformer encoder is then used to calculate the attention weight between each patch and all other patches, dynamically determining the importance of each patch in the global feature, thereby achieving abstraction and fusion of the global features. The feature map of feature X is then expanded into P non-overlapping flattened patches, each of the flattened patches is indexed and numbered, and for each flattened patch, the linear Transformer encoder is applied to encode the relationship between patches to obtain the global features between the patches.

[0015] Furthermore, the classification module in step 2 includes one or more fully connected layers, and finally a Softmax layer for classification, and the classification loss function is the cross entropy loss function , where the number of neurons in the fully connected layer can be adjusted according to the classification task.

[0016] Furthermore, the domain discrimination module in step 2 includes several convolutional layers, fully connected layers and a Softmax function, and a batch normalization layer and an activation function are added between the fully connected layers; the domain discrimination module uses a binary classification method to distinguish whether the palm features are from the source dataset and the target dataset, and then performs feature confusion by gradient reversal to align the edge distribution between the datasets, wherein the adversarial loss function : in, Representative palm images from the source dataset, represents the palm image from the target dataset, S is the number of palm images in the source dataset, and T is the number of palm images in the target dataset. It represents the probability that the domain discrimination module predicts that the palm image comes from the target dataset.

[0017] The palm vein recognition method based on domain adaptation strategy provided by the present invention has the following technical effects: 1. The technical solution provided by this invention can accurately capture the potential connections and differences between different domains through effective learning from a small amount of labeled data, enabling the model to maintain stable and good recognition performance in different application environments and scenarios, while significantly reducing the cost and time of manual labeling. Even when faced with new, unseen target domain data, it can quickly adapt and make accurate recognition judgments, effectively avoiding overfitting and enhancing the model's generalization ability. 2. The lightweight design architecture of the technical solution provided by this invention significantly reduces the computational complexity of the model and reduces reliance on hardware resources, meeting the needs of real-time application scenarios while also reducing device energy consumption and extending device battery life. In sensitive scenarios such as healthcare and finance, protecting data privacy is crucial. The technical solution provided by this invention reduces reliance on large amounts of data, lowering the risk of data leakage and the privacy management costs associated with data collection and storage, providing users with a more secure and reliable palm recognition solution.

[0018] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flow chart of a palm vein recognition method according to a preferred embodiment of the present invention; Figure 2 This is a schematic diagram of palm key points extracted by ROI according to a preferred embodiment of the present invention; Figure 3 This is a schematic diagram of the palm feature extraction module workflow of a preferred embodiment of the present invention; Figure 4 Schematic diagram of depth separation convolution of a preferred embodiment of the present invention; Figure 5 1 is a schematic diagram of the inverse residual structure workflow of a preferred embodiment of the present invention; Figure 6 It is a schematic diagram of the improved inverse residual structure workflow of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0020] The following describes several preferred embodiments of the present invention with reference to the accompanying drawings to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0021] The palm recognition involved in the embodiments of the present invention includes but is not limited to palm print or palm vein recognition, or palm print and palm vein fusion recognition, or palm back and palm back vein recognition, etc., which use palm information for biometric recognition. Deep learning domain adaptation technology is used to solve the problem of reduced recognition performance caused by mismatch between source dataset data and target dataset data of palm print / palm vein data.

[0022] However, existing technologies do not fully consider the following issues: 1) Due to different devices, different usage scenarios (different background information), different near-infrared light sources, different lighting conditions, etc., the collected palm images have differences in contrast, clarity, resolution, etc., resulting in a mismatch between the source dataset data during training and the target dataset data during actual use. The trained model has difficulty coping with these uncontrollable changes in the acquisition environment in actual scenarios, resulting in degraded model performance.

[0023] a) Data differences caused by different collection devices In practical applications of palm vein recognition, the image data for the source dataset (e.g., palm images collected in the training set) and the target dataset (e.g., palm images collected in actual use scenarios) may be acquired using different acquisition devices. Different devices vary in resolution, color reproduction, and noise levels, leading to significant differences in the distribution of features such as pixel values ​​and texture in palm images. For example, palm images captured by a high-quality camera are clear and color-accurate, while images captured by an actual palm vein recognition module may be blurry, exhibit color cast, and exhibit noise due to exposure and ambient light.

[0024] b) Data differences caused by different environmental conditions The source dataset's hand images may have been collected indoors under uniform lighting conditions and against simple backgrounds, while the target dataset's images may have been collected outdoors under complex lighting conditions (such as backlighting or sidelighting) and against complex backgrounds (such as crowds or buildings). This environmental difference can alter the appearance of hand features (such as shadows and contrast), leading to a shift in data distribution. For example, in a backlit environment, certain areas of the palm may be in shadow, resulting in a significant difference in feature distribution compared to a palm image taken indoors under uniform lighting.

[0025] c) Scarcity of labeled data The source dataset may have relatively rich annotated data (for example, some public palm datasets have been identity-annotated). Through domain adaptation technology, we can utilize the annotated data and models of the source dataset to achieve better palm recognition performance when the target dataset has a small amount of annotated data or even no annotated data, thereby reducing the dependence on the large amount of annotated data in the target dataset.

[0026] 2) Data sharing between different institutions or organizations faces legal and ethical challenges due to privacy constraints and data security risks. Palm print and vein data are sensitive personal information, and leaking them can pose serious privacy and security risks. Therefore, when training palm recognition models, data security must be ensured to prevent theft or misuse.

[0027] To address the aforementioned issues, which are not adequately addressed in existing technologies, the technical solutions provided by the present invention primarily involve: combining multiple datasets to construct domain adaptation training and test sets; designing an adversarial domain adaptation network, primarily comprising three deep neural network modules: a palm feature extraction module, a domain discrimination module, and a classification module. The palm feature extraction module extracts multi-dimensional, high-dimensional features from palm images, such as texture, vein distribution, and geometric shape. During model training, the domain discrimination module performs adversarial training against the palm feature extractor in the palm feature extraction module, making it difficult for the domain discrimination module to distinguish whether the input palm image data is from the source dataset or the target dataset. The classification module performs identity recognition and classification on the palm features output by the palm feature extractor. The source and target domain datasets are input into the palm feature extraction module to extract source and target domain palm features. The domain discrimination module discriminates the source and target domain features, and a classifier is used to match the palm features for palm recognition. During forward propagation, the classification loss and domain adversarial loss are calculated. Based on the calculated loss values, the gradient of the loss with respect to the network parameters is calculated using a backpropagation algorithm. During backpropagation, the domain adversarial loss guides the feature extraction module to update parameters toward extracting domain-invariant features, while the classification loss encourages the classification module to accurately classify the data. Through continuous iterative updates, the network gradually adapts to the data distribution of the target domain while maintaining classification performance in the source domain.

[0028] Example 1 like Figure 1 As shown, an embodiment of the present invention provides a palm vein recognition method based on a domain adaptation strategy, comprising the following steps: Step 1: Construct a dataset, including a source dataset and a target dataset. Use all the data in the source dataset and part of the data in the target dataset as the initial training set, and use the remaining data in the target dataset as the initial test set. Then perform ROI extraction to obtain image data of the palm ROI area. Split the initial training set into the formal training set, and split the initial test set into the formal test set. Step 2: Design an adversarial domain adaptation network, which includes a palm feature extraction module, a classification module, and a domain discrimination module; Step 3: Pre-training is performed on the source dataset. The palm feature extraction module extracts palm features from the source dataset and inputs the palm features into the classification module for classification prediction. Based on the classification loss, the back-propagation algorithm is used to update the parameters of the palm feature extraction module and the classification module. Through continuous iterative training, the classification module is able to achieve the pre-set classification performance on the source dataset. Step 4: Conduct adversarial training on the formal training set. Input the formal training set into the pre-trained palm feature extraction module to obtain the corresponding feature representation. Then input the obtained feature representation into the domain discrimination module to determine whether the feature representation comes from the source dataset or the target dataset, and calculate the adversarial loss. At the same time, input the obtained feature representation into the classification module. Calculate the classification loss based on the classification prediction results in the classification module. Calculate the joint loss based on the adversarial loss and the classification loss, and then use the backpropagation algorithm to update the parameters of the palm feature extraction module. Through continuous iterative training, the palm feature extraction module gradually learns domain-invariant features. When the training reaches the preset number of iterations or model threshold, the training ends. Step 5: During the testing phase, the trained palm feature extraction module and classification module are tested using the formal test set to obtain the palm ID and complete the evaluation of the model performance.

[0029] Example 2 Based on Example 1, step 1 includes the following sub-steps: Step 1.1: Use two or more different hand datasets as the source and target datasets, respectively. The source dataset is historical data collected by existing equipment (different acquisition devices and different environments) and contains rich hand annotation information. The target dataset is data collected by newly developed hand acquisition equipment and contains a small amount of annotated data. Use all the data (all images) in the source dataset (source domain) and a portion of the data (a small number of images, one image per category) in the target dataset (target domain) as the initial training set. Use the remaining data in the target dataset as the initial test set. This data is not available during training and is only used to evaluate the model during testing.

[0030] Step 1.2: Preprocess the hand images in the source and target datasets to unify the data format, i.e., adjust the images to a size and format suitable for network input. Data preprocessing methods include, but are not limited to, normalization, enhancement, or resizing. Step 1.3: Add palm segmentation and key point detection tasks to the object detection model, perform palm location and segmentation through ROI extraction, determine the positions of key feature points in the palm, split the initial training set into the formal training set, and split the initial test set into the formal test set.

[0031] In particular, the target detection model in step 1.3 is the YOLO (You Only Look Once) series, Faster R-CNN or SSD, etc.

[0032] In particular, the key feature points in the palm in step 1.3 include 7, namely the base point of the index finger, the base point of the middle finger, the base point of the thumb, the base point of the little finger, the center point of the palm, and the two key points on the left and right where the wrist and palm connect, such as Figure 2 shown.

[0033] Example 3 Based on embodiment 1 or 2, the adversarial domain adaptation network in step 2 includes a palm feature extraction module, a classification module and a domain discrimination module, such as Figure 3 shown.

[0034] The palm feature extraction module is a lightweight network-based module that extracts global and local features from palm images. It includes a depthwise separable convolution module, a multi-level feature extraction module, and a subgraph transformer module. The depthwise separable convolution module performs preliminary feature extraction. The multi-level feature extraction module stacks multiple improved inverted residual blocks (IRBs) and linear transformer encoders. The subgraph transformer module divides the palm ROI image into multiple equally sized subgraphs and uses the transformer model to mine and output a fixed-dimensional global feature vector.

[0035] like Figure 4 As shown in the figure, in the depthwise separable convolution module, the input image first passes through a standard convolution layer and a depthwise convolution layer to perform preliminary feature extraction on the image, such as extracting simple low-level features such as edges and textures. Conventional convolution kernel sizes and stride sizes are typically used, such as a 3×3 kernel with a stride of 2. This reduces computational effort while preserving essential image features.

[0036] In the multi-level feature extraction module, the specific number of stacked layers of multiple improved inverse residual structures and Linear Transformer encoders was determined through ablation experiments and hardware performance. The improved inverse residual structure first expands the number of channels by replacing the original point-by-point convolution with grouped point-by-point convolution. Grouped point-by-point convolution groups the input channels and performs convolution within each group. Compared to point-by-point convolution, which operates on all channels simultaneously, this significantly reduces the number of multiplications and additions. A 3×3 depthwise convolution is then performed on the expanded channel dimension. Depthwise convolution primarily performs convolution operations on each channel independently to extract richer local spatial features. Furthermore, a channel attention module, such as the Squeeze-and-Excitation (SE) module, is combined to perform global average pooling on the features of each channel to obtain global information. A weight coefficient is then generated for each channel to emphasize the features of important channels and suppress information from unimportant channels, thereby improving the model's sensitivity to different features. Finally, grouped point-by-point convolution compresses the number of channels back to their original dimensions, achieving both feature fusion and linear transformation. This structure effectively balances computational effort and feature extraction capabilities while maintaining a lightweight model. Dimensionality increases enable the network to learn more complex features in high-dimensional space, while deep convolution extracts spatial features without significantly increasing computational effort. Dimensionality reduction reduces the number of parameters and computational effort, preventing overfitting from excessive model complexity while also making features more compact and representative. Specifically, the multi-level feature extraction module replaces the ReLU activation function with activation functions such as Mish or Swish. The Mish function, expressed as Mish(x) = x*tanh(softplus(x)), exhibits self-gating properties and can adaptively adjust the output under varying input conditions, improving the model's nonlinear expressiveness while maintaining relatively low computational complexity. The Swish function, expressed as Swish(x) = x*sigmoid(x), also exhibits a degree of adaptability, enabling the model to more effectively learn features during training, particularly when processing complex data.

[0037] In particular, Figure 5 and Figure 6 As shown in the multi-level feature extraction module, CNN and the improved inverted residual block structure (Inverted Residual Blocks) extract features at a certain level. , divide these feature maps into multiple overlapping small patches of equal size (overlap rate between 0% and 50%), and regard each patch as an element in a sequence. Use the linear Transformer encoder to calculate the attention weight between each patch and all other patches, and dynamically determine the importance of each patch in the global feature, so as to achieve effective abstraction and fusion of global features. Expand the input feature map into P non-overlapping flattened patches (Patch). Each patch is clearly indexed and numbered, and the patches are arranged in order from left to right and from top to bottom. In the subsequent processing process, the order of the patches can be determined according to the index to ensure that it corresponds to the spatial position in the original feature map. For each patch, apply the LinearTransformer encoder to encode the relationship between patches to obtain the global features between patches.

[0038] Specifically, the separable self-attention mechanism is used to determine the importance weight of each patch relative to other patches. The separable self-attention mechanism is an improvement to the traditional self-attention mechanism, aiming to reduce computational complexity and improve efficiency in resource-constrained environments.

[0039] The separable self-attention mechanism consists of three branches: input branch (I), key branch (K), and value branch (v).

[0040] Latent Token Generation: First, the input features pass through an embedding layer, converting the input data into a vector representation, resulting in a series of input tokens. These input tokens are then fed into a latent token generation module, which can be a simple linear transformation layer or a more complex neural network structure. Its purpose is to generate a set of latent tokens, which are abstract representations of the original input tokens.

[0041] Context score calculation: Input features pass through input branch I and pass through the linear layer Map the token of each dimension into a scalar to obtain a dimension vector, and then perform a softmax operation to obtain the context score This score reflects the correlation between each input token and the potential token. The higher the score, the stronger the correlation between the token and the potential token.

[0042] Context vector calculation: Input features pass through the key branch K (weight is ) Linear projection gives ,for Each token in is weighted according to its corresponding context score, and then all weighted tokens are summed to obtain a context vector that integrates all input token information. The calculation formula is: .

[0043] Information propagation and output: Input features pass through the value branch (weight is ) is linearly projected into d-dimensional space and then activated by ReLU to obtain .

[0044] Context vector Element-wise multiplication is performed by broadcasting To merge, and Each element of is multiplied accordingly, so that the context information is propagated to each element, achieving The result of the broadcast element multiplication is then passed through the linear layer Perform linear transformation to obtain the final output, completing the re-weighting and information integration of the input token.

[0045] When applying Transformer to each patch for encoding, position encoding information is introduced. Position encoding incorporates the position information of the patch in the original feature map into the model, allowing the model to perceive the relative position of each patch.

[0046] The features encoded by Linear Transformer are folded back to the original spatial dimension, so that the features are restored to a spatial structure similar to the original input feature map. This folding process is the inverse of the previous patching and flattening operations. The features are accurately restored to their original spatial positions based on the patch index and the position information of the pixels within the patch, thus ensuring that the order of the patches and the spatial order of the pixels within the patches are preserved.

[0047] The features and original features , merged through splicing operations. The concatenated features are fused using a convolutional layer. The convolutional layer can perform convolution operations by sliding the convolution kernel across the feature map, further extracting higher-level features from the fused features, achieving deeper feature fusion and making the final features more representative and discriminative.

[0048] In particular, the subgraph Transformer module divides the original image into multiple subgraphs and, leveraging the powerful feature extraction capabilities of the Transformer model, deeply mines the local and global features of the image, thereby achieving a precise understanding and analysis of the image content. The subgraph Transformer module is used only once in the palm feature extraction module. While extracting features, it is able to maintain the original structural information of the palm image as much as possible and automatically learn the long-range dependencies between image subgraphs, which facilitates subsequent feature-based analysis and processing. If the entire image is directly used as input, the high resolution and high dimensionality of the image will result in a huge amount of computation, making the model training and inference process very complex and time-consuming. By dividing the image into small blocks, the dimensionality of each block can be reduced, reducing the amount of data that the model needs to process, thereby reducing computational complexity and improving the model's operating efficiency.

[0049] The working process of the subgraph Transformer module is as follows Figure 3 As shown, specifically: 1) Split the palm ROI image into multiple equally sized sub-images. The size and number of sub-images can be adjusted based on actual needs. Decomposing the palm image into multiple local information helps the Transformer model better capture the local details of the image.

[0050] 2) The segmented sub-images are sequentially fed into the network. Each sub-image undergoes a linear projection before entering the Transformer model. This step converts the pixel values ​​of each sub-image into a fixed-dimensional vector representation through a linear transformation. This converts the spatial information of the image into sequence information suitable for processing by the Transformer model, while also reducing the data dimension and improving the model's computational efficiency.

[0051] 3) Similarly, the Separable Self-Attention mechanism is used as the Transformer structure for feature extraction.

[0052] 4) The Transformer model outputs a fixed-dimensional global feature vector. This global feature vector integrates information from all parts of the image and can reflect the content and characteristics of the entire image.

[0053] The convolutional layer is used to fuse the features extracted by the subgraph transformer module with the first-level features of the multi-level feature extraction module. The features are then fed into the multi-level feature extraction module for further feature extraction.

[0054] Global average pooling is performed on the features extracted by the feature extraction module, compressing the feature map of each channel into a scalar, resulting in a fixed-length feature vector as the final palm feature representation. By stacking multi-level feature extraction modules and combining local and global features, the model can more comprehensively capture image information and improve its understanding of images. Compared to simple CNN modules or self-attention modules, this can simultaneously utilize local details and global semantic associations, resulting in a feature representation that combines rich local and global features.

[0055] Example 4 Based on Example 1, 2 or 3, the classification module in step 2 includes one or more fully connected layers, and finally a Softmax layer for classification. The number of neurons in the fully connected layer can be adjusted according to the specific classification task.

[0056] The classification loss function is the cross entropy loss function , which is used to measure the difference between the output of the classification module and the true classification label.

[0057] In particular, the domain discrimination module in step 2 consists of several convolutional layers, fully connected layers, and a Softmax function. To enhance the discrimination capability, batch normalization layers and activation functions (such as ReLU, ReLU6, etc.) are added between the fully connected layers.

[0058] The domain discrimination module uses a binary classification method to distinguish whether the palm features come from the source dataset and the target dataset, and then performs feature confusion by gradient reversal to align the edge distribution between the datasets. The adversarial loss function : in, Representative palm images come from the source dataset, represents the palm image from the target dataset, S is the number of palm images in the source dataset, T is the number of palm images in the target dataset, represents the probability that the domain discrimination module predicts that the palm image comes from the source dataset, It represents the probability that the domain discrimination module predicts that the palm image comes from the target dataset.

[0059] Adversarial training domain discrimination loss with feature extractor during model training , which ultimately makes it impossible for the domain discrimination module to distinguish whether the data comes from the source domain or the target domain. Through this adversarial approach, the features generated by the palm feature extraction module are made more universal, reducing the differences between domains.

[0060] During the training process, the classification loss and domain adversarial loss are weightedly summed to obtain the final loss function.

[0061] In particular, pre-training in step 3 involves learning key palm features and classification capabilities, including parameter initialization and source-domain data training. Initialization specifically initializes the parameters of the feature extractor, classifier, and discriminator. Using a random initialization method gives the model an initial random state, allowing it to autonomously explore optimal parameters during training. Source-domain data training involves inputting source-domain data into the feature extractor to obtain feature representations, which are then fed into the classifier for classification prediction. Based on the classification loss, the backpropagation algorithm is used to update the parameters of the feature extractor and classifier, ensuring that the classifier achieves good classification performance on the source-domain data. When updating parameters, it is important to select an appropriate optimizer, such as stochastic gradient descent (SGD) or Adam. The choice of optimizer affects the speed and effectiveness of parameter updates. Through continuous iterative training, the classifier achieves good classification performance on the source-domain data, with a gradual decrease in classification loss and improvement in classification accuracy.

[0062] In particular, in the adversarial training of step 4, on the basis of pre-training, the target domain data is introduced, and the feature extraction submodule, domain discrimination module and classification submodule are trained simultaneously.

[0063] The source and target domain data are fed into the pre-trained feature extraction module to obtain corresponding feature representations. These features are then fed into the domain discrimination module, which determines whether the features originate from the source or target domain. Based on the output of the domain discrimination module, an adversarial loss is calculated. The parameters of the feature extraction module and the classification submodule are fixed, and the backpropagation algorithm is used to calculate the gradient based on the adversarial loss to update the parameters of the domain discrimination module.

[0064] At the same time, the classification loss is calculated based on the classification prediction results of the classification submodule and the true label. The parameters of the domain discrimination module and the classification module are fixed, and the adversarial loss and classification loss are calculated according to the optimization objective function. Calculate the total loss, then use the backpropagation algorithm to calculate the gradient and update the parameters of the feature extraction submodule. is the classification loss of the classification module (such as cross entropy loss), It is a hyperparameter that balances the adversarial training loss and other losses. Its value is adjusted through experiments (such as trying different values ​​starting from 0.1) to achieve the best training effect. During the training process, dynamic adjustment As training progresses, gradually increase , so that the model focuses more on feature extraction and classification accuracy in the early stage of training, and pays more attention to domain adaptation in the later stage, enhancing the model's generalization ability for data in different domains.

[0065] Through continuous iterative training, the feature extraction module gradually learns domain-invariant features—features with similar distribution and semantics in the source and target domains—thus reducing the impact of domain differences on model performance. During training, a learning rate decay strategy is employed, gradually reducing the learning rate as training progresses to improve model convergence and stability.

[0066] Training ends and model saves: When the training reaches the preset number of iterations or the model performance stops improving significantly, the training ends. The trained feature extractor and classifier parameters are saved for classification of target domain data or other tasks in real applications.

[0067] The preferred embodiments of the present invention have been described in detail above. It should be understood that numerous modifications and variations based on the concepts of the present invention are possible without inventive effort by those skilled in the art. Therefore, any technical solution that can be derived by one skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A palm vein recognition method based on domain adaptation strategy, characterized in that: The method comprises the following steps: Step 1: Construct a dataset, including a source dataset and a target dataset. Use all the data in the source dataset and part of the data in the target dataset as the initial training set, and use the remaining data in the target dataset as the initial test set. Then perform ROI extraction to obtain image data of the palm ROI area. Split the initial training set into a formal training set, and split the initial test set into a formal test set. Step 2: Design an adversarial domain adaptation network, which includes a palm feature extraction module, a classification module, and a domain discrimination module; Step 3: Pre-training is performed on the source dataset. The palm feature extraction module extracts palm features from the source dataset and inputs the palm features into the classification module for classification prediction. Based on the classification loss, the back propagation algorithm is used to update the parameters of the palm feature extraction module and the classification module. Through continuous iterative training, the classification module is able to achieve a preset classification performance on the source dataset. Step 4: Conduct adversarial training on the formal training set, input the formal training set into the pre-trained palm feature extraction module to obtain the corresponding feature representation, then input the obtained feature representation into the domain discrimination module to determine whether the feature representation comes from the source dataset or the target dataset, and calculate the adversarial loss; at the same time, input the obtained feature representation into the classification module, calculate the classification loss based on the classification prediction result in the classification module; calculate the joint loss based on the adversarial loss and the classification loss, and then use the back propagation algorithm to update the parameters of the palm feature extraction module; through continuous iterative training, the palm feature extraction module gradually learns the domain-invariant features, and the training ends when the preset number of iterations or model threshold is reached; Step 5: During the testing phase, the trained palm feature extraction module and the classification module are tested using the formal test set to obtain the palm ID and complete the evaluation of the model performance.

2. The palm vein recognition method based on domain adaptation strategy according to claim 1, characterized in that: The step 1 includes the following sub-steps: Step 1.1: Use two or more different palm datasets as the source dataset and the target dataset, respectively. The source dataset is historical data collected by an existing device, and the target dataset is data collected by a new palm collection device. Use all the data in the source dataset and part of the data in the target dataset as the initial training set, and use the remaining data in the target dataset as the initial test set. Step 1.2: performing data preprocessing on the source dataset and the target dataset to unify the data formats. The data preprocessing methods include but are not limited to normalization, enhancement, or resizing. Step 1.3: Add palm segmentation and key point detection tasks to the target detection model, perform palm location and segmentation through the ROI extraction, determine the positions of key feature points in the palm, split the initial training set into the formal training set, and split the initial test set into the formal test set.

3. The palm vein recognition method based on domain adaptation strategy according to claim 2, characterized in that: The target detection model in step 1.3 is YOLO series, Faster R-CNN or SSD.

4. The palm vein recognition method based on domain adaptation strategy according to claim 2, characterized in that: The key feature points in the palm in step 1.3 include 7, namely the base point of the index finger, the base point of the middle finger, the base point of the thumb, the base point of the little finger, the center point of the palm, and the two key points on the left and right where the wrist and palm are connected.

5. The palm vein recognition method based on domain adaptation strategy according to claim 1, characterized in that: The palm feature extraction module in step 2 includes a depthwise separable convolution module, a multi-level feature extraction module and a sub-graph Transformer module, wherein the depthwise separable convolution module performs preliminary feature extraction, the multi-level feature extraction module stacks multiple improved inverse residual structures and linear Transformer encoders, and the sub-graph Transformer module divides the image of the palm ROI area into multiple sub-graphs of equal size, and uses the Transformer model to mine and output a global feature vector of fixed dimension.

6. The palm vein recognition method based on domain adaptation strategy according to claim 5, characterized in that: The improved inverse residual structure in the multi-level feature extraction module first expands and groups the number of channels through grouped point-by-point convolution, performs convolution operations within each group, and then performs 3×3 depth convolution on the expanded channel dimension to extract the local spatial features of the channel. At the same time, the channel attention module is combined to perform global average pooling on the features of each channel to obtain the global information of the channel, and then generate a weight coefficient for each channel. Finally, the number of channels is compressed back to the dimension before expansion and grouping through the grouped point-by-point convolution to achieve feature fusion and linear transformation.

7. The palm vein recognition method based on domain adaptation strategy according to claim 6, characterized in that: The multi-level feature extraction module replaces the ReLU activation function with a function including but not limited to a Mish function or a Swish function, wherein the Mish function has a self-gating characteristic and its expression is Mish(x)=x*tanh(softplus(x)); the expression of the Swish function is Swish(x)=x*sigmoid(x).

8. The palm vein recognition method based on domain adaptation strategy according to claim 6, characterized in that: The improved inverse residual structure of the multi-level feature extraction module extracts feature X, and divides the feature map of feature X into multiple overlapping small patches of equal size. Each patch is regarded as an element in a sequence. The linear Transformer encoder is then used to calculate the attention weight between each patch and all other patches, dynamically determining the importance of each patch in the global feature, realizing the abstraction and fusion of the global features. The feature map of feature X is then expanded into P non-overlapping flattened patches, each of the flattened patches is indexed and numbered, and for each flattened patch, the linear Transformer encoder is applied to encode the relationship between patches to obtain the global features between the patches.

9. The palm vein recognition method based on domain adaptation strategy according to claim 1, characterized in that: The classification module in step 2 includes one or more fully connected layers, and finally a Softmax layer for classification. The classification loss function is the cross entropy loss function. , where the number of neurons in the fully connected layer can be adjusted according to the classification task.

10. The palm vein recognition method based on domain adaptation strategy according to claim 1, characterized in that: The domain discrimination module in step 2 includes several convolutional layers, fully connected layers and a Softmax function, and a batch normalization layer and an activation function are added between the fully connected layers; the domain discrimination module uses a binary classification method to distinguish whether the palm features are from the source dataset and the target dataset, and then performs feature confusion by gradient reversal to align the edge distribution between the datasets, wherein the adversarial loss function : in, Representative palm images from the source dataset, represents the palm image from the target dataset, S is the number of palm images in the source dataset, T is the number of palm images in the target dataset, represents the probability that the domain discrimination module predicts that the palm image comes from the source dataset, It represents the probability that the domain discrimination module predicts that the palm image comes from the target dataset.

Citation Information

Patent Citations

  • Augmented reality deep gesture network

    CN112292690A

  • Finger vein recognition training method, test method and related device

    CN114818917A

  • Transfer learning method for improving confrontation-based unsupervised domain adaptation effect

    CN116702878A

  • Gear case fault diagnosis method based on deep transfer learning

    CN116894187A

  • Palm vein feature extraction network training method based on inter-class generated data amplification

    CN118397668A

Cited By

  • Generalized time-frequency positioning method for distributed training scene

    CN121692206A