Palm print feature extraction method based on multi-scale deep semantic segmentation network
By combining a multi-scale deep semantic segmentation network with the CD module and the CBAM attention mechanism, the limitations of traditional palmprint recognition technology in feature extraction accuracy and completeness are overcome, high-precision palmprint feature extraction is achieved, and the objectivity and standardization of traditional Chinese medicine diagnosis are supported.
Patent Information
- Application Number
- CN202510927016.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-21
AI Technical Summary
Traditional image processing technology has limitations in feature extraction accuracy and texture line recognition integrity in palmprint recognition, making it difficult to meet the objectivity and standardization requirements of traditional Chinese medicine diagnosis.
A multi-scale deep semantic segmentation network is adopted, combined with the CD module and the CBAM attention mechanism for multi-scale feature extraction, feature information is transferred through jump connections, and deep separable convolution is introduced in the decoder stage to solve the problem of information loss and enhance feature expression capabilities.
It significantly improves the accuracy and robustness of palmprint feature extraction, provides reliable technical support for TCM diagnosis, and promotes the modernization of TCM diagnosis.
Smart Images

Figure CN120823628A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biometric recognition technology, and more specifically, to the field of palm texture line recognition technology. More specifically, the present invention provides a palmprint feature extraction method based on a multi-scale deep semantic segmentation network. This method innovatively combines a multi-scale deep semantic segmentation network with palmprint feature extraction. By constructing a dedicated palm texture line recognition dataset and designing a novel deep learning model architecture, it achieves accurate palm texture line extraction. Background Art
[0002] Palmprint recognition, as an important research direction in the field of biometric identification, has received widespread attention from academia and industry in recent years. Palm texture lines contain rich individual biometric information and are significantly correlated with human health status. In the theoretical system of Traditional Chinese Medicine, the morphology, direction, and distribution characteristics of palm texture lines are considered to be important indicators reflecting the functional status of internal organs in the human body. This makes palm texture line analysis unique in application value in Traditional Chinese Medicine diagnosis. With the rapid development of medical image processing technology, researchers have begun to apply image analysis methods to the recognition of the three main lines of the palm (life line, wisdom line, and emotion line), and have achieved a series of phased results. However, traditional image processing technology still has obvious limitations in terms of feature extraction accuracy and texture line recognition integrity.
[0003] In recent years, breakthroughs in deep learning technology have brought new opportunities for palmprint recognition. As a major breakthrough in the field of artificial intelligence, deep learning, with its powerful feature learning and representation capabilities, has significantly improved the accuracy and efficiency of palm vein line recognition. Modern palmprint recognition technology builds high-quality annotated datasets and employs advanced semantic segmentation models to accurately segment palm vein lines. This then performs feature extraction and pattern recognition, effectively capturing the unique characteristics of individual palm vein lines. This technological breakthrough not only significantly improves recognition accuracy but also demonstrates excellent robustness and adaptability in practical applications. Notably, advances in palmprint recognition technology provide new technical support for the modernization of Traditional Chinese Medicine (TCM) diagnosis. This technology not only assists TCM practitioners in more accurate health assessments but also facilitates the objectification and standardization of TCM diagnoses. By integrating deep learning with TCM theory, palmprint recognition is expected to develop into an efficient and reliable intelligent diagnostic tool. This will not only help promote the modernization of TCM but also provide a new technological path for modern medical diagnosis. In the future, with the continuous advancement of technology, the application prospects of palmprint recognition in the healthcare field will be even broader. Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the present invention proposes a palmprint feature extraction method based on a multi-scale deep semantic segmentation network. This method achieves accurate extraction of palm texture lines by constructing a multi-scale deep semantic segmentation model, aiming to provide reliable technical support for TCM diagnosis. The innovation of the present invention lies in: first, the accuracy and robustness of palm texture line extraction are significantly improved through the multi-scale deep semantic segmentation network; second, the organic combination of modern computer vision technology and TCM diagnosis theory provides an objective and quantitative analysis basis for TCM diagnosis, thereby promoting the modernization of TCM diagnosis.
[0005] Specifically, the present invention provides the following technical solutions:
[0006] S1. Data preprocessing: Construct a multimodal palmprint image dataset and generate a standardized data format that meets the requirements of deep learning input through data enhancement (including random rotation and Gaussian blur) and standardization.
[0007] S2. Feature extraction: In the encoder stage, the CD module is used to extract multi-scale features from the input image. This mainly uses parallel depthwise convolution with kernels of 1, 3, and 5 to capture local texture features of different receptive fields, and suppresses unimportant areas through the CBAM attention mechanism, thereby focusing on the information of key areas.
[0008] S3, Feature Transfer: The multi-scale extracted features are transferred through two paths. On the one hand, the shallow high-resolution features are transferred to the decoder through skip connections. On the other hand, the feature dimension is reduced through the maximum pooling layer to retain the deep semantic information.
[0009] S4, iterative processing: After performing three feature extraction-transfer iterations, in the fourth iteration, the feature matrix output by the CD module is input into the PE module, and the spatial and channel dimension features are separated by depthwise separable convolution;
[0010] S5, Feature Optimization: The PE module performs a depth-wise separable convolution operation on the received features with a kernel size of 5 to remove redundant feature information and perform feature concatenation on the filtered information and the information transmitted by the skip connection;
[0011] S6, Feature Enhancement: The spliced feature information is input into the CD module again for multi-scale feature extraction to obtain richer feature information. At the same time, the CBAM attention mechanism is used to suppress non-important information and strengthen the feature expression of key areas;
[0012] S7, result generation: Repeat steps S5-S6 three times to finally obtain the feature map segmentation result.
[0013] Preferably, the CD module in step S2 uses deep convolution with different convolution kernel sizes to extract multi-scale palmprint features, and combines the CBAM attention mechanism to focus on important areas. The formula of the CD module is as follows:
[0014] m′ i =σ2(Conv 1×1 (m i )) (1)
[0015]
[0016]
[0017] In formula 1, Conv 1×1 Indicates a 1×1 convolution of the input image, and σ represents the ReLU activation function, which introduces nonlinearity to increase the model's ability to fit complex data. i Represents the result of the image after convolution and activation function, and passes this result to the depth convolution of the next layer, m i represents the input tensor;
[0018] In formula 2, DWConv 1×1 Indicates the use of a 1×1 convolution kernel for depthwise convolution. The function of this convolution kernel is to retain and integrate channel information. CBAM stands for the CBAM attention mechanism, which includes channel attention and spatial attention. It focuses on important areas of the input channel information and maintains interaction with spatial information.
[0019] In formula 3, DWConv 3×3 It indicates that a 3×3 convolution kernel is used for depth convolution. The convolution kernel size is set to 3. This is mainly to capture local channel information and spatial information, and to fuse the feature information from the 1×1 depth convolution to improve the feature expression ability, and then pass the new feature information to the CBAM attention mechanism.
[0020] In formula 4, the summation formula is used, DWConv ksHere, ks is the size of the convolution kernel, and KS defines a set of convolution kernels {1×1, 3×3, 5×5}, where the features of the bottom layer of CDBlock after Conv convolution and ReLU activation are first input to the DWConv5×5 deep convolution layer to complete the global feature extraction under the large receptive field, and then the output of DWConv5×5 is combined with the features from the previous layer branch (the second branch of the 3×3 deep convolution) by element-by-element addition. Since the features of the second branch have been integrated with the output of the first layer branch (1×1 deep convolution) in advance (that is, the result of the first branch DWConv1×1 is first added to the result of the second branch DWConv3×3), the third branch indirectly contains the fine-grained texture information from the first layer 1×1 deep convolution path when aggregating features. In formula 5, This is the final fusion result of the feature maps from the three branches. After feature extraction and layer-by-layer information interaction, the resulting feature map is more accurate and efficient.
[0021] Preferably, the structure of the CD module in step S3 is as follows: the initial features of the CDBlock are provided by the output of the underlying Conv convolution after ReLU activation, and the features are synchronously input to three parallel depth convolution paths (corresponding to the three arrows branching downward from the "ReLU" module in the figure). The three paths capture multi-scale features with depth convolutions of different sizes, corresponding from left to right: Path 1 (DWConv1×1 branch): receives the ReLU output and extracts fine-grained texture features through DWConv1×1; Path 2 (DWConv3×3 branch): receives the ReLU output and extracts medium-scale contextual features through DWConv3×3; Path 3 (DWConv5×5 branch): receives the ReLU output and extracts large receptive field global features through DWConv5×5. The module achieves cross-path feature concatenation through two element-by-element additions: First, the DWConv1×1 output of path 1 is element-wise added to the DWConv3×3 output of path 2, and the fused result is fed into the subsequent CBAM module in path 2. Second, the fused features of path 2 are element-wise added to the DWConv5×5 output of path 3, and the fused result is fed into the subsequent CBAM module in path 3. Features from all three paths are enhanced by CBAM: the DWConv1×1 output of path 1 is directly fed into CBAM; the features of paths 2 and 3 are fed into CBAM after cross-path fusion, achieving weighted attention in both channel and spatial dimensions. The outputs of the three CBAMs are aggregated into the final output of the CD module. The feature information processed by the CD module is then passed to the decoder stage via skip connections. This is primarily because feature information is often lost during downlink processing through the pooling layer. Therefore, skip connections preserve feature information from the current module, allowing for fusion during upsampling, thereby achieving a more comprehensive image reconstruction. Furthermore, it is passed to the pooling layer, further compressing the spatial resolution of the feature map while preserving key information.
[0022] Preferably, in step S4, steps S2 and S3 are repeated three times, mainly to allow the network to gradually extract more advanced features and reduce detailed information, retain the most useful features, thereby enhancing the learning ability of the model, and after multiple repetitions, the network becomes more abstract, thereby providing more meaningful features for the subsequent decoder stage.
[0023] Preferably, in step S5, a PE module is added during the upsampling process. This module mainly includes normalization, ReLU activation function, depthwise separable convolution, and residual connection. Its main function is to solve the information redundancy problem caused by multi-scale feature extraction and effectively restore the spatial detail information that may be lost during the upsampling process. The formula of the PE module is as follows:
[0024]
[0025] In this formula, the encoder's features are channel-adjusted through 1×1 convolution, and the feature map is normalized through the normalization layer. Next, the ReLU function introduces nonlinear changes, enhancing the model's ability to fit complex data. Subsequently, the depthwise separable convolution layer effectively reduces the amount of computation while maintaining computational efficiency. Finally, the 5×5 convolution kernel further enhances the ability to capture spatial information and contextual relationships, thereby helping to restore the texture details of the image. It is then directly fused with the features after the depthwise separable convolution through the residual connection. On the one hand, by retaining the original features of the input, it helps the network avoid excessive transformation of the input features, thereby improving the model's expressive power. On the other hand, it can accelerate convergence and improve training efficiency.
[0026] Preferably, in step S6, the features processed by the PE module are passed through the CD module again. Its function is to use the multi-scale features of the CD module to extract the features after the jump connection and the PE module processing, so as to solve the problem of insufficient detail information in the process of image reconstruction.
[0027] Preferably, in step S7, steps S5 and S6 are repeated three times. Through multiple applications, the model's understanding of global information can be gradually strengthened, and the model can gradually learn and optimize the details of the generation task at each layer, thereby improving the ability to reconstruct feature maps.
[0028] Compared to existing technologies, this method achieves anatomically accurate palm texture line extraction by constructing a high-precision palmprint semantic segmentation dataset and innovatively designing a multi-scale attention fusion network architecture. This method exhibits the following key technical advantages: First, through multi-scale feature extraction and fusion of an attention mechanism, it effectively preserves the detailed features and spatial structure of palmprints. Second, to address the spatial information attenuation problem in traditional codec architectures, a depthwise separable convolution is introduced in the decoding stage to construct a spatial position-preserving unit, effectively enhancing the continuity of ridge line edges through separable feature learning. In a comparative experiment on a self-built palmprint dataset, the peak signal-to-noise ratio (PSNR) of this method was improved by 14.58 dB compared with the traditional image processing method (Gabor filtering + morphological segmentation); compared with the medical image segmentation benchmark model (including U-Net, Swin-UNet and FusionU-UNet), the Dice similarity coefficient (DSC) was increased by 0.87 percentage points, and the mean intersection over union (mIoU) was increased by 7.15 percentage points; therefore, the method proposed in this invention can more accurately extract palm texture features, which not only solves the limitations of traditional palmprint recognition methods in feature extraction accuracy, but also provides a new technical idea for the field of biometric recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0030] Figure 1 A model framework diagram of the palmprint feature extraction method based on a multi-scale deep semantic segmentation network provided in an embodiment of the present invention.
[0031] Figure 2 This is a network structure diagram of the feature extraction CD module provided in an embodiment of the present invention.
[0032] Figure 3 This is a network structure diagram of the feature optimization PE module provided in an embodiment of the present invention.
[0033] Figure 4 This is a schematic diagram of the palm texture lines segmented through model prediction provided by an embodiment of the present invention. In the figure, red represents the life line, green represents the wisdom line, yellow represents the emotion line, blue represents the destiny line, and pink represents the health line. DETAILED DESCRIPTION
[0034] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments described are only a portion of the embodiments of the present invention, and not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0035] Those skilled in the art should be aware that the following specific embodiments or implementations are a series of optimized configurations listed in the present invention to further explain the specific content of the invention, and these configurations can be combined or used in conjunction with each other, unless the present invention explicitly states that some or a specific embodiment or implementation cannot be combined or used in conjunction with other embodiments or implementations. At the same time, the following specific embodiments or implementations are only intended to be optimized configurations and are not to be understood as limiting the scope of protection of the present invention.
[0036] The present invention aims to improve the accuracy and completeness of palm texture line extraction. First, by systematically annotating palmprint images, a high-quality palmprint dataset was constructed, solving the problem of scarcity of palmprint data in the field of deep learning. Secondly, in order to achieve better segmentation effect, a semantic segmentation model was constructed based on the U-Net architecture, and the following innovative improvements were made: the designed CD module was introduced in the encoder stage for multi-scale feature extraction to capture fine palm texture lines; the PE module was added in the decoder stage to eliminate redundant feature information and strengthen the spatial position of the features. The improved model can extract palm texture features comprehensively and multi-scale, significantly improving the accuracy of texture line recognition, while extracting richer texture information to provide a more reliable basis for TCM-assisted diagnosis.
[0037] The core innovation of the present invention is: first, the deep convolution kernels of different scales in the CD module are used to realize multi-scale feature capture, and the CBAM attention mechanism is used to effectively filter out irrelevant information, so that the model focuses on key features; second, the feature information is transferred to the decoder before pooling through jump connections, which effectively solves the problem of information loss in the downsampling process; third, the deep separable convolution operation is introduced in the PE module, and spatial position information is added, which not only solves the problem of information loss in the image reconstruction process, but also improves the feature expression ability by assigning an independent convolution channel to each feature map. In addition, the normalization operation is added to the PE module to ensure training stability, and the ReLU activation function is used to enhance the model's fitting ability for complex features, which significantly improves the detail performance of image reconstruction. The technical advantages of the present invention are reflected in: on the one hand, the multi-scale deep semantic segmentation network is innovatively applied to palm texture line extraction, which greatly improves the recognition accuracy; on the other hand, the segmentation results can be further extracted and applied to TCM auxiliary diagnosis, which not only significantly reduces the workload of medical personnel, but also provides technical support for the objectivity and standardization of TCM diagnosis, and effectively promotes the modernization of TCM. Experimental results show that this method is superior to traditional methods in terms of the completeness and accuracy of texture line extraction, and has important clinical application value. The following is a detailed description of the implementation method of this solution with reference to the accompanying drawings.
[0038] Figure 1 This is a model framework diagram of the palmprint feature extraction method based on a multi-scale deep semantic segmentation network provided by an embodiment of the present invention. The specific structure of the model of this embodiment includes:
[0039] The palm texture line segmentation model proposed in this paper uses an improved U-Net architecture, consisting of a symmetrical encoder-decoder structure. The encoder comprises four sets of "CDBlock × 2 + downsampling" structures (the first three include Maxpool, and the last does not), and the decoder comprises three sets of "Patch Expand + CDBlock × 2 + skip connection" structures, symmetrical with the encoding path. High-precision segmentation is achieved through a cross-layer feature fusion mechanism. The specific structure is described as follows: In the encoder module, the multi-scale feature extraction layer includes each level of the encoder composed of parallel depth convolution, including 3 groups of parallel convolution kernels (the convolution kernel sizes are 1, 3, and 5 respectively), and the multi-scale feature capture of texture lines is realized by splicing features across the receptive field; the downsampling unit includes a 2×2 maximum pooling layer (stride = 2) set at the end of each level of the encoder, the feature map resolution is halved step by step, and the number of channels changes from 32→64→128→256→512, following the classic parameter design of U-Net; the feature transmission path adopts a dual-path transmission mechanism. The horizontal path is a jump connection path for the shallow details of the encoding end to flow back to the decoding end. The connection relationship is strictly symmetrical with the encoding end level: the i-th layer jump connection is the output of the i-th group of the encoder transmitted to the input of the i-th group of the decoder through a horizontal arrow. The vertical path is the core path for downsampling layer by layer and strengthening deep semantics at the encoding end. It is composed of a cascade of CDBlock×2 and Maxpool (pooling) modules, and includes three rounds of downsampling operations: the i-th round extracts local features of the current scale through CDBlock×2, and then passes the features to the next layer through Maxpool (pooling) downsampling; in the decoder module, the feature recovery layer includes: each level of upsampling uses a transposed volume (kernel_size=2, stride=2) to achieve resolution doubling, and the number of channels decreases from 512→256→128→64→32; and the PE module (i.e. Figure 1Patch Expand in the decoder), which contains a convolution with a convolution kernel of 1 for adjusting the channel, and the BatchNorm layer is used alternately with the ReLU activation function, and a residual connection is introduced to alleviate the gradient disappearance problem, and a 5×5 depthwise convolution (Depthwise Conv) and a 1×1 pointwise convolution (Pointwise Conv) are added. The convolution is to enhance the spatial information of the features; the multi-scale feature extraction layer is also introduced in the decoder stage, which mainly processes the output features processed by the PE module and the features concatenated with the jump connection features, and then performs multi-scale feature extraction to achieve the restoration of details when reconstructing the image; the output layer uses dual 3×3 convolution kernels for feature optimization, and finally achieves channel compression through 1×1 convolution, and generates binary segmentation results through the Sigmoid activation function. Experiments show that the improved model can accurately segment the main texture lines of the palm, such as the lifeline, wisdom line, and emotion line. Compared with the segmentation results of the baseline model, the Dice similarity coefficient (DSC) is increased by 11.02 percentage points and the mean intersection over union (mIoU) is increased by 14.82 percentage points. This high-precision segmentation capability provides a reliable technical foundation for the quantitative analysis of traditional Chinese medicine palm diagnosis.
[0040] Figure 2 This is a framework diagram of the CD module of the palmprint feature extraction method based on the multi-scale deep semantic segmentation network provided by the embodiment of the present invention. The specific structure of this module includes:
[0041] Introducing the CD module at the encoder stage (i.e. Figure 1 The CDBlock in the module adopts a multi-scale deep convolution kernel architecture to extract the features of the palm texture line. The specific structure is described as follows: The input feature m is processed by a 1×1 convolution layer and a ReLU activation function (σ2). i Perform nonlinear transformation to generate reference feature m′ i , branch 1 to m′ i After performing 1×1 depthwise separable convolution (DWConv) and optimizing it with the CBAM module, the output is detail-enhanced features. Branch 2 fuses the 1×1 and 3×3 depth convolution results and generates cross-scale features through CBAM Branch 3 aggregates the multi-core (1×1 / 3×3 / 5×5) deep convolution output and generates robust features through CBAM The three-way feature Add element-wise to generate the final output This mechanism effectively solves the information redundancy problem caused by multi-scale feature extraction through the dual weighting of channel attention and spatial attention, enabling the network to adaptively focus on key areas in the image. The feature extraction process of the CD module can be expressed as the following formula:
[0042] m′i =σ2(Conv 1×1 (m i )) (1)
[0043]
[0044] where m′ i Represents the structure after the 1×1 convolution kernel ReLU activation function, m i represents the input tensor, Represents the result of depth convolution with kernel 1 and applying CBAM attention. It represents the result of concatenating the features from the depthwise convolution with a convolution kernel of 1 and the features from the depthwise convolution with a convolution kernel of 3 and performing CBAM processing. It represents the result of concatenating the features from the previous convolution kernel 1 and the convolution kernel 3 with the features from the convolution kernel 5 of this layer and processing them with CBAM. The variable i represents the i-th CDBlock module in the network. Represents the output tensor after the i-th stage.
[0045] The parallel depthwise convolutions are used to extract multi-scale features. The convolution kernels are set to {1×1, 3×3, 5×5}. These kernels capture both local and global information. To prevent feature redundancy in multi-scale feature extraction, CBAM attention is implemented at each layer to focus on important features. Furthermore, to ensure better interaction between feature information at each layer, each layer is passed to the next layer after undergoing depthwise convolution. This prevents the CNN from failing to understand the global context when processing feature information at each layer.
[0046] The formula of the CBAM attention mechanism is as follows:
[0047] M C (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))
[0048] M s (F)=σ(f 5×5 ([AvgPool(F);MaxPool(F)]))
[0049] Where F represents the input feature map, σ() represents the Sigmiod activation function, and M c (F) represents the channel attention weight, M c(F) indicates that the channel attention weight is applied to the original feature map, MLP indicates the fully connected layer, AvgPool() indicates average pooling, MaxPool() indicates maximum pooling, and one of the hyperparameters r is the scaling factor. In order to avoid excessive computer complexity, r is preferably set to 32. s (F) represents the spatial attention weight, f 5×5 The 5×5 here is the size of the convolution kernel. On the one hand, it is able to capture more spatial information, but on the other hand, it cannot be too large, which will easily increase the computational overhead. Therefore, the size of the convolution kernel is set to 5×5.
[0050] Figure 3 This is a PE module framework diagram of the palmprint feature extraction method based on a multi-scale deep semantic segmentation network provided by an embodiment of the present invention. The specific structure of this module is as follows:
[0051] In the decoder stage, in order to better process the feature map transmitted from the encoder, the PE module is added to enhance the spatial position information of the up-sampled features. The specific structure of the module is described as follows: Through the 1×1 convolution layer (Conv 1×1 ) for the input feature u i Perform cross-channel linear transformation to generate benchmark features F base . Batch normalization (BN), ReLU activation function (σ3) and 5×5 depth-wise separable convolution (DSWConv) are performed in sequence to extract high-order spatial-channel features F DSW , F DSW With F base Add them element by element, constrain the output range through secondary batch normalization (BN), and finally generate enhanced features through ReLU activation function (σ3) The core of this module is the addition of depthwise separable convolution. Since the decoder stage is the process of image reconstruction, and reconstruction often focuses on spatial position information, depthwise separable convolution is added to enhance spatial position information. This convolution can also effectively handle the feature redundancy problem caused by multi-scale feature extraction. Its model formula is as follows:
[0052]
[0053] Among them, σ3 represents the ReLU activation function, u i Represents input features, BN represents normalization, DSWConv represents depth-separable convolution, which includes 5×5 depth convolution and 1×1 point-by-point convolution. Represents the enhanced features. In order to enhance the spatial position information while avoiding the problem of excessive computer overhead caused by overly large convolution kernels, the convolution kernel of the depthwise separable convolution is set to 5×5, and a residual connection is added after the feature map passes through the Conv adjustment channel. The residual connection directly splices the feature information with the information after the depthwise separable convolution, avoiding the problem of loss of detail information after the depthwise separable convolution. In order to enable the model to better fit complex data, the ReLU function is selected as the activation function. Compared with the Sigmoid function and the Tanh function, it does not cause the gradient disappearance problem when inputting extreme values, and ReLU will directly map negative inputs to 0, which helps to reduce computational complexity. In addition, the depthwise separable convolution includes depthwise convolution and pointwise convolution. When accepting feature vectors, it first divides independent channels for each feature vector through depthwise convolution, and then integrates the output of the depthwise convolution through pointwise convolution with a convolution kernel of 1×1. The formula is as follows:
[0054]
[0055]
[0056] Where h represents the height of the input feature map, w represents the width of the input feature map, K represents the spatial size of the convolution kernel, and C in Indicates the total number of input channels, where c ranges from 0 <c<C in , X(h+i, w+j, c) represents the value of the cth channel of the input feature map at position (h+i, w+j), K d (C) (i, j) represents the weight of the depth convolution kernel corresponding to the c-th input channel at position (i, j), Represents the output value of the cth channel at position (h,w) after depth convolution. K p (k,c) represents the 1×1 convolution kernel weight that maps the cth input channel to the kth output channel, Represents the final output value of the kth channel at position (h, w) after point-by-point convolution.
[0057] Furthermore, in order to evaluate the effect of palm texture line segmentation, this implementation uses the peak signal-to-noise ratio (PSNR) indicator to verify the difference between the designed multi-scale deep semantic segmentation model and the traditional image processing method for palm texture line segmentation. The results are compared in Table 1.
[0058] Table 1 Peak signal-to-noise ratio of the proposed algorithm and traditional image processing algorithms (unit: dB)
[0059] method PSNR Palmprint recognition using directional line energy feature 35.26 Principal line-based alignment refinement for palmprint recognition 37.42 Palmprint main line advance method based on morphological filtering and neighborhood search 36.74 Palmprint main line extraction based on multi-directional filtering and neighborhood denoising 38.71 Palmprint feature extraction method based on multi-scale deep semantic segmentation network (this method) 53.29
[0060] As can be seen from the table above, compared with traditional image processing algorithms in the existing field, this solution improves the peak signal-to-noise ratio (PSNR) by between 14.58dB and 18.03dB, effectively demonstrating the superior performance of this method in palm texture feature extraction.
[0061] To verify the advantages of this method in deep learning and its superiority in palmprint segmentation tasks, this example selected four evaluation metrics: the Dice coefficient, mean intersection over union (mIoU), mean average precision (mAP), and mean average precision (mPrecision). The proposed model was compared with existing mainstream medical image segmentation models on a palmprint dataset, as shown in Table 2.
[0062] Table 2 Performance comparison of this model and existing mainstream segmentation models on the palmprint dataset
[0063]
[0064]
[0065] The table above shows that our model outperforms other models in terms of mIoU and DSC. Specifically, compared with Transformer-based models (such as Swin-UNet), our model improves DSC, mIoU, and mPrecision by 0.81%, 1.89%, and 2.35%, respectively. Compared with models based on enhanced skip connections (such as Fusion-UNet), our model improves DSC by 0.52% and mIoU by 1.06%. The experimental results demonstrate the superiority of our solution in the palmprint image segmentation task.
[0066] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various non-substantial improvements are made using the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the scope of protection of the present invention.
[0067] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0068] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A palmprint feature extraction method based on a multi-scale deep semantic segmentation network, characterized in that: The method comprises: S1. Data preprocessing: annotate existing palmprint images and construct a palmprint dataset; S2, Feature Extraction: In the encoder stage, the CD module is used to perform multi-scale feature extraction on the images in the input palmprint dataset; S3, feature transfer: The multi-scale extracted features are transferred through two paths, one of which is passed to the decoder through a skip connection, and the other is passed down to the pooling stage; S4, iterative processing: Repeat steps S2-S3 three times. In the fourth iteration, the features processed by the CD module are directly passed to the PE module for upsampling operation; S5, feature optimization: The PE module performs a deep separable operation on the features received in step S4 after being processed by the CD module, removes redundant feature information, and concatenates the filtered feature information with the features transmitted by the jump connection; S6, Feature Enhancement: The concatenated feature information is input into the CD module again for multi-scale feature extraction to obtain richer feature information, while using the attention mechanism to suppress non-important information; S7, result generation: Repeat steps S5-S6 three times to finally obtain the feature map segmentation result.
2. The palmprint feature extraction method based on a multi-scale deep semantic segmentation network according to claim 1 is characterized in that: In step S1, a palmprint image is annotated using an annotation tool, and texture lines are named based on traditional Chinese medicine palmprint knowledge to obtain a palmprint dataset suitable for palmprint segmentation.
3. The palmprint feature extraction method based on a multi-scale deep semantic segmentation network according to claim 1 is characterized in that: The CD module in step S2 uses depthwise convolution with different kernel sizes to extract multi-scale palmprint features and combines the CBAM attention mechanism to focus on important areas. The model formula of the CD module is as follows: m′ i =σ2(Conv 1×1 (m i )) Among them, Conv 1×1 Indicates 1×1 convolution of the input image, σ2 represents the ReLU activation function, m′ i Represents the result of the image after convolution and activation function, m i represents the input tensor; DWConv 1×1 Indicates the use of 1×1 convolution for depth convolution, and CBAM indicates the CBAM attention mechanism operation; DWConv 3×3 Indicates that a 3×3 convolution kernel is used for depth convolution; DWConv ks Where ks is the size of the convolution kernel, and KS defines a set of convolution kernels {1×1, 3×3, 5×5}. The features of the bottom layer of CDBlock after Conv convolution and ReLU activation are first input to the DWConv5×5 deep convolution layer to complete the global feature extraction under the large receptive field. Then, the output of DWConv5×5 is combined with the features of the second branch from the previous 3×3 deep convolution layer by element-by-element addition. Since the features of the second branch have been integrated with the output of the first layer branch in advance, the third branch indirectly contains the fine-grained texture information from the first layer 1×1 deep convolution path during feature aggregation. Represents the result of the final fusion of the feature maps after the three branches.
4. The palmprint feature extraction method based on a multi-scale deep semantic segmentation network according to claim 1, characterized in that: The structure of the CD module is as follows: the initial features of CDBlock are provided by the bottom-level Conv convolution output after ReLU activation, and the features are synchronously input to three parallel depth convolution paths; the three parallel depth convolution paths capture multi-scale features with depth convolutions of different sizes, corresponding from left to right: Path 1: receives ReLU output and extracts fine-grained texture features through DWConv1×1; Path 2: Receives ReLU output and extracts mid-scale context features through DWConv3×3; Path 3: Receive ReLU output and extract large receptive field global features through DWConv5×5; The CD module implements cross-path feature concatenation through two element-by-element additions: the first fusion: the DWConv1×1 output of path 1 is added element-by-element to the DWConv3×3 output of path 2, and the fusion result is input to the subsequent CBAM module of path 2; Second fusion: The fused features of path 2 are added element-by-element to the DWConv5×5 output of path 3, and the fusion result is input to the subsequent CBAM module of path 3; the features of the three paths are all enhanced by CBAM: the DWConv1×1 output of path 1 is directly input to CBAM; the features of paths 2 and 3 are input to CBAM after cross-path fusion; the outputs of the three CBAMs are aggregated into the final output of the CD module.
5. The palmprint feature extraction method based on a multi-scale deep semantic segmentation network according to claim 1 is characterized in that: In step S4, steps S2 and S3 are repeated three times. This is mainly to allow the network to gradually extract more advanced features and reduce detailed information, retaining the most useful features, thereby enhancing the learning ability of the model. Moreover, with multiple repetitions, the network becomes more abstract, thereby providing more meaningful features for the subsequent decoder stage.
6. The palmprint feature extraction method based on a multi-scale deep semantic segmentation network according to claim 1, characterized in that: In step S5, a PE module is added during the upsampling process. The module mainly includes a normalization layer, a ReLU function layer, a depthwise separable convolution layer, and a residual connection layer. The model formula of the PE module is as follows: in, represents the enhanced features, u i represents input features, σ3 represents ReLU activation function, BN represents normalization, DSWConv represents depthwise separable convolution, Conv 1×1 Represents a 1×1 convolution.
7. The palmprint feature extraction method based on a multi-scale deep semantic segmentation network according to claim 6, characterized in that: The PE module adjusts the encoder's features through 1×1 convolution and normalizes the feature map through the normalization layer; then, the ReLU function layer introduces nonlinear changes to enhance the fitting ability; subsequently, the depthwise separable convolution layer effectively reduces the amount of computation while maintaining computational efficiency; finally, the 5×5 convolution kernel is used to further enhance the ability to capture spatial information and contextual relationships; and then the features output by the depthwise separable convolution layer are directly fused through residual connections.
8. The palmprint feature extraction method based on a multi-scale deep semantic segmentation network according to claim 1, characterized in that: In step S6, the features processed by the PE module are passed to the CD module again to utilize the multi-scale features of the CD module to extract the features after the jump connection and the PE module processing, thereby compensating for the problem of insufficient detail information in the image reconstruction process.
9. The palmprint feature extraction method based on a multi-scale deep semantic segmentation network according to claim 1, characterized in that: In step S7, steps S5 and S6 are repeated three times. Through multiple applications, the model's understanding of global information can be gradually strengthened, and the model can gradually learn and optimize the details of the generation task at each layer, thereby improving the ability to reconstruct feature maps.