Palm vein break point detection method based on adaptive noise suppression cooperative coupling feature enhancement
By adopting adaptive noise suppression module and coupling feature enhancement module in palm vein detection, the problems of insufficient background noise suppression and difficulty in multi-scale feature extraction are solved, and the number and position of palm vein flexure points are realized.
Patent Information
- Application Number
- CN202510429803.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing palm vein detection methods have problems such as insufficient background noise suppression, loss of detection target position information, and difficulty in extracting multi-scale features of image, resulting in low accuracy in palm vein detection.
Adaptive noise suppression module and coupling feature enhancement module are adopted to suppress background noise on palm vein images through adaptive noise suppression module, and the coupled feature enhancement module is used to enhance the spatial dependence between channels to capture multi-scale information, thereby achieving accurate detection of the number and location of palm vein vertices.
It effectively suppresses background noise, strengthens position information, captures multi-scale information, and improves the accuracy of palmar vein vertebrae detection.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a palm vein inflection point detection method with adaptive noise suppression and collaborative coupled feature enhancement. Background Art
[0002] With the continuous development of technology, biometric recognition technology is increasingly widely used in various fields of society, especially in the fields of identity authentication and security monitoring. As a unique and difficult-to-forge biometric, palm veins have gradually become an emerging identity verification technology due to their high individual differences, and are applied to multiple scenarios such as financial payment, intelligent device unlocking, and public security. In the financial field, the introduction of palm vein technology can improve the security of transactions and prevent identity theft. In smart homes and health monitoring, palm veins can also provide a convenient and secure individual identification method, enhancing the intelligence and personalized services of devices. In addition, the detailed analysis of palm veins can also be used for health management to help detect the signs of certain diseases at an early stage. Therefore, efficiently performing palm vein recognition and inflection point analysis can provide a safer and smarter living environment for society, and at the same time promote the innovation and application popularization of biometric recognition technology. Existing palm vein detection methods have problems such as insufficient background noise suppression, loss of detection target position information, and difficulty in extracting multi-scale features of images, which lead to low accuracy in palm vein inflection point detection. Summary of the Invention
[0003] In view of the above defects of the prior art, the present invention provides a palm vein inflection point detection method with adaptive noise suppression and collaborative coupled feature enhancement to solve the technical problem of inaccurate palm vein inflection point detection.
[0004] To achieve the above object and other related objects, the present invention provides a palm vein inflection point detection method with adaptive noise suppression and collaborative coupled feature enhancement.
[0005] In an embodiment of the present invention, it includes: obtaining a palm vein image to be processed; inputting the palm vein image into a trained inflection point detection network to obtain the number and position information of palm vein inflection points, and the inflection point detection network includes a cascaded: a shallow feature extraction module for preprocessing the palm vein image; an adaptive noise suppression module for suppressing background noise to improve the signal-to-noise ratio of the image; a coupled feature enhancement module for enhancing the position information between each channel to obtain multi-scale information; and a decoder for processing the enhanced features to obtain the number and position information of the palm vein inflection points.
[0006] In an embodiment of the present invention, the adaptive noise suppression module processes the shallow features output by the shallow feature extraction module according to the following steps: dividing the shallow feature X into G sub-features Xi , where \(i\in\{1,2,\ldots,G\}\); using the GAP algorithm to process the sub - feature \(X\) i to obtain the first feature \(X\) 1i ; using the constant \(\epsilon\) to normalize the first feature \(X\) 1i to obtain the second feature \(X\) 2i ; using the learnable parameters \(\gamma\) and \(\beta\) to transform the second feature \(X\) 2i to obtain the third feature \(X\) 3i ; using the sub - feature \(X\) i and the Sigmoid function to process the third feature \(X\) 3i to obtain the fourth feature \(X\) 4i ; concatenating the fourth feature \(X\) 4i to obtain the output feature \(Y\) of the adaptive noise suppression unit.
[0007] In an embodiment of the present invention, the relationship between each feature is expressed by the following formula: \(X\) 1i =X\) i \(\cdot GAP(X\) i ) ; \(X\) 3i =\(\gamma X\) 2i +\(\beta\); \(X\) 4i =X\) i \(\cdot Sigmoid(X\) 3i ) \(Y = X\) 41 +X\) 42 +\(\cdots+X\) 4G
[0008] In an embodiment of the present invention, the coupled feature enhancement unit includes a learnable graph neural unit and a multi - scale information fusion unit. The learnable graph neural unit enhances the spatial dependence relationship between each channel using a graph structure, and the multi - scale information fusion unit captures multi - scale information using a dual - directed pyramid operation.
[0009] In an embodiment of the present invention, the expression of the learnable graph neural unit is as follows: \(Z = Y\cdot ReLU(Conv1d(A\cdot GAP(Y)))\); \(A = A_0\times A_1+A_2\); \(A_0 = Softmax(Conv1d(GAP(Y)))\); where \(Z\) is the output feature of the learnable graph neural unit, \(Y\) is the output feature of the adaptive noise suppression unit, \(GAP\) is the GAP algorithm, \(Conv1d\) is a one - dimensional convolutional layer, \(ReLU\) is an activation function, \(Softmax\) is a normalization, and \(A_0\), \(A_1\) and \(A_2\) are learnable adjacency matrices.
[0010] In one embodiment of the present invention, the multi-scale information fusion unit includes a first branch and a second branch in parallel. The first branch is composed of a cascaded first convolutional layer, multiple parallel grouped convolutions, and a second convolutional layer. The second branch is composed of a cascaded adaptive average pooling layer, a first convolutional layer, multiple parallel grouped convolutions, a second convolutional layer, and a first upsampling layer. After connecting the output features of the first branch and the second branch, they are sequentially passed through a third convolutional layer and a second upsampling layer to obtain the output features of the multi-scale information fusion unit.
[0011] In one embodiment of the present invention, the multiple parallel grouped convolutions include four grouped convolutions: the first grouped convolution has a kernel size of 9×9 and the number of groups is 16; the second grouped convolution has a kernel size of 7×7 and the number of groups is 8; the third grouped convolution has a kernel size of 5×5 and the number of groups is 4; the fourth grouped convolution has a kernel size of 3×3 and the number of groups is 1.
[0012] In one embodiment of the present invention, the loss function during the training of the inflection point detection network is as follows: K Overall =K Bayesian +λK Count ; where K Overall is the total loss function, K Bayesian is the Bayesian uncertainty loss, K Count is the counting error loss, and λ is a hyperparameter used to balance the weights of the two losses.
[0013] Advantages of the present invention: An adaptive noise suppression and collaborative coupling feature enhancement method for palm vein inflection point detection proposed by the present invention. This method realizes the suppression of background noise in palm vein images through an adaptive noise suppression module, and then uses a coupling feature enhancement module to enhance the spatial dependence relationship between channels to strengthen the position information. At the same time, a pyramid structure is used to increase the kernel size from bottom to top and grouped convolutions are used to reduce the kernel depth to capture multi-scale information, thereby realizing the effective and accurate detection of the number and position of inflection points in palm vein images. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 It is the architecture diagram of the inflection point detection network provided by one embodiment of the present invention; Figure 2 The comparison diagram of palm veins before and after processing provided by an embodiment of the present invention; Figure 3 The architecture diagram of the adaptive noise suppression module provided by an embodiment of the present invention; Figure 4 The architecture diagram of the learnable graph neural unit provided by an embodiment of the present invention; Figure 5 The architecture diagram of the multi-scale information fusion unit provided by an embodiment of the present invention. Detailed implementation manners
[0016] The following describes the implementation manners of the present invention through specific specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. In addition to the specific methods, devices, and materials used in the embodiments, according to the knowledge of those skilled in the art in the technical field and the description of the present invention, any methods, devices, and materials similar or equivalent to the methods, devices, and materials described in the embodiments of the present invention can also be used to implement the present invention.
[0017] It should be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific specific implementation manners, rather than for limiting the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field of the present invention.
[0018] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of these embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0019] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of methods and computer program products that can be implemented according to various embodiments disclosed in the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0020] Please refer to Figure 1 , Figure 1 A palm vein inflection point detection method for adaptive noise suppression and collaborative coupling feature enhancement provided by an embodiment of the present invention includes two steps: First, obtain a palm vein image to be processed, and the palm vein image to be processed is as shown Figure 2 on the left; Second, input the palm vein image into a trained inflection point detection network to obtain the number and position information of the palm vein inflection points. The inflection point detection network includes a cascaded shallow feature extraction module, an adaptive noise suppression module, a coupling feature enhancement module, and a decoder. Among them, the shallow feature extraction module is used to preprocess the palm vein image; the adaptive noise suppression module is used to suppress background noise to improve the signal-to-noise ratio of the image; the coupling feature enhancement module is used to enhance the position information between channels to obtain multi-scale information; the decoder is used to process the enhanced features to obtain the number and position information of the palm vein inflection points, and the processing result is as shown Figure 2 on the right.
[0021] This method realizes the suppression of the background noise of the palm vein image through the adaptive noise suppression module, and then uses the coupling feature enhancement module to enhance the spatial dependence relationship between channels to strengthen the position information. At the same time, the pyramid structure is used to increase the kernel size from bottom to top and group convolution to reduce the kernel depth to capture multi-scale information, so as to realize the effective and accurate detection of the number and position of inflection points in the palm vein image.
[0022] Please refer to Figure 3 , in a specific embodiment of the present invention, the adaptive noise suppression module processes the shallow features output by the shallow feature extraction module according to the following steps (1) to (6), Figure 3 where F in att corresponds to this step.
[0023] (1) Divide the shallow feature X into G sub-features X i , i∈{1,2,…,G}. In this step, the shallow feature X∈R C ×H×W , and perform a high-order decomposition into G sub-feature modules to achieve a more refined encoding and representation of the feature space. This operation aims to mine richer local information through hierarchical segmentation of the original features, while enhancing the model's adaptability and recognition ability to diverse patterns. The obtained X i ∈R C / G×H×W .
[0024] (2) Use the GAP algorithm to solve the sub-feature X i Processing, get the first feature X 1i . GAP (Graph Attention Pooling) is a graph pooling algorithm based on the attention mechanism, which is used in graph neural networks (GNN) to efficiently compress graph structures and retain key information. Its core idea is to select the most important nodes or substructures in the graph through attention scores, thereby generating a more compact graph representation. In a specific embodiment of the present invention, X 1i =X i ·GAP(X i ), that is, using the GAP algorithm to i GAP(X i ), and then use sub-feature X i With GAP(X i ) to obtain the first feature X 1i .
[0025] (3) Use the constant ε to calculate the first feature X 1i Normalize and get the second feature X 2i In a specific embodiment of the present invention, the first feature X is calculated 1i The average value μ i and variance σ i Then, the constant ε can be introduced to avoid the second characteristic X 2i The denominator is 0, and the specific formula is as follows: .
[0026] (4) Use learnable parameters γ and β to adjust the second feature X 2i Transform and get the third feature X 3i The purpose of this step is to ensure that unit conversion can be performed. In a specific embodiment of the present invention, X 3i =γX 2i +β.
[0027] (5) Use sub - feature X i and the Sigmoid function to process the third feature X 3i to obtain the fourth feature X 4i . Use the activation function to generate an attention factor, scale according to the original sub - feature to generate a refined feature. After optimizing each group of sub - features, they will be integrated into the finally optimized feature to keep the feature depth consistent with the input. In a specific embodiment of the present invention, X 4i = X i ·Sigmoid(X 3i ).
[0028] (6) Concatenate the fourth feature X 4i to obtain the output feature Y of the adaptive noise suppression unit. It can be understood that for each sub - feature X i , after being processed through the above steps (2) - (5), a corresponding X 4i can be obtained. Finally, G X 4i will be obtained. Finally, only these fourth features need to be concatenated along the original channel dimension to obtain the overall feature Y, which is expressed by the formula: Y = X 41 +X 42 +…+X 4G .
[0029] Through the processing of the above steps (1) - (6), the image background noise can be suppressed to a great extent, which is beneficial to the accurate extraction of features.
[0030] Please refer to Figure 1 , in a specific embodiment of the present invention, the coupled feature enhancement unit includes a learnable graph neural unit and a multi - scale information fusion unit. The learnable graph neural unit uses the graph structure to enhance the spatial dependence relationship between each channel, and the multi - scale information fusion unit uses the pyramid structure to increase the kernel size from bottom to top and reduce the kernel depth (connectivity) through grouped convolution, that is, to capture multi - scale information through a dual - guided pyramid operation.
[0031] Please refer to Figure 4In a specific embodiment of the present invention, the expression of the learnable graph neural unit is as follows: Z=Y•∙ReLU(Conv1d(A•GAP(Y))); A=A0×A1+A2; A0=Softmax(Conv1d(GAP(Y))); where Z is the output feature of the learnable graph neural unit, Y is the output feature of the adaptive noise suppression unit, GAP is the GAP algorithm, Conv1d is a one-dimensional convolution layer, ReLU is an activation function, Softmax is normalization, and A0, A1, and A2 are learnable adjacency matrices. In this embodiment, three adjacency matrices A0, A1, and A2 are used to represent the connections between vertices. A0 is a diagonal matrix based on self-attention, which is used to suppress useless features and is generated by a one-dimensional convolution layer Conv1d and a Softmax function; A1 contains the vertex itself and needs to be normalized; A2 can customize unique dependencies for different feature vertices, which needs to be optimized during the training process to improve the response between different vertices in order to obtain more adaptive global information.
[0032] See also Figure 5 In a specific embodiment of the present invention, the multi-scale information fusion unit includes a first branch and a second branch in parallel. The first branch is as follows Figure 5 As shown on the left side of the figure, it consists of a cascaded first convolutional layer, multiple parallel group convolutions, and a second convolutional layer; the second branch is shown in Figure 5 As shown on the right side of the figure, it consists of a cascaded adaptive average pooling layer, the first convolution layer, multiple parallel group convolutions, the second convolution layer, and the first upsampling layer; finally, the output features of the first branch and the second branch are connected and then passed through the third convolution layer and the second upsampling layer in sequence to obtain the output features of the multi-scale information fusion unit.
[0033] For the first branch, it has a smaller receptive field and is responsible for tiny objects. It first uses 1×1 convolution to reduce the number of channels to 512, then aggregates several layers with different kernel sizes, and the number of groups G makes the kernels have different connectivity. Each convolution block is followed by a batch normalization layer and a ReLU activation layer.
[0034] The second branch has a similar structure to the first branch. To capture the features of large objects in a global perspective, an adaptive average pooling operation is used at the top to reduce the spatial size of the feature map to 9×9, and the feature map is upsampled to the same resolution as the input by bilinear interpolation at the bottom.
[0035] After the two branches, the features output by the two branches are concatenated, followed by a standard convolutional layer of size 3×3. Finally, we upsample the feature map to the size of the original image. The whole process can improve the robustness of the model to scale changes.
[0036] In a specific embodiment of the present invention, multiple parallel grouped convolutions include four grouped convolutions: the first grouped convolution has a kernel size of 9×9 and the number of groups is 16; the second grouped convolution has a kernel size of 7×7 and the number of groups is 8; the third grouped convolution has a kernel size of 5×5 and the number of groups is 4; the fourth grouped convolution has a kernel size of 3×3 and the number of groups is 1.
[0037] In a specific embodiment of the present invention, a density map is generated through a deep convolutional structure decoding, and the process is as follows: in order to increase the resolution of the feature map, an upsampling operation is adopted; by using layer-by-layer upsampling and convolution, the image size is maintained, and the number of channels is reduced, which not only retains the high-level semantic features but also gradually restores the spatial resolution. This multi-layer convolutional decoder not only has a powerful non-linear mapping ability but also can perform information fusion at multiple scales, thereby improving the accuracy and robustness in the decoding process, especially showing outstanding performance in the accurate recognition and counting of objects in complex environments.
[0038] In a specific embodiment of the present invention, in order to generate a more accurate density map and the accuracy of the inflection point prediction, the loss function during the training of the inflection point detection network is as follows: K Overall =K Bayesian +λK Count ; where K Overall is the total loss function, K Bayesian is the Bayesian uncertainty loss, K Count is the counting error loss, and λ is a hyperparameter used to balance the weights of the two losses. The reliable supervision method improved by the Bayesian uncertainty loss and the counting error loss is used to learn the density probability, and then the counting expectation at each annotation is calculated, which can alleviate the non-uniformity of the density distribution to a certain extent.
[0039] In order to verify the effectiveness of the method in the present invention, the inflection point detection network in the present invention is compared with other mainstream networks. When comparing, the mean absolute error (MAE) and the root mean square error (RMSE) are used to evaluate the performance of the proposed method. The comparison results are shown in the following table.
[0040] Table 1 Comparison table of the effects of the network in the present invention and various networks on testing the same group of datasets
[0041] As can be seen from the above table, the inflection point detection network proposed in the present invention exceeds the mainstream networks in terms of both MAE and RMSE, demonstrating excellent prediction accuracy, being applicable to the tasks of palm vein image datasets, and showing an absolute advantage among multiple traditional and emerging networks.
[0042] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A palm vein breakpoint detection method with adaptive noise suppression and synergistic coupling feature enhancement, characterized in that: include: Acquire a palm vein image to be processed; The palm vein image is input into a trained breakpoint detection network to obtain the number and position information of the palm vein breakpoints. The breakpoint detection network includes a cascade of: A shallow feature extraction module, used for preprocessing the palm vein image; Adaptive noise suppression module, used to suppress background noise to improve the signal-to-noise ratio of the image; The coupling feature enhancement module is used to enhance the position information between each channel to obtain multi-scale information; as well as The decoder is used to process the enhanced features to obtain the number and position information of the palm vein inflection points.
2. The palm vein breakpoint detection method based on adaptive noise suppression and collaborative coupling feature enhancement according to claim 1, characterized in that: The adaptive noise suppression module processes the shallow features output by the shallow feature extraction module according to the following steps: Divide the shallow feature X into G sub-features X i , i∈{1,2,…,G}; The GAP algorithm is used to i Processing, get the first feature X 1i ; Use the constant ε to adjust the first feature X 1i Normalize and get the second feature X 2i ; The second feature X is conditioned on the learnable parameters γ and β 2i Transform and get the third feature X 3i ; Using the sub-feature X i And the Sigmoid function for the third feature X 3i After processing, the fourth feature X is obtained 4i ; Splice the fourth feature X 4i , and obtain the output feature Y of the adaptive noise suppression unit.
3. The palm vein breakpoint detection method based on adaptive noise suppression and collaborative coupling feature enhancement according to claim 2 is characterized in that: The relationship between the features is expressed as follows: X 1i =X i ·GAP(X i ); ; X 3i =γX 2i +b; X 4i =X i ·Sigmoid(X 3i ); Y=X 41 +X 42 +…+X 4G 。 4. The palm vein breakpoint detection method with adaptive noise suppression and synergistic coupling feature enhancement according to claim 1, characterized in that: The coupled feature enhancement unit includes a learnable graph neural unit and a multi-scale information fusion unit. The learnable graph neural unit uses a graph structure to enhance the spatial dependency between each channel, and the multi-scale information fusion unit uses a dual-guided pyramid operation to capture multi-scale information.
5. The palm vein breakpoint detection method with adaptive noise suppression and synergistic coupling feature enhancement according to claim 4, characterized in that: The expression of the learnable graph neural unit is as follows: Z=Y·ReLU(Conv1d(A·GAP(Y))); A=A0×A1+A2; A0=Softmax(Conv1d(GAP(Y))); Wherein, Z is the output feature of the learnable graph neural unit, Y is the output feature of the adaptive noise suppression unit, GAP is the GAP algorithm, Conv1d is the one-dimensional convolutional layer, ReLU is the activation function, Softmax is normalization, and A0, A1 and A2 are learnable adjacency matrices.
6. The palm vein breakpoint detection method with adaptive noise suppression and synergistic coupling feature enhancement according to claim 4, characterized in that: The multi-scale information fusion unit includes a parallel first branch and a second branch, the first branch is composed of a cascaded first convolution layer, multiple parallel grouped convolutions, and a second convolution layer, the second branch is composed of a cascaded adaptive average pooling layer, a first convolution layer, multiple parallel grouped convolutions, a second convolution layer, and a first upsampling layer, and the output features of the first branch and the second branch are connected and then sequentially passed through a third convolution layer and a second upsampling layer to obtain the output features of the multi-scale information fusion unit.
7. The palm vein breakpoint detection method with adaptive noise suppression and synergistic coupling feature enhancement according to claim 6, characterized in that: The multiple parallel grouped convolutions include four grouped convolutions: The convolution kernel size of the first group convolution is 9×9 and the number of groups is 16; The convolution kernel size of the second group convolution is 7×7 and the number of groups is 8; The convolution kernel size of the third group convolution is 5×5 and the number of groups is 4; The convolution kernel size of the fourth group convolution is 3×3 and the number of groups is 1.
8. The palm vein breakpoint detection method with adaptive noise suppression and synergistic coupling feature enhancement according to claim 1, characterized in that: The loss function during the training of the inflection point detection network is as follows: K Overall =K Bayesian +λK Count ; Among them, K Overall is the total loss function, K Bayesian is the Bayesian uncertainty loss, K Count is the counting error loss, and λ is a hyperparameter used to balance the weights of the two losses.
Citation Information
Patent Citations
Sub-mesenteric artery blood vessel reconstruction method based on MIP sequence
CN114897780A
Automatic substrate glass surface defect detection method and system based on machine vision
CN119006469A
Non-contact palm vein anti-attack method based on lightweight network
CN119169674A