Multi-scale Transformer-based Venous Recognition Neural Network Model, Method and System
Through the multi-scale Transformer neural network model, the problem of limited image recognition performance of venous recognition models for different scales is solved. Through the scale transformation and segmentation modules, the accuracy and robustness of venous recognition are improved.
Patent Information
- Application Number
- CN202211591327.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-12-12
AI Technical Summary
When the existing venous recognition model processes venous images of different scales, it ignores the relationship between different sizes, resulting in limited recognition performance.
Using a neural network model based on multi-scale Transformer, the venous pictures are divided into multiple sub-graphs of different sizes through the scale transformation module, and divided into position blocks through the multi-scale segmentation module and linear embedding module. Combining the multi-scale Transformer module and the block convolution module, attention information between different scales and positions is learned.
It improves the recognition performance of the venous recognition model, can calculate relevant information at different scales, enhances the perfection of global information and the friendliness of feature extraction, and improves the recognition accuracy.
Smart Images

Figure CN116229230B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biometric recognition, and particularly relates to a vein recognition neural network model, method and system based on multi-scale Transformer. Background Art
[0002] Many existing classification models can effectively extract feature information such as vein patterns from vein images. Usually, only vein images at one scale are selected for feature extraction and then classification. However, at the vein recognition terminal, due to the non-contact vein acquisition method, the acquired vein images may have different sizes due to different distances from the acquisition terminal. However, a model trained with a unified size is sensitive to images of different scales. Although the image size can be adjusted to one size, this will ignore the mutual relationship between the same image at different sizes, and there is no mutual relationship between different sizes to affect or supplement the high-level semantics of different information positions in the same size, resulting in limited recognition performance of the model. Summary of the Invention
[0003] One of the objectives of the present invention is to provide a vein recognition neural network model based on multi-scale Transformer to solve the technical problem that the existing model ignores the correlation between various sizes, resulting in limited recognition performance.
[0004] The vein recognition neural network model based on multi-scale Transformer in the present invention includes a scale transformation module, a multi-scale segmentation module, a linear embedding module, and a multi-scale Transformer module connected in sequence;
[0005] The scale transformation module is used to scale the vein image I into sub-images I of multiple different sizes n , where n = 1, 2... N. Let h0, w0, and c be the height, width, and number of channels of the vein image I, and h n , w n , c be the height, width, and number of channels of the sub-image I n , there are: h1 > h2 >... h n , w1 > w2 >... w n ;
[0006] The multi-scale segmentation module is used to segment each sub-image I n into position blocks (Patches) of size P×L. For the largest I1 in size, non-overlapping segmentation is adopted; for the remaining sub-images, overlapping segmentation is adopted, and this overlapping segmentation makes each sub-image be segmented into the same number of position blocks of size P×L;
[0007] And, flatten each position block of each sub-image into a sequence P with a length of C = PLcm,n , m = 1, 2…M, n = 1, 2…N; where M is the number of position blocks, and there is:
[0008] M = H × W
[0009]
[0010] The linear embedding module is used to map each sequence P m,n to a feature (Token) T of length D m,n , and splice the features of each sub into a one-dimensional feature sequence I t,n ;
[0011] And, perform learnable position encoding on each feature sequence I t,n respectively;
[0012] And, add a learnable scale embedding sequence E t,n with the same form as the feature sequence I scale , and jointly form a feature sequence set I with the feature sequences of each sub-graph TE ;
[0013] The multi-scale Transformer module includes a scale self-attention calculation part and a spatial self-attention calculation part connected in sequence;
[0014] The scale self-attention calculation part is used to calculate the self-attention weights between position blocks corresponding to the same position on different sub-graphs based on the input feature sequence set I TE , called scale self-attention weights, and obtain the output X TE based on the feature map block set I new ;
[0015] The spatial self-attention calculation part is used to calculate the self-attention weights between position blocks corresponding to different positions on the same sub-graph based on the input X new , called spatial self-attention weights, and obtain the output X new based on X (1) .
[0016] Furthermore, the multi-scale Transformer module further includes a multi-layer perceptron part connected after the spatial self-attention calculation part, including a normalization layer (LN) and a multi-layer perceptron sub-module (MLP) connected in sequence. A Droppath mechanism and a residual connection are introduced in this part, and the output Y of this part is the output of the multi-scale Transformer module.
[0017] Further, the scale self-attention calculation part includes a normalization layer (LN), a scale self-attention sub-module (Scale Attention), and a feed-forward network module (FFN) connected in sequence. A Droppath mechanism and a residual connection are introduced after the feed-forward network module in this part;
[0018] Among them, the scale self-attention sub-module is used to take a total of N + 1 features corresponding to the same patch position in each feature sequence / scale embedding sequence in the input as a group of input sequences, and calculate the self-attention weights between the features in each group of input sequences.
[0019] Further, the spatial self-attention calculation part includes a normalization layer (LN) and a spatial self-attention sub-module (Space Attention) connected in sequence. A Droppath mechanism and a residual connection are introduced after the spatial self-attention sub-module in this part;
[0020] Among them, the spatial self-attention sub-module is used to take the feature sequences corresponding to the same sub-graph or scale embedding patch in the input as a group of input sequences, and calculate the self-attention weights between the feature sequences in each input sequence.
[0021] Further, the self-attention calculation in the multi-scale Transformer module is multi-head self-attention calculation.
[0022] Further, it also includes a patch convolution module;
[0023] At least one of the multi-scale Transformer modules is connected to the patch convolution module to form a multi-scale Transformer-convolution module;
[0024] If the multi-scale Transformer-convolution module contains multiple multi-scale Transformer modules, the multi-scale Transformer modules are cascaded in sequence, and the last multi-scale Transformer module is connected to the patch convolution module;
[0025] When the output Y of the multi-scale Transformer module enters the multi-scale Transformer-convolution module, it transforms into the form of a two-dimensional feature matrix set;
[0026] The patch convolution module includes a first particle convolution stack block, a second particle convolution stack block, and a downsampling layer connected in sequence;
[0027] The first particle convolution stack block is used to, on the one hand, make the input Y pass through a concatenated fully connected layer and a two-dimensional convolution layer with a convolution kernel of 1×1 and a stride of 1 to obtain the output Y (1), where the number of output channels of both the fully connected layer and the convolutional layer is γ < 1 times the number of channels of the input Y;
[0028] Let the input Y first pass through a fully connected - DW convolutional sub - module, in which a DW (Depth - wise) convolutional layer (DW - conv) with a convolution kernel of k×k and a stride of 1 is added on the basis of the fully connected layer, and then pass through a two - dimensional convolutional layer with a convolution kernel of 1×1 and a stride of 1 to obtain the output Y (2) , and the number of output channels of both the fully connected - DW convolutional sub - module and the two - dimensional convolutional layer is γ < 1 times the number of channels of the input Y;
[0029] And, connect the said Y (1) 、Y and Y (3) sequentially in the channel dimension to obtain the output Z;
[0030] The second grain convolution stack block is used to obtain the output Z based on the input Z in the same way as the first grain convolution stack block (1) ;
[0031] The downsampling layer is used to perform downsampling on Z based on a 2×2 convolution kernel (1) , and the number of input channels is half of the number of input channels.
[0032] Furthermore, it includes multiple cascaded multi - scale Transformer - convolutional modules;
[0033] Among them, the neural network form in the subsequent multi - scale Transformer - convolutional module adapts to the output form of the previous multi - scale Transformer - convolutional module;
[0034] And, the number of heads of the multi - head self - attention calculation in the subsequent multi - scale Transformer - convolutional module is 2γ + 1 times that in the previous multi - scale Transformer - convolutional module.
[0035] Furthermore, it includes four cascaded multi - scale Transformer - convolutional modules.
[0036] Another object of the present invention is to provide a vein recognition method, including:
[0037] Step 1: Obtain a vein picture;
[0038] Step 2: Input the vein picture into the aforementioned trained vein recognition neural network model based on multi - scale Transformer to obtain a recognition result.
[0039] Another object of the present invention is to provide a vein recognition system, including a vein picture acquisition module and a recognition module;
[0040] The venous image acquisition module is used to acquire venous images to be recognized;
[0041] The recognition module is internally deployed with the aforementioned trained venous recognition neural network model based on multi-scale Transformer, which is used to obtain recognition results through the venous recognition neural network model according to the input venous images.
[0042] Furthermore, it further includes a training module, which is used to acquire venous image samples for training the venous recognition neural network model;
[0043] And use the venous image samples to train the venous recognition neural network model based on multi-scale Transformer, and then update the parameters in the neural network model.
[0044] The principle and beneficial effects of the present invention are as follows:
[0045] Different from the existing deep learning venous recognition model based on CNN, the present invention proposes a venous recognition neural network model based on Transformer. Among them, the present invention makes a multi-scale improvement to the existing ViT (Vision Transformer) model, transforms the input image into an atlas including sub-images of different sizes through size transformation, and each sub-image is divided / overlapped into position blocks of the same number and the same size. Then, the feature sequences of the blocks at the same position between different scales are used to calculate the connection between the positions at this place on the sub-images of different sizes, so that the model can learn the connection of sizes and indirectly eliminate the sensitivity to different sizes. The model first learns the attention information between different scales at the same position by position, and then learns the attention information between different positions under the traditional unified size by sub-image. At this time, it carries rich inter-scale information, making the global information more perfect and the extracted features more friendly. In order to meet the requirement of being able to calculate the relevant information of different positions at the same time under different scales, a feature map block sequence for memorizing different scale information is added in this model, which has the same form as the feature map blocks of each sub-image. The present invention affects the classification result through the relationship between the same position between different scales and different positions between the same scale, and improves the recognition performance of the model.
[0046] In addition, each image is converted into multiple images of different scales to expand the training samples, thereby improving the recognition performance.
[0047] Since more attention is paid to global attention information in self-attention calculation, in some embodiments of the present invention, a new Patch ConvNN Block is additionally added after the multi-scale Transformer module to extract local information, induce bias, and perform downsampling. The convolution operator allows learning local features by using a local receptive field and sharing weights, while the self-attention mechanism in the Transformer can capture global features. The combination of the two modules can form a complementarity to improve the accuracy of vein recognition.
[0048] In addition, neural network models based on the Transformer usually contain a relatively large number of parameters to be trained. However, in the vein recognition task, there are not a large number of training samples, which may result in the ineffective utilization of the model's capacity. The model is affected by the training conditions and thus has limited improvement in the recognition accuracy of the vein recognition task in practical applications. The strategy of incorporating convolution into the Transformer in the embodiments of the present invention can improve the recognition accuracy from another aspect, which is of great practical significance for the vein recognition task without a large number of training samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic connection structure diagram of the scale transformation module, multi-scale segmentation module, linear embedding module, and multi-scale Transformer module in the embodiments of the present invention.
[0050] Figure 2 It is a schematic diagram of the non-overlapping / overlapping segmentation process for subgraphs of different scales in the embodiments of the present invention.
[0051] Figure 3 It is a schematic diagram of the implementation position parameters and scale feature sequence embedding of the linear embedding module in the embodiments of the present invention.
[0052] Figure 4 It is a schematic logical block diagram of the multi-scale Transformer module in the embodiments of the present invention.
[0053] Figure 5 It is a schematic logical block diagram of the multi-scale Transformer-convolution module in the embodiments of the present invention.
[0054] Figure 6 It is a schematic logical block diagram of the Patch ConvNN Block in the embodiments of the present invention.
[0055] Figure 7 It is a schematic logical block diagram of the first / second stack sub-module in the embodiments of the present invention.
[0056] Figure 8Schematic logic block diagram of the vein recognition neural network model based on multi-scale Transformer in the embodiments of the present invention.
[0057] Figure 9 Flowchart of the vein recognition method in the embodiments of the present invention.
[0058] Figure 10 Schematic block diagram of the vein recognition system in the embodiments of the present invention.
[0059] Figure 11 Schematic block diagram of the vein recognition system in another embodiment of the present invention. Detailed implementation manners
[0060] In this part, data in matrix / vector forms such as input / output pictures, patches, features, etc. are represented in the form of four-dimensional tensors (b, h, w, c). Among them, the first dimension b is the number of matrices / vectors in the set, also known as Batch Size. The second dimension h can be understood as the height dimension or row dimension. The third dimension w can be understood as the width dimension or column dimension. The fourth dimension c can be understood as the channel dimension. For the convenience of understanding, in this part, a single picture (the first dimension b = 1) is used as an example for input. However, in actual applications, the input picture set can be a picture set with the number of pictures being B. In this case, the first dimension of all the following four-dimensional tensors needs to be expanded by B times.
[0061] The vein recognition neural network model based on multi-scale Transformer in this embodiment includes a scale transformation module, a multi-scale segmentation module, a linear embedding module, and a multi-scale Transformer module connected in sequence; the connection manners of these modules are as Figure 1 shown.
[0062] Among them, the scale transformation module (Resize) is used to transform the vein picture I in the form of (1, h0, w0, c) into multiple sub-pictures in different sizes in the form of I n :(1, h n , w n , c), n = 1, 2... N, where h0, w0, and c are the height, width, and number of channels of the vein picture I respectively, and h n , w n , c are the height, width, and number of channels of the sub-picture I n respectively, and h1 > h2 >... h n , w1 > w2 >... w n; For example, for a venous image I in the form of (1, 200, 200, 3), N = 3 sub - graphs of different sizes are obtained through transformation, namely the first sub - graph I1 in the form of (1, 224, 224, 3), the second sub - graph I2 in the form of (1, 168, 168, 3), and the third sub - graph I3 in the form of (1, 112, 112, 3), which form a sub - graph set (Batch) and are input into the subsequent model. The size (Batch size) of this sub - graph set is N = 3.
[0063] The multi - scale segmentation module is used to segment each sub - graph I n into patches of size P×L, as Figure 2 shown. Among them, for the largest - sized I1, non - overlapping segmentation is adopted; for the remaining sub - graphs, overlapping segmentation is adopted. This overlapping segmentation makes each sub - graph be segmented into the same number of position patches of size P×L;
[0064] And, flatten each position patch of each sub - graph into a sequence P of length C = PLc m,n , m = 1, 2…M, n = 1, 2…N; where M is the number of position patches, and there is:
[0065] M = H×W
[0066]
[0067] The linear embedding module (linear Embeding) is used to map each sequence P of length C through a learnable mapping matrix m,n to a feature (Token) T of dimension D m,n , and splice the features of each sub - graph into a feature sequence I in the form of (1, 1, M, D) t,n , and the feature sequences of each sub - graph form a feature sequence set I T : (N, 1, M, D), thereby realizing the shallow - layer feature extraction of each position patch;
[0068] And, perform learnable position encoding on the feature sequences I of each sub - graph t,n respectively. In this embodiment, one - dimensional position encoding is adopted but not limited to it, that is, make I T superimposed with a learnable position parameter set E P : (N, 1, M, D);
[0069] And, add a learnable scale embedding sequence E in the form of (1, 1, M, D) scale , and form a feature sequence set I in the form of (N + 1, 1, M, D) with I T ; TE; To facilitate the input of the Transformer module, the feature sequences in I TE are concatenated into a large feature sequence I TE : (1, 1, (N + 1) × M, D).
[0070] In summary, the feature sequence I TE can be expressed as follows:
[0071]
[0072]
[0073]
[0074] Specifically, the neural network model in this embodiment is set to equivalently implement non - overlapping or overlapping segmentation of each sub - graph and obtain mapped features with the same sequence length from different sub - graphs through a two - dimensional convolutional layer with different strides (Stride) in cooperation with padding operations.
[0075] Taking the above - mentioned input sub - graph set as an example, let the Patch size be 8×8, the convolutional strides corresponding to different sub - graphs be 8, 6, and 4 respectively, and the padding bits be 0, 2, and 4 respectively; the three sub - graphs pass through their respective two - dimensional convolutional layers. The convolutional kernel size corresponding to I1 is 8×8, the stride is 8, the padding is 0, the input channel is c = 3, and the output channel is D = 64, corresponding to non - overlapping segmentation; the convolutional kernel size corresponding to I2 is 8×8, the stride is 6, the padding is 2, the input channel is c = 3, and the output channel is D = 64; the convolutional kernel size corresponding to I3 is 8×8, the stride is 4, the padding is 4, the input channel is c = 3, and the output channel is D = 64. After two - dimensional convolution, flattening two dimensions gives three feature sequences I t,1 、I t,2 and I t,2 .
[0076] In this embodiment, the embedding process of the position parameter and scale parameter sequences is as Figure 3 shown, but not limited to this; first, the feature sequences are concatenated into a large feature sequence in the form of (1, 1, 3·28·28, 64). At this time, the transformation that meets the model requirements for different sizes of the same image is completed; on this basis, first add the learnable position parameter sequence E P : (1, 1, 3·28·28, 64), and then connect the learnable scale parameter block E scale : (1, 1, 28·28, 64), to obtain the feature block I TE:(1, 1, 4·28·28, 64) is used as the input X of the multi-scale Transformer module, that is, a feature sequence composed of 4 groups, each group having 784 features with a dimension of 64, and is used to input the multi-scale Transformer module (MSU-TransformerBlock).
[0077] The multi-scale Transformer module includes a scale self-attention calculation part and a spatial self-attention calculation part connected in sequence;
[0078] The scale self-attention calculation part is used to calculate the self-attention weights between position blocks corresponding to the same position on different subgraphs based on the input X, called scale self-attention weights, and then obtain an output X of the form (1, 1, (N + 1)×M, D) based on the input X and the scale self-attention weights new ;
[0079] The spatial self-attention calculation part is used to calculate the self-attention weights between position blocks corresponding to different positions on the same subgraph based on the input X new and is called spatial self-attention weights, and then obtain an output X of the form (1, 1, (N + 1)×M, D) based on X new and the spatial self-attention weights (1) .
[0080] As Figure 4 shown, in some embodiments, the multi-scale Transformer module further includes a multi-layer perception part connected after the spatial self-attention calculation part, including a normalization layer (LN) and a multi-layer perception sub-module (MLP) connected in sequence. A Droppath mechanism and a residual connection are introduced in this part, and the output Y of this part is the output of the multi-scale Transformer module, in the form of (1, 1, (N + 1)×M, D).
[0081] As Figure 4 shown, in these embodiments, the scale self-attention calculation part includes a normalization layer (LN), a scale self-attention sub-module (ScaleAttention), and a feed-forward network module (FFN) connected in sequence. A Droppath mechanism and a residual connection are introduced after the feed-forward network module in this part; among them, the scale self-attention sub-module is used to take a total of N + 1 features corresponding to the same graph block position in each feature sequence / scale embedding sequence in the input as a group of input sequences, and calculate the self-attention weights between the features in each group of input sequences.
[0082] Specifically, taking the previous input as an example, the input X is transformed into X with a dimension scale:(28·28, 1, 4, 64), that is, four features corresponding to the same patch position in the four feature sequences (including the feature sequences of the three sub - graphs and the scale parameter sequence) form an input sequence. Thus, 28 * 28 input sequences are obtained, and multi - head self - attention calculations are performed separately to obtain the self - attention weights between different scales at the same position;
[0083] X scale First, it passes through the normalization layer (LN), then through the multi - head scale self - attention sub - module (ScaleAttention) for self - attention calculation on the scale. Subsequently, it is processed by the Droppath mechanism (not shown in the figure) and a feed - forward neural network module including a linear layer in sequence, and the output residual R multi-scale maintains the same form as the input X, that is, the form is (1, 1, 28·28·4, 64). Then, through the residual connection, R multi-scale is added to X to obtain the output X of this part. new .
[0084] As Figure 4 shown, in these embodiments, the spatial self - attention calculation part includes a normalization layer (LN) and a spatial self - attention sub - module (Space Attention) connected in sequence. In this part, a Droppath mechanism and a residual connection are introduced after the spatial self - attention sub - module; among them, the spatial self - attention sub - module is used to calculate the self - attention weights between the feature sequences in each input sequence with the feature sequences corresponding to the same sub - graph or scale - embedded patch in the input as a group of input sequences.
[0085] Specifically, taking the previous input as an example, the X new output by the scale self - attention calculation part is transformed into X spatial :(4, 1, 28·28, 64), that is, 28·28 feature sequences (after spatial self - attention calculation, with residuals added to the original feature sequences) in the same feature patch (including the feature patch of the three sub - graphs and the scale parameter patch) are used as 1 input sequence. Thus, 4 input sequences are obtained, and multi - head self - attention calculations are performed separately to obtain the self - attention weights between different positions at the same scale;
[0086] X spatial passes through the normalization layer (LN), then through the multi - head spatial self - attention calculation module (SpaceAttention) for self - attention calculation on the spatial position. Subsequently, the residual R is obtained through the Droppath mechanism. spatial , similarly, in order to be consistent with the input X new , it is transformed into the form of R spatial :(1, 1, 28·28·4, 64), and then through the residual connection, Rspatial Add the output X of this part to X new ; (1) ;
[0087] In these embodiments, the self-attention calculation in the multi-scale Transformer module is multi-head self-attention calculation, but is not limited thereto.
[0088] Taking the aforementioned input as an example, the output X of the spatial self-attention calculation part (1) In this part, after passing through the normalization layer (LN), the multi-layer perceptron sub-module (MLP) and the Droppath mechanism, the residual R is obtained x (1) Add R to X (1) to obtain the output Y: (1, 1, 28·28·4, 64).
[0089] In summary, the output Y of the multi-scale Transformer module can be expressed as follows:
[0090] Y = X (1) + Droppath(MLP(LN(X (1) )))
[0091]
[0092]
[0093] where Droppath(·) represents the processing of the Droppath mechanism, represents the spatial self-attention calculation of multiple heads, d represents the number of heads, LN(·) represents the normalization layer calculation, FFN(·) represents the feed-forward neural network calculation, and MLP(·) represents the multi-layer perceptron calculation;
[0094] In some embodiments, the model further includes a patch convolution module (Patch ConvNN Block);
[0095] As Figure 5 shown, at least one multi-scale Transformer module is connected to the patch convolution module to form a multi-scale Transformer-convolution module;
[0096] If the multi-scale Transformer-convolution module includes multiple multi-scale Transformer modules, the multi-scale Transformer modules are cascaded in sequence, and the last multi-scale Transformer module is connected to the patch convolution module;
[0097] Figure 5The multi-scale Transformer-convolution module in
[0098] When the output Y of the multi-scale Transformer module enters the patch convolution module, its form is transformed into ((N + 1), H, W, D), which can be understood as changing from the form of a set of feature sequences (a set of one-dimensional sequences with features as elements) to the form of a set of feature maps (a set of two-dimensional matrices with feature sequences as elements).
[0099] As Figure 6 shown, the patch convolution module includes a first granular stack block (Granular StackBlock1), a second granular stack block (Granular Stack Block2), and a downsampling layer (Down sample Layer) connected in sequence;
[0100] As Figure 7 shown, the first granular convolution stack block is used to, on the one hand, let the input Y pass through a concatenated fully connected layer and a two-dimensional convolution layer with a convolution kernel of 1×1 and a stride of 1 to obtain the output Y (1) , and the output channel numbers of both the fully connected layer and the convolution layer are γ < 1 times the channel number of the input Y, Y (1) can be expressed as follows:
[0101]
[0102]
[0103] Among them, ReLU(·) represents the ReLU activation function, * represents the convolution operation, W1 is the parameter matrix of the fully connected layer. In the formula, the fully connected layer is equivalent to performing a convolution operation with a 1×1 convolution kernel and a stride of 1 on the input and then passing through the ReLU activation function. W2 is the parameter matrix of the two-dimensional convolution layer with a convolution kernel of 1×1 and a stride of 1, and γ is the reduction factor.
[0104] On the other hand, let the input Y first pass through a fully connected-DW convolution sub-module, in which a depth-wise (DW) convolution layer (DW-conv) with a convolution kernel of k×k (exemplary k = 3 in the figure) and a stride of 1 is added on the basis of the fully connected layer, and then pass through a two-dimensional convolution layer with a convolution kernel of 1×1 and a stride of 1 to obtain the output Y (2) , and the output channel numbers of both the fully connected-DW convolution sub-module and the two-dimensional convolution layer are γ < 1 times the channel number of the input Y, Y (2) can be expressed as follows:
[0105]
[0106]
[0107] Among them, W d is the parameter matrix of the DW convolutional layer.
[0108] And, the Y (1) 、Y and Y (2) are sequentially connected in the channel dimension to obtain the output Z. Here, this operation is called the grain convolution stack, and Z can be expressed as follows:
[0109]
[0110] Among them, Contact(·) represents the connection operation in the channel dimension.
[0111] The second grain convolution stack block is used to obtain the output based on the input Z in the same way as the first grain convolution stack block
[0112] The downsampling layer is used to perform convolution on Z with a 2×2 convolutional kernel (1) with a stride of 2, and the number of output channels is half of the number of input channels, thereby realizing downsampling, and its output is
[0113] Specifically, taking the previous input as an example, the form of the output Y of the multi-scale Transformer module is transformed into the form of a feature map set Y: (4, 28, 28, 64), and is input into the first grain convolution stack block. In this module, on the one hand, Y passes through a concatenated fully connected layer and a two-dimensional convolutional layer with a convolutional kernel of 1×1 and a stride of 1 to obtain the output Y (1) : (4, 28, 28, 32). The number of output channels of both the fully connected layer and the two-dimensional convolutional layer is γ = 0.5 times the number of input channels of Y; on the other hand, Y passes through a concatenated fully connected layer, a DW convolutional layer with a convolutional kernel of 3×3 and a stride of 1, and a two-dimensional convolutional layer with a convolutional kernel of 1×1 and a stride of 1. The number of output channels of the fully connected layer, the DW convolutional layer, and the two-dimensional convolutional layer are all γ = 0.5 times the number of input channels of Y, and the output Y (2) ; : (4, 28, 28, 32), and then Y (1) 、Y and Y (2) are connected in the channel dimension to become the output Z: (4, 28, 28, 128); compared with the input Y, Z (1) doubles in the channel dimension; the output Z of the first grain convolution stack block (1) is input into the second grain convolution stack block and the same operation is performed later. In this way, the output Z of the second grain convolution stack block (1) has the form of (4, 28, 28, 256), and the channel dimension doubles again; then Z (1)In the downsampling layer, a two-dimensional convolution with a 2×2 convolution kernel, a stride of 2, 256 input channels, and 128 output channels is performed to obtain the output Z. (2) :(4, 14, 14, 128).
[0114] As Figure 8 shown, in some embodiments, the model includes multiple sets of cascaded multi-scale Transformer-convolution modules;
[0115] wherein, the neural network form in the subsequent multi-scale Transformer-convolution module is adapted to the output form of the previous multi-scale Transformer-convolution module;
[0116] And, the number of heads in the multi-head self-attention calculation in the subsequent multi-scale Transformer-convolution module is 2γ + 1 times that in the previous multi-scale Transformer-convolution module.
[0117] Figure 8 Four sequentially cascaded modules are exemplarily given in, and the input of each module is the output of the previous module. It should be noted that each module has only one tile convolution module, but there can be multiple levels of cascaded multi-scale Transformer modules. It is not difficult to find that the number of output channels doubles after each pass through the tile convolution module, which poses a challenge to the self-attention calculation in the next layer. Therefore, the number of heads in the scale / space attention calculation module in the multi-scale Transformer module in the next-level module increases with the increase in the number of channels, thereby improving the accuracy of the self-attention calculation. In this example, taking γ = 0.5 as an example, the number of heads in each module is 4, 8, 16, and 32 respectively. On the other hand, due to the change in the input form, the specific settings of each module need to be adaptively changed, and the size of the change is determined by γ as the reduction factor. As Figure 7 stated in, taking γ = 0.5 as an example, after passing through each module, the H and W dimensions of the output feature map set are halved, and the C dimension is doubled. Taking the aforementioned input as an example, after passing through four modules, the output Z (2) 4: (4, 2, 2, 1024) is input to the classification layer (Head layer);
[0118] In the classification layer, if the feature map set output by the last module has not been downsampled to the form of the feature set, such as the aforementioned Z (2) 4: (4, 2, 2, 1024), then the input still needs to pass through the global average pooling layer to obtain a feature set, such as the aforementioned Z (2) 4 obtains Z after passing through the global average pooling layer (3)4: (4, 1, 1, 1024), that is, both the scale parameter sequence and the feature sequences of each sub - figure are summarized into one feature, and then the mean value is calculated among the four features (the first dimension) to obtain the feature, which finally enters the classification layer. For example, the feature Z finally obtained in this embodiment (4) 4: (1, 1, 1, 1024).
[0119] In this example, the final classification layer is a linear layer (fully - connected layer) with an input of 1024 and an output of CL, which takes Z (4) 4: (1, 1, 1, 1024) is input into this linear layer to obtain the classification output K: (1, 1, 1, CL), where CL is the number of categories.
[0120] Input K: (1, 1, 1, CL) into a decision function, such as the Softmax function, to obtain the final vein image recognition (classification) result.
[0121] It is worth noting that the present invention and its embodiments are improvements to the ViT model in the existing literature. Therefore, this article focuses on the differences from the ViT model in the existing literature. Details that already exist in other existing literature or are well - known to those skilled in the art, such as the normalization layer (LN), the feed - forward neural network (FFN), the residual connection, the Droppath mechanism processing, the self - attention calculation mechanism, and the multi - layer perception (MLP) and other technical means, are not elaborated here, or can be referred to in the literature A. DosoViTskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020. and other literature cited by this literature.
[0122] In this embodiment, a vein recognition method is also disclosed, and its process is as Figure 9 shown, including:
[0123] Step 1: Obtain a vein picture;
[0124] Step 2: Input the vein picture into the aforementioned trained vein recognition neural network model based on multi - scale Transformer to obtain the recognition result.
[0125] In this embodiment, a vein recognition system is also disclosed, and this system is as Figure 10As shown, it includes: a venous image acquisition module and an identification module;
[0126] The venous image acquisition module is used to acquire venous images to be identified;
[0127] The identification module is deployed with a pre-trained venous identification neural network model based on multi-scale Transformer, which is used to obtain an identification result according to the input venous image through the venous identification neural network model based on multi-scale Transformer.
[0128] In some other embodiments, as Figure 11 shown, the venous identification system further includes a training module, which is used to acquire venous image samples for training the venous identification neural network model based on multi-scale Transformer in this embodiment;
[0129] and use the venous image samples to train the venous identification neural network model based on multi-scale Transformer, and then update the parameters in the neural network model.
[0130] Experimental example
[0131] In this part, a venous identification neural network model based on multi-scale Transformer (referred to as OUR in the table) as Figure 7 shown is established, and the recognition accuracy rate of the model is trained and tested using venous images in different databases. For comparison, various network models in the prior art are also reproduced and trained and tested in this part. These models and their sources include:
[0132] ResNet: K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
[0133] VGG: K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014.
[0134] FV-CNN: R. Das, E. Piciucco, E. Maiorana, and P. Campisi, “Convolutional neural network for finger-vein-based biometric identification,” IEEE Transactions on Information Forensics and Security, vol. 14, no. 2, pp. 360–373, 2018.
[0135] PV-CNN: H. Qin, M. A. El-Yacoubi, Y. Li, and C. Liu, “Multi-scale and multidirection gan for cnn-based single palm-vein identification,” IEEE Transactions on Information Forensics and Security, vol. 16, pp. 2652–2666, 2021.
[0136] FVRAS-Net: W. Yang, W. Luo, W. Kang, Z. Huang, and Q. Wu, “Fvras-net: An embedded finger-vein recognition and antispoofing system using a unified cnn,” IEEE Transactions on Instrumentation and Measurement, vol. 69, no. 11, pp. 8690–8701, 2020.
[0137] Lightweight CNN: J. Shen, N. Liu, C. Xu, H. Sun, Y. Xiao, D. Li, and Y. Zhang, “Finger vein recognition algorithm based on lightweight deep convolutional neural network,” IEEE Transactions on Instrumentation and Measurement, 2021.
[0138] ViT: A. DosoViTskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020.
[0139] MS-ViT: H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” in Proceedings of the IEEE / CVF International Conference on Computer Vision, 2021, pp. 6824–6835.
[0140] Database 1, “The PolyU multispectral palmprint database”, from The Hong Kong Polytechnic University, contains 6000 palm vein images, including 250 subjects. For each subject, both the left and right palms are collected in two stages, with 6 images collected for each palm in each stage. The average time interval between the two collection stages is 9 days. That is, each volunteer provides 24 images from two palms. All images are collected using near-infrared (NIR). The original palm vein images contain background regions that do not provide discriminative information. Therefore, only the regions of interest are extracted and normalized to images of size 100×100 in the experiment.
[0141] Database 2, “VERA PalmVein”, contains 2200 palm vein images, including 110 subjects. For each subject, both the left and right palms are collected in two stages, with 5 images collected for each palm in each stage. That is, each volunteer provides 20 images from two palms. In the experiment, the images of the regions of interest are extracted and the obtained images are normalized by a preprocessing method.
[0142] Database 3: Tongji University Palmprint Database, which includes 12,000 palm vein images of 300 objects. For each object, images of both the left and right palms are collected in two stages. In each stage, 10 images are collected for each palm. The average time interval between the two collection stages is two months. That is to say, for each volunteer, 40 images from two palms are available. All images are collected in a non-contact manner. Since the images of the region of interest are already included in the database, they can be directly used in the experiment.
[0143] In the experiment, to evaluate the performance of the model, three public databases are each divided into two sets: a training set and a test set. Different palms of the same person are regarded as different classes. So, Database 1 has 500 classifications (L = 500), Database 2 has 220 classifications, and Database 3 has 600 classifications. In the experiment, the palm images collected in the first stage are used as the training set, and the palm images collected in the second stage are used as the test set. Therefore, there are 3,000 images in the training set and test set of Database 3 respectively. Similarly, there are 6,000 images in the training set and test set of Database 2 respectively, and 1,100 images in Database 3.
[0144] For each palm, in the experiment, different numbers of images are selected from the training dataset to train different models, and the test set is used to test the recognition accuracy of the models. For Database 1, 1 to 6 images of each palm are used for training respectively. For Database 2, 2 to 5 images of each palm are used for training respectively. For Database 3, 2, 4, 6, 8, and 10 images of each palm are used for training respectively. Tables 1, 2, and 3 respectively show the recognition accuracies of different models under different training sample quantity conditions based on Database 1, 2, and 3.
[0145] Table 1 Comparison Table of Recognition Accuracies Obtained Based on Database 1
[0146]
[0147] Table 2 Comparison Table of Recognition Accuracies Obtained Based on Database 2
[0148]
[0149] Table 3 Comparison Table of Recognition Accuracies Obtained Based on Database 3
[0150]
[0151]
[0152] As can be seen from Tables 1 - 3, compared with various models in the prior art, the models in this embodiment have higher recognition accuracy in the vast majority of cases. Such good performance is because:
[0153] 1) In this embodiment, the neural network model can not only learn the spatial dependence relationship between position blocks in an image, but also capture information independent of the current image size from images of different scales. Therefore, the neural network model in this embodiment can learn robust feature representation for vein recognition.
[0154] 2) The neural network model in this embodiment incorporates convolution into Transformer. The convolution operator allows learning local features by using a local receptive field and sharing weights, while the self-attention mechanism in Transformer can capture global features. The combination of the two modules can form a complementarity to improve the accuracy of vein recognition.
[0155] 3) Each image is converted into multiple images of different scales to expand the training samples, thereby improving the recognition performance.
[0156] It should be particularly noted that although other Transformer-based models, such as ViT and MS-ViT trained on large-scale data, have shown good performance in many computational vision tasks, in the experiments here, they obtained similar results to CNN-based models. This is because Transformer usually contains more parameters to be trained than CNN. However, in the vein recognition task, there are not a large number of training samples, and the capacity of these models cannot be effectively utilized. 2) Images usually show a strong two-dimensional local structure of spatially correlated adjacent pixels. The CNN architecture allows capturing such local structure by using local receptive fields, sharing weights, and spatial subsampling. Thus, it can be seen that the strategy of incorporating convolution into Transformer in the present invention is very practical for the vein recognition task without a large number of training samples.
[0157] The above embodiments merely exemplarily illustrate the principles and effects of the present invention, rather than limiting the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A venous recognition neural network model system based on a multi-scale Transformer, characterized in that, It includes a scale transformation module, a multi-scale segmentation module, a linear embedding module, and a multi-scale Transformer module connected in sequence, and also includes a tile convolution module; At least one of the multi-scale Transformer modules is connected to the tile convolution module to form a multi-scale Transformer-convolution module group; If the multi-scale Transformer-convolution module group contains multiple multi-scale Transformer modules, each multi-scale Transformer module is cascaded in sequence, and the last multi-scale Transformer module is connected to the tile convolution module; When the output Y of the multi-scale Transformer module enters the tile convolution module, the form of the two-dimensional feature matrix set is transformed; The tile convolution module includes a first particle convolution stack block, a second particle convolution stack block, and a downsampling layer connected in sequence; The scale transformation module is used to scale the vein image I into sub-images I of multiple different sizes n, For n = 1, 2... N, let h0, w0, and c be the height, width, and number of channels of the vein image I, respectively, and h n , w n , and c be the height, width, and number of channels of the sub-image I n , respectively. There are: h1 > h2 >... h n , w1 > w2 >... w n ; The multi-scale segmentation module is used to segment each sub-graph I n into position blocks with a size of P×L. Among them, for the I with the largest size 1, non-overlapping segmentation is adopted; for the remaining sub-graphs, overlapping segmentation is adopted, and this overlapping segmentation makes each sub-graph be segmented into the same number of position blocks with a size of P×L; Also, flatten each position block of each sub - figure into a sequence with a length of C = PLc , m = 1, 2…M, n = 1, 2…N; where M is the number of position blocks, and there is: M = H × W H = , W = ; The linear embedding module is used to map each sequence to features of length D through a learnable mapping matrix E , and splice the features of each sub-graph into a one-dimensional feature sequence I t,n ; And, for each feature sequence I t,n perform learnable positional encoding respectively; In addition, a learnable scale embedding sequence E in the same form as the feature sequence I is added t,n to jointly form a feature sequence set I with the feature sequences of each sub-graph scale ; TE ; The multi-scale Transformer module includes a scale self-attention calculation part and a spatial self-attention calculation part connected in sequence; The scale self-attention calculation part is used to calculate the self-attention weights between position blocks corresponding to the same position on different subgraphs based on the input feature sequence set I TE to obtain the scale self-attention weights, and obtain the output X based on the feature map block set I TE ; new ; The spatial self-attention calculation part is used to calculate the self-attention weights between position blocks corresponding to different positions on the same sub-graph based on the input X new which are called spatial self-attention weights, and obtain the output X new based on X (1) .
2. The system according to claim 1, wherein The multi-scale Transformer module also includes a multi-layer perception part connected after the spatial self-attention calculation part, including a normalization layer and a multi-layer perception sub-module connected in sequence. A Droppath mechanism and a residual connection are introduced in this part, and the output Y of this part is the output of the multi-scale Transformer module.
3. The system according to claim 1, characterized in that, The scale self-attention calculation part includes a normalization layer, a scale self-attention sub-module, and a feed-forward network module connected in sequence. A Droppath mechanism and a residual connection are introduced after the feed-forward network module in this part; Among them, the scale self-attention sub-module is used to take a total of N + 1 features corresponding to the same tile position in each feature sequence / scale embedding sequence in the input as a group of input sequences, and calculate the self-attention weights between the features in each group of input sequences.
4. The system according to claim 1, wherein The spatial self-attention calculation part includes a normalization layer and a spatial self-attention sub-module connected in sequence. A Droppath mechanism and a residual connection are introduced after the spatial self-attention sub-module in this part; Among them, the spatial self-attention sub-module is used to take the feature sequences corresponding to the same sub-graph or scale embedding tile in the input as a group of input sequences, and calculate the self-attention weights between the feature sequences in each input sequence.
5. The system according to claim 1, wherein The self-attention calculation in the multi-scale Transformer module is multi-head self-attention calculation.
6. The system according to claim 1, wherein The first particle volume stack block is used to, on the one hand, obtain an output Y by passing the input Y through a series-connected fully connected layer and a two-dimensional convolutional layer with a convolution kernel of 1×1 and a stride of 1. (1) Among them, the number of output channels of both the fully connected layer and the convolutional layer is γ times the number of channels of the input Y, where γ < 1. On the other hand, let the input Y first pass through a fully-connected depthwise convolution sub-module, in which a depthwise convolution layer with a convolution kernel of k×k and a stride of 1 is added on the basis of the fully-connected layer, and then pass through a two-dimensional convolution layer with a convolution kernel of 1×1 and a stride of 1 to obtain the output Y (2) , and the number of output channels of both the fully-connected depthwise convolution sub-module and the two-dimensional convolution layer is γ times the number of channels of the input Y, where γ < 1; And, connecting the said Y (1) s, Ys, and Ys (3) in sequence in the channel dimension to obtain the output Z; The second particle volume stack block is used to obtain the output Z based on the input Z in the same manner as the first particle volume stack block (1) ; The downsampling layer is used to perform downsampling on Z based on a 2×2 convolutional kernel, and the number of input channels is half of the number of input channels. (1) 7. The system according to claim 6, wherein It includes multiple groups of cascaded multi-scale Transformer-convolution module groups; Among them, the neural network form in the subsequent multi-scale Transformer-convolution module group is adapted to the output form of the previous multi-scale Transformer-convolution module group; Moreover, the number of heads in the multi-head self-attention calculation in the subsequent multi-scale Transformer-convolution module is times that in the previous multi-scale Transformer-convolution module; Preferably, it includes four groups of cascaded multi-scale Transformer-convolution module groups.
8. A vein recognition method, characterized in that, It includes: Step 1: Obtain a vein picture; Step 2: Input the vein image into the trained vein recognition neural network model system based on multi-scale Transformer as described in any one of claims 1-7 to obtain a recognition result.
9. A vein recognition system, characterized in that, It includes a vein image acquisition module and a recognition module; The vein image acquisition module is used to acquire the vein image to be recognized; The recognition module is deployed with a trained vein recognition neural network model system based on multi-scale Transformer as described in any one of claims 1-7, and is used to obtain a recognition result through the vein recognition neural network model system according to the input vein image.
10. The system according to claim 9, characterized in that, It further includes a training module, which is used to acquire vein image samples for training the vein recognition neural network model system; And use the vein image samples to implement training on the vein recognition neural network model system based on multi-scale Transformer, and then update the parameters in the neural network model system.
Citation Information
Patent Citations
Finger vein recognition model training method, recognition method, system and terminal
CN114581965A
Target region segmentation identification method and system of CT image
CN115018809A