Multi-modal image fusion identification method and device, medium and equipment
By combining the feature extraction capabilities of Transformer and CNN, a multi-modal image fusion recognition method with multi-stage dynamic weighted feature fusion is used to solve the problem of the existing technology degradation of recognition performance and image in special circumstances, and the recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510225774.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
The existing recognition technology based on palm lines and palm veins may decline in special circumstances such as lighting, noise and spoofing attacks, and the recognition accuracy may be significantly reduced when the image is incomplete.
The multi-modal image fusion recognition method is adopted, and the global context capture capability of Transformer and the local feature extraction capability of CNN are used, combined with the two-stage dynamic weighted feature fusion, more comprehensive and accurate features are extracted, and a multi-stage dynamic fusion network is trained through the weighted sum of cross entropy loss, full-connection layer weight loss and distance loss.
The recognition accuracy is improved, especially in the case of incomplete images, and it has strong robustness and adaptability, which is better than the recognition accuracy and equal error rate of the prior art.
Smart Images

Figure CN120147801A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image fusion and classification, and particularly to a multi-modal image fusion recognition method, device, medium and equipment. Background Art
[0002] Biometric recognition based on the hand has become a research hotspot and is currently widely used in the field of identity recognition. Single-modal hand features such as fingerprints, finger veins, palm prints, palm veins, and dorsal veins are commonly used in scenarios such as security access control, mobile payment, and medical records. Palm prints and palm veins have attracted much attention due to their rich features and difficulty in forgery. Palm print features include the texture, lines, branch points, etc. on the palm print surface, with large differences between individuals, relatively stable, and not easily changed. Palm vein recognition technology uses near-infrared spectroscopy (NIR) or infrared spectroscopy (IR) radiation imaging technology to capture images of the veins and blood vessels inside the palm. Palm vein images have the uniqueness of vein features and the characteristics of live detection, so palm vein recognition technology has higher security and accuracy.
[0003] Currently, existing recognition technologies based on palm prints and palm veins use the image features of a single modality for recognition. In some special cases, such as lighting, noise, and spoofing attacks, the performance of feature recognition may decline. To overcome this limitation, researchers have explored multi-modal biometric methods to fuse the features of palm prints and palm veins at different levels to comprehensively and effectively improve the performance of the recognition system and achieve effects that cannot be achieved by single-modal biometric recognition. In the past few decades, researchers have proposed many palm print and vein image fusion recognition methods, including traditional image fusion recognition methods and deep learning-based image fusion recognition methods.
[0004] Traditional image fusion recognition methods use manually designed features to perform image fusion at different levels, and often face major challenges in achieving good image fusion performance. It is necessary to carefully design feature extraction and fusion strategies, which not only increases the complexity of the method but may also lead to the loss of complementary information in the original image.
[0005] Most existing deep learning-based image fusion recognition methods are based on the convolutional neural network (CNN) for network model design. Some important shallow features may be lost when passing through the convolutional layer and pooling layer, and there are deficiencies in maintaining global context information during the feature extraction process. When the user's palm may be damaged, affected by lighting, or due to hand posture, etc., resulting in incomplete palm print and palm vein images, the recognition accuracy of the CNN-based image fusion recognition method will be significantly reduced.
[0006] Therefore, how to achieve multi-modal image fusion recognition, improve the recognition accuracy, and thus be more suitable for various actual application scenarios has become a technical problem to be solved currently. Summary of the Invention
[0007] The object of the present invention is to provide a multi-modal image fusion recognition method, device, medium and equipment for realizing multi-modal image fusion recognition and improving the recognition accuracy.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] According to one aspect of the present invention, a multi-modal image fusion recognition device is provided, including: a feature extraction module, a dynamic weighted feature fusion module and a fully connected module; wherein,
[0010] The feature extraction module is used to obtain multi-modal images of different categories and perform shallow feature extraction on the multi-modal images respectively by using a convolutional layer; the dynamic weighted feature fusion module is used to perform feature fusion on the shallow features to obtain shallow fusion features.
[0011] The feature extraction module further performs global feature extraction on the shallow fusion features to obtain global features; the dynamic weighted feature fusion module is used to perform feature fusion on the shallow fusion features and the global features to obtain deep fusion features.
[0012] The feature extraction module is used to perform local feature extraction on the deep fusion features to obtain local features.
[0013] The fully connected module is used to process the local features and identify the category labels of the multi-modal images.
[0014] According to an embodiment of the present invention, the feature extraction module includes an encoding block, a global feature extraction block and a local feature extraction block. The encoding block includes a convolutional layer, a batch normalization layer and an activation layer connected in sequence. The convolutional layer performs shallow feature extraction on the multi-modal images, the batch normalization layer performs batch normalization, and the activation layer performs activation layer processing to obtain the output features of the encoding block carrying the shallow features of the multi-modal images.
[0015] According to an embodiment of the present invention, the dynamic weighted feature fusion module includes a shallow dynamic weighted feature fusion unit and a deep dynamic weighted feature fusion unit; the shallow dynamic weighted feature fusion unit performs element-wise weighted fusion on the output features of the encoding block carrying the shallow features of the multi-modal images, retains all the information in the shallow features of different multi-modal images, and obtains shallow fusion features.
[0016] According to an embodiment of the present invention, the global feature extraction block includes an attention mechanism unit and a Transformer unit; the attention mechanism unit includes a channel attention block and a spatial attention block. The channel attention block is used to capture channel information from the features output by the encoding block, and the spatial attention block is used to capture spatial information from the features output by the encoding block. The channel information and the spatial information are respectively multiplied by the features output by the encoding block and then summed to obtain the output features of the attention mechanism unit;
[0017] The Transformer unit includes a multi-head self-attention unit, a multi-layer perceptron, and multiple layer normalization units. The multi-head self-attention unit divides the output features of the attention mechanism unit into non-overlapping local windows, performs self-attention operations on each window, and then inputs them into layer normalization and a multi-layer perceptron to obtain global features.
[0018] According to an embodiment of the present invention, the deep dynamic weighted feature fusion unit performs channel fusion on the shallow fusion features and the global features to obtain deep fusion features.
[0019] According to an embodiment of the present invention, the local feature extraction block includes a grouped extraction subunit and multiple sequential extraction subunits. The grouped extraction subunit includes multiple parallel grouped convolutional layers and a max pooling layer, and the sequential extraction subunits include a first convolutional layer, a second convolutional layer, and a max pooling layer connected in sequence;
[0020] The multiple parallel grouped convolutional layers respectively perform convolutional processing on the deep fusion features and output them to the max pooling layer;
[0021] The multiple sequential extraction subunits sequentially perform extraction processing on the output of the grouped extraction subunit, and the last sequential extraction subunit outputs local features.
[0022] According to an embodiment of the present invention, the fully connected module includes a first fully connected layer, a second fully connected layer, a batch normalization unit, and a Dropout unit connected in sequence.
[0023] On the other hand, the present invention also provides a multi-modal image fusion recognition method, including the following steps:
[0024] Obtain multi-modal images of different categories;
[0025] Use convolutional layers to perform shallow feature extraction on the multi-modal images respectively, and then perform feature fusion to obtain shallow fusion features;
[0026] Perform global feature extraction on the shallow fusion features to obtain global features;
[0027] Fuse the shallow fusion features and the global features to obtain deep fusion features;
[0028] Extract local features from the deep fusion features to obtain local features;
[0029] Perform a fully connected process on the local features to identify the class label of the multimodal image.
[0030] On the other hand, the present invention also provides a computer storage medium storing instructions, which, when run, implement the multimodal image fusion recognition method described above.
[0031] On the other hand, the present invention also provides a computing device, characterized by including a processor and a communication interface coupled to the processor; the processor is used to run a computer program or instructions to implement the multimodal image fusion recognition method.
[0032] Advantageous Effects
[0033] A multimodal image fusion recognition method, device, medium, and device provided by the present invention have the following advantageous effects compared with the prior art:
[0034] 1. The multimodal image fusion recognition method utilizes the global context capture ability of Transformer and the local feature extraction ability of CNN, combined with two-stage dynamic weighted feature fusion, and the extracted features are more comprehensive and accurate.
[0035] 2. The recognition accuracy and equal error rate of the multimodal image fusion recognition method on the public dataset and on the incomplete dataset are better than those of the prior art, and it has strong robustness and adaptability even in the case of incomplete images.
[0036] 3. The multimodal image fusion recognition method can train a multi-stage dynamic fusion network by adopting a loss function including the weighted sum of cross-entropy loss, fully connected layer weight loss, and distance loss, realize multimodal image fusion recognition, and the obtained network model can effectively fuse and recognize palmprint and vein images, improving the recognition accuracy. Description of the Drawings
[0037] The drawings described here are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0038] Figure 1 is a flowchart of the multimodal image fusion recognition method according to an exemplary embodiment of the present invention;
[0039] Figure 2Schematic diagram of a multi-modal image fusion recognition device according to an exemplary embodiment of the present invention;
[0040] Figure 3 Schematic diagram of a global feature extraction block according to an exemplary embodiment of the present invention;
[0041] Figure 4 Schematic diagram of a local feature extraction block according to an exemplary embodiment of the present invention;
[0042] Figure 5 Variation curve graph of recognition accuracy rate (%) of different methods obtained on Tongji according to an exemplary embodiment of the present invention;
[0043] Figure 6 Variation curve graph of recognition accuracy rate (%) of different methods obtained on CASIA according to an exemplary embodiment of the present invention;
[0044] Figure 7 Variation curve graph of recognition accuracy rate (%) of different methods obtained on IITD_NIR according to an exemplary embodiment of the present invention;
[0045] Figure 8 Variation curve graph of ROC of different methods obtained on Tongji according to an exemplary embodiment of the present invention;
[0046] Figure 9 Variation curve graph of ROC of different methods obtained on CASIA according to an exemplary embodiment of the present invention;
[0047] Figure 10 Variation curve graph of ROC of different methods obtained on IITD_NIR according to an exemplary embodiment of the present invention;
[0048] Figure 11 Schematic diagram of mask operation on the original ROI image according to an exemplary embodiment of the present invention;
[0049] Figure 12 Variation graph of the change amount of recognition accuracy rate (%) on the incomplete Tongji, CASIA and IITD_NIR data sets according to an exemplary embodiment of the present invention;
[0050] Figure 13 Variation graph of the equal error rate (%) change amount on the incomplete Tongji, CASIA and IITD_NIR data sets according to an exemplary embodiment of the present invention. Detailed implementation manners
[0051] For the convenience of clearly describing the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. For example, the first threshold and the second threshold are only used to distinguish different thresholds, and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily mean different.
[0052] It should be noted that in the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0053] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The following at least one (item) or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b or c can mean: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b, c can be single or multiple.
[0054] For the image fusion recognition method based on the CNN architecture, since convolution and pooling may lose some important shallow features, there is insufficient feature extraction in maintaining global context information. When encountering the recognition problem of partial image loss, the performance of this type of method will be greatly reduced. The Transformer network shows powerful performance in natural language processing and computer vision. Through the self-attention mechanism and the global receptive field, the Transformer network can more effectively capture the global context information in the image and achieve more refined feature fusion.
[0055] As Figure 1 shown, a multimodal image fusion recognition method is given, including the following steps:
[0056] Step S1: Obtain multimodal images of different categories;
[0057] Step S2: Use the convolutional layer to perform shallow feature extraction on the multi-modal images respectively, and then perform feature fusion to obtain shallow fusion features;
[0058] Step S3: Perform global feature extraction on the shallow fusion features to obtain global features;
[0059] Step S4: Perform feature fusion on the shallow fusion features and the global features to obtain deep fusion features;
[0060] Step S5: Perform local feature extraction on the deep fusion features to obtain local features;
[0061] Step S6: Perform fully connected processing on the local features to identify the class labels of the multi-modal images.
[0062] The multi-modal images include palmprint images and palmar vein images, and the palmprint images and palmar vein images can be complete images or incomplete images.
[0063] The multi-modal images include fundus images and ophthalmic vein images, and the fundus images and ophthalmic vein images can be complete images or incomplete images.
[0064] The shallow feature extraction is implemented by a convolutional layer; the global feature extraction is implemented by a Transformer network.
[0065] The multi-modal image fusion recognition method utilizes the global context capture ability of the Transformer and the local feature extraction ability of the CNN, combined with two-stage dynamic weighted feature fusion. The method relies on a two-stage dynamic weighted fusion network based on the Transformer and the CNN, and trains the multi-stage dynamic fusion network through the weighted sum of the cross-entropy loss, the fully connected layer weight loss, and the distance loss to achieve multi-modal image fusion recognition, especially for high-precision fusion recognition in the case of complete and incomplete palmprint and palmar vein images. The weighted sum is used to improve the representation ability of the model and prevent overfitting; the combination of the Transformer and the CNN is used to improve the feature extraction ability of palmprint and palmar vein images in the case of complete and incomplete; the two-stage dynamic weighted feature fusion is used to dynamically adjust the fusion weights during the model training process to improve the feature fusion ability of palmprint and palmar vein images. Applying the technical solution of the present invention can achieve effective fusion and recognition of palmprint and vein images, improve the recognition accuracy, and solve the problems of the existing technology in maintaining global context information and low recognition accuracy for the fusion recognition of incomplete palmprint and palmar vein images, so as to be more suitable for various actual application scenarios. In particular, the problem of the fusion of visible light images and infrared images is solved.
[0066] Such as Figure 2As shown, a schematic diagram of a multi-modal image fusion recognition device is given. The multi-modal image fusion recognition device is based on a network architecture of Transformer and CNN, and includes a feature extraction module, a dynamic weighted feature fusion module, and a fully connected module; among them,
[0067] The feature extraction module is used to obtain multi-modal images of different categories, and perform shallow feature extraction on the multi-modal images respectively using convolutional layers; the dynamic weighted feature fusion module is used to perform feature fusion on the shallow features to obtain shallow fusion features;
[0068] The feature extraction module further performs global feature extraction on the shallow fusion features to obtain global features; the dynamic weighted feature fusion module is used to perform feature fusion on the shallow fusion features and the global features to obtain deep fusion features;
[0069] The feature extraction module is used to perform local feature extraction on the deep fusion features to obtain local features;
[0070] The fully connected module is used to process the local features and identify the category labels of the multi-modal images.
[0071] The feature extraction module includes an encoding block, a global feature extraction block, and a local feature extraction block.
[0072] The encoding block includes a convolutional layer, a batch normalization layer, and an activation layer connected in sequence. The convolutional layer performs shallow feature extraction on the multi-modal images, the batch normalization layer performs batch normalization, and the activation layer performs activation layer processing to obtain the output features of the encoding block carrying the shallow features of the multi-modal images. The activation layer can be a ReLU activation layer or a SeLU activation layer.
[0073] For example, in Figure 2 the multi-modal images are a palmprint image and a palm vein image respectively.
[0074] The palmprint image in the dataset is input into the encoder, and the convolutional layer performs shallow feature extraction on the palmprint image. Then, batch normalization and activation layer processing are performed on the extracted shallow features to obtain the output features of the encoding block carrying the shallow features of the palmprint image.
[0075] The palm vein image in the dataset is input into the encoder, and the convolutional layer performs shallow feature extraction on the palm vein image. Then, batch normalization and activation layer processing are performed on the extracted shallow features to obtain the output features of the encoding block carrying the shallow features of the palm vein image.
[0076] The multi-modal images include a palmprint image and a palm vein image. Let I p represent the palmprint image, and let I vRepresents a palm vein image. An encoder is assigned to each multimodal image. After being encoded by the encoders of the encoding block, the palmprint image and the palm vein image are respectively represented as:
[0077]
[0078] Among them, Is the output of the palmprint branch, Is the output of the palm vein branch. Conv(.) represents the convolution operation, B(.) represents the batch normalization process, and R(.) represents the activation operation. When using the ReLU activation layer, it represents the ReLU activation operation.
[0079] The dynamic weighted feature fusion module includes a shallow dynamic weighted feature fusion unit and a deep dynamic weighted feature fusion unit;
[0080] The shallow dynamic weighted feature fusion unit performs element-wise weighted fusion on the output features of the encoding block carrying the shallow features of the multimodal image, retains all the information in the shallow features of different multimodal images, and obtains the shallow fusion features.
[0081] For example, for the shallow features of the palmprint image and the palm vein image, weighted fusion is performed, and weighted fusion is carried out according to the weight values of the two shallow features. For example, the weight values of the shallow features of the palmprint image and the palm vein image are w 1 And 1 - w 1 . The weight value can be initially a random value and the final weight value is obtained through multiple rounds of training.
[0082] The obtained shallow dynamic weighted feature is the shallow fusion feature, denoted as Input 1 , and is represented as:
[0083] Among them, w 1 Represents the weight, which is a learnable parameter and is dynamically learned through the backpropagation algorithm during the model training process, and is used to adjust the weights of the encoder branches corresponding to the multimodal images, such as adjusting the weights of the palmprint branch and the palm vein branch.
[0084] As Figure 3 Shown, the global feature extraction block includes an attention mechanism unit and a Transformer unit;
[0085] The attention mechanism unit includes a channel attention block and a spatial attention block. The channel attention block is used to capture channel information from the output features of the encoding block, and the spatial attention block is used to capture spatial information from the output features of the encoding block. The channel information and the spatial information are respectively multiplied by the output features of the encoding block and then added together to obtain the output features of the attention mechanism unit;
[0086] InFigure 3 In it, the shallow fusion feature Input1 is multiplied by itself after passing through the channel attention operation; the output feature Input1 of the encoding block is multiplied by itself after passing through the spatial attention operation; the two products obtained from the channel attention block and the spatial attention block are added together to obtain the output of the attention mechanism unit, expressed as:
[0087] Output 注意力机制单元 = O 信道注意 × Input 1 + O 空间注意 × Input 1 ,
[0088] Among them, Output 注意力机制单元 represents the output of the attention mechanism unit, O 信道注意 represents the channel attention operation, O 空间注意 represents the spatial attention operation.
[0089] The Transformer unit includes a multi-head self-attention unit, a multi-layer perceptron, and multiple layer normalization units. The multi-head self-attention unit divides the output feature of the attention mechanism unit into non-overlapping local windows, performs self-attention operations on each window, and then inputs it into layer normalization and a multi-layer perceptron to obtain global features.
[0090] The output of the multi-head self-attention is expressed as:
[0091] Output 多头自注意力 = Output 注意力机制单元 + (MSA(LN(Output 注意力机制单元 ))).
[0092] Then it is input into the layer normalization sub-unit and the multi-layer perceptron to obtain the global feature Output 1 , which is expressed by the formula as:
[0093] Output 1 = Output 多头自注意力 + (MLP(LN(Output 多头自注意力 ))).
[0094] Among them, Output 多头自注意力 represents the output of the multi-head self-attention sub-unit, LN represents the layer normalization operation, MSA represents the multi-head self-attention operation, and MLP represents the multi-layer perceptron operation.
[0095] Such as Figure 2As shown, the deep dynamic weighted feature fusion unit performs channel fusion on the shallow fusion feature and the global feature to obtain a deep fusion feature, forming a more representative deep fusion feature representation Input 2 , which is expressed as:
[0096] Input 2 = w 2 × Output 1 ||(1 - w 2 ) × Input 1 , where w 2 represents the weight, which is a learnable parameter and is dynamically learned through the backpropagation algorithm during model training for adjusting the weight of the channel fusion feature, and || represents channel fusion.
[0097] As Figure 4 shown, the local feature extraction block includes a grouped extraction subunit and multiple sequential extraction subunits,
[0098] The grouped extraction subunit includes multiple parallel grouped convolutional layers and a max pooling layer, and the sequential extraction subunit includes a first convolutional layer, a second convolutional layer, and a max pooling layer connected in sequence;
[0099] The multiple parallel grouped convolutional layers respectively perform convolutional processing on the deep fusion feature and then output to the max pooling layer;
[0100] The multiple sequential extraction subunits sequentially perform extraction processing on the output of the grouped extraction subunit, and the last sequential extraction subunit outputs the local feature.
[0101] Then, the obtained local feature is input into the fully connected module for classification. The fully connected module is used for classification and recognition according to the local feature.
[0102] As Figure 1 shown, the fully connected module includes a first fully connected layer, a second fully connected layer, a batch normalization unit, and a Dropout unit connected in sequence.
[0103] The loss function Loss consists of cross-entropy loss, weight loss, and distance loss, and the calculation method is as follows:
[0104]
[0105] Among them, CL represents cross-entropy loss, which is used to improve classification accuracy, λ 1 and λ 2 are non-negative weights, H 1 and H 2 are the weights of the first fully connected layer and the second fully connected layer, |||| 2It represents the calculation of the L2 norm, which is used to avoid overfitting. D is the distance loss, which is used to increase the inter-class distance and decrease the intra-class distance, and its calculation method is as follows:
[0106]
[0107] where N is the batch size, and d i is the Euclidean distance between two samples, Y is the label indicating whether the two samples match, max represents the maximum operation, and m is the set threshold.
[0108] Example 1: Multimodal Image Fusion Recognition Method Based on Palmprint and Palm Vein
[0109] In this example, the Tongji Palmprint and Palm Vein Dataset, CASIA Multispectral Palmprint Dataset, IITD Palmprint Dataset, and PolyU Palm Vein Dataset are used to train a deep learning model. The relevant information of the above datasets is shown in Table 1.
[0110] Table 1: Palmprint and Palm Vein Datasets
[0111]
[0112] The Tongji Palmprint and Palm Vein Dataset includes 24,000 palm images of 300 people collected using a non-contact method. Each palm is collected twice, 10 images are collected each time, and the average time interval between the two collections is 61 days. The image size is 800×600, and the dataset also provides ROI images with a size of 128×128.
[0113] The CASIA Multispectral Palmprint Dataset includes 7,200 palm images of 100 people captured using a self-designed multispectral imaging device. For each hand, two sets of palm images are captured, and the time interval between the two collections is more than one month. 3 samples are collected each time, and these images are collected under six different electromagnetic spectra. The wavelengths of the illuminators corresponding to the six spectra are 460nm, 630nm, 700nm, 850nm, 940nm, and white light, and the image size is 768×576.
[0114] The IITD Palmprint Dataset includes 2,601 images from 230 people collected using a non-contact collection method. Each palm has 5 to 7 samples. In addition to the original images, automatically cropped and normalized palmprint ROI images with a size of 150×150 pixels are also provided.
[0115] The PolyU palm vein dataset is a multi-spectral palmprint database established by the Hong Kong Polytechnic University. Among them, there are 6,000 vein images collected through contact. The image capture is divided into two stages, with an average interval of 9 days between the two stages. Each palm is captured 6 times in each group. This dataset provides ROI images of size 128×128.
[0116] In this embodiment, the above dataset is used to form 3 available palmprint and palm vein datasets, including the Tongji, CASIA, and IITD_NIR datasets. As shown in Table 2, for the Tongji dataset, all palmprint and vein images are used in this embodiment. For the CASIA dataset, palmprint images at 460nm and 630nm are used, and palm vein images at 850nm and 940nm are used. For the IITD_NIR dataset, in this embodiment, IITD and PolyU are combined, and the first 460 palms are selected for each modality, with 5 images for each palm.
[0117] Table 2 Palmprint and Palm Vein Datasets Used in the Experiment
[0118]
[0119] The palmprint and palm vein fusion recognition method includes the following steps:
[0120] S101: Dataset division. In this example, the Tongji, CASIA, and IITD_NIR datasets are used to train the deep learning network model. For the Tongji dataset, the images collected in the first stage are used for the training stage, and the images collected in the second stage are used for the testing stage. For the CASIA dataset, the images at 460nm and 850nm are used for training, and the other two modalities are used for testing. For IITD_NIR, 3 images are used for training and 2 images are used for testing.
[0121] S102: Data preprocessing. Extract the region of interest from the images to obtain ROI images of size 128×128;
[0122] S103: Input the palmprint image and the palm vein image into Figure 2 the multi-modal image fusion recognition device shown. The dimension is represented as Batch_size×3×128×128, where Batch_size is the batch size set for model training and is 128;
[0123] S104: Input the palmprint image and the palm vein image into the encoder respectively, and perform a convolutional unit, a batch normalization unit, and a ReLU activation unit once. The kernel size of the convolutional unit is 3, the stride is 1, and the padding is 1; the input dimension of the encoder unit is 128×3×128×128; the output dimension is 128×64×128×128;
[0124] S105: After obtaining the encoder output, perform shallow feature fusion, and the output dimension is 128×64×128×128;
[0125] S106: Then pass through the attention mechanism unit and the Transformer unit of the global feature extraction module;
[0126] S107: First, perform channel attention and spatial attention operations in the attention mechanism unit, then multiply with the input Input 1 and add the two products to obtain the output of the attention mechanism unit, and the output dimension is 128×64×128×128;
[0127] S108: Then perform global deep feature extraction through a Transformer unit once. The Transformer unit includes a multi-head self-attention unit, a multi-layer perceptron, and several layer normalization units;
[0128] S109: Pass Output 注意力模块 through a layer normalization unit and a multi-head self-attention unit once and then add it to itself to obtain the output Output 多头自注意力 , with the dimension of 128×64×128×128;
[0129] S110: Then pass through a layer normalization unit and a multi-layer perceptron unit once to obtain the output Output 1 , with the dimension of 128×64×128×128;
[0130] S111: Perform fusion of the features before and after the global feature extraction module through weighted channel fusion, and the output dimension is 128×128×128×128;
[0131] S112: Pass the fused features through the local feature extraction unit, which includes four sub-modules. Each sub-module includes two convolutional units and one max pooling unit. The kernel size of the first convolutional layer of each sub-module is 5, the stride is 1, the padding is 3, the kernel size of the second convolutional layer is 3, the stride is 1, the padding is 3, and the kernel size of the max pooling layer is 4, the stride is 2;
[0132] S113: First pass through the first local feature extraction sub-module, repeat the grouped convolutional units twice, with the number of groups being 2, and then pass through a max pooling unit once, and the output dimension is 128×128×64×64;
[0133] S114: Then pass through the second local feature extraction sub-module, repeat the convolutional units twice and a max pooling unit once, and the output dimension is 128×256×32×32;
[0134] S115: Then, it passes through the third local feature extraction sub-module, repeating the convolutional unit twice and the max pooling unit once, with an output dimension of 128×512×16×16;
[0135] S116: Then, it passes through the fourth local feature extraction sub-module, repeating the convolutional unit twice and the max pooling unit once, with an output dimension of 128×512×8×8;
[0136] S117: The local feature Output 2 that has passed through the local feature extraction unit is input into the fully connected module. The fully connected module includes two fully connected layers, a batch normalization unit, and a Dropout unit. For the Tongji dataset, the output dimension is 128×600; for the CASIA dataset, the output dimension is 128×600; for the IITD_NIR dataset, the output dimension is 128×460;
[0137] S118: Calculate the loss value for the output of the multi-modal image fusion recognition device and perform backpropagation to update the model parameters, and the fusion weights w 1 and w 2 are updated accordingly. For each dataset, the model is trained for 500 epochs;
[0138] S119: Calculate the recognition accuracy. During the training process, the test sets of the three datasets are input into the model to obtain the labels of the target images and calculate the recognition accuracy. Save and fix the relevant parameters of the model when the recognition accuracy is the highest, where the fusion weights w 1 and w 2 The weight values on the three datasets when the recognition accuracy is the highest are shown in Table 3
[0139] Table 3: w 1 and w 2 values
[0140]
[0141] S120: Calculate the equal error rate. Input the test sets of the three datasets into the trained model, obtain the feature representations of the images for one-to-one matching, get the true match and false match scores, and calculate the equal error rate;
[0142] S121: Performance analysis. Compare the recognition accuracy and equal error rate of seven deep learning methods, namely VGG16, ResNet18, EEfficientNetV2, PalmCohashNet, PVFNet, CompNet, and CO3Net, on the said dataset. The results are shown in Table 4, where PalmCohashNet and PVFNet are from the literature, and those without results on the dataset are replaced by "-".
[0143] The recognition accuracy and equal error rate on the three datasets are better than the existing methods. The recognition accuracy and ROC performance are as Figures 5 - 10 shown. The recognition accuracy on CASIA is improved by 0.083%, and the equal error rate is reduced by 0.446%. On IITD_NIR, the recognition accuracy is improved by 1.304%, and the equal error rate is reduced by 0.51%.
[0144] Table 4 Comparison results with other methods (%)
[0145]
[0146] Example 2: Recognition method for incomplete multi-modal datasets
[0147] This example further evaluates the recognition performance of the multi-modal fusion recognition method on incomplete datasets. The incomplete Tongji palmprint and palm vein dataset, CASIA multi-spectral palmprint dataset, IITD palmprint dataset, and PolyU palm vein dataset are used to simulate various challenges that may be encountered in real recognition scenarios, such as hand occlusion, pose change, or illumination change, which will inevitably lead to incomplete captured images.
[0148] The recognition method for the incomplete multi-modal dataset includes the following steps:
[0149] S201: Dataset division. In this example, three test datasets, Tongji, CASIA, and IITD_NIR, are used to test the best model trained in Example 1. For the Tongji dataset, the images collected in the second stage are used as the test dataset. For the CASIA dataset, the 630nm and 940nm images are used as the test dataset. For IITD_NIR, 2 images are used as the test dataset;
[0150] S202: Image incompleteness processing. Mask operation is performed on the ROI images of the dataset. As Figure 11 shown, the left side is the original ROI image, and the right side is the image processed by masking. The same masking operation is applied to each pair of palmprint and palm vein images in all test datasets. The calculation method is as follows:
[0151] I p_incomplete(i,j,c) = I p(i,j,c) × Mask (i,j)
[0152] I v_incomplete(i,j,c) = I v(i,j,c) × Mask (i,j)
[0153] where, I p_incomplete(i,j,c)Denote the pixel value of the palmprint image after mask processing at position (i, j) and channel c as I p(i,j,c) Denote the pixel value of the input palmprint image at the same position and the same channel as I v_incomplete(i,j,c) Denote the pixel value of the palm vein image after mask processing at position (i, j) and channel c as I v(i,j,c) Denote the pixel value of the input palm vein image at the same position and the same channel. Mask represents a mask matrix of size 128×128, Mask (i,j) Denote the element at position (i, j). The value of the element on the matrix is 0 or 1. The elements with value 0 are concentrated at the four corners and the edges of the matrix, and the rest of the elements are 1. Mask operations are performed on each channel of the original palmprint and palm vein ROI images, and finally merged back into the RGB image;
[0154] S203: Recognition accuracy calculation. Input the three mutilated test data sets into the model trained in Embodiment 1 to obtain the labels of the target images and calculate the recognition accuracy;
[0155] S204: Equal error rate calculation. Input the three mutilated test data sets into the model trained in Embodiment 1, obtain the feature representations of the images for one-to-one matching, obtain the true matching and false matching scores, and calculate the equal error rate;
[0156] S205: Performance analysis. Compare the recognition accuracy and equal error rate of five deep learning methods, namely VGG16, ResNet18, EEfficientNetV2, CompNet and CO3Net, on the said data set. The results are shown in Table 5. The recognition performance of these methods on the incomplete data set has decreased. Figures 12 - 13 The bar chart in shows the performance changes of the accuracy and equal error rate on the mutilated Tongji, CASIA and IITD_NIR data sets. It can be seen from the figure that the method proposed by the present invention has the smallest performance change. On the three incomplete data sets, the accuracy has decreased by 7.117%, 18.333% and 5.761% respectively, and the equal error rate has changed by 8.033%, 18.191% and 5.548% respectively. This shows that even in the case of mutilated images, the method of the present invention has strong robustness and adaptability.
[0157] Table 5 Comparison results with other methods on the mutilated data set (%)
[0158]
[0159] In addition, according to an exemplary embodiment of the present invention, a computer-readable storage medium storing a computer program may also be provided. The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to execute a multi-modal image fusion recognition method according to an exemplary embodiment of the present invention. The computer-readable recording medium is any data storage device that can store data read by a computer system. Examples of the computer-readable recording medium include: read-only memory, random access memory, compact disc read-only memory, magnetic tape, floppy disk, optical data storage device, and carrier wave (such as data transmission via the Internet through a wired or wireless transmission path).
[0160] In addition, according to an exemplary embodiment of the present invention, a computing device may also be provided. The computing device includes a processor and a memory. The memory is used to store a computer program. The computer program is executed by the processor to cause the processor to execute a computer program of a multi-modal image fusion recognition method according to an exemplary embodiment of the present invention.
[0161] Although the present invention has been described in conjunction with various embodiments, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the drawings, the disclosure content, and the like. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions listed in the specification. Certain measures are recited in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0162] Although the present invention has been described in connection with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present invention. Accordingly, the present specification and the drawings are merely exemplary descriptions of the present invention and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A multimodal image fusion recognition device, characterized in that: include: Feature extraction module, dynamic weighted feature fusion module and full connection module; among them, The feature extraction module is used to obtain multimodal images of different categories and perform shallow feature extraction on the multimodal images using convolutional layers; the dynamic weighted feature fusion module is used to perform feature fusion on shallow features to obtain shallow fusion features; The feature extraction module further performs global feature extraction on the shallow fusion feature to obtain the global feature; the dynamic weighted feature fusion module is used to perform feature fusion on the shallow fusion feature and the global feature to obtain the deep fusion feature; A feature extraction module is used to extract local features from deep fusion features to obtain local features; The fully connected module is used to process the local features and identify the category labels of the multimodal images.
2. The multimodal image fusion recognition device according to claim 1, characterized in that: The feature extraction module includes a coding block, a global feature extraction block and a local feature extraction block. The coding block includes a convolution layer, a batch normalization layer, and an activation layer connected in sequence. The convolution layer performs shallow feature extraction on the multimodal image, the batch normalization layer performs batch normalization, and the activation layer performs activation layer processing to obtain coding block output features carrying shallow features of the multimodal image.
3. The multimodal image fusion recognition device according to claim 2, characterized in that: The dynamic weighted feature fusion module includes a shallow dynamic weighted feature fusion unit and a deep dynamic weighted feature fusion unit; The shallow dynamic weighted feature fusion unit performs element-wise weighted fusion on the output features of the coding block carrying the shallow features of the multimodal image, retains all information in the shallow features of different multimodal images, and obtains shallow fusion features.
4. The multimodal image fusion recognition device according to claim 3, characterized in that: The global feature extraction block includes an attention mechanism unit and a Transformer unit; The attention mechanism unit includes a channel attention block and a spatial attention block. The channel attention block is used to capture channel information from the output features of the coding block, and the spatial attention block is used to capture spatial information from the output features of the coding block. The channel information and the spatial information are respectively multiplied by the output features of the coding block and then added to obtain the output features of the attention mechanism unit. The Transformer unit includes a multi-head self-attention unit, a multi-layer perceptron and multiple layer normalization units. The multi-head self-attention unit divides the output features of the attention mechanism unit into non-overlapping local windows, performs a self-attention operation on each window, and then inputs the layer normalization and multi-layer perceptron to obtain global features.
5. The multimodal image fusion recognition device according to claim 4, characterized in that: The deep dynamic weighted feature fusion unit performs channel fusion on the shallow fusion features and the global features to obtain deep fusion features.
6. The multimodal image fusion recognition device according to claim 4, characterized in that: The local feature extraction block includes a group extraction subunit and a plurality of sequential extraction subunits. The group extraction subunit includes a plurality of parallel group convolutional layers and a maximum pooling layer, and the sequential extraction subunit includes a first convolutional layer, a second convolutional layer, and a maximum pooling layer connected sequentially; The multiple parallel grouped convolutional layers respectively perform convolution processing on the deep fusion features and output them to the maximum pooling layer; The multiple sequential extraction subunits extract and process the outputs of the grouping extraction subunits in sequence, and the last sequential extraction subunit outputs a local feature.
7. The multimodal image fusion recognition device according to claim 1, characterized in that: The fully connected module includes a first fully connected layer, a second fully connected layer, a batch normalization unit and a Dropout unit which are sequentially connected.
8. A multimodal image fusion recognition method based on the device described in claims 1-7, characterized in that: The following steps are involved: Acquire multimodal images of different categories; The convolutional layers are used to extract shallow features from multimodal images, and then feature fusion is performed to obtain shallow fusion features; Perform global feature extraction on shallow fusion features to obtain global features; Fusing the shallow fusion feature with the global feature to obtain a deep fusion feature; Perform local feature extraction on the deep fusion features to obtain local features; The local features are fully connected to obtain category labels of the multimodal images.
9. A computer storage medium, characterized in that The computer storage medium stores instructions, and when the instructions are executed, the multimodal image fusion recognition method described in claim 8 is implemented.
10. A computing device, characterized in that: It comprises a processor and a communication interface coupled to the processor; the processor is used to run a computer program or instruction to implement the multimodal image fusion recognition method described in claim 8.
Citation Information
Patent Citations
Deep learning method of multi-modal image visibility detection model based on shallow fusion
CN111738314A
Multi-modal remote sensing data classification method fusing global and local information
CN116863247A
Multi-modal emotion recognition method and system based on hybrid fusion and attention mechanism, storage medium and terminal
CN117349442A
Multi-modal emotion recognition method based on staged attention mechanism
CN118378128A
Drunk driving detection method and system based on sensor and machine vision
CN118457218A
Cited By
Multi-modal video fusion method, device and equipment for wide-area articulated naturality web, and medium
CN120510485A