Multi-modal Image Recognition Method Based on the Missing Modality of Dynamic Hint Information
By introducing a fusion module of dynamic prompt information in multimodal image recognition, the problems of modal missing and neglected sample properties are solved, the robustness and performance of the recognition task are improved, adapting to complex environments and reducing overfitting.
Patent Information
- Application Number
- CN202410716518.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-06-04
AI Technical Summary
When dealing with the missing modality, the existing multimodal recognition method ignores the full interaction between the remaining modalities, treats all samples equally, and ignores the difficulty of the samples.
A multimodal image recognition method based on missing modes of dynamic prompt information is proposed. By constructing an overall network model, it includes a modal shared feature extraction network, a modal private feature extraction network, and a fusion module based on dynamic prompt information. This method performs random mode missing during training, and uses dynamic prompt information to flexibly adjust the fusion strategy to generate features of missing modes.
It improves the robustness and performance of multimodal recognition tasks, can better adapt to the variable conditions in real life, reduces the risk of overfitting, and dynamically adjusts the fusion strategy according to the difficulty of the sample, improving the model's ability to cope with complex situations.
Smart Images

Figure CN118692156B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a multimodal image recognition method for missing modalities based on dynamic prompt information. Background Art
[0002] In practical applications, a single modality often fails to provide sufficient information to comprehensively understand and describe an object or a scene. Taking the visual recognition task as an example, relying solely on images may not capture important information related to context, sound, or other perceptual modalities. In addition, single-modal methods usually have a strong dependence on specific environmental conditions and are also highly sensitive to noise and interference. In different environments or backgrounds, the model performance may be significantly affected because it cannot obtain additional information from other perceptual channels. These deficiencies of single modalities have prompted researchers to turn their attention to multimodal methods, by integrating information from different perceptual channels, to overcome the limitations of single modalities and improve the system performance and robustness. This shift aims to break through the limitations of a single perceptual mode and better adapt to the complex and variable conditions in the real world.
[0003] For example, Mehraj et al. proposed a multimodal biometric recognition system based on a multi-level hybrid feature fusion mechanism. First, a pre-trained network with transfer learning capabilities was used to extract feature vectors, which were then fused with handcrafted feature vectors using a histogram of oriented gradients feature descriptor (Mehraj H, Mir A H. “A Multi-Biometric System Based on Multi-Level Hybrid Feature Fusion” Herald of the Russian Academy of Sciences, vol. 91 no. 2 pp. 176-196. 2021). Luo et al. proposed a feature fusion network for joint iris-periocular feature recognition based on a multi-attention mechanism. Spatial attention and channel attention were used in the feature extraction module to effectively learn the most important features and suppress unnecessary features. In addition, a joint attention mechanism was introduced in the feature fusion module, which could adaptively fuse features to obtain better iris-periocular feature representations (Luo Z, Li J, Zhu Y. “A Deep Feature Fusion Network Based on Multiple Attention Mechanisms for Joint Iris-Periocular Biometric Recognition” IEEE Signal Processing Letters, vol. 28 pp. 1060-1064. 2021). Existing multimodal fusion methods utilize the correlation between multimodal information to fuse existing modal information. However, they ignore the easy and difficult samples in the recognition task, treating each sample equally and neglecting the nature of the samples themselves.
[0004] In addition, in a multimodal scenario, modality missing becomes an important and common challenge. Modality missing means that the system cannot obtain all information, which may affect the comprehensive understanding of tasks such as recognition. For example, in a scenario that simultaneously contains images, text, and sounds, if one of the modalities is missing, the system may not be able to fully understand the entire context. Moreover, there are potential correlations between different modalities, and the missing of one modality may make it difficult for the model to learn a complete semantic representation. Modality missing may make the system more sensitive to noise and interference, reducing the robustness of the system. Most existing modality missing methods ignore the sufficient interaction between the remaining modalities when generating the missing modality, thus limiting the improvement of the performance of multimodal recognition tasks. Summary of the Invention
[0005] The present invention proposes a multi-modal image recognition method for missing modalities based on dynamic prompt information, which can effectively solve the problems of ignoring the sufficient interaction between the remaining modalities and ignoring the nature of the samples themselves when generating missing modalities in multi-modal missing tasks, and treating each sample equally.
[0006] The technical solution to implement the present invention is as follows: A multi-modal image recognition method for missing modalities based on dynamic prompt information, comprising the following steps:
[0007] Step 1: Collect a number of palmprint and palm vein images to construct a multi-modal palmprint and palm vein image dataset, preprocess the above dataset, and then divide it into a training set and a test set according to 1:1.
[0008] Step 2: Construct an overall network model:
[0009] The overall network model includes a common feature extraction network for modalities, a private feature extraction network for modalities, and a fusion module based on dynamic prompt information.
[0010] Step 3: Train the overall network model:
[0011] Use the training set to train the overall network model. During training, perform random modality missing to obtain the missing modality and the non-missing modality. After feature extraction, obtain the feature of each modality of the overall network model.
[0012] Step 3-1: Use the training set to train the overall network model. During training, perform random modality missing to obtain the missing modality and the non-missing modality.
[0013] Step 3-2: Use the common feature extraction network for modalities to extract features from the non-missing modality to obtain the common feature of the non-missing modality, and use the private feature extraction network for modalities to extract features from the non-missing modality to obtain the private feature of the non-missing modality.
[0014] Step 3-3: According to the difficulty of different samples, use the fusion module based on dynamic prompt information to flexibly adjust the fusion strategy to obtain the first fusion feature; on the basis of the private feature of the non-missing modality, fuse it with the first fusion feature to generate the feature of the missing modality.
[0015] Step 3-4: Fuse the private feature of the non-missing modality and the feature of the missing modality to obtain the second fusion feature.
[0016] Step 4: Input the second fusion feature into the classification layer to obtain the predicted label, calculate the difference between the predicted label and the true label through the loss function, and backpropagate the error to update the model parameters, thereby optimizing the trained overall network model, improving the recognition accuracy of the network model, and obtaining the final network model, enabling it to classify samples more accurately.
[0017] Step 5: Use the test set to evaluate the final network model and test its accuracy and error rate.
[0018] Compared with the prior art, the significant advantages of the present invention are as follows:
[0019] (1) The multi-modal image recognition method for missing modalities based on dynamic hint information proposed by the present invention enables the model to be more robust and better adapt to the changing conditions in real life in the scenario where both training and testing stages are missing. And the model learns more comprehensive feature representations in the case of missing modalities, and is more likely to learn more robust and general representations for different modalities during training, reducing the risk of overfitting.
[0020] (2) The present invention proposes a fusion module for dynamic hint information, which enhances the sufficient interaction between the remaining modalities and improves the performance of the multi-modal recognition task. And it can dynamically select the fusion module according to the difficulty level of the recognition samples to fuse the modality information that is not missing. This personalized adaptability enables the system to be more optimized for the characteristics of different samples and improves the model's ability to handle various complex situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of the multi-modal image recognition method for missing modalities based on dynamic hint information of the present invention.
[0022] Figure 2 is a framework diagram of the overall network model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] In the present invention, descriptions such as "first" and "second" are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0025] The technical solutions between various embodiments of the present invention can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0026] Next, the specific implementation methods, as well as the technical difficulties and innovative points of this invention, will be further introduced in combination with this design example.
[0027] The present invention first extracts common features and private features of non-missing multi-modalities, and applies a fusion module based on dynamic prompt information to fuse non-missing multi-modal information on the basis of the private features of each modality. Then, the common features and the fused information are used to complete the missing modality features. In the fusion module of dynamic prompt information, according to the difficulty of identifying samples, the fusion module is dynamically selected to fuse non-missing modality information. This personalized adaptability enables the system to better optimize according to the characteristics of different samples and improve the model's ability to handle various complex situations.
[0028] Combined Figures 1 to 2 , the multi-modal image recognition method for missing modalities based on dynamic prompt information according to the present invention includes the following steps:
[0029] Step 1: Preprocess the data set. Through the Transforms class in PyTorch, perform scale transformation on the divided data set to adjust it to a unified size (for example, 224*224), and then perform consistent normalization processing. Divide the normalized data set into a training set and a test set according to 1:1.
[0030] Step 2: Construct the overall network model:
[0031] The overall network model includes a common feature extraction network for modalities, a private feature extraction network for modalities, and a fusion module based on dynamic prompt information.
[0032] Step 3: Use the common feature extraction network for modalities and the private feature extraction network for modalities to extract common features and private features from non-missing multi-modal images respectively, and correspondingly obtain the common features of non-missing modalities and the private features of non-missing modalities, specifically as follows:
[0033] Step 3-1: Use three Transformer blocks to extract the common features of non-missing modalities. Each Transformer block includes two normalization layers, an attention layer, and a feed-forward neural network layer to obtain the common features of non-missing modalities, namely X1 and Y1.
[0034] Step 3-2: Use the first two layers of Resnet34 to extract the private features of the non-missing modalities, including two sets of convolutional layers, a batch normalization layer, a ReLU activation function, a max pooling layer, and two Resnet residual layers, to obtain the private features of the non-missing modalities, namely X and Y.
[0035] Step 3-3: According to the difficulty of different samples, use the fusion module based on dynamic hint information to flexibly adjust the fusion strategy, while avoiding using the same fusion strategy for all samples, effectively reducing the computational overhead, and obtaining the first fusion feature; on the basis of the private features of the non-missing modalities, fuse with the first fusion feature to generate the features of the missing modality, specifically as follows:
[0036] Construct a simple fusion unit to normalize the private features X and Y of the non-missing modalities, then perform a concatenation operation on the two normalized features, and finally use the ReLU activation function to non-linearly transform the fusion feature to obtain the fusion feature Fusion 1 :
[0037] Fusion 1 = Cat(Norm(X), Norm(Y)) (1)
[0038] Fusion 1 = ReLU(Fusion 1 ) (2)
[0039] In the above formula, Cat represents the concatenation operation, Norm represents the normalization operation, and ReLU represents the activation function;
[0040] Construct an intra-modal feature enhancement fusion unit to normalize the private features X and Y of the non-missing modalities, then create a random affine transformation object to perform an affine transformation on the two normalized features to obtain the first affine feature Affine X and the second affine feature Affine Y :
[0041] Affine X = Random_A(Norm(X)) (3)
[0042] Affine Y = Random_A(Norm(Y)) (4) In the above formula, Norm represents the normalization operation; Random_A represents the random affine transformation operation;
[0043] Multiply the first affine feature Affine X by the normalized X feature for weighting to obtain the first weighted feature Mult X, multiply the second Affine feature by the normalized Y feature for weighting to obtain the second weighted feature Mult Y , as follows: Y , specifically as follows:
[0044] Mult X = matmul(Norm(X), Affine X ) (5)
[0045] Mult Y = matmul(Norm(Y), Affine Y ) (6)
[0046] In the above formula, matmul represents the element-wise multiplication operation.
[0047] Then add the first weighted feature to the normalized X feature to obtain the third weighted feature Add X , add the second weighted feature to the normalized Y feature to obtain the fourth weighted feature Addy, specifically as follows:
[0048] Add X = Mult X + Norm(X) (7)
[0049] Add Y = Mult Y + Norm(Y) (8)
[0050] Finally, perform a concatenation operation on the third weighted feature and the fourth weighted feature, and use the ReLU activation function for non-linearity to obtain the fused feature Fusion with enhanced features 2 :
[0051] Fusion 2 = ReLU(Cat(Add X , Add Y )) (9)
[0052] In the above formula, Cat(·) represents the concatenation operation.
[0053] Construct a dynamic spatial-channel fusion unit, which includes a spatial attention unit, a channel attention unit, and a fusion unit.
[0054] First, perform two 1x1 convolution operations on X and Y to obtain the first feature X 1 and the second feature Y 1 , and perform adaptive weighting on the spatial and channel units. The specific formula is as follows:
[0055] X 1= Conv(X) (10)
[0056] Y 1 = Conv(Y) (11)
[0057] In the spatial attention unit, convolutional operations at three scales are used to forcibly capture the spatial relationships of the modalities, obtaining the spatial feature F 1 , and the specific formula is as follows:
[0058] X mul = matmul((Conv5(X)), (Conv7(X))) (12)
[0059] F 1 = X mul + Conv3(X) (13)
[0060] In the above formula, Conv5() represents a 5x5 convolutional operation, Conv7() represents a 7x7 convolutional operation, Conv3() represents a 3x3 convolutional operation, and matmul() represents a feature multiplication operation; X mul represents the spatially enhanced feature.
[0061] In the channel attention unit, the spatial information of the Y feature is captured to obtain the first channel feature Φ[1:C]; specifically, first, the feature is split into two parts along the channel dimension to obtain the first split feature y 1 [1:c] and the second split feature y 2 [c + 1:C]; the specific formula is as follows:
[0062] y 1 [1:c] = Split(1:c) (14)
[0063] y 2 [c + 1:C] = Split(c + 1:C) (15)
[0064] In the above formula, Split represents the splitting operation along the channel dimension, c represents the c-th channel, and C represents the total number of channels of the feature.
[0065] The y 1 [1:c] is first passed through a depthwise separable convolutional residual block to obtain the enhanced first feature, that is The formula is as follows:
[0066]
[0067] In the above formula, IR represents a depthwise separable convolutional residual block.
[0068] The enhanced first feature is added to the second split feature to obtain the third feature Next, the third feature passes through a depthwise separable convolutional residual block and is multiplied by the first segmentation feature to obtain the fourth feature, that is Finally, the third feature and the fourth feature are concatenated to obtain the channel feature Φ[1:C], as shown in the following formula:
[0069]
[0070] In the above formula, IR represents the depthwise separable convolutional residual block; mul represents the element-wise multiplication operation, and Cat represents the concatenation operation.
[0071] In the fusion unit, the spatial feature and the channel feature are fused to obtain the spatial-channel fusion feature Fusion 3 , as shown in the following formula:
[0072] Fusion 3 = ReLU(Cat(F 1 , Φ[1:C])) (21)
[0073] In the above formula, ReLU represents the activation function, Cat represents the concatenation operation, F 1 is the spatial feature, and Φ[1:C] is the channel feature.
[0074] Use dynamic soft routing to perform adaptive routing selection on the simple fusion unit, the intra-modal feature enhancement fusion unit, and the dynamic spatial-channel fusion unit to obtain the hint information of the missing modality.
[0075] First, each channel of Fusion 1 , Fusion 2 and Fusion 3 is made row-sparse, that is, L2 regularization is performed along the spatial dimension, then each channel is made column-sparse, that is, L2 regularization is performed along the channel dimension, and then the sum is taken from the feature dimension to obtain the weighted feature A i , and the formula is as follows:
[0076]
[0077] In the above formula, A i represents the weight result of the three selection units, c, h, and w respectively represent the channel, height, and width sizes of the feature; |||| represents the L2 regularization operation; i represents the i-th fusion feature, j represents the j-th channel, and k represents the element at the k-th position.
[0078] Then, weight assignment is performed on the weighted feature A i to obtain the weight feature weight i , and the formula is as follows:
[0079]
[0080] Finally, multiply the weight feature and the fusion feature Fusion i to perform weighting and obtain the feature R i , and then perform a concatenation operation to obtain the first fusion feature Fusions; as shown in the following formula:
[0081] R i = weight i * Fusion i i = 1, 2, 3 (24)
[0082] Fusions = Cat(R 1 , R 2 , R 3 ) (25)
[0083] Based on the common features extracted in step 3-2, concatenate and fuse them with the first fusion feature passing through the fusion module based on dynamic hint information to generate the feature Z of the missing modality:
[0084] Z = Cat(Depth(X1), Adj(Norm(Fusions))) (26)
[0085] In the above formula, Cat represents the concatenation operation, Depth represents the separable convolution operation on the feature X1, Norm represents the normalization operation, and Adj represents the channel adjustment operation on the feature.
[0086] Step 3-4: Fuse the private features of the non-missing modality and the features of the missing modality to obtain the second fusion feature.
[0087] Step 4: Input the second fusion feature into the classification layer of the overall network model to obtain the predicted label, calculate the difference between the predicted label and the true label through the loss function, and backpropagate the error to update the model parameters, thereby optimizing the overall network model during training, improving the recognition accuracy of the network model, and obtaining the final network model, enabling it to perform sample classification more accurately.
[0088] Step 5: Use the test set to evaluate the final network model and test its accuracy and error rate.
[0089] Example 1
[0090] The multi-modal image recognition method for missing modalities based on dynamic hint information according to the present invention includes the following steps:
[0091] Step 1: Take biometric recognition as an example. Palm, palmprint, and palm vein images of 290 individuals are collected. Each person provides 10 images for each modality, and an image database for three-modal biometric recognition is established. The divided dataset is subjected to scale transformation through the Transforms class in PyTorch and adjusted to a unified size (e.g., 224*224), then consistent normalization processing is performed, and the database is divided into a training set and a test set. The ratio of the number of palm images in the training set to the test set is 1:1.
[0092] Step 2: Construct the overall network model:
[0093] The overall network model includes a common feature extraction network for modalities, a private feature extraction network for modalities, and a fusion module based on dynamic cue information.
[0094] Step 3: Use the common feature extraction network for modalities and the private feature extraction network for modalities to extract common features and private features from the non-missing multi-modal images respectively, and correspondingly obtain the common features of the non-missing modalities and the private features of the non-missing modalities, specifically as follows:
[0095] Step 3-1: Use three Transformer blocks to extract the common features of the non-missing modalities. Each Transformer block includes two normalization layers, an attention layer, and a feed-forward neural network layer to obtain the common features of the non-missing modalities, namely X1 and Y1;
[0096] Step 3-2: Use the first two layers of Resnet34 to extract the private features of the non-missing modalities, including two groups of convolutional layers, a batch normalization layer, a ReLU activation function, a max pooling layer, and two Resnet residual layers to obtain the private features of the non-missing modalities, namely X and Y.
[0097] Step 3-3: According to the difficulty of different samples, use the fusion module based on dynamic cue information to flexibly adjust the fusion strategy, and at the same time avoid using the same fusion strategy for all samples, effectively reducing the computational overhead to obtain the first fusion feature; on the basis of the private features of the non-missing modalities, fuse with the first fusion feature to generate the features of the missing modalities, specifically as follows:
[0098] Construct a simple fusion unit, perform normalization operations on the private features X and Y of the non-missing modalities, then perform a concatenation operation on the two normalized features, and finally use the ReLU activation function to non-linearly process the fusion feature to obtain the fusion feature Fusion 1 :
[0099] Fusion 1 = Cat(Norm(X), Norm(Y)) (1)
[0100] Fusion 1 = ReLU(Fusion 1 ) (2)
[0101] In the above formula, Cat represents the concatenation operation, Norm represents the normalization operation, and ReLU represents the activation function;
[0102] Construct an intra-modal feature enhancement fusion unit. Perform a normalization operation on the private features X and Y of the non-missing modality. Then create a random affine transformation object and perform an affine transformation on the two normalized features to obtain the first affine feature Affine X and the second affine feature Affine Y :
[0103] Affine X = Random_A(Norm(X)) (3)
[0104] Affine Y = Random_A(Norm(Y)) (4) In the above formula, Norm represents the normalization operation; Random_A represents the random affine transformation operation;
[0105] Multiply the first affine feature Affine X by the normalized X feature for weighting to obtain the first weighted feature Mult X , multiply the second affine feature Affine Y by the normalized Y feature for weighting to obtain the second weighted feature Mult Y , specifically as follows:
[0106] Mult X = matmul(Norm(X), Affine X ) (5)
[0107] Mult Y = matmul(Norm(Y), Affine Y ) (6)
[0108] In the above formula, matmul represents the element-wise multiplication operation;
[0109] Then add the first weighted feature to the normalized X feature to obtain the third weighted feature Add X , add the second weighted feature to the normalized Y feature to obtain the fourth weighted feature Addy, specifically as follows:
[0110] Add X = Mult X + Norm(X) (7)
[0111] Add Y = Mult Y + Norm(Y) (8)
[0112] Finally, connect the third weighted feature and the fourth weighted feature, and use the ReLU activation function for non - linearization to obtain the fused feature Fusion with enhanced features 2 :
[0113] Fusion 2 = ReLU(Cat(Add X , Add Y )) (9)
[0114] In the above formula, Cat(·) represents the connection operation.
[0115] Construct a dynamic spatial - channel fusion unit, which includes a spatial attention unit, a channel attention unit, and a fusion unit.
[0116] First, perform two 1x1 convolution operations on X and Y to obtain the first feature X 1 and the second feature Y 1 , and perform adaptive weighting on the spatial and channel units. The specific formula is as follows:
[0117] X 1 = Conv(X) (10)
[0118] Y 1 = Conv(Y) (11)
[0119] In the spatial attention unit, use convolution operations of three scales to forcibly capture the spatial relationship of the modality to obtain the spatial feature F 1 , and the specific formula is as follows:
[0120] X mul = matmul((Conv5(X)), (Conv7(X))) (12)
[0121] F 1 = X mul + Conv3(X) (13)
[0122] In the above formula, Conv5() represents a 5x5 convolution operation, Conv7() represents a 7x7 convolution operation, Conv3() represents a 3x3 convolution operation, matmul() represents a feature multiplication operation; X mul represents the spatially enhanced feature;
[0123] The spatial information of feature Y is captured in the channel attention unit to obtain the first channel feature Φ[1:C]; specifically, the feature is first split into two parts from the channel dimension to obtain the first segmentation feature y 1 [1:c] and the second segmentation feature y 2 [c+1:C]; the specific formula is as follows:
[0124] y 1 [1:c]=Split(1:c) (14)
[0125] y 2 [c+1:C]=Split(c+1:C) (15)
[0126] In the above formula, Split represents the splitting operation from the channel dimension, c represents the cth channel, and C represents the total number of channels of the feature.
[0127] y 1 [1:c] First, a separable convolution depth residual block is passed to obtain the enhanced first feature, that is, The formula is as follows:
[0128]
[0129] In the above formula, IR represents the depth residual block based on separable convolution.
[0130] Add the enhanced first feature to the second segmentation feature to obtain the third feature Next, the third feature passes through a separable convolutional depth residual block and is multiplied with the first segmentation feature to obtain the fourth feature, i.e. Finally, the third feature and the fourth feature are connected to obtain the channel feature Φ[1:C], as shown in the following formula:
[0131]
[0132] In the above formula, IR represents the depth residual block based on separable convolution; mul represents the element multiplication operation, and Cat represents the connection operation;
[0133] In the fusion unit, the spatial features and channel features are fused to obtain the spatial channel fusion feature Fusion 3 , as shown below:
[0134] Fusion 3 =ReLU(Cat(F 1 , Φ[1:C])) (21)
[0135] In the above formula, ReLU represents the activation function, Cat represents the connection operation, and F 1is the spatial feature, and Φ[1:C] is the channel feature.
[0136] Use dynamic soft routing to perform adaptive routing selection on the simple fusion unit, the intra-modal feature enhancement fusion unit, and the dynamic spatial-channel fusion unit to obtain the hint information of the missing modality.
[0137] First, perform row sparsity on each channel of Fusion 1 , Fusion 2 and Fusion 3 , that is, perform L2 regularization along the spatial dimension, then perform column sparsity on each channel, that is, perform L2 regularization along the channel dimension, and then sum from the feature dimension to obtain the weighted feature A i , and the formula is as follows:
[0138]
[0139] In the above formula, A i represents the weight results of the three selection units, c, h, and w respectively represent the channel, height, and width sizes of the feature; |||| represents the L2 regularization operation; i represents the i-th fusion feature, j represents the j-th channel, and k represents the element at the k-th position.
[0140] Then, perform weight assignment on the weighted feature A i to obtain the weighted feature weight i , and the formula is as follows:
[0141]
[0142] Finally, multiply the weighted feature by the fusion feature Fusion i for weighting to obtain the feature R i , and then perform a concatenation operation to obtain the first fusion feature Fusions; as shown in the following formula:
[0143] R i = weight i * Fusion i i = 1, 2, 3 (24)
[0144] Fusions = Cat(R 1 , R 2 , R 3 ) (25)
[0145] Based on the common features extracted in step 3-2, concatenate and fuse with the first fusion feature passing through the fusion module based on the dynamic hint information to generate the feature Z of the missing modality:
[0146] Z = Cat(Depth(X1), Adj(Norm(Fusions))) (26)
[0147] In the above formula, Cat represents the concatenation operation, Depth represents the separable convolution operation on feature X1, Norm represents the normalization operation, and Adj represents the channel adjustment operation on the feature.
[0148] Step 3-4: Fuse the private features of the non-missing modality and the features of the missing modality to obtain the second fused feature.
[0149] Step 4: Input the second fused feature into the classification layer of the overall network model to obtain the predicted label. Calculate the difference between the predicted label and the true label through the loss function, and backpropagate the error to update the model parameters, so as to optimize the overall trained network model, improve the recognition accuracy of the network model, and obtain the final network model, enabling it to classify samples more accurately.
[0150] Step 5: Use the test set to evaluate the final network model and test its accuracy and error rate.
[0151] The present invention conducts experiments on the NVIDIA GeForce GTX 1650 GPU. It is implemented in Python 3.6.9 using the PyTorch framework. The network is pre-trained on the ImageNet database and the network parameters are fine-tuned on the dataset. The Adam optimizer is used during the training process. The learning rate, batch size, training epochs, and weight decay of the network are set to 1e -3 、4、300 and 1e -3 。 We compare the method proposed in the present invention with biometric recognition and algorithms related to modality missing, including ShaSpecNet, DVMAN, and DENet, as the comparison models.
[0152] Table 1 Comparison experiment results
[0153]
[0154] According to the experimental results shown in Table 1, the method proposed in the present invention performs stably under different missing rates and can achieve the optimal recognition effect, with good model robustness.
Claims
1. A multimodal image recognition method based on missing modality of dynamic prompt information, characterized in that: include: Step 1: Collect a number of palm print and palm vein images to construct a multimodal palm print and palm vein image dataset, pre-process the dataset, and then divide it into a training set and a test set in a 1:1 ratio; Step 2: Build the overall network model: The overall network model includes a common feature extraction network for each modality, a private feature extraction network for each modality, and a fusion module based on dynamic prompt information; Step 3: Train the overall network model: Step 3-1, use the training set to train the overall network model, perform random mode loss during training, and obtain missing modes and non-missing modes; Step 3-2, using the common feature extraction network of the modality to extract features from the non-missing modality, to obtain the common features of the non-missing modality, and using the private feature extraction network of the modality to extract features from the non-missing modality, to obtain the private features of the non-missing modality; Step 3-3, according to the difficulty of different samples, the fusion strategy is flexibly adjusted using the fusion module based on dynamic prompt information to obtain the first fusion feature; Based on the private features of the non-missing modality, the first fused features are fused to generate features of the missing modality; Among them: a simple fusion unit is constructed, the private features X and Y of the non-missing mode are normalized, and then the two normalized features are connected, and finally the ReLU activation function is used to nonlinearize the fused features to obtain the fused features; Construct an intra-modality feature enhancement fusion unit, normalize the private features X and Y of the non-missing modality, then create a random affine transformation object, perform an affine transformation on the two normalized features, and obtain the first affine feature and the second affine feature: The first affine feature is multiplied by the normalized X feature to obtain a first weighted feature, and the second affine feature is multiplied by the normalized Y feature to obtain a second weighted feature. Then, the first weighted feature is added to the normalized X feature to obtain the third weighted feature, and the second weighted feature is added to the normalized Y feature to obtain the fourth weighted feature; Finally, the third weighted feature is connected with the fourth weighted feature, and the ReLU activation function is used for nonlinearization to obtain the fusion feature after feature enhancement, and a dynamic space-channel fusion unit is constructed. The dynamic space-channel fusion unit includes a spatial attention unit, a channel attention unit and a fusion unit; Step 3-4: Fuse the private features of the non-missing modality with the features of the missing modality to obtain the second fused features. Step 4: Input the second fusion feature into the classification layer of the overall network model to obtain the predicted label, calculate the difference between the predicted label and the true label through the loss function, and back-propagate the error to update the model parameters, thereby optimizing the trained overall network model, improving the recognition accuracy of the network model, and obtaining the final network model; Step 5: Use the test set to evaluate the final network model and test its accuracy and error rate.
2. The multimodal image recognition method based on missing modality of dynamic prompt information according to claim 1, characterized in that: In step 1, a number of palm print and palm vein images are collected to construct a multimodal palm print and palm vein image dataset, and the dataset is preprocessed and then divided into a training set and a test set, as follows: Several palm print and palm vein images are collected to construct a multimodal palm print and palm vein image dataset. The dataset is scaled using the Transforms class in PyTorch to adjust it to a uniform size, and then normalized consistently. The normalized dataset is divided into a training set and a test set at a ratio of 1:
1.
3. The multimodal image recognition method based on missing modality of dynamic prompt information according to claim 1, characterized in that: In step 3-2, the common feature extraction network of the modality is used to extract features from the non-missing modality to obtain the common features of the non-missing modality, and the private feature extraction network of the modality is used to extract features from the non-missing modality to obtain the private features of the non-missing modality, as follows: The common feature extraction network is composed of three Transformer blocks, each of which includes two normalization layers, an attention layer, and a feedforward neural network layer. The common features of the non-missing modalities are extracted through the common feature extraction network to obtain the common features of the non-missing modalities, namely X1; The private feature extraction network consists of the first two layers of Resnet34, including two groups of convolutional layers, a batch normalization layer, a ReLU activation function, a maximum pooling layer, and two Resnet residual layers. The private feature extraction network is used to extract the private features of the non-missing modality to obtain the private features of the non-missing modality, namely X and Y.
Citation Information
Patent Citations
Incomplete multi-modal medical image learning method
CN117218453A