Multi-modal image fusion and identification method
By introducing multi-head attention mechanism and learnable parameter adjustment mechanism in multi-modal image fusion technology, dynamically fusion and adjustment of modal features are solved, and the problems of poor fusion effect and uncaptured modal dependencies in the existing technology are achieved, achieving more efficient and flexible feature representation and recognition performance.
Patent Information
- Application Number
- CN202510181350.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-24
AI Technical Summary
The existing multimodal image fusion technology is difficult to dynamically adjust the information content of different modal features, resulting in poor fusion effect and failure to fully capture the complex dependencies between modals, affecting the recognition performance.
The multi-head attention mechanism and the learnable parameter adjustment mechanism are adopted to dynamically fuse the characteristics of different modalities, automatically allocate weights by calculating the correlation between modalities, and adjust the contribution of each modality through learnable parameters to capture complex dependencies and adapt to different scenarios.
It significantly improves the richness and accuracy of feature representation, enhances the adaptability and generalization ability of the model, and improves the accuracy and robustness of multimodal image fusion and recognition.
Smart Images

Figure CN120198753A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image fusion, and specifically relates to a multi-modal image fusion and recognition method. Background Art
[0002] With the progress of technology, multi-modal image fusion and recognition technology has shown great application potential in many fields, such as security monitoring, autonomous driving, medical diagnosis, etc. Multi-modal images, including visible light images, infrared images, radar images, etc., each carry unique information and can complementarily describe the same scene or target. However, how to effectively fuse these different modal image data to achieve accurate and robust recognition has always been a difficult point in the technical field; multi-modal image fusion and recognition technology refers to the technology of effectively combining and comprehensively processing multiple modal image information obtained from different sensors or the same sensor under different conditions. This technology improves the clarity, integrity, and reliability of images by extracting and integrating the feature information of different image sources, thereby achieving more accurate recognition and analysis of target objects. Multi-modal image fusion methods include pixel-level, feature-level, and decision-level fusion, and recognition technology involves multiple fields such as pattern recognition, machine learning, and deep learning.
[0003] However, existing technologies often adopt fixed fusion strategies, such as simple weighted averaging or feature splicing, and cannot dynamically adjust according to the actual information content of different modal features. This results in over-reliance on modalities with less information and neglect of modalities with rich information during the fusion process, thus affecting the fusion effect.
[0004] Failure to fully consider the complex dependence relationships between modalities: At the same time, there may be complex non-linear dependence relationships between different modal features. The lack of capturing these relationships will result in incomplete feature representations after fusion and affect the recognition performance. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-modal image fusion and recognition method to solve the above-mentioned problems.
[0006] The technical solution adopted by the present invention is as follows: A multi-modal image fusion and recognition method, which includes the following steps:
[0007] S1: Collect multi-modal image data, and preprocess the images of each modality, including denoising, normalization, and alignment operations, to ensure the consistency of different modal images in terms of space and intensity;
[0008] S2: Feature extraction: Use a pre-trained deep learning model to extract high-level features of each modality image respectively; for each modality, the extracted features include local features and global features to capture the details and overall information of the image;
[0009] S3: Design a cross-modal attention mechanism for dynamically fusing features of different modalities; by calculating the correlation between features of different modalities, automatically assign weights to pay more attention to the modality with rich information during the fusion process; use the multi-head attention mechanism to capture the complex dependencies between modalities and adjust the contribution degree of each modality through learnable parameters;
[0010] S4: Feature fusion: fuse the features weighted by the cross-modal attention mechanism; the fusion methods include weighted summation, concatenation or methods based on graph neural networks to ensure that the fused features can fully express the complementarity of multi-modal information;
[0011] S5: Feature enhancement: further enhance the fused features, using self-attention mechanism or convolutional neural network to enhance the expressive power of features, highlight important feature information, and suppress redundant information;
[0012] S6: Classifier design: design a multi-task classifier for classifying the fused features; the classifier includes fully connected layers and Softmax layers to output the probability distribution of each class; for different tasks, design different classification heads;
[0013] S7: Loss function design: design a multi-task loss function, combining cross-entropy loss and contrastive loss to ensure that the model can achieve good performance on multiple tasks; the loss function should consider the contribution degree of different modalities and balance the losses of each task through weighting;
[0014] S8: Model training and optimization: use the backpropagation algorithm to train the model and optimize the loss function; adopt Adam and SGD optimizers and combine with the learning rate decay strategy to ensure that the model can converge quickly and avoid overfitting during the training process;
[0015] S9: Post-processing and result analysis: post-process the output of the model to improve the recognition accuracy; finally, conduct visual analysis on the recognition results, evaluate the performance of the model under different modalities, and further adjust the model parameters or structure according to the results.
[0016] In a preferred embodiment, in step S1, multi-modal image data is first collected from multiple sensors, specifically including high-resolution visible light images, thermally sensitive infrared images, and radar images with strong penetration; for each modality of images, detailed preprocessing operations are performed: an adaptive median filter is used to denoise the visible light images to preserve edge details; histogram equalization is applied to the infrared images to enhance contrast; speckle noise suppression is performed on the radar images using a Lee filter; then, spatial alignment of different modality images is achieved through affine transformation and interpolation algorithms to ensure their precise correspondence at the pixel level; finally, the pixel values of all images are normalized to the range of 0 to 1 to eliminate intensity differences between different sensors.
[0017] In a preferred embodiment, in step S2, a pre-trained ResNet-50 model is used as a feature extractor. First, the visible light images are input into the ResNet-50 model, and 2048-dimensional local feature vectors are extracted through the convolutional and pooling layers of the model. These vectors capture the edge and texture detail information of the images; at the same time, 1024-dimensional global feature vectors are extracted using the global average pooling layer of the model. These vectors represent the overall information of the shape and contour of the images; similarly, the same operations are performed on the infrared images and radar images to extract the corresponding modality feature vectors respectively.
[0018] In a preferred embodiment, in step S3, the specific steps include:
[0019] S3-1: Initialization of multi-head attention mechanism:
[0020] Suppose there are N modalities, and each modality is represented as M_i (i = 1, 2,..., N);
[0021] For each modality M_i, multiple query (Q), key (K), and value (V) matrices are initialized, denoted as Q_i, K_i, V_i respectively;
[0022] These matrices are obtained from the features of each modality through linear transformation, that is, Q_i = W_Q * M_i, K_i = W_K * M_i, V_i = W_V * M_i, where W_Q, W_K, W_V are learnable weight matrices;
[0023] S3-2: Calculate intra-modal attention:
[0024] For each modality M_i, calculate the self-attention score:
[0025] A_i = softmax(Q_i * K_i^T / sqrt(d_k))
[0026] Among them, d_k is the dimension of the key, which is used to scale the dot product result to prevent gradient vanishing or explosion;
[0027] Calculate the feature representation after self-attention:
[0028] O_i = A_i * V_i;
[0029] S3-3: Calculate cross-modal attention:
[0030] Calculate the attention score between modalities, the attention from modality M_i to modality M_j:
[0031] A_i->j = softmax(Q_i * K_j^T / sqrt(d_k))
[0032] Calculate the feature representation after cross-modal attention:
[0033] O_i->j = A_i->j * V_j
[0034] S3-4: Multi-head attention fusion:
[0035] Concatenate or average the outputs of multiple heads to obtain the fused features of each modality:
[0036] O_i^final = Concat(O_i, O_i->1, O_i->2,..., O_i->N) # Concatenation method
[0037] Or
[0038] O_i^final = Mean(O_i, O_i->1, O_i->2,..., O_i->N) # Averaging method
[0039] S3-5: Learnable parameter adjustment: Introduce learnable parameter γ_i to adjust the contribution degree of each modality:
[0040] F = γ_1 * O_1^final + γ_2 * O_2^final +... + γ_N * O_N^final
[0041] These parameters are learned through backpropagation during training.
[0042] In a preferred embodiment, in step S4, a cross-modal feature fusion method based on the attention mechanism is adopted; first, the feature vectors extracted from different modalities are concatenated to form a long vector; then, a multi-layer perceptron is used to learn the weight relationship between different modality features; specifically, the input of the MLP is the concatenated feature vector, and the output is the weight of each modality feature; these weights reflect the importance of different modality features in the recognition task; then, the weights are multiplied by the corresponding modality feature vectors and summed to obtain the fused feature vector; this method can dynamically adjust the contribution degrees of different modality features and achieve more effective feature fusion.
[0043] In a preferred embodiment, in step S5, a support vector machine is selected as the classifier because it performs well in dealing with high-dimensional features and small sample problems; first, the fused feature vector is divided into a training set and a validation set; then, the training set data is used to train the SVM model, and the optimal model parameters, including the kernel function and the penalty coefficient, are selected through cross-validation; during the training process, the gradient descent algorithm is used to optimize the model parameters to minimize the classification error.
[0044] In a preferred embodiment, in step S6, the trained SVM model is evaluated using the validation set; the evaluation metrics include accuracy, recall, and F1 score; specifically, the matching degree between the prediction result of the model on the validation set and the true label is calculated to obtain the accuracy; the recall reflects the ability of the model to identify positive samples; the F1 score is the harmonic mean of the accuracy and the recall, comprehensively reflecting the performance of the model; in addition, a confusion matrix is also drawn to visually show the recognition effect of the model on different categories.
[0045] In a preferred embodiment, in step S7, if it is found that the recognition effect of the model on a certain category is not good, try to increase the training samples of this category, or adjust the kernel function and penalty coefficient of the SVM; in addition, more complex models are also considered to improve the feature extraction and classification capabilities of the model; during the optimization process, continuously monitor the performance metrics on the validation set to ensure that the model does not overfit.
[0046] In a preferred embodiment, in step S8, non-maximum suppression is used to remove redundant detection boxes. The specific approach is as follows: for each detection box, calculate its intersection over union (IoU) with other detection boxes. If the IoU exceeds the set threshold, then keep the detection box with the highest confidence and suppress other detection boxes; in addition, a confidence threshold is also set to filter out the detection results below this threshold, thereby reducing false recognition.
[0047] In a preferred embodiment, in step S9, a depthwise separable convolutional layer is introduced into the feature extraction network, and its parameters include the convolutional kernel size, the feature map size, the number of input channels, and the number of output channels (N); such a convolutional layer can effectively reduce the number of parameters and the amount of computation, while enhancing the model's ability to capture detailed features.
[0048] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:
[0049] 1. In the present invention, by introducing the multi-head attention mechanism, the model's ability to capture complex dependencies between modalities is significantly enhanced. In multi-modal image fusion and recognition tasks, image data of different modalities often contain complementary information. For example, visible light images are rich in details, infrared images are thermally sensitive, and radar images have penetration power. The multi-head attention mechanism can simultaneously focus on the features of multiple modalities and learn the internal connections between the features of different modalities through the self-attention mechanism. This design enables the model to more flexibly fuse multi-modal information, effectively improving the richness and accuracy of feature representation. For example, in complex scenarios, the model can use the thermally sensitive information of infrared images to assist in target recognition in visible light images, thereby improving the accuracy and robustness of recognition.
[0050] 2. In the present invention, the learnable parameter adjustment mechanism provides strong support for the dynamic optimization of the contribution degrees of each modality. In practical applications, the importance of image data of different modalities may vary in different scenarios. Through the learnable parameters, the model can automatically adjust the contribution degrees of each modality during the training process, so that the modality features that are more important in a specific scenario are strengthened. This dynamic adjustment mechanism not only improves the adaptability of the model but also enhances the generalization ability of the model to diverse scenarios. For example, in nighttime or low-light conditions, infrared images may provide more effective information than visible light images, and the model will automatically increase the weight of the infrared modality through learning, thereby maintaining high recognition performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a schematic diagram of the process principle of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] Embodiment:
[0054] Referring to Figure 1 , a multi-modal image fusion and recognition method includes the following steps:
[0055] S1: Collect multimodal image data (such as visible light, infrared, radar, etc.), and preprocess the images of each modality, including operations such as denoising, normalization, alignment, etc., to ensure the consistency of different modality images in space and intensity.
[0056] S2: Feature extraction: Use pre-trained deep learning models (such as ResNet, VGG, etc.) to extract high-level features of each modality image respectively. For each modality, the extracted features include local features and global features to capture the details and overall information of the image.
[0057] S3: Design a cross-modal attention mechanism for dynamically fusing features of different modalities. By calculating the correlation between features of different modalities, automatically assign weights to pay more attention to the modality with rich information during the fusion process. Use the multi-head attention mechanism to capture the complex dependencies between modalities, and adjust the contribution degree of each modality through learnable parameters.
[0058] S4: Feature fusion: Fuse the features weighted by the cross-modal attention mechanism. The fusion method can adopt weighted summation, concatenation or graph neural network-based methods to ensure that the fused features can fully express the complementarity of multimodal information.
[0059] S5: Feature enhancement: Further enhance the fused features, using self-attention mechanism or convolutional neural network (CNN) to enhance the expression ability of features, highlight important feature information, and suppress redundant information.
[0060] S6: Classifier design: Design a multi-task classifier for classifying the fused features. The classifier can include fully connected layers, Softmax layers, etc., and output the probability distribution of each class. For different tasks (such as object recognition, scene classification, etc.), different classification heads can be designed.
[0061] S7: Loss function design: Design a multi-task loss function, combining cross-entropy loss, contrastive loss, etc., to ensure that the model can achieve good performance on multiple tasks. The loss function should consider the contribution degree of different modalities and balance the losses of each task through weighting.
[0062] S8: Model training and optimization: Use the backpropagation algorithm to train the model and optimize the loss function. Optimizers such as Adam, SGD, etc. can be adopted, and combined with the learning rate decay strategy to ensure that the model can converge quickly during training and avoid overfitting.
[0063] S9: Post - processing and result analysis: Post - process the output of the model, such as non - maximum suppression or threshold screening, to improve the recognition accuracy. Finally, perform visual analysis on the recognition results, evaluate the performance of the model in different modalities, and further adjust the model parameters or structure according to the results.
[0064] 2. A multi - modal image fusion and recognition method as claimed in claim 1, wherein: In step S1, first collect multi - modal image data from multiple sensors, specifically including high - resolution visible light images, thermally sensitive infrared images, and radar images with strong penetration. For each modality of images, perform detailed pre - processing operations: Use an adaptive median filter to denoise the visible light image to preserve edge details; apply histogram equalization to the infrared image to enhance contrast; perform speckle noise suppression on the radar image using the Lee filter. Then, achieve spatial alignment of different modality images through affine transformation and interpolation algorithms to ensure their precise correspondence at the pixel level. Finally, normalize the pixel values of all images to the range of 0 to 1 to eliminate intensity differences between different sensors.
[0065] In step S2, use a pre - trained ResNet - 50 model as the feature extractor. First, input the visible light image into the ResNet - 50 model, and extract 2048 - dimensional local feature vectors through the convolutional and pooling layers of the model. These vectors capture details such as edges and textures of the image. At the same time, use the global average pooling layer of the model to extract 1024 - dimensional global feature vectors, which represent overall information such as the shape and contour of the image. Similarly, perform the same operations on the infrared image and the radar image to extract the corresponding modality feature vectors respectively.
[0066] In step S3, the specific steps include:
[0067] S3 - 1: Initialization of multi - head attention mechanism:
[0068] Suppose there are N modalities, and each modality is represented as M_i (i = 1, 2, …, N).
[0069] For each modality M_i, initialize multiple query (Q), key (K), and value (V) matrices, denoted as Q_i, K_i, V_i respectively.
[0070] These matrices are obtained from the features of each modality through linear transformation, that is, Q_i = W_Q * M_i, K_i = W_K * M_i, V_i = W_V * M_i, where W_Q, W_K, W_V are learnable weight matrices.
[0071] S3 - 2: Calculate intra - modality attention:
[0072] For each modality \(M_i\), calculate the self-attention scores:
[0073] \(A_i=\text{softmax}(Q_i * K_i^T / \sqrt{d_k})\)
[0074] where \(d_k\) is the dimension of the key, used to scale the dot product result to prevent gradient vanishing or explosion.
[0075] Calculate the feature representation after self-attention:
[0076] \(O_i = A_i * V_i\);
[0077] S3-3: Calculate cross-modal attention:
[0078] Calculate the attention scores between modalities, e.g., the attention from modality \(M_i\) to modality \(M_j\):
[0079] \(A_{i\rightarrow j}=\text{softmax}(Q_i * K_j^T / \sqrt{d_k})\)
[0080] Calculate the feature representation after cross-modal attention:
[0081] \(O_{i\rightarrow j}=A_{i\rightarrow j} * V_j\)
[0082] S3-4: Multi-head attention fusion:
[0083] Concatenate or average the outputs of multiple heads to obtain the fused features for each modality:
[0084] \(O_i^{\text{final}}=\text{Concat}(O_i, O_{i\rightarrow 1}, O_{i\rightarrow 2},..., O_{i\rightarrow N})\) # Concatenation method
[0085] Or
[0086] \(O_i^{\text{final}}=\text{Mean}(O_i, O_{i\rightarrow 1}, O_{i\rightarrow 2},..., O_{i\rightarrow N})\) # Averaging method
[0087] S3-5: Learnable parameter adjustment: Introduce learnable parameters \(\gamma_i\) to adjust the contribution degree of each modality:
[0088] \(F = \gamma_1 * O_1^{\text{final}}+\gamma_2 * O_2^{\text{final}}+...+\gamma_N * O_N^{\text{final}}\)
[0089] These parameters are learned through backpropagation during training.
[0090] In step S4, a cross-modal feature fusion method based on the attention mechanism is adopted. First, the feature vectors extracted from different modalities are concatenated to form a long vector. Then, a multi-layer perceptron (MLP) is used to learn the weight relationships between the features of different modalities. Specifically, the input of the MLP is the concatenated feature vector, and the output is the weight of each modality feature. These weights reflect the importance of different modality features in the recognition task. Next, the weights are multiplied by the corresponding modality feature vectors and summed to obtain the fused feature vector. This method can dynamically adjust the contribution degrees of different modality features and achieve more effective feature fusion.
[0091] In step S5, a support vector machine (SVM) is selected as the classifier because it performs well in dealing with high-dimensional features and small-sample problems. First, the fused feature vector is divided into a training set and a validation set. Then, the SVM model is trained using the training set data, and the optimal model parameters, such as the kernel function and penalty coefficient, are selected through cross-validation. During the training process, the gradient descent algorithm is used to optimize the model parameters to minimize the classification error.
[0092] In step S6, the trained SVM model is evaluated using the validation set. The evaluation metrics include accuracy, recall, and F1-score. Specifically, the matching degree between the prediction results of the model on the validation set and the true labels is calculated to obtain the accuracy; the recall reflects the ability of the model to identify positive samples; the F1-score is the harmonic mean of the accuracy and recall, comprehensively reflecting the performance of the model. In addition, a confusion matrix is drawn to visually show the recognition effect of the model on different categories.
[0093] In step S7, if it is found that the recognition effect of the model on a certain category is not good, one can try to increase the training samples of this category, or adjust the kernel function and penalty coefficient of the SVM. In addition, more complex models, such as deep neural networks, can also be considered to improve the feature extraction and classification capabilities of the model. During the optimization process, continuously monitor the performance metrics on the validation set to ensure that the model does not overfit.
[0094] In step S8, non-maximum suppression (NMS) is adopted to remove redundant detection boxes. The specific approach is as follows: for each detection box, calculate its intersection over union (IoU) with other detection boxes. If the IoU exceeds the set threshold (such as 0.5), then retain the detection box with the highest confidence and suppress other detection boxes. In addition, a confidence threshold is also set to filter out the detection results below this threshold, thereby reducing misidentifications.
[0095] In step S9, a depthwise separable convolutional layer is introduced into the feature extraction network, and its parameters include the convolutional kernel size (Dk), the feature map size (Df), the number of input channels (M), and the number of output channels (N). This convolutional layer can effectively reduce the number of parameters and the computational amount, while enhancing the model's ability to capture detailed features.
[0096] As can be seen from the above: In the present invention, by introducing the multi-head attention mechanism, the model's ability to capture complex interdependencies between modalities is significantly enhanced. In multi-modal image fusion and recognition tasks, image data of different modalities often contain complementary information, such as the rich details of visible light images, the thermal sensitivity of infrared images, and the penetration power of radar images. The multi-head attention mechanism can simultaneously focus on the features of multiple modalities and learn the internal relationships between the features of different modalities through the self-attention mechanism. This design enables the model to more flexibly fuse multi-modal information, effectively improving the richness and accuracy of feature representation. For example, in complex scenarios, the model can utilize the thermal-sensitive information of infrared images to assist in target recognition in visible light images, thereby improving the accuracy and robustness of recognition.
[0097] In the present invention, the learnable parameter adjustment mechanism provides strong support for the dynamic optimization of the contribution degrees of each modality. In practical applications, the importance of image data of different modalities may vary in different scenarios. Through the learnable parameters, the model can automatically adjust the contribution degrees of each modality during the training process, enabling the modality features that are more important in a specific scenario to be strengthened. This dynamic adjustment mechanism not only improves the model's adaptability but also enhances the model's generalization ability to diverse scenarios. For example, in nighttime or low-light conditions, infrared images may provide more effective information than visible light images, and the model will automatically increase the weight of the infrared modality through learning, thereby maintaining a high recognition performance.
[0098] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article, or device comprising the element.
[0099] The foregoing description enables those skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal image fusion and recognition method, characterized in that: The steps include: S1: Collect multimodal image data and preprocess the images of each modality, including denoising, normalization, and alignment operations to ensure that the images of different modalities are consistent in space and intensity; S2: Feature extraction: Use a pre-trained deep learning model to extract high-level features of each modality image. For each modality, the extracted features include local features and global features to capture the details and overall information of the image. S3: Design a cross-modal attention mechanism to dynamically fuse features from different modalities; automatically assign weights by calculating the correlation between features from different modalities, so that more attention is paid to the information-rich modality during the fusion process; capture the complex dependencies between modalities by using a multi-head attention mechanism, and adjust the contribution of each modality through learnable parameters; S4: Feature fusion: The features weighted by the cross-modal attention mechanism are fused; the fusion method adopts weighted summation, splicing or graph neural network-based methods to ensure that the fused features can fully express the complementarity of multimodal information; S5: Feature enhancement: Further enhance the fused features, use self-attention mechanism or convolutional neural network to enhance the expressiveness of features, highlight important feature information, and suppress redundant information; S6: Classifier design: Design a multi-task classifier to classify the fused features; the classifier includes a fully connected layer and a Softmax layer, and outputs the probability distribution of each category; different classification heads are designed for different tasks; S7: Loss function design: Design a multi-task loss function, combining cross entropy loss and contrast loss to ensure that the model can achieve good performance on multiple tasks; the loss function should consider the contribution of different modalities and balance the loss of each task in a weighted manner; S8: Model training and optimization: Use the back propagation algorithm to train the model and optimize the loss function; use Adam and SGD optimizers, combined with the learning rate decay strategy, to ensure that the model can converge quickly during training and avoid overfitting; S9: Post-processing and result analysis: Post-process the output of the model to improve recognition accuracy; finally, visualize and analyze the recognition results to evaluate the performance of the model in different modes, and further adjust the model parameters or structure based on the results.
2. A multimodal image fusion and recognition method as claimed in claim 1, characterized in that: In the step S1, multimodal image data are first collected from a variety of sensors, specifically including high-resolution visible light images, heat-sensitive infrared images, and radar images with strong penetration; detailed preprocessing operations are performed on images of each modality: an adaptive median filter is used to denoise the visible light image to retain edge details; histogram equalization is applied to the infrared image to enhance contrast; speckle noise is suppressed on the radar image using a Lee filter; then, spatial alignment of images of different modalities is achieved through affine transformation and interpolation algorithms to ensure their precise correspondence at the pixel level; finally, the pixel values of all images are normalized to a range of 0 to 1 to eliminate intensity differences between different sensors.
3. The multimodal image fusion and recognition method according to claim 1, characterized in that: In step S2, a pre-trained ResNet-50 model is used as a feature extractor. First, the visible light image is input into the ResNet-50 model, and 2048-dimensional local feature vectors are extracted through the convolution layer and pooling layer of the model. These vectors capture the edge and texture detail information of the image; at the same time, the global average pooling layer of the model is used to extract 1024-dimensional global feature vectors. These vectors represent the overall shape and contour information of the image; similarly, the same operation is performed on the infrared image and the radar image to extract the feature vectors of the corresponding modality respectively.
4. The multimodal image fusion and recognition method according to claim 1, characterized in that: In step S3, the specific steps include: S3-1: Initialization of multi-head attention mechanism: Assume there are N modes, each mode is represented by Mi (i = 1, 2, ..., N); For each modality Mi, initialize multiple query (Q), key (K), and value (V) matrices, denoted as Qi, Ki, and Vi respectively; These matrices are obtained from the features of each modality through linear transformation, i.e., Q_i=W_Q*M_i, K_i=W_K*M_i, V_i=W_V*M_i, where W_Q, W_K, W_V are learnable weight matrices; S3-2: Calculate intra-modal attention: For each modality Mi, calculate the self-attention score: A_i=softmax(Q_i*K_i^T / sqrt(d_k)) Where d_k is the dimension of the key, which is used to scale the dot product result to prevent the gradient from disappearing or exploding; Calculate the feature representation after self-attention: O_i=A_i*V_i; S3-3: Calculating cross-modal attention: Calculate the attention score between modalities, from modality Mi to modality Mj: A_i->j=softmax(Q_i*K_j^T / sqrt(d_k)) Calculate the feature representation after cross-modal attention: O_i->j=A_i->j*V_j S3-4: Multi-head attention fusion: The outputs of multiple heads are concatenated or averaged to obtain the fusion features of each modality: O_i^final=Concat(O_i,O_i->1,O_i->2,...,O_i->N)#Splicing method or O_i^final=Mean(O_i,O_i->1,O_i->2,...,O_i->N)#Averaging method S3-5: Learnable parameter adjustment: Introduce learnable parameters γ_i to adjust the contribution of each mode: F=γ_1*O_1^final+γ_2*O_2^final+...+γ_N*O_N^final These parameters are learned during training via back-propagation.
5. The multimodal image fusion and recognition method according to claim 1, characterized in that: In the step S4, a cross-modal feature fusion method based on the attention mechanism is adopted; first, the feature vectors extracted from different modalities are spliced to form a long vector; then, the weight relationship between different modal features is learned through a multi-layer perceptron; specifically, the input of the MLP is the spliced feature vector, and the output is the weight of each modal feature; these weights reflect the importance of different modal features in the recognition task; then, the weight is multiplied by the feature vector of the corresponding modality, and the sum is taken to obtain the fused feature vector; this method can dynamically adjust the contribution of different modal features to achieve more effective feature fusion.
6. The multimodal image fusion and recognition method according to claim 1, characterized in that: In step S5, a support vector machine is selected as a classifier because it performs well in processing high-dimensional features and small sample problems; first, the fused feature vector is divided into a training set and a validation set; then, the training set data is used to train the SVM model, and the optimal model parameters, including the kernel function and the penalty coefficient, are selected through cross-validation; during the training process, the gradient descent algorithm is used to optimize the model parameters to minimize the classification error.
7. The multimodal image fusion and recognition method according to claim 1, characterized in that: In step S6, the trained SVM model is evaluated using a validation set; the evaluation indicators include accuracy, recall and F1 score; specifically, the accuracy is obtained by calculating the degree of match between the prediction results of the model on the validation set and the true label; the recall rate reflects the ability of the model to identify positive samples; the F1 score is the harmonic mean of the accuracy and recall rate, which comprehensively reflects the performance of the model; in addition, a confusion matrix is drawn to intuitively show the recognition effect of the model on different categories.
8. The multimodal image fusion and recognition method according to claim 1, characterized in that: In step S7, if it is found that the recognition effect of the model on a certain category is not good, try to increase the training samples of this category, or adjust the kernel function and penalty coefficient of the SVM; in addition, consider using a more complex model to improve the feature extraction and classification capabilities of the model; during the optimization process, continuously monitor the performance indicators on the validation set to ensure that the model does not overfit.
9. The multimodal image fusion and recognition method according to claim 1, characterized in that: In step S8, non-maximum suppression is used to remove redundant detection frames. Specifically, for each detection frame, its intersection over union (IoU) with other detection frames is calculated. If the IoU exceeds a set threshold, the detection frame with the highest confidence is retained and other detection frames are suppressed. In addition, a confidence threshold is set to filter out detection results below the threshold, thereby reducing misidentification.
10. The multimodal image fusion and recognition method according to claim 1, characterized in that: In step S9, a depth-separable convolutional layer is introduced into the feature extraction network, and its parameters include convolution kernel size, feature map size, number of input channels and number of output channels (N); this convolutional layer can effectively reduce the number of parameters and the amount of calculation, while enhancing the model's ability to capture detailed features.
Citation Information
Cited By
Three-mode unsupervised industrial anomaly detection method based on reconstruction network
CN120388024A
A trimodal unsupervised industrial anomaly detection method based on reconstruction network
CN120388024B
Visible light, infrared and IQ signal fusion individual identification method based on cross-modal cross attention
CN121095685A
PCB defect detection method and system based on multi-mode deep learning
CN121211210A
Image detection system and method based on multi-modal feature fusion
CN121616930A