Ground target identification method and system based on multi-modal image feature fusion
Through the multimodal image feature fusion method, using a lightweight network and a three-level attention mechanism, the recognition problem of traditional single-modal sensors in complex environments is solved, and high-precision and low-overhead ground target recognition is achieved.
Patent Information
- Application Number
- CN202510754683.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional single-modal sensors have limitations when identifying ground targets in complex environments. The difficulty in aligning heterogeneous data, imbalanced modal contributions, and insufficient computational efficiency and generalization lead to low and unstable recognition accuracy.
A multimodal image feature fusion method is adopted. By improving the network by lightweight MobileNetV3, EfficientNet-B4 and ResNet-50, combined with the channel-space-modality three-level attention mechanism and the sub-modal pre-training strategy, feature decoupling and efficient parallel computing are achieved, and the modal contribution weights are adaptively calibrated to enhance robustness and adaptability.
It significantly improves the accuracy and robustness of ground target recognition, reduces computational overhead, and enhances the decision-making ability and scalability of the model in complex environments.
Smart Images

Figure CN120747684A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of radar signal processing technology, and in particular to a ground target recognition method and system based on multimodal image feature fusion. Background Art
[0002] With the rapid development of remote sensing technology and the increasing demand for intelligent military capabilities, ground target recognition is becoming increasingly important in areas such as national defense security, disaster monitoring, and traffic management. Traditional recognition methods often rely on single-modal sensor data, such as visible light, infrared, or synthetic aperture radar (SAR) imagery. However, single data sources face significant limitations in complex environments: visible light images are susceptible to changes in lighting and weather conditions, infrared imaging has nighttime detection capabilities but low resolution, and SAR images can penetrate clouds and fog but are subject to speckle noise. This is particularly true in scenarios such as battlefield reconnaissance and urban security, where targets are often camouflaged, obscured, or dynamically interfered with. Multimodal information fusion is urgently needed to achieve information complementarity and overcome the performance bottlenecks of single sensors.
[0003] In recent years, multimodal object recognition has gradually become a research hotspot, but its technical implementation still faces multiple challenges:
[0004] 1. Challenges in aligning heterogeneous data: Visible light, infrared, and SAR imagery differ significantly in their imaging principles, resolution, and feature dimensions. For example, visible light reflects the color and texture of a target's surface, infrared captures the distribution of thermal radiation, and SAR characterizes the target's geometry through electromagnetic wave scattering. Traditional fusion methods, such as pixel-level stitching and decision-level voting, struggle to effectively correlate cross-modal features, leading to information redundancy or conflict.
[0005] 2. Imbalanced modal contributions: Existing methods often perform a simple weighted average or concatenation of multimodal features, ignoring the dynamic changes in modal reliability in different scenarios. For example, the weight of infrared data should be significantly increased in night missions, while SAR data should dominate decision-making when obscured by clouds and fog.
[0006] 3. Insufficient computational efficiency and generalization: End-to-end joint training of multimodal networks easily leads to gradient competition, resulting in unstable model convergence; the high cost of labeling large-scale multimodal datasets restricts the model's generalization ability in small sample scenarios. Summary of the Invention
[0007] The purpose of the present invention is to propose a ground target recognition method and system based on multimodal image feature fusion with high accuracy, strong robustness, low computational overhead, strong adaptability and strong scalability.
[0008] The technical solution to achieve the purpose of the present invention is: a ground target recognition method based on multimodal image feature fusion, comprising the following steps:
[0009] Step 1: Obtain infrared images, visible light images, and SAR images of ground targets;
[0010] Step 2: Perform image amplification on the images of the three modalities and add classification labels;
[0011] Step 3: For data of different modalities, different neural network models are used for network training to achieve preliminary classification of the target;
[0012] Step 4: Remove the classification heads of the three pre-trained neural network models, fix the feature extraction network parameters, and use the pre-trained feature extraction module to extract features of different modal data;
[0013] Step 5: The features extracted from different modalities are fused using the channel-space-modality three-level attention fusion module;
[0014] Step 6: Retrain the neural network model using the multimodal images and labels to obtain a trained neural network model;
[0015] Step 7: Input the remote sensing image to be identified into the trained neural network model and output the classification result.
[0016] A ground target recognition system based on multimodal image feature fusion is used to implement the ground target recognition method based on multimodal image feature fusion. The system includes the first to seventh modules, and the functions of each module are as follows:
[0017] The first module obtains infrared images, visible light images, and SAR images of ground targets;
[0018] The second module performs image amplification on the three modal images and adds classification labels;
[0019] The third module uses different neural network models to train the network for data of different modalities to achieve preliminary classification of the target;
[0020] In the fourth module, the classification heads of the three pre-trained neural network models are removed, the feature extraction network parameters are fixed, and the pre-trained feature extraction module is used to extract features of different modal data;
[0021] The fifth module fuses the features extracted from different modalities using a three-level attention fusion module: channel-space-modality.
[0022] The sixth module uses multimodal images and labels to retrain the neural network model to obtain a trained neural network model;
[0023] The seventh module inputs the remote sensing image to be identified into the trained neural network model and outputs the classification result.
[0024] Compared with the existing technology, the present invention has the following significant advantages: (1) Through heterogeneous network design, a lightweight MobileNetV3, a composite scaling EfficientNet-B4 and a non-local attention ResNet-50 improved network are constructed for visible light, infrared and SAR modalities respectively, which significantly reduces cross-modal interference and realizes feature decoupling and efficient parallel computing; (2) A channel-space-modality three-level attention mechanism is adopted to adaptively calibrate the multimodal contribution weights, solve the problem of feature redundancy and spatial offset in traditional methods, and enhance the decision robustness in complex environments; (3) A sub-modal pre-training strategy is adopted in combination with gradient isolation and orthogonal constraints to ensure that each network focuses on intra-modal feature mining and reduce model parameter redundancy; (4) Adaptive data enhancement technology and frequency domain detail enhancement module are adopted to effectively improve the network's tolerance to noise, occlusion and deformation, ensuring efficient fusion and stable representation of multi-source heterogeneous data; (5) This architecture has low computational overhead, high adaptability and strong scalability, providing a general technical foundation for multimodal perception tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 The figure is a flow chart of a ground target recognition method based on multimodal image feature fusion according to the present invention.
[0026] Figure 2 is a visible light image of a ground target obtained in an embodiment of the present invention.
[0027] Figure 3 This is an infrared image of a ground target obtained in an embodiment of the present invention.
[0028] Figure 4 is a SAR image of a ground target obtained in an embodiment of the present invention. DETAILED DESCRIPTION
[0029] It is easy to understand that, based on the technical solution of the present invention, those skilled in the art can imagine various embodiments of the present invention without changing the essential spirit of the present invention. Therefore, the following specific embodiments and drawings are merely illustrative of the technical solution of the present invention and should not be regarded as the entire invention or as limiting or defining the technical solution of the present invention.
[0030] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.
[0031] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0032] like Figure 1 As shown, the ground target recognition method based on multimodal image feature fusion of the present invention includes the following steps:
[0033] Step 1: Obtain infrared images, visible light images, and SAR images of ground targets;
[0034] Step 2: Perform image amplification on the images of the three modalities and add classification labels;
[0035] Step 3: For data of different modalities, different neural network models are used for network training to achieve preliminary classification of the target;
[0036] Step 4: Remove the classification heads of the three pre-trained neural network models, fix the feature extraction network parameters, and use the pre-trained feature extraction module to extract features of different modal data;
[0037] Step 5: The features extracted from different modalities are fused using the channel-space-modality three-level attention fusion module;
[0038] Step 6: Retrain the neural network model using the multimodal images and labels to obtain a trained neural network model;
[0039] Step 7: Input the remote sensing image to be identified into the trained neural network model and output the classification result.
[0040] As a specific example, the image amplification of the three modal images described in step 2 and the addition of classification labels are as follows:
[0041] Step 2.1: For visible light images, use RandAugment technology to perform preprocessing of random combination rotation, color dithering and occlusion simulation;
[0042] Step 2.2: For infrared images, perform preprocessing by histogram equalization and thermal radiation noise injection;
[0043] Step 2.3: Preprocess the SAR image by suppressing speckle noise and enhancing geometric deformation.
[0044] Step 2.4: Independently label the preprocessed data to maintain spatial consistency of multimodal labels.
[0045] As a specific example, in step 3, different neural network models are used to perform network training for data of different modalities to achieve preliminary classification of the target, as follows:
[0046] Step 3.1: For the infrared image network, a lightweight MobileNetV3 model combined with a channel attention module is used for network training to achieve preliminary classification of thermal radiation distribution and contour features.
[0047] Step 3.2: For the visible light image network, use EfficientNet-B4 as the backbone and use a compound scaling strategy to balance computational efficiency and feature expression capabilities to extract texture details of visible light images and fuse high-frequency texture details through the multi-scale feature pyramid (MSFP).
[0048] Step 3.3: For the SAR image network, based on the improved network of ResNet-50, a non-local attention layer is added to the network to strengthen the geometric structure representation. The geometric structure representation is strengthened by cross-attention and combined with multi-scale dilated convolution to optimize the target scattering center positioning error.
[0049] As a specific example, the classification heads of the three pre-trained neural network models described in step 4 are removed, and the feature extraction network parameters are fixed. Specifically, a framework of sub-modal pre-training and global fine-tuning is adopted:
[0050] Phase 1: Freeze the feature extraction network and only train the classification head;
[0051] The second stage: unfreeze some shallow network parameters, jointly optimize the fusion module and classifier, and gradually increase the proportion of difficult samples.
[0052] As a specific example, the features extracted from different modalities described in step 5 are fused using the channel-space-modality three-level attention fusion module, as follows:
[0053] Step 5.1, channel-level calibration: Dynamically assign weights to each modal feature channel through the SE module to suppress redundant information, as follows:
[0054] Step 5.1.1: The feature maps extracted from infrared, visible light, and SAR images are fed into the global average pooling layer to compress the spatial dimensions and generate channel weights. The access weights are then fed into a two-layer fully connected network to generate channel weights through dimensionality reduction, dimensionality increase, and sigmoid activation.
[0055] Step 5.1.2: Multiply the generated channel weights by the original feature map of the corresponding modality element by element to enhance key channel information, suppress redundant channel features, and complete the adaptive calibration of the channel dimension;
[0056] Step 5.2: Spatial alignment: Use deformable convolution to achieve spatial alignment of cross-modal feature maps, as follows:
[0057] Step 5.2.1: Input the channel-calibrated infrared, visible light, and SAR image features into the deformable convolution module. The convolution operation dynamically generates irregular sampling offsets to compensate for the spatial misalignment between different modalities.
[0058] Step 5.2.2: Based on the generated offset, perform irregular grid sampling on each modal feature map to enhance the spatial adaptability of the feature map and output a multimodal feature representation with spatial alignment;
[0059] Step 5.3, Modal Interaction Decision: Construct a cross-modal attention matrix and automatically adjust the modal contribution weights according to the environmental context, as follows:
[0060] Step 5.3.1: Concatenate the spatially aligned multimodal feature maps and input them into the global average pooling layer to extract contextual features of different modalities.
[0061] Step 5.3.2: Construct a cross-modal attention matrix based on the extracted contextual features, where the rows and columns of the matrix correspond to different modalities, and calculate the inter-modal dependency weights through a fully connected layer and softmax.
[0062] Step 5.3.3: Perform weighted fusion of multimodal features according to the attention matrix, dynamically adjust the contribution ratio of infrared, visible light, and SAR image features, and output the final fusion features.
[0063] As a specific example, the neural network model is retrained using multimodal images and labels in step 6 to obtain a trained neural network model, as follows:
[0064] Step 6.1: Preprocess the remote sensing image to be identified, and simultaneously obtain visible light, infrared, SAR and other modal image data of the same area. Divide the multimodal images and corresponding labels into training set, validation set and test set, and distribute them in a ratio of 8:1:1. Normalize the multimodal data and map the pixel values to the interval [0,1].
[0065] Step 6.2: Input the preprocessed multimodal images into the pretrained neural network model. The texture features of the visible light image are extracted using the EfficientNet-B4 backbone network, the thermal radiation features of the infrared image are extracted using the modified MobileNetV3 network, and the structural features of the SAR image are extracted using a multi-scale feature pyramid. Feature fusion is performed using the channel-space-modality three-level attention fusion module.
[0066] Step 6.3: The fused features are mapped to the ground target category space through the fully connected layer, and the Softmax activation function is applied to output the probability distribution of each category. The category with the highest probability is selected as the preliminary recognition result.
[0067] Step 6.4: Set the cross-entropy loss function as the model optimization objective to measure the difference between the predicted class probability distribution and the true label; select the Adam optimizer, set the initial learning rate to 0.001, and set the training hyperparameters, including the number of training epochs (epochs) to 50 and the batch size (batch size) to 32.
[0068] Step 6.5: After training is complete, use the test set to perform a final evaluation on the saved optimal model. Calculate metrics such as precision, recall, and mean average precision (mAP) to comprehensively measure the model's performance in the multimodal ground object recognition task.
[0069] The present invention also provides a ground target recognition system based on multimodal image feature fusion, which is used to implement the ground target recognition method based on multimodal image feature fusion. The system includes the first to seventh modules, and the functions of each module are as follows:
[0070] The first module obtains infrared images, visible light images, and SAR images of ground targets;
[0071] The second module performs image amplification on the three modal images and adds classification labels;
[0072] The third module uses different neural network models to train the network for data of different modalities to achieve preliminary classification of the target;
[0073] In the fourth module, the classification heads of the three pre-trained neural network models are removed, the feature extraction network parameters are fixed, and the pre-trained feature extraction module is used to extract features of different modal data;
[0074] The fifth module fuses the features extracted from different modalities using a three-level attention fusion module: channel-space-modality.
[0075] The sixth module uses multimodal images and labels to retrain the neural network model to obtain a trained neural network model;
[0076] The seventh module inputs the remote sensing image to be identified into the trained neural network model and outputs the classification result.
[0077] In one embodiment, the present invention also provides a mobile terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the ground target recognition method based on multimodal image feature fusion when executing the program.
[0078] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0079] Example
[0080] This example uses the publicly available WHU-OPT-SAR dataset, which is free and includes seven major ground objects: farmland, urban areas, rural areas, water bodies, forests, roads, and other objects. Images are available in three modalities: visible light, infrared, and SAR. The three ground objects are selected as the dataset, and the training and test sets are divided into an 8:2 ratio.
[0081] Figure 2 、 Figure 3 、 Figure 4 They are the visible light image, infrared image, and SAR image of the ground target obtained respectively. Table 1 is a comparison chart of the ground space target recognition results of single-modal, dual-modal, and multi-modal image feature fusion in an embodiment of the present invention.
[0082] Table 1
[0083]
[0084] As shown in Table 1, the highest recognition rate using single-modal information is only 87.43%. Using dual-modal fusion for target recognition significantly surpasses single-modal recognition, reaching 94.67%. Trimodal fusion achieves the highest recognition rate, achieving 96.74%. This demonstrates that multimodal image joint learning can effectively learn the correlations between different modalities, significantly improving target recognition rates.
[0085] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A ground target recognition method based on multimodal image feature fusion, characterized in that: The following steps are involved: Step 1: Obtain infrared images, visible light images, and SAR images of ground targets; Step 2: Perform image amplification on the images of the three modalities and add classification labels; Step 3: For data of different modalities, different neural network models are used for network training to achieve preliminary classification of the target; Step 4: Remove the classification heads of the three pre-trained neural network models, fix the feature extraction network parameters, and use the pre-trained feature extraction module to extract features of different modal data; Step 5: The features extracted from different modalities are fused using the channel-space-modality three-level attention fusion module; Step 6: Retrain the neural network model using the multimodal images and labels to obtain a trained neural network model; Step 7: Input the remote sensing image to be identified into the trained neural network model and output the classification result.
2. The ground target recognition method based on multimodal image feature fusion according to claim 1, characterized in that: The three modal images described in step 2 are augmented and classified as follows: Step 2.1: For visible light images, use RandAugment technology to perform preprocessing of random combination rotation, color dithering and occlusion simulation; Step 2.2: For infrared images, perform preprocessing by histogram equalization and thermal radiation noise injection; Step 2.3: Preprocess the SAR image by suppressing speckle noise and enhancing geometric deformation. Step 2.4: Independently label the preprocessed data to maintain spatial consistency of multimodal labels.
3. The ground target recognition method based on multimodal image feature fusion according to claim 1, characterized in that: In step 3, different neural network models are used to train the network for data of different modalities to achieve preliminary classification of the target, as follows: Step 3.1: For the infrared image network, a lightweight MobileNetV3 model combined with a channel attention module is used for network training to achieve preliminary classification of thermal radiation distribution and contour features. Step 3.2: For the visible light image network, use EfficientNet-B4 as the backbone and use a compound scaling strategy to balance computational efficiency and feature expression capabilities to extract texture details of visible light images and fuse high-frequency texture details through the multi-scale feature pyramid (MSFP). Step 3.3: For the SAR image network, based on the improved network of ResNet-50, a non-local attention layer is added to the network to strengthen the geometric structure representation. The geometric structure representation is strengthened by cross-attention and combined with multi-scale dilated convolution to optimize the target scattering center positioning error.
4. The ground target recognition method based on multimodal image feature fusion according to claim 1, characterized in that: Remove the classification heads from the three pre-trained neural network models described in step 4, fix the feature extraction network parameters, and specifically adopt the framework of sub-modal pre-training-global fine-tuning: Phase 1: Freeze the feature extraction network and only train the classification head; The second stage: unfreeze some shallow network parameters, jointly optimize the fusion module and classifier, and gradually increase the proportion of difficult samples.
5. The ground target recognition method based on multimodal image feature fusion according to claim 1, characterized in that: The features extracted from different modalities described in step 5 are fused using the channel-space-modality three-level attention fusion module, as follows: Step 5.1, channel-level calibration: Dynamically assign weights to each modal feature channel through the SE module to suppress redundant information; Step 5.2: Spatial alignment: Use deformable convolution to achieve spatial registration of cross-modal feature maps; Step 5.3, modal interaction decision: Construct a cross-modal attention matrix and automatically adjust the modal contribution weight according to the environmental context.
6. The ground target recognition method based on multimodal image feature fusion according to claim 5, characterized in that: The specific steps for channel-level calibration described in step 5.1 are as follows: Step 5.1.1: The feature maps extracted from infrared, visible light, and SAR images are fed into the global average pooling layer to compress the spatial dimensions and generate channel weights. The access weights are then fed into a two-layer fully connected network to generate channel weights through dimensionality reduction, dimensionality increase, and sigmoid activation. Step 5.1.2: Multiply the generated channel weights element-by-element with the original feature map of the corresponding modality to enhance key channel information, suppress redundant channel features, and complete the adaptive calibration of the channel dimension.
7. The ground target recognition method based on multimodal image feature fusion according to claim 5, characterized in that: The specific steps for spatial alignment described in step 5.2 are as follows: Step 5.2.1: Input the channel-calibrated infrared, visible light, and SAR image features into the deformable convolution module. The convolution operation dynamically generates irregular sampling offsets to compensate for the spatial misalignment between different modalities. Step 5.2.2: Based on the generated offset, perform irregular grid sampling on each modal feature map to enhance the spatial adaptability of the feature map and output a multimodal feature representation with spatial position alignment.
8. The ground target recognition method based on multimodal image feature fusion according to claim 5, characterized in that: The modal interaction decision described in step 5.3 is as follows: Step 5.3.1: Concatenate the spatially aligned multimodal feature maps and input them into the global average pooling layer to extract contextual features of different modalities. Step 5.3.2: Construct a cross-modal attention matrix based on the extracted contextual features, where the rows and columns of the matrix correspond to different modalities, and calculate the inter-modal dependency weights through a fully connected layer and softmax. Step 5.3.3: Perform weighted fusion of multimodal features according to the attention matrix, dynamically adjust the contribution ratio of infrared, visible light, and SAR image features, and output the final fusion features.
9. The ground target recognition method based on multimodal image feature fusion according to claim 1, characterized in that: In step 6, the neural network model is retrained using multimodal images and labels to obtain a trained neural network model, as follows: Step 6.1: Preprocess the remote sensing image to be identified. Simultaneously acquire visible light, infrared, and SAR image data for the same area. Divide the multimodal images and corresponding labels into training, validation, and test sets in an 8:1:1 ratio. Normalize the multimodal data and map pixel values to the [0, 1] interval. Step 6.2: Input the preprocessed multimodal images into the pretrained neural network model. Texture features are extracted from visible light images using the EfficientNet-B4 backbone network, thermal radiation features are extracted from infrared images using the improved MobileNetV3 network, and structural features are extracted from SAR images using a multi-scale feature pyramid. Feature fusion is then performed using a channel-space-modality three-level attention fusion module. Step 6.3: The fused features are mapped to the ground object category space through a fully connected layer. The Softmax activation function is applied to output the probability distribution of each category, and the category with the highest probability is selected as the preliminary recognition result. Step 6.4: Set the cross entropy loss function as the model optimization objective to measure the difference between the predicted class probability distribution and the true label; select the Adam optimizer, set the initial learning rate to 0.001, and set the training hyperparameters, including the number of training epochs to 50 and the batch size to 32; Step 6.5: After training is complete, use the test set to perform a final evaluation on the saved optimal model and calculate the precision, recall, and mean average precision (mAP) to measure the performance of the model in the multimodal ground object recognition task.
10. A ground target recognition system based on multimodal image feature fusion, characterized in that: The system is used to implement the ground target recognition method based on multimodal image feature fusion as described in any one of claims 1 to 9. The system includes the first to seventh modules, and the functions of each module are as follows: The first module obtains infrared images, visible light images, and SAR images of ground targets; The second module performs image amplification on the three modal images and adds classification labels; The third module uses different neural network models to train the network for data of different modalities to achieve preliminary classification of the target; In the fourth module, the classification heads of the three pre-trained neural network models are removed, the feature extraction network parameters are fixed, and the pre-trained feature extraction module is used to extract features of different modal data; The fifth module fuses the features extracted from different modalities using a three-level attention fusion module: channel-space-modality. The sixth module uses multimodal images and labels to retrain the neural network model to obtain a trained neural network model; The seventh module inputs the remote sensing image to be identified into the trained neural network model and outputs the classification result.
Citation Information
Cited By
Detection mechanism guided multi-mode element learning remote sensing reconnaissance target identification method
CN121074378A
A multi-modal meta-learning remote sensing reconnaissance target recognition method guided by a detection mechanism
CN121074378B
SAR image cross-view generation method, system and device based on 3D Gaussian ellipsoid scatterer and medium
CN122289540A