High-resolution remote sensing image accurate classification system and method based on deep learning multi-modal fusion
By adopting multimodal data optimization preprocessing, deep fusion model design, automated annotation and semi-supervised learning, lightweight model and hardware acceleration technology, and strategies to enhance interpretability in a multimodal fusion system for high-resolution remote sensing images, the problems of data heterogeneity, computing efficiency, overfitting, labeling difficulties, real-time and interpretability are solved, and efficient, real-time and interpretable remote sensing image classification is achieved.
Patent Information
- Application Number
- CN202510012475.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-27
AI Technical Summary
Existing high-resolution remote sensing image multimodal fusion system based on deep learning has challenges in terms of data heterogeneity, high computing resource consumption, lack of labeled data, overfitting, real-time and interpretability.
Through the optimization preprocessing of multimodal data, deep fusion model design, automated annotation and semi-supervised learning, lightweight model and hardware acceleration technology, and strategies to enhance interpretability, data heterogeneity, computing efficiency, overfitting, labeling difficulties, real-time and interpretability are solved.
While maintaining high classification accuracy, it improves computing efficiency, reduces manual intervention, and enhances the transparency and generalization capabilities of the model, meeting the needs of practical applications.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of precise classification systems for remote sensing images, and specifically to a precise classification system and method for high-resolution remote sensing images based on deep learning multi-modal fusion. Background Art
[0002] A precise classification system for high-resolution remote sensing images based on deep learning multi-modal fusion is a technology that uses deep learning models to combine multiple data sources (such as optical images, radar data, infrared images, etc.). By effectively fusing information from different modalities, the classification accuracy of remote sensing images can be improved. This method constructs a multi-modal neural network, fuses image data obtained from different sensors, uses a convolutional neural network (CNN) to extract spatial features, and combines features from other modalities for joint learning. The system can identify more detailed land cover classes in high-resolution remote sensing images, overcome the limitations of single-modal data, improve the accuracy and robustness of classification, and is widely used in fields such as urban planning, agricultural monitoring, and environmental assessment.
[0003] Although the precise classification system for high-resolution remote sensing images based on deep learning multi-modal fusion has great potential in practical applications, there are still some technical drawbacks and challenges, mainly including the following aspects:
[0004] 1. Heterogeneity and inconsistency of data
[0005] Feature differences in different-modal data: Different types of remote sensing data (such as optical images, radar data, infrared images, etc.) have different spatial resolutions, spectral characteristics, and noise characteristics. These differences may lead to inconsistent information during the fusion process, thereby affecting the final classification effect.
[0006] Data alignment problem: Images of different modalities often have deviations in time, space, and angle. How to accurately align these data is a technical challenge. Especially in high-resolution remote sensing images, alignment errors may lead to poor fusion effects, thereby affecting the classification results.
[0007] 2. High computational resource consumption
[0008] Model complexity: Deep learning models, especially deep neural networks for multi-modal fusion, usually involve a large number of parameters and calculations. The data volume of high-resolution remote sensing images is huge, and the demand for computational resources during training and inference is extremely high. High-performance hardware devices (such as GPU clusters) may be required, which increases the cost and deployment difficulty of the system.
[0009] Long training time: Due to the large number of model parameters and training data volume, the training time of deep learning models is often long. Especially in the case of multi-modal data fusion, the model requires more complex calculations and optimization strategies.
[0010] 3. Lack of sufficient labeled data
[0011] Difficulty in data annotation: The annotation of high - resolution remote sensing images requires a large amount of manual intervention and is costly, especially for the annotation of multi - modal data, which requires comprehensive annotation of different types of data and is particularly difficult for large - scale data sets. The quality of data annotation directly affects the training effect of the model.
[0012] Imbalance of labeled data: In practical applications, different types of ground objects may be unevenly distributed in remote sensing images, resulting in a scarcity of samples for some classes and causing imbalance in training data. This will cause the deep - learning model to be biased towards the classes with a larger amount of data during classification, affecting the classification accuracy.
[0013] 4. Overfitting problem
[0014] Model complexity and overfitting: Deep - learning models, especially multi - modal fusion models, have many parameters and are complex, making them prone to overfitting on the training set, especially when the sample size is small or the data noise is large. Overfitting will lead to poor generalization ability of the model on new data and affect the performance in practical applications.
[0015] Cross - domain adaptation problem: Due to the regional differences in remote sensing data, the same model may show large performance fluctuations on data from different geographical regions or different time periods, resulting in insufficient generalization ability.
[0016] 5. Computational efficiency and real - time problems
[0017] Challenges in real - time applications: The processing of high - resolution remote sensing images usually requires a large amount of computing resources, making real - time processing difficult. Especially in application scenarios that require quick classification results (such as post - disaster assessment, real - time monitoring, etc.), deep - learning models may have difficulty meeting real - time requirements.
[0018] Slow inference speed: When deep - learning models perform inference, the data processing volume for high - resolution images is large. Especially in the case of multi - modal fusion, the data dimension after image fusion is higher and the computational burden is heavier.
[0019] 6. Poor model interpretability
[0020] Black - box problem: Deep - learning models are usually regarded as "black boxes", and their decision - making processes are not easy to explain. In remote sensing image classification tasks, especially in application scenarios where clear evidence is required for important decisions (such as environmental monitoring, disaster prediction, etc.), the interpretability of the model is crucial. The lack of transparency of deep - learning models may limit their practical applications in some fields.
[0021] 7. Impact of noise and incomplete data
[0022] Noise interference: Various noises (such as atmospheric noise, sensor noise, etc.) often exist in remote sensing images. These noises can affect the image quality and further influence the classification accuracy of deep learning models. Multimodal fusion methods may not be able to fully handle these noises, and especially when dealing with low-quality or missing data, the performance may decrease significantly.
[0023] Data missing problem: In some cases, different modality data may be missing or unavailable. How to perform effective fusion and classification in the case of missing data is also an urgent problem to be solved.
[0024] 8. Processing bottlenecks of high-resolution images
[0025] Data processing bottleneck: High-resolution remote sensing images have large file sizes and huge amounts of data, requiring a large amount of storage space and memory during processing. Even with strong hardware support, the data processing speed may still be limited, especially when dealing with large-scale remote sensing image data.
[0026] Therefore, we propose a precise classification system and method for high-resolution remote sensing images based on deep learning multimodal fusion. Summary of the invention
[0027] To achieve the above object, the present invention provides the following technical solutions: A precise classification system and method for high-resolution remote sensing images based on deep learning multimodal fusion, including the following steps:
[0028] S1: Multimodal data preprocessing and fusion optimization
[0029] S1.1: Standardization and normalization: Perform standardization and normalization processing on remote sensing data of different modalities to eliminate the influence caused by differences such as resolution and spectral range between different modalities. For example, for optical images, the contrast can be enhanced through histogram equalization; for radar images, denoising processing can be performed;
[0030] S1.2: Image registration and alignment: Use image registration algorithms (such as feature point matching, mutual information method, etc.) to ensure the spatial alignment of different modality data. For modality images with mismatched resolutions, use super-resolution reconstruction technology to improve the spatial resolution of low-resolution data and ensure the consistency of multimodal images in space and time;
[0031] S1.3: Multiscale convolutional feature fusion: Use a multiscale convolutional neural network (CNN) to extract image features at different scales, which can better capture the detailed information in the image. Especially in high-resolution remote sensing images, different modalities may contain different spatial information;
[0032] S1.4: Deep Fusion and Channel Attention Mechanism: Introduce the channel attention mechanism (such as SE-Net, etc.) into the deep learning network. By adaptively adjusting the importance of different modality features, reduce the influence of noise interference and enhance the representation ability of key features. Merge the information of different modalities into a feature space through deep fusion (e.g., concatenation or weighted fusion) to improve the classification performance.
[0033] S2: Improved Deep Learning Model Design
[0034] S2.1: Dual-Stream Network: Design two or more network branches for remote sensing data of different modalities (e.g., one branch for processing optical images and another for processing radar images). Each network independently learns the features of its respective modality, and finally fuse the features through a fusion module (such as a weighted fusion layer). This design can maintain the independence of each modality information in the network and avoid mixed noise;
[0035] S2.2: Joint Learning of Image and Semantic Features: To enhance classification accuracy, adopt joint learning of image features and semantic information (such as labels of ground object types, regional information, etc.). By introducing a pre-trained knowledge graph or semantic information system, the model can better understand the semantic relationships of different ground objects and reduce misclassification;
[0036] S2.3: Dropout and Data Augmentation: Use Dropout regularization during training to reduce the risk of network overfitting. In addition, data augmentation methods (such as rotation, scaling, affine transformation, etc.) can be used to increase the diversity of training data, thereby improving the generalization ability of the model;
[0037] S2.4: Hybrid Data Augmentation: Combine different modality data for hybrid augmentation (e.g., data augmentation of optical images and radar images), making the training data more diverse and reducing the impact of data imbalance.
[0038] S3: Automatic Annotation and Semi-Supervised Learning
[0039] S3.1: Weak Supervised Learning: Adopt weak supervised learning methods to reduce the dependence on manually annotated data. Coarse labels can be generated through clustering algorithms or image segmentation methods, and then fine-tuned with a small amount of high-quality annotated data to reduce the annotation cost;
[0040] S3.2: Transfer Learning: For remote sensing images in different regions or at different times, adopt transfer learning methods. Use the labeled data of one region to train a basic model, and then adjust the model through transfer learning to adapt to the new geographical region. This can reduce the need for training data in the new region and avoid the bottleneck of data annotation;
[0041] S3.3: Self-Supervised Feature Learning: Through self-supervised learning methods, the model can automatically extract potential features from unlabeled data. By leveraging pixel-level similarity or contrast learning between images, the training of feature representations is carried out to enhance the robustness of the model, especially significantly improving the classification performance in the absence of labeled data.
[0042] S4: Improving Real-Time Performance and Computational Efficiency
[0043] S4.1: Model Compression and Quantization: On the premise of ensuring classification accuracy, through methods such as pruning, knowledge distillation, and model quantization, the computational amount and storage requirements of the model are reduced, the inference speed and real-time performance are improved, and it adapts to the real-time inference requirements on embedded or edge devices;
[0044] S4.2: Edge Computing and Distributed Processing: The model inference task is assigned to edge computing nodes or distributed computing environments to reduce the burden and latency of the central server and ensure real-time interaction of data between the ground station and edge devices;
[0045] S4.3: TensorRT and Hardware Acceleration: Utilize optimization inference engines such as TensorRT to optimize for hardware (such as GPUs, TPUs, FPGAs, etc.), significantly improving the inference speed of image classification. At the same time, select deep learning frameworks that support hardware acceleration, such as ONNX, TensorFlow Lite, etc., to reduce inference latency and ensure real-time classification.
[0046] S5: Interpretability and Model Transparency
[0047] S5.1: Interpretable Deep Learning Models: Use interpretability enhancement techniques (such as Grad-CAM, LIME, etc.) so that the model can show the image regions or modal features it focuses on during classification. This not only helps users understand the decision-making process of the model but also contributes to improving the transparency and reliability of the model;
[0048] S5.2: Integrated Model Explanation: Combine the prediction results of different modalities, use ensemble learning methods (such as model fusion) to improve classification accuracy, and at the same time provide the contribution degrees and explanations of each modality to help analyze which modal features play a decisive role in the final classification.
[0049] S6: Anomaly Detection and Noise Processing
[0050] S6.1: Multi-Modal Noise Modeling and Denoising: Design noise modeling and denoising methods for the noise characteristics of different modal images. For example, use convolutional neural networks for end-to-end image denoising to remove atmospheric interference noise in optical images or echo noise in radar images;
[0051] S6.2: Noise Adaptive Learning: Through an adaptive filtering method, dynamically adjust the noise processing strategy to ensure the robustness and accuracy of the model in complex environments.
[0052] S7: System Testing and Optimization
[0053] S7.1: Multi-region Testing: Test the performance of the classification model in different geographical regions to ensure that the model has good generalization ability in multiple scenarios. If the performance is poor in a specific region, further optimize the model through transfer learning;
[0054] S7.2: Time Series Data Validation: Use time series data to test the model to ensure that it can handle seasonal changes or data at different time periods. Through learning time series data, improve the stability and accuracy of the model over a long time span.
[0055] S8: Release and Continuous Update
[0056] S8.1: Continuous Monitoring and Feedback: Once the system is put into use, through continuous monitoring and user feedback, update the model in a timely manner to address possible performance degradation issues;
[0057] S8.2: Model Update and Adaptive Training: According to new remote sensing data and user requirements, continuously update and optimize the model, and use the Incremental Learning method for adaptive training.
[0058] Preferably, in step S1, the pixel values of all optical images are normalized to the range [0,1], and standardized using the mean and standard deviation of each image to make the images have the same dynamic range. For images with different spectral bands (such as multi-spectral images), each band is normalized separately. Radar images usually require logarithmic transformation (Log Transformation) to compress the dynamic range, and at the same time, filters (such as mean filter, Gaussian filter, etc.) are used for noise suppression. Image contrast can be enhanced through histogram equalization, especially in low-contrast areas, to enhance the classification effect. Use feature point-based methods (such as SIFT, SURF, ORB, etc.) or frequency-domain-based mutual information methods for registration of different modality images. By calculating the similarity, obtain the optimal transformation matrix. For images with different spatial resolutions, use interpolation methods (such as bilinear interpolation, cubic interpolation) for resampling to ensure that multi-modal images are aligned within the same spatial range. Use super-resolution reconstruction (such as SRGAN based on deep learning) to improve the spatial resolution of low-resolution modality images.
[0059] Preferably, in step S1, a multi-scale convolutional neural network is designed, and different convolutional kernel sizes are used to extract different-scale features of the image. Combine convolutional kernels of different sizes such as 1x1 convolution, 3x3 convolution, and 5x5 convolution to extract global and local features, and use "pyramid pooling" or "dilated convolution" to further capture large-scale features and enhance the recognition ability of complex structures in remote sensing images. Channel attention mechanism (SE-Net): By performing channel attention mechanism processing on the feature maps of each modality, automatically adjust the weights of each modality and reduce the influence of useless features. The weight of each channel is calculated by computing its global average pooling feature and calculating the attention weight through a fully connected layer. Depth fusion layer: Design a weighted fusion layer, and dynamically adjust the fusion ratio according to the importance of each modality. An adaptive fusion method can be adopted, such as using self-attention mechanism (Self-attention) to calculate the importance of different modality features and perform weighted fusion.
[0060] Preferably, in step S2, two independent sub-networks are designed, and each sub-network processes one modality data. For example, one sub-network is used to process optical images, and the other is used to process radar images. The convolutional layers in each sub-network are responsible for extracting the features of their respective modalities, and finally, these features are merged through a fusion module. In the fusion module, features of multiple modalities are fused in a concatenation or weighted summation manner to further improve the classification accuracy. In the deep neural network, a special semantic information module is added to obtain prior knowledge of ground object categories. For example, combine the geographical information of remote sensing images (such as land cover type, climate information, etc.) as additional input information, and input it into the network together with the image features for joint training. Use the multi-task learning (Multi-Task Learning) framework to perform tasks such as ground object recognition and object boundary detection while classifying images, so as to help the model better understand the semantic relationship of ground objects.
[0061] Preferably, in step S2, during the training process, randomly discard a part of neurons (Dropout) to prevent the model from over-relying on certain specific features. Adopt image data augmentation methods such as rotation, translation, scaling, random cropping, etc. to increase the diversity of training samples and reduce the risk of overfitting. For multi-modal data, enhance it by randomly combining different modal data. For example, randomly select a part of optical images and radar images for stitching to generate new training samples and increase the robustness of the model.
[0062] Preferably, in step S3, some unlabeled data and some labeled data are used to train by means of semi-supervised learning. For example, an autoencoder is used to learn the features of unlabeled images, fine-tuning is performed using a small amount of labeled data, a clustering algorithm (such as K-means) is used to cluster the unlabeled data, and then the clustering results are corrected by manual annotation or semi-automated methods. A pre-trained network (for example, a model pre-trained based on ImageNet or other remote sensing datasets) is used, and fine-tuning is performed to adapt to the new dataset. For images in different geographical regions, the trained model can be migrated to the new region and trained with a small amount of new labeled data to save time and resources.
[0063] Preferably, in step S3, a self-supervised learning method, such as jigsaw puzzle or contrastive learning of images, is adopted to enable the network to automatically learn meaningful feature representations from unlabeled data. A contrastive loss function (such as SimCLR) is used to train the model, so that the feature vector distances between similar images or image regions are relatively close, thereby enhancing the model's understanding of local details of the images.
[0064] Preferably, in step S4, pruning technology is used to remove redundant network weights, reducing computational complexity and storage requirements. During the pruning process, importance assessment can be performed based on the L1 norm or gradient of each convolutional kernel. Quantization is used to convert floating-point weights into fixed-point weights, reducing memory occupancy and accelerating the inference process. The trained deep learning model is deployed to edge devices (such as drones, satellites, etc.) for real-time inference, reducing data transmission and latency. A distributed training framework (such as TensorFlow Distributed, Horovod, etc.) is used for large-scale training, and multiple machines are used for parallel training to improve computational efficiency. TensorRT is used for model optimization, and the trained model is converted into an optimized format for efficient inference on the GPU. Inference is accelerated by optimizing memory usage, reducing the amount of computation, and increasing parallelism, supporting acceleration on specific hardware platforms (such as NVIDIA Jetson, Google Coral, etc.), and using the Tensor Core and other acceleration functions of the hardware for inference.
[0065] Preferably, in step S5, techniques such as Grad-CAM are used to visualize the activation map of a certain layer in the convolutional neural network, helping to understand the image regions that the model focuses on when classifying remote sensing images. Combining methods such as LIME (Local Interpretable Model-agnostic Explanations) or SHAP (SHapley Additive exPlanations), explanations for the model prediction results are obtained to help users understand how the model makes classification decisions. Using ensemble learning methods (such as random forest, XGBoost, etc.), the prediction results of each modality are fused, and the contribution of each modality to the final decision is explained through the feature importance analysis of the ensemble model.
[0066] Preferably, in step S6, for radar images, a denoising network based on a convolutional neural network (CNN) or U-Net structure can be used to denoise the images.
[0067] Compared with the prior art, the present invention provides a high-resolution remote sensing image precise classification system and method based on deep learning multi-modal fusion, having the following beneficial effects:
[0068] The high-resolution remote sensing image precise classification system and method based on deep learning multi-modal fusion, by adopting optimized preprocessing of multi-modal data, deep fusion model design, automated annotation and semi-supervised learning, lightweight model and hardware acceleration technology, and strategies to enhance interpretability, aims to solve problems such as data heterogeneity, computational efficiency, overfitting, annotation difficulty, real-time performance, and interpretability existing in existing methods. Through these optimizations, the system can improve computational efficiency, reduce manual intervention, and enhance the transparency and generalization ability of the model while maintaining high classification accuracy, meeting the actual application requirements. Specific Embodiments
[0069] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] Embodiment
[0071] Embodiment of the high-resolution remote sensing image precise classification system and method based on deep learning multi-modal fusion
[0072] The high-resolution remote sensing image precise classification system and method based on deep learning multi-modal fusion includes the following steps:
[0073] S1: Multi-modal data preprocessing and fusion optimization
[0074] S1.1: Standardization and normalization: Standardize and normalize remote sensing data of different modalities to eliminate the impacts brought by differences such as resolution and spectral range between different modalities. For example, for optical images, histogram equalization can be used to enhance the contrast; for radar images, denoising processing can be carried out.
[0075] S1.2: Image registration and alignment: Use image registration algorithms (such as feature point matching, mutual information method, etc.) to ensure the spatial alignment of data of different modalities. For modality images with mismatched resolutions, super-resolution reconstruction technology is used to improve the spatial resolution of low-resolution data and ensure the consistency of multi-modal images in space and time.
[0076] S1.3: Multi-scale convolutional feature fusion: Use a multi-scale convolutional neural network (CNN) to extract image features at different scales, which can better capture the detailed information in the images. Especially in high-resolution remote sensing images, different modalities may contain different spatial information.
[0077] S1.4: Deep fusion and channel attention mechanism: Introduce a channel attention mechanism (such as SE-Net, etc.) into the deep learning network. By adaptively adjusting the importance of features of different modalities, the impact of noise interference is reduced, and the representation ability of key features is enhanced. Through deep fusion (for example, concatenation or weighted fusion), the information of different modalities is merged into a feature space to improve the classification performance.
[0078] S2: Improved deep learning model design
[0079] S2.1: Dual-Stream Network: Design two or more network branches for remote sensing data of different modalities (for example, one branch processes optical images and the other processes radar images). Each network independently learns the features of its respective modality, and finally the features are fused through a fusion module (such as a weighted fusion layer). This design can maintain the independence of information of each modality in the network and avoid mixed noise.
[0080] S2.2: Joint learning of image and semantic features: To enhance the classification accuracy, joint learning of image features and semantic information (such as labels of ground object types, regional information, etc.) is adopted. By introducing a pre-trained knowledge graph or semantic information system, the model can better understand the semantic relationships of different ground objects and reduce misclassification.
[0081] S2.3: Dropout and data augmentation: By using Dropout regularization during the training process, the risk of network overfitting is reduced. In addition, data augmentation methods (such as rotation, scaling, affine transformation, etc.) can be used to increase the diversity of training data, thereby improving the generalization ability of the model.
[0082] S2.4: Hybrid Data Augmentation: Combine different modality data for hybrid augmentation (e.g., data augmentation for optical images and radar images), making the training data more diverse and reducing the impact of data imbalance.
[0083] S3: Automated Annotation and Semi-Supervised Learning
[0084] S3.1: Weak Supervised Learning: Adopt weak supervised learning methods to reduce the dependence on manually annotated data. Coarse labels can be generated through clustering algorithms or image segmentation methods and then finely adjusted with a small amount of high-quality annotated data to reduce the annotation cost.
[0085] S3.2: Transfer Learning: For remote sensing images in different regions or time periods, adopt transfer learning methods. Use the data of the annotated region to train the base model, and then adjust the model through transfer learning to adapt to the new geographical region. This can reduce the need for training data in the new region and avoid the bottleneck of data annotation.
[0086] S3.3: Self-Supervised Feature Learning: Through self-supervised learning methods, the model can automatically extract potential features from unannotated data. Use pixel-level similarity or contrast learning between images to train the feature representation and improve the robustness of the model. Especially when there is a lack of annotated data, it can significantly improve the classification performance.
[0087] S4: Improving Real-Time Performance and Computational Efficiency
[0088] S4.1: Model Compression and Quantization: On the premise of ensuring classification accuracy, reduce the computational amount and storage requirements of the model through methods such as pruning, knowledge distillation, and model quantization, improve the inference speed and real-time performance, and adapt to the real-time inference requirements on embedded or edge devices.
[0089] S4.2: Edge Computing and Distributed Processing: Allocate the model inference task to edge computing nodes or distributed computing environments to reduce the burden and latency of the central server and ensure the real-time interaction of data between the ground station and edge devices.
[0090] S4.3: TensorRT and Hardware Acceleration: Utilize optimization inference engines such as TensorRT to optimize for hardware (such as GPU, TPU, FPGA, etc.), significantly improving the inference speed of image classification. At the same time, select deep learning frameworks that support hardware acceleration, such as ONNX, TensorFlow Lite, etc., to reduce the inference latency and ensure real-time classification.
[0091] S5: Interpretability and Model Transparency
[0092] S5.1: Explainable Deep Learning Model: Use interpretability enhancement techniques (such as Grad-CAM, LIME, etc.) so that the model can show the image regions or modal features it focuses on during classification. This not only helps users understand the model's decision-making process but also contributes to improving the transparency and reliability of the model;
[0093] S5.2: Ensemble Model Explanation: Combine the prediction results of different modalities, and use ensemble learning methods (such as model fusion) to improve the classification accuracy, and at the same time provide the contribution degrees and explanations of each modality to help analyze which modal features play a decisive role in the final classification.
[0094] S6: Anomaly Detection and Noise Processing
[0095] S6.1: Multimodal Noise Modeling and Denoising: Design noise modeling and denoising methods according to the noise characteristics of different modality images. For example, use convolutional neural networks for end-to-end image denoising to remove atmospheric interference noise in optical images or echo noise in radar images;
[0096] S6.2: Noise Adaptive Learning: Dynamically adjust the noise processing strategy through adaptive filtering methods to ensure the robustness and accuracy of the model in complex environments.
[0097] S7: System Testing and Optimization
[0098] S7.1: Multi-region Testing: Test the performance of the classification model in different geographical regions to ensure that the model has good generalization ability in multiple scenarios. If the performance is poor in a specific region, further optimize the model through transfer learning;
[0099] S7.2: Time Series Data Validation: Use time series data to test the model to ensure that it can handle seasonal changes or data at different time periods. Through the learning of time series data, improve the stability and accuracy of the model over a long time span.
[0100] S8: Release and Continuous Update
[0101] S8.1: Continuous Monitoring and Feedback: Once the system is put into use, through continuous monitoring and user feedback, update the model in a timely manner to solve possible performance degradation problems;
[0102] S8.2: Model Update and Adaptive Training: Continuously update and optimize the model according to new remote sensing data and user needs, and use incremental learning methods for adaptive training.
[0103] Specifically, in step S1, the pixel values of all optical images are normalized to the range of [0, 1], and standardized using the mean and standard deviation of each image to make the images have the same dynamic range. For images with different spectral bands (such as multispectral images), each band is normalized separately. Radar images usually require logarithmic transformation (Log Transformation) to compress the dynamic range, and at the same time, filters (such as mean filtering, Gaussian filtering, etc.) are used for noise suppression. The image contrast can be enhanced through histogram equalization, especially in low-contrast areas, to enhance the classification effect. The registration of different modality images is performed using feature point-based methods (such as SIFT, SURF, ORB, etc.) or mutual information method based on the frequency domain. By calculating the similarity, the optimal transformation matrix is obtained. For images with different spatial resolutions, interpolation methods (such as bilinear interpolation, cubic interpolation) are used for resampling to ensure that multimodal images are aligned within the same spatial range. Super-resolution reconstruction (such as SRGAN based on deep learning) is used to improve the spatial resolution of low-resolution modality images.
[0104] Specifically, in step S1, a multi-scale convolutional neural network is designed, and different convolutional kernel sizes are used to extract different scale features of the image. Combine convolutional kernels of different sizes such as 1x1 convolution, 3x3 convolution, and 5x5 convolution to extract global and local features, and use "pyramid pooling" or "dilated convolution" to further capture large-scale features and enhance the recognition ability of complex structures in remote sensing images. Channel attention mechanism (SE-Net): By processing the feature maps of each modality through the channel attention mechanism, the weights of each modality are automatically adjusted to reduce the influence of useless features. The weight of each channel is calculated by computing its global average pooling feature and calculating the attention weight through a fully connected layer. Depth fusion layer: Design a weighted fusion layer, and dynamically adjust the fusion ratio according to the importance of each modality. An adaptive fusion method can be adopted, such as using self-attention mechanism (Self-attention) to calculate the importance of different modality features and perform weighted fusion.
[0105] Specifically, in step S2, two independent sub-networks are designed, and each sub-network processes one type of modal data. For example, one sub-network is used to process optical images, and the other is used to process radar images. The convolutional layers in each sub-network are responsible for extracting the features of their respective modalities. Finally, these features are merged through a fusion module. In the fusion module, the features of multiple modalities are fused in a concatenation or weighted summation manner to further improve the classification accuracy. In the deep neural network, a dedicated semantic information module is added to obtain prior knowledge of the ground object categories. For example, the geographical information of the remote sensing image (such as land cover type, climate information, etc.) is combined as additional input information and input into the network together with the image features for joint training. The multi-task learning framework is used to perform tasks such as ground object recognition and object boundary detection while classifying the images, thereby helping the model better understand the semantic relationships of the ground objects.
[0106] Specifically, during the training process in step S2, a part of the neurons are randomly discarded (Dropout) to prevent the model from over-relying on certain specific features. The image data augmentation method is adopted, such as rotation, translation, scaling, random cropping, etc., to increase the diversity of the training samples and reduce the risk of overfitting. For multi-modal data, different modalities of data are randomly combined for augmentation. For example, a part of the optical image and the radar image are randomly selected and spliced to generate new training samples, increasing the robustness of the model.
[0107] Specifically, in step S3, unlabeled data and some labeled data are used to train using the semi-supervised learning method. For example, an autoencoder is used to learn the features of the unlabeled images, and the small sample of labeled data is used for fine-tuning. The clustering algorithm (such as K-means) is used to cluster the unlabeled data, and then the clustering results are corrected through manual annotation or semi-automated methods. A pre-trained network (such as a model pre-trained based on ImageNet or other remote sensing datasets) is used, and fine-tuning is performed to adapt to the new dataset. For images in different geographical regions, the trained model can be migrated to the new region and trained with a small amount of new labeled data to save time and resources.
[0108] Specifically, in step S3, self-supervised learning methods such as Image Jigsaw Puzzle or Contrastive Learning are adopted to enable the network to automatically learn meaningful feature representations from unlabeled data. The model is trained using a contrastive loss function (such as SimCLR) so that the feature vector distances between similar images or image regions are relatively close, thereby enhancing the model's understanding of local image details.
[0109] Specifically, in step S4, Pruning technology is used to remove redundant network weights, reducing computational complexity and storage requirements. During the pruning process, importance assessment can be based on the L1 norm or gradient of each convolutional kernel. Through Quantization, floating-point weights are converted into fixed-point weights, reducing memory occupancy and accelerating the inference process. The trained deep learning model is deployed to edge devices (such as drones, satellites, etc.) for real-time inference, reducing data transmission and latency. A distributed training framework (such as TensorFlow Distributed, Horovod, etc.) is used for large-scale training, leveraging multiple machines for parallel training to improve computational efficiency. TensorRT is used for model optimization, converting the trained model into an optimized format for efficient inference on the GPU. Inference is accelerated by optimizing memory usage, reducing the amount of computation, and increasing parallelism, supporting acceleration on specific hardware platforms (such as NVIDIA Jetson, Google Coral, etc.), and leveraging the hardware's Tensor Core and other acceleration capabilities for inference.
[0110] Specifically, in step S5, techniques such as Grad-CAM are used to visualize the activation map of a certain layer in the convolutional neural network, helping to understand the image regions that the model focuses on when classifying remote sensing images. Combined with methods such as LIME (Local Interpretable Model) or SHAP (SHapley Additive exPlanations), explanations for the model's prediction results are obtained to help users understand how the model makes classification decisions. An ensemble learning method (such as Random Forest, XGBoost, etc.) is used to fuse the prediction results of each modality, and the contribution of each modality to the final decision is explained through feature importance analysis of the ensemble model.
[0111] Specifically, for radar images in step S6, a denoising network based on a convolutional neural network (CNN) or U-Net structure can be used to denoise the images.
[0112] Through the above technical solutions, in the present invention, by adopting optimized preprocessing of multi-modal data, deep fusion model design, automated annotation and semi-supervised learning, lightweight model and hardware acceleration technology, and strategies to enhance interpretability, it aims to solve problems such as data heterogeneity, computational efficiency, overfitting, annotation difficulties, real-time performance, and interpretability existing in existing methods. Through these optimizations, the system can improve computational efficiency, reduce manual intervention, and enhance the transparency and generalization ability of the model while maintaining high classification accuracy, meeting the requirements of practical applications.
[0113] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion, characterized by: The following steps are involved: S1: Multimodal data preprocessing and fusion optimization S1.1: Standardization and normalization: Standardize and normalize remote sensing data of different modes to eliminate the impact of differences in resolution, spectral range, etc. between different modes. For example, for optical images, contrast can be enhanced by histogram equalization; for radar images, denoising can be performed; S1.2: Image registration and alignment: Use image registration algorithms (such as feature point matching, mutual information method, etc.) to ensure spatial alignment of data from different modalities. For modal images with mismatched resolutions, use super-resolution reconstruction technology to improve the spatial resolution of low-resolution data to ensure the consistency of multi-modal images in space and time; S1.3: Multi-scale convolutional feature fusion: A multi-scale convolutional neural network (CNN) is used to extract image features at different scales, which can better capture the detailed information in the image, especially in high-resolution remote sensing images, where different modalities may contain different spatial information; S1.4: Deep fusion and channel attention mechanism: Introduce channel attention mechanism (such as SE-Net, etc.) in deep learning network, reduce the influence of noise interference and enhance the representation ability of key features by adaptively adjusting the importance of different modal features. Through deep fusion (for example, concatenation or weighted fusion), information of different modalities is merged into one feature space to improve classification performance. S2: Improved deep learning model design S2.1: Dual-Stream Network: Design two or more network branches for remote sensing data of different modalities (for example, one for processing optical images and the other for processing radar images). Each network independently learns the features of its own modality, and finally fuses the features through a fusion module (such as a weighted fusion layer). This design can maintain the independence of each modality information in the network and avoid mixed noise; S2.2: Joint learning of image and semantic features: In order to enhance classification accuracy, image features and semantic information (such as labels of feature types, regional information, etc.) are jointly learned. By introducing pre-trained knowledge graphs or semantic information systems, the model can better understand the semantic relationships between different features and reduce misclassification; S2.3: Dropout and data enhancement: By using Dropout regularization during training, the risk of network overfitting can be reduced. In addition, data enhancement methods (such as rotation, scaling, affine transformation, etc.) can be used to increase the diversity of training data, thereby improving the generalization ability of the model; S2.4: Hybrid data enhancement: Combine different modal data for hybrid enhancement (for example, data enhancement of optical images and radar images) to make the training data more diverse and reduce the impact of data imbalance. S3: Automated Labeling and Semi-Supervised Learning S3.1: Weakly supervised learning: Weakly supervised learning methods are used to reduce the reliance on manually labeled data. Rough labels can be generated through clustering algorithms or image segmentation methods, and then fine-tuned with small samples of high-quality labeled data to reduce labeling costs; S3.2: Transfer learning: For remote sensing images from different regions or different time periods, transfer learning methods are used to train the basic model using the labeled regional data, and then the model is adjusted to adapt to the new geographical area through transfer learning. This can reduce the demand for training data in new regions and avoid the bottleneck of data annotation; S3.3: Self-supervised feature learning: Through self-supervised learning methods, the model can automatically extract potential features from unlabeled data. It uses pixel-level similarity or contrast learning between images to train feature representation and improve the robustness of the model, especially when there is a lack of labeled data, which can significantly improve classification performance. S4: Improving real-time performance and computing efficiency S4.1: Model compression and quantization: Under the premise of ensuring classification accuracy, reduce the model's computational workload and storage requirements through pruning, knowledge distillation, and model quantization, improve inference speed and real-time performance, and adapt to real-time inference requirements on embedded or edge devices; S4.2: Edge computing and distributed processing: Assign model reasoning tasks to edge computing nodes or distributed computing environments to reduce the burden and latency of central servers and ensure real-time interaction of data between ground stations and edge devices; S4.3: TensorRT and hardware acceleration: Use TensorRT and other optimized inference engines to optimize hardware (such as GPU, TPU, FPGA, etc.) to significantly improve the inference speed of image classification. At the same time, choose deep learning frameworks that support hardware acceleration, such as ONNX, TensorFlow Lite, etc., to reduce inference latency and ensure real-time classification. S5: Interpretability and Model Transparency S5.1: Interpretable deep learning models: Use interpretability enhancement techniques (such as Grad-CAM, LIME, etc.) to enable the model to display the image regions or modality features it focuses on during classification. This not only helps users understand the decision-making process of the model, but also helps improve the transparency and reliability of the model; S5.2: Integrated model interpretation: Combine the prediction results of different modalities and use integrated learning methods (such as model fusion) to improve classification accuracy, while providing the contribution and interpretation of each modality to help analyze which modal features play a decisive role in the final classification. S6: Anomaly Detection and Noise Processing S6.1: Multimodal noise modeling and denoising: Design noise modeling and denoising methods based on the noise characteristics of different modal images. For example, use convolutional neural networks for end-to-end image denoising to remove atmospheric interference noise in optical images or echo noise in radar images; S6.2: Noise adaptive learning: Dynamically adjust the noise processing strategy through adaptive filtering methods to ensure the robustness and accuracy of the model in complex environments. S7: System testing and optimization S7.1: Multi-region testing: Test the performance of the classification model in different geographic regions to ensure that the model has good generalization ability in multiple scenarios. If the performance is not good in a specific region, further optimize the model through transfer learning; S7.2: Time series data validation: Use time series data to test the model to ensure that it can cope with seasonal changes or data from different time periods. By learning from time series data, the stability and accuracy of the model over a long period of time can be improved. S8: Release and Continuous Update S8.1: Continuous monitoring and feedback: Once the system is put into use, the model should be updated in a timely manner through continuous monitoring and user feedback to solve possible performance degradation problems; S8.2: Model update and adaptive training: Continuously update and optimize the model based on new remote sensing data and user needs, and use incremental learning methods for adaptive training.
2. The high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion according to claim 1, characterized in that: In the step S1, the pixel values of all optical images are normalized to the range of [0,1], and the mean and standard deviation of each image are used for standardization so that the images have the same dynamic range. For images with different spectral bands (such as multispectral images), each band is normalized separately. Radar images usually need to be logarithmically transformed (Log Transformation) to compress the dynamic range, and filters (such as mean filtering, Gaussian filtering, etc.) are used for noise suppression. The image contrast can be enhanced by histogram equalization, especially in low-contrast areas, to enhance the classification effect. The registration of different modal images is performed using feature point-based methods (such as SIFT, SURF, ORB, etc.) or frequency-domain-based mutual information methods. By calculating the similarity, the optimal transformation matrix is obtained. For images with different spatial resolutions, interpolation methods (such as bilinear interpolation and cubic interpolation) are used for resampling to ensure that multimodal images are aligned within the same spatial range. Super-resolution reconstruction (such as SRGAN based on deep learning) is used to improve the spatial resolution of low-resolution modal images.
3. The high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion according to claim 1, characterized in that: In the step S1, a multi-scale convolutional neural network is designed, and different convolution kernel sizes are used to extract different scale features of the image. Convolution kernels of different sizes such as 1x1 convolution, 3x3 convolution and 5x5 convolution are combined to extract global and local features, and "pyramid pooling" or "dilated convolution" is used to further capture large-scale features and enhance the recognition ability of complex structures in remote sensing images. Channel attention mechanism (SE-Net): By processing the feature map of each modality with the channel attention mechanism, the weight of each modality is automatically adjusted to reduce the impact of useless features. The weight of each channel is calculated by calculating its global average pooling feature, and the attention weight is calculated through a fully connected layer. Deep fusion layer: A weighted fusion layer is designed to dynamically adjust the fusion ratio according to the importance of each modality. An adaptive fusion method can be used, such as using a self-attention mechanism to calculate the importance of different modal features and weighted fusion.
4. The high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion according to claim 1, characterized in that: Two independent subnetworks are designed in step S2, and each subnetwork processes a modal data. For example, one subnetwork is used to process optical images, and the other is used to process radar images. The convolution layer in each subnetwork is responsible for extracting the features of each modality, and finally these features are merged through a fusion module. In the fusion module, the features of multiple modalities are fused in a concatenation or weighted summation manner to further improve the classification accuracy. In the deep neural network, a special semantic information module is added to obtain the prior knowledge of the object category. For example, the geographic information of the remote sensing image (such as land cover type, climate information, etc.) is combined as additional input information, and is input into the network together with the image features for joint training. Using the multi-task learning framework, tasks such as object recognition and object boundary detection are performed while classifying the image, thereby helping the model to better understand the semantic relationship of the object.
5. The high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion according to claim 1, characterized in that: In step S2, during the training process, a part of neurons is randomly discarded (Dropout) to prevent the model from over-relying on certain specific features. Image data enhancement methods such as rotation, translation, scaling, random cropping, etc. are used to increase the diversity of training samples and reduce the risk of overfitting. For multimodal data, data of different modalities are randomly combined for enhancement. For example, a part of the optical image and the radar image is randomly selected for splicing to generate new training samples and increase the robustness of the model.
6. The high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion according to claim 1, characterized in that: In the step S3, unlabeled data and some labeled data are used to train using a semi-supervised learning method. For example, an autoencoder is used to learn the features of unlabeled images, and a small sample of labeled data is used for fine-tuning. A clustering algorithm (such as K-means) is used to cluster the unlabeled data, and then the clustering results are corrected by manual annotation or semi-automatic methods, using a pre-trained network (such as a model pre-trained based on ImageNet or other remote sensing data sets), and fine-tuning is used to adapt to the new data set. For images in different geographical areas, the trained model can be migrated to the new area and trained with a small amount of new labeled data, saving time and resources.
7. The high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion according to claim 1, characterized in that: In step S3, a self-supervised learning method, such as Jigsaw Puzzle or Contrastive Learning, is used to allow the network to automatically learn meaningful feature representations from unlabeled data, and a contrast loss function (such as SimCLR) is used to train the model so that the feature vectors between similar images or image regions are closer, thereby enhancing the model's understanding of local details of the image.
8. The high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion according to claim 1, characterized in that: In step S4, pruning technology is used to remove redundant network weights to reduce computational complexity and storage requirements. During the pruning process, importance evaluation can be performed based on the L1 norm or gradient of each convolution kernel, and floating-point weights are converted into fixed-point weights through quantization to reduce memory usage and accelerate the reasoning process. The trained deep learning model is deployed to edge devices (such as drones, satellites, etc.) for real-time reasoning, reducing data transmission and latency, using distributed training frameworks (such as TensorFlow Distributed, Horovod, etc.) for large-scale training, using multiple machines for parallel training to improve computational efficiency, using TensorRT for model optimization, and converting the trained model into an optimized format for efficient reasoning on the GPU. Reasoning is accelerated by optimizing memory usage, reducing computational complexity, and increasing parallelism, supporting acceleration on specific hardware platforms (such as NVIDIA Jetson, Google Coral, etc.), and using the hardware's Tensor Core and other acceleration functions for reasoning.
9. The high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion according to claim 1, characterized in that: In step S5, the activation map of a certain layer in the convolutional neural network is visualized using technologies such as Grad-CAM to help understand the image area that the model focuses on when classifying remote sensing images. In combination with methods such as LIME (local interpretable model) or SHAP (SHapley Additive exPlanations), an explanation of the model prediction results is obtained to help users understand how the model makes classification decisions. The prediction results of each modality are integrated using ensemble learning methods (such as random forest, XGBoost, etc.), and the contribution of each modality to the final decision is explained through feature importance analysis of the ensemble model.
10. The high-resolution remote sensing image accurate classification system and method based on deep learning multimodal fusion according to claim 1, characterized in that: In step S6, for the radar image, a denoising network based on a convolutional neural network (CNN) or a U-Net structure may be used to denoise the image.
Citation Information
Cited By
Target multi-dimensional detection method based on deep learning multi-modal fusion technology
CN120339645A
A multi-dimensional target detection method based on deep learning multimodal fusion technology
CN120339645B
Real-time target detection and intelligent identification system and method based on unmanned aerial vehicle image
CN120411837A
Attention mechanism-based crop suitability dynamic evaluation method
CN120449112A
Large-computing-power SAR real-time imaging and target recognition system based on FPGA + GPU architecture
CN120595254A